Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
246,890 characters · 47 sections · 237 citation commands
Quantile Time Series Regression Models Revisited
The study of systemic risk remains a complex issue which occurs due to financial connectedness and interdependence in markets. Specifically, the increased level of connectivity and interdependence (see for example, billio2012econometric, diebold2012better, diebold2014network among others), can lead to the phenomenon of correlated defaults of financial institutions (see, duffie2009frailty). Specifically, the dependent variable of interest represents a portfolio loss (e.g., large default losses on portfolios of corporate debt). According to duffie2009frailty\footnote{The framework of duffie2009frailty provide a robust statistical estimation procedure in which the practitioner can construct the distribution of default times and rates based on the frailty variable and unobserved heterogeneity in the model.}:
Risk measures such as the CoVaR provide a representation of systemic risk in financial markets. Thus, motivated by the above aspects in these lecture series, we focus on studying existing estimation and inference methodologies for quantile time series models. We begin by reviewing relevant theory for quantile processes (moderate deviation principles e.g., see mao2019moderate) as well as estimation of quantile risk measures based on distribution functions. Therefore, we are interested to study of applications of time series regression models based on a conditional quantile functional form in both stationary (e.g., see he2020inference, katsouris2021optimal, katsouris2023statistical, escanciano2010specification) and nonstationary (e.g., see qu2008testing, lee2016predictive, xiao2009quantile, katsouris2023structural) time series models. Although we focus on quantile time series regression models we discuss relevant estimation aspects to quantile regression (see, koenker1978regression, koenker2002inference) in general such as the studies of victor2005extremal, portnoy2012nearly and daouia2022extremile among others.
A directly relevant framework to both the statistics as well as the econometrics literature is concerned with M-estimation techniques. Specifically, the M-estimation approach is a robust statistical methodology (e.g., robust to outliers and heavy-tails in data) pioneered by huber1996robust\footnote{In terms of M-tests these have been introduced in the literature for linear models such as in the paper of sen1982m, sen1986asymptotic and sen1987preliminary. See also el2001asymptotic.}. In the econometrics literature it is often used to refer to any estimator based on the maximization or minimization of a criterion function under the assumption of a pair of stationary time series. Thus, for the case of nonstationary time series (integrated, nearly-integrated or using the local-to-unity parametrization i.e., see phillips2007limit), more recently the literature has seen growing attention (see, xiao2012robust). Suppose that we are working with a parametrized model $( \mathbb{M}, \boldsymbol{\theta} )$. The range of the parameter-defining mapping $\boldsymbol{\theta}$ will be a parameter space $\Theta \in \mathbb{R}^k$. Let $Q^n ( \boldsymbol{y}, \boldsymbol{\theta} )$ denote the value of some criterion function, where $\boldsymbol{y}^n$ is a sample of $n$ observations on one or more dependent variables, and $\boldsymbol{\theta} \in \Theta$. Usually, $Q^n$ will depend on exogenous or predetermined variables as well as dependent variables $\boldsymbol{y}^n$.
Our main interest is to determine the asymptotic behaviour of the regression quantile process as defined below for different modelling environments.
Consider a nondecreasing function $G \in D [a,b]$ and define
for any $p \in \mathbb{R}$. Moreover, suppose we are interested to the limit behaviour of a test statistic of the following form
Then, the rejection region for testing the null hypothesis $H_0$ against $H_1$ is
where $c$ is a positive constant.
Moreover, the probability $\alpha_n$ of Type I Error and the probability $\beta_n$ of Type II Error are given by the following expressions
such that $a_n$ is the probability of false rejection.
Therefore, it holds that
Thus, we obtain the following moderate deviation rates
The above definitions are particularly useful for establishing moderate and large deviation principles for expression (ref) along with suitable rate functions depending on the properties of the underline stochastic process under consideration (see, victor2005extremal, gao2011delta, portnoy2012nearly, mao2019moderate and daouia2022extremile). Furthermore, from the nonstationary time series perspective, katsouris2022asymptotic develops a framework for moderate deviation principles from the unit boundary in nonstationary quantile autoregressions (see, also knight1991limit and references therein).
More recently, de2023conditional consider the connection between conditional quantiles and expectation operators which is quite useful for portfolio optimization purposes\footnote{Although for quantiles and conditional quantiles to be used in portfolio optimization problems further regularity conditions are needed to ensure that the statistical problem is well-defined (see, katsouris2021optimal). The statistical elicitability and properties of such risk measures is discussed by fissler2016higher, fissler2021elicitability, fissler2023backtesting. Moreover, patton2019dynamic, lin2023portfolio and corradi2023out investigate the estimation of these systemic risk measures and their applications to forecasting and backtesting.} and other optimization problems in economics and finance. Given a random variable in a probability space $( \Omega, \mathcal{F}, \mathbb{P} )$ and a $\sigma-$algebra $\mathcal{G} \subset \mathcal{F}$, we want to define the conditional quantile map
where $\rho_{\tau} ( \cdot )$ is the check function for some $\tau \in (0,1)$ (see, koenker1978regression, koenker1987estimation) defined as $\rho_{\tau} ( \cdot ) := ( \tau - 1 ) x \cdot \mathbf{1} \left\{ x < 0 \right\} + \tau x \cdot \mathbf{1} \left\{ x \geq 0 \right\}$.
\paragraph{Monotonization}
Another application for modelling quantile functions when the information set includes of many covariates considers the monotonization property of quantile operators. Suppose that $\mathcal{U}$ is a bounded closed interval such as $\mathcal{U} = [ u_L, u_U ]$ with $0 < u_L < u_U < 1$. The conditional quantile function $\mathcal{Q}_{Y | X} (u|x)$ is monotonically nondecreasing in $u$. However, the plug-in estimator $\mathcal{ \hat{Q} }_{Y | X} (u|x)$ constructed is not necessarily so. Let $F_{Y | X} ( y | X)$ to denote the conditional distribution of $Y$ given $X$. Then, for every realization of $X$, the map $y \mapsto F_{Y | X} ( y | X )$ is twice continuously differentiable with
In other words, quantile regression is a means of modelling the conditional quantile function. More precisely, the particular regression methodology has the ability to capture differentiated effects of the explanatory variables on various conditional quantiles of the dependent variable. Generally speaking, conditional quantiles of $Y$ given $X$ can provide additional information not captured by conditional mean in describing the relationship between $Y$ and $X$. Another important aspect we consider is the classical result of Huber for models with nonstandard conditions such as quantile regression models. Thus, the first order condition (FOC) defined as the right derivative of the objective function plays a key role in deriving the asymptotic theory of estimators for the quantile model. In particular, we can show that the parameter vector estimator solves these FOC and then apply a Bahadur representation for the estimator.
Denote with $Z_n = o_p(1)$ the random sequence $Z_n \to_p 0$, where it denotes convergence in probability. Then, the monotonicity condition $\frac{\partial }{ \partial \tau } x^{\prime} \beta (\tau) = \big[ f_{Y|X} \big( F^{-1}_{Y|X} (\tau) \big) \big]^{-1} \geq \frac{1}{b}$, implies that $x^{\prime} \beta (\tau)$ must grow by at least a rate of $\delta_n / b$ from $t_k$ to $t_{k+1}$. Using a Bahadur representation, $\big( \hat{\beta} ( \tau_k ) - \beta ( \tau_k ) \big) = \mathcal{O}_p \left( n^{-1/2} \right)$ uniformly in $k$ and so $\left\{ \hat{\beta}_k \right\}_{k=1}^M$ must also be strictly monotonic with probability tending to one.
The estimation of risk measures such as the Value at Risk, (VaR) and Conditional Value at Risk, (CoVaR) is an important aspect in risk management and portfolio allocation problems. The VaR is computed as the quantile of the loss distribution at a given confidence level, while the CoVaR (or expected shortfall) is the expected loss conditional on a certain level of VaR (see, acharya2012capital, acemoglu2015systemic). Various studies in the literature investigate the properties of these risk measures as well as related econometric applications (see, white2015var, blasques2016spillover, hardle2016tenet, chen2019tail). In particular, the CoVaR has recently been the preferred risk measure of examination due to the fact that it captures the conditional event of one financial institution being under stress given the event of another financial institution being at its Value at risk (see, tobias2016covar). The challenge that the econometrician faces is that this risk measure is not elicitable (see, fissler2016higher, patton2019dynamic), which means that there is no direct loss function which CoVaR is the solution to the minimization of the expected loss, and therefore a common practise is the use of a two-step estimation procedure for estimation and forecasting purposes.
Since risk measures depend on the state of the economy (see, tobias2016covar, cai2008nonparametric), then a common practise is to consider an information set, which contains both macroeconomic and financial variables. We shall denote with $Y_t$ the risk or loss of an asset portfolio, for the corresponding information set, denoted as $\mathcal{F}_{t-1}$. In particular, cai2008nonparametric considers the conditional $\tau-$quantile of the adapted sequence $\left\{ Y_{t} | \mathcal{F}_{t-1} \right\}_{t=1}^n$ where $\tau \in (0,1)$. Specifically, the one-period ahead VaR at confidence level $\tau$ is denoted by $Q_{ \tau }( Y_{t} | \mathcal{F}_{t-1} )$. Assume that we have available data $\{ X_t, Y_t \}$ for $t=1,...,n$ which are assumed to be generated from a stationary process.
Let $Q_{\tau} (x)$ be the CVaR\footnote{Note that CVaR and CoVaR denote two different risk measures in the literature. CVaR implies the estimation of VaR given the information set, while CoVaR denotes the risk measure proposed by tobias2016covar. Moreover, CES in the literature has similar definition as the CoVaR, however the CoVaR estimation requires to use the two-step quantile regression procedure as proposed by tobias2016covar (see, also patton2019dynamic).} which can be expressed as $Q_{\tau}(x) = S^{-1} ( \tau |x )$ where $S( \tau |x ) = 1 - F( y | x )$ and $F( y | x )$ is the conditional CDF of $Y_t$ given $X_t = x$. Then, cai2008nonparametric propose the nonparametric estimation of $Q_{\tau} (x)$ can be constructed as $\widehat{Q_{\tau}} ( x ) = \widehat{S}^{-1} ( \tau | x )$, where $\widehat{S}^{-1} ( \tau | x )$ is a nonparametric estimation of $S^{-1} ( \tau | x )$. Then the CES denoted as $\mu_{\tau} (x)$ is formulated as below
where $f( y | x)$ represents the conditional PDF of $Y_t$ given $X_t = x$. Therefore, to estimate $\mu_{\tau} (x)$, one can use the plugging-in method as below
Further details on a nonparametric framework suitable for estimating risk measures such as the expected shortfall are given in the study of martins2018nonparametric.
In summary suppose that $Q( \tau | x ) := Q \left( \tau | \underline{X}_j = x \right)$ denotes the conditional $\tau-$th quantile of $Y_j$ given $\underline{X}_j = x$, which is the main conditional function we are interested to estimate and forecast. An equivalent expression in terms of the probability space, is written as below
We estimate the conditional quantile function (CQF) via the following loss function
where $\tau \in (0,1)$, is a specific quantile level, and $\rho_{\tau}( u ) = u \big( \tau - \mathbf{1} { \left\{ u < 0 \right\} } \big)$ is the check function. Then conditional quantile function is estimated via
Furthermore, a linear approximation to the CQF is provided by the QR parameter $\beta ( \tau )$, which solves the population minimization problem described below:
under the assumption of integrability and uniqueness of the solution. Therefore, the QR parameter $\beta ( \tau )$ provides a summary statistic for the CQF. Then, the corresponding QR estimator has the following form
Note that equivalently, the QR estimator $\hat{ \beta } ( \tau )$ is also the generalized method of moments (GMM) estimator based on the unconditional moment restriction given by
Moreover, when the CQF is modelled via a linear to the regressors function, such that $Q( Y | X ) = X^{\prime} \beta ( \tau )$ or $F_Y ( X^{\prime} \beta ( \tau ) \big| X ) = \tau$, then the coefficient $\beta ( \tau )$ satisfies the conditional moment restriction given by
with almost surely convergence. Further relevant literature to the estimation of risk measures such as the VaR and CoVaR include the framework of xu2016model, who consider a semi-parametric index model with multiple covariates for the joint estimation of the VaR and Expected Shortfall. Moreover, white2015var) consider a multivariate specification for the joint estimation of quantile risk measures.
Therefore, to assess the predictive accuracy of these risk measures, the literature has proposed the methodology of backtesting since the CoVaR is considered to be an unobservable quantity. In particular, some notable studies that propose backtesting methodologies for the VaR are presented by escanciano2010backtesting and escanciano2011robust. Specifically, escanciano2010backtesting, consider the forecast evaluation problem using backtesting methodologies robust to estimation risk. Therefore, it is of paramount importance to develop powerful tests for the correct specification of parametric conditional quantiles over a possibly continuous range of quantiles of interest and under general conditions on the underlying data-generating process.
Denote the conditional CDF of $Y$ given $X = x$ by $F_{Y|X} (. | x)$ and its conditional quantile at $\uptau \in (0,1)$ by $Q ( \uptau | x )$, that is,
where $Q( \uptau | x )$ is modelled as a general nonlinear function of $x$ and $\uptau$. We fix $x$ and treat $Q( \uptau | x )$ as a process in $\uptau$, where $\uptau \in \mathcal{T} = [ \lambda_1, \lambda_2 ]$ with $0 < \lambda_1 \leq \lambda_2 < 1$. In this section we follow the framework of portnoy2012nearly. Denote with $\dot{\phi}(t)$ the conditional characteristic function of the random variable
given $x_i$. Moreover, let $f_i(y)$ and $F_i(y)$ denote the conditional density and CDF of $Y_i$ given $x_i$.
Then, portnoy2012nearly considers the proof for the Hungarian construction developed inductively. More precisely, it follows that the coverage probability may be computed using only two terms of the Taylor series expansion for the normal CDF:
In other words, portnoy2012nearly employs the “Hungarian” construction of komlos1975approximation who provide an alternative expansion for the one-sample quantile process with nearly the root-n error rate (see, portnoy2012nearly). Notice that establishing nearly root$-n$ approximations of quantile-based processes is instrumental for avoiding nearly singular designs (see, randles1982asymptotic, knight2001comparing, knight2008shrinkage as well as bhattacharya2020quantile among others). Thus, motivated by these considerations portnoy2012nearly develops a framework to provide an increased accuracy for conditional inference beyond that provided by the traditional Bahadur representation. Specifically, the focus is to provide a theoretical justification for an error bound of nearly root-n order uniformly in $\uptau$. Define with
where $\big[1 / f \big( F^{-1} ( \uptau ) \big) \big]$, corresponds to the sparsity function. Next, we consider a bivariate approximation for the joint density of one regression quantile and the difference between this one and a second regression quantile (properly normalized for the difference in $\uptau-$values). Let $\epsilon \leq \uptau_1 \leq 1 - \epsilon$ for some $\epsilon > 0$, and let $\uptau_2 = \uptau_1 + a_n$ with $a_n > c n^{-b}$ for some $b < 1$. Define with
The following theorem provides the “Hungarian” construction:
Quantile regression is a flexible and powerful approach which allows us to model the quantiles of the conditional distribution of a response variable given a set of covariates. Regression quantile estimators can be viewed as M-estimators and standard asymptotic inference is readily available based on likelihood-ratio, Wald and score-type test statistics. However, these statistics require the estimation of the sparsity function and this can lead to nonparametric density estimation. The most appealing feature of QR methods is that they allow estimation of the effect of covariates on many points of the outcome distribution, including the tails as well as the center of the distribution. On the other hand, these do not capture the full distribution impact of a variable unless the variable affects all quantiles of the outcome distribution in the same way. Overall QR methods have the ability to capture heterogeneous effects. In the framework proposed by chernozhukov2009finite, the authors show that valid finite sample confidence regions can be constructed for parameters of a model defined by quantile restrictions under minimal assumptions. The proposed approach makes use of the fact that the estimating equations that correspond to conditional quantile restrictions are are conditionally pivotal; that is, conditional on the exogenous regressors and instruments, the estimating equations are pivotal in finite samples.
\paragraph{Uniform Convergence Rates}
We first establish uniform convergence rates of $\hat{Q}_x ( \uptau )$. The following Bahadur representation of the linear quantile regression estimator $\hat{ \boldsymbol{\beta} }$ is employed in subsequent proofs. Recall that we denote with $Q_{ \boldsymbol{x} }( \uptau )$, where $\uptau \in (0,1)$, the conditional $\uptau-$quantile of Y given $\boldsymbol{x}$ (e.g., see ota2019quantile). Notice that we have that the conditional quantile function with respect to the quantile index $\uptau$ coincides with the reciprocal of the conditional density at $Q_{ \boldsymbol{x} }( \uptau )$ such that
where $s_{ \boldsymbol{x} } \left( \uptau \right)$ denotes the sparsity function. The estimation of the sparsity function will affect the finite-sample performance of the test statistics if not consistently estimated. A key quantity for establishing related asymptotic theory results is the estimation of the following matrix
Notice that $U_t = F( Y_t | \boldsymbol{X}_t )$ where $F( y | \boldsymbol{X} )$ is the conditional distribution function of Y given $\boldsymbol{X}$. Therefore, to prove the technical lemma we begin by expanding the following expression
Therefore, observe that
Therefore, by the Taylor expansion we obtain that
uniformly in $\uptau \in [ \epsilon / 2, - \epsilon / 2 ]$. It remains to show that
within the interval $\big[ \epsilon / 2, - \epsilon /2 \big]$. Thus, since we have $\sqrt{n}-$consistency within the region $\big[ \epsilon / 2, - \epsilon /2 \big]$, that is, $\left\lVert \widehat{\beta} - \beta \right\rVert_{ \big[ \epsilon / 2, - \epsilon /2 \big] }$, for any $M_n \to \infty$ sufficiently slowly such that
Therefore, we consider the following function class as below
where $\mathbb{S}^{d-1} = \left\{ x \in \mathbb{R}^d : \left\lVert x \right\rVert = 1 \right\}$.
\paragraph{Quantile Regression under Misspecification}
Specifically, the Theorem presented in angrist2006quantile provides the asymptotic distribution for the stochastic sequence which involves the model parameter. Then, the mapping $\uptau \mapsto \boldsymbol{\beta}(\uptau)$ is continuous by the implicit function theorem and relevant assumptions. In fact, because $\boldsymbol{\beta}(\uptau)$ solves
Moreover, it holds that $\frac{ d \boldsymbol{\beta}(\uptau) }{ d\uptau } = J(\uptau)^{-1} \mathbb{E} \left[ X \right]$. Hence, $\uptau \mapsto \mathbb{G}_n \big[ \phi_{\uptau} \big( Y - X^{\prime} \boldsymbol{\beta}(\uptau) \big) X \big]$, is stochastically equicontinuous over $\mathcal{T}_{\iota}$ for the metric given by
Thus, stochastic equicontinuity of $\uptau \mapsto \mathbb{G}_n \big[ \phi_{\uptau} \big( Y - X^{\prime} \boldsymbol{\beta}(\uptau) \big) X \big]$ and a multivariate central limit theorem
where $z(\uptau)$ is a Gaussian process with covariance function $]=\Sigma(.,.)$ which implies that
Therefore, it holds that
Consider that the conditional density $f_Y( y | X = x )$ exists, and is bounded and uniformly continuous in $y$ and uniformly in $x$ over the support of $X$. Define with
is positive definite for all $\uptau \in \mathcal{T}_{\iota}$. Then, the quantile regression process is uniformly consistent,
and $J(\uptau):= \sqrt{n} \left( \hat{\boldsymbol{\beta}}(\uptau) - \boldsymbol{\beta}(\uptau) \right)$, converges in distribution to a zero mean Gaussian process $z(\uptau)$, where $z(\uptau)$ is defined by its covariance function $\Sigma( \uptau_1, \uptau_2 ) := \mathbb{E} \big[ z(\uptau_1), z(\uptau_2)^{\prime} \big]$ with
Thus, if the model is correctly specified, that is, $Q_{\uptau} (Y|X) = X^{\prime} \boldsymbol{\beta}(\uptau)$ almost surely, then $\Sigma( \uptau_1, \uptau_2 )$ simplifies to the following expression
Notice that the proof of this theorem in angrist2006quantile, proceeds by establishing the uniform consistency of the QR process $\uptau \mapsto \hat{\boldsymbol{\beta}}(\uptau)$. Furthermore, note that the class of functions
is Donsker, the estimating equation for the QR process, is given by the following expression
is expanded in $\hat{\boldsymbol{\beta}}(\uptau)$ around the true parameter value $\boldsymbol{\beta}(\uptau)$ since
uniformly in $\uptau \in \mathcal{T}_{\iota}$. Then, the conclusion of the theorem follows by the central limit theorem for empirical processes indexed by Donsker classes of functions. In summary, the above theorem establishes joint asymptotic normality for the entire QR process. Moreover, the theorem allows for misspecification and imposes a little structure on the underlying conditional quantile function, such as smoothness of $Q_{\uptau} (Y|X)$ in $X$. Therefore, the result in the theorem states that the limiting distribution of the QR process (and of any single QR coefficient) will, in general, be affected by misspecification. In particular, the covariance function that describes the limiting distribution is generally different from the covariance function that arises under correct specification.
Following the framework of qu2015nonparametric (see also xu2013nonparametric), we start by reviewing the idea underlying the local linear regression. More precisely, for a given $\uptau \in (0,1)$, the method assumes that $\mathcal{Q} ( \uptau | x )$ is a smooth function of $x$ and exploits the following first-order Taylor approximation:
The local linear estimator of $\mathcal{Q} ( \uptau | x )$, denote by $\hat{\alpha} ( \uptau )$, is determined by
where $\rho_{\uptau} (u) = u \big( \uptau - I \left( u < 0 \right) \big)$ is the check function, $K$ is a kernel function and $h_{n, \uptau}$ is a bandwidth parameter that depends on $\uptau$.
In particular, the local linear regression has several advantages over the local constant fit such that: (i) the bias of $\hat{\alpha} ( \uptau)$ is not affected by the value of $f_X^{\prime} (x)$ and $\partial \mathcal{Q} ( \uptau | x ) / \partial x$, (ii) it is of the same order irrespective of whether $x$ is a boundary point, and (iii) plug-in data-driven bandwidth selection does not require estimating the derivatives of the marginal density, therefore is relatively simple to implement.
Define with $u_i( \uptau ) = y_i - \alpha( \uptau ) - ( x_i - x )^{\prime} \beta ( \uptau )$, where $\alpha(\uptau) \in \mathbb{R}$ and $\beta(\uptau) \in \mathbb{R}^d$ are some candidate parameter values. Let
Then, $u_i( \uptau )$ can be decomposed as
where $z_{i, \uptau}^{\prime} = \big( 1, \left( x_i - x \right)^{\prime} / h_{n, \uptau} \big)$.
This decomposition is useful because it breaks $u_i ( \uptau )$ into three components: the true residual, the error due to the Taylor approximation and the error caused by replacing the unknown parameter values in the approximation with some estimates.
A Bahadur representation of quantile-dependent model coefficients based on M-estimators is essential for establishing asymptotic theory results. Various studies obtain such representations under different modelling conditions. he1996general, consider a sequence of variables $\left\{ x_i \right\}_{ i = 1}^n$, that are independent but not necessarily identically distributed, while wu2005bahadur considers a Bahadur representation for quantile processes of dependent sequences and ren2020local the case of near-epoch dependence.
Thus, following the framework proposed by he1996general, suppose that there exists $\theta_0$ such that
for some score function $\psi$. Therefore, we consider any $M-$estimator $\hat{\theta}_n$ of $\theta_0$ which satisfies
\paragraph{M-estimators with smooth score functions}
Consider the simplest case, where $\widehat{\theta}_n$ is defined through the following moment condition
where $\phi$ is a Lipschitz continuous function. Denote with $Q_n = \sum_{i=1}^n z_i z_i^{\prime}$ and denote with $x_i = \left( y_i, z_i \right)$ and $\psi \left( x_i, \theta \right) = \phi \left( y_i - z_i^{\prime} \theta \right) z_i$. Notice that part of $x_i$ has a degenerate point-mass distribution.
Consider the joint density for $\hat{\beta} ( \uptau_1 )$ and $\hat{\beta} ( \uptau_2 )$. One apparent complication is that one needs $a_n \equiv \left| \uptau_1 - \uptau_2 \right| \to 0$, making the covariance matrix tend to singularity. Therefore, the particular approach requires that $D_n$ converges at a rate depending on $a_n$.
Relevant studies include among others galvao2020unbiased.
Under the above assumptions, the quantile regression estimator has Bahadur representation given by the following expression
where $\psi_{\tau}( . ) = \tau - \mathbf{1} \left\{ . < 0 \right\}$ and $\psi_{ \boldsymbol{ \beta } }( . ) \equiv \boldsymbol{x}_t \psi_{\tau}( . )$. Thus, the quantile regression estimator has asymptotic normal distribution given by
where $D = \frac{ \displaystyle \uptau (1 - \uptau) }{ \displaystyle f^2 \left( F^{-1} (\uptau) \right) } Q^{-1}$.
Following xiao2012time, a simple high-level assumption that we make on the QAR process is monotonicity of the functional form of the model. The monotonicity of the conditional quantile functions allows specific forms for the $\theta$ functions. Moreover, statistical inference can be conducted but the limiting distribution needs to be modified to accommodate the possible misspecification. Additionally, in various occasions it is necessary to employ a bootstrap resampling scheme to approximate the large sample properties of estimators and test statistics (see, hahn1995bootstrapping, hagemann2017cluster, galvao2021hac, galvao2023bootstrap).
Quantile regression can be applied to regression models with dependent errors. In particular, consider the following linear model
where $X_t$ and $u_t$ are $k$ and $1-$dimensional weakly dependent stationary random variables, $\left\{ X_t \right\}$ and $\left\{ u_t \right\}$ are independent with each other, $\mathbb{E} \left( u_t \right) = 0$. Denote with $F_u(.)$ the distribution function of $u_t$, then conditional on $X_t$, the $\uptau-$th quantile of $Y_t$ is given by the following expression
where $\boldsymbol{\theta} (\uptau ) = \left( \alpha + F_u^{-1} ( \uptau), \boldsymbol{\beta} \right)^{\prime}$. Then, the vector of parameters, $\boldsymbol{\theta} (\uptau )$, can be estimated by solving the optimization problem below
Define with $u_{t \uptau} = Y_t - \theta( \uptau )^{\prime} \boldsymbol{Z}_t$, we have that $\mathbb{E} \left[ \psi_{\uptau} ( u_{t \uptau} ) | X_t \right] = 0$. Moreover, under assumptions on moments and weak dependence on $( X_t, u_t )$,
where $\boldsymbol{\Sigma} ( \uptau )$ is the long-run covariance matrix of $\boldsymbol{Z}_t \psi_{ \uptau } ( u_{t \uptau} )$ defined by
Then, the quantile regression estimator has the following asymptotic representation:
As a result,
Thus, the above results may be extended to the case where other elements in $\boldsymbol{\theta} ( \uptau )$ are also $\uptau-$dependent. Statistical inference based on $\widehat{\boldsymbol{\theta}} ( \uptau )$ requires estimation of the covariance matrices $\boldsymbol{\Sigma}_z$ and $\boldsymbol{\Sigma} ( \tau )$. The matrix $\boldsymbol{\Sigma}_z$ can be easily estimated by its sample analogue
while $\boldsymbol{\Sigma} ( \uptau )$ may be estimated following the HAC estimation literature. Define with $\widehat{u}_{t \tau} = Y_t - \widehat{ \boldsymbol{\theta} } ( \uptau )^{\prime} \boldsymbol{Z}_t$, we may estimate $\boldsymbol{\Sigma} ( \uptau )$ by
where $k(.)$ is the lag window defined on [-1,1] with $k(0) = 1$ and $M$ is the bandwidth parameter satisfying the property that $M \to \infty$ and $M / n \to 0$ as the sample size $n \to \infty$. For instance, the asymptotic properties for regression quantiles with $m-$dependent errors is examined by portnoy1991asymptotic. Notice that the delta method is also employed by girard2021extreme.
Overall, when using Wald-type statistics for testing linear restrictions in quantile regression models one follows the following steps. Consider $\uptau \in \mathcal{T}_{\iota}$, where $\mathcal{T}_{\iota } = [ \iota , 1 - \iota ]$ such that $0 < \iota < 1/ 2$. Define with
where $\hat{f}_{i,m,n} \left( \xi_t ( \uptau ) \right)$ is the conditional density estimator embedding a regression quantile spacing. Then, the following form of the Wald statistic follows
The limit distribution of the test statistic $\mathcal{W}_{n,m} ( \uptau )$ is given by the Corollary below. Denote with $\mathcal{W}_{n,m} ( \uptau )$ and recall the following definition $ \boldsymbol{\Delta}_m ( \uptau ) \equiv \underset{ \mathsf{u} }{ \mathsf{arg \ min} } \ \mathcal{Z} \left( \uptau, \mathsf{u} \right)$, where $\mathcal{Z} \left( \uptau, \mathsf{u} \right)$ is a criterion function.
Therefore, the particular expression is regarded as a stochastic process in the space $\mathcal{D} \left( [0,1] \right)^d$ of $\mathbb{R}^d-$valued right-continuous functions with left-hand limits on $[0,1]$. Then, based on existing results on the weak convergence of regression quantile processes in $\mathcal{D} \left( [0,1] \right)^d$, we can establish the limit of the sup-Wald test statistic as given by the following Corollary.
Following the framework of galvao2014testing, denote with $\left( y_t, q_t, \boldsymbol{x}_t^{\prime} \right)^{\prime}$ be a triple of scalar dependent variable $y_t$, a scalar threshold variable $q_t$ and a vector $d$ of explanatory variables $\boldsymbol{x}_t$. Moreover, denote with $\mathcal{Q}_{ y_t } \left( \uptau | \mathcal{F}_{t-1} \right)$ the conditional $\tau-$quantile of $y_t$ given the $\sigma-$algebra $\mathcal{F}_{t-1}$, where $\uptau \in (0,1)$, that is a random quantile within the compact set $(0,1)$. Then, consider testing the null hypothesis
against the alternative hypothesis
where $\mathcal{T}:= [ \uptau_L, \uptau_U ]$ is a bounded closed interval in $(0,1)$ and $\gamma_0$ is the threshold parameter. More precisely, the null hypothesis assumes that the conditional quantile function is linear in $\boldsymbol{x}_t$ uniformly over a given range of quantiles, while the alternative hypothesis assumes that the conditional quantile regression function follows a threshold model at some quantile.
Therefore, to differentiate the alternative from the null hypothesis, we assume that $\boldsymbol{\theta}_1 ( \uptau_0 ) \neq \boldsymbol{\theta}_2 ( \uptau_0 )$. However, for notation convenience we formulate the null hypothesis with different notation. Denote with $\boldsymbol{\beta}_{(1)} ( \uptau_0 ) = \boldsymbol{ \theta }_1 ( \uptau_0 )$ and $\boldsymbol{\beta}_{(2)} ( \uptau_0 ) = \boldsymbol{ \theta }_2 ( \uptau_0 ) - \boldsymbol{ \theta }_1 ( \uptau_0 )$.
Thus, the alternative hypothesis can be expressed as below
where $\boldsymbol{z}_t ( \gamma ) = \big( \boldsymbol{x}_t^{\prime}, I \left( q_t \leq \gamma \right) \boldsymbol{x}_t^{\prime} \big)^{\prime}$ and $\boldsymbol{\beta} ( \uptau_0 ) := \big( \boldsymbol{\beta}_{(1)}^{\prime} ( \uptau_0 ), \boldsymbol{\beta}_{(2)}^{\prime} ( \uptau_0 ) \big)^{\prime}$. Therefore, using this notation we can write the null hypothesis as below
regardless of the value of $\gamma \in \Gamma$.
Thus, given $( \uptau, \gamma ) \in \mathcal{I} \times \Gamma$, let $\hat{ \boldsymbol{\beta} } ( \uptau, \gamma )$ be the estimator defined by
In other words, $\hat{ \boldsymbol{\beta} } ( \uptau, \gamma )$ represents the quantile regression estimator when we treat $\boldsymbol{z}_t ( \gamma )$ as explanatory variables. Furthermore, when $\mathcal{H}_0$ is true, under suitable regulatory conditions, $\hat{ \boldsymbol{\beta} }_2 ( \uptau, \gamma )$ converges in probability to $\mathbf{0}$ for each $( \uptau, \gamma ) \in \mathcal{T} \times \Gamma$. On the other hand, when $\mathcal{H}_1$ is true, then $\hat{\boldsymbol{\beta}}_2 ( \uptau_0, \gamma_0 )$ converges in probability to $\hat{\boldsymbol{\beta}}_{(2)} ( \uptau_0 )\neq 0$. However, we know a priori neither the quantile $\uptau_0$ where the linearity breaks down nor the true value of the threshold parameter $\gamma_0$ at that quantile. Therefore, it is reasonable to reject $\mathcal{H}_0$ if the magnitude of $\hat{ \boldsymbol{\beta} }_2 ( \uptau_0, \gamma_0 )$ is suitably large for some $( \uptau, \gamma ) \in \mathcal{T} \times \Gamma$.
A natural choice is to test $\mathcal{H}_0$ against $\mathcal{H}_1$ by the supremum of the Wald process as below
where $\boldsymbol{V}_{22} ( \uptau, \gamma )$ is the asymptotic covariance matrix of $\sqrt{n} \hat{\boldsymbol{\beta}}_{(2)} ( \uptau, \gamma )$ under $\mathcal{H}_0$.
Therefore, when $\mathcal{H}_0$ is true, $\boldsymbol{\beta}_{(1)} ( \uptau ) = \boldsymbol{\beta}_{(1)}^{*} ( \uptau )$ and $ \mathcal{Q}_{ y_t } \left( \uptau | \mathcal{F}_{t-1} \right) = \boldsymbol{x}_t^{\prime} \boldsymbol{\beta}_{(1)}^{*} ( \uptau )$ for all $\uptau \in \mathcal{T}$. When $\mathcal{H}_0$ is not true, $\boldsymbol{\beta}_{(1)}^{*} ( \uptau )$ it can be interpreted as the coefficient vector of the best linear predictor of the conditional quantile function against a certain weighted mean-squared loss function. In other words, the condition given by Assumption (ref) above demonstrates that the limiting null distribution of $SW_n$ depends on the probability limit of the quantile regression estimator $\hat{\boldsymbol{\beta}}_1 ( \uptau, \gamma )$ under the null hypothesis.
Thus, under these conditions, the map $\uptau \in \mathcal{T}^{*} \mapsto \boldsymbol{\beta}_{(1)}^{*} ( \uptau )$ is continuously differentiable by the implicit function theorem. Furthermore, Assumption (ref) guarantees that the matrices $\boldsymbol{\Omega}_0 \left( \gamma_1, \gamma_2 \right)$ and $\boldsymbol{\Omega}_1 \left( \uptau, \gamma \right)$ do not degenerate for each fixed $\gamma \in \Gamma$ and $( \uptau, \gamma ) \in \mathcal{T} \times \Gamma$, respectively. Due to the computational property of the quantile regression estimate we can select $\hat{ \boldsymbol{\beta} } ( \uptau, \gamma )$ in such a way that the path $( \uptau, \gamma ) \mapsto \hat{ \boldsymbol{\beta} } ( \uptau, \gamma )$ is bounded. Therefore, we can also assume that the path
is bounded over $( \uptau, \gamma ) \in \mathcal{T} \times \Gamma$.
Denote with $\ell^{\infty} \left( \mathcal{T} \times \Gamma \right)$ the space of all bounded functions on $\mathcal{T} \times \Gamma$ equipped with the uniform topology, and $\left( \ell^{\infty} \left( \mathcal{T} \times \Gamma \right) \right)^{2d}$ to denote the $(2d)-$product space of $\ell^{\infty} \left( \mathcal{T} \times \Gamma \right)$ equipped with the product topology. Here, we denote with $\boldsymbol{\beta}^{*} ( \uptau ):= \left( \boldsymbol{\beta}_{(1)}^{*} ( \uptau )^{\prime}, \boldsymbol{0} ^{\prime} \right)^{\prime} \in \mathbb{R}^{2d}$.
We introduce the local objective function as below
where $\boldsymbol{u} \in \mathbb{R}^{2d}$ and $( \uptau, \gamma ) \in \mathcal{T} \times \Gamma$. Furthermore, we observe that the normalized random quantity $\sqrt{n} \left( \hat{ \boldsymbol{\beta} } ( \uptau, \gamma ) - \boldsymbol{\beta}^{*} ( \uptau ) \right)$ minimizes $\mathcal{Z}_n \left( \boldsymbol{u}, \uptau, \gamma \right)$ with respect to $\boldsymbol{u}$ for each fixed $( \uptau, \gamma ) \in \mathcal{T} \times \Gamma$. Thus, to prove Theorem 1, we utilize Theorem 2 of kato2009asymptotics. More specifically, since $\mathcal{Z}_n \left( \boldsymbol{u}, \uptau, \gamma \right)$ is now convex in $\boldsymbol{u}$, by Theorem 2 of kato2009asymptotics\footnote{Notice that Theorem 2 of kato2009asymptotics guarantees that there is no need to establish the uniform $n^{- 1/ 2}$ rate of $\hat{ \boldsymbol{\beta} } \left( \uptau, \gamma \right)$ in the course of establishing the weak convergence of the process.}, it is sufficient to prove the following proposition.
Using knight1998limiting identity we obtain
Then, we decompose the stochastic process $\mathcal{Z}_n \left( \boldsymbol{u}, \uptau, \gamma \right)$ into three parts as below:
We denote with
Since we have that
Therefore, by the dominated convergence theorem, we have that
Furthermore, we define with
where $( \uptau, \upgamma ) \in \mathcal{T} \times \Gamma$ and $s \in \mathbb{R}$. Therefore, we shall show that
which leads to
The main idea of the local linear fit is to approximate the $q_{\tau}(z)$ in a neighbourhood of $x$ by a linear function such that
where
Therefore, we define an estimator by setting $q_{\tau} (x)$ at $x = \left( x_1,..., x_p \right)^{\prime} \in \mathbb{R}^p$. Thus, locally estimating $\big( q_{\tau} (x), \dot{q}_{\tau}(x) \big)$ is equivalent to estimating $\widehat{q}_{\tau} (x) \equiv \widehat{\alpha}_0$ and $\widehat{ \dot{q} }_{\tau} (x) \equiv \widehat{\alpha}_1$. Then,
where $K_h (x) = h_n^{ -p } K ( x / h_n )$, $K$ is a kernel function on $\mathbb{R}^p$, and $h_n > 0$ is the bandwidth.
Assume that $\left\{ \left( Y_t, X_t \right) \right\}$ is a stationary multivariate time series on a probability space $\left( \Omega, \mathcal{F}, P \right)$, where $X_t$ and $Y_t$ are random variables. Notice that $X_t$ may consist of both the lags of endogenous and/or exogenous variables. In particular, we are interested in the $\tau-$conditional quantile function of $Y_t$ given $X_t = x$,
Based on the powerful took of the weak Bahadur representation, we can establish the asymptotic distribution of the local linear quantile regression estimates under near-epoch dependence. The following lemmas are helpful for proving the asymptotic normality result. Suppose that
Furthermore, the usual Cramer-Wold device will be adopted. For all $c := \left( c_0, c_1^{\prime} \right)^{\prime} \in \mathbb{R}^{ p + 1}$. Define,
Then, the expectation and asymptotic variance of the above expression is provided below.
Then, the asymptotic variance is given by
\paragraph{Full and sub-model estimations}
We consider the following partitioned form as defined by yuzbacsi2017pretest:
where $p = p_1 + p_2$ and $\boldsymbol{\beta}_1$, $\boldsymbol{\beta}_2$ parameters are of order $p_1$ and $p_2$ respectively and $\boldsymbol{x}_i = \left( \boldsymbol{x}_{1i}^{\prime}, \boldsymbol{x}_{2i}^{\prime} \right)$ and $\epsilon_{i}$'s are the error terms. The conditional quantile function for the response variable $y_i$ is written in the form:
The null hypothesis of interest can be formulated as below
Considering the full model (FM) versus the sub-model (SM) the testing hypothesis is evaluated based on the following Wald statistic
where $\boldsymbol{D}_{ij}$ for $i,j \in \left\{ 1,2 \right\}$ is the $(i,j)-$th partition of the $\boldsymbol{D}$ matrix and $\boldsymbol{D}^{ij}$ is the $(i,j)-$th partition of the $\boldsymbol{D}^{-1}$ matrix such that
Denote with $\omega := \sqrt{ \uptau( 1 - \uptau)} \big/ f \big( F^{-1} ( \uptau ) \big)$, where the term $f \big( F^{-1} ( \uptau ) \big)$ is called the sparsity parameter or quantile density parameter and the sensitivity of the test statistic naturally depends on this parameter.
The full model (FM) quantile regression estimator is obtained by
The sub-model (SM) quantile regression estimator is given by
The sub-model (SM) quantile regression estimator is obtained by
Let $\left\{ K_n \right\}$ be a sequence of local alternatives given by
where $\boldsymbol{\kappa} = \left( \kappa_1,..., \kappa_{p_2} \right)^{\prime} \in \mathbb{R}^{p_2}$ is a fixed vector. When $\boldsymbol{\kappa} = \boldsymbol{0}_{ p_2 }$, then the null hypothesis is true. Furthermore, we consider the following proposition to establish the asymptotic properties of the estimators.
where $\chi^2_{\nu} \left( \Delta \right)$ is a non-central chi-square distribution with $\nu$ degrees of freedom and non-centrality parameter $\Delta$.
where $\boldsymbol{\delta} = \boldsymbol{D}_{11}^{-1} \boldsymbol{D}_{12} \boldsymbol{\kappa}$, $\boldsymbol{\Phi} = \omega^2 \boldsymbol{D}_{11}^{-1} \boldsymbol{D}_{12} \boldsymbol{D}_{22.1}^{-1} \boldsymbol{D}_{21} \boldsymbol{D}_{11}^{-1}$, $\boldsymbol{\Sigma}_{12} = - \omega^2 \boldsymbol{D}_{12} \boldsymbol{D}_{21} \boldsymbol{D}_{11}^{-1}$ and $\boldsymbol{\Sigma}^{\star} = \boldsymbol{\Sigma}_{21} + \omega^2 \boldsymbol{D}_{11.2}^{-1}$.
Further details on the use of Stein-type estimators in statistical applications can be found in the studies of nkurunziza2012shrinkage and chen2016class. An application of restricted estimators for structural break testing is proposed by nkurunziza2021inference.
The econometric framework for estimation and inference for quantile autoregression is presented by koenker2004unit, koenker2006quantile and galvao2009unit (see, also koenker2017handbook).
Consider the following quantile autoregressive model
where $\rho_{\tau} ( \mathsf{u} ) = \mathsf{u} \big( \tau - \mathbf{1} \left\{ \mathsf{u} < 0 \right\} \big)$ is the check function.
The following unit root t-ratio test statistic holds
Notice that the asymptotic distribution of the above t-ratio test statistic is nonstandard since $B_{\psi}$ and $B_{\mu}$ are correlated. However, it can be decomposed into a linear combination of two independent parts. Specifically the following decomposition holds
Consider the ADF regression model. Denote the $\sigma-$field generated by $\left\{ u_s, s \leq t \right\}$ by $\mathcal{F}_t$, then conditional on $\mathcal{F}_{t-1}$, the $\uptau-$th conditional quantile of $Y_t$ is given by
Moreover, denote with $\alpha_0 ( \uptau ) = \mathcal{Q}_{u}( \uptau )$, $\alpha_j ( \uptau )$, $j = 1,...,p, p = q+1$, and define with
we have that $\mathcal{Q}_{ Y_t } \left( \uptau | \mathcal{F}_{t-1} \right) = X_t^{\prime} \alpha ( \uptau )$. Then, the unit root quantile autoregressive model can be estimated
Denote with $w_t = \Delta Y_t$, $u_{t \uptau} = Y_t - X_t^{\prime} \alpha ( \uptau )$, under the unit root hypothesis and other regularity assumptions
where
is the long-run covariance matrix of the bivariate Brownian motion and can be written as $\boldsymbol{\Sigma}_0 ( \uptau ) + \boldsymbol{\Sigma}_1 ( \uptau ) + \boldsymbol{\Sigma}_1 ( \uptau )^{\prime}$, where
In addition we have that,
Notice that the random function $\displaystyle n^{- 1 / 2} \sum_{t=1}^{ \floor{nr} } \psi_{\uptau} \left( u_{t \uptau} \right)$ converges to a two-parameter process such that $B_{ \psi }^{\uptau} (r) = B_{\psi} ( \uptau, r )$, which is partially a Brownian bridge in the sense that for fixed $r$, $B_{ \psi }^{\uptau} (r) = B_{ \psi } ( \uptau, r)$ is a rescaled Brownian bridge, while for each $\uptau$, $\displaystyle n^{- 1 / 2} \sum_{t=1}^{ \floor{nr} } \psi_{\uptau} \left( u_{t \uptau} \right)$ converges weakly to a Brownian motion with variance equal to $\tau (1 - \uptau)$. Moreover, for each fixed pair $( \uptau_0, r )$ we have that
Let $\widehat{\alpha}( \uptau ) = \left( \widehat{\alpha}_0( \uptau ), \widehat{\alpha}_1( \uptau ) ,..., \widehat{\alpha}_p( \uptau ) \right)$ and $D_n = \mathsf{diag} \left( \sqrt{n}, n , \sqrt{n}, ... , \sqrt{n} \right)$.
Now define with $B^{\mu}_w (r) = B_w (r) - \int_0^1 B_w$ is a demeaned Brownian motion. Therefore, we can derive for any fixed $\uptau$, the test statistic $t_n( \uptau )$, that is, the quantile regression counterpart of the well-known ADF t-ratio for a unit root. Thus, it can be shown that the limiting distribution of $t_n( \uptau )$ is nonstandard and depends on nuisance parameters $\left( \sigma_w^2, \sigma_{w \psi } ( \uptau) \right)$ as $B_w$ and $B_{ \psi }^{\uptau}$ are correlated Brownian motions.
Notice that the limiting distribution of the t-ratio $t_n( \uptau )$ can be decomposed as a linear combination of two (independent) distributions, with weights determined by a long-run (zero frequency) correlation coefficient that can be consistently estimated. Following phillips1990statistical we have that
where $\lambda_{ \omega \psi} ( \uptau ) = \sigma_{ w \psi } ( \uptau )/ \sigma_w^2$ and $\boldsymbol{B}_{\psi.w}^{ \uptau }$ is a Brownian motion with variance $\sigma^2_{ \psi . w} (\uptau) = \sigma^2_{\psi} ( \uptau ) - \sigma^2_{ w \psi } ( \uptau ) / \sigma_w^2$ and is independent of $\boldsymbol{B}^{\mu}_w$.
Therefore, the limiting distribution of the t-ratio $t_n ( \uptau )$ can be decomposed as below
For notation convenience we can also rewrite the Brownian motions $\boldsymbol{B}_w(r)$ and $\boldsymbol{B}^{\uptau}_{ \psi . w }(r)$:
and for the corresponding demeaned Brownian motions as below
where $\boldsymbol{W}_{1}(s)$ and $\boldsymbol{W}_{2}(s)$ are standard Brownian motions and are independent stochastic processes. Note also that $\sigma^2_{\psi} ( \uptau ) = \uptau ( 1 - \uptau)$, and the limiting distribution of $t_n( \uptau )$ can be written as below
where
Consider the scalar-valued random variable $y_t$ that follows the nonlinear time series model (see, uematsu2019nonstationary)
where $g : \mathbb{R} \times \mathbb{R}^{\ell} \to \mathbb{R}$ is a known regression function and the error term $u_t$ is a zero-mean stationary process. We denote with $g ( x_t , \beta ) \equiv g_t (\beta)$. Furthermore, the regressor $x_t$ follows an $I(1)$ time series process as defined below
where $x_0 = 0$ and the innovation sequence $v_t$ is assumed to be stationary with mean zero.
Then, the $\uptau-$th quantile NQR estimator $\hat{\theta}_n (\uptau) = \left( \hat{\alpha}_n ( \uptau ), \hat{\beta}_n(\uptau) \right)$ is obtained by the following minimization problem:
where $\rho_{\uptau} ( \mathsf{u} ) = \mathsf{u} \big( \uptau - \mathbf{1} \left( \mathsf{u} < 0 \right) \big)$ is the check function and $\psi_{\uptau} ( \mathsf{u} ) = \uptau - \mathbf{1} \left( \mathsf{u} < 0 \right)$. Next, we impose some parametric assumptions regarding the distribution of the error term $u_t$.
Let $\alpha_{0 \uptau} := \alpha_0 + F^{-1} ( \uptau )$ and define the new parameter vector $\theta_{0 \uptau} = \big( \alpha_{0 \uptau}, \beta_{0 \uptau}^{\prime} \big)^{\prime}$. We then, rewrite the error term as below
Notice that $\mathbb{E} \psi_{\uptau} ( u_{t \uptau} ) = 0$ and $\mathcal{Q}_{ u_{t \uptau} } ( \uptau ) = 0$, where $\mathcal{Q}_{ u_{t \uptau} } ( \uptau ) = 0$ is the $\uptau-$th quantile of $u_{t \uptau}$. Therefore, in view of $\theta_{0 \uptau}$ and $u_{t \uptau}$, the argument of the check function can be written as below
Then, for the error terms $u_{t \uptau}$ and $v_t$, we construct two partial sum processes as below
Following the framework of uematsu2019nonstationary, we employ the proof of Lemma 7.5 as below. In particular, we denote with
where
Notice that to avoid technical problems in taking conditional expectations, we consider truncation of $\upgamma_t(\lambda)$ at some finite number $m > 0$ and denote with $W_{nm} (\lambda) = \sum_{t=1}^n w_{tm} ( \lambda )$, where
Further define with
and
Therefore, using the notation above, we will derive the probability limit of $W_n(\lambda)$ through the following steps. First, we compute the limiting variable of
by considering the simultaneous limits $n \to \infty$ and $m \to \infty$. Next, we check the asymptotic equivalence of $W_{nm} (\lambda)$ and $\bar{W}_{nm} (\lambda)$. Finally, we verify that the asymptotic equivalence of $W_n(\lambda)$ and $W_{nm}(\lambda)$, that is, the effect of truncation by $I_{tm}(\lambda)$ is asymptotically negligible.
The predictive regression model has been widely used for investigating the predictability of asset returns. Applications include estimation and inference methodologies for predictive regression model based on a parametric (see, kostakis2015Robust) or nonparametric (see, kasparis2015nonparametric) conditional mean functional form. The current literature mainly covers linear conditional mean autoregressive and predictive regression models as well as dynamic conditional mean predictive regression models. When one is interested for testing the presence of quantile predictability, related studies have considered the implementation of the model based on the conditional quantile functional form. Related studies include galvao2009unit, maynard2023inference, lee2016predictive, fan2019predictive, katsouris2022asymptotic, cai2022new and liu2023unified. The inclusion of model intercepts in both the predictive regression model as well as the autoregressive model either based on a conditional mean or on a conditional quantile functional forms imposes additional challenges when developing inference techniques for the $\beta$ parameter based on the usual likelihood ratio test or Wald-type statistics; especially when the autoregressive coefficient is nearly-integrated (see, liu2019asymptotic, wang2022testing, yang2023unified).
Various studies in the literature employ the likelihood-based approach to investigate the properties and asymptotic theory of corresponding estimators and test statistics for nonstationary autoregressive processes and predictive regression models. These likelihood approaches include the weighted least squares, the empirical likelihood methodology as well as the profile likelihood approach. In particular, zhu2014predictive, liu2019empirical, liu2019unified and yang2021unified consider an application of the empirical likelihood approach which induces uniform inference across the different degrees of persistence and confidence regions for the coefficients of the linear predictive regression. Moreover, chen2013uniform propose a uniform inference approach (see, also chen2009bias\footnote{Related studies include chen2009restricted, chen2012restricted as well as zhao2000restricted.}) which is valid regardless of the persistence of regressors using the weighted least squares approximate restricted likelihood (WLSRL).
Furthermore, liu2019unified propose predictability tests that correspond to testing the null hypothesis of no statistical significance for the model coefficients of a predictive regression model with additional covariates the difference of the predicting variable. Specifically, these predictability tests are based on the unified empirical likelihood methodology that rejects the null hypothesis $H_0: \beta_1 = \beta_2 = 0$ or testing the hypothesis $H_0: \beta_1 = 0$ and $H_0: \beta_2 = 0$. Moreover, under the presence of known nonzero model intercept in the predictive regression model using weighted score equations and the sample-splitting methodology it can be proved that nuisance parameter-free inference holds regardless of the persistence properties of regressors. However, implementing the sample-split approach into a multivariate nonstationary vector of regressors remains a computational challenging task as this would require to split the data into blocks equal to the dimension of the vector-valued regressor and then use each block to construct score equations with respect to each variable in the vector. A similar difficulty exists when one considers the IVX filtation to a multivariate or even high-dimensional setting (see, XuGuo2022).
Both methodologies fundamentally are trying to solve a similar problem as the presence of model intercept in the predictive regression model causes a nonstandard limiting distribution results especially when $X_t$ is nonstationary. To overcome this problem these two different methods consider a different approach. On the one hand the empirical likelihood method takes the first difference of the paired data such that $\left\{ Y_{t+1} - Y_t \right\}_{t=1}^n$ and $\left\{ X_{t+1} - X_t \right\}_{t=1}^n$ but a new issue arises as the corresponding error sequence $\left\{ U_{t+1} - U_t \right\}_{t=1}^n$ violates the independence assumption. Therefore, to fix this issue a split sample methodology is applied which means that the data is splited into two parts, the first difference transformation is applied based on a larger lag (i.e., the value of $m = \floor{n/2}$) which ensures $m-$dependence holds (i.e., approximate or asymptotic independence as $m \to \infty$), and lastly the empirical likelihood estimation is applied. The asymptotic theory for the empirical likelihood confidence interval for $\beta_0$ is then proved to follow standard distribution results (i.e., converges to a $\chi^2-$distribution).
On the other hand, the IVX filtration proposed by magdalinos2009econometric operates differently but still with the aim to solve a similar issue, to avoid having nonstandard asymptotic results under the presence of nonstationary regressors. More specifically, the IVX filtration is an endogenous instrumentation procedure which is used as an alternative to the OLS estimator for the unknown model coefficient $\beta_0$. In particular, the IVX instruments are constructed with the main consideration of obtaining less persistence regressors than the original series $\left\{ X_t \right\}$ without having to consider a first-difference transformation based on a large value of the lag as well as the application of sample splitting to ensure that the distribution of the innovations of the predictive regression have independent realizations. This methodology implies that asymptotic theory for the unknown model coefficient follows a mixed gaussian distribution and the corresponding Wald test follow a $\chi^2-$distribution.
In summary, both methodologies have proved to produce uniform inference for unit root moderate deviations in the case of a local-to-unity autoregressive specification. The empirical likelihood plus weighted least squares induces uniform inference regadless of persistence properties and the presence of an intercept in the autoregressive specification of regressors. Similarly the IVX instrumentation also induces uniform inference regardless of nonstationary properties (as in kostakis2015Robust) while more recently is has been proved by Magdalinos2022uniform that is indeed also robust to the presence of an intercept as well as more general nonstationarity properties. Similar aspects hold for conditional quantile functional forms. In particular, lee2016predictive develops a framework for the implementation of the IVX instrumentation in quantile predictive regression models which is found to have standard asymptotic theory while more recently liu2023unified proved that their approach is indeed robust to both the presence of intercept in the autoregressive specification and nonstationarity of the $\left\{ X_t \right\}$ process. Currently, katsouris2022asymptotic investigates the extension of the framework proposed by Magdalinos2022uniform, to quantile autoregressions and quantile predictive regression models. Lastly, another aspect is the development of a framework for testing the presence of structural breaks in quantile predictive regression models, which is proposed by katsouris2023structural (see, also qu2008testing).
We first focus on the quantile predictive regression model which can be extended within a dynamic setting. Therefore, within the framework proposed by katsouris2021optimal, katsouris2023statistical, we consider the estimation procedure of VaR of a financial institution which is given by expression under the additional assumption that the predictors are generated via a local-to-unity parametrization. Before doing that we study carefully the framework proposed by lee2016predictive.
To begin with the standard conditional mean predictive regression model is given by
with $\mathbb{E} \left( u_{0t} | \mathcal{F}_{t-1} \right) = 0$. Furthermore, the $p-$dimensional vector of predictors $x_{t-1}$ is generated via the following LUR specification
where $n$ is the sample size, and $C = diag \left( c_1,...,c_p \right)$ is the matrix for the degree of persistence. For the purpose of this paper, we consider that predictors employed for the estimation of the risk measures, belong only to one of the two persistence classes below
The innovation structure allows for linear process dependence for $u_{xt}$ and imposes a conditionally homoscedastic martingale difference sequence condition for $u_{0t}$ as is a standard practise in the related predictive regression literature. Under the related regulatory conditions, the usual FCLT holds (as per phillips1992asymptotics)
Next, we consider the corresponding quantile predictive regression model, which requires that
Furthermore, in order to define the innovation structure of the quantile predictive regression system, we define the piecewise derivative of the loss function such that $\psi_{\tau} ( u ) = \tau - \mathbf{1} \left\{ u < 0 \right\}$.
Thus, we have that $u_{0 t\tau} = u_{0t} - F^{-1}_{ u_0} (\tau)$ where $F^{-1}_{ u_0} (\tau)$ denotes the unconditional $\tau-$quantile of $u_{0t}$. Then, the corresponding FCLT for the quantile predictive regression system is
Therefore, the following assumption holds.
Note in the case that we allow for persistence regressors in the QR specification, then nonstandard limit theory applies, thus is worth examining the particular case. We follow the derivations in lee2016predictive.
where $\phi_{\tau}( u ) = u \left( \tau - \mathbf{1}_{ \left\{ u < 0 \right\} } \right)$, for $\tau \in (0,1)$, represents the asymmetric QR function and $X_{t-1} = \left( 1, x_{t-1}^{\prime} \right)^{\prime}$ includes the intercept and the set of regressors $x_{t-1}$. Following lee2016predictive, we use the following normalization matrices according to the persistence properties of the regressors, that is,
We consider the instrumental variable $Z_{tn}$ which is based on the IVX methodology proposed by phillipsmagdal2009econometric. The IVX instrument is constructed as
After demeaning the QR specification we obtain the transformed model without model intercept
Then, to consider the corresponding limiting distribution of the IVX-QR estimators, we use the following embedded normalizations as below
and using the notation $\alpha \wedge \delta = \text{min} \left( \alpha, \delta \right)$,
The next theorem gives the limit theory of the IVX-QR estimator under various degrees of persistence (see, lee2016predictive and katsouris2023structural).
The above Theorem shows that we obtain a mixed normal limiting distribution regardless of the persistence properties of the predictors (see, Theorem 3.1 in lee2016predictive).
Consider the following partial sum process
where $\omega^{*}_{\psi}$ is the long-run variance of $\psi_{\uptau} \left( \epsilon_{j \uptau} \right)$ and the partial sum process follows an invariance principle and converges weakly to a standard Brownian motion $W(r)$. Choosing a continuous functional $h(.)$ that measures the fluctuation of $Y_n(r)$, then a robust test for cointegration can be constructed based on $h \big( Y_n(r) \big)$. By the continuous mapping theorem, under regularity conditions and the null of cointegration,
In principle, any metric that measures the fluctuations in $Y_n(r)$ is a natural candidate for the functional $h$. The classical KS typr or CvM type measures are of particular interest. A robust test for cointegration can then be constructed based on
where $\widehat{ \omega }^{*}_{\psi}$ is a consistent estimator of $\omega^{*}_{\psi}$.
Consider the following cointegration model
where $\boldsymbol{x}_t$ is a $k-$dimensional vector of integrated regressors, $\boldsymbol{z}_t = \left( 1, \boldsymbol{x}_t^{\prime} \right)^{\prime}$ and $u_t$ is mean zero stationary innovation sequence. The quantile regression estimator of the cointegrating vector can be obtained by solving the following (where $p = k +1$)
where $\rho_{ \uptau } (u) = u \big( \uptau - \mathbf{1} \left\{ \mathsf{u} < 0 \right\} \big)$.
In order to derive the limit distribution of the quantile regression estimator of the cointegrating vector we follow the approach as presented in xiao2009quantile. Let $f(.)$ and $F_{y|x}(.)$ be the pdf and the CDF of $u_t$ and denote the inverse of the check function with $\psi_{ \uptau } (u) = \big[ \uptau - \mathbf{1} \left\{ u < 0 \right\} \big]$. The following assumptions set the background theory in order to support the econometric analysis of this section.
Furthermore, we partition $\boldsymbol{\Omega}_{ww}$ as below
Note that the asymptotic behaviour of $n^{-1} \sum_{t=1}^n \boldsymbol{x}_t \psi_{ \uptau } \big( u_t ( \uptau ) \big)$. Under Assumption (ref) the following asymptotic result holds
where $\lambda_{ v \psi_{ \uptau } } $ is the one-sided long-run covariance between $\boldsymbol{v}_t$ and $\psi_{ \uptau } \big( u_t ( \uptau ) \big)$.
Due to the nonstationarity in $\boldsymbol{x}_t$, the two model parameters in the parameter vector $\widehat{ \boldsymbol{\theta} } ( \uptau ) = \left( \hat{ \boldsymbol{\alpha} }(\uptau), \hat{ \boldsymbol{\beta} }(\uptau)^{\prime} \right)^{\prime}$ have different rates of convergence. In particular, the estimate of the cointegrating vector $\hat{ \boldsymbol{\beta} }(\uptau)^{\prime}$ converges at rate $n$, while the intercept $\hat{ \boldsymbol{\alpha} }(\uptau)$ converges at rate $\sqrt{n}$. Thus, in order to account for the different convergence rates among the model intercept and the estimate of the cointegrating vector we use the normalization matrix $\mathcal{ \boldsymbol{D} }_n = \mathsf{diag} \left( \sqrt{n} , n \boldsymbol{I}_k \right)$, where $\boldsymbol{I}_k$ is the $k \times k$ identity matrix. The limit distribution of the quantile regression estimator for the cointegration model is given by the Theorem (ref).
In this section, we are interested to propose a robust econometric framework for robust estimation and inference for the quantile regression in the case of a cointegrating model. Firstly, we note that the asymptotic behaviour of quantile regression-based inference procedures depends on the limit distribution of $\widehat{ \boldsymbol{\theta} } ( \tau )$. However, as shown by Theorem (ref), the limiting processes $\boldsymbol{B}_v ( r )$ and $B_{ \psi_{\uptau} } ( r )$ are correlated Brownian motions whenever contemporaneous correlation between $v_t$ and $\psi_{ \uptau } \big( u_t( \uptau ) \big)$ exists.
Moreover, super-consistency, $\widehat{ \boldsymbol{\theta} } ( \uptau )$ is second-order biased and the miscentering effect in the limit distribution is reflected in $\Delta_{v \psi}$. Consequently, the distribution of the test based on the quantile regression residual will be dependent on nuisance parameters. Therefore, in order to restore the asymptotic nuisance parameter free property of the inference procedure, we need to modify the original quantile regression estimator so that we obtain a mixed normal limit distribution. Usually in the literature, there are two approaches one can take to achieve the restoration of the the asymptotic nuisance parameter free property. For example, one approach is the implementation of a nonparametric fully-modification of the original quantile regression estimator, and a second approach is the parametrically augmented quantile regression implementation using leads and lags. In order to develop a fully-modified quantile cointegrating regression estimator one can follow the approach proposed by phillips1990statistical.
We first decompose the asymptotic distribution given by (ref) into the following two components:
where
such that $B_{\psi . v}$ is a Brownian motion with variance $\omega^2_{ \psi . v}$ as defined above. Moreover, note that $\boldsymbol{B}_{\psi . v}$ is independent of $\boldsymbol{B}_v(r)$ and the first term in the above decomposition, $\left[ \int_0^1 \boldsymbol{B}_v^{\mu} \boldsymbol{B}_v^{\prime \mu} \right]^{-1} \left[ \int_0^1 \boldsymbol{B}_v^{\mu} d\boldsymbol{B}_{\psi . v} \right]$, is a mixed Gaussian variate. The basic idea of fully-modification on $\widehat{ \boldsymbol{\beta} } ( \uptau )$ is to construct a nonparametric correction to remove the second term in the above decomposition.
Therefore, to facilitate the nonparametric correction, we consider the following kernel estimators of $\Omega_{vv}$, $\Omega_{v \psi}$, $\lambda_{v \psi}$, $\lambda_{vv}$, such that
where $k(.)$ is the kernel function defined on the interval $[-1,1]$ with $k(0) = 1$, and $M$ the bandwidth parameter satisfying the property that $M \to \infty$ and $M / n \to 0$. The sample covariance matrices are
We define the following nonparametric fully modified quantile regression estimators as below
and $\widehat{\lambda}_{v \psi}^{+} = \widehat{\lambda}_{v \psi} - \widehat{\lambda}_{v v} \widehat{\Omega}_{vv}^{-1} \widehat{\Omega}_{v\psi} $.
Similar to the fully modified OLS estimators, the fully modified quantile regression estimator of the cointegrating vector has a mixed normal distribution in the limit.
Note that the fully modified quantile regression estimator and the resulting asymptotic mixed normal asymptotic distribution facilitates statistical inference (such as the use of the Wald test) based on quantile cointegrating regression. Therefore, the classical inference problem of linear restrictions on the cointegrating vector $\beta$ is such that $\mathbb{H}_0: \mathcal{R} \boldsymbol{\beta} = r$, where $\mathcal{R}$ denotes an $(q \times k)$ matrix of linear restrictions and $r$ denotes a $q-$ dimensional vector.
Thus, under the null hypothesis, and the assumptions of Theorem (ref), we have that
where $\mathcal{N} \big( \boldsymbol{0}, \boldsymbol{I}_q \big)$ represents a $q-$dimensional standard Normal distribution.
Define with $\mathcal{ \boldsymbol{Q} }_x = \sum_{t=1}^n ( \boldsymbol{x}_t - \bar{\boldsymbol{x}} ) ( \boldsymbol{x}_t - \bar{\boldsymbol{x}} )^{\prime}$. Then, a the classical Wald test in the case of quantile cointegrating regression can be constructed via the following expression
where $f_{y|x} \left( \widehat{ F_{y|x}^{-1} ( \uptau)} \right)$ and $\widehat{ \omega}_{\psi.v}$ are consistent estimators of $f_{y|x} \left( F_{y|x}^{-1} ( \uptau) \right)$ and $\omega_{\psi.v}$ respectively. Then, the asymptotic distribution of the Wald statistic is summarized via the following Theorem.
where $\chi_q^2$ is a centred Chi-square random variable with $q-$degrees of freedom.
Next, we show that
In particular, the above result holds because of the following
Then, notice that by billingsley1968convergence we have that
which implies that
Similarly, we can show that
Therefore, we have that
As a result we obtain that
Therefore, by the convexity Lemma of pollard1991asymptotics and arguments presented by knight1998limiting, notice that the functional $\mathcal{Z}_n(v)$ is minimized at $\widehat{v} = \boldsymbol{D}_n \big( \widehat{\boldsymbol{\theta}}(\uptau) - \boldsymbol{\theta}(\uptau) \big)$ while $\mathcal{Z}(v)$ is minimized at
by Lemma A of knight1989limit we have that
Therefore, by Theorem 1, the asymptotic distribution of the random measure $n \big( \widehat{\beta}(\uptau) - \beta(\uptau) \big)$ can be expressed as below
Furthermore, we have that
are consistent estimates of $\boldsymbol{\Omega}_{vv}$, $\boldsymbol{\Omega}_{v \psi_{\uptau} }$, $\lambda_{v \psi_{\uptau} }$ and $\lambda_{vv}$. Thus, we obtain the following weak convergence
The asymptotic distribution of $\widehat{\beta}( \uptau )$ in quantile cointegrating regressions is mixture normal. Another interesting problem in the quantile cointegration regression model is the hypothesis test on constancy of the cointegrating vector $\boldsymbol{\beta} ( \uptau ) = \bar{\boldsymbol{\beta} }$ , over $\uptau \in \mathcal{T}_{\iota}$, where $\bar{ \boldsymbol{\beta} }$ is a vector of unknown constants. A natural preliminary candidate for testing constancy of the cointegrating vector is a standardized version of $\left( \widehat{ \boldsymbol{\beta} } - \bar{\boldsymbol{\beta}} \right)$. Under the null hypothesis, we have that
Denote with $\widehat{\boldsymbol{\beta}}$ as a preliminary estimator of $\bar{ \boldsymbol{\beta} }$ and consider the following process
Then, under the null hypothesis $\mathcal{H}_0: \boldsymbol{\beta} (\uptau) = \bar{\boldsymbol{\beta}}$ it holds that
where $B^{*}_{\epsilon}$ is the Brownian motion limit for the partial sum process of $\epsilon_t$. Therefore, we may test varying-coefficient behaviour based on the KS statistic. Notice that for the matrix $\boldsymbol{V}_{xx} = \displaystyle \int_{0}^{\infty} e^{r \boldsymbol{C}_p} \boldsymbol{\Omega}_{xx} e^{r \boldsymbol{C}_p} dr$, and thus by applying integration by parts we can obtain the following formula $\boldsymbol{C} \boldsymbol{V}_{xx} + \boldsymbol{V}_{xx} \boldsymbol{C} = - \boldsymbol{\Omega}_{xx}$ (see, for example expression (49) in magdalinos2009econometric).
In general, lets suppose we have that
Define with such that $u_t(\uptau) = u_t - F^{-1}(\uptau)$. Then, for the error terms of the quantile predictive regression model the following invariance principles hold
In addition, suppose that the following conditions hold:
Based on the martingale difference sequence assumption on $u_t$, it follows that $U_n^{\psi} ( \lambda, \uptau ) \overset{ d }{ \to } U^{\psi} ( \lambda, \uptau )$, where $U^{\psi} ( \lambda, \uptau )$ is viewed as a Brownian motion with variance $\lambda \omega_{\psi}(\uptau)^2 = \lambda \uptau ( 1 - \uptau)$ for a fixed value of $\uptau$. Thus, for each fixed pair $( \lambda, \uptau )$, it holds that $U^{\psi} ( \lambda, \uptau )$ is distributed as $\mathcal{N} \big( 0, \lambda \omega_{\psi}(\uptau)^2 \big)$. A standard assumption in the nonstationary time series analysis is that $V_n$ converge weakly jointly with $U_n^{\psi}$ to a vector Brownian motion. Notice that for the development of the asymptotic theory we keep the argument $\uptau$ fixed (see, also cho2015quantile). Another example is the framework proposed by cho2015quantile who discuss quantile cointegration in ADL models. In that case, it holds
Therefore, one may be interested to develop a framework for detecting multiple breaks in quantile piecewise regression models or quantile cointegrated regressions. For instance if the number of break points $m$ is given, then estimating their locations and the $(m + 1)$ piecewise quantile autoregressive models at a specific quantile $\uptau \in (0,1)$ can be done via solving
Further related studies on structural break testing for quantile regressions include qu2008testing and fanchange2023 within a stationary time series environment and katsouris2023structural within a nonstationary time series environment. Although both of these streams of literature correspond to structural break testing on quantile-dependent coefficients within a full-sample period. On the other hand, a different application would correspond to an implementation of the monitoring framework within based on a quantile regression, in the case of risk measures (see the recent study by hoga2023monitoring).
The next two sections cover briefly key results from two additional relevant topics for quantile regression models, that is, (i) estimation in high dimensional quantile regressions and, (ii) specification testing in quantile regressions. Therefore, developing powerful tests for the correct specification\footnote{Recall that an omnibus test is a statistical test for which the alternative hypothesis is a negation of the null hypothesis. For example, the particular family of tests allows to test the monotonicity of conditional moments. On the other hand, a different approach in the literature, that is, the unconditional backtesting can be decomposed into an estimation risk component and a model risk component which implies that under correct specification of the parametric VaR model, almost surely, the model risk vanishes (see, escanciano2010backtesting). In contract, under misspecification, the model risk does not vanish and has a negligible effect on the unconditional test.} of parametric conditional quantiles over a possibly continuous range of quantiles and under general conditions on the underlying data-generating process is a good model validity practice. The main idea of regression model misspecification is explained by stinchcombe1998consistent as below:
Let $\mathcal{S} := \left\{ f(., \theta ) : \mathbb{R}^k \to \mathbb{R} | \theta \in \Theta \right\}$, $ \Theta \subset \mathbb{R}^p, p \in \mathbb{N}$. The model is correctly specified for $\mathbb{E} \left( Y | X \right)$ when $f( X , \theta_0 )$ is a version of $\mathbb{E} \left( Y | X \right)$ for some $\theta_0 \in \Theta$. For instance, $\hat{\theta}_n$ can be a quantile regression estimator. Denote with $\epsilon:= Y - f(X, \theta^{*} )$, the correct specification of $\mathcal{S}$ for $\mathbb{E} \left( Y | X \right)$ is equivalent to $e := \mathbb{E} \left( \epsilon | X \right) = 0$, almost surely. Under suitable conditions on $Y$ and $f$, $e$ is an integrable function of $X$, that is, $e \in L^p ( X ) := L^p ( \Omega, \sigma(X), P )$ for some $p \in [1, \infty]$.
Following the framework of belloni2011 consider a response variable $y$ and $p-$dimensional covariates $x$ such that the $u-$th conditional quantile function of $y$ given $x$ is
where $\mathcal{U} \subset (0,1)$ is a compact set of quantile indices. Recall that the $u-$th conditional quantile $F_{ y_i | x_i }^{-1} ( u| x_i )$ is the inverse of the conditional distribution function $F_{ y_i | x_i }^{-1} ( y | x_i )$ of $y_i$ given $x_i$. Furthermore, we consider the case where the dimension $p$ of the model is large, possibly much larger than the available sample size $n$, but the true model $\beta (u)$ has a sparse support
having only $s_u \leq s \leq n / \mathsf{log} ( n \cup p )$ nonzero components for all $u \in \mathcal{U}$. Therefore, the corresponding population coefficient $\beta (u)$ is known to minimize the criterion function
In other words, given a random sample $\left\{ (y_1, x_1),..., (y_n, x_n) \right\}$, the quantile regression estimator of $\beta(u)$ is defined as a minimizer of the empirical analoge given by
The main challenge of the statistical problem under examination is that in high-dimensional settings, particularly when $p \geq n$, ordinary quantile regression is generally inconsistent, which motivates the use of penalization in order to remove all, or at least nearly all, regressors whose populations coefficients are zero, thereby possibly restoring consistency. Then, the $\ell_1-$penalized quantile regression estimator $\widehat{\beta}(u)$ is a solution to the following optimization problem:
where $\hat{\sigma}^2_j = \mathbb{E}_n \big[ x_{ij}^2 \big]$, which ensures that the conditional variance of the error term has a bounded variance and thus excluding infitine variance cases as the related distribution function of the innovation sequence generating the data mechanism under examination. To show the result of consistency, it suffices to show that for any $\epsilon > 0$, there exists a sufficiently large $C$ such that
In other words, this inequality implies that with probability at least $1 - \epsilon$, there is a local minimizer $\tilde{\beta} ( \uptau )$ within the shrinking ball $\big\{ \beta ( \uptau ) + a_n c, \left\lVertc\right\rVert_2 = C \big\}$ such that $\left\lVert \tilde{\beta} ( \uptau ) - \beta ( \uptau ) \right\rVert_2 = \mathcal{O}_p ( a_n )$. Therefore, the proof can be obtained by showing that the following term is positive
\paragraph{Proof.}
Therefore, using the similar arguments in $I_1$, we have that
Consequently, by the Cauchy-Schwarz inequality, and $t > k$, we have that
Further applications of high dimensional quantile time series regression models are studied by belloni2023high while in a cross-sectional setting relevant frameworks are proposed by he2013quantile, carlier2016vector, he2021smoothed, lee2023complete and zhang2023bootstrap.
Some further applications of quantile time series regressions are discussed in felix2023some.
Following the framework of he2021smoothed, consider a univariate response variable $y \in \mathbb{R}$ and a $p-$dimensional covariate vector, the primary goal here is to learn the effect of $\boldsymbol{x}$ on the distribution of $y$. Let $F_{y| \boldsymbol{x}}$ be the conditional distribution function of y given $\boldsymbol{x}$. The dependence between $y$ and $\boldsymbol{x}$ is then fully characterized by the conditional quantile functions of $y$ given $\boldsymbol{x}$, denoted as $F^{-1}_{y| \boldsymbol{x}}$, for $0 < \uptau < 1$.
Consider a linear quantile regression model at a given $\uptau \in (0,1)$, that is, the $\uptau-$the conditional quantile function is
where $\boldsymbol{\beta}_0(\uptau) = \big( \beta_1^{*},..., \beta_p^{*} \big)^{\prime} \in \mathbb{R}^p$ is the true quantile regression coefficient.
Let $Q( \boldsymbol{\beta} ) = \mathbb{E} \big[ \widehat{Q}(\boldsymbol{\beta}) \big]$ be the population quantile loss function. Under mild conditions, $Q(.)$ is twice differentiable and strongly convex in a neighbourhood of $\boldsymbol{\beta}$ with Hessian matrix such that
is the random noise and $f_{\boldsymbol{\epsilon | \boldsymbol{x} } }(.)$ is the conditional density of $\epsilon$ given $\boldsymbol{x}$.
On the other hand, by the first-order condition the population parameter $\boldsymbol{0}$ satisfies the moment condition
Moreover, the Bahadur representation can be used to establish the limiting distribution of the estimator or its functionals. We consider a fundamental statistical inference problem for testing the linear hypothesis $\mathcal{H}_0: \langle \boldsymbol{a}, \boldsymbol{b}_0 \rangle = 0$, where $\boldsymbol{a} \in \mathbb{R}^p$ is a deterministic vector that satisfies that defines a linear functional of interest. It is then natural to consider a test statistic that depends on $\sqrt{n} \langle \boldsymbol{a}, \widehat{\boldsymbol{b}}_h \rangle$. Based on the non-asymptotic result (finite sample) of the following theorem, we establish a Berry-Esseen bound for the linear projection of the conquer estimator.
where $\sigma_h^2 = \sigma_h^2(\boldsymbol{a}) = \boldsymbol{a}^{\prime} \boldsymbol{J}_h^{-1} \mathbb{E} \big[ \big\{ \mathcal{K}_h ( - \epsilon ) - \uptau \big\}^2 \boldsymbol{x} \boldsymbol{x}^{\prime} \big] \boldsymbol{J}_h^{-1} \boldsymbol{a}$, where $\Phi(.)$ denotes the standard normal distribution function. Moreover, it holds that
An advantage of convolution smoothing is that it facilitates conditional density estimation for the quantile regression process. Assume that $Q_y ( \uptau | \boldsymbol{\mathcal{X}} ) = F^{-1}_{y| \boldsymbol{\mathcal{X}} } = \langle \boldsymbol{\mathcal{X}}, \boldsymbol{\beta}_0 (\uptau) \rangle$ for all $\uptau \in (0,1)$. Under mild regularity conditions,
In particular, the inverse conditional density function plays an important role in, for example, the study of quantile treatment effects through modelling inverse prosperity scores. Therefore, by the linear conditional quantile model assumption, we have that
Recall that the conquer estimator $\hat{\boldsymbol{\beta}}_h(\uptau)$ satisfies the first-order condition such that $\nabla \widehat{Q}_h \big( \widehat{\boldsymbol{\beta}}_h(\uptau) \big) = \boldsymbol{0}$. Therefore, taking the partial derivative with respect to $\uptau$ on both sides, it follows by the chain rule that
Consequently, the inverse functions $1 / f_{y_i | \boldsymbol{x}_i} \big( \boldsymbol{x}_i \boldsymbol{\beta}_0(\uptau) \big)$ can be directly estimated by $\boldsymbol{x}_i^{\prime} \frac{ \partial \widehat{\boldsymbol{\beta}}_h(\uptau) }{ \partial \uptau }$.
In this section, we follow carefully the framework proposed by escanciano2010specification which corresponds to the class of consistent parametric specification tests for dynamic conditional quantiles. A recent relevant framework is proposed by horvath2022consistent. The conditional moment restriction test approach considers directly the model estimates (for the corresponding model functional form we impose), which is a more flexible method allowing to consider e.g., different estimation methods (parametric versus nonparametric such as the NW estimation method or the nonparametric series estimation).
Let $W_{t-1} = \left( Y_{t-1}, Y_{t-2}, Z^{\prime}_{t-1}, Z^{\prime}_{t-2} \right)$ be the set of variables included in the quantile regression model for estimating the CoVaR and VaR risk measures. Assuming that the conditional distribution of $Y_t$ given $W_{t-1}$ is continuous, we define the $a-$th conditional VaR of $Y_t$ given $W_{t-1}$ as the $\mathcal{F}_{t-1}-$measurable function $q_{\alpha}\left( W_{t-1} \right)$ satisfying the following probability statement
For parametric VaR inference, one assumes the existence of a parametric family of functions
and proceeds to make VaR forecasts using the model $\mathcal{M}$. Statistical inference is based on the crucial assumption that $q_{\alpha} \in \mathcal{M}$, which implies that there exists some $\theta_0 \in \Theta$ such that $m_{\alpha} \left( W_{t-1}, \theta_0 \right) = q_{\alpha} \left( W_{t-1} \right)$ almost surely. For instance, in parametric model the nuisance parameter $\theta_0$ belongs to $\Theta$, with $\Theta$ a compact set in a Euclidean space $\mathbb{R}^p$. More specifically, parametric VaR models are a commonly used methodology in the literature, since the the functional form $m_{\alpha} \left( W_{t-1}, \theta_0 \right)$ along with the parameter $\theta_0$ provides a parsimonious representation of the VaR risk measure based on the available information set of the decision maker at the particular time period.
Here we present some examples which demonstrate the choice of the family of models required for the parametric estimation of the VaR. A suitable model is the linear quantile regression (LQR)
Following the framework of escanciano2010specification we can obtain some insights regarding the specification testing methodology. In particular, the proposed test statistics are based on the fact that $q \in \mathcal{M}$ is characterized by the infinite set of conditional moment restrictions given as below
Note that the moment condition given by expression (ref) is equivalent to the condition of expression (ref). Therefore, the test statistic and hypothesis testing of interest is
against the nonparametric alternatives given by
where $\Psi \left( \epsilon \right) = \mathbf{1} \left( \epsilon \leq 0 \right) - \alpha$. Therefore, to simplify notation we denote with $\Psi_{ \alpha, t } \left( \theta \right) \equiv \Psi_{ \alpha } \big( Y_t - m \left( \mathcal{I}_{t-1}, \theta \right) \big)$ and $m_{t-1} \left( \theta \right) \equiv m \left( \mathcal{I}_{t-1}, \theta \right)$. Thus, under the null hypothesis, and assuming that a continuity condition for $m( . )$ holds, then $m_{t-1} \left( \theta_0 \right)$ is identified as the $\alpha-$th quantile of the conditional distribution of $Y_t$ given $\mathcal{I}_{t-1}$, for all $\alpha$. Then, to characterize $\mathbb{H}_0$ by the infinite number of unconditional moment restrictions we write
Therefore, given a sample $\left\{ \left( Y_t, \mathcal{I}_{t-1}^{\prime} \right)^{\prime}: 1 \leq t \leq n \right\}$ and a parameter value $\theta$, one can consider the following quantile-marked empirical process indexed by $x \in \mathbb{R}^d$
in order to construct a formal specification testing procedure based on the functional form proposed by escanciano2010specification. A commonly used estimator for the unknown parameter $\theta_0$ is the quantile regression estimator, which can be obtained as the solution of the corresponding minimization problem with $\rho_{\alpha} \left( \epsilon \right) = - \Psi_{ \alpha } \left( \epsilon \right) \epsilon $.
We briefly discuss the main results of the asymptotic theory proposed by escanciano2010specification. More precisely, in order to establish the limit distribution of the quantile-marked empirical process, under the null hypothesis, $\mathbb{H}_0$, one needs the following notation. Define the family of conditional distributions $F_x( y ) : = \mathbb{P} \big( Y_t \leq y | \mathcal{I}_{t-1} = x \big)$. Moreover, let
where $\left\{ ( Y_t, Z_t^{\prime} )^{\prime} : t \in \mathbb{Z} \right\}$ is a strictly stationary and ergodic process. Then, under the null hypothesis, $\mathbb{H}_0$, we have that $\left\{ \Psi_{\alpha, t} \left( \theta_0 \right), \mathcal{F}_t \right\}$ is a mds \ $\forall$ \ $\alpha \in \Pi$, such that, $\Pi \subset [0,1]$. Furthermore, it holds that the finite-dimensional distribution of $R_n$ converge to those of a multivariate normal distribution with mean a zero mean vector and variance-covariance matrix given by the covariance function as expressed below
where $v_1 = ( x_1^{\prime}, \alpha_1 )$ and $v_2 = ( x_2^{\prime}, \alpha_1 )$ represent generic elements of $\Pi$.
Moreover, the following condition holds
Then, $Q_n ( \alpha )$ converges weakly to a Gaussian process $Q(.)$ with zero mean and covariance function
Denote with $\mathcal{T}_m := \left\{ \alpha_j \right\}_{j = 1}^{ \alpha}$ the points in the grid which represent a set of different quantile levels, with $\alpha_1 < ... < \alpha_m$. Let $W_{\text{exp}}$ be the $n \times n$ matrix with $w_{\text{exp}, t, s} = \text{exp} \left( -\frac{1}{2} \left| \mathcal{I}_{t-1} - \mathcal{I}_{s-1} \right|^2 \right)$ and let $\psi$ be the $n \times n$ matrix with typical elements given by $\psi_{ij} = \psi_{\alpha j } \left( Y_i - m \left( \mathcal{I}_{i-1}, \theta_n \right) \right)$. Hence, the CvM test statistic is computed as below
where $\psi_{.j}$ denotes the $j-$th column of $\Psi$, with $m$ fixed points such as $\left\{ \alpha_j \right\}_{j = 1}^{ \alpha}$ are deterministic.
A different perspective is presented in the framework proposed by firpo2022gmm which in practise it corresponds to a misspecification testing methodology in conditional quantile regression models based on a GMM estimation approach. Specifically, consider the linear quantile regression model (QR) as
where $Y \in \mathcal{Y} \subset \mathbb{R}$ is the outcome variable, $X \in \mathcal{X} \subset \mathbb{R}^K$ is a $K-$dimensional vector of covariates and $Q_{Y|X}$ is the conditional quantile function of $Y$ given $X$ and $\tau \in (0,1)$ is the fixed quantile. Thus, a quantile functional form such as the slope vector or a subset of its components can be written as below
Moreover, by imposing the above restriction on the first $K_1$ components of $\beta$ such that
where $X_1 \in \mathbb{R}^{K_1}$ are the first components of $X$ and $X_2 \in \mathbb{R}^{K_2}$ are the remaining components. Therefore, if $X = ( X_1, X_2^{\top} )^{\top}$, where $X_1$ is a scalar variable, we can compute $\widehat{\theta}$ by minimizing the weighted distance between the QR estimator $\beta_1 (\tau)$ and parametrization $\mathsf{g}_1(\theta, \tau)$ using the following objective function
Then, firpo2022gmm impose the following assumptions to study the limiting properties of $\widehat{\theta}_{GMM}$.
First we show the uniform convergence of $\widehat{Q}_n (\theta)$ to $Q(\theta)$. Notice that
Next, we establish the asymptotic normality of the MD-QR estimator. In particular, by the first order condition, the estimator $\widehat{\theta}_{MD}$ satisfies the following condition
Moreover, using a Taylor expansion of $\mathsf{g} \left( \widehat{\theta}_{MD}, \tau \right)$ around $\theta_0$ we obtain
where $\bar{\theta}$ is a line segment between $\widehat{\theta}_{MD}$ and $\theta_0$. Therefore, we obtain that
By Theorem 3 in angrist2006quantile, it follows that $\sqrt{n} \left( \widehat{\beta}(\tau) - \beta (\tau) \right) \overset{d}{\to} \boldsymbol{Z}_{\beta}( \tau)$, is a mean-zero Gaussian process in $\ell^{\infty} ( \mathcal{T}, \boldsymbol{R}^K )$ which is the set of bounded functions that map $\mathcal{T}$ into $\mathbb{R}^K$ with covariance kernel $\Sigma ( \tau, \tau^{\prime} )$. Therefore, the parametric rate $\sqrt{n}$ times the above term, converges in distribution to
by the continuous mapping theorem for functionals. Since $\boldsymbol{Z}$ is a linear functional of $\boldsymbol{Z}_{\beta}$, a mean-zero Gaussian process, is also mean-zero and Gaussian. Its variance is equal to
Moreover, firpo2022gmm propose a minimum distance quantile regression (MD-QR) estimator as an alternative to GMM estimators. The MD-QR is defined by minimizing the integrated weighted distance, over a subset of quantiles $\mathcal{T}$, between the standard QR estimator and the parametric model for some of the coefficients of interest. Let $\widehat{\beta}(\tau)$ denote the standard QR estimator. The MD-QR estimator of $\theta$
where $\widehat{W}(\tau) \in \mathbb{R}^{K \times K}$ is a weight function that converges in probability to a positive-definite valued function $W(\tau)$. Notice that if one is interested about the correct specification of the function $\mathsf{g} (.,.)$ which is used to model some of the components of $\beta(\tau)$. Since we are using the GMM objective function, hence a simple specification test can be applied.
Proof.
and by the delta method, we have that
Therefore, we can write the normalized objective function as below
since under correct specification it holds that $\textcolor{blue}{\beta(\tau) \equiv \mathsf{g}(\theta_0, \tau)}$ for all $\tau \in \mathcal{T}$. Therefore, by an application of the functional delta method and by the consistency of $\widehat{W}(\tau)$ for $W(\tau)$, we obtain
A relevant framework is proposed by tu2022nonparametric. In particular, their model specification testing approach is constructed by a comparison of the nonparametric estimator of $\mathsf{g} ( x,z )$ and the corresponding parametric estimator of $\mathsf{g}_0 ( x, z ; \theta_0 )$ under the null hypothesis. The following test statistic is used
where $\pi_1 (x)$ and $\pi_2 (z)$ are positive integrable weight functions and the parametric quantile estimator $\hat{\theta}_n$
and $p_n( x, z )$ is a weighting function. Furthermore, due to the presence of nonstationary regressors as well as the kernel density estimators involved makes deriving the asymptotic theory quite challenging. Therefore, we consider an alternative test statistic to $R_n$ which is a corresponding smooth version. Recall that the quantile estimator $\widehat{\mathsf{g}}(x,z)$ of $\mathsf{g}(x,z)$ such that
where $\psi_{\tau} = \tau - \mathbf{1} \left\{ t < 0 \right\}$, and there exists a smooth version $\widetilde{\mathsf{g}}_0 \left( x, z; \theta \right)$ of $\mathsf{g}_0 \left( x, z; \theta \right)$ such that
Therefore, a smoothed version of $R_n$, with $\mathsf{g}_0 (x,z; \widehat{\theta}_n )$ replaced by $\widetilde{\mathsf{g}}_0 (x,z; \widehat{\theta}_n )$ is
and, if we let $p_n( x,z ) = \left\{ \sum_{t=1}^n K_1 \left( \frac{x_t - x}{h_1} \right) K_2 \left( \frac{z_t - z}{h_2} \right) \right\}^2$. Define with
Thus, $\frac{ d_n }{n h_1 h_2} \mathcal{S}_n$ converges in distribution to a local time of an Ornstein-Uhlenbeck process.
In this Section, we present the main results of the framework proposed by wang2012specification who construct a statistical methodology for parametric specification testing in cointegrating models (see, also kasparis2012dynamic).
Consider the nonlinear cointegrating regression model given by
where $u_t$ is a stationary error process, and $x_t$ is a nonstationary regressor. The testing hypothesis of interest is expressed as below
To test the null hypothesis, we use the following kernel-smoothed test statistic
involving the parametric regression residuals $\hat{u}_{t+1} = y_{t-1} - f( x_t, \hat{\theta} )$, where $K(x)$ is a nonnegative real kernel function and $h$ the window bandwidth satisfying $h \equiv h_n \to 0$ as $n \to \infty$. Note that the difficulty for developing an asymptotic theory for $S_n$ stems from the presence of the kernel weights $K \left( (x_t - x_s ) / h \right)$. The behaviour of these weights depends on the self intersection properties of $x_t$ in the sample. Therefore, to establish asymptotics for $S_n$, we need to account for the related limit theory which involves self-intersection local time of a Gaussian process.
These class of specification testing methodologies as in the study of duffy2021estimation (among others) compare the functional forms of the parametric against the alternative hypothesis of a correctly specified nonparametric functional form (currently the literature corresponds to the conditional mean functional form but extensions to a conditional quantile functional form can potentially be possible).
\paragraph{Parametric Regression}
Consider the OLS estimator of $\left( \mu, \gamma \right)$ of the regression model given by
where the unknown model parameters are obtained as below:
where $f$ has a known nonlinear functional form. The single regressor included in the nonparametric predictive regression model is assumed to be predetermined or $\mathcal{F}_{t-1}-$measurable. Then, duffy2021estimation show that if $x_t$ is stationary, the OLS estimator is asymptotically normal, whereas in the case that $x_t(n)$ exhibits strong dependence structure, then the OLS estimator has a non-standard limiting distribution. Specifically, the authors assume a linear array of the form
where $\left\{ \xi_t \right\}_{ t \in \mathbb{Z} }$ satisfies the following assumption.
Furthermore, based on the local-to-unity parametrization duffy2021estimation consider in their framework two distinct classes of persistence processes, that is, nearly integrated (NI) processes and mildly integrated (MI) processes. Both MI and NI processes can be defined in terms of triangular arrays
where $v_t$ is a stationary process and $\kappa_n > 0$ with $\kappa_n < \infty$, so that the autoregressive coefficient becomes approximates to unity as $n$ grows. Both NI and MI processes describe highly persistent autoregressive processes, which have a root in the vicinity of unity. As a matter of fact both processes have been extensively used to study the behaviour of inferential procedures under local departures from unit roots and robust inferential procedures (e.g., see phillips2007limit and kostakis2015Robust). The crucial difference between NI and MI processes concerns the assumed growth rate of the sequence $\kappa_n$. In particular, NI processes are defined by $\kappa_n / n \to c \neq 0$, with the consequence that $n^{-1 / 2 } x_{ \floor{ n r}}$ converges weakly to an Ornstein-Uhlenbeck procees. On the other hand, MI processes have $\kappa_n / n \to 0$, which implies that the process $x_t(n)$ moves closer to the neighbourhood of stationarity, which implies that FCLT cannot be used to derive the asymptotics of functionals of these processes duffy2021estimation.
\paragraph{Nonparametric regression}
Consider the nonparametric estimation of the predictive regression
The unknown function is estimated nonparametrically via the Nadaraya-Watson (NW) estimator
and the corresponding local linear (LL) estimator is expressed as
where $K_{th} (x) := K \left( \frac{x_{t-1} - x}{ h_n } \right)$. Then, the testing hypothesis of interest is that the true regression function $m$ belongs to a certain parametric family. More specifically, the null hypothesis is formulated as
where $f$ is a known function. The alternative hypothesis implies that no such $\mu$ and $\gamma$ exist.
In other words, the specification test proposed by duffy2021estimation allows to test the null hypothesis of parametric fit under the presence of NI regressor, which closely resembles that of gao2009specification, who propose a studentized U-statistic formed of kernel-weighted OLS regression residuals. The specification test proposed by duffy2021estimation tests the null hypothesis $\mathbb{H}_0$ by comparing the parametric OLS and the corresponding nonparametric (kernel) estimates of the $m$ function. In particular, the model specified in (ref) can be estimated nonparametrically at each $x$ as below
Denote with $\left( \hat{\mu}, \hat{\gamma} \right)$ the OLS estimates of $\left( \mu, \gamma \right)$, then the statistical distance corresponds to a comparison between the fit provided by the parametric and local nonparametric estimates of the model via an ensemble of $t$statistics of the following form
where $\hat{\sigma}_u^2 (x): = \bigg( \sum_{t=2}^n K_{th} (x) \bigg)^{-1} \sum_{t=2}^n \big[ y_t - \hat{\mu} - \hat{\gamma} f(x_{t-1} \big] K_{th} (x)$. Therefore, under the null hypothesis of no misspecification both the parametric and nonparametric estimators converge to identical limits and so for each $x \in \mathbb{R}$, the following limit is obtained
Following escanciano2009econometrics who consider specification testing in the context of nonlinear cointegrating regressions with nonlinear dependence. Specifically, the near-epoch dependence condition when imposed on the error term of the regression model is considered to impose a form of nonlinear dependence in time series regressions models.
In particular, the above conditions ensure that
According to escanciano2009econometrics, the above theorem provides sufficient conditions for cointegrated variables to be generated by a nonlinear error correction model. Furthermore, the fully modified OLS and FM-OLS estimation approach involves a two-step procedure in which in the first step the following econometric specification $\Delta y_t = \psi_1 \Delta x_t + \gamma z_{t-1} + v_t$, is estimated by OLS. Moreover, in the second step semiparametric corrections are made for the serial correlations of the residuals $z_t$ and for the endogeneity of the $x-$regressors. Under general conditions the fully modified estimator is asymptotically efficient. A particular case of the more general form of a nonlinear parametric cointegration is of the form $y_t = f( x_t, \beta ) + v_t$, where $x_t$ is a $( p \times 1 )$ vector of $I(1)$ regressors, $v_t$ is a zero-mean stationary error term and $f( x_t, \beta )$ a smooth function of the process $x_t$, known up to the finite-dimensional parameter vector $\beta$. Thus, in the classical asymptotic theory, the properties of estimators of $\beta$, i.e., the nonlinear least squares estimator (NLSE), depend on the specific class of functions where $f( x_t, \beta )$ belongs. The rate of convergence of the NLSE is class-specific and in some cases involves random scaling.