Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
98,799 characters · 11 sections · 87 citation commands
Predictive Quantile Regression with Mixed Roots and Increasing Dimensions: The ALQR Approach
Predictive quantile regression (QR) identifies the impact of predictors on a set of conditional quantiles of a response variable. It provides richer information on the heterogeneous distributional prediction. For example, the conditional quantile prediction of stock returns receives much attention in finance since the tail quantile information has a crucial role in measuring risk. Many economic state variables are employed to predict stock returns and the number of candidate predictors is often large. When a large number of predictors are available, researchers encounter the inevitable model selection issue. A good model selection can improve forecasting performance but the opposite can also occur. Considering the importance of model selection in practice, we need constructive guidance for empirical applications.
In this paper we propose the adaptive lasso for predictive quantile regression (ALQR). Although there exists a large volume of literature on predictive mean regression of equity returns (see, e.g.\ campbell1987stock, fama1988dividend, hodrick1992dividend, cenesizoglu2012return, andersen2020pricing among others), predictive QR is relatively understudied. cenesizoglu2008distribution is an early paper on predictive QR, and maynard2011inference, lee2016predictive, fan2019predictive, gungor2019exact, and cai2022new recently develop inference methods in predictive QR with nonstationarity and heteroskedasticity. The proposed method is different from these approaches. We consider predictive QR with an increasing number of mixed root predictors and address two important problems raised in the stock returns data. First, the prediction power of each predictor can vary over different quantiles. By adapting the lasso, we allow the model selection based on real data not by the researcher's discretion. Second, the predictors widely used in predicting equity returns are composed of stationary, local unit root, and cointegrated processes. For example, we plot the time-series of two predictors, dividend price ratio (dp) and default yield spread (dfr), in Figure (ref). Even a simple eyeballing test easily confirms their different levels of persistence. (We conduct more informative estimation procedures in Section (ref)). We provide a unified adaptive lasso framework that allows those mixed root predictors. We show that the estimator converges to the true parameter value at different convergence rates and that the faster convergence rates make the adaptive lasso more efficient in the model selection.
Naturally, some technical challenges arise. Since the seminal paper of tibshirani1996regression, the lasso has been intensively studied in various fields of statistical analysis. However, most studies have been focusing on the i.i.d.\ sample and it has not been a long time since more studies have been conducted with dependent data (see the references in the related literature section below). Furthermore, we allow nonstationary predictors, which impose an additional difficulty in the formal analysis of the proposed adaptive lasso. We tackle these issues by considering a simple model that only contains unit-root predictors first. Once establishing the desired properties of ALQR in this model, we generalize the model so that it includes all stationary, local unit root, and cointegrated predictors.
The contributions of this paper are two-fold. First, to the best of our knowledge, this is the first paper to study predictive QR with an increasing number of mixed root predictors. We propose the adaptive lasso for predictive quantile regression (ALQR) and derive the convergence rates, model selection consistency and asymptotic distributions of ALQR under some regularity conditions. As a by-product, we also prove that the standard QR estimator is consistent under both mixed roots (including the local unit roots) and the increasing dimension of the predictors, which is new in the literature. Second, we conduct an empirical analysis of the stock returns data and find that ALQR can improve prediction performance over existing alternatives across different quantiles. We apply the ALQR method along with the existing alternatives to the data set. The results confirm that ALQR shows better prediction performance across different quantiles, particularly at higher quantiles. As illustrated in Section 3, ALQR can be readily applicable to other applications.
The rest of the paper is organized as follows. This section finishes with a review of the relevant literature. In section 2, we introduce predictive QR models formally and define the ALQR estimator. In section 3, we investigate the performance of ALQR using the out-of-sample quantile prediction problem of stock returns. We study the theoretical properties of ALQR in sections 4 and 5. In section 7, we conduct some Monte Carlo simulation experiments. Section 7 concludes. All technical proofs are relegated to the appendix.
\paragraph{Related Literature} The lasso has been extensively studied for cross-sectional data. Recently, there has been development in lasso procedures with dependent data. basu2015regularized exploit the spectral properties of the stationary time series design and investigate the regularity condition of the lasso that leads to the non-asymptotic bounds and the consistency results. kock2015oracle investigate the oracle property of the lasso in a stationary vector autoregression model. adamek2020lasso provide an inference procedure based on the debiased/desparsified lasso under the near-epoch dependence assumption. chernozhukov2021lasso propose a penalty selection algorithm with weakly dependent data and the post-selection inference procedure. wu2016performance and wong2020lasso analyze the lasso with non-sub-Gaussian processes that allow a heavy-tail distribution. medeiros2016l1 show the asymptotic properties of the adaptive lasso for stationary high-dimensional time series models. For the cointegrated models, kock2016consistent shows the oracle property of the adaptive lasso in the autoregression model. Interestingly, he finds that the unit root test can be incorporated into the model selection procedure by adopting the Dickey-Fuller form of autoregression. liao2015automated propose a shrinkage estimator with multiple penalty terms to select the rank of cointegration and estimate the parameters simultaneously in a vector error correction model. liang2019determination, zhang2019identifying, and onatski2018alternative investigate the same model in the high-dimensional setting.
koo2020high recently use lasso to improve the prediction of stock returns. It is shown that lasso significantly reduces forecasting mean squared errors even with a mixture of stationary, unit-root, and cointegrated variables. However, the conventional lasso method may not have model selection consistency and the oracle property as shown by meinshausen2004consistent and fan2001variable. The adaptive lasso proposed by zou2006adaptive improves the performance of the lasso. Instead of imposing the same penalty weight on all candidate parameters, the adaptive lasso penalizes each parameter proportionally to the inverse of its initial estimate. With a proper choice of the tuning parameter ${\Greekmath 0115} _{n}$, the adaptive penalty weights for the irrelevant variables approach infinity, whereas those for the relevant variables converge to constants. lee2021lasso apply the adaptive lasso to a predictive mean regression framework. Similar to koo2020high, predictors are allowed to have different degrees of persistence and cointegration. lee2021lasso find that the adaptive lasso and a newly proposed twin adaptive lasso outperform the alternative methods in terms of predictor selection consistency and out-of-sample mean squared errors.
Some effort has been also made to investigate model selection and model estimation in QR under the i.i.d.\ samples. The $\ell_1$-penalized method in the QR framework has been studied for high-dimensional data analysis (see, e.g.\ portnoy1984asymptotic, portnoy1985asymptotic, knight2000asymptotics, koenker_2005, li20081, lee2018oracle and belloni2011_L1). To overcome the problem of inconsistent model selection (fan2014adaptive; wang2012quantile), recent studies have further considered the adaptive lasso in QR. wu2009variable discuss how to conduct model selection for QR models using SCAD and the adaptive lasso method. zheng2013adaptive establish the oracle property for an adaptive lasso QR model with heterogeneous error sequences. zheng2015globally study a globally adaptive lasso method for ultra high-dimensional QR models.
\paragraph{Notation} Let $\Vert \cdot \Vert$ and $\Vert \cdot \Vert_0$ denote $\ell_2$-norm and $\ell_0$-norm, respectively. For a matrix $S$, $\Vert S \Vert$ represents the spectral norm. Let $F_a(\cdot)$ and $f_a(\cdot)$ denote a cumulative distribution function (CDF) and a probability density function (pdf) of a generic random variable $a$. Let $S>0$ denote a generic positive definite matrix $S$. Let ${\Greekmath 0115}_{\min }\left( S\right) $ and ${\Greekmath 0115} _{\max }\left( S\right) $ denote the smallest and largest eigenvalues of $S$. We use $O_p(1)$ and $o_p(1)$ when a sequence is bounded in probability and converges to zero in probability, respectively. The $O(1)$ and $o(1)$ denote the non-stochastic counterparts.
In this section, we introduce the predictive quantile regression model and the adaptive lasso for quantile regression (ALQR). We will develop the theory for the ALQR in two steps in Section (ref). For a better exposition of the theory, we also propose the model and its estimator in two separate cases: (i) unit-root predictors; and (ii) mixed-root predictors.
Consider a predictive QR model with unit-root predictors:
where $Q_{y_{t}}\left( {\Greekmath 011C} |\mathcal{F}_{t-1}\right) $ is the conditional ${\Greekmath 011C}$-quantile of $y_{t}$, $\mathcal{F}_{t}=\left\{ z_{j}\right\} _{j=-\infty }^{t}$ with $z_{j}=(y_{j},x_{j}^{\prime })^{\prime }$ is a natural filtration, $x_{t}$ is a $p_{n}$-dimensional vector of unit-root predictors with a stationary $O_{p}(1)$ initialization of $v_{0}=\sum_{j=0}^{\infty }D_{vj}{\Greekmath 010F} _{-j}$ following the innovation structure below, and ${\Greekmath 010C}_{0{\Greekmath 011C}}$ is the corresponding true parameter vector.
The innovation of unit root predictors, $v_{t}$, follows a linear process, which is commonly assumed in the predictive regression literature (see phillips2013predictive, phillips2016robust, cai2022new, lee2021lasso for a few recent papers):
where $mds$ denotes martingale difference sequences with respect to the natural filtration.
We allow the dimension of predictors to increase, i.e. $p_n\rightarrow \infty$ as $n \rightarrow \infty$. To make notation simple, we omit subscript $n$ and use $p$ unless it may cause any confusion. We define $u_{t{\Greekmath 011C} }:=y_{t}-Q_{y_{t}}\left( {\Greekmath 011C} |\mathcal{F}_{t-1}\right)$, the deviation of $y_t$ from the conditional ${\Greekmath 011C}$-quantile. If the CDF of $u_{t{\Greekmath 011C}}$ is continuous, we have
where $F_{u_{t{\Greekmath 011C} }}^{-1}$ is the inverse CDF of $u_{t{\Greekmath 011C} } $. It holds by construction that $\Pr \left( u_{t{\Greekmath 011C} }<0|\mathcal{F}_{t-1}\right) ={\Greekmath 011C}$, which is equivalent to $F_{u_{t{\Greekmath 011C} }}^{-1}\left( {\Greekmath 011C} | \mathcal{F}_{t-1}\right) = 0$. To see the equivalence to the unconditional ${\Greekmath 011C}$-quantile, we note that
where $\mathbf{1}(\cdot )$ denotes an indicator function. The equations above imply that both conditional and unconditional ${\Greekmath 011C}$-quantiles of $u_{t{\Greekmath 011C}}$ are zero for any given ${\Greekmath 011C}$. However, it does not mean that two distribution functions are the same, i.e. $F_{u_{t{\Greekmath 011C}_1 }}^{-1}\left( {\Greekmath 011C}_2 | \mathcal{F}_{t-1}\right) \neq F_{u_{t{\Greekmath 011C}_1 }}^{-1}\left( {\Greekmath 011C}_2 \right)$ for ${\Greekmath 011C}_1 \neq {\Greekmath 011C}_2$ in general.
Let ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} }):= {\Greekmath 011C} - \mathbf{1} \left( u_{t{\Greekmath 011C} }<0\right)$. It is easy to find that ${\Greekmath 0120} _{{\Greekmath 011C} }(u_{t{\Greekmath 011C} })$ is uncorrelated to any predetermined regressor. That is, ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })$ is uncorrelated with any past innovations, $v_{t-j}$, for $j\geq 1$:
by the law of iterated expectations and the fact that $E\left( {\Greekmath 0120}_{{\Greekmath 011C} }(u_{t{\Greekmath 011C} })|\mathcal{\ F}_{t-1}\right) ={\Greekmath 011C} -\Pr \left( u_{t{\Greekmath 011C}}<0|\mathcal{F}_{t-1}\right) =0$. In our settings, however, we allow the QR-induced regression errors ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })$ to be contemporaneously correlated with the innovations of unit root sequences. This is commonly assumed in cointegration and predictive regression literature, inducing a potential second order bias arising from the one-sided correlation, see, e.g., xiao2009quantile.
We now define the ALQR that minimizes the penalized QR objective function as follows:
where ${\Greekmath 011A} _{{\Greekmath 011C} }(u):=u({\Greekmath 011C} -\mathbf{1}(u<0))$. Following zou2006adaptive, we define the tuning parameter of the penalty term as
where ${\Greekmath 0115}_n$ is a sequence converging to infinity, ${\Greekmath 0121} _{j}:=|{\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}|^{{\Greekmath 010D} }$ with a positive constant ${\Greekmath 010D}$, and ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$ is a first-step consistent estimator for ${{\Greekmath 010C}}_{0{\Greekmath 011C} ,j}$. The penalty ${\Greekmath 0115}_{n,j}$, which contrasts with the penalty in the standard lasso, is an adaptive weight and it has an inverse relationship with ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$. The idea of using a consistent first-step estimator as an adaptive weight in ${\Greekmath 0115}_{n,j}$ is to avoid over-penalizing important predictors (or under-penalizing irrelevant predictors). In this paper, we use the ordinary QR estimator (koenker1978regression) below as a first-step estimator:
The consistency of the ordinary QR estimator is provided in Section (ref).
We now introduce a QR model with mixed-root predictors. We assume the predictors of the model have different degrees of persistence: $I(0)$, local unit roots, and cointegration. This model is the most relevant in practice. The model in the previous section can be seen as a special case of it.
Let $z_{t}$, $x_{t}^{c}$, and $x_{t}$ be vectors of stationary,\footnote{In this paper, we use covariance (weak) stationarity unless noted otherwise. The linear process assumption below implies covariance stationarity of $z_{t}$. In addition, the first element of $z_{t}$ is $1$ so that ${\Greekmath 010C} _{0{\Greekmath 011C},1 }^{z}$ plays a role of the intercept. } cointegrated, and local unit-root predictors whose lengths are $p_z$, $p_c$, and $p_x$, respectively. Let $p:= p_z+p_c+p_x$. Given a sample of $\{y_{t},z_{t},x_{t}^{c},x_{t}\}_{t=1}^{n}$, the QR model with mixed roots is defined as follows:
The cointegrated system in $x_{t}^{c}$ has the triangular representation by phillips1991optimal: for $x_{1t}^c$ and $x_{2t}^c$ whose dimensions are $p_1\times 1$ and $p_2\times 1$, respectively, \begingroup\allowdisplaybreaks
\endgroup where $R_{2}=I_{p_{2}}+c_{2}/n$ with $c_{2}=\mathrm{diag}\left( \check{c} _{1},\ldots,\check{c}_{p_{2}}\right)$, $L$ is the lag operator, and the vector $v_{1t}^{c}$ is the vector of $I$(0) cointegrating residuals. The cointegrated regressors $x_{t}^{c}$ are decomposed into $x_{1t}^c$ and $x_{2t}^c$. The local-to-unity parameter of $x_{2t}^c$ is assumed to be $\check{c}_{j}\in \left( -\infty ,\infty \right)$ for $j=1,\ldots,p_2$. Thus, the cointegrated system in $x_t^c$ includes both stationary and nonstationary local unit root regions. Using (ref), we can easily characterize the cointegration relations in $x_{t}^{c}$ by the matrices $A\equiv(I_{p_{1}},-A_{1})$ and $A_{1}$.
The near integrated process $x_{t}=(x_{t1},...,x_{tp_{x}})$ is defined in a similar way:
where $R_{x}=I_{p_{x}}+c_{x}/n$ with $c_{x}=\mathrm{diag}\left(\tilde{c} _{1},...,\tilde{c}_{p_{x}}\right)$ and $\tilde{c}_{j}\in \left( -\infty ,\infty \right) $. Again, the local-to-unity specification includes the unit root process as a special case when $\tilde{c}_{j}=0$.
Let $X_{t}:=(z_{t}^{\prime },x_{t}^{c \prime},x_{t}^{\prime })^{\prime }$ and ${\Greekmath 010C} ^{\ast }:=({\Greekmath 010C} ^{z^{\prime }},{\Greekmath 010C} ^{c^{\prime }},{\Greekmath 010C} ^{x^{\prime }})^{\prime }$. Define a $p_c$-dimensional vector $v_{t}^{c}:=({v_{1t}^{c}}^{\prime },{v_{2t}^{c}}^{\prime })^{\prime }$ and a $p$-dimensional vector $e_{t}:=\left( z_{t}^{\prime },{v_{t}^{c}}^{\prime },v_{t}^{\prime }\right)^{\prime }$. Abusing notation on ${\Greekmath 010F}_t$, we assume that $e_t$ follows a linear process:
Note that this assumption implies a covariance stationary $O_{p}(1)$ initialization of $e_{0}=\left( z_{0}^{\prime },{ v_{0}^{c}}^{\prime },v_{0}^{\prime }\right) ^{\prime }=\sum_{j=0}^{\infty }D_{ej}{\Greekmath 010F} _{-j}$.
Similar to Section (ref), the QR-induced regression errors, ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })$, are uncorrelated to any predetermined regressor but allowed to be contemporaneously correlated with the innovations of unit root sequences $x_{2t}^{c}$ and $x_{t}$, the stationary predictor $z_t$ as well as the cointegrating residuals $v_{2t}^{c}$.
We define the ALQR for Model (ref) as follows:
where ${\Greekmath 0115} _{n,j}={{\Greekmath 0115} _{n}}/{\Greekmath 0121} _{j}$, ${\Greekmath 0121} _{j}=|{\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}|^{{\Greekmath 010D} }$, ${\Greekmath 010D} $ is a constant, and ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$ is a first-step consistent estimator. We recommend using the ordinary QR estimator for ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$. The consistency of QR will be discussed in Section (ref).
In this section, we consider the quantile prediction problem of stock returns and illustrate the usefulness of the proposed ALQR method. We use an updated version of the data set in welch2008comprehensive, which ranges from January 1952 to December 2019. It is composed of 816 monthly observations of the US financial and macroeconomic variables. Our goal is to predict the quantiles of the excess stock returns ($y_t$) using 12 predictors. The variable names are summarized in Table (ref). See also welch2008comprehensive for more details on these variables.
We first check the mixed-root property of the predictors. In Figures (ref)--(ref), we plot each predictor using monthly observations. The predictors (dp, dy, ep, bm, dfy, ntis, lty, and tbl) collected in Figure (ref) are highly persistent and have quite different patterns from the other four predictors (svar, dfr, ltr and infl) in Figure (ref). The first-order autoregression coefficients reported in Table (ref) also suggest that predictors have heterogeneous degrees of persistence. We also conduct the Johansen test for cointegration. The results show that the cointegrating rank is $3$ in all of the 804-month rolling windows.\footnote{The Johansen cointegration test results are presented in the appendix, Table (ref).} Thus, the data fit into the mixed-root model structure discussed in Section (ref).
We next investigate the performance of ALQR in the quantile prediction problem of stock returns. For comparison, we also apply three alternative estimators in the literature: the QR estimator without selecting predictors (QR), the quantile lasso (LASSO), and the unconditional quantile without using any predictor (QUANT). For the tuning parameter of LASSO and ALQR, we use both the Bayesian information criteria (BIC) and the generalized information criteria (GIC), which are explained in detail in Section (ref). We evaluate the performance of quantile prediction using the final prediction error (FPE) and the out-of-sample $R^2$ following lu2015jackknife. The FPE measures the out-of-sample quantile prediction errors:
where $\hat{y}_s$ is a prediction for ${\Greekmath 011C}$-quantile of $y_s$ and $S$ is the number of out-of-sample predictions. Note that FPE(${\Greekmath 011C}$) averages the quantile loss function. It is different from the standard mean squared prediction error. At each quantile ${\Greekmath 011C}$, a smaller FPE implies a better quantile prediction. The second measure of the performance is the out-of-sample $R^2$:
where $\bar{y}_{s}$ is the unconditional ${\Greekmath 011C}$-quantile of $y$ (QUANT) from the training sample. Note that the out-of-sample $R^2$ measures the performance of each method relative to QUANT. For example, a positive $R^2$ indicates that the prediction error is smaller than that of QUANT, and a larger $R^2$ implies a better prediction. By definition, $R^2$ of QUANT is always zero.
Tables (ref)--(ref) summarize the prediction results with BIC. They are based on the one-step-ahead prediction for the last 12 and 24 periods of the sample, respectively. The results with GIC are similar and left in the appendix. We consider 5 different quantiles, ${\Greekmath 011C}=0.05, 0.1, 0.5, 0.9$, and $0.95$. For each quantile, the best performance results, i.e., the smallest FPE and the highest $R^2$, are marked in bold. In addition to performance measures, we also report the average number of selected predictors and the corresponding tuning parameter value ${\Greekmath 0115}$ for LASSO and ALQR.
Overall, ALQR shows quite satisfactory results. First, ALQR performs the best in prediction across all quantiles, except for two designs: ${\Greekmath 011C}=0.95$ at the 12-period experiment and ${\Greekmath 011C}=0.05$ at the 24-period experiment. In the former design, LASSO performs slightly better than ALQR but the difference is negligible. In the latter design, QR performs better than the alternatives. Second, the relative performance of ALQR to QUANT improves substantially for higher quantiles. Looking at $R^2$, we find that ALQR predicts 44% (${\Greekmath 011C}=0.9$) and 32% (${\Greekmath 011C}=0.95$) better than QUANT at the 12-period design and 28% (${\Greekmath 011C}=0.9$) and 31% (${\Greekmath 011C}=0.95$) at the 24-period design, respectively. QR does not show particularly better prediction performance than QUANT, which is in line with the well-known result in the mean stock return prediction literature (see, e.g., welch2008comprehensive). Third, ALQR selects fewer predictors than LASSO on average. This result is consistent with the report in literature that the standard lasso overselects relevant predictors in practice (See, e.g., wasserman2009high). In Section (ref), we prove the oracle properties of ALQR including model selection consistency. In Section (ref), we provide further numerical evidence that the number of selected variables by ALQR is closer to the true sparsity via Monte Carlo simulations.
Finally, we make some remarks from the empirical perspective. In Figure (ref), we plot the coefficient estimates of ALQR at the 12-period experiment. To make the graph readable, we divide the predictors into two groups (persistent and stationary) and restrict the quantiles into ${\Greekmath 011C}=0.1, 0.5$, and $0.9$. First, we observe that predictors have heterogeneous effects across quantiles. This provides useful information on predicting the distribution of stock returns. For high stock returns (${\Greekmath 011C}=0.9$), default yield spread (dfy), dividend price ratio (dp), dividend yield ratio (dy), stock variance (svar), and inflation (infl) show larger effects. However, long term yield (lty), net equity expansion (ntis), treasury bill rates (tbl), and stock variance (svar) show large coefficients for low stock returns (${\Greekmath 011C}=0.1$). Second, the magnitude of the coefficients at the median is relatively smaller than that at both tails. This result coincides with the previous findings in the literature (see, e.g., welch2008comprehensive and fan2019predictive) that it is more difficult to predict the center part of the stock return distribution. The positive but small $R^{2}$'s of ALQR at ${\Greekmath 011C}=0.5$ in Tables (ref)--(ref) also reflect such an aspect. Third, the coefficient of stock variance (svar) swings from a large positive value at ${\Greekmath 011C}=0.9$ to a large negative number at ${\Greekmath 011C}=0.1$. Since the return distribution would spread out when the market is volatile, this result is consistent with the stylized fact in the financial market.
This section investigates the asymptotic properties of ALQR. We show that the oracle properties (fan2001variable) of the adaptive lasso remain valid in the QR framework when predictors have both increasing dimensions and various degrees of persistence. This result provides theoretical support for applying the ALQR method to stock returns prediction in Section 3. In Section 4.1, we discuss the case that all predictors are $I(1)$. Section 4.2 then generalizes the results to the case of mixed root predictors.
Consider the QR model in Section (ref) with unit-root predictors only. The following assumptions provide regularity conditions.
For the bound and rate conditions, we define $A_{(t,p)}:=E\left[f_{t-1}(0)x_{t-1}x_{t-1}^{\prime }\right]$ and $B_{(t,p)}:=E\left[ {\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })^{2}x_{t-1}x_{t-1}^{\prime }\right]$.
Note that, for each $t$ and $p$, ${\Greekmath 0115}_{\max}(A_{(t,p)}/p)$ and ${\Greekmath 0115}_{\max}(B_{(t,p)}/p)$ are bounded above since $|f_{t-1}(\cdot)|$ and $|{\Greekmath 0120}_{{\Greekmath 011C}}(\cdot)|$ are uniformly bounded and the variance of $x_{t-1}$ is finite by the design of $v_{t}$. Let $\overline{c}_{A(t,p)}<\infty$ and $\overline{c}_{B(t,p)}<\infty$ be these finite bounds. We next define the following uniform bounds over $t$ for each $n$: $\underline{c}_{A(n,p)}:=\min_{1\leq t\leq n}\underline{c}_{A(t,p)}$, $\overline{c}_{A(n,p)}:=\max_{1\leq t\leq n}\overline{c}_{A(t,p)}$, and $\overline{c}_{B(n,p)}:=\max_{1\leq t\leq n}\overline{c}_{B(t,p)}$.
We make some remarks on these assumptions. Assumption $f$ is a standard restriction in the QR literature on the conditional density of regression residuals. Assumptions $L1$ and $U1$ modify the conventional conditions from stationary QR literature to allow for serial dependence in predictors and regression errors. In particular, we modify Assumption $A.2$ of lu2015jackknife to accommodate nonstationary $x_t$. Based on the model settings in Section (ref), we find that
for $t=\left\lfloor rn\right\rfloor $ with $r\in \left( 0,1\right) $, as $ n\rightarrow \infty$. Therefore, we impose a similar set of restricted eigenvalue conditions for $\frac{E\left[ x_{t-1}x_{t-1}^{\prime }\right] }{t}$. This is a natural extension of the existing conditions in the i.i.d.\ or stationary QR setup to a nonstationary time series model with an increasing dimension. We allow the upper bounds $\overline{c}_{A(t,p)}$, $\overline{c}_{B(t,p)}$ and the lower bounds $\underline{c}_{A(t,p)}$, $\underline{c}_{B(t,p)}$ to depend on the dimension of predictors, hence the upper bounds can diverge to infinity and the lower bounds can converge to zero when $p\rightarrow\infty$ as $n\rightarrow \infty$. However, Assumption $U1$ (ii) imposes further restrictions on the upper bounds: The rate of convergence in $p$ is slow enough so that $(n/p^{\Greekmath 010B})$-consistency of the ALQR estimator can be achieved. Specifically, Assumption $L1$ prevents high correlation between predictor innovations. Note that ${\Greekmath 0115}_{\min}(E[x_{t-1}x_{t-1}]/t)=0$ if there exits perfect multicollinearity between predictor innovations. Also, Assumption $U1$ restricts the strength of the contemporaneous correlation in $v_t$ by limiting nonzero off-diagonal terms in $\Sigma_{vv}$.\footnote{ In the i.i.d.\ mean regression model with highly correlated regressors, it is well known that the lasso performs worse than the ridge estimator. To accommodate the lasso with highly correlated regressors, researchers have also developed variants of the lasso such as the elastic net zou2005regularization and the group lasso yuan2006model. As shown in Figure (ref) in the appendix, however, most predictor innovations in our empirical application are not highly correlated to each other and satisfy this assumption. Furthermore, we investigate the performance of the proposed ALQR estimator when predictor innovations are highly correlated in the simulation experiments. } In Assumption $U1$ (i), the condition $0<{\Greekmath 010B} {\Greekmath 0110} <1$ imposes $a_{n}=\frac{p^{{\Greekmath 010B} }}{n}=\frac{n^{{\Greekmath 010B} {\Greekmath 0110} }}{n}\rightarrow 0$, indicating the restriction on the growth rate of the number of parameters. In the appendix, we discuss the technicality of these restrictions on $p$ and ${\Greekmath 0115} $ in detail. Assumption ${\Greekmath 0115} 1$ collects a set of rate conditions on ${\Greekmath 0115}_n$ to show the asymptotic properties of the ALQR estimators. Assumption ${\Greekmath 0115} 1$ (i) is comparable to the condition of ${\Greekmath 0115}_n/\sqrt{n}\rightarrow 0$ in a stationary time series model with fixed dimensions. This condition allows $(n/p^{\Greekmath 010B})$-rate under the lasso variable selection. Assumption ${\Greekmath 0115} 1$ (ii) presents a necessary condition for model selection consistency of ALQR, which depends on the consistency of the initial estimator for ${\Greekmath 010C}_{0{\Greekmath 011C}}$ (as shown in the appendix, proof of Theorem (ref)). Note that this condition is developed given that the ordinary QR estimator defined in ((ref)) is used as an initial estimator for ${\Greekmath 010C} _{0{\Greekmath 011C} }$. It can be generalized further since we do not require a certain rate of consistency for the initial estimator. If another consistent estimator is used, condition (ii) may be adjusted to the corresponding rate of consistency. Assumption (ref) is for the minimum signal strength of the coefficients in the true active set, which we adopted from \citet*{sherwood2016partially}.
Next we show that the ordinary QR estimator is consistent and can be a qualified initial estimator for ALQR.
It is well-known that the rate of convergence is $\frac{1}{n^{1/2}}$ with fixed $p$ and $\frac{p^{1/2}}{n^{1/2}}$ with increasing $p$ when predictors are $I$(0). This indicates that the information loss (or the degrees of freedom) to estimate the increasing number of parameters is $p^{1/2}$. It is also known that when the predictors are $I$(1), the rate of convergence becomes $\frac{1}{n}$ with fixed $p$. The super-consistency is caused by the stronger signal of $X^{\prime }X$, where $X$ is a $n\times p$ matrix of predictors. To allow for increasing $p$, Lemma (ref) extends the existing result for nonstationary QR and finds that the rate becomes $\frac{p^{{\Greekmath 010B} } }{n}=\frac{p^{1/2}\cdot p^{{\Greekmath 010B}-1/2 }}{n}$, where the additional rate loss $p^{{\Greekmath 010B} -1/2}$ comes from the increasing singularity of $E\left[ x_{t}x_{t}^{\prime }\right] $ as summarized in Assumptions $L1$ and $U1$. In contrast, we have $0<\left[ {\Greekmath 0115} _{\min }\left( E\left[ x_{t}x_{t}^{\prime }\right] \right) \right] <\left[ {\Greekmath 0115} _{\max }\left( E\left[ x_{t}x_{t}^{\prime }\right] \right) \right] <\infty $ for $I$(0) predictors with $p/n\rightarrow 0$.
Taking the QR estimator of Lemma (ref) as an initial estimator for ${\Greekmath 010C} _{0{\Greekmath 011C} }$, we show the consistency of ALQR and model selection consistency under unit roots.
Theorem (ref) confirms that both the ALQR estimator and the QR estimator of Lemma (ref) are $(n/p^{\Greekmath 010B})$-consistent. With a proper choice of the penalty parameter ${\Greekmath 0115}_n$, the rate of consistency for the ALQR estimate is not affected. As we show in the proof of Theorem (ref), this estimation rate is possible under Assumption ${\Greekmath 0115} 1$.
Theorem (ref) shows that the ALQR procedure is an oracle procedure. When the predictors with diverging dimensions are all unit root, the adaptive lasso is at least as good as other “oracle" procedures in terms of variable selection. This is a major difference from the standard lasso, which does not enjoy the oracle properties. As in zou2006adaptive, with a proper choice of ${\Greekmath 0115}_n$, the adaptive lasso procedure can simultaneously estimate relevant coefficients and remove unimportant ones with probability approaching unity. In this study, we further develop the result of zou2006adaptive and generalize the adaptive lasso procedure to a nonstationary QR model. Note that the consistency in model selection can be achieved only under Assumption ${\Greekmath 0115}1$.
In this section, we consider a general QR model discussed in Section (ref):
where $X_{t}=(z_{t}^{\prime },{\ x_{1t}^{c}}^{\prime },{x_{2t}^{c}}^{\prime },x_{t}^{\prime })^{\prime }$ and ${\Greekmath 010C} _{0{\Greekmath 011C} }=({\Greekmath 010C} _{0{\Greekmath 011C} }^{z^{\prime }},{\Greekmath 010C} _{0,1{\Greekmath 011C} }^{c^{\prime }},{\Greekmath 010C} _{0,2{\Greekmath 011C} }^{c^{\prime }},{\Greekmath 010C} _{0{\Greekmath 011C} }^{x^{\prime }})^{\prime }$ are $p$-dimensional vectors. This model allows predictors to have heterogeneous degrees of persistence as well as increasing dimensions. The mixed-root predictors in $X_t$ include stationary ($z_{t}$), local unit root ($x_{t}$), and cointegrated ($x_{1t}^c$, $x_{2t}^c$) processes.
To capture the cointegration relation between $x_{1t}^{c}$ and $x_{2t}^{c}$, we define a transformation matrix $H$ as follows:
Using $H$, the original model can be rewritten as:
where $\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }:=H{\Greekmath 010C} _{0{\Greekmath 011C} }=({\Greekmath 010C} _{0{\Greekmath 011C} }^{z \prime},{\Greekmath 010C} _{0,1{\Greekmath 011C} }^{c \prime}, (A_{1}^{\prime }{\Greekmath 010C} _{0,1{\Greekmath 011C} }^{c}+{\Greekmath 010C} _{0,2{\Greekmath 011C} }^{c})', {\Greekmath 010C} _{0{\Greekmath 011C} }^{x \prime})'$ and $\tilde{x}_{t-1}^{\prime }:=X_{t-1}^{\prime }H^{-1}=(z_{t}^{\prime },{\ v_{1t}^{c}}^{\prime },{x_{2t}^{c}}^{\prime},x_{t}^{\prime })\equiv\left( w_{t}^{(0)\prime },w_{t}^{(1)\prime }\right)$. The $w_{t}^{(0)}=(z_{t}^{\prime },{v_{1t}^{c}}^{\prime })^{\prime }\ $ is a $(p_z+p_1)$-dimensional vector that collects all $I(0)$ processes, and $\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(0)}$ is the corresponding $(p_{z}+p_{1}) $-dimensional QR regression coefficient. The $w_{t}^{(1)}=({x_{2t}^{c}}^{\prime },x_{t}^{\prime })^{\prime }$ is a $(p_2+p_x)$-dimensional vector of all $I(1)$ processes, $\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(1)}$ is the corresponding $(p_{2}+p_{x})$ -dimensional QR regression coefficient. Since the cointegration rank of this system is $p_1$, we let $r:=p_{z}+p_{1}$ be the number of $I(0)$ predictors in $\tilde{x}_{t-1}$. Note that the discussion in Section (ref) can be considered as a special case of the mixed root model with $r=0$. In this section, we generalize the results in Section (ref) and examine a more practical model with $r>0$.
The following assumptions provide additional regularity conditions for mixed roots. Abusing notation slightly, we will use the same notation for the sequences of bounds. Let $M_{n}$ be a diagonal matrix of the form
Assumption $L2$, $U2$, and ${\Greekmath 0115}2$ are analogies to Assumption $L1$, $U1$ and ${\Greekmath 0115}1$ in Section (ref).1. They are essential regularity conditions used to show ALQR's consistency in parameter estimation and model selection under mixed roots. Compared with Assumption ${\Greekmath 0115} 1$, Assumption ${\Greekmath 0115} 2$ are more restrictive on the choice of ${\Greekmath 0115}_n$, because ${\Greekmath 0115}_n$ of Assumption ${\Greekmath 0115} 2$ accommodates both stationary and local unit root predictors.
We first show the consistency of QR for the transformed model in (ref).
In a mixed root model, Lemma (ref) shows that the QR estimators for the $I(0)$ predictors and $I(1)$ predictors have different convergence rates. As we discussed in the previous section, the faster convergence rate for the $I(1)$ predictors comes from the stronger signal of $E\left[w_{t}^{(1)}w_{t}^{(1) \prime}\right]$. For the later use of these rates, we define $a_{n}^{(0)}:=(p^{{\Greekmath 010B} }/\sqrt{n})$ and $a_{n}^{(1)}:=(p^{{\Greekmath 010B} }/n)$.
Based on the definition of the transformed model, we know that Lemma (ref) can further imply consistency of the QR estimator in the original model (ref). Since all parameters are identical in model (ref) and (ref) except $\tilde{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c}=A_{1}^{\prime }\hat{{\Greekmath 010C}}_{1{\Greekmath 011C} }^{c}+\hat{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c}$, the consistency of QR estimator in (ref) can be verified if we can show that $\hat{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c} $ is an $a_n^{(0)}$-consistent estimator for ${\Greekmath 010C}_{0,2{\Greekmath 011C}}^{c}$. To see the convergence rate of $\hat{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c}$, we have
Thus, the $I$(0) rate dominates, and $\left\Vert \hat{{\Greekmath 010C}}_{2{\Greekmath 011C}}^{c}-{\Greekmath 010C} _{0,2{\Greekmath 011C} }^{c}\right\Vert =O_{p}\left( a_{n}^{(0)}\right) $.
These results are summarized in the following corollary.
Note that ${{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(0)}$ is the $(p_{z}+p_{1}+p_{2}) $-dimensional QR regression coefficient, and ${{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(1)}$ is the $p_{x}$ -dimensional QR regression coefficient, as given in Equation ((ref)). We now investigate the oracle properties of ALQR with mixed roots.
It is worth emphasizing that the ALQR procedure does not require practitioners to conduct any pretest to identify different types of predictors. With the proper choice of ${\Greekmath 0115}_n$ along with consistent initial estimates of the parameters, ALQR can simultaneously estimate relevant regression coefficients and provide the consistent variable selection for the mixed root QR model.
In this section, we derive the asymptotic distributions of the proposed ALQR estimators. In QR literature, the Convexity Lemma (pollard1991asymptotics) typically provides a convenient limit theory, bypassing the stochastic equicontinuity argument, see Section 4 of koenker_2005. Unfortunately, for the increasing dimension, we cannot use the Convexity Lemma which is defined in $\mathbb{R} ^{p}$ with a fixed $p<\infty $. Thus we modify the classical chaining argument of ruppert1980trimmed to our ALQR\ framework. lu2015jackknife recently modified the proof of ruppert1980trimmed to show the asymptotic distribution of Jackknife Model Averaging QR with iid regressors of increasing dimensions. We follow a similar proof strategy but substantially refine the results to prove the distributional limit theory with weakly dependent regressors of increasing dimensions, along with the adaptive lasso penalty functions. For the asymptotic distribution with non-cointegrated local unit root predictors, we assume that the number of predictors ($p_{x}$) is fixed. Note that $p=p_{z}+p_{1}+p_{2}+p_{x}$ is still allowed to increase, even with this additional restriction of $ p_{x}<\infty $.
We now state the additional assumptions we need to prove the asymptotic distributions of the ALQR estimators.
Assumption (ref) is a strong mixing condition with the specific mixing rates to use the Bernstein Inequality of merlevede2009bernstein. Assumption (ref) is the moment condition we adopt from lu2015jackknife. The moment condition is stronger than the typical assumptions when the number of the regressors is fixed. Assumption (ref) is another set of combined rate conditions required to prove the asymptotic theories of ALQR estimators. Finally, Assumption (ref) restricts the number of non-cointegrated local unit root predictors, which is also assumed in koo2020high who study the mean regressions with increasing dimensions. The distributional QR limit theory with the increasing number of (non-cointegrated) unit-root regressors is an open question, and we leave it for a future research.
We are now ready to provide the asymptotic distributions of ALQR estimators for the coefficients in the true active set, which we denote $\hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(i),ALQR\ast }(1)$ with $i=1$ and $2$ for $I(0)$ and $I(1)$ predictors, respectively. Let $R$ be a generic constant $l\times p_{i\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }}$matrix, where $l$ is fixed and $p_{i}$ is defined conformably below. Notation $\Longrightarrow $ indicates the convergence in distribution.
The result of Theorem (ref)-(i) is in line with Theorem 2 of medeiros2016l1 who studied the asymptotic distributions of adaptive lasso estimators in mean regressions. It is also interesting to see that there is an additional nuisance parameter $A_{1}$ (the cointegrating vector) in the asymptotic distribution of ALQR with the cointegrated predictors.
In this section, we conduct a set of Monte Carlo simulation studies to evaluate the forecasting performance of the proposed ALQR method. As is motivated by the data patterns of the stock returns application in Section (ref), we construct a simulation environment that has 12 predictors. As seen in Table (ref), we allow for mix-root predictors and let 8 of them be persistent.
Based on the monthly observations of the stationary predictors (svar, dfr, infl and ltr), we calibrate the structure of $z_t$ by estimating a VAR($p$) model with a Bayesian information criterion (BIC). An estimated VAR(2) model is obtained as follows:
where $u_{zt}\sim N(0,1)$. To identify the cointegrating relations among persistent predictors, we apply the Johansen test. The estimated cointegrating rank is $3$. Thus, we generate a set of cointegrating predictors from the following model:
where $v_{t}^{c}\sim N(0,1)$. For unit-root predictors, we use $i.i.d.$\ $N(0,1)$ innovations with stationary initializations.
We consider the following scenarios with different numbers of zero coefficients:\newline (1) 6 non-zero coefficients:
(2) 8 non-zero coefficients:
(3) 12 non-zero coefficients (no zero coefficient):
The sample size is set to be $n=1000$. The last $12$ periods are used for out-of-sample prediction evaluation and the remaining periods are used for in-sample estimation. The quantiles of interest are $0.05$, $0.1$, $0.5$, $0.9$, $0.95$. The prediction performance is measured by the final prediction error (FPE) and the out-of-sample $R^2$ defined in (ref) and (ref), respectively. The performance measures are computed by averaging over $1000$ replications.
We briefly discuss the choice of tuning parameter ${\Greekmath 0115}$. For ALQR, we recommend using the Bayesian information criterion (BIC) proposed by wang2007unified or the generalized information criterion (GIC) proposed by fan2013tuning and zheng2015globally. The objective function for choosing ${\Greekmath 0115}$ is defined as
where $\hat{{\Greekmath 010C}}_{{\Greekmath 011C}}({\Greekmath 0115})$ is ALQR given ${\Greekmath 0115}$, $\Gamma _{n}$ is a positive sequence converging to $0$, and $\hat{S}({\Greekmath 0115}):=\Vert\hat{\mathcal{A}}_{n} \Vert_0$ is the number of active predictors selected by $\hat{{\Greekmath 010C}}_{{\Greekmath 011C}}({\Greekmath 0115})$. We set $\Gamma _{n}={\log(n)}/{n}$ for BIC, and $\Gamma_{n}={\log(p)} \cdot \log(\log(n))/{n}$ for GIC and select $\hat{{\Greekmath 0115}}$ that minimizes the information criterion (objective) function. We use the same procedures for LASSO. Different approaches like the $k$-fold cross-validation and the Akaike information criterion (AIC) are available but it is known that they sometimes fail to effectively identify the true model (see, e.g. shao1997asymptotic; wang2007tuning; zhang2010regularization). zheng2013adaptive also discussed that the statistical properties of the $k$-fold CV have not been well understood for high-dimensional regression with heavy-tailed errors, where QR is often applied. On the contrary, wang2007robust show the model selection consistency of BIC for the fixed dimension and wang2009shrinkage do for the increasing dimension with $p<n$. Also, fan2013tuning and zheng2015globally show that GIC can identify the underlying true model consistently with probability approaching $1$.
Overall, the performance of ALQR is satisfactory and confirms the theory developed in the previous section. Tables (ref)--(ref) summarize the simulation results from Scenarios 1--3. The upper and lower panels of each table contain the results of using BIC and GIC, respectively. First, when there are zero coefficients in the model (Scenarios 1 and 2), ALQR performs better than its alternatives in terms of FPE (or $R^2$). As predicted by existing theory in the literature, both LASSO and ALQR show better performance than QR. In the meantime, ALQR performs uniformly, though slightly, better than LASSO. When all predictors have non-zero coefficients (Scenario 3), LASSO performs slightly better than ALQR over all quantiles except ${\Greekmath 011C}=0.5$. Second, the simulation results confirm the model selection consistency of ALQR derived in Theorem (ref). When we take a look at the average number of selected predictors in each table, the results of ALQR are quite close to the number of the true non-zero coefficients (6, 8, and 12 in each scenario). As known in the literature, LASSO mostly overselects them in Scenarios 1--2. Even in Scenario 3 where there are no zero predictors, ALQR selects most of them successfully. These results show the robustness of ALQR in terms of model selection. Finally, both BIC and GIC perform well under different designs and we do not see much difference between these two methods.
To investigate the robustness of the simulation results above, we further conduct simulation experiments with predictors with highly correlated innovations. In section (ref), the oracle properties of the ALQR estimator are developed under the assumption that most predictor innovations are not highly correlated to each other. Although our empirical application satisfies this assumption, we provide additional simulation results in Tables (ref)-(ref) of the appendix to demonstrate the robustness of ALQR in various settings of correlated predictor innovations. Specifically, we generate the $I$(1) predictors in Scenario 1-3 with correlated innovations. We consider correlation ${\Greekmath 011A}=$0.1, 0.5, and 0.9. We find that when the true DGP includes zero coefficients, ALQR is better than other alternative methods in most of the settings. Only when DGP has no zero coefficient ALQR is at risk of underperforming across quantiles, though the performance of the other methods is only slightly better than ALQR. An interesting finding is that ridge regression, commonly recommended when many predictors are highly correlated, has a mixed performance across quantiles. Based on the simulation results, we recommend using ALQR, especially in the case of sparse data. But ridge regression and other methods may be used if sparsity is not found in the data and the quantile of interest is around the center of the distribution.
In sum, the numerical experiments confirm that ALQR can provide satisfactory prediction performance in finite samples. The results are robust over different quantiles and simulation designs. Therefore, we can expect similar results in other quantile prediction applications when the predictors have mixed roots and are composed of $I(0)$, $I(1)$, and cointegrated processes.
In this paper, we show that the adaptive lasso for quantile regression (ALQR) is attractive in forecasting with stationary and nonstationary predictors as well as cointegrated predictors. The framework is general enough to include mixed roots but ALQR does not require any researchers' knowledge on the specific structure of each predictor nor the order of integration. In this general framework, we show that ALQR preserves the oracle properties. These advantages offer substantial convenience and robustness to empirical researchers working with quantile prediction using time series data.
We have focused on the case where the number of covariates, $p$, is allowed to grow as sample size increases, although $p$ is smaller than $n$. This framework justifies a wide range of practical applications in economics, such as the stock return quantile prediction in this paper. It would be an interesting future research to allow $p$ to be even larger than $n$, which has not been studied in a general time series framework with mixed roots.