EconBase
← Back to paper

Predictive Quantile Regression with Mixed Roots and Increasing Dimensions: The ALQR Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

98,799 characters · 11 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Predictive Quantile Regression with Mixed Roots and Increasing Dimensions: The ALQR Approach

abstractIn this paper we propose the adaptive lasso for predictive quantile regression (ALQR). Reflecting empirical findings, we allow predictors to have various degrees of persistence and exhibit different signal strengths. The number of predictors is allowed to grow with the sample size. We study regularity conditions under which stationary, local unit root, and cointegrated predictors are present simultaneously. We next show the convergence rates, model selection consistency, and asymptotic distributions of ALQR. We apply the proposed method to the out-of-sample quantile prediction problem of stock returns and find that it outperforms the existing alternatives. We also provide numerical evidence from additional Monte Carlo experiments, supporting the theoretical results. Keywords: adaptive lasso, cointegration, forecasting, oracle property, quantile regression \ JEL classification: C22, C53, C61

Introduction

Predictive quantile regression (QR) identifies the impact of predictors on a set of conditional quantiles of a response variable. It provides richer information on the heterogeneous distributional prediction. For example, the conditional quantile prediction of stock returns receives much attention in finance since the tail quantile information has a crucial role in measuring risk. Many economic state variables are employed to predict stock returns and the number of candidate predictors is often large. When a large number of predictors are available, researchers encounter the inevitable model selection issue. A good model selection can improve forecasting performance but the opposite can also occur. Considering the importance of model selection in practice, we need constructive guidance for empirical applications.

In this paper we propose the adaptive lasso for predictive quantile regression (ALQR). Although there exists a large volume of literature on predictive mean regression of equity returns (see, e.g.\ campbell1987stock, fama1988dividend, hodrick1992dividend, cenesizoglu2012return, andersen2020pricing among others), predictive QR is relatively understudied. cenesizoglu2008distribution is an early paper on predictive QR, and maynard2011inference, lee2016predictive, fan2019predictive, gungor2019exact, and cai2022new recently develop inference methods in predictive QR with nonstationarity and heteroskedasticity. The proposed method is different from these approaches. We consider predictive QR with an increasing number of mixed root predictors and address two important problems raised in the stock returns data. First, the prediction power of each predictor can vary over different quantiles. By adapting the lasso, we allow the model selection based on real data not by the researcher's discretion. Second, the predictors widely used in predicting equity returns are composed of stationary, local unit root, and cointegrated processes. For example, we plot the time-series of two predictors, dividend price ratio (dp) and default yield spread (dfr), in Figure (ref). Even a simple eyeballing test easily confirms their different levels of persistence. (We conduct more informative estimation procedures in Section (ref)). We provide a unified adaptive lasso framework that allows those mixed root predictors. We show that the estimator converges to the true parameter value at different convergence rates and that the faster convergence rates make the adaptive lasso more efficient in the model selection.

figure[figure omitted — 601 chars of source]

Naturally, some technical challenges arise. Since the seminal paper of tibshirani1996regression, the lasso has been intensively studied in various fields of statistical analysis. However, most studies have been focusing on the i.i.d.\ sample and it has not been a long time since more studies have been conducted with dependent data (see the references in the related literature section below). Furthermore, we allow nonstationary predictors, which impose an additional difficulty in the formal analysis of the proposed adaptive lasso. We tackle these issues by considering a simple model that only contains unit-root predictors first. Once establishing the desired properties of ALQR in this model, we generalize the model so that it includes all stationary, local unit root, and cointegrated predictors.

The contributions of this paper are two-fold. First, to the best of our knowledge, this is the first paper to study predictive QR with an increasing number of mixed root predictors. We propose the adaptive lasso for predictive quantile regression (ALQR) and derive the convergence rates, model selection consistency and asymptotic distributions of ALQR under some regularity conditions. As a by-product, we also prove that the standard QR estimator is consistent under both mixed roots (including the local unit roots) and the increasing dimension of the predictors, which is new in the literature. Second, we conduct an empirical analysis of the stock returns data and find that ALQR can improve prediction performance over existing alternatives across different quantiles. We apply the ALQR method along with the existing alternatives to the data set. The results confirm that ALQR shows better prediction performance across different quantiles, particularly at higher quantiles. As illustrated in Section 3, ALQR can be readily applicable to other applications.

The rest of the paper is organized as follows. This section finishes with a review of the relevant literature. In section 2, we introduce predictive QR models formally and define the ALQR estimator. In section 3, we investigate the performance of ALQR using the out-of-sample quantile prediction problem of stock returns. We study the theoretical properties of ALQR in sections 4 and 5. In section 7, we conduct some Monte Carlo simulation experiments. Section 7 concludes. All technical proofs are relegated to the appendix.

\paragraph{Related Literature} The lasso has been extensively studied for cross-sectional data. Recently, there has been development in lasso procedures with dependent data. basu2015regularized exploit the spectral properties of the stationary time series design and investigate the regularity condition of the lasso that leads to the non-asymptotic bounds and the consistency results. kock2015oracle investigate the oracle property of the lasso in a stationary vector autoregression model. adamek2020lasso provide an inference procedure based on the debiased/desparsified lasso under the near-epoch dependence assumption. chernozhukov2021lasso propose a penalty selection algorithm with weakly dependent data and the post-selection inference procedure. wu2016performance and wong2020lasso analyze the lasso with non-sub-Gaussian processes that allow a heavy-tail distribution. medeiros2016l1 show the asymptotic properties of the adaptive lasso for stationary high-dimensional time series models. For the cointegrated models, kock2016consistent shows the oracle property of the adaptive lasso in the autoregression model. Interestingly, he finds that the unit root test can be incorporated into the model selection procedure by adopting the Dickey-Fuller form of autoregression. liao2015automated propose a shrinkage estimator with multiple penalty terms to select the rank of cointegration and estimate the parameters simultaneously in a vector error correction model. liang2019determination, zhang2019identifying, and onatski2018alternative investigate the same model in the high-dimensional setting.

koo2020high recently use lasso to improve the prediction of stock returns. It is shown that lasso significantly reduces forecasting mean squared errors even with a mixture of stationary, unit-root, and cointegrated variables. However, the conventional lasso method may not have model selection consistency and the oracle property as shown by meinshausen2004consistent and fan2001variable. The adaptive lasso proposed by zou2006adaptive improves the performance of the lasso. Instead of imposing the same penalty weight on all candidate parameters, the adaptive lasso penalizes each parameter proportionally to the inverse of its initial estimate. With a proper choice of the tuning parameter ${\Greekmath 0115} _{n}$, the adaptive penalty weights for the irrelevant variables approach infinity, whereas those for the relevant variables converge to constants. lee2021lasso apply the adaptive lasso to a predictive mean regression framework. Similar to koo2020high, predictors are allowed to have different degrees of persistence and cointegration. lee2021lasso find that the adaptive lasso and a newly proposed twin adaptive lasso outperform the alternative methods in terms of predictor selection consistency and out-of-sample mean squared errors.

Some effort has been also made to investigate model selection and model estimation in QR under the i.i.d.\ samples. The $\ell_1$-penalized method in the QR framework has been studied for high-dimensional data analysis (see, e.g.\ portnoy1984asymptotic, portnoy1985asymptotic, knight2000asymptotics, koenker_2005, li20081, lee2018oracle and belloni2011_L1). To overcome the problem of inconsistent model selection (fan2014adaptive; wang2012quantile), recent studies have further considered the adaptive lasso in QR. wu2009variable discuss how to conduct model selection for QR models using SCAD and the adaptive lasso method. zheng2013adaptive establish the oracle property for an adaptive lasso QR model with heterogeneous error sequences. zheng2015globally study a globally adaptive lasso method for ultra high-dimensional QR models.

\paragraph{Notation} Let $\Vert \cdot \Vert$ and $\Vert \cdot \Vert_0$ denote $\ell_2$-norm and $\ell_0$-norm, respectively. For a matrix $S$, $\Vert S \Vert$ represents the spectral norm. Let $F_a(\cdot)$ and $f_a(\cdot)$ denote a cumulative distribution function (CDF) and a probability density function (pdf) of a generic random variable $a$. Let $S>0$ denote a generic positive definite matrix $S$. Let ${\Greekmath 0115}_{\min }\left( S\right) $ and ${\Greekmath 0115} _{\max }\left( S\right) $ denote the smallest and largest eigenvalues of $S$. We use $O_p(1)$ and $o_p(1)$ when a sequence is bounded in probability and converges to zero in probability, respectively. The $O(1)$ and $o(1)$ denote the non-stochastic counterparts.

Model and the ALQR Estimator

In this section, we introduce the predictive quantile regression model and the adaptive lasso for quantile regression (ALQR). We will develop the theory for the ALQR in two steps in Section (ref). For a better exposition of the theory, we also propose the model and its estimator in two separate cases: (i) unit-root predictors; and (ii) mixed-root predictors.

QR Model with Unit-Root Predictors

Consider a predictive QR model with unit-root predictors:

align[align omitted — 331 chars of source]

where $Q_{y_{t}}\left( {\Greekmath 011C} |\mathcal{F}_{t-1}\right) $ is the conditional ${\Greekmath 011C}$-quantile of $y_{t}$, $\mathcal{F}_{t}=\left\{ z_{j}\right\} _{j=-\infty }^{t}$ with $z_{j}=(y_{j},x_{j}^{\prime })^{\prime }$ is a natural filtration, $x_{t}$ is a $p_{n}$-dimensional vector of unit-root predictors with a stationary $O_{p}(1)$ initialization of $v_{0}=\sum_{j=0}^{\infty }D_{vj}{\Greekmath 010F} _{-j}$ following the innovation structure below, and ${\Greekmath 010C}_{0{\Greekmath 011C}}$ is the corresponding true parameter vector.

The innovation of unit root predictors, $v_{t}$, follows a linear process, which is commonly assumed in the predictive regression literature (see phillips2013predictive, phillips2016robust, cai2022new, lee2021lasso for a few recent papers):

equation*[equation* omitted — 273 chars of source]
equation*[equation* omitted — 172 chars of source]
equation*[equation* omitted — 372 chars of source]
equation*[equation* omitted — 114 chars of source]

where $mds$ denotes martingale difference sequences with respect to the natural filtration.

We allow the dimension of predictors to increase, i.e. $p_n\rightarrow \infty$ as $n \rightarrow \infty$. To make notation simple, we omit subscript $n$ and use $p$ unless it may cause any confusion. We define $u_{t{\Greekmath 011C} }:=y_{t}-Q_{y_{t}}\left( {\Greekmath 011C} |\mathcal{F}_{t-1}\right)$, the deviation of $y_t$ from the conditional ${\Greekmath 011C}$-quantile. If the CDF of $u_{t{\Greekmath 011C}}$ is continuous, we have

align*[align* omitted — 276 chars of source]

where $F_{u_{t{\Greekmath 011C} }}^{-1}$ is the inverse CDF of $u_{t{\Greekmath 011C} } $. It holds by construction that $\Pr \left( u_{t{\Greekmath 011C} }<0|\mathcal{F}_{t-1}\right) ={\Greekmath 011C}$, which is equivalent to $F_{u_{t{\Greekmath 011C} }}^{-1}\left( {\Greekmath 011C} | \mathcal{F}_{t-1}\right) = 0$. To see the equivalence to the unconditional ${\Greekmath 011C}$-quantile, we note that

align*[align* omitted — 252 chars of source]

where $\mathbf{1}(\cdot )$ denotes an indicator function. The equations above imply that both conditional and unconditional ${\Greekmath 011C}$-quantiles of $u_{t{\Greekmath 011C}}$ are zero for any given ${\Greekmath 011C}$. However, it does not mean that two distribution functions are the same, i.e. $F_{u_{t{\Greekmath 011C}_1 }}^{-1}\left( {\Greekmath 011C}_2 | \mathcal{F}_{t-1}\right) \neq F_{u_{t{\Greekmath 011C}_1 }}^{-1}\left( {\Greekmath 011C}_2 \right)$ for ${\Greekmath 011C}_1 \neq {\Greekmath 011C}_2$ in general.

Let ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} }):= {\Greekmath 011C} - \mathbf{1} \left( u_{t{\Greekmath 011C} }<0\right)$. It is easy to find that ${\Greekmath 0120} _{{\Greekmath 011C} }(u_{t{\Greekmath 011C} })$ is uncorrelated to any predetermined regressor. That is, ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })$ is uncorrelated with any past innovations, $v_{t-j}$, for $j\geq 1$:

align*[align* omitted — 503 chars of source]

by the law of iterated expectations and the fact that $E\left( {\Greekmath 0120}_{{\Greekmath 011C} }(u_{t{\Greekmath 011C} })|\mathcal{\ F}_{t-1}\right) ={\Greekmath 011C} -\Pr \left( u_{t{\Greekmath 011C}}<0|\mathcal{F}_{t-1}\right) =0$. In our settings, however, we allow the QR-induced regression errors ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })$ to be contemporaneously correlated with the innovations of unit root sequences. This is commonly assumed in cointegration and predictive regression literature, inducing a potential second order bias arising from the one-sided correlation, see, e.g., xiao2009quantile.

We now define the ALQR that minimizes the penalized QR objective function as follows:

equation[equation omitted — 438 chars of source]

where ${\Greekmath 011A} _{{\Greekmath 011C} }(u):=u({\Greekmath 011C} -\mathbf{1}(u<0))$. Following zou2006adaptive, we define the tuning parameter of the penalty term as

equation*[equation* omitted — 214 chars of source]

where ${\Greekmath 0115}_n$ is a sequence converging to infinity, ${\Greekmath 0121} _{j}:=|{\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}|^{{\Greekmath 010D} }$ with a positive constant ${\Greekmath 010D}$, and ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$ is a first-step consistent estimator for ${{\Greekmath 010C}}_{0{\Greekmath 011C} ,j}$. The penalty ${\Greekmath 0115}_{n,j}$, which contrasts with the penalty in the standard lasso, is an adaptive weight and it has an inverse relationship with ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$. The idea of using a consistent first-step estimator as an adaptive weight in ${\Greekmath 0115}_{n,j}$ is to avoid over-penalizing important predictors (or under-penalizing irrelevant predictors). In this paper, we use the ordinary QR estimator (koenker1978regression) below as a first-step estimator:

equation[equation omitted — 365 chars of source]

The consistency of the ordinary QR estimator is provided in Section (ref).

QR Model with Mixed Roots

We now introduce a QR model with mixed-root predictors. We assume the predictors of the model have different degrees of persistence: $I(0)$, local unit roots, and cointegration. This model is the most relevant in practice. The model in the previous section can be seen as a special case of it.

Let $z_{t}$, $x_{t}^{c}$, and $x_{t}$ be vectors of stationary,\footnote{In this paper, we use covariance (weak) stationarity unless noted otherwise. The linear process assumption below implies covariance stationarity of $z_{t}$. In addition, the first element of $z_{t}$ is $1$ so that ${\Greekmath 010C} _{0{\Greekmath 011C},1 }^{z}$ plays a role of the intercept. } cointegrated, and local unit-root predictors whose lengths are $p_z$, $p_c$, and $p_x$, respectively. Let $p:= p_z+p_c+p_x$. Given a sample of $\{y_{t},z_{t},x_{t}^{c},x_{t}\}_{t=1}^{n}$, the QR model with mixed roots is defined as follows:

equation[equation omitted — 308 chars of source]

The cointegrated system in $x_{t}^{c}$ has the triangular representation by phillips1991optimal: for $x_{1t}^c$ and $x_{2t}^c$ whose dimensions are $p_1\times 1$ and $p_2\times 1$, respectively, \begingroup\allowdisplaybreaks

align[align omitted — 145 chars of source]

\endgroup where $R_{2}=I_{p_{2}}+c_{2}/n$ with $c_{2}=\mathrm{diag}\left( \check{c} _{1},\ldots,\check{c}_{p_{2}}\right)$, $L$ is the lag operator, and the vector $v_{1t}^{c}$ is the vector of $I$(0) cointegrating residuals. The cointegrated regressors $x_{t}^{c}$ are decomposed into $x_{1t}^c$ and $x_{2t}^c$. The local-to-unity parameter of $x_{2t}^c$ is assumed to be $\check{c}_{j}\in \left( -\infty ,\infty \right)$ for $j=1,\ldots,p_2$. Thus, the cointegrated system in $x_t^c$ includes both stationary and nonstationary local unit root regions. Using (ref), we can easily characterize the cointegration relations in $x_{t}^{c}$ by the matrices $A\equiv(I_{p_{1}},-A_{1})$ and $A_{1}$.

The near integrated process $x_{t}=(x_{t1},...,x_{tp_{x}})$ is defined in a similar way:

equation[equation omitted — 135 chars of source]

where $R_{x}=I_{p_{x}}+c_{x}/n$ with $c_{x}=\mathrm{diag}\left(\tilde{c} _{1},...,\tilde{c}_{p_{x}}\right)$ and $\tilde{c}_{j}\in \left( -\infty ,\infty \right) $. Again, the local-to-unity specification includes the unit root process as a special case when $\tilde{c}_{j}=0$.

Let $X_{t}:=(z_{t}^{\prime },x_{t}^{c \prime},x_{t}^{\prime })^{\prime }$ and ${\Greekmath 010C} ^{\ast }:=({\Greekmath 010C} ^{z^{\prime }},{\Greekmath 010C} ^{c^{\prime }},{\Greekmath 010C} ^{x^{\prime }})^{\prime }$. Define a $p_c$-dimensional vector $v_{t}^{c}:=({v_{1t}^{c}}^{\prime },{v_{2t}^{c}}^{\prime })^{\prime }$ and a $p$-dimensional vector $e_{t}:=\left( z_{t}^{\prime },{v_{t}^{c}}^{\prime },v_{t}^{\prime }\right)^{\prime }$. Abusing notation on ${\Greekmath 010F}_t$, we assume that $e_t$ follows a linear process:

equation*[equation* omitted — 112 chars of source]
equation*[equation* omitted — 567 chars of source]
equation*[equation* omitted — 172 chars of source]
equation*[equation* omitted — 374 chars of source]
equation*[equation* omitted — 128 chars of source]

Note that this assumption implies a covariance stationary $O_{p}(1)$ initialization of $e_{0}=\left( z_{0}^{\prime },{ v_{0}^{c}}^{\prime },v_{0}^{\prime }\right) ^{\prime }=\sum_{j=0}^{\infty }D_{ej}{\Greekmath 010F} _{-j}$.

Similar to Section (ref), the QR-induced regression errors, ${\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })$, are uncorrelated to any predetermined regressor but allowed to be contemporaneously correlated with the innovations of unit root sequences $x_{2t}^{c}$ and $x_{t}$, the stationary predictor $z_t$ as well as the cointegrating residuals $v_{2t}^{c}$.

We define the ALQR for Model (ref) as follows:

equation[equation omitted — 355 chars of source]

where ${\Greekmath 0115} _{n,j}={{\Greekmath 0115} _{n}}/{\Greekmath 0121} _{j}$, ${\Greekmath 0121} _{j}=|{\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}|^{{\Greekmath 010D} }$, ${\Greekmath 010D} $ is a constant, and ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$ is a first-step consistent estimator. We recommend using the ordinary QR estimator for ${\tilde{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}}$. The consistency of QR will be discussed in Section (ref).

Quantile Prediction of Stock Returns

In this section, we consider the quantile prediction problem of stock returns and illustrate the usefulness of the proposed ALQR method. We use an updated version of the data set in welch2008comprehensive, which ranges from January 1952 to December 2019. It is composed of 816 monthly observations of the US financial and macroeconomic variables. Our goal is to predict the quantiles of the excess stock returns ($y_t$) using 12 predictors. The variable names are summarized in Table (ref). See also welch2008comprehensive for more details on these variables.

table[table omitted — 1,220 chars of source]

We first check the mixed-root property of the predictors. In Figures (ref)--(ref), we plot each predictor using monthly observations. The predictors (dp, dy, ep, bm, dfy, ntis, lty, and tbl) collected in Figure (ref) are highly persistent and have quite different patterns from the other four predictors (svar, dfr, ltr and infl) in Figure (ref). The first-order autoregression coefficients reported in Table (ref) also suggest that predictors have heterogeneous degrees of persistence. We also conduct the Johansen test for cointegration. The results show that the cointegrating rank is $3$ in all of the 804-month rolling windows.\footnote{The Johansen cointegration test results are presented in the appendix, Table (ref).} Thus, the data fit into the mixed-root model structure discussed in Section (ref).

figure[figure omitted — 890 chars of source]
figure[figure omitted — 673 chars of source]
commentTo compare the performances of prediction at multiple quantiles, we employ six different strategies: QR, lasso QR, Alasso QR with GIC, Alasso QR with BIC, JMA and QRIC, the same set of methods discussed in the previous section. For each method, we conduct 12 out-of-sample one-step-ahead predictions, using 779-month rolling windows for in-sample estimation. The prediction performance is evaluated by FPE and R$^2$, which are defined in equation ( (ref)) and ((ref)).

We next investigate the performance of ALQR in the quantile prediction problem of stock returns. For comparison, we also apply three alternative estimators in the literature: the QR estimator without selecting predictors (QR), the quantile lasso (LASSO), and the unconditional quantile without using any predictor (QUANT). For the tuning parameter of LASSO and ALQR, we use both the Bayesian information criteria (BIC) and the generalized information criteria (GIC), which are explained in detail in Section (ref). We evaluate the performance of quantile prediction using the final prediction error (FPE) and the out-of-sample $R^2$ following lu2015jackknife. The FPE measures the out-of-sample quantile prediction errors:

eqnarray[eqnarray omitted — 214 chars of source]

where $\hat{y}_s$ is a prediction for ${\Greekmath 011C}$-quantile of $y_s$ and $S$ is the number of out-of-sample predictions. Note that FPE(${\Greekmath 011C}$) averages the quantile loss function. It is different from the standard mean squared prediction error. At each quantile ${\Greekmath 011C}$, a smaller FPE implies a better quantile prediction. The second measure of the performance is the out-of-sample $R^2$:

eqnarray[eqnarray omitted — 225 chars of source]

where $\bar{y}_{s}$ is the unconditional ${\Greekmath 011C}$-quantile of $y$ (QUANT) from the training sample. Note that the out-of-sample $R^2$ measures the performance of each method relative to QUANT. For example, a positive $R^2$ indicates that the prediction error is smaller than that of QUANT, and a larger $R^2$ implies a better prediction. By definition, $R^2$ of QUANT is always zero.

Tables (ref)--(ref) summarize the prediction results with BIC. They are based on the one-step-ahead prediction for the last 12 and 24 periods of the sample, respectively. The results with GIC are similar and left in the appendix. We consider 5 different quantiles, ${\Greekmath 011C}=0.05, 0.1, 0.5, 0.9$, and $0.95$. For each quantile, the best performance results, i.e., the smallest FPE and the highest $R^2$, are marked in bold. In addition to performance measures, we also report the average number of selected predictors and the corresponding tuning parameter value ${\Greekmath 0115}$ for LASSO and ALQR.

Overall, ALQR shows quite satisfactory results. First, ALQR performs the best in prediction across all quantiles, except for two designs: ${\Greekmath 011C}=0.95$ at the 12-period experiment and ${\Greekmath 011C}=0.05$ at the 24-period experiment. In the former design, LASSO performs slightly better than ALQR but the difference is negligible. In the latter design, QR performs better than the alternatives. Second, the relative performance of ALQR to QUANT improves substantially for higher quantiles. Looking at $R^2$, we find that ALQR predicts 44% (${\Greekmath 011C}=0.9$) and 32% (${\Greekmath 011C}=0.95$) better than QUANT at the 12-period design and 28% (${\Greekmath 011C}=0.9$) and 31% (${\Greekmath 011C}=0.95$) at the 24-period design, respectively. QR does not show particularly better prediction performance than QUANT, which is in line with the well-known result in the mean stock return prediction literature (see, e.g., welch2008comprehensive). Third, ALQR selects fewer predictors than LASSO on average. This result is consistent with the report in literature that the standard lasso overselects relevant predictors in practice (See, e.g., wasserman2009high). In Section (ref), we prove the oracle properties of ALQR including model selection consistency. In Section (ref), we provide further numerical evidence that the number of selected variables by ALQR is closer to the true sparsity via Monte Carlo simulations.

commentThe {\color{black}ALQR} method shows the best forecasting performance among all methods: it produces the smallest one-step-ahead prediction errors and thus the largest R$^2$ at every quantile of interest (except for ${\Greekmath 011C}=0.05$ with a $24$-period average). (2) Compared to the unconditional quantiles, {\color{black}ALQR} can greatly reduce forecast errors by selecting a set of informative predictors. The difference of forecasting errors of these two methods becomes more evident as we move from lower quantiles to the upper quantiles. (3) ALQR works better than the unconditional quantiles in predicting the average market performance. In Figure (ref), we demonstrate that a set of stationary and nonstationary predictors selected by ALQR can be used to improve the predicting performance at the median. For example, we find that $tbl$, $infl$, and $svar$ are negatively associated with the market returns at ${\Greekmath 011C}=0.5$. On the other hand, some other variables such as $dfy$ are positively related to the median of market returns. (4) ALQR selects fewer predictors than lasso. In Tables (ref), (ref), (ref) and (ref), we show that the optimal number of predictors selected by ALQR is always smaller than the number by lasso. This result is consistent with existing literature on the conservative model selection of lasso. In Section (ref) and Section (ref), we will prove ALQR's consistent variable selection and its oracle properties, and provide further evidences that the number of selected variables by ALQR is closer to the true sparsity in a simulation environment that mimics our empirical scenario.
table[table omitted — 1,766 chars of source]
table[table omitted — 1,773 chars of source]

Finally, we make some remarks from the empirical perspective. In Figure (ref), we plot the coefficient estimates of ALQR at the 12-period experiment. To make the graph readable, we divide the predictors into two groups (persistent and stationary) and restrict the quantiles into ${\Greekmath 011C}=0.1, 0.5$, and $0.9$. First, we observe that predictors have heterogeneous effects across quantiles. This provides useful information on predicting the distribution of stock returns. For high stock returns (${\Greekmath 011C}=0.9$), default yield spread (dfy), dividend price ratio (dp), dividend yield ratio (dy), stock variance (svar), and inflation (infl) show larger effects. However, long term yield (lty), net equity expansion (ntis), treasury bill rates (tbl), and stock variance (svar) show large coefficients for low stock returns (${\Greekmath 011C}=0.1$). Second, the magnitude of the coefficients at the median is relatively smaller than that at both tails. This result coincides with the previous findings in the literature (see, e.g., welch2008comprehensive and fan2019predictive) that it is more difficult to predict the center part of the stock return distribution. The positive but small $R^{2}$'s of ALQR at ${\Greekmath 011C}=0.5$ in Tables (ref)--(ref) also reflect such an aspect. Third, the coefficient of stock variance (svar) swings from a large positive value at ${\Greekmath 011C}=0.9$ to a large negative number at ${\Greekmath 011C}=0.1$. Since the return distribution would spread out when the market is volatile, this result is consistent with the stylized fact in the financial market.

figure[figure omitted — 185 chars of source]
commentOur empirical analysis using QR-based methods shows that QR can be an informative technique for predicting the financial market returns. It allows to study a wider range of predictability beyond the mean of the return distribution. For example, in Figure (ref), we provide a clear picture of how a set of state variables can effectively contribute to the prediction of the return distribution. For ${\Greekmath 011C}=0.5$, (median of the market returns), the majority of the predictors have (near) zero coefficients. Although a few number of predictors such as $svar$, $infl$, and $dfy$ are selected, the absolute values of their estimated coefficients are considerably smaller than the effective predictors at both tails. This result is consistent with the finding in literature such as welch2008comprehensive and fan2019predictive, suggesting the difficulty in predicting the center part of the return distribution. However, it does not necessarily imply that the predictors are uninformative in predicting the other parts of the return distribution. For investors or policy makers, it would be also useful to find informative predictors in the bad time (lower tail) or the good time (upper tail) of the market. At ${\Greekmath 011C}=0.1$ and $ {\Greekmath 011C}=0.9$, our results indicate a different story of the market with large negative or positive returns. For ${\Greekmath 011C}=0.9$, $dp$, $dy$, and $dfy$ are effective nonstationary predictor. Among stationary predictors, $svar$ and $infl$ are also informative predictors for predicting the upper tail. In particular, the effect of stock variance ($svar$) now becomes positive and much larger than the coefficient at the median. For ${\Greekmath 011C}=0.1$, a different set of predictors is effective in predicting the lower tails of return distribution. This may due to the asymmetry of predictors' ability to predict market returns. The economic story behind the results could be addressed. For example, the negative coefficients of $svar$ at ${\Greekmath 011C}=0.1$ implies that increase in the market volatility can shift the lower tails further downward, thus leading to a higher probability of even larger negative returns. Meanwhile, for the upper tail at ${\Greekmath 011C}=0.9$, positive coefficients of $svar$ indicate that a larger stock volatility can also shift the upper tails further upward, thus increasing the prospects of subsequent larger positive returns.
comment\begin{figure}[thbp] \caption{The estimated AR(1) coefficients of predictors: (1)} \end{figure} \begin{figure}[thbp] \caption{The estimated AR(1) coefficients of predictors: (2)} \end{figure}

Oracle Properties of ALQR Estimators

This section investigates the asymptotic properties of ALQR. We show that the oracle properties (fan2001variable) of the adaptive lasso remain valid in the QR framework when predictors have both increasing dimensions and various degrees of persistence. This result provides theoretical support for applying the ALQR method to stock returns prediction in Section 3. In Section 4.1, we discuss the case that all predictors are $I(1)$. Section 4.2 then generalizes the results to the case of mixed root predictors.

Unit-Root Predictors

Consider the QR model in Section (ref) with unit-root predictors only. The following assumptions provide regularity conditions.

assumption[Assumption $f$] (i) The distribution function of $u_{t{\Greekmath 011C} }$, $F(\cdot )$, has a continuous density $f(\cdot )$ with $f(a)>0$ on $\{a:0<F(a)<1\}$. (ii) The derivative of conditional distribution function $F_{t-1}(a)=Pr[u_{t {\Greekmath 011C} }<a|\mathcal{F}_{t-1}]$, which we denote $f_{t-1}(\cdot )$, is continuous and uniformly bounded above by a finite constant $c_{f}$. (iii) For any sequence ${\Greekmath 0110} _{n}\rightarrow F^{-1}({\Greekmath 011C} )$, $ f_{t-1}({\Greekmath 0110} _{n})$ is uniformly integrable, and $E[f_{t-1}^{1+{\Greekmath 0111} }(F^{-1}({\Greekmath 011C} ))]<\infty $ for some ${\Greekmath 0111} >0$.

For the bound and rate conditions, we define $A_{(t,p)}:=E\left[f_{t-1}(0)x_{t-1}x_{t-1}^{\prime }\right]$ and $B_{(t,p)}:=E\left[ {\Greekmath 0120} _{{\Greekmath 011C}}(u_{t{\Greekmath 011C} })^{2}x_{t-1}x_{t-1}^{\prime }\right]$.

assumption[Assumption $L1$] For each $t$ and $p$, there exist some constants $ \underline{c}_{A(t,p)}$ and $\underline{c}_{B(t,p)}$ such that $0<\underline{c}_{A(t,p)}\leq {\Greekmath 0115} _{\min }(\frac{A_{(t,p)}}{t})$ and $0<\underline{c}_{B(t,p)}\leq {\Greekmath 0115} _{\min }(\frac{B_{(t,p)}}{t})$.

Note that, for each $t$ and $p$, ${\Greekmath 0115}_{\max}(A_{(t,p)}/p)$ and ${\Greekmath 0115}_{\max}(B_{(t,p)}/p)$ are bounded above since $|f_{t-1}(\cdot)|$ and $|{\Greekmath 0120}_{{\Greekmath 011C}}(\cdot)|$ are uniformly bounded and the variance of $x_{t-1}$ is finite by the design of $v_{t}$. Let $\overline{c}_{A(t,p)}<\infty$ and $\overline{c}_{B(t,p)}<\infty$ be these finite bounds. We next define the following uniform bounds over $t$ for each $n$: $\underline{c}_{A(n,p)}:=\min_{1\leq t\leq n}\underline{c}_{A(t,p)}$, $\overline{c}_{A(n,p)}:=\max_{1\leq t\leq n}\overline{c}_{A(t,p)}$, and $\overline{c}_{B(n,p)}:=\max_{1\leq t\leq n}\overline{c}_{B(t,p)}$.

assumption[Assumption $U1$] For some ${\Greekmath 010B}>0$, (i) $p=n^{{\Greekmath 0110} }$ with $0<{\Greekmath 010B} {\Greekmath 0110} <1$; (ii) $\frac{\overline{c}_{A(n,p)}^{1/2}n^{1/2}}{p^{{\Greekmath 010B} }}\vee \frac{\overline{c} _{B(n,p)}^{1/2}}{p^{{\Greekmath 010B} }}=o(\underline{c}_{A(n,p)})$; (iii)$\frac{ p^{3/2}}{n^{2+{\Greekmath 010B} {\Greekmath 0110}}\overline{c}_{A(n,p)}^{1/2}}=\frac{n^{(3/2){\Greekmath 0110} }}{n^{2+{\Greekmath 010B} {\Greekmath 0110} }\overline{c}_{A(n,p)}^{1/2}}=o(1).$
assumption[Assumption $\protect{\Greekmath 0115} 1$] ${\Greekmath 0115} _{n}$ satisfies that (i) $\frac{{\Greekmath 0115} _{n}n^{\frac{1}{2}{\Greekmath 0110} }}{n^{1+{\Greekmath 010B} {\Greekmath 0110} }\underline{c}_{A(n,p)}}\rightarrow 0$; (ii) $\frac{{\Greekmath 0115} _{n}n^{(1-{\Greekmath 010B} {\Greekmath 0110} ){\Greekmath 010D} }}{n^{2+{\Greekmath 010B} {\Greekmath 0110} }\overline{c}_{A(n,p)}^{1/2}}\rightarrow \infty $, with ${\Greekmath 010D} >0.$
assumptionLet $q_{n}=\left\Vert \mathcal{A}_{n}\right\Vert_0 $ be the size of the true active set $\mathcal{A}_{n}$. There exist positive constants $c_{q}$, $ c_{{\Greekmath 010C} }$ and $c_0$ such that (i) $q_{n}=O(n^{c_{q}})$ with $c_{q}<1/3;$(ii) $2c_{q}<c_{{\Greekmath 010C} }<1$ and (iii) \begin{equation*} n^{(1-c_{{\Greekmath 010C} })/2}\cdot \min_{1\leq j\leq q_{n}}\left\vert {\Greekmath 010C} _{0j}\right\vert \geq c_0. \end{equation*}

We make some remarks on these assumptions. Assumption $f$ is a standard restriction in the QR literature on the conditional density of regression residuals. Assumptions $L1$ and $U1$ modify the conventional conditions from stationary QR literature to allow for serial dependence in predictors and regression errors. In particular, we modify Assumption $A.2$ of lu2015jackknife to accommodate nonstationary $x_t$. Based on the model settings in Section (ref), we find that

equation*[equation* omitted — 88 chars of source]

for $t=\left\lfloor rn\right\rfloor $ with $r\in \left( 0,1\right) $, as $ n\rightarrow \infty$. Therefore, we impose a similar set of restricted eigenvalue conditions for $\frac{E\left[ x_{t-1}x_{t-1}^{\prime }\right] }{t}$. This is a natural extension of the existing conditions in the i.i.d.\ or stationary QR setup to a nonstationary time series model with an increasing dimension. We allow the upper bounds $\overline{c}_{A(t,p)}$, $\overline{c}_{B(t,p)}$ and the lower bounds $\underline{c}_{A(t,p)}$, $\underline{c}_{B(t,p)}$ to depend on the dimension of predictors, hence the upper bounds can diverge to infinity and the lower bounds can converge to zero when $p\rightarrow\infty$ as $n\rightarrow \infty$. However, Assumption $U1$ (ii) imposes further restrictions on the upper bounds: The rate of convergence in $p$ is slow enough so that $(n/p^{\Greekmath 010B})$-consistency of the ALQR estimator can be achieved. Specifically, Assumption $L1$ prevents high correlation between predictor innovations. Note that ${\Greekmath 0115}_{\min}(E[x_{t-1}x_{t-1}]/t)=0$ if there exits perfect multicollinearity between predictor innovations. Also, Assumption $U1$ restricts the strength of the contemporaneous correlation in $v_t$ by limiting nonzero off-diagonal terms in $\Sigma_{vv}$.\footnote{ In the i.i.d.\ mean regression model with highly correlated regressors, it is well known that the lasso performs worse than the ridge estimator. To accommodate the lasso with highly correlated regressors, researchers have also developed variants of the lasso such as the elastic net zou2005regularization and the group lasso yuan2006model. As shown in Figure (ref) in the appendix, however, most predictor innovations in our empirical application are not highly correlated to each other and satisfy this assumption. Furthermore, we investigate the performance of the proposed ALQR estimator when predictor innovations are highly correlated in the simulation experiments. } In Assumption $U1$ (i), the condition $0<{\Greekmath 010B} {\Greekmath 0110} <1$ imposes $a_{n}=\frac{p^{{\Greekmath 010B} }}{n}=\frac{n^{{\Greekmath 010B} {\Greekmath 0110} }}{n}\rightarrow 0$, indicating the restriction on the growth rate of the number of parameters. In the appendix, we discuss the technicality of these restrictions on $p$ and ${\Greekmath 0115} $ in detail. Assumption ${\Greekmath 0115} 1$ collects a set of rate conditions on ${\Greekmath 0115}_n$ to show the asymptotic properties of the ALQR estimators. Assumption ${\Greekmath 0115} 1$ (i) is comparable to the condition of ${\Greekmath 0115}_n/\sqrt{n}\rightarrow 0$ in a stationary time series model with fixed dimensions. This condition allows $(n/p^{\Greekmath 010B})$-rate under the lasso variable selection. Assumption ${\Greekmath 0115} 1$ (ii) presents a necessary condition for model selection consistency of ALQR, which depends on the consistency of the initial estimator for ${\Greekmath 010C}_{0{\Greekmath 011C}}$ (as shown in the appendix, proof of Theorem (ref)). Note that this condition is developed given that the ordinary QR estimator defined in ((ref)) is used as an initial estimator for ${\Greekmath 010C} _{0{\Greekmath 011C} }$. It can be generalized further since we do not require a certain rate of consistency for the initial estimator. If another consistent estimator is used, condition (ii) may be adjusted to the corresponding rate of consistency. Assumption (ref) is for the minimum signal strength of the coefficients in the true active set, which we adopted from \citet*{sherwood2016partially}.

Next we show that the ordinary QR estimator is consistent and can be a qualified initial estimator for ALQR.

lemma(Consistency of QR Estimator with unit-root predictors) Under Assumptions (ref)--(ref), the QR estimator defined in ((ref)) satisfies \begin{equation*} \left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{QR}-{\Greekmath 010C} _{0{\Greekmath 011C} }\right\Vert =O_{p}\left( \frac{p^{{\Greekmath 010B} }}{n}\right) . \end{equation*}

It is well-known that the rate of convergence is $\frac{1}{n^{1/2}}$ with fixed $p$ and $\frac{p^{1/2}}{n^{1/2}}$ with increasing $p$ when predictors are $I$(0). This indicates that the information loss (or the degrees of freedom) to estimate the increasing number of parameters is $p^{1/2}$. It is also known that when the predictors are $I$(1), the rate of convergence becomes $\frac{1}{n}$ with fixed $p$. The super-consistency is caused by the stronger signal of $X^{\prime }X$, where $X$ is a $n\times p$ matrix of predictors. To allow for increasing $p$, Lemma (ref) extends the existing result for nonstationary QR and finds that the rate becomes $\frac{p^{{\Greekmath 010B} } }{n}=\frac{p^{1/2}\cdot p^{{\Greekmath 010B}-1/2 }}{n}$, where the additional rate loss $p^{{\Greekmath 010B} -1/2}$ comes from the increasing singularity of $E\left[ x_{t}x_{t}^{\prime }\right] $ as summarized in Assumptions $L1$ and $U1$. In contrast, we have $0<\left[ {\Greekmath 0115} _{\min }\left( E\left[ x_{t}x_{t}^{\prime }\right] \right) \right] <\left[ {\Greekmath 0115} _{\max }\left( E\left[ x_{t}x_{t}^{\prime }\right] \right) \right] <\infty $ for $I$(0) predictors with $p/n\rightarrow 0$.

Taking the QR estimator of Lemma (ref) as an initial estimator for ${\Greekmath 010C} _{0{\Greekmath 011C} }$, we show the consistency of ALQR and model selection consistency under unit roots.

theorem{(Consistency of ALQR under Unit Roots)} Under Assumptions (ref)--(ref), the ALQR estimator defined in ((ref)) satisfies \begin{equation*} \left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{ALQR}-{\Greekmath 010C} _{0{\Greekmath 011C} }\right\Vert =O_{p}\left( \frac{p^{{\Greekmath 010B} }}{n}\right) . \end{equation*}

Theorem (ref) confirms that both the ALQR estimator and the QR estimator of Lemma (ref) are $(n/p^{\Greekmath 010B})$-consistent. With a proper choice of the penalty parameter ${\Greekmath 0115}_n$, the rate of consistency for the ALQR estimate is not affected. As we show in the proof of Theorem (ref), this estimation rate is possible under Assumption ${\Greekmath 0115} 1$.

theorem{(Model Selection Consistency under Unit Roots)} Let $\hat{\mathcal{A}}_{n}:=\{j:\hat{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}^{ALQR}\neq 0\}$ and $\mathcal{A}_{0}:=\{j:{\Greekmath 010C} _{0{\Greekmath 011C} ,j}\neq0\}$. Under Assumptions (ref)--(ref), it holds that \begin{equation*} \Pr (j\in \hat{\mathcal{A}}_{n})\longrightarrow 0,\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for $j\notin \mathcal{A}_{0}$}, \end{equation*} as $n\rightarrow \infty $.

Theorem (ref) shows that the ALQR procedure is an oracle procedure. When the predictors with diverging dimensions are all unit root, the adaptive lasso is at least as good as other “oracle" procedures in terms of variable selection. This is a major difference from the standard lasso, which does not enjoy the oracle properties. As in zou2006adaptive, with a proper choice of ${\Greekmath 0115}_n$, the adaptive lasso procedure can simultaneously estimate relevant coefficients and remove unimportant ones with probability approaching unity. In this study, we further develop the result of zou2006adaptive and generalize the adaptive lasso procedure to a nonstationary QR model. Note that the consistency in model selection can be achieved only under Assumption ${\Greekmath 0115}1$.

Mixed-Root Predictors

In this section, we consider a general QR model discussed in Section (ref):

align[align omitted — 674 chars of source]

where $X_{t}=(z_{t}^{\prime },{\ x_{1t}^{c}}^{\prime },{x_{2t}^{c}}^{\prime },x_{t}^{\prime })^{\prime }$ and ${\Greekmath 010C} _{0{\Greekmath 011C} }=({\Greekmath 010C} _{0{\Greekmath 011C} }^{z^{\prime }},{\Greekmath 010C} _{0,1{\Greekmath 011C} }^{c^{\prime }},{\Greekmath 010C} _{0,2{\Greekmath 011C} }^{c^{\prime }},{\Greekmath 010C} _{0{\Greekmath 011C} }^{x^{\prime }})^{\prime }$ are $p$-dimensional vectors. This model allows predictors to have heterogeneous degrees of persistence as well as increasing dimensions. The mixed-root predictors in $X_t$ include stationary ($z_{t}$), local unit root ($x_{t}$), and cointegrated ($x_{1t}^c$, $x_{2t}^c$) processes.

To capture the cointegration relation between $x_{1t}^{c}$ and $x_{2t}^{c}$, we define a transformation matrix $H$ as follows:

equation*[equation* omitted — 222 chars of source]

Using $H$, the original model can be rewritten as:

align[align omitted — 370 chars of source]

where $\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }:=H{\Greekmath 010C} _{0{\Greekmath 011C} }=({\Greekmath 010C} _{0{\Greekmath 011C} }^{z \prime},{\Greekmath 010C} _{0,1{\Greekmath 011C} }^{c \prime}, (A_{1}^{\prime }{\Greekmath 010C} _{0,1{\Greekmath 011C} }^{c}+{\Greekmath 010C} _{0,2{\Greekmath 011C} }^{c})', {\Greekmath 010C} _{0{\Greekmath 011C} }^{x \prime})'$ and $\tilde{x}_{t-1}^{\prime }:=X_{t-1}^{\prime }H^{-1}=(z_{t}^{\prime },{\ v_{1t}^{c}}^{\prime },{x_{2t}^{c}}^{\prime},x_{t}^{\prime })\equiv\left( w_{t}^{(0)\prime },w_{t}^{(1)\prime }\right)$. The $w_{t}^{(0)}=(z_{t}^{\prime },{v_{1t}^{c}}^{\prime })^{\prime }\ $ is a $(p_z+p_1)$-dimensional vector that collects all $I(0)$ processes, and $\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(0)}$ is the corresponding $(p_{z}+p_{1}) $-dimensional QR regression coefficient. The $w_{t}^{(1)}=({x_{2t}^{c}}^{\prime },x_{t}^{\prime })^{\prime }$ is a $(p_2+p_x)$-dimensional vector of all $I(1)$ processes, $\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(1)}$ is the corresponding $(p_{2}+p_{x})$ -dimensional QR regression coefficient. Since the cointegration rank of this system is $p_1$, we let $r:=p_{z}+p_{1}$ be the number of $I(0)$ predictors in $\tilde{x}_{t-1}$. Note that the discussion in Section (ref) can be considered as a special case of the mixed root model with $r=0$. In this section, we generalize the results in Section (ref) and examine a more practical model with $r>0$.

The following assumptions provide additional regularity conditions for mixed roots. Abusing notation slightly, we will use the same notation for the sequences of bounds. Let $M_{n}$ be a diagonal matrix of the form

equation*[equation* omitted — 131 chars of source]
assumption[Assumption$~L2$] There exist $\underline{c}_{A(n,p)}$ and $\underline{c}_{B(n,p)}$ such that \begin{align*} 0<c_{A(n,p)} & \leq {\Greekmath 0115} _{\min }\left( \frac{M_{n}^{\prime}\sum_{t=1}^{n}E\left[ f_{t-1}(0)\tilde{x}_{t-1}\tilde{x}_{t-1}^{\prime }\right] M_{n}}{n^{2}}\right) \end{align*}and \begin{align*} 0<c_{B(n,p)} & \leq {\Greekmath 0115} _{\min }\left( \frac{M_{n}^{\prime}\sum_{t=1}^{n}E\left[ {\Greekmath 0120} _{{\Greekmath 011C} }(u_{t{\Greekmath 011C} })^{2}\tilde{x}_{t-1}\tilde{x}_{t-1}^{\prime }\right] M_{n}}{n^{2}}\right) \end{align*} $\,$ as $n\rightarrow \infty $.
assumption[Assumption$~U2$] There exist $\overline{c}_{A(n,p)}$ and $\overline{c}_{B(n,p)}$ such that \begin{align*} {\Greekmath 0115} _{\max }\left( \frac{M_{n}^{\prime}\sum_{t=1}^{n}E\left[ f_{t-1}(0)\tilde{x}_{t-1}\tilde{x}_{t-1}^{\prime }\right] M_{n}}{n^{2}}\right) \leq \overline{c}_{A(n,p)}<\infty \end{align*} and \begin{align*} {\Greekmath 0115} _{\max }\left(\frac{M_{n}^{\prime }\sum_{t=1}^{n}E\left[ {\Greekmath 0120} _{{\Greekmath 011C} }(u_{t{\Greekmath 011C} })^{2}\tilde{x}_{t-1}\tilde{x}_{t-1}^{\prime }\right] M_{n}}{n^{2}}\right) \leq \overline{c}_{B(n,p)}<\infty \end{align*} $\,$ as $n\rightarrow \infty. $\newline
assumption[Assumption$~\protect{\Greekmath 0115} 2$] When ${\Greekmath 0115} _{n}\rightarrow \infty$ as $n \rightarrow \infty$, (i) ${\Greekmath 0115}_{n}\frac{n^{\frac{1}{2}{\Greekmath 0110} }}{n^{\frac{1}{2}+{\Greekmath 010B} {\Greekmath 0110} }\underline{c}_{A(n,p)}}\rightarrow 0$; (ii) ${\Greekmath 0115}_{n} \frac{n^{\left( 1/2-{\Greekmath 010B} {\Greekmath 0110} \right) {\Greekmath 010D} }}{n^{2+{\Greekmath 010B} {\Greekmath 0110} }\overline{c}_{A(n,p)}^{1/2}}\rightarrow \infty .$

Assumption $L2$, $U2$, and ${\Greekmath 0115}2$ are analogies to Assumption $L1$, $U1$ and ${\Greekmath 0115}1$ in Section (ref).1. They are essential regularity conditions used to show ALQR's consistency in parameter estimation and model selection under mixed roots. Compared with Assumption ${\Greekmath 0115} 1$, Assumption ${\Greekmath 0115} 2$ are more restrictive on the choice of ${\Greekmath 0115}_n$, because ${\Greekmath 0115}_n$ of Assumption ${\Greekmath 0115} 2$ accommodates both stationary and local unit root predictors.

We first show the consistency of QR for the transformed model in (ref).

lemma(Consistency of the transformed QR Estimator under Mixed Roots) Under Assumptions (ref), (ref), (ref) and (ref), we have \begin{align*} \left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(0),QR}-\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(0)}\right\Vert =O_{p}\left( \frac{p^{{\Greekmath 010B} }}{\sqrt{n}}\right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }\left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(1),QR}-\tilde{{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(1)}\right\Vert =O_{p}\left( \frac{p^{{\Greekmath 010B} }}{n}\right). \end{align*}

In a mixed root model, Lemma (ref) shows that the QR estimators for the $I(0)$ predictors and $I(1)$ predictors have different convergence rates. As we discussed in the previous section, the faster convergence rate for the $I(1)$ predictors comes from the stronger signal of $E\left[w_{t}^{(1)}w_{t}^{(1) \prime}\right]$. For the later use of these rates, we define $a_{n}^{(0)}:=(p^{{\Greekmath 010B} }/\sqrt{n})$ and $a_{n}^{(1)}:=(p^{{\Greekmath 010B} }/n)$.

Based on the definition of the transformed model, we know that Lemma (ref) can further imply consistency of the QR estimator in the original model (ref). Since all parameters are identical in model (ref) and (ref) except $\tilde{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c}=A_{1}^{\prime }\hat{{\Greekmath 010C}}_{1{\Greekmath 011C} }^{c}+\hat{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c}$, the consistency of QR estimator in (ref) can be verified if we can show that $\hat{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c} $ is an $a_n^{(0)}$-consistent estimator for ${\Greekmath 010C}_{0,2{\Greekmath 011C}}^{c}$. To see the convergence rate of $\hat{{\Greekmath 010C}}_{2{\Greekmath 011C} }^{c}$, we have

eqnarray*[eqnarray* omitted — 842 chars of source]

Thus, the $I$(0) rate dominates, and $\left\Vert \hat{{\Greekmath 010C}}_{2{\Greekmath 011C}}^{c}-{\Greekmath 010C} _{0,2{\Greekmath 011C} }^{c}\right\Vert =O_{p}\left( a_{n}^{(0)}\right) $.

These results are summarized in the following corollary.

corollary(Consistency of the QR Estimator under Mixed Roots) Under Assumptions (ref), (ref), (ref) and (ref),we have \begin{align*} \left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(0),QR\ast }-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(0)}\right\Vert =O_{p}\left( a_{n}^{(0)} \right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }\left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(1),QR\ast }-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(1)}\right\Vert =O_{p}\left( a_{n}^{(1)}\right). \end{align*}

Note that ${{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(0)}$ is the $(p_{z}+p_{1}+p_{2}) $-dimensional QR regression coefficient, and ${{\Greekmath 010C}}_{0{\Greekmath 011C} }^{(1)}$ is the $p_{x}$ -dimensional QR regression coefficient, as given in Equation ((ref)). We now investigate the oracle properties of ALQR with mixed roots.

theorem{(Consistency of ALQR under Mixed Roots)} Under Assumptions (ref), (ref), (ref)-(ref), we have \begin{align*} \left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(0),ALQR\ast }-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(0)}\right\Vert =O_{p}\left( a_{n}^{(0)} \right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }\left\Vert \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(1),ALQR\ast }-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(1)}\right\Vert =O_{p}\left( a_{n}^{(1)}\right). \end{align*}
theorem{(Model Selection Consistency under Mixed Roots)} Under Assumptions (ref), (ref), (ref)-(ref), it holds that \begin{equation*} \Pr (j\in \hat{\mathcal{A}}_{n})\longrightarrow 0,\quad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{for $j\notin \mathcal{A}_{0}$}, \end{equation*} where $\hat{\mathcal{A}}_{n}=\{j:\hat{{\Greekmath 010C}}_{{\Greekmath 011C} ,j}^{ALQR\ast }\neq 0\}$ and $\mathcal{A}_{0}=\{j:{\Greekmath 010C} _{0{\Greekmath 011C} ,j}\neq 0\}$.

It is worth emphasizing that the ALQR procedure does not require practitioners to conduct any pretest to identify different types of predictors. With the proper choice of ${\Greekmath 0115}_n$ along with consistent initial estimates of the parameters, ALQR can simultaneously estimate relevant regression coefficients and provide the consistent variable selection for the mixed root QR model.

Asymptotic Distributions of ALQR Estimators

In this section, we derive the asymptotic distributions of the proposed ALQR estimators. In QR literature, the Convexity Lemma (pollard1991asymptotics) typically provides a convenient limit theory, bypassing the stochastic equicontinuity argument, see Section 4 of koenker_2005. Unfortunately, for the increasing dimension, we cannot use the Convexity Lemma which is defined in $\mathbb{R} ^{p}$ with a fixed $p<\infty $. Thus we modify the classical chaining argument of ruppert1980trimmed to our ALQR\ framework. lu2015jackknife recently modified the proof of ruppert1980trimmed to show the asymptotic distribution of Jackknife Model Averaging QR with iid regressors of increasing dimensions. We follow a similar proof strategy but substantially refine the results to prove the distributional limit theory with weakly dependent regressors of increasing dimensions, along with the adaptive lasso penalty functions. For the asymptotic distribution with non-cointegrated local unit root predictors, we assume that the number of predictors ($p_{x}$) is fixed. Note that $p=p_{z}+p_{1}+p_{2}+p_{x}$ is still allowed to increase, even with this additional restriction of $ p_{x}<\infty $.

We now state the additional assumptions we need to prove the asymptotic distributions of the ALQR estimators.

assumptionThe $p_{z}\times1$ vector of stationary predictors $ z_{t}=\sum_{j=0}^{\infty }D_{zj}{\Greekmath 010F} _{t-j}$ satisfies the following condition from Corollary 2 of withers1981conditions \begin{equation*} \Vert D_{zj} \Vert=O\left( e^{-vj}\right) \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ with }v>0 \end{equation*} along with the conditions (1), (2), (5) of withers1981conditions, provided in the appendix for brevity, see ((ref))-((ref)) in Section (ref) of the appendix. Then $z_{t}$ is strong mixing with the strong mixing number ${\Greekmath 010B} (j)=O(e^{-v{\Greekmath 0115} j})$ with another constant ${\Greekmath 0115} >0$.
assumption$\sup_{j\geq 1}E(z_{t-1,j}^{8})\leq c_{z}$ for some $c_{z}\leq \infty $ .
assumptionAs $n\rightarrow \infty $, $\underline{c}_{B}^{-1/2}\overline{c} _{A}^{1/2}p^{2}n^{-1/2}\rightarrow 0$, and $n^{-3}p^{12}\underline{c} _{B}^{-4}\rightarrow 0$.
assumptionWe require $p_{x}$ (the number of non-cointegrated local unit root predictors) to be finite: $p_{x}<\infty$.

Assumption (ref) is a strong mixing condition with the specific mixing rates to use the Bernstein Inequality of merlevede2009bernstein. Assumption (ref) is the moment condition we adopt from lu2015jackknife. The moment condition is stronger than the typical assumptions when the number of the regressors is fixed. Assumption (ref) is another set of combined rate conditions required to prove the asymptotic theories of ALQR estimators. Finally, Assumption (ref) restricts the number of non-cointegrated local unit root predictors, which is also assumed in koo2020high who study the mean regressions with increasing dimensions. The distributional QR limit theory with the increasing number of (non-cointegrated) unit-root regressors is an open question, and we leave it for a future research.

We are now ready to provide the asymptotic distributions of ALQR estimators for the coefficients in the true active set, which we denote $\hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(i),ALQR\ast }(1)$ with $i=1$ and $2$ for $I(0)$ and $I(1)$ predictors, respectively. Let $R$ be a generic constant $l\times p_{i\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }}$matrix, where $l$ is fixed and $p_{i}$ is defined conformably below. Notation $\Longrightarrow $ indicates the convergence in distribution.

theorem{(Asymptotic Distributions of ALQR Estimators)} When Assumptions (ref), (ref), (ref)-(ref), (ref)-(ref) hold, (i) For I(0) active predictors (the first $p_{z}+p_{1}$ predictors; hence $p_{i}=p_{z}+p_{1}$), the first $q_{n}$ non-zero elements has the following limit theory:\ \begin{equation*} \sqrt{n}R\left[ \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(0),ALQR\ast }(1)-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(0)}(1)\right] \Longrightarrow N\left( 0,RA_{(t,p)}^{-1}B_{(t,p)}A_{(t,p)}^{-1}R^{\prime }\right) \end{equation*} (ii) For the middle $p_{2}$ predictors (the parameters with respect to the active cointegrated predictors; hence $p_{i}=p_{2}$), the first $ q_{n}$ non-zero elements has the following limit theory: \begin{equation*} \sqrt{n}R\left[ \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(1),ALQR\ast }(1)-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(1)}(1)\right] \Longrightarrow N\left( 0,RA_{1}^{\prime }A_{(t,p)}^{-1}B_{(t,p)}A_{(t,p)}^{-1}A_{1}R^{\prime }\right) \end{equation*} (iii) For the last $p_{x}$ (non-cointegrated) active local unit root predictors, with $p_{x}<\infty $, the first $q_{n}$ non-zero elements has the following limit theory:\ \begin{equation*} n\left[ \hat{{\Greekmath 010C}}_{{\Greekmath 011C} }^{(1),ALQR\ast }(1)-{\Greekmath 010C} _{0{\Greekmath 011C} }^{(1)}(1) \right] \Longrightarrow \left( M_{{\Greekmath 010C} _{{\Greekmath 011C} }}^{x}\right) ^{-1}G_{{\Greekmath 011C} }^{x}, \end{equation*} where \begin{equation*} G_{{\Greekmath 011C} }^{x}=\int J_{x}^{c}(r)dB_{{\Greekmath 0120} _{{\Greekmath 011C} }}\relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{, and }M_{{\Greekmath 010C} _{{\Greekmath 011C} }}^{x}=f_{u}(0)^{-1}\int J_{x}^{c}(r)J_{x}^{c}(r)^{\prime }dr \end{equation*} and $B_{{\Greekmath 0120} _{{\Greekmath 011C} }}:=BM({\Greekmath 011C} (1-{\Greekmath 011C} ))$ is a Brownian motion and $ J_{x}^{c}(r)=\int_{0}^{r}e^{(r-s)C}dB_{x}(s)$ is an Ornstein-Uhlenbeck (OU) process.

The result of Theorem (ref)-(i) is in line with Theorem 2 of medeiros2016l1 who studied the asymptotic distributions of adaptive lasso estimators in mean regressions. It is also interesting to see that there is an additional nuisance parameter $A_{1}$ (the cointegrating vector) in the asymptotic distribution of ALQR with the cointegrated predictors.

Monte Carlo Simulations

In this section, we conduct a set of Monte Carlo simulation studies to evaluate the forecasting performance of the proposed ALQR method. As is motivated by the data patterns of the stock returns application in Section (ref), we construct a simulation environment that has 12 predictors. As seen in Table (ref), we allow for mix-root predictors and let 8 of them be persistent.

Based on the monthly observations of the stationary predictors (svar, dfr, infl and ltr), we calibrate the structure of $z_t$ by estimating a VAR($p$) model with a Bayesian information criterion (BIC). An estimated VAR(2) model is obtained as follows:

eqnarray*[eqnarray* omitted — 367 chars of source]

where $u_{zt}\sim N(0,1)$. To identify the cointegrating relations among persistent predictors, we apply the Johansen test. The estimated cointegrating rank is $3$. Thus, we generate a set of cointegrating predictors from the following model:

equation*[equation* omitted — 320 chars of source]

where $v_{t}^{c}\sim N(0,1)$. For unit-root predictors, we use $i.i.d.$\ $N(0,1)$ innovations with stationary initializations.

We consider the following scenarios with different numbers of zero coefficients:\newline (1) 6 non-zero coefficients:

equation*[equation* omitted — 235 chars of source]

(2) 8 non-zero coefficients:

equation*[equation* omitted — 211 chars of source]

(3) 12 non-zero coefficients (no zero coefficient):

equation*[equation* omitted — 184 chars of source]

The sample size is set to be $n=1000$. The last $12$ periods are used for out-of-sample prediction evaluation and the remaining periods are used for in-sample estimation. The quantiles of interest are $0.05$, $0.1$, $0.5$, $0.9$, $0.95$. The prediction performance is measured by the final prediction error (FPE) and the out-of-sample $R^2$ defined in (ref) and (ref), respectively. The performance measures are computed by averaging over $1000$ replications.

table[table omitted — 2,669 chars of source]
table[table omitted — 2,671 chars of source]
table[table omitted — 2,710 chars of source]

We briefly discuss the choice of tuning parameter ${\Greekmath 0115}$. For ALQR, we recommend using the Bayesian information criterion (BIC) proposed by wang2007unified or the generalized information criterion (GIC) proposed by fan2013tuning and zheng2015globally. The objective function for choosing ${\Greekmath 0115}$ is defined as

equation[equation omitted — 262 chars of source]

where $\hat{{\Greekmath 010C}}_{{\Greekmath 011C}}({\Greekmath 0115})$ is ALQR given ${\Greekmath 0115}$, $\Gamma _{n}$ is a positive sequence converging to $0$, and $\hat{S}({\Greekmath 0115}):=\Vert\hat{\mathcal{A}}_{n} \Vert_0$ is the number of active predictors selected by $\hat{{\Greekmath 010C}}_{{\Greekmath 011C}}({\Greekmath 0115})$. We set $\Gamma _{n}={\log(n)}/{n}$ for BIC, and $\Gamma_{n}={\log(p)} \cdot \log(\log(n))/{n}$ for GIC and select $\hat{{\Greekmath 0115}}$ that minimizes the information criterion (objective) function. We use the same procedures for LASSO. Different approaches like the $k$-fold cross-validation and the Akaike information criterion (AIC) are available but it is known that they sometimes fail to effectively identify the true model (see, e.g. shao1997asymptotic; wang2007tuning; zhang2010regularization). zheng2013adaptive also discussed that the statistical properties of the $k$-fold CV have not been well understood for high-dimensional regression with heavy-tailed errors, where QR is often applied. On the contrary, wang2007robust show the model selection consistency of BIC for the fixed dimension and wang2009shrinkage do for the increasing dimension with $p<n$. Also, fan2013tuning and zheng2015globally show that GIC can identify the underlying true model consistently with probability approaching $1$.

Overall, the performance of ALQR is satisfactory and confirms the theory developed in the previous section. Tables (ref)--(ref) summarize the simulation results from Scenarios 1--3. The upper and lower panels of each table contain the results of using BIC and GIC, respectively. First, when there are zero coefficients in the model (Scenarios 1 and 2), ALQR performs better than its alternatives in terms of FPE (or $R^2$). As predicted by existing theory in the literature, both LASSO and ALQR show better performance than QR. In the meantime, ALQR performs uniformly, though slightly, better than LASSO. When all predictors have non-zero coefficients (Scenario 3), LASSO performs slightly better than ALQR over all quantiles except ${\Greekmath 011C}=0.5$. Second, the simulation results confirm the model selection consistency of ALQR derived in Theorem (ref). When we take a look at the average number of selected predictors in each table, the results of ALQR are quite close to the number of the true non-zero coefficients (6, 8, and 12 in each scenario). As known in the literature, LASSO mostly overselects them in Scenarios 1--2. Even in Scenario 3 where there are no zero predictors, ALQR selects most of them successfully. These results show the robustness of ALQR in terms of model selection. Finally, both BIC and GIC perform well under different designs and we do not see much difference between these two methods.

To investigate the robustness of the simulation results above, we further conduct simulation experiments with predictors with highly correlated innovations. In section (ref), the oracle properties of the ALQR estimator are developed under the assumption that most predictor innovations are not highly correlated to each other. Although our empirical application satisfies this assumption, we provide additional simulation results in Tables (ref)-(ref) of the appendix to demonstrate the robustness of ALQR in various settings of correlated predictor innovations. Specifically, we generate the $I$(1) predictors in Scenario 1-3 with correlated innovations. We consider correlation ${\Greekmath 011A}=$0.1, 0.5, and 0.9. We find that when the true DGP includes zero coefficients, ALQR is better than other alternative methods in most of the settings. Only when DGP has no zero coefficient ALQR is at risk of underperforming across quantiles, though the performance of the other methods is only slightly better than ALQR. An interesting finding is that ridge regression, commonly recommended when many predictors are highly correlated, has a mixed performance across quantiles. Based on the simulation results, we recommend using ALQR, especially in the case of sparse data. But ridge regression and other methods may be used if sparsity is not found in the data and the quantile of interest is around the center of the distribution.

In sum, the numerical experiments confirm that ALQR can provide satisfactory prediction performance in finite samples. The results are robust over different quantiles and simulation designs. Therefore, we can expect similar results in other quantile prediction applications when the predictors have mixed roots and are composed of $I(0)$, $I(1)$, and cointegrated processes.

Conclusion

In this paper, we show that the adaptive lasso for quantile regression (ALQR) is attractive in forecasting with stationary and nonstationary predictors as well as cointegrated predictors. The framework is general enough to include mixed roots but ALQR does not require any researchers' knowledge on the specific structure of each predictor nor the order of integration. In this general framework, we show that ALQR preserves the oracle properties. These advantages offer substantial convenience and robustness to empirical researchers working with quantile prediction using time series data.

We have focused on the case where the number of covariates, $p$, is allowed to grow as sample size increases, although $p$ is smaller than $n$. This framework justifies a wide range of practical applications in economics, such as the stock return quantile prediction in this paper. It would be an interesting future research to allow $p$ to be even larger than $n$, which has not been studied in a general time series framework with mixed roots.

appendix