Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
90,706 characters · 11 sections · 126 citation commands
Inference in Predictive Quantile Regressions
\thispagestyle{empty}
In this paper, we develop asymptotic theory in the context of predictive quantile regressions with nearly integrated regressors. Beginning with influential work by Shiller84, Campbell&Shiller88a,Campbell&Shiller88b, Fama&French88 and Hod92, there has been extensive literature on testing whether a variety of proposed predictors can forecast mean stock returns. This has implications, not only for the risk neutral market efficiency hypothesis, but also for portfolio analysis. Indeed, subsequent empirical work debates the ability of investors to use predictors, such as dividend or earning price ratios, to create dynamic asset allocation strategies that outperform the market Goyal:Welch:2008,Campbell&Thompson2008.
While most empirical literature has focused exclusively on predictive means or variances, the portfolio decision often depends on the entire return distribution. Likewise, the tails of the distribution are of particular interest to risk managers and are also important to policymakers, who must consider the worst case, as well as baseline, forecast scenarios. Cenesizoglu:Timmermann:08 employ the quantile regression method introduced by KoenkerBassett78 to extract a richer set of return predictions. They find that a number of predictors have little information for the center of the distribution yet have important and often asymmetric implications for the tails.
One reason that the ongoing debate over predictive mean regression has lasted so long is that the limiting distribution of the standard $t$-statistic is nonstandard. Firstly, the predictor variables, such as dividend yields, dividend price and earning price ratios, are strongly autocorrelated. Secondly, although pre-determined, these predictors are not strictly exogenous because their innovations are often highly correlated with the error term in the predictive regression. Consequently, tests using the standard normal critical values will over-reject the null hypothesis of non-predictability, as is found by Mankiw&Shapiro86,Stambaugh86,Cavanagh/Elliott/Stock:95,stambaugh99.
Much attention has been devoted to overcoming such size distortions in predictive mean regressions, resulting in a rich literature. Perhaps the most popular approach has been the use of an explicit local-to-unity specification for the predictor.\footnote{The literature on mean predictive tests is too extensive to provide a full review here. Other prominent approaches include the IVX approach PhillipsMagdelliano2009wp,KostakisMagdalinosStamatogiannis:2015, nearly ElliotMullerWatson15 and conditionally optimal tests JanssonMoreira06, linear projection methods Cai14joe and inference based on small sample distributions in parametric models Nelson&Kim93,stambaugh99,Lewellen04, to name just a few.} Cavanagh/Elliott/Stock:95 propose corrected critical values based on a local-to-unity model with known values of the local-to-unity parameter ($c$). Since this parameter cannot be consistently estimated, they propose feasible inference methods using a Bonferroni bound and confidence interval on $c$ based on stock91. cy06 develop an efficient test of predictability for a known local-to-unity parameter $c$. Since their correction depends on $c$, a refined Bonferroni bounds procedure is employed for feasible inference. Hjalmarsson:07 notes that the cy06 procedure can be interpreted as a local-to-unity version of the fully modified estimator of Phillips&Hansen90, and HjalmarssonJFE2011 proposes a generalization to long horizon returns.
In contrast to this large literature on predictive mean regression, we are aware of no theoretical work prior to our original working paper version MaynardShimotsuWang2011 that establishes valid econometric inference methods in quantile predictive regression with persistent regressors. In this paper, we develop proper inference methods for short-horizon predictive quantile regressions with nearly integrated regressors. This paper makes three main contributions. First, we derive the limit distribution of the quantile regression coefficients by generalizing results of xiao:09, who derives inference in a quantile regression with cointegrated time series, to the local-to-unity setting. Second, we derive the asymptotics of heteroskedasticity and autocorrelation consistent (HAC) covariance matrix estimate and $t$-statistic. In contrast to predictive mean regression, the error terms in predictive quantile regression can be serially correlated. For example, when the stock return contains a GARCH component, the quantiles of the stock return are serially correlated because large returns are followed by large returns. Therefore, it is essential to use a HAC covariance matrix estimate and a HAC $t$-statistic. Existing literature in predictive quantile regression, such as Lee2016, FanLee19joe, and cai23joe, assume the error terms are serially uncorrelated. Consequently, their asymptotic results no longer hold, for example, when the stock return contains a GARCH component. As in the case of predictive mean regression, the limiting distribution of the standard and HAC $t$-statistics are nonstandard, and the standard inference procedures are unreliable when predictors are both persistent and endogenous. Third, we provide an inference procedure that is valid both when the predictor is a local-to-unity process and when the predictor is stationary. When the largest root of the predictor lies in the near unit root range, we provide a fully modified bias correction to the quantile regression estimator. This is equivalent to a quantile version of the correction in cy06. Phillips14 has proven that the predictive tests of Cavanagh/Elliott/Stock:95 and cy06 become invalid if the predictor is stationary. To address this problem, we follow in the spirit of ElliotMullerWatson15 and switch to a standard predictive quantile regression HAC $t$-test with a slightly conservative critical value when the largest root lies in the stationary range. We refer to this as a switching-FM predictive quantile regression test. Our Monte Carlo simulations verify that the switching-FM quantile regression test has good size and power both when the predictor is a local-to-unity process and when the predictor is stationary.
Subsequent to MaynardShimotsuWang2011, Lee2016 develops the IVXQR test that uses a mildly integrated instrument generated by filtering the original predictor. Like our switching-FM test, the IVXQR test avoids the problems noted by Phillips14. FanLee19joe extend the IVXQR test to allow for heteroskedasticity and suggest a bootstrap inference. Recently, cai23joe develop a new test, $t^w$, that uses an auxiliary regressor formed by a weighted combination of an exogenous simulated nonstationary process and a bounded transformation of the original regressor. In cai23joe's simulation results, their $t^w$ test has better finite sample size and power than the IVXQR test. In our simulations, we find that the switching-FM test has higher power than the $t^w$ test with nearly comparable size, and the $t^w$ test has a modest size advantage at tail quantiles. On the other hand, the switching-FM test is designed for a single predictor. The tests of Lee2016, FanLee19joe, and cai23joe have the distinctive advantage of generalizing easily to a multi-predictor setting. GungorLugo2019 develop a maximized Monte Carlo approach to exact finite sample inference in predictive quantile regression, but at the cost of requiring i.i.d.\ return innovations.
The predictive quantile regression is more distantly related to sign, sign-rank, and directional tests of predictability Campbell&Dufour95,Campbell&Dufour97,GungorLugo2020 and to the (cross) quantilogram Linton:Whang:2007, HanLintonTaksushiWang16, LeeLintonWhang20. More broadly, our results contribute to a rapidly developing literature in quantile regression for time series data. Although too numerous to survey here, developments include quantile autoregression KoenkerXiao2006,ChenKoenkerXiao2009, dynamic quantile models EngleManganelli2004,GourierouxJasiak2008, unit root quantile autoregression KoenkerXiao2004,Galvo2009, and quantile cointegration xiao:09,cho15joe.
The remainder of the paper is organized as follows. Section 2 establishes the framework of the problem and develops the asymptotic theory for predictive quantile regression under a local-to-unity specification. In Section 3, a switching-FM predictive quantile test is proposed. In Section 4, results from our simulation study are reported. In Section 5, the techniques are applied to test the predictability of the stock return distribution using three commonly employed valuation predictors. Section 6 concludes the paper. The appendix provides proofs, and tables are included at the end.
In matters of notation, let $Q_{a_{t}}(\tau)$ and $Q_{a_{t}}(\tau|\mathcal{F})$ denote the unconditional and conditional $\tau$-quantile of $a_t$ conditional on $\mathcal{F}$. Let $\|A\|_r = (\sum_{ij}E|a_{ij}|^r)^{1/r}$ denote the $L^r$-norm. Let $:=$ denote “equals by definition.” Let $\Rightarrow$ denote weak convergence of the associated probability measures. Let $\equiv$ denote equality in distribution. Let $BM(\Omega)$ denote a Brownian motion with covariance matrix $\Omega$. Let $[x]$ denote the largest integer less than or equal to $x$. Let $MN(0,\Omega)$ denote the mixed normal distribution with variance $\Omega$. Continuous stochastic processes such as Brownian motion $B(r)$ on $[0,1]$ are usually written simply as $B$, and integrals $\int$ are understood to be taken over the interval $[0,1]$, unless specified otherwise. Let $I\{\cdot\}$ denote the indicator function. All limits below are taken as $T \to \infty$ unless stated otherwise.
We model the conditional $\tau$-quantile of $y_t$ as
where $y_t$ is typically a financial return, $x_t$ is a predictor, such as earnings or dividend price ratio, $\mathcal{F}_{t-1}$ is the information contained in the lags of $x_t$, $\gamma(\tau):= (\gamma_0(\tau),\gamma_1(\tau))'$, and $z_t:= (1,x_t)'$. In model ((ref)), the dependence of $\gamma_{1}(\tau)$ on $\tau$ allows the impact of $x_{t-1}$ to vary across the quantiles of $y_t$.
As noted by Cenesizoglu:Timmermann:08, the predictive quantile model is robust to outliers and encompasses a number of other empirical models for financial returns. For example, if $x_{t}$ is a variable with predictive content for volatility, such as squared returns or realized volatility, we may consider a model of the form $y_t = b_0 + b_1 x_{t-1} + (c_0 + c_1 x_{t-1} ) e_{t}$, where $e_t$ is independent of $\mathcal{F}_{t-1}$. The predictive quantile for $y_t$ then takes the form $Q_{y_t}(\tau | \mathcal{F}_{t-1}) = b_0 + c_0 Q_{e_t}(\tau) + (b_1 + c_1 Q_{e_t}(\tau)) x_{t-1}$. The predictive quantile model also encompasses a random-coefficient model KoenkerXiao2006
where $e_t \sim U[0,1]$ is independent of $\mathcal{F}_{t-1}$. Provided that the right hand side is monotone increasing in $e_t$, the predictive quantile for $y_t$ in this model is $Q_{y_t}(\tau | \mathcal{F}_{t-1}) = b_0(\tau) + b_1 (\tau) x_{t-1}$.
Next, we consider the data-generating process for the predictor. Since most predictors employed in practice are highly persistent, we model the regressor $x_t$ as a near-unit root process. Specifically, we assume that
where $x_0 = o_p(T^{1/2})$, $v_t$ is a mean-zero stationary process, and $T$ is the sample size. A number of prior studies have used this framework to model predictors such as earnings and dividend price ratios, which are highly persistent but a priori stationary on economic grounds.
The standard quantile regression coefficient estimates are given by
where $\rho_\tau(u):=u\psi_\tau(u)$ with $\psi_\tau(u):=\tau-I(u<0)$ as in KoenkerBassett78. When $\tau=0.5$, ((ref)) gives the least absolute deviation estimator. Define $\mathcal{F}_{t-1} := \sigma\{x_{t-j}, j\geq 1\}$, and
where the second equality follows from (ref). Since $\psi_\tau(u)=\tau-I(u<0)$, we have $E\left[\psi_\tau(u_{t\tau})\middle| \mathcal{F}_{t-1} \right]=0$ and $\text{var}\left[\psi_\tau(u_{t\tau})\middle| \mathcal{F}_{t-1} \right]=\tau(1-\tau)$.
In the literature, it has been recognized that the model $Q_{y_{t}}(\tau|\mathcal{F}_{t-1}) = \gamma_{0} + \gamma_{1}(\tau)x_{t-1}$ with quantile-varying $\gamma_{1}(\tau)$ poses difficulties for asymptotic analysis with local-to-unity regressors. When $x_{t-1}$ is a local-to-unity process and $\gamma_{1}(\tau)$ varies with $\tau$, $u_{t\tau}$ contains a local-to-unity component for some $\tau$ because, if $ \gamma_1(\tau_1) \neq \gamma_1(\tau_2)$ for some $\tau_1 \neq \tau_2$, at least one of $u_{t\tau_1}=y_{t}-\gamma_0(\tau_1) - \gamma_1(\tau_1) x_{t-1}$ or $u_{t\tau_2} = y_{t}-\gamma_0(\tau_2) - \gamma_1(\tau_2) x_{t-1}$ contains a local-to-unity component. Consequently, the current literature assumes $\gamma_{1}(\tau)=\gamma_1$ for all $\tau \in (0,1)$ either explicitly or implicitly; see Lee2016, FanLee19joe, and cai23joe.
In view of this, we explicitly impose $\gamma_{1}(\tau)=\gamma_1$ for all $\tau \in (0,1)$ as in xiao:09 and assume $u_{t\tau}$ is stationary in Assumption (ref) below. This rules out some interesting models, such as the random-coefficient model ((ref)). We address this problem in Section (ref) by considering models for which $\gamma_1(\tau)$ varies in $\tau$ but only locally. Specifically, we will analyze the power of our tests under the model
where $e_t$ is independent of $\mathcal{F}_{t-1}$, $b(\cdot)$ is increasing, and $\zeta>0$ is non-random. In this model, the predictive quantile for $y_t$ is $Q_{y_t}(\tau | \mathcal{F}_{t-1}) = \gamma_0 (\tau) + \gamma_1 x_{t-1} + T^{\kappa-1}b (Q_{e_t}(\tau)) |x_{t-1}+\zeta|$, where $\gamma_0(\tau) = \gamma_0 + Q_{e_t}(\tau)$, and $x_{t-1}$ has a local quantile-varying effect, $T^{\kappa-1}b (Q_{e_t}(\tau)) |x_{t-1}+\zeta|$. Section (ref) shows that our test statistic rejects $H_0: \gamma_{1}(\tau)=0$ with probability approaching one when $\gamma_1=0$ and $b(Q_{e_t}(\tau)) \neq 0$. In other words, our test can detect the existence of a local quantile-varying predictive component.
We collect the assumptions. Let $U_t(\tau) := (\psi_\tau(u_{t\tau}),v_t)'$.
Assumption (ref) is similar to Assumption 2.1(i) in Lee2016 and Assumption A2(i) in cai23joe. Assumption (ref) is essentially the same as Assumption 1 in Hansen:92. Note that the condition $\lim_{n\to\infty}n^{-1}E(V_{n}V_{n}') =\Omega<\infty$ in Hansen:92 is satisfied because we assume $U_t(\tau)$ is stationary. We assume $\beta \geq 3$ because our proof uses Theorem 4.2 of Hansen:92. Assumption (ref) imposes the mixing condition on both $u_{t \tau}$ and directly on $f_{u_\tau, t-1}(0)$.\ For GARCH($p,q$) and augmented GARCH models, carrascochen02et show conditions under which the squared residuals $u_t^2$ and the latent conditional volatility process $\sigma_{t+1}^2$ are jointly mixing. Then, $f_{u_\tau, t-1}(0)$ is also mixing because we can write the conditional density of $u_t$ as a finite lag function of $(u_t^2, \sigma_{t+1}^2)$.
It follows from Theorem 4.4 of Hansen:92 and Assumption (ref) that
The first result in (ref) corresponds to Assumption A of xiao:09.
The following proposition provides the limiting distribution of the predictive quantile regression estimator in ((ref)). Define $D_T :=\text{diag}(T^{1/2},T)$.
It follows from Proposition (ref) that
The asymptotic distribution is nonstandard. When $c=0$, it specializes the result of the quantile cointegrating regression xiao:09 to the case of predictive regression. The extension to $c<0$ was first derived in MaynardShimotsuWang2011 under the additional assumption $Q_{u_t}(\tau|\mathcal{F}_{t-1})=Q_{u_t}(\tau)$. Lee2016 derives the asymptotic distribution when $x_t = (1+c/T)x_{t-1}+v_t$ with both positive $c$ and $x_t = (1+b/T^\alpha)x_{t-1}+v_t$ with $\alpha \in (0,1)$ (mildly integrated $x_t$ and mildly explosive $x_t$). FanLee19joe derive the asymptotic distribution of the predictive quantile regression estimator when $y_t-\gamma_0(\tau)$ contains a conditionally heteroskedastic error of the form $\sigma_t \varepsilon_t$, under the restriction that $\psi_\tau(u_{t\tau})$ is a martingale difference sequence.
As in the case of cointegrating regression, some further insight into the bias can be gained from projecting $B_{\psi}$ onto $B_v$ Phillips89ET. Conformable to $(B_\psi,B_v)$, we partition $\Omega_\tau$ into\footnote{We thank Ji Hyung Lee for pointing out a typo in MaynardShimotsuWang2011 (our earlier working paper), which had $\delta$ in place of $\delta_{\tau}$.} \[ \Omega_{\tau} =
=
, \] where $\delta_\tau:= \omega_{\psi v}/(\omega_{\psi}\omega_{v})$. If $\psi_\tau(u_{t\tau})$ is serially uncorrelated, then $\omega_{\psi}^2$ is simplified to $\tau(1-\tau)$.
In general, $\psi_\tau(u_{t\tau})$ is serially correlated. For example, when $y_t$ follows a GARCH process, the stochastic process $I(y_t < Q_{y_t}(\tau))$ is serially correlated for $\tau\neq 0.5$ because one large value of $y_{t-1}$ is likely to be followed by another large value of $y_t$. Define $\omega_{\psi.v}^2 := \omega_{\psi}^2 - \omega_{v}^{-2} \omega_{\psi v}^2= \omega_{\psi}^2(1-\delta_\tau^2)$ and $B_{\psi. v} := \omega_{\psi.v}^{-1}(B_{\psi} - \omega_{v}^{-2} \omega_{\psi v} B_v)= \omega_{\psi.v}^{-1}(B_{\psi} - \omega_{v}^{-1} \omega_{\psi}\delta_\tau B_v)$, then $B_{\psi. v}$ is $BM(1)$ and independent of $B_v$. Using the decomposition $B_\psi = \omega_{v}^{-2} \omega_{\psi v} B_v + \omega_{\psi.v} B_{\psi.v}$, we may express (ref) as
The first stochastic integral inside the square brackets is the local-to-unity generalization of the (demeaned) Dickey-Fuller distribution and contributes a downward (upward) second-order bias to the estimate of $\widehat{\gamma}_1(\tau)$ for $\delta_{\tau} >0$ ($\delta_{\tau}<0$). The extent of the bias depends on both $\delta_{\tau}$ and on $c$. The second term in brackets is mixed normal, and normal conditional on $\mathcal{F}_v = \sigma\left( B_v(r), 0\leq r \leq 1\right)$. As in the case of linear predictive regression, there is no endogeneity term because $E\left[\psi_\tau(u_{t\tau})\middle| \mathcal{F}_{t-1} \right]=0$. The distribution of the estimator depends on $\tau$ through both $\delta_{\tau}$ and $B_{\psi.v}$.
We provide HAC standard errors, denoted by $\text{se}(\widehat{\gamma}_1)$, and then derive the asymptotic distribution of the HAC $t$-statistic $(\widehat{\gamma}_1(\tau)-\gamma_1) / \text{se}(\widehat{\gamma}_1)$ when $x_t$ is local-to-unity as in (ref). Define $\Delta_{fz}(\tau) := E[ f_{u_{t \tau},t-1}(0) z_{t-1} z_{t-1}']$, and define the long-run variance of $w_{t \tau}:=z_{t-1} \psi_\tau(u_{t\tau})$ as $\Sigma(\tau) := \sum_{\ell =-\infty}^{\infty}\Gamma(\ell)$, where $\Gamma(\ell):= E[w_{(t +\ell)\tau}w_{t \tau}' ]$, suppressing the dependence of $\Gamma(\ell )$ on $\tau$. The HAC standard error, $\text{se}(\widehat{\gamma}_1)$, of $\widehat \gamma_1(\tau)$ is defined as the square root of the $(2,2)$th element of $T^{-1} \widehat \Delta_{fz}^{-1}(\tau) \widehat \Sigma(\tau) \widehat \Delta_{fz}^{-1}(\tau)$, where \[ \widehat \Delta_{fz}(\tau) := \frac{1}{Th} \sum_{t=1}^T \phi \left(\frac{\widehat u_{t\tau}}{h}\right)z_{t-1} z_{t-1}', \quad \widehat \Sigma(\tau) := \sum_{\ell =-m}^m k \left(\frac{ \ell}{m} \right) \widehat \Gamma(\ell ), \] $\widehat{u}_{t\tau} := y_t - \widehat{\gamma}(\tau)'z_{t-1}$, $\phi(\cdot)$ and $k(\cdot)$ are kernels, $h$ is the bandwidth, and $m$ is the lag length. The autocovariance estimate $\widehat \Gamma(\ell )$ is computed as \[ \widehat \Gamma(\ell ) :=
\] where $\widehat{w}_{t \tau}:=z_{t-1} \psi_\tau(\widehat u_{t\tau})$. When $x_t$ is stationary and follows
with a fixed $\phi \in (-1,1)$, xiao12hdbk and galvao23jasa show $\sqrt{T}(\widehat \gamma(\tau) - \gamma(\tau)) \to_d N(0,\Delta_{fz}(\tau)^{-1} \Sigma(\tau) \Delta_{fz}(\tau)^{-1})$ and $(\widehat{\gamma}_1(\tau)-\gamma_1) / \text{se}(\widehat{\gamma}_1) \rightarrow_d N(0,1)$.
We introduce some additional assumptions for analyzing the asymptotics of $\text{se}(\widehat{\gamma}_1)$.
Many probability density functions, including the standard normal density, satisfy Assumption (ref)(a). Define $\mathcal{F}^*_{t-1}:= \sigma\{y_{t-j},x_{t-j}, j\geq 1\}$, and let $f_{u_{t \tau},t-1}^*(x)$ denote the conditional density of $y_t - \gamma_0(\tau)$ conditional on $\mathcal{F}^*_{t-1}$.
The following proposition shows the null limiting distribution of the standard HAC $t$-statistic for testing $H_0: \gamma_1(\tau)=\gamma_1$.\footnote{When $\psi_\tau(\widehat u_{t\tau})$ is serially uncorrelated, this result is originally derived in Proposition 2 of MaynardShimotsuWang2011. A similar result is also shown in Lee2016. } The null limiting distribution depends on $c$ and $\delta_{\tau}$.
In the stock return predictability example, it is appropriate to model many predictors, such as the dividend price ratio, as near unit root processes as in (ref) with $c<0$. In this case, as shown above, the limiting distribution of $t_{\gamma_1}(\tau)$ is nonstandard and dependent on the nuisance parameters $c$ and $\delta_{\tau}$. When $\delta_\tau \neq 0$, the predictability test with standard normal critical values tends to over-reject the null hypothesis. This is a problem in practice because financial data, such as prices and dividends, do not satisfy strict exogeneity. The over-rejection is especially severe when the residual cross-correlation is large.
For a given value of $c$, the first term in (ref), which causes the asymptotic bias in $\widehat{\gamma}_1(\tau)$, can be removed using a specialization of the fully modified (FM) approach to the predictive quantile regression framework. Define
where $\phi = 1+c/T$, and $\widehat{f_{u_\tau}(0)}$, $\widehat{\omega}_\psi$, $\widehat{\omega}_v$, $\widehat{\delta}_{\tau}$, and $\widehat{\lambda}_{vv}$ are consistent estimators of $f_{u_\tau}(0)$, $\omega_\psi$, $\omega_v$, $\delta_{\tau}$, and $\lambda_{vv} := \sum_{k=1}^\infty E v_0 v_k=1/2 (\omega_v^2 - Ev_0^2)$, respectively. Because $T^{-1}\sum_{t=1}^T x_{t-1}^{\mu} (x_t -\phi x_{t-1}) \Rightarrow \int J_c^\mu dB_{v}+\lambda_{vv}$, the terms in brackets remove the first term in (ref).
In our simulation and empirical application, we use\ a standard kernel density estimator $\widehat{f_{u_\tau}(0)}=1/(Th)\sum_{t=1}^T k(\widehat{u}_{t\tau,t-1}/h)$, where the bandwidth $h$ is chosen by Silverman:86's rule of thumb, and we estimate $\omega_\psi$, $\omega_v$, $\delta_{\tau}$, and $\lambda_{vv}$ nonparametrically. See Section (ref) for details.
Define the standard error for the fully modified estimator as $\mbox{se}(\widehat{\gamma}_1(\tau,c)^{+}) := \\ (\widehat{\omega}_{\psi.v} / \widehat{f_{u_\tau}(0)}) (\sum_{t=1}^T (x_{t-1}^{\mu})^2 )^{-1/2}$, where $\widehat{\omega}_{\psi.v} = (\widehat\omega_\psi^2(1-\widehat{\delta}_{\tau}^2))^{1/2}$. The following proposition shows that the fully modified estimator has a mixed normal asymptotic distribution and the associated $t$-statistic has a standard normal null asymptotic distribution.
The following corollary provides the asymptotic distribution of the fully modified $t$-statistic for testing $H_0: \gamma_1(\tau)=0$ under a local alternative.
Consider testing $H_0: \gamma_1(\tau)=0$ against $H_A:\gamma_1(\tau) \neq 0$. If the value of $c$ is known, the test that rejects $H_0$ when $|t_{\gamma_1}(\tau,c)^+|>z_{1-\alpha}$ has the asymptotic size $\alpha$. In practice, $c$ is unknown and cannot be consistently estimated. We follow the approach of cy06 and obtain a conservative testing procedure employing Bonferroni bounds. Inverting the GLS-ADF unit root test of Elliott&Rothenberg&Stock96 on $x_t$ in the spirit of stock91 yields a first-stage confidence interval for $c$ with confidence level $\alpha_1$, which we refer to as $\mbox{CI}_c(\alpha_1)$. A Bonferroni test rejects $H_0: \gamma_1(\tau)=0$ in favor of $H_A:\gamma_1(\tau) \neq 0$ if $\max_{c^*\in \mbox{CI}_c(\alpha_1)} |t_{\gamma_1}(\tau,c^*)^+| \geq z_{1-\alpha_2}$, where $\alpha_2$ is chosen so that $\alpha_1+\alpha_2 \leq \alpha$. By the Bonferroni inequality, the asymptotic size of this test is no greater than $\alpha$.\footnote{Equivalently, we can form a confidence interval $\mbox{CI}_{\gamma_1} (\alpha_2,\tau,c)$ of $\gamma_1(\tau)$ with level $1-\alpha_2$ for each value of $c$, define $\mbox{CI}_\gamma(\alpha,\tau):=\cup_{c^*\in \mbox{CI}_c(\alpha_1)}\mbox{CI}_{\gamma_1} (\alpha_2,\tau,c^*)$, and reject $H_0$ when $\mbox{CI}_\gamma(\alpha,\tau)$ does not contain $0$.}
Phillips14 points out that the stock91 confidence interval becomes invalid with asymptotic coverage probability zero when $x_t$ is stationary, leading to the invalidity of Bonferroni tests that depend on it. Indeed, when $\delta_{\tau}<0$ and $c$ is large negative, we find that the Bonferroni predictive quantile test becomes extremely conservative against the right-sided alternative and extremely oversized against the left-sided alternative. On the other hand, we found the test to work quite well in the near unit root range. We also noticed similar size problems when $c$ takes very large positive (explosive) values, even though such values are considered implausibly large by most of the predictive regression literature. In view of this, we assume $c \leq 4$ henceforth and rule out implausibly large positive values of $c$.\footnote{cy06 assume $c \leq 5$.}
A useful observation is that the persistence ranges for which the Bonferroni test breaks down are also the ranges in which the standard predictive quantile tests work reasonably well. Table (ref) shows the $5$ and $95$ percentiles of $Z(c,\delta_\tau)$ defined in ((ref)) for selected values of $(c,\delta_\tau)$. These percentiles are simulated by approximating $B_v(r)$ by $T^{-1/2}\sum_{t=1}^{[Tr]}v_t$ for $v_t\sim i.i.d.\ N(0,1)$ using $1,000,000$ replications with $T=10,000$. As $c \to -\infty$, $Z(c,\delta_\tau)$ converges to $N(0,1)$ Phillips14. When $c \ll 0$, the 5 and 95 percentiles of $Z(c,-1)$ are similar to those of $N(0,1)$, and neither quantile changes very much as $\delta_\tau$ changes. Therefore, when $c \ll 0$, we can use the 5 percentile of a $N(0,1)$ and the 95 percentile of $Z(c^*,-1)$ for $c<c^*$ as conservative critical values for the standard $t$-test without sacrificing much power.
Because the density of $Z(c,\delta_\tau)$ is asymmetric, henceforth we consider separately the right-tailed test of $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)>0$ with level $(1-\alpha_2/2)$ and the left-tailed test of $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)<0$ with level $(1-\alpha_2/2)$. To preserve the good properties of the test in the near unit root range while addressing the problems that occur outside it, we propose a switching version of the quantile FM predictive test in the spirit of ElliotMullerWatson15. First, consider the right-tailed test of $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)>0$. Fix a switching threshold $\overline c_{L}<0$. If the first-stage confidence interval $\mbox{CI}_c(\alpha_1)=[\underline c, \overline c]$ lies entirely within the near unit root range, i.e., $\underline{c} > \overline c_L$ we employ only the Bonferroni FM test. When $\mbox{CI}_c(\alpha_1)$ lies entirely outside of the near unit root region, i.e., $\overline{c}< \overline c_L$, we use only the HAC $t$-test with the critical value from $Z(\overline c_L,-1)$. Finally, when $\mbox{CI}_c(\alpha_1)$ lies only partly in the near unit root region, i.e., $\underline{c} \leq \overline c_L \leq \overline{c}$, we employ both tests and reject the null hypothesis only if both tests reject. In the left-tailed test of $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)<0$, we fix $\underline c_{L}<0$, which can be different from $\overline c_L$, and proceed similarly to the right-tailed test but use the critical value from a $N(0,1)$ both in the Bonferroni FM test and the $t$-test. This is because the 5 percentile of the $N(0,1)$ serves as the conservative critical value for the $t$-test.
Consider testing $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)>0$. Define $\overline\varphi_{FM}(\tau,\alpha_1,\alpha_2):=I\{\min_{c^*\in \mbox{CI}_c(\alpha_1)} t_{\gamma_1}(\tau,c^*)^+ \geq z_{1-\alpha_2/2}\}$ and $\overline\varphi_{t}(\tau,\alpha_2,\overline c_L):=I\{ t_{\gamma_1}(\tau) \geq z_{1-\alpha_2/2}(\overline c_L)\}$, where $t_{\gamma_1}(\tau) $ and $t_{\gamma_1}(\tau,c)^+$ are defined as in Propositions (ref) and (ref) with $\gamma_1=0$, and $z_{1-\alpha_2/2}(c)$ is the $100(1-\alpha_2/2)$ percentile of $Z(c,-1)$. Define the test statistic
The switching-FM test rejects $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)>0$ if $\overline\varphi(\tau, \alpha_1, \alpha_2, \overline c_L)=1$. For testing $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)<0$, defining $\underline\varphi_{FM}(\tau,\alpha_1,\alpha_2):=I\{\max_{c^*\in \mbox{CI}_c(\alpha_1)} t_{\gamma_1}(\tau,c^*)^+ \leq z_{\alpha_2/2}\}$ and $\underline\varphi_{t}(\tau,\alpha_2,\underline c_L):=I\{ t_{\gamma_1}(\tau) \leq z_{\alpha_2/2}\}$ and defining $\underline\varphi(\tau, \alpha_1, \alpha_2, \underline c_L)$ as in (ref) gives the switching-FM test statistic. The two-tailed switching-FM test rejects $H_0:\gamma_1(\tau)=0$ against $H_A: \gamma_1(\tau)\neq 0$ if $\overline\varphi(\tau, \alpha_1, \alpha_2, \overline c_L) + \underline\varphi(\tau, \alpha_1, \alpha_2, \underline c_L)=1$.
The following proposition shows that the asymptotic size of the one-tailed and two-tailed switching-FM test does not exceed $\alpha_1 + \alpha_2/2$ and $\alpha_1 + \alpha_2$, respectively.
In mean predictive regression, cy06 consider the uniformly most powerful (UMP) test of $H_0:\gamma_1=0$ against $H_1:\gamma_1>0$ when $c$ is known, and their $Q$-test takes a union of the UMP test over $c$ using the Bonferroni method. ElliotMullerWatson15 establish the optimal test against an alternative model that integrates $c$ and $\gamma_1$ with respect to a user-chosen probability distribution $F$. Figure 4 of ElliotMullerWatson15 shows that their test is more powerful than cy06's $Q$-test for many values of $c$, in particular when $c$ is close to 0, but the $Q$-test is more powerful for certain values of $c$. Hjalmarsson:07 notes that the $Q$-test can be interpreted as a local-to-unity version of the fully modified $t$-test of Phillips&Hansen90. Therefore, when $u_t$ follows a Laplace distribution, our fully-modified quantile regression test would be asymptotically optimal when $c$ is known.
In mean predictive regression inference, Cavanagh/Elliott/Stock:95 and cy06 find that the Bonferroni test, as described above, tends to be excessively conservative. They simulate the asymptotic distribution of their Bonferroni test statistic and adjust the first-stage confidence level $\alpha_1$ to mitigate their test's conservative nature.
Similar to cy06, we simulate the asymptotic distribution of the switching-FM test statistic under the null hypothesis with a large $T$ ($T=5,000$) and adjust the first-stage confidence level $\overline\alpha_1$ for the right-tailed test and $\underline\alpha_1$ for the left-tailed test so that the test is slightly conservative. We fix $\alpha_2$. Let $\tilde \alpha_2 = \alpha_2 - \epsilon$ for a small $\epsilon>0$, and let $Q_c:=[\underline q, \overline q]$ define a region for $c$. Then, we select $\overline\alpha_1$ and $\underline\alpha_1$ over the grid $\{0.01, 0.02, \ldots, 0.98\}$ so that
holds for all $c \in Q_c$ and with equality for some $c \in Q_c$. The following proposition shows that, when $\underline q$ is chosen sufficiently large negative, this version of the one-tailed switching-FM test has asymptotic size no larger than $\alpha_2/2$ both when $x_t$ is stationary and when $x_t$ is local-to-unity, including the case $c <\underline q$.
In practice, one needs to choose the value of $\overline c_L$ and $\underline c_L$. Making $\overline c_L$ more negative makes the test less conservative for $c < \overline c_L$ because $z_{1-\alpha_2/2}(\overline c_L)$ becomes smaller. On the other hand, this makes the Bonferroni FM test more conservative for $c > \overline c_L$. We set $\alpha_2 = 0.1$, $\epsilon = 0.04$, $Q_c=[-120,4]$ and choose the value of $\overline c_L$ from $\{-120,-110,-100,\ldots,-30\}$ to minimize the average under-rejection\footnote{The under-rejection is calculated as the average of the difference between 0.05 and the rejection frequency. As seen in Proposition (ref), this difference is always positive.} over $c \in\{-200,-180,-160, \ldots, 0\}$ and $\delta_\tau \in \{-0.797, -0.598, -0.399, -0.199, 0\}$,\footnote{These values of $\delta_\tau$ correspond to the correlation coefficient $\delta \in \{-0.999, -0.75, -0.50, -0.25, 0\}$ between $u_t$ and $v_t$ when $\tau =0.5$.} where $\overline\alpha_1$, which depends on $\overline c_L$, is chosen to satisfy (ref). The under-rejection probability is approximated using simulations with $T=5,000$ and $10,000$ replications, where $(u_t,v_t)$ is drawn from a bivariate normal distribution. We thus select $\overline c_L$ to minimize average under-rejection both inside and below $Q_c$ while complying with the requirements of Proposition (ref) to avoid any over-rejection. This procedure gives $\overline c_L = -90$. We obtain $\underline c_L = -100$ by a similar procedure.
Table (ref) shows the resulting adjusted significance levels $\overline{\alpha}_1$ and $\underline{\alpha}_1$ used for the confidence intervals on $c$ for the right and left-tailed predictive tests, respectively. Table (ref) also shows the values used for $\delta$ and the corresponding values of $\delta_{\tau}$. $\delta$ and $\delta_{\tau}$ are closely related, although $\delta_{\tau}$ is generally smaller in magnitude than $\delta$. We use the value of $\delta_{\tau}$ when employing these lookup tables. Both $\underline{\alpha}_1$ and $\overline{\alpha}_1$ are smaller when $|\delta_{\tau}|$ is larger and the persistent regressor problem is worse. Even then, however, they are well above five percent, suggesting that the adjustment can help to reduce the conservativeness of the Bonferroni test procedure.
In this section, we analyze the asymptotic power of the switching-FM test under local quantile-varying alternatives with local-to-unity $x_t$. In order to facilitate asymptotic analysis, we model $y_t$ as a random-coefficient process similar to KoenkerXiao2006:
where $e_t$ is independent of $\mathcal{F}_{t-1}$, $b(\cdot)$ is weakly monotone increasing with $\sup_{x,y} [b(x)-b(y)]/(x-y) = M_b < \infty$, and $\zeta>0$ is non-random. When $\zeta$ is sufficiently large, $x_{t-1}+\zeta>0$ holds for all $t$ in the observed data, and the absolute value hardly matters in practice.
The predictive quantile for $y_t$
where $\gamma_0(\tau) = \gamma_0 + Q_{e_t}(\tau)$, has a local quantile-varying component. Define $e_{t\tau}:=e_t - Q_{e_t}(\tau)$. Under ((ref)), we have
Because $\kappa<1/2$, $e_{t\tau}$ dominates the right hand side, and $u_{t\tau}$ behaves like an $I(0)$ process. We show that, with some restrictions on $\kappa>0$, the switching-FM test rejects $H_0:\gamma_1(\tau)=0$ with probability approaching one.
We collect assumptions. Assumption (ref) corresponds to Assumption 2.1(ii) in Lee2016.
The following proposition shows the asymptotic distribution of the QR estimator under local alternatives ((ref)). Define $D_T^\kappa := \text{diag}(T^{1/2-\kappa},T^{1-\kappa})$ and $\gamma(\tau):=(\gamma_0(\tau), \gamma_1)'$.
The asymptotic distribution is a functional of $b(Q_{e_t}(\tau))$ and $J_c(r)$. Under the local alternative with $\gamma_1=0$, $\widehat{\gamma}_1(\tau)$ converges to 0 at a slower rate $T^{\kappa-1}$ than $T^{-1}$.
The following proposition shows the consistency of the switching-FM test under local alternatives ((ref)). The additional assumption $T^{\kappa-1/2}/h \to 0$ is necessary to control the convergence rate of the standard error. When one uses the optimal bandwidth $h = C T^{-1/5}$, the restriction on $\kappa$ becomes $\kappa \in (0,3/10)$, which is fairly weak.
Our Monte Carlo study has three primary objectives. First, we examine the extent of the size distortion of a standard predictive quantile $t$-test. Second, we study the size performance of the proposed switching-FM test across different levels of persistence and serial correlation, including for cases of volatility persistence, leverage, and fat-tails. Finally, we compare the small sample size and power of the switching-FM test to the $t^w$ test of cai23joe.\footnote{Comparisons of the $t^w$ test to Lee2016's IVXQR test can be found in cai23joe, where $t^w$ is generally found to have better size and higher power than IVXQR.} We evaluate test power under both the traditional linear regression alternative and under two random coefficient models: one from Section (ref), in which only the shoulder and tail quantiles are predictable, and a second specification from cai23joe, in which predictability is stronger in the upper quantiles than in the lower quantiles.
When implementing the switching-FM test, we estimate $\omega_\psi$, $\omega_v$, $\delta_{\tau}$, and $\lambda_{vv}$ by kernel-based heteroskedasticity and autocorrelation consistent (HAC) estimators. We use a Gaussian kernel for $\phi(\cdot)$ and the Bartlett kernel for $k(\cdot)$. Because $\psi_\tau(u_{t\tau})$ has persistent serial correlation when $y_t$ follows a GARCH process, we apply a heterogeneous VAR (HVAR)-prewhitening Corsi09jfemets to $U_t=(\psi_\tau(u_{t\tau}),v_t)'$. Specifically, we first fit a HVAR model $U_t = \beta_0 + \Phi^{(m)} U_{t-1}+ \Phi^{(q)} U_{t-1}^{(q)} + \Phi^{(y)} U_{t-1}^{(y)} + \varepsilon_t$ to $U_t$, where $U_{t-1}^{(q)} = (1/3)\sum_{j=1}^3 U_{t-j}$ and $U_{t-1}^{(y)} = (1/12)\sum_{j=1}^{12} U_{t-j}$. Because $v_t$ has a weak serial correlation in predictive regression models, and $\psi_\tau(u_{t\tau})$ is uncorrelated with lagged $v_t$'s by its definition, we impose zero restrictions on the $(1,2)$th element of $\Phi^{(m)}$, $\Phi^{(q)}$, and $\Phi^{(y)}$ and the second row of $\Phi^{(q)}$ and $\Phi^{(y)}$. After estimating the long-run variance matrix $\Omega_\varepsilon$ of $\varepsilon_t$ by a HAC estimator with the Andrews91 bandwidth choice, we obtain $\widehat \Omega_\tau$ by recoloring $\widehat{\Omega}_\varepsilon$ with the estimate of $\Phi^{(m)}$, $\Phi^{(q)}$, and $\Phi^{(y)}$. $\lambda_{vv}$ is estimated similarly. The bandwidth $h$ in $\widehat \Delta_{fz}$ is chosen by Silverman:86's rule of thumb.
We simulate the predictor from an autoregressive model of order one, as in (ref), employing $x_0=0$ as the starting value. We consider the odd decile values of the quantile level $\tau = 0.1, 0.3,\ldots, 0.9$ for sample sizes of $800$ and $1600$. All simulations are based on 10,000 Monte Carlo replications. Table (ref) shows null rejection rates of nominal five percent tests of $H_0:\gamma_1(\tau)\leq 0$ against $H_A:\gamma_1(\tau)>0$ when $y_t$ is generated by $y_t=e_t$ and $(v_{t},e_{t})$ follows an i.i.d.\ bivariate normal distribution with means equal to zero, unit variances, and correlation $\delta$. We select $\delta<0$ because empirical estimates of $\delta$ using valuation predictors, such as the earning price ratio or dividend price ratio, are generally negative. Similarly, we conduct one-sided versions of all the predictive tests since $\gamma_1>0$ is the relevant alternative hypothesis in empirical work using valuation predictors. We compare three tests: the conventional quantile regression $t$-test without HAC,\footnote{We use a non-HAC version of the standard $t$-statistic in Table (ref) because $e_t$'s in (ref) are serially independent.} the switching FM-test, and cai23joe's $t^w$ test. The quantile levels ($\tau$) vary across the table columns. We vary the local-to-unity parameter ($c$) across rows, including both the standard near unit range $-25<c\leq 0$ and $c=-200$, which entails more stationary behavior. The endogeneity is strong in the top two panels ($\delta=-0.95$) but moderate in the bottom two panels ($\delta=-0.5$). We increase the sample size from $T=800$ in the left panels to $T=1600$ in the right panels.
The results in Table (ref) indicate that the size problem in conventional predictive quantile regressions can be non-trivial, with rejection rates as high as $0.289$ even for a sample size of $1600$. This confirms both our original finding in MaynardShimotsuWang2011 and subsequent results in Lee2016 and cai23joe. The degree of size distortion depends heavily on both the local-to-unity parameter $c$ and the residual correlation $\delta$ (which impacts $\delta_{\tau }$). The table also confirms that the size distortion gradually dissipates as $c$ becomes more negative. This finding supports the use of the switching-FM test. It is also in line with Lee2016's theoretical finding that the distribution of quantile regression estimator becomes standard normal in the mildly integrated case.
The rejection rates of both the switching-FM test and the $t^w$ test are much closer to their target five percent level. The accurate rejection rates for $c=-200$ also address the critique of Phillips14. In both tests, we do observe a slight over-sizing in the outermost deciles when $T=800$. However, this improves when the sample size increases to $T=1600$. Overall, the size performance of the two tests is similar. The only discernible differences are that the $t^w$ test performs somewhat better in the outermost quantiles when $c\leq-5$, while the switching-FM test is slightly less conservative when $c=-25$.
In Table (ref), we examine the size of the switching-FM test when $y_t=e_t$ and $e_t$ has conditional heteroskedasticity, leverage, and fat-tails. We do not include the $t^w$ test in this table since it is not designed for the case when $\psi_\tau(u_{t\tau})$ has autocorrelation. When $e_t$ follows a GARCH process and correlates with $v_t$, the outer population quantiles of $y_t$ can be predicted using $x_{t-1}$. This places us outside of the null hypothesis. Therefore, we consider the following two models; (A) $e_t$ follows a GJR-GARCH(1,1)-t($\nu$) process that is independent of $v_t$, and (B) $e_t$ has fat tails, is correlated with $v_t$, but has no conditional heteroskedasticity. In model A, we use the estimated parameters of a GJR-GARCH(1,1)-t($\nu$) model fitted to the monthly stock returns from our empirical application and generate $e_t$ as $e_t = \sigma_t\varepsilon_{t}$ where
and $(\varepsilon_{t},v_t)$ is drawn from mutually independent and i.i.d.\ $t$-distributions with $\nu$ degrees of freedom. In model B, $(v_t,e_t)$ are jointly drawn from an i.i.d.\ multivariate $t$-distribution with degrees of freedom $\nu$ and correlation $\delta=-0.95$. All innovations are rescaled to have unit variances.
Table (ref) has two panels. Panels A and B report the results with models A and B, respectively. Using our empirical returns, we estimate $8.64$ degrees of freedom in our GJR-GARCH-$t$ residual. Rounding this value down, we set $\nu=8$ in Panel A. In Panel B, we use $\nu=3$, which rounds down our estimate of $3.67$ when fitting a $t$-distribution without GJR-GARCH to our empirical returns. In both panels of Table (ref), the size of the switching-FM test is close to the i.i.d.\ case when $0.3 \leq \tau \leq 0.7$. In Panel A, there is some finite sample size distortion in the outer deciles, but the over-rejection is not severe even for $T=800$ and improves for $T=1600$. Therefore, the switching-FM test performs satisfactorily for a realistically calibrated GJR-GARCH-$t$ model that allows for volatility, leverage, and fat-tails. The finite sample size distortion in the tail is a bit stronger in Panel B, which is not surprising given the very fat tails of Panel B. Nonetheless, this size distortion also improves considerably when increasing the sample size from $800$ to $1600$.
In Table (ref), we examine the finite sample power of the switching-FM test and compare its performance to that of the $t^w$ test under the traditional linear alternative $y_t = \gamma_1 x_{t-1} + e_t$ with Gaussian errors as in Table (ref). We consider the local alternative $\gamma_1 =\gamma_1^*/T$, analyzed in Corollary (ref), for $\gamma_1^* \in \{5, 10, 15, 20, 25\}$ using $T=800$ (left-side panels) and $T=1600$ (right-side panels). The value of $\tau$ is varied across the rows. To save space, we show only the median and outer deciles. The unshaded columns (2 and 8) provide null rejection rates ($\gamma_1^*=0$), while the shaded columns (Columns 3--7 and 9--13) provide rejection rates under the alternative hypothesis that $\gamma_1^*>0$. The top, middle, and bottom panels show results for three different values of the local-to-unity parameter: $c=-5$ (Panels A--B), $c=-10$ (Panels C--D), and $c = -25$ (Panels E--F). We show only results for $\delta=-0.95$ since overall power comparisons for $\delta=-0.50$ are similar. As can be seen from Table (ref), both tests perform well. Their power increases reliably as $\gamma^*$ increases. Their power remains approximately constant as $T$ increases and $\gamma^*/T$ shrinks, demonstrating power against the local alternative. While the $t^w$ test has a modest size advantage in the tails when $c=-5$, overall, the FM-switching test has higher power across all six panels, with substantial differences in some cases.
In Table (ref), we next compare power under an alternative hypothesis for which there is predictability in the tails and shoulders but no median predictability. We again test $H_0:\gamma(\tau)\leq 0$ against $H_A:\gamma(\tau)>0$ using both the switching-FM and $t^w$ tests with $y_t$ generated according to (ref). To exclude median predictability we set $\gamma_1=0$, but allow for predictability at other quantiles by setting $\kappa=0.25$, $b(e_t) = be_t$, and $\zeta = 25$. We vary the value of $b$ across columns. Under this specification, $\gamma(\tau)=0$ holds for either $\tau=0.5$ or $b=0$, whereas $\gamma(\tau)>0$ holds only when both $\tau>0.5$ and $b>0$. Thus, the null rejection rates are shown in the unshaded regions corresponding to the union of $b=0$ and $\tau=0.5$ of each sub-panel. Finite sample power is shown in the shaded regions for which $\tau>0.5$ and $b>0$. Since we employ one-sided tests, we do not show results for $\tau<0.5$, for which $\gamma(\tau)<0$.
Not surprisingly, the size results in column 2 for $b=0$ are similar to those in the previous tables. The null rejection rates in the rows associated with $\tau=0.5$ when $b>0$ are new to this table. In this conditionally heteroskedastic case, increases in $x_{t-1}$ are associated with increased residual variance. An example of $x_{t-1}$ would be a volatility predictor such as the realized variance. Encouragingly, the size of both tests remains stable across the columns. In this model, we move away from the null hypothesis by increasing either $b$ or $\tau$. Reassuringly, the power of both the tests improves with an increase in either parameter. Comparing the left and right panels, we also see that the power of both tests increases when the sample size increases, which corroborates the consistency of the switching-FM test shown in Proposition (ref). Lastly, we note that the power of the switching-FM test again exceeds that of the $t^w$ test.
Finally, in Table (ref), we simulate $y_t$ from the same random coefficient model employed by cai23joe. Specifically, we generate $x_t$ from the autoregressive model in (ref) and generate $y_t$ from the random coefficient model:
where the $(v_{t},e_{t})$ are i.i.d.\ innovations from a bivariate normal distribution with means equal to zero, unit variances, and correlation $\delta$. The model implies a value of $\gamma_1(\tau)= T^{-1}\gamma^*\left(Q_{e_t}(\tau)+3\right)$ in (ref) (see cai23joe). Under this alternative, for $\gamma^*>0$, an increase in $x_{t-1}$ can impact all the quantiles of $y_t$, but it has a larger positive impact on the upper quantiles of $y_t$ than it does on the lower quantiles. Indeed, the results in Table (ref) confirm that the power of both tests is strongest at the ninth decile and weakest at the first decile. However, in all cases, the power of the tests increases with both $\gamma^*$ and with sample size. The switching-FM test is generally more powerful than the $t^w$ test, while previous results in cai23joe show that the $t^w$ test is more powerful than the IVXQR test.
To summarize our Monte Carlo results, standard predictive quantile regression $t$-tests can suffer severe over-rejection when the predictor is both persistent and endogenous. The switching-FM predictive quantile test proposed here successfully solves this problem. Furthermore, it also provides good size when the predictor is far less persistent, thereby addressing the critique of Phillips14. Overall, the size results for the switching-FM test are quite good, matching those of the $t^w$ test at all but the outermost quantiles. Our simulations further confirm the robustness of the switching-FM test to conditional heteroskedasticity and fat-tails. Finally, the switching-FM test is found to have better power than the $t^w$ test in both the linear alternative and in two different random coefficient models. Since cai23joe have previously shown the $t^w$ test to have higher power than the IVXQR test, we find these results promising. On the other hand, both the $t^w$ and IVXQR tests are readily extended to multiple predictors, whereas the switching-FM test is applicable for only a single predictor.
We apply the switching-FM predictive method developed above to test for predictability at different points in the stock return distribution. We employ monthly data from Goyal:Welch:2008 updated to 2015 and focus on the valuation based predictors for which the predictive regression problem is most pertinent.\footnote{While some of the other predictors employed by Goyal:Welch:2008 are also persistent, estimates of $\delta$ generally indicated little endogeneity, implying little distortion from the standard quantile tests. } Namely, we separately employ lagged values of the dividend price ratio ($dp_t$), the earnings price ratio ($ep_t$), and the book-to-market ratio ($bm_t$) as univariate predictors. We test their ability to predict the center, shoulders, and tails of the return distribution, using value weighted monthly excess returns including dividends on the S&P 500 from the Center for Research on Security Prices (CRSP). The data runs from January 1926 until December 2015.\footnote{We thank Amit Goyal for use of his publicly available data and for answers to several queries, including confirmation that the CRSP monthly return series with dividends is unavailable prior to 1926.}
Compared to the vast literature on mean prediction, there have been relatively few applications of predictive quantile regression. Cenesizoglu:Timmermann:08 employed predictive quantile methods with data on 16 predictors from Goyal:Welch:2008 ending in 2005. However, they used standard predictive quantile tests without addressing their size distortion. In MaynardShimotsuWang2011, we revisited these results using an earlier version of our Bonferroni test. Subsequent studies have employed the IVXQR Lee2016,FanLee19joe, the maximized Monte Carlo test GungorLugo2019, and a test based on an auxiliary regressor cai23joe.
In Table (ref), we first provide preliminary indications of the persistence and endogeneity of the three valuation predictors. In Row 2, we provide the $t$-statistic for the GLS-ADF unit root test of Elliott&Rothenberg&Stock96 using BIC to select the lag length. At the 5% significance level, we can reject a unit root in the earnings price ratio but fail to reject for the other two predictors. Under the assumption of a local-to-unity model for the predictors, we next construct a 95% confidence interval ($c_L^{0.95},c_U^{0.95}$) for the local-to-unity parameter by inverting the GLS-ADF $t$-test following the approach of stock91. The lower bounds $c_L^{0.95}$ (Row 4) range between $-10$ and $-20$, confirming that all three variables have near unit roots. The upper bound $c_U^{0.95}$ (Row 5) for the earnings price ratio is only very slightly below zero, whereas it just slightly exceeds zero for the other two predictors. The values of $\phi$ implied by the lower and upper bounds on the confidence region for $c$ are provided in Rows 6--7. Overall, these results confirm that all three predictors are well-modeled as near-unit root processes.
The final row of Table (ref) provides the sample correlation coefficient between $\widehat{\varepsilon}_{t}$ and $\widehat{u}_{t}$, where $\widehat{\varepsilon}_{t}$ is the residual from the ADF regression on $x_t$, using BIC to select lag-lengths, and $\widehat{u}_{t}$ is the residual from regressing the stock return on the predictor. The estimated residual correlations are large and negative for all three predictors. This combination of persistence and endogeneity suggests nontrivial size distortion in standard quantile predictive regression.
Table (ref) provides the empirical results from the switching-FM predictive quantile regression tests. To assess the ability of the valuation predictors to predict at different points in the return distribution, including the center, shoulders, and tails, we conduct the test at each of the nine deciles shown in the top row. The three panels display the test results for the log dividend price ratio, earnings price ratio, and book-to-market ratio.
The first two rows of each panel in Table (ref) show the standard quantile regression slope coefficient and its non-HAC $t$-statistic, without bias or size correction. As discussed earlier, this standard, non-HAC $t$-statistic suffers from two problems: serial correlation in $\psi_\tau(u_{t\tau})$ and a nonstandard asymptotic distribution. The HAC version of the standard $t$-statistic in the third row addresses the first problem, but its distribution is still nonstandard. Row 4 of each panel reports estimates of the long-run residual correlation, $\delta_{\tau}=\omega_{\psi v}/(\omega_{\psi}\omega_{v})$, that are large enough to induce substantial size distortion in the standard HAC quantile $t$-test. Rows 5--6 of each panel of Table (ref) report the first-stage confidence interval on $c$ computed using the adjusted significance levels $\underline{\alpha}_1$ and $\overline{\alpha}_1$ in Table (ref) corresponding to $\widehat{\delta}_\tau$.\footnote{Since $\underline{\alpha}_1$ and $\overline{\alpha}_1$ are adjusted upwards (see Section (ref)), they are tighter than the 95% confidence intervals $(c_L^{0.95},c_U^{0.95})$ from Table (ref). They also depend on the quantile level $\tau$.} For this dataset, all of our first-stage confidence intervals on $c$ lie within the near unit root range.
The final two rows report the Bonferroni confidence interval for the slope coefficient $\gamma_1(\tau)$ obtained by inverting the switching-FM test. When the lower bound is positive, i.e., $\underline{\gamma}_1(\tau)>0$, the switching-FM test rejects the null hypothesis of no predictability against $H_A: \gamma_1(\tau)>0$ at the 5% significance level. Similarly, an upper bound below zero, $\overline{\gamma}_1(\tau)<0$, implies a left-sided rejection at the 5% significance level. The cases for which either occurs are marked in bold.
The test results for the dividend-price ratio show an interesting pattern across quantiles. Goyal:Welch:2008 have argued that in-sample mean-predictability from the dividend price ratio is heavily reliant on observations from the oil crisis period of the early 1970s and disappears in later samples. Similarly, our switching-FM predictive quantile test is unable to reject the null hypothesis of no median predictability using the dividend price ratio (top panel). Given the robustness of quantile regression, this lends additional support to Goyal:Welch:2008's earlier results.
It would nonetheless be premature to conclude that the dividend-price ratio lacks useful predictive content for returns. In fact, the switching-FM test shows it to be predictive for the upper shoulder of the return distribution. An increase in the dividend price ratio corresponds to a larger right shoulder and tail for the return distribution. Since the left shoulder is insignificant, the overall predictive pattern differs from that of a pure increase in volatility but is in line with the informal notion that a low market valuation (low price relative to dividend) may set the stage for future market rallies. For example, valuations may be low during a “bear” market, and the right shoulder of the distribution may reflect the possibility of a strong market recovery.
Our results also underline the importance of robustifying inference to both the serial correlation of the quantile-regression-induced residuals and the size distortion resulting from the persistence and endogeneity of the predictors. The standard non-robust $t$-statistic in the second row naively indicates predictability in the tails and/or shoulders using all three predictors. However, after applying HAC standard errors in row 3, the significance is lost in several cases, including the left-tail for the dividend price ratio and the right-tail for the book-to-market ratio. Finally, after using the switching-FM test, only the right shoulder of the dividend price ratios remains significant. Consequently, using a standard predictive quantile test, even with HAC standard errors, can greatly exaggerate the evidence of quantile predictability. Thus, the use of both robust standard errors and proper inference procedures is essential in empirical applications involving predictive quantile regression.
This paper develops inference in predictive quantile regressions with a nearly nonstationary regressor. We derive the limiting distributions of the quantile regression coefficient and its corresponding HAC $t$-statistic under a local-to-unity specification for the predictor. The asymptotic analysis suggests size distortion using standard tests of quantile predictability. Our simulations indicate that when the predictor is both persistent and endogenous, the size distortion can be nearly as serious as that of predictive (mean) regression. MaynardShimotsuWang2011 was the first to identify these issues. Lee2016 has subsequently generalized these results to the mildly integrated and explosive cases.
One of the challenges to correcting inference in predictive regression is the inability to consistently estimate the local-to-unity parameter. A popular solution in the mean regression case has been the use of a Bonferroni bound in conjunction with a first-stage bound on this parameter. While this works well in practice when the predictor has a near unit root, Phillips14 has recently shown that it becomes invalid when the predictor is stationary. We, therefore, propose a switching-fully modified (FM) predictive quantile regression test that uses a Bonferroni bound over an FM style bias-corrected quantile regression estimator when the predictor has a near unit root and switches to a standard (quantile) test with slightly conservative critical values when the predictor is stationary.
Our simulations indicate that this method works well in practice over a wide range of values for the largest autoregressive root, including values slightly below one (near unit root) and far below one (stationary). They also demonstrate that while both our test and the $t^w$ predictive quantile test of cai23joe perform well in the finite sample, the switching-FM test can compare favorably in terms of power while maintaining roughly similar finite sample size. Previous results by cai23joe show their $t^w$ test to have better size and power than IVXQR. On the other hand, both the $t^w$ and IVXQR tests have the practical advantage of easily generalizing to multivariate predictors and the $t^w$ test has a modest size advantage at tail quantiles.
We test the predictability of three heavily employed valuation predictors, the dividend price ratio, earnings price ratio, and book-to-market ratio, for monthly returns on the S&P 500. Our data spans 1927-2015, which includes the financial crisis and its aftermath, a particularly interesting period when considering the quantiles of the return distribution. We find no significant evidence of predictability in the center of the distribution using even the standard predictive quantile test, and, for two of our three predictors, findings of predictability at other quantiles from the standard test are overturned by our switching-FM test. Nonetheless, we find significant evidence that the dividend price ratio is predictive for the right shoulder of the return distribution.
Extensions of the switching-FM approach to the case of multiple predictors and joint tests could provide an interesting but non-trivial direction for future research. The design of a suitable switching rule with multiple predictors could prove challenging when some regressors are persistent and others are not. The practical implementation of the Bonferroni bound, especially the adjustments to the first-stage confidence interval needed to keep it from being overly conservative, could also be considerately complicated in even a bivariate case. In view of our empirical finding of in-sample predictability at certain non-central quantiles, the development of tests for the out-of-sample predictive power of quantile prediction with persistent and endogenous regressors could provide a second promising direction for future research.