Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,220 characters · 19 sections · 68 citation commands
Testing for Stationary or Persistent Coefficient Randomness in Predictive Regressions
\baselineskip= 6mm
In the applied finance and economics literature, stock return predictability has been widely discussed. Predictability analysis often involves regressing stock returns on the lag of another financial/economic variable and investigating whether the coefficient on the variable is significantly different from zero (a significantly non-zero coefficient amounts to evidence for predictability). Despite a large body of literature on this topic, researchers have reported mixed evidence of predictability. Recently, several researchers have suspected this is due to the time-varying nature of predictability (i.e., the coefficient on the predictor may be nonzero in some periods of time but may be zero in other periods) and tried to estimate its time path guidolinTimeVaryingStock2013, devpuraStockReturnPredictability2018, rahmanPredictivePowerDividend2019, hammamiUnderstandingTimevaryingShorthorizon2020. In such studies, the coefficient at each time point is estimated in a rolling-window fashion, that is, by using data over limited sample periods, and authors plot the time path of its estimate, assessing whether the coefficient is time-varying and identifying the period when predictability exists (if any).
The authors, however, conduct this analysis without statistically confirming the assumption that the coefficient on the predictor is time-varying. They often argue it is time-varying just because its estimated trajectory is seemingly not constant over time. However, the coefficient estimates can be seemingly time-varying even if the coefficient is, in fact, constant over the full sample. This is because the coefficient at each time point is estimated by using a subsample, which tends to be small rahmanPredictivePowerDividend2019, and thus the standard errors of the coefficient estimates tend to be large. This implies that even if the coefficient is actually constant over time, the time variations of its estimates can be so large that only a graphical inspection erroneously supports the hypothesis of time-varying predictability.
A more decent approach is to statistically test for the constancy of the coefficient before estimating it or its trajectory. If the null of constant coefficient is rejected, the rolling-window estimation may be applied to identify the period when predictability exists (otherwise, the usual method based on the full sample should be used to test whether the coefficient is zero).
Testing procedures for coefficient constancy in predictive regressions have recently been discussed by georgievTestingParameterInstability2018. They propose several bootstrap-based tests under the assumption that the predictor is highly persistent, and find that a test based on nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test performs well when the coefficient on the predictor experiences an abrupt break or follows a near unit root AR(1) process. This result is due to the fact that nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test is designed to be the locally most powerful against martingale coefficient processes, which accommodate random walk and processes with abrupt breaks. However, in empirical applications, the coefficient process may not be a martingale, and it is unclear whether the LM test retains good power in such a case.
Models with non-martingale random coefficients are widely seen in the literature. swamyLinearPredictionEstimation1980 propose regression models with stationary random coefficients, gives them a justification from statistical and economic perspectives, and discuss prediction and estimation methods. linTestingParameterConstancy1999 derive the LM test for coefficient constancy against the alternative of stationary random coefficient, and fuSpecificationTestsTimevarying2023 propose a test to distinguish between stationary and persistent variations in parameters. Stationary-coefficient models are also employed in empirical studies. For instance, chiangForecastingTreasuryBill1991 estimate a model where spot interest rate is regressed on forward interest rate, and coefficients are assumed to be possibly stationary stochastic processes. leeRevisitingDemandMoney2013 estimate the money demand function with stationary random coefficients.
When stationary (or $\mathrm{I}(0)$) random coefficient is assumed as the alternative hypothesis, we can naturally test the null of constant coefficient by using a Wald-type test based on a linear model derived from the original predictive regression. This testing procedure is essentially the same as the one employed by nishiTestingCoefficientRandomness2023, who considers AR(1) models with i.i.d. random coefficient and shows that the Wald test performs well in such a case. In this article, we show that the Wald test performs better than Nyblom's LM test in predictive regressions with stationary (rather than i.i.d.) random coefficient. The Wald test is, however, less powerful than the LM test when the coefficient process follows a near unit root (or $\mathrm{I}(1)$) process.\footnote{One can also consider intermediate cases where the coefficient process is of $\mathrm{I}(d)$, where $d\in (0,1)$, but we do not consider such cases and will say that the coefficient process is stationary if it is $\mathrm{I}(0)$.}
The power comparison between the LM and Wald tests is made under the sequence of local alternatives. As shown later, the appropriate localizing rate of the alternative hypotheses depends on the persistence of the random coefficient and which test (the LM or Wald) is considered. When the random coefficient is $\mathrm{I}(0)$, the LM test has a nontrivial asymptotic power against the sequence of local alternatives approaching the null of zero randomness at the rate of $T^{-1/2}$, while the appropriate rate of the local alternatives for the Wald test is $T^{-3/4}$, where $T$ is the sample size. This implies that the Wald test can detect smaller variations in the random coefficient than the LM test when the random coefficient is $\mathrm{I}(0)$. In contrast, when the random coefficient is $\mathrm{I}(1)$, the LM test has nontrivial asymptotic power against the sequence of local alternatives of order $T^{-3/2}$, but the Wald test against that of order $T^{-5/4}$. Thus, the LM test detects smaller coefficient variations than the Wald test when the random coefficient is $\mathrm{I}(1)$.
Because one test does not dominate the other in terms of power, we consider combining the LM and Wald tests to obtain new tests that have good power properties irrespective of the persistence of the random coefficient. Our proposal is simply to take the sum and product of the two statistics. We will investigate the power properties of four tests in total.
The asymptotic null distributions of the Wald and LM test statistics are found to depend on nuisance parameters under our setting (and so do the combinations of these two tests). To implement these tests without the knowledge about the nuisance parameters, we base them on subsampling, in which critical values are approximated by using a number of test statistics calculated from subsamples andrewsHybridSizeCorrectedSubsampling2009.\footnote{Another way to circumvent the nuisance-parameter problem is to use bootstrap as georgievTestingParameterInstability2018 do. Because different approaches require different mathematical assumptions and discussions, we focus on subsampling. The reason why we use subsampling instead of bootstrap is explained in Remark 3 in Section 2.} We discuss the validity of the subsampling approach in our setting by some theoretical and numerical analyses.
Finally, we revisit devpuraStockReturnPredictability2018's (devpuraStockReturnPredictability2018) study. They determine whether predictability is time-varying based on a sequence of p-values associated with the rolling-window coefficient estimates. They do this, however, without any adjustment for multiple testing, resulting in uncontrolled size. In this article, we investigate the coefficient constancy by applying the tests considered in this article. Our result mostly reverses devpuraStockReturnPredictability2018's (devpuraStockReturnPredictability2018) conclusion.
The remainder of this article is structured as follows. In Section 2, we define the Wald test statistic assuming that the random coefficient is an $\mathrm{I}(0)$ process. We analyze and compare the asymptotic behavior of the Wald and LM tests under this setting. In Section 3, these two tests are compared under the assumption that the coefficient process is a near unit root, and two combination tests of the Wald and LM tests are proposed. Section 4 proposes to base these tests on subsampling. In Section 5, the model is extended to the heteroskedastic case, and we define modified test statistics. Section 6 conducts Monte Carlo simulations to investigate finite-sample behavior of the tests, and Section 7 gives a real data application. Section 8 concludes the article. All mathematical proofs are relegated to Appendix A.
In this section, we consider a predictive regression where the coefficient on the persistent predictor is expressed by the sum of a constant $\beta$ and stationary random variable $s_{\beta,t}$:
where $\mathrm{E}[s_{\beta,t}]=0$, $\mathrm{Var}(s_{\beta,t})=1$, and $s_{x,0}=O_p(1)$. Here, the predictor, $x_{t}$, is modeled as a persistent local-to-unity AR(1) process, which is a commonly employed assumption in the literature campbellEfficientTestsStock2006. Note that $\beta$ and $\omega_\beta^2$ denote the mean and variance of the coefficient on $x_{t-1}$, respectively. Therefore, coefficient constancy is equivalent to $\omega_\beta^2=0$ in our setting. To simplify the discussion in what follows, we set $\alpha_y=\alpha_x=s_{x,0}=0$, resulting in
with $x_t = (1+c_xT^{-1})x_{t-1} + \varepsilon_{x,t} \ (x_0=0)$.
We suppose the following assumptions hold.
In Assumption (ref)(i), the disturbance $\varepsilon_{y,t}$ of predictive regression (ref) is correlated with the innovation $\varepsilon_{x,t}$ that drives $x_{t}$, with correlation coefficient $\gamma$. Assumption (ref)(ii) specifies that the innovation process is stationary and weakly serially dependent, while all the innovations except $s_{\beta,t}$ are restricted to be serially uncorrelated by Assumption (ref)(iv). Assumption (ref)(iv) also implies conditional homoskedasticity of $\varepsilon_{y,t}$, i.e., $\mathrm{E}[\varepsilon_{y,t}^2|\mathcal{F}_{t-1}] = \sigma_{y}^2$. The case where $\varepsilon_{y,t}$ is conditionally heteroskedastic is considered in Section 5.
Define $\sigma_\eta^2 \coloneqq \mathrm{E}[\eta_t^2]$ and $\sigma_{\beta,lr}^2 \coloneqq \mathrm{E}[s_{\beta,1}^2] + 2\sum_{i=1}^{\infty}\mathrm{E}[s_{\beta,1}s_{\beta,1+i}]$. Define also the partial sum process $(W_{x,T}, W_{y,T}, W_{\eta,T},W_{\beta,lr,T})'$ on $[0,1]$ by $W_{i,T}(r) \coloneqq T^{-1/2}\sigma_i^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}\varepsilon_{i,t}$, $i=x,y$, $W_{\eta,T}(r) \coloneqq T^{-1/2}\sigma_\eta^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}\eta_t$ and $W_{\beta,lr,T}(r) \coloneqq T^{-1/2}\sigma_{\beta,lr}^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}s_{\beta,t}$. Then, it follows from the functional central limit theorem (FCLT) that under Assumption (ref), as $T \to \infty$,
in the Skorokhod space $D_4[0,1]$, where $(W_{x},W_{y},W_{\eta},W_{\beta,lr})'$ is a vector Brownian motion. Furthermore, by the standard continuous mapping argument, we also have $J_{x,T}(r) \coloneqq T^{-1/2}\sigma_x^{-1}x_{\lfloor Tr \rfloor} \Rightarrow J_x(r)$, where $J_x$ is an Ornstein-Uhlenbeck (OU) process satisfying $dJ_x(r) = c_xJ_x(r)dr + dW_x(r)$.
In some analysis, we rely on the following assumption instead of Assumption (ref):
In the sequel, we let $\sigma_{y\beta} \coloneqq \mathrm{Cov}(\varepsilon_{y,t},s_{\beta,t})$.
Our aim here is to construct a Wald-type test for coefficient constancy in (ref), or $\omega_\beta^2=0$. To make our point clear, suppose for now $\beta$ is known. Define $z_t(\beta) \coloneqq y_t-\beta x_{t-1} = \omega_\beta s_{\beta,t}x_{t-1} + \varepsilon_{y,t}$, and obtain
where $\delta \coloneqq 2\omega_\beta\sigma_{y\beta}$ and $\xi_t \coloneqq (\varepsilon_{y,t}^2-\sigma_y^2) + 2\omega_\beta x_{t-1}(s_{\beta,t}\varepsilon_{y,t} - \sigma_{y\beta}) + \omega_\beta^2 x_{t-1}^2(s_{\beta,t}^2 - 1)$. Viewing (ref) as a linear model with $\xi_t$ playing the role of the disturbance, nishiTestingCoefficientRandomness2023 proposes to test $\mathrm{H_0}:\omega_\beta^2=0$ by the Wald test for $(\delta, \omega_\beta^2) = (0,0)$, since $\omega_\beta^2=0$ if and only if $(\delta,\omega_\beta^2) = (0,0)$. Moreover, the LM test of linTestingParameterConstancy1999, which is designed to detect stationary random coefficient, can be derived as the (squared) t-test for $\omega_\beta^2=0$ under the restriction $\delta=0$ (see Appendix B). Therefore, it is natural in our setting to employ model (ref) as the basis for constructing a test statistic. To define the Wald test explicitly, we first regress the variables on a constant, obtaining
where $\widetilde{z_t^2}(\beta) \coloneqq z_t^2(\beta) - T^{-1}\sum_{t=1}^{T}z_t^2(\beta)$, $\tilde{x}_{1,t-1} \coloneqq x_{t-1} - T^{-1}\sum_{t=1}^Tx_{t-1}$, $\tilde{x}_{2,t-1} \coloneqq x_{t-1}^2 - T^{-1}\sum_{t=1}^{T}x_{t-1}^2$, and $\tilde{\xi}_t \coloneqq \xi_t - T^{-1}\sum_{t=1}^T\xi_t$. Letting $Z_2(\beta) \coloneqq (\widetilde{z_1^2}(\beta),\ldots,\widetilde{z_T^2}(\beta))'$, $X_1 \coloneqq (\tilde{x}_{1,0},\ldots,\tilde{x}_{1,T-1})'$, $X_2 \coloneqq (\tilde{x}_{2,0},\ldots,\tilde{x}_{2,T-1})'$, and $\Xi \coloneqq (\tilde{\xi}_1,\ldots,\tilde{\xi}_t)'$, we arrive at
where $X\coloneqq (X_1,X_2)$ and $\theta\coloneqq (\delta,\omega_\beta^2)'$. Then, the Wald test statistic is
where $\hat{\theta}_T(\beta)$ and $\hat{\sigma}_{\tilde{\xi}}^2(\beta)$ are the OLS coefficient and variance estimators from (ref), respectively: $\hat{\theta}_T(\beta) = (X'X)^{-1}X'Z_2(\beta)$ and $\hat{\sigma}_{\tilde{\xi}}^2(\beta) = T^{-1}Z_2(\beta)'M_XZ_2(\beta)$ with $M_X \coloneqq I - X(X'X)^{-1}X'$.
The Wald test defined above uses the true value of $\beta$. Because $\beta$ is usually unknown in applications, we use $\mathrm{W}_T(\hat{\beta})$, where $\hat{\beta}$ is the OLS estimator of $\beta$: $\hat{\beta} =\sum_{t=1}^{T}y_t x_{t-1}/\sum_{t=1}^{T}x_{t-1}^2$. The use of $\hat{\beta}$ makes the Wald test invariant to the value of $\beta$ and feasible without any knowledge about $\beta$ (in particular, about whether it is zero or not).
We also consider nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test, which is defined by
where $\hat{\sigma}_T^2 \coloneqq T^{-1}\sum_{t=1}^{T}z_t^2(\hat{\beta})$.\footnote{georgievTestingParameterInstability2018 also investigate the performance of the $\mathrm{Sup}F$ test proposed by andrewsTestsParameterInstability1993. Although this test is known to be comparable in terms of power with Nyblom's LM test when parameters are $\mathrm{I}(1)$, it is not designed to detect (non)stationary randomness in parameters. Therefore, Nyblom's LM test seems more appropriate as the test to be compared with our Wald test, as argued by hansenTestsParameterInstability2002.}
There are two points worth mentioning in Theorem (ref). First, the asymptotic distributions of the Wald and LM test statistics are derived under different local alternatives. The Wald test statistic has the limiting alternative distribution when $\omega_\beta=gT^{-3/4}$, while the LM test does when $\omega_\beta=gT^{-1/2}$. Because the former specification captures a sequence of local alternatives closer to the null of $\omega_\beta=0$, it implies the Wald test can detect smaller variations in random coefficient than the LM test, as long as the random coefficient is $\mathrm{I}(0)$. This is intuitively because the Wald test is constructed under the assumption that the random coefficient follows a stationary process, while the derivation of the LM test is based on the assumption that the coefficient process is nonstationary. This finding is in line with the results of linTestingParameterConstancy1999. They verify through simulation that their test, which is designed to detect stationary random coefficient, outperforms nyblomTestingConstancyParameters1989's LM test when the random coefficient is a stationary AR process. Theorem (ref) gives a theoretical explanation for their simulation results.\footnote{linTestingParameterConstancy1999's (linTestingParameterConstancy1999) LM test has the asymptotic alternative distribution when $\omega_\beta=gT^{-3/4}$, as our Wald test does. See Appendix B for details.}
The second point is that the asymptotic null distributions ($g=0$) of the Wald and LM test statistics are dependent on nuisance parameters. Specifically, the Wald test depends on the correlation between $W_x$ and $W_\eta$, and the LM test on the correlation between $W_x$ and $W_y$. Moreover, these tests depend on $c_x$, the persistence parameter of $x_t$. Therefore, information on the nuisance parameters is required to perform these tests. In Section 4, we propose to base them on subsampling to solve this problem.
In this section, we consider a predictive regression model where the random part of the coefficient is near $\mathrm{I}(1)$:
where $x_t = (1+c_xT^{-1})x_{t-1} + \varepsilon_{x,t} \ (x_0=0)$. We impose the following condition on $\varepsilon_{\beta,t}$.
georgievTestingParameterInstability2018 prove that, under model (ref), the LM test statistic has the asymptotic alternative distribution when $\omega_\beta = gT^{-3/2}$. The next theorem establishes that the Wald test statistic has the asymptotic alternative distribution if $\omega_\beta = gT^{-5/4}$. To state this precisely, we need the following notation: let $J_\beta(r)$ be an OU process satisfying $dJ_\beta(r) = c_\beta J_\beta(r)dr + dW_\beta(r)$, where $W_\beta(\cdot)$ is the standard Brownian motion to which $W_{\beta,T}(\cdot) \coloneqq T^{-1/2}\sum_{t=1}^{\lfloor T\cdot \rfloor}\varepsilon_{\beta,t}$ weakly converges.
Based on Theorems (ref) and (ref) and the result of georgievTestingParameterInstability2018, we observe that the convergence rate of $\omega_\beta$ under which the tests have nontrivial asymptotic power depends on the nature of random coefficient $s_{\beta,t}$. We summarize this in Table (ref). When the random part $s_{\beta,t}$ of the coefficient is $\mathrm{I}(0)$, the Wald test has a nontrivial power for $\omega_\beta=gT^{-3/4}$, which describes a sequence of local alternatives closer to the null of $\omega_\beta=0$ than $\omega_\beta=gT^{-1/2}$ for the LM case. However, in the case of $\mathrm{I}(1)$ $s_{\beta,t}$, the convergence rate of $\omega_\beta$ associated with the nontrivial power of the LM test is $gT^{-3/2}$, which converges to the null of $\omega_\beta=0$ faster than $gT^{-5/4}$ for the Wald test. This implies that the Wald test dominates the LM test in terms of power when $s_{\beta,t}$ is $\mathrm{I}(0)$, while the LM test is more powerful than the Wald test when $s_{\beta,t}$ is $\mathrm{I}(1)$.
Given the above observation, one ideal testing procedure is to use the Wald test when $\mathrm{I}(0)$ random coefficient is suspected and use the LM test when $\mathrm{I}(1)$ random coefficient is probable. However, the persistence of the coefficient (if it is random) is often unknown, and thus it will be difficult to rely on such a procedure in practice. To detect coefficient randomness efficiently whether it is $\mathrm{I}(0)$ or $\mathrm{I}(1)$, we propose two combination tests of the Wald and LM, one being defined as the sum of them and the other as the product. Specifically, we consider the following test statistics:
We shall call $\mathrm{S}_T$ the Sum test statistic and $\mathrm{P}_T$ the Product test statistic. The motivation of these two tests is simple. When $s_{\beta,t}$ is $\mathrm{I}(0)$ and $\omega_\beta=gT^{-3/4}$, $\mathrm{W}_T(\hat{\beta})$, $\mathrm{S}_T$ and $\mathrm{P}_T$ weakly converge to alternative distributions, but $\mathrm{LM}_T$ to the null distribution. In contrast, if $s_{\beta,t}$ is $\mathrm{I}(1)$ and $\omega_\beta=gT^{-3/2}$, $\mathrm{LM}_T(\hat{\beta})$, $\mathrm{S}_T$ and $\mathrm{P}_T$ have nontrivial asymptotic power but $\mathrm{W}_T(\hat{\beta})$ does not. In other words, the Sum and Product tests are more influenced by the more powerful of the Wald and LM tests, and thus they are expected to perform well in both the $\mathrm{I}(0)$ and $\mathrm{I}(1)$ $s_{\beta,t}$ cases.\footnote{A more preferable way to combine the LM and Wald tests will be to weight those test statistics based on some information on the persistence of the random coefficient. For example, the weighting function might be a function of the test statistic proposed by fuSpecificationTestsTimevarying2023, which is aimed at distinguishing between stationary and persistent parameters. However, their statistic cannot be directly used in our setting because their analysis is based on the assumption that regressor $x_t$ is weakly stationary.} The finite-sample performances of these tests are examined in Section 6.
As observed in previous sections, tests we have considered are dependent on nuisance parameters. In the literature, the subsampling approach has been a popular solution to nuisance parameter problems romanoSubsamplingIntervalsAutoregressive2001,choiSubsamplingVectorAutoregressive2005,andrewsHybridSizeCorrectedSubsampling2009. The method allows us to approximate the sampling distribution of a statistic by calculating a number of the statistics based on subsets of the full sample. The idea behind this strategy is that, in large samples, the sampling distribution of subsample statistics will be close to that of the full sample statistic, and hence the empirical distribution of a number of subsample statistics can approximate it.
To fix ideas, take the LM test as an example. To implement the subsampling method, one has to set the subsample size $b < T$, the number of data used to calculate each subsample statistic. Then, the first subsample statistic is calculated using the subsample ranging from $t=1$ to $t=b$, the second one is based on $t=2,\ldots,b+1$, and continue until the last statistic based on $t=T-b+1,\ldots,T$ is calculated. Letting $\mathrm{LM}_{j,b}$ denote the subsample LM statistic based on $b$ data ranging from $t=j$, the subsample test statistics can be written as
where $\hat{\beta}_{j,b} \coloneqq \sum_{t=j}^{j+b-1}y_tx_{t-1}/\sum_{t=j}^{j+b-1}x^2_{t-1}$ is the $\beta$ estimate based on $t=j,\ldots,j+b-1$ and $\hat{\sigma}^2_{j,b} \coloneqq b^{-1}\sum_{t=j}^{j+b-1}z^2_t(\hat{\beta}_{j,b})$. If the significance level is $\alpha$, then the critical value (denoted by $c_{\mathrm{LM}}(1-\alpha)$) is defined by the $1-\alpha$ quantile of the empirical distribution function
where $1\{\cdot\}$ is the indicator function. The null is rejected if the full sample test statistic $\mathrm{LM}_T$ exceeds the critical value. The Wald, Sum and Product tests can be based on the subsampling approach in the same way.
When implementing the subsampling method, one has to choose the subsample size $b$. One necessary condition is $1/b + b/T\to 0$ as $T\to \infty$ romanoSubsamplingIntervalsAutoregressive2001. One reasonable way to choose $b$ is to compute the empirical distribution $L_{b,T}$ for a number of $b$ in some set $B$ and select $b$ that leads to the most `appropriate' $L_{b,T}$ in some sense. In the literature, several ways to define $B$ and the appropriateness criterion have been proposed. We here follow jachSubsamplingInferenceMean2012 and set, for some $q \in (0,1)$, $B = \{b_j \coloneqq \lfloor q^jT \rfloor|j=j_{\mathrm{max}},j_{\mathrm{max}}-1,\ldots,j_{\mathrm{min}}\}$, with $j_{\mathrm{min}}<j_{\mathrm{max}} \ (j_{\mathrm{min}},j_{\mathrm{max}} \in \mathbb{N})$, and select $b$ such that $b^{\mathrm{sel}} \coloneqq \operatorname*{\arg\!\min}_{b_j \in B} \ \mathrm{KS}(L_{b_j,T},L_{b_{j+1},T})$, where $\mathrm{KS}(x,y)$ denotes the Kolmogorov-Smirnov (KS) distance between distributions $x$ and $y$.
The idea behind this selection rule is that $L_{b,T}$, viewed as a function of $b$, should vary only slightly in an appropriate range of $b$, and hence values of $b$ leading to small KS distances should be selected. After preliminary experiments, we set $(q,j_{\mathrm{min}},j_{\mathrm{max}}) = (0.85,\lfloor \log_{0.85}(0.27) \rfloor,\lfloor \log_{0.85}(5/T^{0.9}) \rfloor)$ for the LM test and $(q,j_{\mathrm{min}},j_{\mathrm{max}}) = (0.95,\lfloor \log_{0.95}(0.3) \rfloor,\lfloor \log_{0.95}(15/T^{0.9}) \rfloor)$ for the Wald, Sum and Product tests, so that we consider $b$'s up to about 30% of the sample size $T$.
andrewsHybridSizeCorrectedSubsampling2009 consider subsampling-based inference in nonregular models, including regression models with nearly integrated regressors, and provide sufficient conditions for subsampling-based tests to work. Although it will take a complicated analysis to check whether their sufficient conditions hold in our setting, the (in)validity of the subsampling approach can be conjectured by checking the monotonicity in $c_x$ of the critical value of the asymptotic null distributions. Here, we follow Andrews and Guggenberger's argument to discuss the validity of the subsampling approach in our setting.
For the sake of exposition, let $S_T$ denote a generic test statistic calculated using the full sample and $S_{j,b}, \ j =1,\ldots,T-b+1,$ denote its subsampling counterparts, which are functions of predictor $x_t = (1+c_x/T)x_{t-1} + \varepsilon_{x,t}$. Suppose that, under the null hypothesis, $S_T$ has an asymptotic distribution $P_{c_x}$, dependent on $c_x$. The aim of subsampling is to approximate the critical value from $P_{c_x}$ by the empirical distribution of $S_{j,b}$. However, due to the local-to-unity nature of the predictor, subsample statistics $S_{j,b}$ weakly converge to $P_0$, the asymptotic distribution of $S_T$ obtained under $c_x=0$. This is because $x_t = (1+c_x/T)x_{t-1} + \varepsilon_{x,t} = (1+c_x(b/T)/b)x_{t-1} + \varepsilon_{x,t} = (1+o(1)/b)x_{t-1} + \varepsilon_{x,t}$ (since $b/T \to 0$ as $T\to \infty$), so that under the asymptotics of $b\to \infty$ (rather than $T \to \infty$) $S_{j,b}$'s weakly converge to $P_0$. Therefore, the validity of subsampling is determined by the relation between the critical values from $P_{c_x}$ (the weak limit of $S_T$) and $P_0$ (the weak limit of $S_{j,b}$).
Figures (ref) and (ref) depict the 0.95 quantiles of the asymptotic distributions of the LM and Wald statistics, respectively, as a function of $c_x$. The quantiles of the LM statistic are computed for $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t}) = 0,-0.3,-0.6,-0.9$, and the quantiles of the Wald statistic are computed for $\mathrm{Corr}(\varepsilon_{x,t},\eta_t) = 0,-0.3,-0.6,-0.9$.\footnote{Quantiles are insensitive to the sign of the correlations, and thus we consider negative correlations only.} Computations are based on 20000 simulations where $(W_x,J_x,W_y,W_\eta)$ are approximated by normalized sums of 1000 steps. For the LM test, the asymptotic 0.95 quantile is minimized at $c_x=0$, and as $c_x$ tends to $-\infty$, it increases and converges to 0.47, the asymptotic 0.95 quantile obtained when $x_t$ is $\mathrm{I}(0)$ hansenTestingParameterInstability1992, for any value of $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$. Note that the critical value taken from the empirical distribution of the subsample LM statistics asymptotically corresponds to that obtained under $c_x=0$. Therefore, if the true value of $c_x$ is negative, the critical value based on the subsample LM statistics is smaller than the 0.95 quantile of the asymptotic distribution of the full-sample LM statistic for any $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$, resulting in over-rejection. To quantify the degree of the over-rejection, we show in Figure (ref) the asymptotic sizes of the subsampling LM test with significance level 0.05. It indicates that the LM test will be oversized for almost all the values of $c_x$, irrespective of the value of $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$.
To solve over-rejection problems associated with subsampling tests, andrewsHybridSizeCorrectedSubsampling2009 propose a hybrid test, where the referenced critical value is simply the maximum of the subsampling critical value and the critical value obtained under $c_x = -\infty$. In the case of the LM test with 0.05 significance level, the hybrid critical value is $c^*_{\mathrm{LM}}(0.95) \coloneqq \mathrm{max}\{c_{\mathrm{LM}}(0.95), 0.47\}$, where $c_{\mathrm{LM}}(0.95)$ is the subsampling critical value. Since the asymptotic 0.95 quantile of the full sample LM statistic under the null hypothesis is below 0.47 irrespective of the true value of $c_x$, this makes the LM test a conservative one. In what follows, we will use this hybrid critical value for the LM test.
For the Wald test, the asymptotic 0.95 quantile (Figure (ref)) is maximized at $-c_x=0$ irrespective of the value of $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$ and is decreasing in $-c_x$. This implies that the subsampling Wald test will asymptotically have size equal to or less than 0.05 (see Figure (ref)). Strictly speaking, the quantile is not maximized at $-c_x=0$ when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$, and thus the Wald test is slightly oversized in this case. However, we can show that the asymptotic null distribution of the Wald statistic is exactly the chi-square distribution with two degrees of freedom when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$, irrespective of the value of $c_x$. This is because $W_x$ and $W_\eta$ are independent in this case, and thus $J_x$ and $W_\eta$ also are. This implies that the asymptotic null distribution of $W_T(\hat{\beta})$ given in Theorem (ref) (with $g=0$), if conditional on $W_x$, is chi square with two degrees of freedom, and so is it unconditionally. Therefore, the theoretical asymptotic 0.95 quantile of the Wald statistic is 5.99 for any value of $c_x$, and the type 1 error is exactly 0.05 for all $c_x$. The slight discrepancy between the theoretical result and the result observed in Figure (ref) is only due to the approximation error. In summary, the Wald test, if based on subsampling, is expected to be valid in that its size is equal to or below the nominal level.
The analysis for the Sum and Product tests is somewhat complicated because they depend on both $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$ and $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$. Moreover, they also depend on $\mathrm{Corr}(\varepsilon_{y,t},\eta_t)$. To gain some insight, we consider the case where $\varepsilon_{y,t}$ is symmetric, which is satisfied when, e.g., $\varepsilon_{y,t}$ is normally distributed. This yields $\mathrm{Corr}(\varepsilon_{y,t},\eta_t)=0$, and thus we can focus on the effect of $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$ and $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$ on the asymptotic size. We consider $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})=0,-0.3,-0.6,-0.9$ and $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0,-0.3$.\footnote{We do not consider the cases with $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=-0.6,-0.9$ because these values can cause the variance-covariance matrix of $(\varepsilon_{x,t}, \varepsilon_{y,t}, \eta_t)'$ to be an indefinite matrix.} In Figures (ref) and (ref), we plot the asymptotic quantiles and size of the Sum and Product tests as functions of $-c_x$ and $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})=0,-0.3,-0.6,-0.9$, fixing the value of $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$.
In Figure (ref), we plot the asymptotic 0.95 quantiles of the Sum test statistic when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$. It attains the minimum value when $-c_x=0$ and increases as $-c_x$ gets larger, although in a non-monotonic way. This leads to a slight over-rejection (see Figure (ref)). However, we suspect that this result is mostly due to the approximation error for the asymptotic null distribution of the Sum test statistic, as in the case of the Wald test. Indeed, we found that the type 1 error is closer to or below the nominal level if we increase the steps for the approximation of $(W_x,J_x,W_y,W_\eta)$ and the number of replications. When $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=-0.3$ (Figure (ref)), the asymptotic 0.95 quantile is decreasing in $-c_x$ in general and attains the maximum at $-c_x=0$. Therefore, the asymptotic size is equal to or below the nominal 0.05 level.
Figure (ref) plots the asymptotic 0.95 quantiles of the Product statistic when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$. It attains the minimum value when $-c_x=0$ and increases as $-c_x$ gets larger, and thus the Product test is asymptotically oversized (see Figure (ref)). When $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=-0.3$, the same results are obtained (Figures (ref) and (ref)). Therefore, the Product test should not be based on subsampling, even in the special case (where $\varepsilon_{y,t}$ is symmetric).
The finite-sample size and power of the subsampling LM, Wald, and Sum tests will be investigated in Section 6.
In this section, we extend our discussions thus far to the conditionally heteroskedastic case, which is more plausible in many applications. Incorporating the conditional heteroskedasticity requires us to modify the tests considered in the preceding sections.
The heteroskedasticity-robust version of the LM test is derived by hansenTestingParameterInstability1992, which is of the form
Subsampling counterparts are defined similarly.
For the Wald test, the usual modification involving the use of HC variance estimator is insufficient in our setting. The problem is that, when predictive regression error $\varepsilon_{y,t}$ is conditionally heteroskedastic, then $\varepsilon_{y,t}^2$ is serially correlated, which induces serial correlation in $\eta_t = \varepsilon_{y,t}^2 - \sigma_y^2$ and violates Assumption (ref)(iv).
To solve this problem, we follow nishiStochasticLocalModerate2024 to modify the Wald test. To facilitate the discussion, the following regularity condition is imposed on $\varepsilon_{x,t}$, $v_t$ and $\eta_t$.
Note that $\varepsilon_{y,t}$ and $\eta_t$ are stationary and $\alpha$-mixing with the same mixing coefficients as $(\varepsilon_{x,t},v_t)'$ under Assumption (ref)(i). Assumption (ref)(i) allows for weak serial correlation in $\eta_t$, while we keep the no serial correlation assumption for $\varepsilon_{x,t}$ and $v_t$ (and thus $\varepsilon_{y,t}$) in (ii) (cf. Assumption (ref)(iv)). Under this condition, conditional heteroskedasticity in $\varepsilon_{y,t}$ may be permitted. In Assumption (ref)(iii), we assume that $\varepsilon_{x,t}$ and $\eta_t$ have a finite 6th moment. This condition is a technical one to guarantee that Theorem 3.1 of liangWeakConvergenceStochastic2016 holds. This assumption might be weakened if a more sophisticated asymptotic theory is available.
Under Assumption (ref), we have $\Bigl(T^{-1/2}\sigma_{x}^{-1}\sum_{t=1}^{\lfloor T\cdot \rfloor}\varepsilon_{x,t}, T^{-1/2}\sigma_{\eta,lr}^{-1}\sum_{t=1}^{\lfloor T\cdot \rfloor}\eta_t\Bigr)' \Rightarrow \Bigl(W_x(\cdot), W_{\eta,lr}(\cdot)\Bigr)'$, where $(W_x, W_{\eta,lr})'$ is a vector Brownian motion. Furthermore, the one-sided long-run covariance between $\varepsilon_{x,t}$ and $\eta_t$, $\Lambda_{x\eta} \coloneqq \sum_{t=1}^{\infty}E[\varepsilon_{x,1}\eta_{t+1}]$, exists.
Under Assumption (ref), we can show that the asymptotic null distribution of the Wald test statistic becomes
see Appendix A for the proof. In view of this expression, we modify $\mathrm{W}_T(\hat{\beta})$ in the following way:
where $\hat{\sigma}_{\tilde{\xi},lr}^2(\hat{\beta})$ and $\hat{\Lambda}_{x\eta,T}$ are consistent estimators of $\sigma_{\eta,lr}^2$ and $\Lambda_{x\eta}$, respectively. In this article, the long-run (co)variance estimators are constructed nonparametrically using the Parzen kernel, i.e.,
where $\hat{\tilde{\xi}}_t$ is the OLS residual from (ref), and
Here, $l$ denotes the truncation number, for which we set $l=\lfloor 4(T/100)^{1/3}\rfloor$ in this article.
Under the null $\mathrm{H}_0:\omega_\beta=0$, a standard argument will show that these estimators are consistent because $\hat{\beta}-\beta=O_p(T^{-1})$. It will be tedious and much involved to establish the consistency of these estimators under the alternative hypothesis. However, we expect that they are consistent under the sequences of local alternatives considered in this article because the convergence rates of those local alternatives are sufficiently fast that the presence of random coefficient does not affect the consistency of the short-run variance estimator $\hat{\sigma}_{\tilde{\xi}}^2(\hat{\beta})$ (see Lemmas (ref) and (ref)).
The following theorem establishes that $W_T^{*}(\hat{\beta})$ has the same asymptotic null distribution as $W_T(\hat{\beta})$ has under Assumption (ref) (up to the difference between $W_{\eta}$ and $W_{\eta,lr}$).
In this section, Monte Carlo simulations are conducted to investigate the finite-sample performance of the tests. The test statistics considered are $\mathrm{LM}_T^*$, $\mathrm{W}_T^*(\hat{\beta})$ and their sum and product. These tests are based on subsampling with the choice of subsample size explained in Section 4.2. To see whether the heteroskedasticity-robust subsampling tests behave well, we use the data generating process given by (ref) along with
Here, we set $(\alpha_y,\beta,\alpha_x,s_{x,0},s_{\beta,0})=(0,0,0,0,0)$, $\rho_x=0.95$ and $\gamma=-0.8$.\footnote{We calculate the Wald statistic using demeaned $y_t$ and $x_t$, allowing for nonzero location parameters (see Remark (ref)). We also do this when conducting an empirical analysis in Section 7.} For the persistence of $s_{\beta,t}$, we consider $\rho_\beta \in \{0.6,0.8,0.9,0.98\}$. The coefficient variance $\omega_\beta^2$ is taken as $\omega_\beta^2 \in \{0.01,0.05,0.1,0.5\}$. For the driver process, we specify that $v_t$ is GARCH(1,1):
where $(u_{x,t},u_{v,t},u_{\beta,t})' \sim \mathrm{i.i.d.} \ N(0,\mathrm{diag}(1,0.36,1))$, and $\sigma_{v,t}^2=c_0+c_1\sigma_{v,t-1}^2+c_2v_{t-1}^2$. For the GARCH parameters, we consider the following configurations:
All the tests are implemented under 5% significance level.
Empirical sizes are displayed in Table (ref). Because we have found that the Product test is oversized if based on subsampling (see Section 4.3.2), we do not consider the Product test here. For DGP1, the sizes of all the tests are around 0.02-0.04, which indicates that these tests are valid in that their sizes are below the nominal level. The sizes for DGP2 and DGP3 are similar, so the same comment applies.
As noted in Remark (ref), our approach does not necessarily imply $b^{\mathrm{sel}}/T\to 0$ as $T\to \infty$, which is a necessary condition for subsampling to work. Therefore, we report in Table (ref) the medians of $b^{\mathrm{sel}}/T$ to verify whether the condition holds. Since the median of $b^{\mathrm{sel}}/T$ moves toward zero as $T$ increases for all the tests and DGPs, the condition $b^{\mathrm{sel}}/T\to 0$ seems to hold.
In Appendix C, we conduct further experiments where the source of GARCH effects for $\varepsilon_{y,t}$ is those for $\varepsilon_{x,t}$. The GARCH parameters are chosen such that the simulated variables mimic the real data used in Section 7. The result is similar to that for DGPs 1-3.
To investigate the power properties of the tests, we check the size-adjusted power and the power of subsampling-based tests, under $T=100$ and DGP1. The size-adjusted powers are displayed in Table (ref). For the case of $\rho_\beta=0.6$, $s_{\beta,t}$ is a stationary AR(1) process, and the Wald test outperforms the LM test, as expected. Furthermore, the Sum and Product tests are as powerful as the Wald test, with the Sum test being the best. When $\rho_\beta=0.8$, $s_{\beta,t}$ is more persistent, and hence the LM test can perform better than the Wald test (along with the Sum and Product tests) for small $\omega_\beta^2$. However, for moderate to large $\omega_\beta^2$ ($= 0.05,0.1,0.5$), the LM test is dominated by the others. The best-performing test in these cases is the Product test. A similar pattern in power properties is observed in the case of more persistent $s_{\beta,t}$ with $\rho_\beta=0.9$, but the performance of the LM test gets better. When $\rho_\beta=0.98$, $s_{\beta,t}$ is almost a unit root, and the LM test performs the best and the Wald test worst for all $\omega_\beta^2$ considered, as expected from our theoretical analysis. As for the combination tests, the powers of the Sum test are quite similar to those of the Wald test, while the Product test behaves similarly to the LM test, ranking second or first.
Table (ref) shows the powers when tests are based on subsampling. The product test, which should not be based on subsampling, is not considered here. The pattern of the power of the tests is the same as that of the size-adjusted power: the Wald and Sum tests perform well when $s_{\beta,t}$ is stationary, with the Sum test attaining a higher power, and the LM test is the best test when $s_{\beta,t}$ is strongly persistent. Note that the Sum test is at least the second best test in all cases (because the Product test is not considered here).
This section illustrates how the tests considered thus far can be applied to real data analysis. Our analytical framework is exactly the same as that employed by devpuraStockReturnPredictability2018, who investigate whether 14 financial variables have a time-varying predictive ability for the U.S. stock market excess returns. They do this by first estimating predictive coefficient at each time point by recursive-window estimation, calculating p-values associated with the predictive coefficient estimates, and then concluding predictability is time-varying if 25-90% of the p-values are less than a nominal level (0.05, say). A problem of this approach, however, is that it conducts statistical tests sequentially without any adjustment for multiple testing, resulting in uncontrolled size. Moreover, the “25-90%" criterion is arbitrary. To get around these problems, we use the tests considered in this article (except the Product test). Note that these tests, if based on subsampling, generally do not suffer from the issues considered by devpuraStockReturnPredictability2018, namely, persistence, long-run endogeneity and conditional heteroskedasticity of predictors.
Our working regression is as in Section 6 with the U.S stock market excess returns used as dependent variable, and each of the 14 variables as a predictor. Specifically, the predictors are book-to-market ratio (BM), dividend payout ratio (DE), default yield spread (DFY), default return spread (DFR), dividend-price ratio (DP), dividend yield (DY), earnings-to-price (EP), inflation (INFL), long-term bond return (LTR), long-term bond yield (LTY), net equity expansion (NTIS), stock variance (SVAR), T-bill rate (TBL) and term spread (TMS); see devpuraStockReturnPredictability2018 for the detailed description of the variables. All the data are monthly, spanning 1927:1-2014:12 ($T=1055$).\footnote{We use data from 1927:1 to 2014:11 for predictors and from 1927:2 to 2014:12 for the dependent variable.} The data are taken from Amit Goyal's web page (\url{https://sites.google.com/view/agoyal145}). The estimated GARCH(1,1) effects for predictors are displayed in Table (ref), and the testing results in Table (ref).
Before discussing the testing results, it should be recalled that the Wald-based tests may suffer from the identification problem when the predictor is a stationary GARCH process (see Remark (ref)). Therefore, we should not conduct the Wald-based tests when such a predictor is used. According to Table (ref), DFR, INFL, LTR and SVAR are such predictors, so we do not perform the Wald and Sum tests when these variables are used as the predictor. We note here that chenTestingSmoothStructural2012's (chenTestingSmoothStructural2012) tests for parameter instability, which are constructed under the stationarity of $x_t$, may be applied to these predictors (as long as the assumptions on which their tests rely are satisfied).
devpuraStockReturnPredictability2018 determine that 7 predictors (EP, DFY, LTR, LTY, NTIS, SVAR and TBL) have time-varying predictive ability. According to our results, the LM test does not reject the null of constant predictability for any predictors at 5% level.\footnote{For DFR, INFL, LTR and SVAR, we also performed the standard LM test using the critical value from Table 1 of hansenTestingParameterInstability1992, which is calculated under the assumption of $\mathrm{I}(0)$ regressor. However, the standard LM test also failed to reject the null for these predictors.} In sharp contrast, the Wald and Sum tests reject the null for 5 predictors (BM, DE, DP, DY, DFY). Neither of the tests reject the null for EP, LTY, NTIS, TBL and TMS, 4 of which are judged to have time-varying predictive ability by devpuraStockReturnPredictability2018.
In this article, we proposed a Wald-type test for coefficient constancy that is designed to detect stationary random coefficient. We analyzed power properties of the Wald test and nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test under the sequence of local alternatives that takes different forms depending on whether the random coefficient is stationary or integrated and which test is considered. We found that the Wald test outperforms the LM test when the coefficient process is stationary, while the LM test is superior to the Wald test when the coefficient process is a (near) unit root. We also considered two combination tests of the Wald and LM, which take their sum and product. Our simulation study revealed the LM test is the best when random coefficient is highly persistent, the Wald and Sum tests are the most powerful tests when coefficient process is stationary, and the Product test has good power properties irrespective of the persistence of random coefficient, with the best performance for moderate persistence. The message from this result is that the best test for coefficient randomness differs from context to context, and the persistence of the random coefficient determines which test is the best one.
To deal with the nuisance-parameter problem associated with the above tests, we proposed to base them on subsampling. However, there may be a better way (in terms of performance and/or computation) to tackle the nuisance-parameter problem than the subsampling approach. A promising approach will be the IVX method proposed by phillipsEconometricInferenceVicinity2009. It has been shown in the literature kostakisRobustEconometricInference2015,phillipsRobustEconometricInference2016 that the IVX approach leads to valid inference in predictive regressions, irrespective of regressors' persistence and long-run endogeneity. The use of the IVX method would also make Wald-based tests feasible when the predictor is GARCH and has an autoregressive root distant from unity.
We applied the tests to the U.S. stock returns data. The result obtained from this exercise mostly contradicted the conclusion of an earlier research that sequentially applies hypothesis tests to investigate the coefficient constancy. This implies that earlier findings based only on sequential estimation of predictability might be misleading.
\titleformat*{\section}