EconBase
← Back to paper

Testing for Stationary or Persistent Coefficient Randomness in Predictive Regressions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,220 characters · 19 sections · 68 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Testing for Stationary or Persistent Coefficient Randomness in Predictive Regressions

\baselineskip= 6mm

titlingpage\begin{abstract} This study considers tests for coefficient randomness in predictive regressions. Our focus is on how tests for coefficient randomness are influenced by the persistence of random coefficient. We show that when the random coefficient is stationary, or I(0), nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test loses its optimality (in terms of power), which is established against the alternative of integrated, or I(1), random coefficient. We demonstrate this by constructing a test that is more powerful than the LM test when the random coefficient is stationary, although the test is dominated in terms of power by the LM test when the random coefficient is integrated. The power comparison is made under the sequence of local alternatives that approaches the null hypothesis at different rates depending on the persistence of the random coefficient and which test is considered. We revisit an earlier empirical research and apply the tests considered in this study to the U.S. stock returns data. The result mostly reverses the earlier finding. \end{abstract} Keywords: Coefficient constancy test, Predictive regression, Random coefficient JEL Codes: C12, C22, C58

Introduction

In the applied finance and economics literature, stock return predictability has been widely discussed. Predictability analysis often involves regressing stock returns on the lag of another financial/economic variable and investigating whether the coefficient on the variable is significantly different from zero (a significantly non-zero coefficient amounts to evidence for predictability). Despite a large body of literature on this topic, researchers have reported mixed evidence of predictability. Recently, several researchers have suspected this is due to the time-varying nature of predictability (i.e., the coefficient on the predictor may be nonzero in some periods of time but may be zero in other periods) and tried to estimate its time path guidolinTimeVaryingStock2013, devpuraStockReturnPredictability2018, rahmanPredictivePowerDividend2019, hammamiUnderstandingTimevaryingShorthorizon2020. In such studies, the coefficient at each time point is estimated in a rolling-window fashion, that is, by using data over limited sample periods, and authors plot the time path of its estimate, assessing whether the coefficient is time-varying and identifying the period when predictability exists (if any).

The authors, however, conduct this analysis without statistically confirming the assumption that the coefficient on the predictor is time-varying. They often argue it is time-varying just because its estimated trajectory is seemingly not constant over time. However, the coefficient estimates can be seemingly time-varying even if the coefficient is, in fact, constant over the full sample. This is because the coefficient at each time point is estimated by using a subsample, which tends to be small rahmanPredictivePowerDividend2019, and thus the standard errors of the coefficient estimates tend to be large. This implies that even if the coefficient is actually constant over time, the time variations of its estimates can be so large that only a graphical inspection erroneously supports the hypothesis of time-varying predictability.

A more decent approach is to statistically test for the constancy of the coefficient before estimating it or its trajectory. If the null of constant coefficient is rejected, the rolling-window estimation may be applied to identify the period when predictability exists (otherwise, the usual method based on the full sample should be used to test whether the coefficient is zero).

Testing procedures for coefficient constancy in predictive regressions have recently been discussed by georgievTestingParameterInstability2018. They propose several bootstrap-based tests under the assumption that the predictor is highly persistent, and find that a test based on nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test performs well when the coefficient on the predictor experiences an abrupt break or follows a near unit root AR(1) process. This result is due to the fact that nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test is designed to be the locally most powerful against martingale coefficient processes, which accommodate random walk and processes with abrupt breaks. However, in empirical applications, the coefficient process may not be a martingale, and it is unclear whether the LM test retains good power in such a case.

Models with non-martingale random coefficients are widely seen in the literature. swamyLinearPredictionEstimation1980 propose regression models with stationary random coefficients, gives them a justification from statistical and economic perspectives, and discuss prediction and estimation methods. linTestingParameterConstancy1999 derive the LM test for coefficient constancy against the alternative of stationary random coefficient, and fuSpecificationTestsTimevarying2023 propose a test to distinguish between stationary and persistent variations in parameters. Stationary-coefficient models are also employed in empirical studies. For instance, chiangForecastingTreasuryBill1991 estimate a model where spot interest rate is regressed on forward interest rate, and coefficients are assumed to be possibly stationary stochastic processes. leeRevisitingDemandMoney2013 estimate the money demand function with stationary random coefficients.

When stationary (or $\mathrm{I}(0)$) random coefficient is assumed as the alternative hypothesis, we can naturally test the null of constant coefficient by using a Wald-type test based on a linear model derived from the original predictive regression. This testing procedure is essentially the same as the one employed by nishiTestingCoefficientRandomness2023, who considers AR(1) models with i.i.d. random coefficient and shows that the Wald test performs well in such a case. In this article, we show that the Wald test performs better than Nyblom's LM test in predictive regressions with stationary (rather than i.i.d.) random coefficient. The Wald test is, however, less powerful than the LM test when the coefficient process follows a near unit root (or $\mathrm{I}(1)$) process.\footnote{One can also consider intermediate cases where the coefficient process is of $\mathrm{I}(d)$, where $d\in (0,1)$, but we do not consider such cases and will say that the coefficient process is stationary if it is $\mathrm{I}(0)$.}

The power comparison between the LM and Wald tests is made under the sequence of local alternatives. As shown later, the appropriate localizing rate of the alternative hypotheses depends on the persistence of the random coefficient and which test (the LM or Wald) is considered. When the random coefficient is $\mathrm{I}(0)$, the LM test has a nontrivial asymptotic power against the sequence of local alternatives approaching the null of zero randomness at the rate of $T^{-1/2}$, while the appropriate rate of the local alternatives for the Wald test is $T^{-3/4}$, where $T$ is the sample size. This implies that the Wald test can detect smaller variations in the random coefficient than the LM test when the random coefficient is $\mathrm{I}(0)$. In contrast, when the random coefficient is $\mathrm{I}(1)$, the LM test has nontrivial asymptotic power against the sequence of local alternatives of order $T^{-3/2}$, but the Wald test against that of order $T^{-5/4}$. Thus, the LM test detects smaller coefficient variations than the Wald test when the random coefficient is $\mathrm{I}(1)$.

Because one test does not dominate the other in terms of power, we consider combining the LM and Wald tests to obtain new tests that have good power properties irrespective of the persistence of the random coefficient. Our proposal is simply to take the sum and product of the two statistics. We will investigate the power properties of four tests in total.

The asymptotic null distributions of the Wald and LM test statistics are found to depend on nuisance parameters under our setting (and so do the combinations of these two tests). To implement these tests without the knowledge about the nuisance parameters, we base them on subsampling, in which critical values are approximated by using a number of test statistics calculated from subsamples andrewsHybridSizeCorrectedSubsampling2009.\footnote{Another way to circumvent the nuisance-parameter problem is to use bootstrap as georgievTestingParameterInstability2018 do. Because different approaches require different mathematical assumptions and discussions, we focus on subsampling. The reason why we use subsampling instead of bootstrap is explained in Remark 3 in Section 2.} We discuss the validity of the subsampling approach in our setting by some theoretical and numerical analyses.

Finally, we revisit devpuraStockReturnPredictability2018's (devpuraStockReturnPredictability2018) study. They determine whether predictability is time-varying based on a sequence of p-values associated with the rolling-window coefficient estimates. They do this, however, without any adjustment for multiple testing, resulting in uncontrolled size. In this article, we investigate the coefficient constancy by applying the tests considered in this article. Our result mostly reverses devpuraStockReturnPredictability2018's (devpuraStockReturnPredictability2018) conclusion.

The remainder of this article is structured as follows. In Section 2, we define the Wald test statistic assuming that the random coefficient is an $\mathrm{I}(0)$ process. We analyze and compare the asymptotic behavior of the Wald and LM tests under this setting. In Section 3, these two tests are compared under the assumption that the coefficient process is a near unit root, and two combination tests of the Wald and LM tests are proposed. Section 4 proposes to base these tests on subsampling. In Section 5, the model is extended to the heteroskedastic case, and we define modified test statistics. Section 6 conducts Monte Carlo simulations to investigate finite-sample behavior of the tests, and Section 7 gives a real data application. Section 8 concludes the article. All mathematical proofs are relegated to Appendix A.

The Case of $\mathrm{I}(0)$ Random Coefficient

Model and assumptions

In this section, we consider a predictive regression where the coefficient on the persistent predictor is expressed by the sum of a constant $\beta$ and stationary random variable $s_{\beta,t}$:

align[align omitted — 264 chars of source]

where $\mathrm{E}[s_{\beta,t}]=0$, $\mathrm{Var}(s_{\beta,t})=1$, and $s_{x,0}=O_p(1)$. Here, the predictor, $x_{t}$, is modeled as a persistent local-to-unity AR(1) process, which is a commonly employed assumption in the literature campbellEfficientTestsStock2006. Note that $\beta$ and $\omega_\beta^2$ denote the mean and variance of the coefficient on $x_{t-1}$, respectively. Therefore, coefficient constancy is equivalent to $\omega_\beta^2=0$ in our setting. To simplify the discussion in what follows, we set $\alpha_y=\alpha_x=s_{x,0}=0$, resulting in

align[align omitted — 128 chars of source]

with $x_t = (1+c_xT^{-1})x_{t-1} + \varepsilon_{x,t} \ (x_0=0)$.

remgeorgievTestingParameterInstability2018 consider cases where the intercept term as well as the slope parameter may vary over time. If the intercept term is the sum of constant $\alpha_y$ and stationary random variable $\omega_\alpha s_{\alpha,t}$, then the random part is absorbed into the disturbance term, and thus $\omega_\alpha$ is not identified. In empirical studies, however, researchers are often interested only in the time variation of the slope parameter. Therefore, we may focus on models where only the slope parameter possibly exhibits time variation. Interested readers are referred to georgievTestingParameterInstability2018 for marginal and joint tests for time variation in the intercept and slope (when they are nonstationary).

We suppose the following assumptions hold.

asm\begin{itemize} • $\varepsilon_{y,t}/\sigma_y = \gamma(\varepsilon_{x,t}/\sigma_x) + v_t$, where $\sigma_i^2 \coloneqq \mathrm{Var}(\varepsilon_{i,t})$, $i=x,y$, $\mathrm{E}[\varepsilon_{x,t}]=\mathrm{E}[v_t]=\mathrm{E}[\varepsilon_{x,t}v_t]=0$ and $\gamma \in (-1,1)$ is the correlation coefficient between $\varepsilon_{y,t}$ and $\varepsilon_{x,t}$. • $(\varepsilon_{x,t},v_t,s_{\beta,t})$ is stationary and $\alpha$-mixing with mixing coefficients $\alpha(m) = O(m^{-\frac{r}{r-2}})$ for some $r>2$. • $\mathrm{E}[|\varepsilon_{x,t}|^{4(r-1)+\delta}] + \mathrm{E}[|v_t|^{4(r-1)+\delta}] + \mathrm{E}[|s_{\beta,t}|^{4(r-1)+\delta}] < \infty$ for some $\delta>0$. • $(\varepsilon_{x,t},v_t,\eta_t)$, where $\eta_t \coloneqq \varepsilon_{y,t}^2 - \sigma_y^2$, is a martingale difference sequence (m.d.s.) with respect to some filtration $\mathcal{F}_t$ to which $(\varepsilon_{x,t}, v_t, s_{\beta,t})$ is adapted. \end{itemize}

In Assumption (ref)(i), the disturbance $\varepsilon_{y,t}$ of predictive regression (ref) is correlated with the innovation $\varepsilon_{x,t}$ that drives $x_{t}$, with correlation coefficient $\gamma$. Assumption (ref)(ii) specifies that the innovation process is stationary and weakly serially dependent, while all the innovations except $s_{\beta,t}$ are restricted to be serially uncorrelated by Assumption (ref)(iv). Assumption (ref)(iv) also implies conditional homoskedasticity of $\varepsilon_{y,t}$, i.e., $\mathrm{E}[\varepsilon_{y,t}^2|\mathcal{F}_{t-1}] = \sigma_{y}^2$. The case where $\varepsilon_{y,t}$ is conditionally heteroskedastic is considered in Section 5.

Define $\sigma_\eta^2 \coloneqq \mathrm{E}[\eta_t^2]$ and $\sigma_{\beta,lr}^2 \coloneqq \mathrm{E}[s_{\beta,1}^2] + 2\sum_{i=1}^{\infty}\mathrm{E}[s_{\beta,1}s_{\beta,1+i}]$. Define also the partial sum process $(W_{x,T}, W_{y,T}, W_{\eta,T},W_{\beta,lr,T})'$ on $[0,1]$ by $W_{i,T}(r) \coloneqq T^{-1/2}\sigma_i^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}\varepsilon_{i,t}$, $i=x,y$, $W_{\eta,T}(r) \coloneqq T^{-1/2}\sigma_\eta^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}\eta_t$ and $W_{\beta,lr,T}(r) \coloneqq T^{-1/2}\sigma_{\beta,lr}^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}s_{\beta,t}$. Then, it follows from the functional central limit theorem (FCLT) that under Assumption (ref), as $T \to \infty$,

align[align omitted — 118 chars of source]

in the Skorokhod space $D_4[0,1]$, where $(W_{x},W_{y},W_{\eta},W_{\beta,lr})'$ is a vector Brownian motion. Furthermore, by the standard continuous mapping argument, we also have $J_{x,T}(r) \coloneqq T^{-1/2}\sigma_x^{-1}x_{\lfloor Tr \rfloor} \Rightarrow J_x(r)$, where $J_x$ is an Ornstein-Uhlenbeck (OU) process satisfying $dJ_x(r) = c_xJ_x(r)dr + dW_x(r)$.

In some analysis, we rely on the following assumption instead of Assumption (ref):

asmAssumption (ref) holds with (ii) and (iii) replaced by the following conditions: \begin{itemize} • $(\varepsilon_{x,t},v_t,s_{\beta,t})$ is stationary and $\alpha$-mixing with mixing coefficients $\alpha(m) = O(m^{-\frac{r}{r-2}})$ for some $2<r<2.4$. • $\mathrm{E}[|\varepsilon_{x,t}|^{3r}] + \mathrm{E}[|v_t|^{3r}] + \mathrm{E}[|s_{\beta,t}|^{3r}] < \infty$. \end{itemize}

Wald and LM tests and asymptotic distributions

In the sequel, we let $\sigma_{y\beta} \coloneqq \mathrm{Cov}(\varepsilon_{y,t},s_{\beta,t})$.

Our aim here is to construct a Wald-type test for coefficient constancy in (ref), or $\omega_\beta^2=0$. To make our point clear, suppose for now $\beta$ is known. Define $z_t(\beta) \coloneqq y_t-\beta x_{t-1} = \omega_\beta s_{\beta,t}x_{t-1} + \varepsilon_{y,t}$, and obtain

align[align omitted — 115 chars of source]

where $\delta \coloneqq 2\omega_\beta\sigma_{y\beta}$ and $\xi_t \coloneqq (\varepsilon_{y,t}^2-\sigma_y^2) + 2\omega_\beta x_{t-1}(s_{\beta,t}\varepsilon_{y,t} - \sigma_{y\beta}) + \omega_\beta^2 x_{t-1}^2(s_{\beta,t}^2 - 1)$. Viewing (ref) as a linear model with $\xi_t$ playing the role of the disturbance, nishiTestingCoefficientRandomness2023 proposes to test $\mathrm{H_0}:\omega_\beta^2=0$ by the Wald test for $(\delta, \omega_\beta^2) = (0,0)$, since $\omega_\beta^2=0$ if and only if $(\delta,\omega_\beta^2) = (0,0)$. Moreover, the LM test of linTestingParameterConstancy1999, which is designed to detect stationary random coefficient, can be derived as the (squared) t-test for $\omega_\beta^2=0$ under the restriction $\delta=0$ (see Appendix B). Therefore, it is natural in our setting to employ model (ref) as the basis for constructing a test statistic. To define the Wald test explicitly, we first regress the variables on a constant, obtaining

align[align omitted — 147 chars of source]

where $\widetilde{z_t^2}(\beta) \coloneqq z_t^2(\beta) - T^{-1}\sum_{t=1}^{T}z_t^2(\beta)$, $\tilde{x}_{1,t-1} \coloneqq x_{t-1} - T^{-1}\sum_{t=1}^Tx_{t-1}$, $\tilde{x}_{2,t-1} \coloneqq x_{t-1}^2 - T^{-1}\sum_{t=1}^{T}x_{t-1}^2$, and $\tilde{\xi}_t \coloneqq \xi_t - T^{-1}\sum_{t=1}^T\xi_t$. Letting $Z_2(\beta) \coloneqq (\widetilde{z_1^2}(\beta),\ldots,\widetilde{z_T^2}(\beta))'$, $X_1 \coloneqq (\tilde{x}_{1,0},\ldots,\tilde{x}_{1,T-1})'$, $X_2 \coloneqq (\tilde{x}_{2,0},\ldots,\tilde{x}_{2,T-1})'$, and $\Xi \coloneqq (\tilde{\xi}_1,\ldots,\tilde{\xi}_t)'$, we arrive at

align[align omitted — 71 chars of source]

where $X\coloneqq (X_1,X_2)$ and $\theta\coloneqq (\delta,\omega_\beta^2)'$. Then, the Wald test statistic is

align[align omitted — 131 chars of source]

where $\hat{\theta}_T(\beta)$ and $\hat{\sigma}_{\tilde{\xi}}^2(\beta)$ are the OLS coefficient and variance estimators from (ref), respectively: $\hat{\theta}_T(\beta) = (X'X)^{-1}X'Z_2(\beta)$ and $\hat{\sigma}_{\tilde{\xi}}^2(\beta) = T^{-1}Z_2(\beta)'M_XZ_2(\beta)$ with $M_X \coloneqq I - X(X'X)^{-1}X'$.

The Wald test defined above uses the true value of $\beta$. Because $\beta$ is usually unknown in applications, we use $\mathrm{W}_T(\hat{\beta})$, where $\hat{\beta}$ is the OLS estimator of $\beta$: $\hat{\beta} =\sum_{t=1}^{T}y_t x_{t-1}/\sum_{t=1}^{T}x_{t-1}^2$. The use of $\hat{\beta}$ makes the Wald test invariant to the value of $\beta$ and feasible without any knowledge about $\beta$ (in particular, about whether it is zero or not).

We also consider nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test, which is defined by

align[align omitted — 156 chars of source]

where $\hat{\sigma}_T^2 \coloneqq T^{-1}\sum_{t=1}^{T}z_t^2(\hat{\beta})$.\footnote{georgievTestingParameterInstability2018 also investigate the performance of the $\mathrm{Sup}F$ test proposed by andrewsTestsParameterInstability1993. Although this test is known to be comparable in terms of power with Nyblom's LM test when parameters are $\mathrm{I}(1)$, it is not designed to detect (non)stationary randomness in parameters. Therefore, Nyblom's LM test seems more appropriate as the test to be compared with our Wald test, as argued by hansenTestsParameterInstability2002.}

remWhen $(\alpha_y,\alpha_x,s_{x,0}) \neq (0,0,0)$ in (ref), $y_t$ and $x_{t-1}$ are replaced with their demeaned counterparts in (ref) and the definitions of $z_t(\beta)$ and $\hat{\beta}$.
thm\begin{itemize} • Suppose Assumption (ref) holds. If $\omega_\beta^2 = g^2T^{-3/2}$ and $\mathrm{Corr}(\varepsilon_{y,t},s_{\beta,t}) = \sigma_{y\beta}/\sigma_y = mT^{-1/4}$, then we have, under model (ref), \begin{align} \mathrm{W}_T(\hat{\beta}) \Rightarrow &\Bigg\{ \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)dW_\eta(r) \\ \int_{0}^{1}\tilde{J}_{x,2}(r)dW_\eta(r) \end{pmatrix} + \frac{\sigma_x}{\sigma_\eta} \begin{pmatrix} g^2\sigma_x\int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr+2gm\sigma_y\int_{0}^{1}\tilde{J}_{x,1}(r)^2dr \\ g^2\sigma_x\int_{0}^{1}\tilde{J}_{x,2}(r)^2dr+2gm\sigma_y\int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr \end{pmatrix}\Biggr\}' \notag \\ &\times \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)^2dr & \int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr \\ \int_{0}^{1}\tilde{J}_{x,2}(r)\tilde{J}_{x,1}(r)dr & \int_{0}^{1}\tilde{J}_{x,2}(r)^2dr \end{pmatrix}^{-1} \notag \\ &\times \Bigg\{ \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)dW_\eta(r) \\ \int_{0}^{1}\tilde{J}_{x,2}(r)dW_\eta(r) \end{pmatrix} + \frac{\sigma_x}{\sigma_\eta} \begin{pmatrix} g^2\sigma_x\int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr+2gm\sigma_y\int_{0}^{1}\tilde{J}_{x,1}(r)^2dr \\ g^2\sigma_x\int_{0}^{1}\tilde{J}_{x,2}(r)^2dr+2gm\sigma_y\int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr \end{pmatrix}\Biggr\}, \end{align} where $\tilde{J}_{x,1}(r) \coloneqq J_x(r) - \int_{0}^{1}J_x(s)ds$ and $\tilde{J}_{x,2}(r) \coloneqq J_x(r)^2 - \int_{0}^{1}J_x(s)^2ds$.\footnote{The localization $\mathrm{Corr}(\varepsilon_{y,t},s_{\beta,t}) = mT^{-1/4}$ is just a technical condition to prevent test statistics from diverging and is not of interest in itself; see the proof of Theorem (ref) and nishiTestingCoefficientRandomness2023.} • Suppose Assumption (ref) holds. If $\omega_\beta^2 = g^2T^{-1}$ and $\mathrm{Corr}(\varepsilon_{y,t},s_{\beta,t}) = mT^{-1/4}$, then we have, under model (ref), \begin{align} \mathrm{LM}_T &\Rightarrow \frac{\int_{0}^{1}\bigl[\int_{0}^{s}J_x(r)dA(r) + g(\sigma_x\sigma_{\beta,lr}/\sigma_y)\int_{0}^{s}J_x(r)dB(r)\bigr]^2ds}{\bigl(1+g^2(\sigma_x^2/\sigma_y^2)\int_{0}^{1}J_x(r)^2dr\bigr)\int_{0}^{1}J_x(r)^2dr}, \\ \intertext{where} A(r) &\coloneqq W_y(r) - \frac{\int_{0}^{1}J_x(s)dW_y(s)}{\int_{0}^{1}J_x(s)^2ds}\int_{0}^{r}J_x(s)ds \\ \intertext{and} B(r) \coloneqq \Biggl(\int_{0}^{r}J_x(s)dW_{\beta,lr}(s) -& \frac{\int_{0}^{1}J_x(s)^2dW_{\beta,lr}(s)}{\int_{0}^{1}J_x(s)^2ds}\int_{0}^{r}J_x(s)ds\Biggr) + \frac{\Lambda_{x\beta}}{\sigma_x\sigma_{\beta,lr}}\Biggl(r-\frac{2\int_{0}^{1}J_x(s)ds}{\int_{0}^{1}J_x(s)^2ds}\int_{0}^{r}J_x(s)ds\Biggr) \end{align} with $\Lambda_{x\beta} \coloneqq \sum_{i=1}^{\infty}\mathrm{E}[\varepsilon_{x,1}s_{\beta,1+i}]$. \end{itemize}

There are two points worth mentioning in Theorem (ref). First, the asymptotic distributions of the Wald and LM test statistics are derived under different local alternatives. The Wald test statistic has the limiting alternative distribution when $\omega_\beta=gT^{-3/4}$, while the LM test does when $\omega_\beta=gT^{-1/2}$. Because the former specification captures a sequence of local alternatives closer to the null of $\omega_\beta=0$, it implies the Wald test can detect smaller variations in random coefficient than the LM test, as long as the random coefficient is $\mathrm{I}(0)$. This is intuitively because the Wald test is constructed under the assumption that the random coefficient follows a stationary process, while the derivation of the LM test is based on the assumption that the coefficient process is nonstationary. This finding is in line with the results of linTestingParameterConstancy1999. They verify through simulation that their test, which is designed to detect stationary random coefficient, outperforms nyblomTestingConstancyParameters1989's LM test when the random coefficient is a stationary AR process. Theorem (ref) gives a theoretical explanation for their simulation results.\footnote{linTestingParameterConstancy1999's (linTestingParameterConstancy1999) LM test has the asymptotic alternative distribution when $\omega_\beta=gT^{-3/4}$, as our Wald test does. See Appendix B for details.}

The second point is that the asymptotic null distributions ($g=0$) of the Wald and LM test statistics are dependent on nuisance parameters. Specifically, the Wald test depends on the correlation between $W_x$ and $W_\eta$, and the LM test on the correlation between $W_x$ and $W_y$. Moreover, these tests depend on $c_x$, the persistence parameter of $x_t$. Therefore, information on the nuisance parameters is required to perform these tests. In Section 4, we propose to base them on subsampling to solve this problem.

remgeorgievTestingParameterInstability2018 propose to include $\Delta x_t$ in the right hand side of (ref) to make the LM test independent of the correlation between $W_x$ and $W_y$, which renders the use of bootstrap valid in their setting. As for the Wald test, however, a Bonferroni approach is needed to make the test free from the correlation between $W_x$ and $W_\eta$ nishiTestingCoefficientRandomness2023. Because the joint use of the bootstrap and Bonferroni methods is impractical, we will rely on subsampling, which does not require the elimination of the correlations.

The Case of $\mathrm{I}(1)$ Random Coefficient

In this section, we consider a predictive regression model where the random part of the coefficient is near $\mathrm{I}(1)$:

align[align omitted — 250 chars of source]

where $x_t = (1+c_xT^{-1})x_{t-1} + \varepsilon_{x,t} \ (x_0=0)$. We impose the following condition on $\varepsilon_{\beta,t}$.

asm$\varepsilon_{\beta,t}$ is an m.d.s. and $\mathrm{Var}(\varepsilon_{\beta,t})=1$.

georgievTestingParameterInstability2018 prove that, under model (ref), the LM test statistic has the asymptotic alternative distribution when $\omega_\beta = gT^{-3/2}$. The next theorem establishes that the Wald test statistic has the asymptotic alternative distribution if $\omega_\beta = gT^{-5/4}$. To state this precisely, we need the following notation: let $J_\beta(r)$ be an OU process satisfying $dJ_\beta(r) = c_\beta J_\beta(r)dr + dW_\beta(r)$, where $W_\beta(\cdot)$ is the standard Brownian motion to which $W_{\beta,T}(\cdot) \coloneqq T^{-1/2}\sum_{t=1}^{\lfloor T\cdot \rfloor}\varepsilon_{\beta,t}$ weakly converges.

thmSuppose Assumption (ref) with $s_{\beta,t}$ replaced by $\varepsilon_{\beta,t}$ and Assumption (ref) hold. If $\omega_\beta = gT^{-5/4}$, then we have, under model (ref), \begin{align} \mathrm{W}_T(\hat{\beta}) \Rightarrow &\Bigg\{ \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)dW_\eta(r) \\ \int_{0}^{1}\tilde{J}_{x,2}(r)dW_\eta(r) \end{pmatrix} + \frac{g^2\sigma_x^2}{\sigma_\eta} \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)J_x(r)^2Q(r)^2dr \\ \int_{0}^{1}\tilde{J}_{x,2}(r)J_x(r)^2Q(r)^2dr \end{pmatrix}\Biggr\}' \notag \\ &\times \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)^2dr & \int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr \\ \int_{0}^{1}\tilde{J}_{x,2}(r)\tilde{J}_{x,1}(r)dr & \int_{0}^{1}\tilde{J}_{x,2}(r)^2dr \end{pmatrix}^{-1} \notag \\ &\times \Bigg\{ \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)dW_\eta(r) \\ \int_{0}^{1}\tilde{J}_{x,2}(r)dW_\eta(r) \end{pmatrix} + \frac{g^2\sigma_x^2}{\sigma_\eta} \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)J_x(r)^2Q(r)^2dr \\ \int_{0}^{1}\tilde{J}_{x,2}(r)J_x(r)^2Q(r)^2dr \end{pmatrix}\Biggr\}, \end{align} where $Q(r) \coloneqq J_\beta(r) - (\int_{0}^{1}J_x(r)^2dr)^{-1}\int_{0}^{1}J_x(r)^2J_\beta(r)dr$.

Based on Theorems (ref) and (ref) and the result of georgievTestingParameterInstability2018, we observe that the convergence rate of $\omega_\beta$ under which the tests have nontrivial asymptotic power depends on the nature of random coefficient $s_{\beta,t}$. We summarize this in Table (ref). When the random part $s_{\beta,t}$ of the coefficient is $\mathrm{I}(0)$, the Wald test has a nontrivial power for $\omega_\beta=gT^{-3/4}$, which describes a sequence of local alternatives closer to the null of $\omega_\beta=0$ than $\omega_\beta=gT^{-1/2}$ for the LM case. However, in the case of $\mathrm{I}(1)$ $s_{\beta,t}$, the convergence rate of $\omega_\beta$ associated with the nontrivial power of the LM test is $gT^{-3/2}$, which converges to the null of $\omega_\beta=0$ faster than $gT^{-5/4}$ for the Wald test. This implies that the Wald test dominates the LM test in terms of power when $s_{\beta,t}$ is $\mathrm{I}(0)$, while the LM test is more powerful than the Wald test when $s_{\beta,t}$ is $\mathrm{I}(1)$.

remNyblom's LM test is also the locally most powerful test against a structural break. Suppose that $s_{\beta,t} = 1\{t>\lfloor\tau_0T \rfloor\}$, where $1\{\cdot\}$ is the indicator function, and $\tau_0 \in(0,1)$. Then, it can be shown that our Wald test has a nontrivial asymptotic power when $\omega_\beta=gT^{-3/4}$, whereas Nyblom's LM test statistic has a local alternative distribution under $\omega_\beta=gT^{-1}$, as shown by georgievTestingParameterInstability2018. This implies that Nyblom's LM test is more powerful than our Wald test when parameter instability is due to a structural break.\footnote{Note that our analysis is about the local-alternative, or small-variation case. As pointed out by perronUsefulnessLackThereof2016, different results might be obtained for the tests studied in this article, if we consider large variations in parameters and allow for lagged dependent variables as regressors and serial correlation in the error term. We exclude such a case under model (ref) and Assumptions (ref) and (ref).} Because the asymptotic distribution of the Wald statistic can be derived in a similar way to that in the case of $\mathrm{I}(1)$ $s_{\beta,t}$, the detail is omitted.

Given the above observation, one ideal testing procedure is to use the Wald test when $\mathrm{I}(0)$ random coefficient is suspected and use the LM test when $\mathrm{I}(1)$ random coefficient is probable. However, the persistence of the coefficient (if it is random) is often unknown, and thus it will be difficult to rely on such a procedure in practice. To detect coefficient randomness efficiently whether it is $\mathrm{I}(0)$ or $\mathrm{I}(1)$, we propose two combination tests of the Wald and LM, one being defined as the sum of them and the other as the product. Specifically, we consider the following test statistics:

align[align omitted — 172 chars of source]

We shall call $\mathrm{S}_T$ the Sum test statistic and $\mathrm{P}_T$ the Product test statistic. The motivation of these two tests is simple. When $s_{\beta,t}$ is $\mathrm{I}(0)$ and $\omega_\beta=gT^{-3/4}$, $\mathrm{W}_T(\hat{\beta})$, $\mathrm{S}_T$ and $\mathrm{P}_T$ weakly converge to alternative distributions, but $\mathrm{LM}_T$ to the null distribution. In contrast, if $s_{\beta,t}$ is $\mathrm{I}(1)$ and $\omega_\beta=gT^{-3/2}$, $\mathrm{LM}_T(\hat{\beta})$, $\mathrm{S}_T$ and $\mathrm{P}_T$ have nontrivial asymptotic power but $\mathrm{W}_T(\hat{\beta})$ does not. In other words, the Sum and Product tests are more influenced by the more powerful of the Wald and LM tests, and thus they are expected to perform well in both the $\mathrm{I}(0)$ and $\mathrm{I}(1)$ $s_{\beta,t}$ cases.\footnote{A more preferable way to combine the LM and Wald tests will be to weight those test statistics based on some information on the persistence of the random coefficient. For example, the weighting function might be a function of the test statistic proposed by fuSpecificationTestsTimevarying2023, which is aimed at distinguishing between stationary and persistent parameters. However, their statistic cannot be directly used in our setting because their analysis is based on the assumption that regressor $x_t$ is weakly stationary.} The finite-sample performances of these tests are examined in Section 6.

Subsampling Approach

Basics

As observed in previous sections, tests we have considered are dependent on nuisance parameters. In the literature, the subsampling approach has been a popular solution to nuisance parameter problems romanoSubsamplingIntervalsAutoregressive2001,choiSubsamplingVectorAutoregressive2005,andrewsHybridSizeCorrectedSubsampling2009. The method allows us to approximate the sampling distribution of a statistic by calculating a number of the statistics based on subsets of the full sample. The idea behind this strategy is that, in large samples, the sampling distribution of subsample statistics will be close to that of the full sample statistic, and hence the empirical distribution of a number of subsample statistics can approximate it.

To fix ideas, take the LM test as an example. To implement the subsampling method, one has to set the subsample size $b < T$, the number of data used to calculate each subsample statistic. Then, the first subsample statistic is calculated using the subsample ranging from $t=1$ to $t=b$, the second one is based on $t=2,\ldots,b+1$, and continue until the last statistic based on $t=T-b+1,\ldots,T$ is calculated. Letting $\mathrm{LM}_{j,b}$ denote the subsample LM statistic based on $b$ data ranging from $t=j$, the subsample test statistics can be written as

align[align omitted — 195 chars of source]

where $\hat{\beta}_{j,b} \coloneqq \sum_{t=j}^{j+b-1}y_tx_{t-1}/\sum_{t=j}^{j+b-1}x^2_{t-1}$ is the $\beta$ estimate based on $t=j,\ldots,j+b-1$ and $\hat{\sigma}^2_{j,b} \coloneqq b^{-1}\sum_{t=j}^{j+b-1}z^2_t(\hat{\beta}_{j,b})$. If the significance level is $\alpha$, then the critical value (denoted by $c_{\mathrm{LM}}(1-\alpha)$) is defined by the $1-\alpha$ quantile of the empirical distribution function

align[align omitted — 86 chars of source]

where $1\{\cdot\}$ is the indicator function. The null is rejected if the full sample test statistic $\mathrm{LM}_T$ exceeds the critical value. The Wald, Sum and Product tests can be based on the subsampling approach in the same way.

The choice of $b$

When implementing the subsampling method, one has to choose the subsample size $b$. One necessary condition is $1/b + b/T\to 0$ as $T\to \infty$ romanoSubsamplingIntervalsAutoregressive2001. One reasonable way to choose $b$ is to compute the empirical distribution $L_{b,T}$ for a number of $b$ in some set $B$ and select $b$ that leads to the most `appropriate' $L_{b,T}$ in some sense. In the literature, several ways to define $B$ and the appropriateness criterion have been proposed. We here follow jachSubsamplingInferenceMean2012 and set, for some $q \in (0,1)$, $B = \{b_j \coloneqq \lfloor q^jT \rfloor|j=j_{\mathrm{max}},j_{\mathrm{max}}-1,\ldots,j_{\mathrm{min}}\}$, with $j_{\mathrm{min}}<j_{\mathrm{max}} \ (j_{\mathrm{min}},j_{\mathrm{max}} \in \mathbb{N})$, and select $b$ such that $b^{\mathrm{sel}} \coloneqq \operatorname*{\arg\!\min}_{b_j \in B} \ \mathrm{KS}(L_{b_j,T},L_{b_{j+1},T})$, where $\mathrm{KS}(x,y)$ denotes the Kolmogorov-Smirnov (KS) distance between distributions $x$ and $y$.

The idea behind this selection rule is that $L_{b,T}$, viewed as a function of $b$, should vary only slightly in an appropriate range of $b$, and hence values of $b$ leading to small KS distances should be selected. After preliminary experiments, we set $(q,j_{\mathrm{min}},j_{\mathrm{max}}) = (0.85,\lfloor \log_{0.85}(0.27) \rfloor,\lfloor \log_{0.85}(5/T^{0.9}) \rfloor)$ for the LM test and $(q,j_{\mathrm{min}},j_{\mathrm{max}}) = (0.95,\lfloor \log_{0.95}(0.3) \rfloor,\lfloor \log_{0.95}(15/T^{0.9}) \rfloor)$ for the Wald, Sum and Product tests, so that we consider $b$'s up to about 30% of the sample size $T$.

remOne drawback of the above approach is that it does not guarantee $b^{\mathrm{sel}}/T \to 0$ as $T\to \infty$, since $\max B/T \to q^{j_{\mathrm{min}}} >0$. To check whether this condition holds, we will investigate the behavior of $b^{\mathrm{sel}}/T$ as $T$ increases in the Monte Carlo simulation conducted in Section 6.

Validity of subsampling

andrewsHybridSizeCorrectedSubsampling2009 consider subsampling-based inference in nonregular models, including regression models with nearly integrated regressors, and provide sufficient conditions for subsampling-based tests to work. Although it will take a complicated analysis to check whether their sufficient conditions hold in our setting, the (in)validity of the subsampling approach can be conjectured by checking the monotonicity in $c_x$ of the critical value of the asymptotic null distributions. Here, we follow Andrews and Guggenberger's argument to discuss the validity of the subsampling approach in our setting.

For the sake of exposition, let $S_T$ denote a generic test statistic calculated using the full sample and $S_{j,b}, \ j =1,\ldots,T-b+1,$ denote its subsampling counterparts, which are functions of predictor $x_t = (1+c_x/T)x_{t-1} + \varepsilon_{x,t}$. Suppose that, under the null hypothesis, $S_T$ has an asymptotic distribution $P_{c_x}$, dependent on $c_x$. The aim of subsampling is to approximate the critical value from $P_{c_x}$ by the empirical distribution of $S_{j,b}$. However, due to the local-to-unity nature of the predictor, subsample statistics $S_{j,b}$ weakly converge to $P_0$, the asymptotic distribution of $S_T$ obtained under $c_x=0$. This is because $x_t = (1+c_x/T)x_{t-1} + \varepsilon_{x,t} = (1+c_x(b/T)/b)x_{t-1} + \varepsilon_{x,t} = (1+o(1)/b)x_{t-1} + \varepsilon_{x,t}$ (since $b/T \to 0$ as $T\to \infty$), so that under the asymptotics of $b\to \infty$ (rather than $T \to \infty$) $S_{j,b}$'s weakly converge to $P_0$. Therefore, the validity of subsampling is determined by the relation between the critical values from $P_{c_x}$ (the weak limit of $S_T$) and $P_0$ (the weak limit of $S_{j,b}$).

The LM and Wald tests

Figures (ref) and (ref) depict the 0.95 quantiles of the asymptotic distributions of the LM and Wald statistics, respectively, as a function of $c_x$. The quantiles of the LM statistic are computed for $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t}) = 0,-0.3,-0.6,-0.9$, and the quantiles of the Wald statistic are computed for $\mathrm{Corr}(\varepsilon_{x,t},\eta_t) = 0,-0.3,-0.6,-0.9$.\footnote{Quantiles are insensitive to the sign of the correlations, and thus we consider negative correlations only.} Computations are based on 20000 simulations where $(W_x,J_x,W_y,W_\eta)$ are approximated by normalized sums of 1000 steps. For the LM test, the asymptotic 0.95 quantile is minimized at $c_x=0$, and as $c_x$ tends to $-\infty$, it increases and converges to 0.47, the asymptotic 0.95 quantile obtained when $x_t$ is $\mathrm{I}(0)$ hansenTestingParameterInstability1992, for any value of $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$. Note that the critical value taken from the empirical distribution of the subsample LM statistics asymptotically corresponds to that obtained under $c_x=0$. Therefore, if the true value of $c_x$ is negative, the critical value based on the subsample LM statistics is smaller than the 0.95 quantile of the asymptotic distribution of the full-sample LM statistic for any $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$, resulting in over-rejection. To quantify the degree of the over-rejection, we show in Figure (ref) the asymptotic sizes of the subsampling LM test with significance level 0.05. It indicates that the LM test will be oversized for almost all the values of $c_x$, irrespective of the value of $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$.

To solve over-rejection problems associated with subsampling tests, andrewsHybridSizeCorrectedSubsampling2009 propose a hybrid test, where the referenced critical value is simply the maximum of the subsampling critical value and the critical value obtained under $c_x = -\infty$. In the case of the LM test with 0.05 significance level, the hybrid critical value is $c^*_{\mathrm{LM}}(0.95) \coloneqq \mathrm{max}\{c_{\mathrm{LM}}(0.95), 0.47\}$, where $c_{\mathrm{LM}}(0.95)$ is the subsampling critical value. Since the asymptotic 0.95 quantile of the full sample LM statistic under the null hypothesis is below 0.47 irrespective of the true value of $c_x$, this makes the LM test a conservative one. In what follows, we will use this hybrid critical value for the LM test.

For the Wald test, the asymptotic 0.95 quantile (Figure (ref)) is maximized at $-c_x=0$ irrespective of the value of $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$ and is decreasing in $-c_x$. This implies that the subsampling Wald test will asymptotically have size equal to or less than 0.05 (see Figure (ref)). Strictly speaking, the quantile is not maximized at $-c_x=0$ when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$, and thus the Wald test is slightly oversized in this case. However, we can show that the asymptotic null distribution of the Wald statistic is exactly the chi-square distribution with two degrees of freedom when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$, irrespective of the value of $c_x$. This is because $W_x$ and $W_\eta$ are independent in this case, and thus $J_x$ and $W_\eta$ also are. This implies that the asymptotic null distribution of $W_T(\hat{\beta})$ given in Theorem (ref) (with $g=0$), if conditional on $W_x$, is chi square with two degrees of freedom, and so is it unconditionally. Therefore, the theoretical asymptotic 0.95 quantile of the Wald statistic is 5.99 for any value of $c_x$, and the type 1 error is exactly 0.05 for all $c_x$. The slight discrepancy between the theoretical result and the result observed in Figure (ref) is only due to the approximation error. In summary, the Wald test, if based on subsampling, is expected to be valid in that its size is equal to or below the nominal level.

The Sum and Product tests

The analysis for the Sum and Product tests is somewhat complicated because they depend on both $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$ and $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$. Moreover, they also depend on $\mathrm{Corr}(\varepsilon_{y,t},\eta_t)$. To gain some insight, we consider the case where $\varepsilon_{y,t}$ is symmetric, which is satisfied when, e.g., $\varepsilon_{y,t}$ is normally distributed. This yields $\mathrm{Corr}(\varepsilon_{y,t},\eta_t)=0$, and thus we can focus on the effect of $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})$ and $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$ on the asymptotic size. We consider $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})=0,-0.3,-0.6,-0.9$ and $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0,-0.3$.\footnote{We do not consider the cases with $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=-0.6,-0.9$ because these values can cause the variance-covariance matrix of $(\varepsilon_{x,t}, \varepsilon_{y,t}, \eta_t)'$ to be an indefinite matrix.} In Figures (ref) and (ref), we plot the asymptotic quantiles and size of the Sum and Product tests as functions of $-c_x$ and $\mathrm{Corr}(\varepsilon_{x,t},\varepsilon_{y,t})=0,-0.3,-0.6,-0.9$, fixing the value of $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)$.

In Figure (ref), we plot the asymptotic 0.95 quantiles of the Sum test statistic when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$. It attains the minimum value when $-c_x=0$ and increases as $-c_x$ gets larger, although in a non-monotonic way. This leads to a slight over-rejection (see Figure (ref)). However, we suspect that this result is mostly due to the approximation error for the asymptotic null distribution of the Sum test statistic, as in the case of the Wald test. Indeed, we found that the type 1 error is closer to or below the nominal level if we increase the steps for the approximation of $(W_x,J_x,W_y,W_\eta)$ and the number of replications. When $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=-0.3$ (Figure (ref)), the asymptotic 0.95 quantile is decreasing in $-c_x$ in general and attains the maximum at $-c_x=0$. Therefore, the asymptotic size is equal to or below the nominal 0.05 level.

Figure (ref) plots the asymptotic 0.95 quantiles of the Product statistic when $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=0$. It attains the minimum value when $-c_x=0$ and increases as $-c_x$ gets larger, and thus the Product test is asymptotically oversized (see Figure (ref)). When $\mathrm{Corr}(\varepsilon_{x,t},\eta_t)=-0.3$, the same results are obtained (Figures (ref) and (ref)). Therefore, the Product test should not be based on subsampling, even in the special case (where $\varepsilon_{y,t}$ is symmetric).

remThe hybrid approach proposed by andrewsHybridSizeCorrectedSubsampling2009 is not directly applicable to the Sum and Product tests. This is because the asymptotic null distributions of these test statistics depend on the correlation between the LM and Wald tests, even when $c_x= -\infty$, which leads to violation of Assumption K of andrewsHybridSizeCorrectedSubsampling2009. Size correction methods for these tests will be a topic for future works.

The finite-sample size and power of the subsampling LM, Wald, and Sum tests will be investigated in Section 6.

Conditionally Heteroskedastic Case

In this section, we extend our discussions thus far to the conditionally heteroskedastic case, which is more plausible in many applications. Incorporating the conditional heteroskedasticity requires us to modify the tests considered in the preceding sections.

LM test

The heteroskedasticity-robust version of the LM test is derived by hansenTestingParameterInstability1992, which is of the form

align[align omitted — 157 chars of source]

Subsampling counterparts are defined similarly.

Wald test

For the Wald test, the usual modification involving the use of HC variance estimator is insufficient in our setting. The problem is that, when predictive regression error $\varepsilon_{y,t}$ is conditionally heteroskedastic, then $\varepsilon_{y,t}^2$ is serially correlated, which induces serial correlation in $\eta_t = \varepsilon_{y,t}^2 - \sigma_y^2$ and violates Assumption (ref)(iv).

To solve this problem, we follow nishiStochasticLocalModerate2024 to modify the Wald test. To facilitate the discussion, the following regularity condition is imposed on $\varepsilon_{x,t}$, $v_t$ and $\eta_t$.

asm\begin{itemize} • $(\varepsilon_{x,t},v_t)'$ is stationary and $\alpha$-mixing with mixing coefficients $\alpha(m) = O(m^{-\gamma})$ for some $\gamma>6$. Moreover, $E[|\varepsilon_{x,t}|^{6}] +E[|v_t|^6] + E[|\eta_t|^6] < \infty$. • $(\varepsilon_{x,t},v_t)'$ is an m.d.s. with respect to some filtration $\mathcal{F}_t$ to which $(\varepsilon_{x,t},v_t)'$ is adapted. • The long-run variance-covariance matrix of $(\varepsilon_{x,t}, \eta_t)'$ is positive definite: \begin{align} \begin{pmatrix} \sigma_{x}^2 & \sigma_{x\eta, lr} \\ \sigma_{x\eta, lr} & \sigma_{\eta,lr}^2 \end{pmatrix} > 0, \end{align} where $\sigma_x^2=\mathrm{E}[\varepsilon_{x,t}^2]$, $\sigma_{\eta,lr}^2\coloneqq E[\eta_1^2]+2\sum_{t=1}^{\infty}E[\eta_1\eta_{t+1}]$, and $\sigma_{x\eta,lr}\coloneqq E[\varepsilon_{x,1}\eta_1]+\sum_{t=1}^{\infty}E[\varepsilon_{x,1}\eta_{t+1}] + \sum_{t=1}^{\infty}E[\eta_1\varepsilon_{x,t+1}]$. \end{itemize}

Note that $\varepsilon_{y,t}$ and $\eta_t$ are stationary and $\alpha$-mixing with the same mixing coefficients as $(\varepsilon_{x,t},v_t)'$ under Assumption (ref)(i). Assumption (ref)(i) allows for weak serial correlation in $\eta_t$, while we keep the no serial correlation assumption for $\varepsilon_{x,t}$ and $v_t$ (and thus $\varepsilon_{y,t}$) in (ii) (cf. Assumption (ref)(iv)). Under this condition, conditional heteroskedasticity in $\varepsilon_{y,t}$ may be permitted. In Assumption (ref)(iii), we assume that $\varepsilon_{x,t}$ and $\eta_t$ have a finite 6th moment. This condition is a technical one to guarantee that Theorem 3.1 of liangWeakConvergenceStochastic2016 holds. This assumption might be weakened if a more sophisticated asymptotic theory is available.

Under Assumption (ref), we have $\Bigl(T^{-1/2}\sigma_{x}^{-1}\sum_{t=1}^{\lfloor T\cdot \rfloor}\varepsilon_{x,t}, T^{-1/2}\sigma_{\eta,lr}^{-1}\sum_{t=1}^{\lfloor T\cdot \rfloor}\eta_t\Bigr)' \Rightarrow \Bigl(W_x(\cdot), W_{\eta,lr}(\cdot)\Bigr)'$, where $(W_x, W_{\eta,lr})'$ is a vector Brownian motion. Furthermore, the one-sided long-run covariance between $\varepsilon_{x,t}$ and $\eta_t$, $\Lambda_{x\eta} \coloneqq \sum_{t=1}^{\infty}E[\varepsilon_{x,1}\eta_{t+1}]$, exists.

Under Assumption (ref), we can show that the asymptotic null distribution of the Wald test statistic becomes

align[align omitted — 806 chars of source]

see Appendix A for the proof. In view of this expression, we modify $\mathrm{W}_T(\hat{\beta})$ in the following way:

align[align omitted — 495 chars of source]

where $\hat{\sigma}_{\tilde{\xi},lr}^2(\hat{\beta})$ and $\hat{\Lambda}_{x\eta,T}$ are consistent estimators of $\sigma_{\eta,lr}^2$ and $\Lambda_{x\eta}$, respectively. In this article, the long-run (co)variance estimators are constructed nonparametrically using the Parzen kernel, i.e.,

align[align omitted — 398 chars of source]

where $\hat{\tilde{\xi}}_t$ is the OLS residual from (ref), and

align[align omitted — 166 chars of source]

Here, $l$ denotes the truncation number, for which we set $l=\lfloor 4(T/100)^{1/3}\rfloor$ in this article.

remAnother way to construct $\hat{\Lambda}_{x\eta,T}$ would be to replace $\Delta x_{t-i}$ with $\hat{\varepsilon}_{x,t-i}$, the residual obtained from the regression of $x_t$ on $(1,x_{t-1})$. This choice is valid in our setting, where $x_t$ is modeled as (near) $\mathrm{I}(1)$. However, if $x_{t}$ is $\mathrm{I}(0)$, the correction involving $\hat{\Lambda}_{x\eta,T}$ should be asymptotically negligible. In such a case, we need $\hat{\Lambda}_{x\eta,T}=o_p(T^{-1/2})$. As shown in Theorem 4.1 and Lemma 8.1 of phillipsFullyModifiedLeast1995, this is achieved by using (ref) and setting $l = cT^{\kappa}$ for some $c>0$ and $\kappa \in(1/4,2/3)$.

Under the null $\mathrm{H}_0:\omega_\beta=0$, a standard argument will show that these estimators are consistent because $\hat{\beta}-\beta=O_p(T^{-1})$. It will be tedious and much involved to establish the consistency of these estimators under the alternative hypothesis. However, we expect that they are consistent under the sequences of local alternatives considered in this article because the convergence rates of those local alternatives are sufficiently fast that the presence of random coefficient does not affect the consistency of the short-run variance estimator $\hat{\sigma}_{\tilde{\xi}}^2(\hat{\beta})$ (see Lemmas (ref) and (ref)).

The following theorem establishes that $W_T^{*}(\hat{\beta})$ has the same asymptotic null distribution as $W_T(\hat{\beta})$ has under Assumption (ref) (up to the difference between $W_{\eta}$ and $W_{\eta,lr}$).

thmSuppose Assumption (ref) and the null $\mathrm{H}_0:\omega_{\beta}=0$ hold. Then, we have \begin{align} W_T^{*}(\hat{\beta}) \Rightarrow &\begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)dW_{\eta,lr}(r) \\ \int_{0}^{1}\tilde{J}_{x,2}(r)dW_{\eta,lr}(r) \end{pmatrix}' \begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)^2dr & \int_{0}^{1}\tilde{J}_{x,1}(r)\tilde{J}_{x,2}(r)dr \\ \int_{0}^{1}\tilde{J}_{x,2}(r)\tilde{J}_{x,1}(r)dr & \int_{0}^{1}\tilde{J}_{x,2}(r)^2dr \end{pmatrix}^{-1}\begin{pmatrix} \int_{0}^{1}\tilde{J}_{x,1}(r)dW_{\eta,lr}(r) \\ \int_{0}^{1}\tilde{J}_{x,2}(r)dW_{\eta,lr}(r) \end{pmatrix} . \end{align}
remWhen $x_t$ is $\mathrm{I}(0)$, it is necessary that $E[x_{t-1}\eta_t] = E[x^2_{t-1}\eta_t]=0$ holds to identify the parameters in (ref), from which the Wald-based tests are derived. For example, these conditions do not hold when $\varepsilon_{x,t}$ is conditionally heteroskedastic and $\gamma \neq 0$, where $\gamma$ is the correlation coefficient between $\varepsilon_{x,t}$ and $\varepsilon_{y,t}$.
remThe above modification made to the Wald test does not allow for the unconditional heteroskedasticity in $\varepsilon_{y,t}$. To see this, recall that the Wald test is based on regression (ref). If $\varepsilon_{y,t}$ is unconditionally heteroskedastic and $E[\varepsilon_{y,t}^2] = \sigma_{y,t}^2$, say, then the appropriate regression model becomes \begin{align} z_t^2(\beta) = \sigma_{y,t}^2 + \delta x_{t-1} + \omega_\beta^2x_{t-1}^2 + \xi_t. \end{align} Because the Wald test is based on the OLS estimator from (ref), the time-varying intercept, $\sigma_{y,t}^2$, cannot be eliminated by the usual demeaning approach as in (ref). To accommodate the unconditional heteroskedasticity in $\varepsilon_{y,t}$, a different construction will be required for the Wald test. Note that the heteroskedasticity-robust LM test, $\mathrm{LM}_T^*$, allows for the unconditional heteroskedasticity (see georgievTestingParameterInstability2018).

Monte Carlo Simulation

In this section, Monte Carlo simulations are conducted to investigate the finite-sample performance of the tests. The test statistics considered are $\mathrm{LM}_T^*$, $\mathrm{W}_T^*(\hat{\beta})$ and their sum and product. These tests are based on subsampling with the choice of subsample size explained in Section 4.2. To see whether the heteroskedasticity-robust subsampling tests behave well, we use the data generating process given by (ref) along with

align[align omitted — 239 chars of source]

Here, we set $(\alpha_y,\beta,\alpha_x,s_{x,0},s_{\beta,0})=(0,0,0,0,0)$, $\rho_x=0.95$ and $\gamma=-0.8$.\footnote{We calculate the Wald statistic using demeaned $y_t$ and $x_t$, allowing for nonzero location parameters (see Remark (ref)). We also do this when conducting an empirical analysis in Section 7.} For the persistence of $s_{\beta,t}$, we consider $\rho_\beta \in \{0.6,0.8,0.9,0.98\}$. The coefficient variance $\omega_\beta^2$ is taken as $\omega_\beta^2 \in \{0.01,0.05,0.1,0.5\}$. For the driver process, we specify that $v_t$ is GARCH(1,1):

align[align omitted — 246 chars of source]

where $(u_{x,t},u_{v,t},u_{\beta,t})' \sim \mathrm{i.i.d.} \ N(0,\mathrm{diag}(1,0.36,1))$, and $\sigma_{v,t}^2=c_0+c_1\sigma_{v,t-1}^2+c_2v_{t-1}^2$. For the GARCH parameters, we consider the following configurations:

itemize• DGP1 (weak case): $(c_0,c_1,c_2) = (0.6,0.2,0.2)$. • DGP2 (moderate case): $(c_0,c_1,c_2) = (0.4,0.3,0.3)$. • DGP3 (strong case): $(c_0,c_1,c_2) = (0.2,0.4,0.4)$.

All the tests are implemented under 5% significance level.

Size

Empirical sizes are displayed in Table (ref). Because we have found that the Product test is oversized if based on subsampling (see Section 4.3.2), we do not consider the Product test here. For DGP1, the sizes of all the tests are around 0.02-0.04, which indicates that these tests are valid in that their sizes are below the nominal level. The sizes for DGP2 and DGP3 are similar, so the same comment applies.

As noted in Remark (ref), our approach does not necessarily imply $b^{\mathrm{sel}}/T\to 0$ as $T\to \infty$, which is a necessary condition for subsampling to work. Therefore, we report in Table (ref) the medians of $b^{\mathrm{sel}}/T$ to verify whether the condition holds. Since the median of $b^{\mathrm{sel}}/T$ moves toward zero as $T$ increases for all the tests and DGPs, the condition $b^{\mathrm{sel}}/T\to 0$ seems to hold.

In Appendix C, we conduct further experiments where the source of GARCH effects for $\varepsilon_{y,t}$ is those for $\varepsilon_{x,t}$. The GARCH parameters are chosen such that the simulated variables mimic the real data used in Section 7. The result is similar to that for DGPs 1-3.

Power comparison

To investigate the power properties of the tests, we check the size-adjusted power and the power of subsampling-based tests, under $T=100$ and DGP1. The size-adjusted powers are displayed in Table (ref). For the case of $\rho_\beta=0.6$, $s_{\beta,t}$ is a stationary AR(1) process, and the Wald test outperforms the LM test, as expected. Furthermore, the Sum and Product tests are as powerful as the Wald test, with the Sum test being the best. When $\rho_\beta=0.8$, $s_{\beta,t}$ is more persistent, and hence the LM test can perform better than the Wald test (along with the Sum and Product tests) for small $\omega_\beta^2$. However, for moderate to large $\omega_\beta^2$ ($= 0.05,0.1,0.5$), the LM test is dominated by the others. The best-performing test in these cases is the Product test. A similar pattern in power properties is observed in the case of more persistent $s_{\beta,t}$ with $\rho_\beta=0.9$, but the performance of the LM test gets better. When $\rho_\beta=0.98$, $s_{\beta,t}$ is almost a unit root, and the LM test performs the best and the Wald test worst for all $\omega_\beta^2$ considered, as expected from our theoretical analysis. As for the combination tests, the powers of the Sum test are quite similar to those of the Wald test, while the Product test behaves similarly to the LM test, ranking second or first.

Table (ref) shows the powers when tests are based on subsampling. The product test, which should not be based on subsampling, is not considered here. The pattern of the power of the tests is the same as that of the size-adjusted power: the Wald and Sum tests perform well when $s_{\beta,t}$ is stationary, with the Sum test attaining a higher power, and the LM test is the best test when $s_{\beta,t}$ is strongly persistent. Note that the Sum test is at least the second best test in all cases (because the Product test is not considered here).

Empirical Application

This section illustrates how the tests considered thus far can be applied to real data analysis. Our analytical framework is exactly the same as that employed by devpuraStockReturnPredictability2018, who investigate whether 14 financial variables have a time-varying predictive ability for the U.S. stock market excess returns. They do this by first estimating predictive coefficient at each time point by recursive-window estimation, calculating p-values associated with the predictive coefficient estimates, and then concluding predictability is time-varying if 25-90% of the p-values are less than a nominal level (0.05, say). A problem of this approach, however, is that it conducts statistical tests sequentially without any adjustment for multiple testing, resulting in uncontrolled size. Moreover, the “25-90%" criterion is arbitrary. To get around these problems, we use the tests considered in this article (except the Product test). Note that these tests, if based on subsampling, generally do not suffer from the issues considered by devpuraStockReturnPredictability2018, namely, persistence, long-run endogeneity and conditional heteroskedasticity of predictors.

Our working regression is as in Section 6 with the U.S stock market excess returns used as dependent variable, and each of the 14 variables as a predictor. Specifically, the predictors are book-to-market ratio (BM), dividend payout ratio (DE), default yield spread (DFY), default return spread (DFR), dividend-price ratio (DP), dividend yield (DY), earnings-to-price (EP), inflation (INFL), long-term bond return (LTR), long-term bond yield (LTY), net equity expansion (NTIS), stock variance (SVAR), T-bill rate (TBL) and term spread (TMS); see devpuraStockReturnPredictability2018 for the detailed description of the variables. All the data are monthly, spanning 1927:1-2014:12 ($T=1055$).\footnote{We use data from 1927:1 to 2014:11 for predictors and from 1927:2 to 2014:12 for the dependent variable.} The data are taken from Amit Goyal's web page (\url{https://sites.google.com/view/agoyal145}). The estimated GARCH(1,1) effects for predictors are displayed in Table (ref), and the testing results in Table (ref).

Before discussing the testing results, it should be recalled that the Wald-based tests may suffer from the identification problem when the predictor is a stationary GARCH process (see Remark (ref)). Therefore, we should not conduct the Wald-based tests when such a predictor is used. According to Table (ref), DFR, INFL, LTR and SVAR are such predictors, so we do not perform the Wald and Sum tests when these variables are used as the predictor. We note here that chenTestingSmoothStructural2012's (chenTestingSmoothStructural2012) tests for parameter instability, which are constructed under the stationarity of $x_t$, may be applied to these predictors (as long as the assumptions on which their tests rely are satisfied).

devpuraStockReturnPredictability2018 determine that 7 predictors (EP, DFY, LTR, LTY, NTIS, SVAR and TBL) have time-varying predictive ability. According to our results, the LM test does not reject the null of constant predictability for any predictors at 5% level.\footnote{For DFR, INFL, LTR and SVAR, we also performed the standard LM test using the critical value from Table 1 of hansenTestingParameterInstability1992, which is calculated under the assumption of $\mathrm{I}(0)$ regressor. However, the standard LM test also failed to reject the null for these predictors.} In sharp contrast, the Wald and Sum tests reject the null for 5 predictors (BM, DE, DP, DY, DFY). Neither of the tests reject the null for EP, LTY, NTIS, TBL and TMS, 4 of which are judged to have time-varying predictive ability by devpuraStockReturnPredictability2018.

Conclusion

In this article, we proposed a Wald-type test for coefficient constancy that is designed to detect stationary random coefficient. We analyzed power properties of the Wald test and nyblomTestingConstancyParameters1989's (nyblomTestingConstancyParameters1989) LM test under the sequence of local alternatives that takes different forms depending on whether the random coefficient is stationary or integrated and which test is considered. We found that the Wald test outperforms the LM test when the coefficient process is stationary, while the LM test is superior to the Wald test when the coefficient process is a (near) unit root. We also considered two combination tests of the Wald and LM, which take their sum and product. Our simulation study revealed the LM test is the best when random coefficient is highly persistent, the Wald and Sum tests are the most powerful tests when coefficient process is stationary, and the Product test has good power properties irrespective of the persistence of random coefficient, with the best performance for moderate persistence. The message from this result is that the best test for coefficient randomness differs from context to context, and the persistence of the random coefficient determines which test is the best one.

To deal with the nuisance-parameter problem associated with the above tests, we proposed to base them on subsampling. However, there may be a better way (in terms of performance and/or computation) to tackle the nuisance-parameter problem than the subsampling approach. A promising approach will be the IVX method proposed by phillipsEconometricInferenceVicinity2009. It has been shown in the literature kostakisRobustEconometricInference2015,phillipsRobustEconometricInference2016 that the IVX approach leads to valid inference in predictive regressions, irrespective of regressors' persistence and long-run endogeneity. The use of the IVX method would also make Wald-based tests feasible when the predictor is GARCH and has an autoregressive root distant from unity.

We applied the tests to the U.S. stock returns data. The result obtained from this exercise mostly contradicted the conclusion of an earlier research that sequentially applies hypothesis tests to investigate the coefficient constancy. This implies that earlier findings based only on sequential estimation of predictability might be misleading.

table[table omitted — 398 chars of source]
table[table omitted — 876 chars of source]
table[table omitted — 858 chars of source]
table[table omitted — 1,717 chars of source]
table[table omitted — 1,584 chars of source]
table[table omitted — 1,421 chars of source]
table[table omitted — 923 chars of source]
figure[figure omitted — 642 chars of source]
figure[figure omitted — 686 chars of source]
sidewaysfigure\begin{subfigure}{0.47\textwidth} \caption{Asymptotic 0.95 quantiles ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=0$)} \end{subfigure} \begin{subfigure}{0.47\textwidth} \caption{Asymptotic 0.95 quantiles ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=-0.3$)} \end{subfigure} \begin{subfigure}{0.47\textwidth} \caption{Asymptotic sizes under the 0.05 nominal level ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=0$)} \end{subfigure} \begin{subfigure}{0.47\textwidth} \caption{Asymptotic sizes under the 0.05 nominal level ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=-0.3$)} \end{subfigure} \caption{Asymptotic behavior of the Sum test}
sidewaysfigure\begin{subfigure}{0.47\textwidth} \caption{Asymptotic 0.95 quantiles ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=0$)} \end{subfigure} \begin{subfigure}{0.47\textwidth} \caption{Asymptotic 0.95 quantiles ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=-0.3$)} \end{subfigure} \begin{subfigure}{0.47\textwidth} \caption{Asymptotic sizes under the 0.05 nominal level ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=0$)} \end{subfigure} \begin{subfigure}{0.47\textwidth} \caption{Asymptotic sizes under the 0.05 nominal level ($\mathrm{Corr}(\varepsilon_{x,t}, \eta_t)=-0.3$)} \end{subfigure} \caption{Asymptotic behavior of the Product test}

\titleformat*{\section}