Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
60,047 characters · 10 sections · 68 citation commands
Testing for Coefficient Randomness in Local-to-Unity Autoregressions
\baselineskip= 6mm
In this paper, we consider the random-coefficient autoregressive (RCA) model,
where $(\varepsilon_t,v_t)'$ is a random vector with $\mathbb{E}[\varepsilon_t] = \mathbb{E}[v_t] = 0$ and $\mathbb{V}[v_t]=1$. Model (ref) is a generalization of the usual AR(1) model, in that the autoregressive coefficient fluctuates over time around its mean $\rho$ with variance $\omega^2$, instead of being constant at $\rho$. Since nicholls1982RandomCoefficient, much attention has been paid to the estimation and inference theory of model (ref); see, for example, hwang2005ExplosiveRandomCoefficientb, aue2011QuasiLikelihoodEstimation, horvath2019Testingrandomnessa and references therein.
One aspect of the inferential theory for (ref) that has attracted econometricians and statisticians is how to test the hypothesis of $\omega^2=0$, that is, how to test the null hypothesis of the usual autoregressive specification against the RCA alternative. nicholls1982RandomCoefficient and lee1998Coefficientconstancya proposed test statistics for testing $H_0: \omega^2=0$ under the stationarity condition $\rho^2+\omega^2<1$. nagakura2009Testingcoefficienta showed the lee1998Coefficientconstancya test statistic has the same asymptotic null distribution when $\rho=1$ as when $|\rho|<1$.\footnote{nagakura2009Asymptotictheory also showed the lee1998Coefficientconstancya test is consistent when $\rho^2+\omega^2>1$ and some other conditions hold. However, he did not show the test statistic has the same limiting null distribution when $\rho>1$ as when $|\rho|\leq1$.} The condition, $\rho^2+\omega^2\leq1$, which these tests rest on, however, is restrictive and limits their applicability in practice, in view of the empirical analyses by hill2014UnifiedIntervalb and horvath2019Testingrandomnessa. They applied model (ref) to several macroeconomic variables and estimated $\rho$ to be near 1 for most of the variables, with $\rho$ estimated to be greater than 1 for some. For instance, the $\rho$ estimates obtained by horvath2019Testingrandomnessa took values between 0.9884 and 1.0021. Their results indicate that nonstationary RCA models where $\rho^2+\omega^2$ and $\rho$ may be greater than 1 with $\rho$ near unity are empirically relevant, and thus testing procedures are required that are valid under these conditions. horvath2019Testingrandomnessa proposed a test statistic for $H_0: \omega^2=0$ valid under the nonstationarity conditions (as well as under the stationarity conditions).
Another insight earlier studies provide is that the variation of the autoregressive root is smaller (if it exists) than assumed in the extant literature. For example, according to the empirical analysis conducted by horvath2019Testingrandomnessa, 4 variables out of 7 were estimated to have positive variance ($\omega^2>0$) in their coefficient, which is between $7.2\times10^{-5}$ and $8.2\times10^{-3}$. However, horvath2019Testingrandomnessa studied the finite-sample power of their test only for $\omega^2\geq2.5\times10^{-1}$, which are of larger magnitude than observed in practice (as long as macroeconomic variables are concerned). \footnote{horvath2022ChangepointDetection applied the RCA model to a cryptocurrency index and estimated $\omega^2$ to be around $10^{-1}$.} Therefore, it is unclear whether their test performs well when $\omega^2$ is of small magnitude that seems to be typical in empirical studies. In fact, their test has almost no distinguishing power when $\omega^2$ is relatively small, as our simulation studies will reveal.
Given these observations, we consider model (ref) with $\rho$ near unity and $\omega^2$ near zero and employ the following near unit root RCA model:
where $\rho_T=1+a/T$ and $\omega_T=c/T^{3/4}$. When $\omega_T^2=0$, (ref) is the popular (conventional) local-to-unity AR(1) model. This specification has been useful in analyzing the power of unit root tests elliott1996Efficienttestsa and in developing estimation and inference theory for models involving persistent variables phillips1987Unifiedasymptotica,stock1991Confidenceintervals,campbell2006Efficienttests,phillips2014ConfidenceIntervals. The local-to-zero variance $\omega_T^2=c^2/T^{3/2}$ has been employed under $\rho_T=1$ by mccabe1998Powertestsa and nishi2022StochasticLocal to derive and compare the local asymptotic power functions of unit root tests of $\omega^2=0$ against $\omega^2>0$. Introducing the local-to-zero variance provides us with a convenient framework in which we can evaluate power functions of several tests against the alternatives that are close to the null of $\omega^2=0$ and seem relevant for empirical applications.\footnote{For example, when $T=200$ and $0<c^2\leq50$, the variance $\omega_T^2$ takes values between 0 and $1.8\times10^{-2}$.} Model (ref) integrates these two local-to-unit-root parametrizations, thereby rendering itself an empirically relevant random-coefficient model.
In the literature, testing procedures for $H_0: \omega^2=0$ have often been studied in the special case with $\rho_T=1$. Such a model is called stochastic unit root (STUR) model. Test statistics for $H_0: \omega^2=0$ in the STUR model have been proposed by earlier studies, including mccabe1995Testingtimea, leybourne1996Caneconomica and distaso2008Testingunit. Moreover, mccabe1998Powertestsa and nishi2022StochasticLocal analyzed the power properties of several tests for $H_0: \omega^2=0$ under the STUR modelling. On the one hand, the STUR model is a generalization of the pure unit root model, which is a main reason authors have paid much attention to it. But on the other hand, the assumption of $\rho_T$ exactly being unity is a restrictive condition that is unlikely to hold in empirical analysis. Our model, (ref), generalizes the STUR model by allowing $\rho_T$ to take a value different from unity (but near it), which is a more realistic assumption given the empirical analyses conducted by earlier studies.
nishi2022StochasticLocal revealed that, under the STUR ($\rho_T=1$) modelling with $\mathrm{Corr}(\varepsilon_t,v_t)=0$, the test for $H_0: \omega^2=0$ proposed by lee1998Coefficientconstancya and nagakura2009Testingcoefficienta (hereafter, LN) has a higher local asymptotic power function than other tests. We will demonstrate that this is also the case when $\rho_T = 1+a/T$ and $\mathrm{Corr}(\varepsilon_t,v_t)=0$. In this paper, however, it will be shown that the LN test can perform poorly when $\mathrm{Corr}(\varepsilon_t,v_t)\neq0$, for both the cases $\rho_T=1$ and $\rho_T=1+a/T$. We will therefore propose several tests for $H_0: \omega^2=0$ whose power properties are robust to the value of $\mathrm{Corr}(\varepsilon_t,v_t)$. One of those tests turns out to be more powerful than the LN test for moderate to large values of $\mathrm{Corr}(\varepsilon_t,v_t)$. To the best of our knowledge, this study is the first to investigate how the value of $\mathrm{Corr}(\varepsilon_t,v_t)$ affects the power properties of tests for $H_0: \omega^2=0$, with the only exception of su2012Examiningpowera, who analyzed through simulation this effect in finite samples under $\rho=1$.
The other issue that this paper tackles is how to remove the effect of nuisance parameters from the null distributions of the test statistics mentioned above. As pointed out by nagakura2009Testingcoefficienta, the limiting null distribution of the LN test statistic depends on the correlation between $\varepsilon_t$ and $\varepsilon_t^2-\sigma_{\varepsilon}^2$, $\mathrm{Corr}(\varepsilon_t,\varepsilon_t^2-\sigma_{\varepsilon}^2)$, and so do the test statistics proposed in this article. If the true value of $\rho_T$ is known, the nuisance parameter can be removed straightforwardly by a similar way to the modification proposed by nagakura2009Testingcoefficienta, but this is not the case when the true $\rho_T$ is unknown. This problem is caused by the fact that the localizing parameter $a$ cannot be consistently estimated, as will be made clear later. As a solution for this problem, we propose a testing procedure based on the Bonferroni approach as campbell2006Efficienttests and phillips2014ConfidenceIntervals did in the context of predictive regressions.
The remainder of this paper is organized as follows. In Section 2, we consider the local-to-unity model (ref) with the true value of $\rho_T$ (or equivalently $a$) known, to uncover the effect $\mathrm{Corr}(\varepsilon_t,v_t)$ has on power properties of tests for $H_0: \omega^2=0$ and propose new tests that have power functions robust to this effect. We also propose a modification to make these tests independent of $\mathrm{Corr}(\varepsilon_t,\varepsilon_t^2-\sigma_{\varepsilon}^2)$, a nuisance parameter, under the null. In Section 3, we consider model (ref) with unknown $\rho_T$. Because the modification proposed in the preceding section cannot be directly applied in this case due to the fact that $a$ is not consistently estimable, we base our tests on the Bonferroni approach by constructing a confidence interval for $a$. Section 4 analyzes the finite-sample properties of our tests and compares them with those of existing tests. In Section 5, we apply our tests to real data. Section 6 concludes.
In this section, we begin our analysis under the assumption that the true value of $\rho_T$ is known. This assumption will be relaxed in the next section. Our analysis for model (ref) in this and the next section is conducted under the following assumption.
Define $\eta_t \coloneqq \varepsilon_t^2 - \sigma_{\varepsilon}^2$ and $\sigma_\eta^2 \coloneqq \mathbb{E}[\eta_t^2]$. Define also the partial sum process $(W_{\varepsilon,T}, W_{\eta,T})'$ on $[0,1]$ by $W_{\varepsilon,T}(r) \coloneqq T^{-1/2}\sigma_{\varepsilon}^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}\varepsilon_t$ and $W_{\eta,T}(r) \coloneqq T^{-1/2}\sigma_\eta^{-1}\sum_{t=1}^{\lfloor Tr\rfloor}\eta_t$. Then, it follows from the functional central limit theorem (FCLT) that under Assumption (ref), as $T \to \infty$
in the Skorokhod space $D[0,1]$, where $(W_{\varepsilon},W_{\eta})'$ is a vector Brownian motion with the covariance coefficient
with $\psi \coloneqq \mathbb{E}[\varepsilon_t\eta_t]/(\sigma_{\varepsilon}\sigma_\eta)$. Note that $W_\varepsilon$ and $W_\eta$ are not necessarily independent because of the covariance $\psi$, and the Brownian motion $W_\eta$ satisfies the following equality in distribution:
where $W_1$ is a standard Brownian motion independent of $W_\varepsilon$. Following the argument by nishi2022StochasticLocal, we can show that under Assumption (ref), the standardized process $T^{-1/2}y_{\lfloor T\cdot \rfloor}$ on $[0,1]$ constructed from (ref) weakly converges to the Ornstein-Uhlenbeck (OU) process $\sigma_{\varepsilon}J_a(\cdot)$, where $J_a$ solves $dJ_a(r) = aJ_a(r)dr + dW_\varepsilon(r)$. We can also construct consistent estimators of variances, namely, $\sigma_{\varepsilon}^2$ and $\sigma_\eta^2$. Define $z_t(\rho_T) \coloneqq y_t - \rho_T y_{t-1}(=\omega_Tv_ty_{t-1}+\varepsilon_t)$. Then, the estimators $\hat{\sigma}_{\varepsilon,T}^2(\rho_T) \coloneqq T^{-1} \sum_{t=1}^{T}z_t^2(\rho_T)$ and $\hat{\sigma}^2_{\eta,T}(\rho_T) \coloneqq T^{-1}\sum_{t=1}^{T}\{z_t^2(\rho_T) - \hat{\sigma}_{\varepsilon,T}^2(\rho_T)\}^2$ are consistent for $\sigma_{\varepsilon}^2$ and $\sigma_\eta^2$, respectively, which is proven in Appendix B.
Several tests of $H_0: \omega_T^2=0$ for model (ref) have been proposed by earlier work such as LN. The LN test statistic is defined by
nishi2022StochasticLocal found that the LN test has a high power function under the assumption that $\sigma_{\varepsilon v}=0$ and $\rho_T=1$.
As pointed out by nishi2022StochasticLocal, the LN test, which was originally derived as a locally best invariant test, can also be obtained as the t-test for $H_0: \omega_T^2=0$ under $\sigma_{\varepsilon v}=0$. To see this, note that from model (ref), a simple calculation gives
where $\xi_t \coloneqq \omega_T^2y_{t-1}^2(v_t^2-1) + 2\omega_Ty_{t-1}\varepsilon_tv_t + (\varepsilon_t^2-\sigma_{\varepsilon}^2)$. Because $\mathbb{E}[\xi_t]=0$ and $\mathbb{E}[y_{t-1}^2\xi_t]=0$ under Assumption (ref) with $\sigma_{\varepsilon v}=0$, model (ref) may be viewed as the linear regression model with $\xi_t$ playing the role of the disturbance, and the LN test statistic is obtained as the t-test for $H_0: \omega_T^2=0$ (with the variance estimator under the null used).
In view of this observation, the LN test seems to be a natural test for $H_0: \omega_T^2=0$. However, this may not be the case when $\sigma_{\varepsilon v}\neq0$, because it results in $\mathbb{E}[y_{t-1}^2\xi_t]=2\omega_T\sigma_{\varepsilon v}\mathbb{E}[y_{t-1}^3]\neq0$, that is, endogeneity. In fact, as we will shortly show, the power of the LN test is crucially affected by the value of $\sigma_{\varepsilon v}$, and the greater the value of $\sigma_{\varepsilon v}$ (in absolute value), the more poorly the LN test performs. Because we consider the localized model (ref), the setup for analyzing the influence $\sigma_{\varepsilon v}$ has is also a localized one. To derive relevant local asymptotic distributions, we localize the correlation (rather than the covariance $\sigma_{\varepsilon v}$) between $\varepsilon_t$ and $v_t$ in the following way:
Here, $q$ is interpreted as the correlation coefficient in the limit as $T\to \infty$. Under this localization, the LN test statistic has the following asymptotic distribution:
Note that when $q=0$ and $\rho_T=1$ (or $a=0$), the limit distribution reduces to the one nishi2022StochasticLocal derived (their equation (19)). nishi2022StochasticLocal found that the LN test performs better than other tests when $q=0$. To see how the value of $q$ alters the LN test's power properties, we simulate the asymptotic distribution in (ref) by 100,000 replications. The replications are based on $\varepsilon_t \sim \mathrm{i.i.d} \ N(0,1)$, so that $\sigma_\varepsilon^2=1$, $\sigma_\eta^2=2$ and $\psi=\mathbb{E}[\varepsilon_t\eta_t]/(\sigma_\varepsilon\sigma_\eta)=0$. Note that the effect of $q$ on the limit distribution is symmetric when $\psi=0$; that is, $q=\bar{q}$ and $q=-\bar{q}$ produce the identical distribution. This is because $(W_\varepsilon,W_\eta)' \stackrel{d}{=}(-W_\varepsilon,W_\eta)'$ when $\psi=0$, and hence $(J_a,W_\eta)' \stackrel{d}{=} (-J_a,W_\eta)'$ in this case. Thus, in this simulation, and in the replications conducted later where $\psi=0$ holds, we only consider positive values of $q$. Figure (ref) shows the power functions of the LN test for $q=0,1,2,3$ under $a=0$. One noticeable feature is that as $q$ gets larger, the power function gets lower and flatter for $c^2\geq 5$, while the power becomes greater for $c^2\leq5$. Since the power gains over $c^2\leq5$ are obviously outweighed by the power losses over $c^2\geq5$, we conclude that the LN test performs poorly when $q$ is large (in absolute value).\footnote{In our unreported simulations, we also found a similar tendency in the power function of distaso2008Testingunit's LM test, which is based on the assumption of $\varepsilon_t$ and $v_t$ being independent. The results are omitted to save space.}
The LN test's poor performance for large $q$ could be attributed to the endogeneity $\mathbb{E}[y_{t-1}^2\xi_t]=2c\sigma_\varepsilon qT^{-1}\mathbb{E}[y_{t-1}^3] \neq 0$ in (ref). One solution for this endogeneity is to augment model (ref) by adding $y_{t-1}$ as a regressor,
where $\delta_T \coloneqq 2c\sigma_\varepsilon qT^{-1}$ and $\xi^*_t \coloneqq \omega_T^2y_{t-1}^2(v_t^2-1) + 2\omega_Ty_{t-1}(\varepsilon_tv_t - \sigma_{\varepsilon v}) + (\varepsilon_t^2-\sigma_{\varepsilon}^2)$. Because this augmented model is free from the endogeneity, that is, $\mathbb{E}[\xi_t^*] = \mathbb{E}[y_{t-1}\xi_t^*] = \mathbb{E}[y_{t-1}^2\xi_t^*]=0$ under Assumption (ref), we propose using the t-test for $H_0: \omega_T^2=0$ in model (ref). We also propose using the Wald test for $(\delta_T,\omega_T^2)=(0,0)$, because $\omega_T^2=0$ if and only if $(\delta_T,\omega_T^2)=(0,0)$. To express the t and Wald test statistics, first regress $z_t^2(\rho_T)$, $y_{t-1}$ and $y_{t-1}^2$ on a constant:
where $\widetilde{z_t^2}(\rho_T) \coloneqq z_t^2(\rho_T) - T^{-1}\sum_{t=1}^{T}z_t^2(\rho_T)$, and $\widetilde{y_{t-1}}$, $\widetilde{y_{t-1}^2}$ and $\widetilde{\xi_t^*}$ are defined similarly. Let $\widetilde{Z_2}(\rho_T) \coloneqq (\widetilde{z_1^2}(\rho_T),\widetilde{z_2^2}(\rho_T),\ldots,\widetilde{z_T^2}(\rho_T))'$, $\widetilde{X_1} \coloneqq (\widetilde{y_{0}},\widetilde{y_{1}},\ldots,\widetilde{y_{T-1}})'$, $\widetilde{X_2} \coloneqq (\widetilde{y_{0}^2},\widetilde{y_{1}^2},\ldots,\widetilde{y_{T-1}^2})'$, and $\widetilde{\Xi^*} \coloneqq (\widetilde{\xi_1^*},\widetilde{\xi_2^*},\ldots,\widetilde{\xi_T^*})'$. Then, model (ref) can be expressed in matrix notation as
where $\widetilde{X} \coloneqq (\widetilde{X_1},\widetilde{X_2})$ and $\theta_T \coloneqq (\delta_T,\omega_T^2)'$. The t and Wald test statistics are then given by
where $M_1 \coloneqq I_T - \widetilde{X_1}(\widetilde{X_1}'\widetilde{X_1})^{-1}\widetilde{X_1}'$, and $\hat{\sigma}_{\widetilde{\xi^*}}^2(\rho_T)$ and $\hat{\theta}_T(\rho_T)$ are the OLS variance and coefficient estimators of (ref), respectively. We shall call the t and Wald tests in model (ref) augmented t and Wald tests, respectively. The limiting distributions of these augmented test statistics are collected in the following theorem.
There are two points worth mentioning. First, the augmented t test statistic in model (ref) has the asymptotic null distribution independent of $q$. This is a direct result of the augmentation, which is intended to remove the endogeneity from the linearized model (ref). When performing the augmented t test, we regress the endogenous regressor $\widetilde{y_{t-1}^2}$ (and $\widetilde{z_t^2}(\rho_T)$) on $\widetilde{y_{t-1}}$. In the limit, this projection amounts to replacing $\widetilde{J_{a,2}}$ in the limiting distribution of $\mathrm{LN}_T$ with $Q_a$, the residual from the linear projection of $\widetilde{J_{a,2}}$ on $\widetilde{J_{a,1}}$ in the Hilbert space (see, for example, phillips1990AsymptoticProperties).
The second point is that the limiting null distributions of the augmented t and Wald test statistics are standard normal and chi square with 2 degrees of freedom, respectively, when $\psi=\mathbb{E}[\varepsilon_t\eta_t]/(\sigma_\varepsilon \sigma_\eta)=0$. This will be seen immediately upon noting $J_a$ and $W_\eta$ are independent when $\psi=0$ (due to the independence between $W_\varepsilon$ and $W_\eta$), and the limit distributions under $c^2=0$ conditional on $W_\varepsilon$ are standard normal and chi square with 2 degrees of freedom (and so are they unconditionally). Along the lines of this argument, the asymptotic null distribution of the LN test statistic is seen to be standard normal when $\psi=0$.
Unless $\psi=0$, the test statistics discussed thus far have asymptotic null distributions dependent on $\psi$. Actually, this dependence stems from the strong persistence of the regressors $y_{t-1}$ and $y_{t-1}^2$ and the long run endogeneity present in the linearized models (ref) and (ref). To illustrate this, consider model (ref) and the LN test. Under the null of $\omega_T^2=0$, the model reduces to
where $\eta_t=\varepsilon_t^2-\sigma_{\varepsilon}^2$ and $y_{t} = \rho_T y_{t-1} + \varepsilon_t$. This model is free from endogeneity because $\mathbb{E}[\eta_t] = \mathbb{E}[y_{t-1}^2\eta_t]=0$. However, the regressor $y_{t-1}^2$ is, in a sense, “endogenous" in the limit as $T\to\infty$, because of the correlation $\psi$ between its innovation $\varepsilon_t$ and the disturbance $\eta_t=\varepsilon_t^2-\sigma_{\varepsilon}^2$. Indeed, the LN test statistic becomes
where $\widetilde{J_{a,2}}(r)$ ($\approx T^{-1}\widetilde{y_{\lfloor Tr \rfloor}^2}$) is correlated with the differential $dW_\eta(r)$ ($\approx T^{-1/2}\eta_{\lfloor Tr \rfloor}$). This correlation originates from that between $W_\varepsilon$ and $W_\eta$, which is denoted by $\psi$. Therefore, the correlation between the regressor's innovation $\varepsilon_t$ and the regression disturbance $\eta_t$ affects the test statistic's behavior in the limit.
Interestingly, the situation we are in is analogous to the one that has been considered in the literature on predictive regressions. In predictive regressions, a main aim is typically to investigate whether stock returns ($r_t$) can be predicted by another economic or financial variable ($x_t$). To test the predictability of $r_t$, predictive regressions involve regressing $r_t$ on a constant and the lag of $x_t$:
Here, $\beta$ represents the predictability of $r_t$; the stock return is not predictable by $x_t$ if $\beta=0$. In the literature on predictive regressions, it has been well known that the usual t test for the hypothesis $\beta=0$ can be misleading when the regressor $x_t$ is persistent, or has a generating mechanism of the form $x_t = \rho_{x,T}x_{t-1} + u_{x,t}$ with $\rho_{x,T}= 1+a_x/T$, and its innovation $u_{x,t}$ is correlated with the regression disturbance $u_{r,t}$. The problem arising in such a case is that the asymptotic null distribution of the t statistic depends on $\mathrm{Corr}(u_{x,t},u_{r,t})$ and is not standard normal unless $\mathrm{Corr}(u_{x,t},u_{r,t})=0$ campbell2006Efficienttests. Our situation here is essentially the same: the regressor $y_{t-1}^2$ in (ref) is persistent with the mechanism $y_t=\rho_Ty_{t-1}+\varepsilon_t$ and its innovation $\varepsilon_t$ is correlated with the regression disturbance $\eta_t$, which results in the test statistic having the limiting null distribution dependent on the correlation $\psi=\mathrm{Corr}(\varepsilon_t,\eta_t)$.
For the predictive regression model, there is an extensive literature on this problem, and numerous solutions have been proposed campbell2006Efficienttests,phillips2013Predictiveregression,phillips2014ConfidenceIntervals,kostakis2015RobustEconometric. For our case, fortunately, one of those solutions can be applied. Specifically, following campbell2006Efficienttests, we modify the test statistics (LN, augmented t and augmented Wald) so that their asymptotic null distributions are standard normal and chi square with 2 degrees of freedom, irrespective of the value of $\psi$. To explain the idea, take the LN test as an example. Note that under the null of $c^2=0$, $z_t^2(\rho_T)$ in the numerator of the LN test statistic may be asymptotically expressed as
given the distributional equivalence (ref). This observation leads us to propose the following modification to remove the effect of $\psi$:
where $\hat{\psi}_T(\rho_T) \coloneqq T^{-1}\sum_{t=1}^{T}z_t(\rho_T)\{z_t^2(\rho_T) - \hat{\sigma}_{\varepsilon,T}^2(\rho_T)\}/(\hat{\sigma}_{\varepsilon,T}(\rho_T)\hat{\sigma}_{\eta,T}(\rho_T))$ is an estimator of $\psi$.\footnote{In fact, campbell2006Efficienttests proposed the modification based on the optimality argument.} In Appendix B, we show that $\hat{\psi}_T(\rho_T)$ is consistent. Note that $z_t(\rho_T)$ is used as a proxy for $T^{1/2}\times \sigma_{\varepsilon}dW_\varepsilon$. With the replacement given in (ref), the modified LN test statistic is defined by
The modified augmented t and Wald tests are based on the following regression model:
where $\widetilde{Z_2^*}(\rho_T) \coloneqq (\widetilde{z_1^{2*}}(\rho_T),\widetilde{z_2^{2*}}(\rho_T),\ldots,\widetilde{z_T^{2*}}(\rho_T))'$ with $\widetilde{z_t^{2*}}(\rho_T) \coloneqq z_t^{2*}(\rho_T) - T^{-1}\sum_{t=1}^{T}z_t^{2*}(\rho_T)$. The modified augmented test statistics are
where $\hat{\sigma}_{\widetilde{\xi^{**}}}^2(\rho_T)$ and $\hat{\theta}_T^{*}(\rho_T)$ are the OLS variance and coefficient estimators of (ref), respectively.
It should be noticed that the modified test statistics have pivotal asymptotic null distributions (standard normal and chi square with 2 degrees of freedom), thanks to the independence between $J_a$ and $W_1$. Also note that the limiting distributions are unaffected by our modification when $\psi=0$ (cf. Theorems (ref) and (ref)).
Figure (ref) compares the asymptotic power functions of the modified LN, augmented t and augmented Wald tests under $q=0,1,2,3$ and $a=0$ with $\varepsilon_t \sim \mathrm{i.i.d} \ N(0,1)$. When $q=0,1$, the LN test performs best, and the Wald test's power function is slightly below the LN test's. As for the comparison between the augmented t and Wald tests, the Wald test has better power properties than the t test. When $q=2,3$, the Wald test outperforms the LN test (and the augmented t test) with the greater dominance by the Wald test for larger $q$. To translate these results into finite-sample ones, consider for example the case of $T=200$, in which case $-3.76<q<3.76$. Based on the results from the local asymptotic case, it is expected that the augmented Wald test performs more poorly than the LN test if $|q|\leq 1$, or $|\mathrm{Corr}(\varepsilon_t,v_t)|\leq0.266$ in this case, and outperforms the LN test otherwise. To see whether this reasoning gives a good approximation of finite-sample results, we present the size-adjusted power functions for $T=200$ in Figure (ref). The calculation of these power functions are based on 20,000 replications where $\varepsilon_t \sim \mathrm{i.i.d} \ N(0,1)$, $v_t \sim \mathrm{i.i.d} \ N(0,1)$, $\mathrm{Corr}(\varepsilon_t,v_t)\in\{0,0.25,0.5,0.75\}$ and $\omega_T^2\leq0.0177$ so that $c^2\leq50$ under $T=200$. Figure (ref) shows the local asymptotic analysis can well predict the finite-sample results: the LN test performs slightly better than the augmented Wald test when $|\mathrm{Corr}(\varepsilon_t,v_t)|\leq0.25$, but the latter test outperforms the former otherwise.
We also compute the asymptotic power functions under $a\in\{-5,-10\}$, to investigate the effect of the $a$ value on the power properties. The computed power functions are displayed in Figures (ref) and (ref). When $a<0$, the general pattern of the power properties is the same as when $a=0$: the LN test performs best for small $q$, and the Wald test performs best for moderate to large $q$. However, the powers of all the tests we consider get lower as $a$ deviates from 0 (as long as $a<0$).\footnote{According to the results of our simulation studies given later, deviations from 0 by positive $a$ seem to lead to higher powers.} A similar tendency of the power properties of tests for $H_0: \omega_T^2=0$ has been observed through simulation by earlier work such as nagakura2009Testingcoefficienta and horvath2019Testingrandomnessa. Nonetheless, even when $a<0$, the Wald test's power function is increasing in $q$ while the LN test's is decreasing in $q$. This renders the Wald test preferable in empirical applications, where in general the degree of the correlation between the random coefficient and disturbance is unknown to practitioners.
Although all the above results are based on the assumption that the true $\rho_T$ is known so that the tests discussed so far are infeasible, they suggest the potential of the augmentation to enhance the ability to detect the nonzero variance of the autoregressive root.
In this section, we consider model (ref) with the true $\rho_T$ unknown. To deal with the uncertainty about $\rho_T$, we use the OLS estimator $\hat{\rho}_T$ of $\rho_T$, which is defined by $\hat{\rho}_T\coloneqq \sum_{t=1}^{T}y_{t-1}y_t/\sum_{t=1}^{T}y_{t-1}^2$. In Appendix C, we show that $\hat{\rho}_T$ is $T-$consistent, and that other estimators defined in the preceding section such as $\hat{\sigma}_{\varepsilon,T}^2(\rho_T)$ remain consistent if they are computed using $\hat{\rho}_T$ instead of the true $\rho_T$. Given these consistency results, one may expect that the asymptotic behaviors of $\mathrm{LN}_T^*(\hat{\rho}_T)$, $t_{\hat{\omega}_T^2}^*(\hat{\rho}_T)$ and $W_T^*(\hat{\rho}_T)$ are the same as those of $\mathrm{LN}_T^*(\rho_T)$, $t_{\hat{\omega}_T^2}^*(\rho_T)$ and $W_T^*(\rho_T)$. Unfortunately, however, this is not the case. Indeed, it can be shown that the limiting null distributions of the test statistics, if calculated using $\hat{\rho}_T$, are no longer normal or chi square. This is essentially because $a$ is not consistently estimable, which has been well known in the literature campbell2006Efficienttests.
To deal with this problem, following cavanagh1995InferenceModels and campbell2006Efficienttests, we base our tests on the Bonferroni approach by using a confidence interval for $\rho_T$ instead of its point estimate $\hat{\rho}_T$. The Bonferroni-based test consists of two steps: first, construct a confidence interval for $\rho_T$, and then repeat either of the tests considered above using all the hypothetical $\rho_T$ values belonging to the confidence interval. Specifically, letting $S_T(\rho_T)$ denote either the modified LN, augmented t or augmented Wald test statistic (calculated using $\rho_T$), the testing procedure based on the Bonferroni approach is described as follows:
We shall call tests following these two steps Bonferroni tests.
Although the Bonferroni test defined above is a valid test with significance level $\alpha=\alpha_1+\alpha_2$, the test's type 1 errors (dependent on $(a,\psi)$) can be quite smaller than $\alpha$, as pointed out by cavanagh1995InferenceModels and campbell2006Efficienttests. To make the type 1 errors close to the desired nominal level, $\tilde{\alpha}$ (say), they proposed using a pair $(\alpha_1,\alpha_2)$ with which the Bonferroni test's type 1 errors are close to the given $\tilde{\alpha}$. Theoretically, we would find such numerous pairs $(\alpha_1,\alpha_2)$ by changing both the values of $\alpha_1$ and $\alpha_2$. To simplify the searching process, we fix $\alpha_2$ and set $\alpha_2=\tilde{\alpha}$, following campbell2006Efficienttests. In this study, we consider the case $\tilde{\alpha}=\alpha_2=0.05$, giving the Bonferroni test with significance level 0.05. Then, for each $\psi \in (-1,1)$, we numerically find the value of $\alpha_1(\psi)$ such that under the null $P_{a,\psi}\Bigl(\bigcap_{\bar{\rho}_T\in \mathrm{CI}(\alpha_1(\psi))}\Bigl\{S_T(\bar{\rho}_T)>cv_{\alpha_2}\Bigr\}\Bigr)\leq\tilde{\alpha}=0.05$, and this probability is as close to 0.05 as possible, for all $a$ on some grid. The simulation process to find the $\alpha_1(\psi)$ values is described in Appendix A.
Table (ref) displays the significance level $\alpha_1(\psi)$ of the confidence interval for $\rho_T$ along with corresponding intervals of $|\psi|$ values.\footnote{We assign $\alpha_1$ values to each interval of $\psi$ instead of each single $\psi$ (on some grid) for computational ease of the Bonferroni-test.} Given the good asymptotic performance by the modified Wald test, we report in Table (ref) $\alpha_1$ values for the case of $S_T(\rho_T)$ being the modified augmented Wald test statistic. When performing the Bonferroni-Wald test, first estimate $\psi$ by its consistent estimator $\hat{\psi}(\hat{\rho}_T)$, and then select the value of $\alpha_1$ based on Table (ref). For example, if $|\hat{\psi}(\hat{\rho}_T)|=0.17$, $\alpha_1=0.31$ is selected. With the selected $\alpha_1$ value, one can perform the Bonferroni-Wald test following the two-step testing procedure outlined before. The finite-sample performances of the Bonferroni-Wald test are investigated through simulation in the next section.
As described in Appendix A, the simulation exercise to determine the values of $\alpha_1$ is based on the asymptotic procedure. Thus, we need to verify whether the Bonferroni-Wald test we have proposed performs well in finite samples.
Following nagakura2009Testingcoefficienta, we employ three data generating mechanism to evaluate empirical sizes: (i) $\varepsilon_t\sim \mathrm{i.i.d} \ N(0,1)$, so that $\psi=0$, (ii) $\varepsilon_t \sim \mathrm{i.i.d} \ (\chi^2(10)-10)/\sqrt{20}$, so that $\psi=0.5$, and (iii) $\varepsilon_t \sim \mathrm{i.i.d} \ (\chi^2(1)-1)/\sqrt{2}$, so that $\psi=0.756$. For each case, we set $y_0=0$. The sample size we use is $T\in\{200,500,1000\}$, and the number of replications is 5,000. The $\rho_T=\rho$ values are fixed across $T$, and we consider $\rho\in [0.7,1.01]$. The simulation results are collected in Figure (ref). The general pattern is that the empirical rejection rates under $H_0$ are relatively small when $\rho<1$ and tend to be greater than the nominal level 0.05 when $\rho$ is near unity. As for the normal case (Figure (ref)(a)), rejection rates are stable around 0.05 over $\rho\in[0.7,1.01]$, with them approaching 0.05 as $T$ increases. As for the chi square cases (Figure (ref)(b) and (c)), rejection rates stay around the nominal level, but they get farther away from 0.05 when $|\psi|$ is larger, $T$ is smaller and $\rho$ is near 1. In particular, when $\varepsilon_t \sim \mathrm{i.i.d} \ (\chi^2(1)-1)/\sqrt{2}$ and $T=200$, the rejection rates can be as large as 0.09 (around $\rho=1$), although they approach 0.05 as $T$ increases. A similar tendency can be observed in the finite-sample behavior of modified LN tests proposed by nagakura2009Testingcoefficienta according to his simulation results. One possible cause of this phenomenon would be the finite-sample bias involved in the estimation of $\psi$ when the $\psi$ value is large. One may be advised to use a more conservative Bonferroni-Wald test (with smaller $\alpha_1$ values) when the $|\psi|$ estimate is large and the sample size is not so large.
Next, we investigate tests' ability to detect the nonzero variance in the autoregressive root. The simulation design is as follows: $T=200$, $\varepsilon_t\sim \mathrm{i.i.d} \ N(0,1)$ and $v_t \sim \mathrm{i.i.d} \ N(0,1)$ with $\mathrm{Corr}(\varepsilon_t,v_t)\in \{0,0.25,0.5,0.75\}$, $\rho\in\{1.01,1,0.98,0.95\}$, and $\omega^2 \in (0,0.01]$ (corresponding to $c^2 \in (0,28.28]$). The tests we consider here are the Bonferroni-Wald test, the infeasible modified Wald test (calculated using the true $\rho$), and one of the modified LN tests proposed by nagakura2009Testingcoefficienta, which is denoted by $\widetilde{G}_{T,1}$ in his notation. The infeasible Wald test is taken as a benchmark, and thereby we can evaluate the power loss originating from using the confidence interval for $\rho$ to perform the Bonferroni-Wald test. nagakura2009Testingcoefficienta's modified LN test statistic is designed to converge in distribution under $H_0$ to the standard normal irrespective of the $\psi$ value, under the data generating mechanism with $\rho \in (-1,1]$ fixed (i.e., independent of $T$). Because nagakura2009Testingcoefficienta did not show its asymptotic null distribution remains standard normal when $\rho>1$, we do not perform it for the case $\rho=1.01$. We also consider the test proposed recently by horvath2019Testingrandomnessa (hereafter HT). The HT test is a randomized one, and their test statistic $\Theta_{T,R}$ (in their notation) converges in distribution to the chi square with one degree of freedom, for almost all realizations.\footnote{The HT test needs a tuning parameter, $x \in(0,0.5)$, to be performed, and they stated their test's performance is insensitive to the choice of the $x$ value and set $x=0.1$. However, in our simulation, their test's performance is somehow affected by the choice of $x$. Thus, we set $x=0.38$, to obtain simulation results similar to those of HT.} The HT test is not originally proposed under the local-to-unity specification, but it will be informative to practitioners to reveal the performance of the test under this setting.
The simulation results are shown in Figures (ref) through (ref). For the case $\rho=1.01$, where all but the LN test are performed, the infeasible and Bonferroni-Wald tests have good power, and the discrepancy between their power functions is small. The latter result is due to the refinement on the construction of the confidence interval for $\rho$. In contrast, the power function of the HT test stays around the nominal level 0.05 over $\omega^2 \in (0,0.01]$, which implies it has almost no distinguishing power for these alternatives. Indeed, HT conjectured (based on some theoretical analysis) that their test would have no power when $1-\rho_T=O(T^{-1})$ and $\omega_T^2 = O(T^{-1})$, which is of larger magnitude than our local-to-zero variance $\omega_T^2=O(T^{-3/2})$.\footnote{nishi2022StochasticLocal showed that under the STUR specification, the LN and some other tests are consistent when $\omega_T^2=O(T^{-1})$, from which they concluded that this case should be regarded as capturing stochastic “moderate" departures from a unit root rather than local departures.} Our simulation results corroborate their statement.
Turning to the cases $\rho \leq 1$, it is noticeable that for each $\rho$, the LN test performs well for small values of $\mathrm{Corr}(\varepsilon_t,v_t)$, but Wald type tests outperform the LN test otherwise. In particular, the greater value of $\mathrm{Corr}(\varepsilon_t,v_t)$ leads to the greater dominance by the Wald type tests. The power function of HT test, again, stays around the nominal level for these cases. Overall, the Bonferroni-Wald test performs better than the LN and HT tests for moderate to large values of $\mathrm{Corr}(\varepsilon_t,v_t)$ (irrespective of the $\rho$ value) and performs almost as efficiently as its ideal, infeasible counterpart.
In this section, we apply the Bonferroni-Wald test along with the LN and HT tests\footnote{Based on our simulation results, we set the tuning parameter $x=0.38$ for the HT test.} to several U.S. macroeconomic and financial time series, following hill2014UnifiedIntervalb and HT. The dataset includes CPI, real GDP, industrial production, M2, S&P500, the 3 month Treasury bill rate and the unemployment rate. We take the logarithm of the first 5 series before detrending all the series. All data have been extracted from Federal Reserve Economic Data. The data description and testing results are displayed in Tables (ref) and (ref), respectively.
Before discussing main results, it should be recalled from our simulation results that the Bonferroni-Wald test can be oversized when $T$ is not large, $\psi$ is large and $\rho$ is near unity (see Section 4). Thus, we need to check whether all the three conditions hold for our series or not. According to Table (ref), all the $\rho$ estimates are near unity, as expected. We also have obtained moderate estimates of $\psi$ for the CPI and S&P500 series, but the sample sizes for them are large enough that it is unlikely the Bonferroni-Wald test is oversized for these series. As for GDP, the sample size is $T=292$ but the $\psi$ estimate is near zero, and hence we expect the Bonferroni-Wald test is not severely oversized for this series. Overall, all the three conditions for the potential oversize problem do not jointly hold.
For GDP and T-bill rate, we have obtained consistent results from all the three tests: the null of $H_0:\omega^2=0$ is not rejected for these series. This coincidence could be viewed as evidence for nonrandomness of the autoregressive root for these series. As for CPI, industrial production and S&P500, it is observed that the Bonferroni-Wald and LN tests reject the null while the HT test does not. This result will be attributed to the fact that the former two tests are much more powerful than the latter one. Finally, as for M2 and unemployment rate, only the Bonferroni-Wald test rejects the null. This outcome will be due to the fact that the Bonferroni-Wald test tends to be the most powerful among the three tests.
In Table (ref), we also present the $\omega^2$ estimates proposed by horvath2019Testingrandomnessa (denoted by $\hat{\omega}^2_\mathrm{HT}$). They showed this estimator is consistent under the non-local RCA models (with $\rho$ and $\omega^2$ fixed). The values of $\hat{\omega}^2_\mathrm{HT}$ seem to support our testing results. For instance, we have a negative $\hat{\omega}^2_\mathrm{HT}$ for T-bill rate, for which none of the tests reject the null. Moreover, for the other series, $\hat{\omega}^2_\mathrm{HT}$ takes values around $10^{-4}$ to $10^{-3}$, magnitudes of the coefficient randomness with which the HT test tends to be powerless under the sample size given in Table (ref). For example, for M2, the sample size is $T=2044$, and $\hat{\omega}^2_\mathrm{HT}$ is $3.12\times10^{-4}$, translating into $\hat{c}^2=\hat{\omega}^2_\mathrm{HT}\times T^{3/2}=28.83$ estimate of the localizing coefficient $c^2$. According to our simulation results, the HT test has almost no power against the alternative of this magnitude, hence the testing result given in Table (ref).
Given these findings, our empirical application illustrates the merit in choosing our Bonferroni-Wald test over existing ones.
Given the results of empirical analyses conducted by earlier studies, the local-to-unity RCA models, which extend the STUR modelling, are empirically relevant. Under this setting, we can analyze the effect of the correlation between the random coefficient and disturbance on the power properties of tests for coefficient randomness. Theoretical and simulation analyses reveal that tests proposed by earlier studies can perform poorly when the degree of the correlation is moderate to large and the coefficient randomness is local to zero, while the augmented-Wald test we have proposed performs well even in such cases. Our test is also independent of the nuisance parameter $\psi$, the correlation between the disturbance and its square, so that it is implementable without the knowledge about the value of $\psi$. To deal with the uncertainty about the mean $\rho$ of the autoregressive root, we have proposed using a confidence interval for $\rho$, leading to the Bonferroni-Wald test, where the significance level for the confidence interval is selected according to the value of the $\psi$ estimate. Embedding this selection process into the Bonferroni-Wald test helps stabilize the test's size and improve the test's power.
Several directions for future research are possible. First, from the similarity in the construction of test statistics between our model and predictive regressions, it is expected that the theory developed by numerous studies on predictive regressions can be applied to testing for coefficient randomness in local-to-unity autoregressions. For example phillips2013Predictiveregression,phillips2016Robusteconometric proposed the use of the so-called IVX procedure for predictability testing kostakis2015RobustEconometric. This procedure leads to good size and power properties and also requires less computational burden for implementation than the Bonferroni approach employed by campbell2006Efficienttests. The use of the IVX approach may facilitate testing for coefficient randomness in local-to-unity autoregressions. The analysis of tests' performance when $\rho$ is distant from unity will also be of interest. In such a case, more preferable tests might be available than the Bonferroni-Wald test, which is based on the local-to-unity asymptotics.