Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
80,409 characters · 7 sections · 69 citation commands
Robust Inference on Infinite and Growing Dimensional Time Series Regression
\newgeometry{left=1.25in, right=1.25in,bottom=1.25in,top=1.25in} \setstretch{1.25}
\setstretch{1.5}
\allowdisplaybreaks
\allowdisplaybreaks
This paper develops asymptotically valid tests for inference on infinite-order and growing dimensional time series regression models, revealing the presence of an hitherto undetected nonlinear serial dependence or high-order long-run variance (HLV) factor. This factor depends on the model error and regressors in a nonlinear fashion, and can appear in limit distributions when the data exhibit dependence and the number of restrictions grows. Chow tests and tests for linear restrictions are both covered. Our theory, simulations and empirical results show the deleterious effect of ignoring the HLV term, and we propose a testing procedure that is robust to its presence. This is shown to possess desirable finite sample properties. While the HLV factor is revealed by our increasing dimension asymptotics, it can contaminate inference even in multiple regressions with a moderate number of covariates. Such specifications are ubiquitous in practice. Thus, the findings and recommendations of this paper are important for practitioners wishing to make correct inferences when data are dependent.
Models of infinite or growing dimension have been widely studied in the recent econometric literature, reflecting modern applications with rich sets of variables. For example, the asset pricing literature has suggested hundreds of potential risk factors to explain returns, see feng2019taming. With a larger number of observations accumulating over time, it is natural to include more of these variables as covariates even without resorting to penalized estimation methods. In fact, an attitude that permits the number of covariates to grow as a function of sample size is tacitly adopted in the literature. In a survey, koenker1988asymptotic observed that the number of regressors in empirical work increases as the sample size $n$ increases, roughly like $n^{1/4}$, suggesting that practitioners implicitly treat model complexity as a function of sample size. Finally, nonparametric methods such as series estimation have found wide applications in the economics and finance literature, see e.g. jorda2005estimation, chen2007, Chen2015. These methods involve the approximation of an infinite-order model with a sequence of growing dimensional models. Taken together, this proliferation of models highlights the importance of developing appropriate techniques for their study.
Our approach is to develop tests for null hypotheses that involve a growing number of restrictions $p$ in time series regression, with $p$ increasing slower than sample size. As a leading example, we consider the Chow test, due to Chow1960, to test for a structural break at a prescribed time. This has the advantage of being a simple exclusion restriction test with wide applicability. After examining the key issues in this simple context, we present results for the testing of general linear restrictions. This extends specification tests with slowly growing $p$, see e.g. Hong1995 and Gupta2018, to time series regression. However, our testing problem is distinct from the so-called `many restrictions' setting in e.g. Calhoun2011, Anatolyev2012,Anatolyev2019, kline2020leave, amongst others, where the number of restrictions grows proportionally to sample size.
We derive the asymptotic distribution of the Chow test Wald statistic centered by $ p $ and normalized by $ \sqrt{2p} $. This yields asymptotic normality with an unknown asymptotic variance $ \mathcal{V} $, which we term the HLV, provided that $ p $ meets certain growth conditions. The HLV factor $ \mathcal{V} $ captures high-order autocovariances of the regressors and disturbances, echoing the long-run variance that appears in fixed dimensional time series regression, and vanishes under simplifying assumptions that remove these high-order autocovariances. The new HLV factor $ \mathcal{V} $ does not appear in fixed $ p $ asymptotic regimes, nor does it appear in the independent data setting of Hong1995, who use the same transformation and obtain asymptotic standard normality.
We robustify the Chow test against the HLV by a random scaling because of numerical evidence that asymptotic normal inference based on consistent estimators of HLV performs poorly in finite samples, reported in Section (ref). The random scaling is motivated by heteroskedasticity autocorrelation robust (HAR) inference, see e.g. kiefer2002heteroskedasticity, Sun2014, and lazarus2018har, just to name a few. The resulting asymptotic distribution is pivotal. However, unlike conventional HAR inference, where the standard Wiener process characterizes mixed normality, our limit distribution is represented by two dependent centered Gaussian processes $ W(r) $ and $ \bar{W}(r) $ such that $ EW(r)^2 = r^2 $ and $ E\bar{W}(r)^2 = (1-r)^2 $. Although pivotal, the asymptotic distribution depends on the location of the hypothesized break date and thus we provide R code to compute the p-values. Similarly, we robustify the general linear restrictions Wald test to the HLV and provide suitable R code.
Finite sample bias in the Wald statistic, or in quadratic statistics more generally, is a serious issue when $ p $ is large. See e.g. kline2020leave for more discussion and a bias correction proposal that works well even when $ p $ is proportional to the sample size but under independent sampling. Our simulations document that the problem is even worse in time series regression. Thus, we propose a bootstrap bias correction which imposes the null hypothesis in the resampling so as not to sacrifice power unduly. Even a small number of bootstrap iterations appear sufficient to reduce the bias, making computation easily manageable. Based on these findings, we recommend a bias-corrected and HLV-robust test to practitioners.
In simulations for a range of settings across regression with many covariates, long AR fits and sieve regression, we demonstrate that our statistic exhibits excellent size control without sacrificing power excessively. Failure to correct for the HLV can seriously affect inference, in general leading to over-rejection and often severely so. Such a pattern is shown to persist for the two types of tests that we provide: Chow tests and exclusion restrictions. In an empirical example based on Hamilton2003,Hamilton2009, we show that using our bias-corrected and HLV-robust tests can yield inferences that lead to new conclusions when considering the relation between oil prices and economic activity.
The paper is organized as follows. Section (ref) introduces the model and the Chow test, along with some basic assumptions and examples. In Section (ref) we provide an asymptotic theory while Section (ref) introduces our HLV-robust and bias-corrected test statistic. The testing of general linear restrictions is covered in Section (ref). Section (ref) contains a Monte Carlo study of finite sample performance, and Section (ref) demonstrates our test with real data. All the proofs of theorems and lemmas are collected in two further appendices, the second of which is available online. Throughout the paper, cross-referenced items prefixed with `S' can be found in this online supplementary appendix. An R-package to reproduce the simulations and empirical example is available in the replication files.
We consider the issue of testing for a structural break at a known point in the conditional mean function of $y_{t}$ given the information available up to $ t-1$, ie $E\left( y_{t}|\mathcal{F}_{t-1}\right) , $ where $\mathcal{F}_{t-1}$ denotes the filtration up to time $t-1$. In nonparametric regression, $\mathcal{F}_{t-1}$ typically consists of a finite number of observable covariates $z_{t}$. In the context of the infinite order autoregressive AR$\left( \infty \right) $ model, $\mathcal{F}_{t-1}$ is the collection of all the lagged dependent variables, $\left\{y_{t-j}\right\}_{j\geq 1}$. Alternatively, it can be viewed as a genuine high-dimensional regression model which may contain an infinite number of covariates and their lags. We allow for array structure but we do not introduce further notation to denote it unless necessary.
Given a sample of size $n$, we estimate the unknown regression function via a growing-dimensional (or truncated) linear regression
where $x_{nt}$ and $\beta _{n}$ are $p$-dimensional vectors and $ p\rightarrow \infty $ as $n\rightarrow \infty $ to estimate $E\left( y_{t}| \mathcal{F}_{t-1}\right) $ consistently. To be more precise, let \[ \varepsilon _{t}=y_{t}-E\left( y_{t}|\mathcal{F}_{t-1}\right) , \] $\beta _{n}$ be the best linear predictor of $y_{t}$ given $x_{nt}$, and $r_{nt}=E\left( y_{t}|\mathcal{F}_{t-1}\right) -x_{nt}^{\prime }\beta _{n}$. Then, $ e_{nt}=r_{nt}+\varepsilon _{t}$. Throughout the paper, let $C$ ($c$) denote a generic positive and finite constant, arbitrarily large (small) but independent of $n$, and `a.s.' stand for `almost surely'. Introduce the following assumptions:
The theory presented in the paper may not hold if in fact we only have $E\left(x_{nt}\varepsilon_t\right)=0$ as the long run variance of $x_{nt}\varepsilon_t$ will then appear in the type of quadratic statistics that we consider.
We discuss this assumption on the negligibility of the approximation error in more detail in Section (ref), where specific examples are introduced. The subscript $n$ will now usually be dropped, although we will emphasize this occasionally to remind the reader of the $n$-dependence of certain quantities.
Introduce a potential structural break for these models at a given time, say $ t=[n \gamma] $, $\gamma \in \Gamma \subset (0,1)$, with $\Gamma$ compact and $[\cdot]$ denoting the integer part of the argument. That is, $ \beta = {\beta _{1}\ \ {\text{if }}t/n\leq \gamma }$ and $\beta={\beta _{2}\ \ {\text{if }}t/n>\gamma }. $ We write the model as
where $\delta _{1}=\beta _{1}$, $\delta _{2}=\beta _{2}-\beta _{1}$, and $1\left\{\cdot\right\}$ denotes the indicator function. Consider the Wald test for the exclusion restriction $ \delta_{2} = 0 $, namely the Chow test for the presence of a structural break at a known date.
Let $ \hat{\delta }\left( \gamma \right) $ and $ \hat{e}_t\left( \gamma \right) $ denote the OLS estimate and the OLS residuals, respectively, and $x_{t}\left( \gamma \right) :=\left(x_{t}^{\prime },x_{t}^{\prime }1\left\{ t/n>\gamma \right\} \right) ^{\prime}$. Also, let $ \hat{M}\left( \gamma \right) ={n}^{-1}\sum_{t=1}^{n}x_{t}\left( \gamma \right) x_{t}\left( \gamma \right) ^{\prime }, $ and $\hat{\Omega}\left( \gamma \right) $ denote an estimator of $ E\varepsilon _{t}^{2}x_{t}\left( \gamma \right) x_{t}^{\prime }\left( \gamma \right) $. For instance, $\hat{\Omega}\left( \gamma \right) $ can be set as $n^{-1}\sum_{t=1}^{n}x_{t}\left( \gamma \right) x_{t}\left( \gamma \right) ^{\prime }\hat{e}_{t}\left( \gamma \right) ^{2}$ (the Eicker-White formula) or, assuming conditional homoskedasticity, it can be $\hat{\sigma}\left( \gamma \right) ^{2} \hat{M}\left( \gamma \right) $, where $\hat{\sigma}^{2}\left( \gamma \right) =n^{-1}\sum_{t=1}^{n}\hat{e}_{t}\left( \gamma \right) ^{2}$. The choice depends on the case being considered. Then, the Wald statistic for the familiar Chow test is defined as
where $R=\left( 0_{p\times p}:I_{p}\right) $ is a selection matrix.
When the dimension $ p $ of $ x_t $ grows with the sample size $ n $, the Wald statistic diverges as it is approximately chi-squared distributed with $p$ degrees of freedom. Thus, a conventional approach, as used e.g. by DeJong1994 and Hong1995 in the cross-sectional (independent data) framework is to introduce a new centering and scaling to define
since the mean and variance of a chi-square distribution with $ p $ degrees of freedom are $ p $ and $ 2p $, respectively. Furthermore, it has been established that the standard normal approximation of $ \mathcal{Q}_n $ is valid in their settings. Subsequent sections investigate how this conventional approach fails in the context of growing or infinite dimensional time series models, mirroring the failure of time series inference procedures without heteroskedasticity and autocorrelation correction or robustification.
This section provides the asymptotic distribution of the Chow test statistic under the null and also shows that the statistic has non-trivial power against local alternatives at an appropriate nonparametric rate.
There has been some recent interest in the so-called many regressor setting where $p $ is allowed to be proportional to $n $, see e.g. cattaneo2018inference and kline2020leave. We do not permit such a large $p $ as our hypothesis of interest concerns a $p $-dimensional restriction and the design matrix of time series data faces more difficulties in satisfying the rank condition. In this regard, Chen2001 provide an interesting example from an ANOVA design where the weak convergence of the empirical distribution of residuals from the linear regression with growing dimension fails when the dimension $p$ is of order $n^{1/3}$. They compare various growth conditions for $p$ in the literature and conclude that $p^{3}\log ^{2}p=o\left( n\right) $ is nearly necessary for a general stochastic design. Heuristically, a\ hypothesis represented through the empirical distribution function imposes an infinite number of restrictions, like our structural break testing also does, and valid testing of such a hypothesis demands a tighter control on the growth rate of $p$.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{Asymptotic Null Distribution}
Define $\left\Vert A\right\Vert =\left\{ \overline{\lambda }(A^{\prime }A)\right\} ^{\frac{1}{2}}$ for a generic matrix $A$, where $\underline{ \lambda }$ (respectively $\overline{\lambda }$) denotes the smallest (largest) eigenvalue of a symmetric nonnegative definite matrix. Any limit stated as `$n\rightarrow\infty$' is taken as both $ n $ and $ p $ grow to infinity simultaneously unless specified otherwise. We also introduce the $p\times p$ non-stochastic matrix sequences $M$ and $ \Omega $ and define
Several factors determine the bound $\varkappa _{p}$ for nonparametric series regression. It is proportional to $\sqrt{p/n}$ or $ p/\sqrt{n} $ up to logarithmic factors with iid data, depending on the choice of basis functions. For dependent data, the mixing decay rate also contributes to $ \varkappa_{p} $. The exact rate $v_{p}$ depends on a particular example. We formally introduce our examples of multiple linear regression, AR($\infty$) and nonparametric sieve regression in Section (ref). Primitive conditions and expressions for $\varkappa _{p}$ and $v_p$ are given in Propositions (ref) and (ref) in Appendix (ref), using the results of peligrad1982invariance, Newey1997, gonccalves2007bootstrapping and Chen2015.
Recall that the eigenvalues of the Kronecker product of two symmetric matrices are the products of their eigenvalues, and $ \gamma $ is bounded away from zero and one. Thus, $M(\gamma)$ and $\Omega(\gamma)$ inherit the eigenvalue restrictions on $M$ and $\Omega$ in Assumption (ref) $(ii)$ and $(iii)$, up to positive constants.
To develop the distributional limit of ${\mathcal{Q}}_{n}(\gamma )$ where both $ n $ and $ p $ diverge simultaneously, we introduce more conditions. Now, for convenience we let $\xi_{t}=\Omega^{-1/2}x_{t}\varepsilon_{t}$, $ \mathcal{G}_t $ denote a filtration for $ \xi_{t} $, $\Upsilon_t=E\left(\xi_{t} \xi_{t}'|\mathcal{G}_{t-1}\right)$, and $\Xi_s=\sum_{t_1=1}^{s-1}\sum_{t_2=1}^{s-1} \xi_{t_1} \xi_{t_2}'$. The filtration $ \mathcal{G}_t $ need not be $ \mathcal{F}_t $ but a simpler one as long as it makes $ \xi_t $ a mds. Indeed, some conditions may be easier to verify under simpler filtrations. The next assumption introduces the HLV factor $\mathcal{V}$ formally.
The first condition can be met if moments of $ \overline{\lambda}\left(\Upsilon_t\right) $ of an order higher than $ 1/\nu $ are bounded for all $ t $. The restriction on the summability rate of $\mathrm{cov}\left(\mathrm{tr}\left(\Upsilon_{t}\Xi_{t}\right),\mathrm{tr}\left(\Upsilon_{s}\Xi_{s}\right)\right)$ is related to the dimension $p$. To gain some insight, consider the case where the conditional moment $\Upsilon_{t}$ is homogeneous, so that $\Upsilon_{t}=I_{p}$ for all $t$. Then, some tedious algebra yields that $\mathrm{cov}\left(\mathrm{tr}\left(\Upsilon_{t}\Xi_{t}\right),\mathrm{tr}\left(\Upsilon_{s}\Xi_{s}\right)\right)= O(n^2 p) $ uniformly over all $s,t$ with $s < t $. This implies that the double sum of the covariances is $O\left(n^4p\right)$ and thus meets the required condition as $ p\to \infty $. Our assumption says that more generally this double sum over covariances must be $o\left(n^4p^2\right)$ as $n,p\rightarrow\infty$.
Also, note that under the special case where $\left\{ x_{t}\varepsilon _{t}\right\} $ is an iid sequence, we have $ {\mathcal{V}}=\lim_{m,p\rightarrow \infty } {m}^{-1}\sum_{t=1}^{m-1} p^{-1}{\mathrm{tr}}E\left( \xi_{m}\xi_{m}^{\prime }\right) E\left( \xi_{t}\xi_{t}^{\prime }\right) =1, $ thus ${\mathcal{V}}$ is an extra factor that appears in the limit due to nonlinear dependence in the data. In particular, it captures a {high-order serial correlation} of $\xi_{t}$, while $\xi_{t}$ itself does not have serial correlation since it is a martingale difference sequence.
For mean zero random variables $a_{1i},a_{2j},a_{3k},a_{4l}$, let $\mathrm{cum}_{ijkl}\left(a_{1i},a_{2j},a_{3k},a_{4l}\right)$ denote the fourth cumulant.
This assumption controls the temporal dependence in $\left\{x_t\varepsilon_t\right\}$ and is discussed in Andrews1991a, for example, wherein sufficient conditions for it to hold are also provided. The following theorem establishes distributional convergence for a given $ \gamma $.
Theorem (ref) highlights the distinctive feature of testing growing number of restrictions in time series regressions. Unlike the independent cross-sectional case, we have to robustify the test against the HLV term $ \mathcal{V} $. The provenance of this term can be illustrated by some formulae, details of which are contained in the full proofs. These proofs first establish (Theorem (ref)) that
where
and $\psi_t(\gamma)=\left({1}\left(t/n\leq \gamma\right)-\gamma\right)/\sqrt{\gamma(1-\gamma)}$. Also note that $ n^{-1}\sum_{t=1}^n\psi_t(\gamma)^2 \to 1 $. Thus, just as for the familiar Wald statistic, we have a quadratic form structure for $\mathcal{R}_n(\gamma)$. When $p$ is fixed and there is no approximation error, we note that ((ref)) has also been established by Andrews1993, Cho2017 and Sun2021.
This then yields the approximation
where \[ \mathcal{S}_n(\gamma)=\frac{2}{\sqrt{2p}}\frac{1}{n}\sum_{t=2}^n\psi_t(\gamma)\xi_t'\sum_{s<t}\psi_s(\gamma)\xi_s=\frac{\sqrt{2}}{\sqrt{n}}\sum_{t=2}^n v_t(\gamma), \] say, by Lemma (ref). Then $\mathcal{V}=\lim_{n,p\rightarrow\infty}2n^{-1}\sum_{s,t=2}^n cov \left(v_s(\gamma),v_t(\gamma)\right)$, i.e. the limiting variance of $\mathcal{S}_n(\gamma)$. Note that the $v_t(\gamma)$ are defined as products of terms of the type $ x_t \varepsilon_t $ and the cumulative sum of their lags, implying that the variances of the $ v_t(\gamma) $ themselves contain high-order covariance terms. This explains why we call $ \mathcal{V} $ a HLV despite the mds property of the $ v_t (\gamma)$, which implies that $\{v_t(\gamma)\}$ is uncorrelated.
The next section establishes that the test based on $ \mathcal{Q}_n $ has nontrivial local power under suitable sequences of local alternatives, following which we study more detailed characteristics of $ \mathcal{V} $ and develop a HLV-robust test.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{Local Alternatives}
We consider a sequence of local alternatives that converge to the null at $ p^{1/4}/\sqrt{n}$-rate to study the local power properties of the test. This is slower than the usual $1/\sqrt{n}$ parametric rate and has been employed by a number of other authors, e.g. DeJong1994, Hong1995, Gupta2018. It is a cost of the growing-dimensional nature of the problem. Our sequence of local alternatives is:
where $\tau $ is a unit length $p\times 1$ vector.
Note that the noncentrality term is positive, implying nontrivial power of the test since the critical region is formed by $ \mathcal{Q}_n (\gamma) $ being greater than equal to a critical value. Also, $ \left\vert \tau ^{\prime }M\Omega ^{-1}M\tau \right\vert \leq \left\Vert \tau \right\Vert \left\Vert M\right\Vert ^{2}\left\Vert \Omega ^{-1}\right\Vert =\overline{\lambda }(M)^{2}/\underline{\lambda }(\Omega ) $ for any $n$. As Assumption (ref) assumes that the numerator $ \bar{\lambda}(M) $ is bounded but the denominator may not be bounded away from zero, $\tau _{\infty }=\lim_{n\rightarrow \infty }\tau ^{\prime}M\Omega ^{-1}M\tau$ may diverge to positive infinity to imply more power.
In this section we provide a detailed study of the HLV $\mathcal{V}$ that our analysis has discovered. In particular, we present some alternative representations of $\mathcal{V}$ that shed more light on its structure.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{Discussion} \sloppy We first examine the relevance of $\mathcal{V}$. Specifically, we analyze the `pre-limiting' quantity $ \mathcal{V}_n =2var\left(n^{-1}\sum_{t=2}^{n}\xi_{t}'p^{-1/2}\text{\ensuremath{\sum_{s=1}^{t-1}\xi_{s}}}\right)$. This can be rewritten as \[ \mathcal{V}_n =2n^{-1}\sum_{i=1}^{n-1}\left(\gamma\left(i,0\right)(n-i)/{n}+2\sum_{j=1}^{n-i}\gamma\left(i,j\right)(n-i-j)/{n}\right), \] where \[ \gamma\left(i,j\right)=p^{-1}E\left(\xi_{t}'\xi_{t-i}\xi_{t}'\xi_{t-i-j}\right)=p^{-1}E\left(x_{t}'\Omega^{-1}\text{\ensuremath{x_{t-i}}}x_{t}'\Omega^{-1}\text{\ensuremath{x_{t-i-j}\varepsilon_{t}^2\varepsilon_{t-i}\varepsilon_{t-i-j}}}\right). \] This is a high-order autocovariance and captures a nonlinear serial dependence in the sequence $ x_t \varepsilon_t $, which disappears entirely for $ j>0 $ in independent cross sectional data. We encounter $\mathcal{V}_n\rightarrow \mathcal{V}\neq 1$ when $n^{-1}\sum_{i=1}^{n-1}\sum_{j=1}^{n-i}\gamma\left(i,j\right)(n-i-j)/{n}$ has a nonzero limit, with terms arising that are fourth-order cross-moments of the $\varepsilon_{t}$ and $ x_t $. Thus, the behaviour of such cross-moments is the key to obtaining non-unity $\mathcal{V}$. Robinson1991, studying time series specification testing, encountered a term of the form $ E\left(\varepsilon_t^2 \varepsilon_{t-i} \varepsilon_{t-i-j} \right)$, somewhat different from ours albeit also of cross-moment type, but imposed conditions that nullify it when $ j>0 $.
The Wald statistic is a quadratic form in the moment process. To establish the limit of the Wald statistic when the number of variables (i.e., the number of moments) grows with the sample size, we need to account for the variance of the quadratic form, hence the appearance of fourth-order dependence of a certain type in the moment process. A form of fourth-order dependence has also been encountered in HAR testing, see e.g. Lobato2002. In Section (ref), we present some figures to show how $\mathcal{V}$ can vary for various designs and deviate significantly from unity.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{HLV-Robust Test Statistic}
This section propose a random scaling approach to robustify our test statistic against the unknown HLV term $\mathcal{V}$. We opt for this because our numerical experiments in Section (ref) (specifically Figure (ref) and its discussion therein) reveal poor finite sample performance of the standard sample variance of $ q_t = \left( np\right)^{-1/2} x_{t}'\hat{\Omega}^{-1} \hat{e}_{t} \text{\ensuremath{\sum_{s=1}^{t-1}x_{s}\hat{e}_{s}}} $. The presence of the cumulative sums $ \sum_{s=1}^{t-1} $ and the estimated quantities $ \hat{\Omega}^{-1}$ and $ \hat{e}_{s} $ in the construction of $ q_t $ are likely to contribute to poor finite sample behavior of its sample variance.
The random scaling approach has been employed when consistent estimators of the asymptotic variance perform poorly in finite samples. For instance, heteroskedasticity and autocorrelation consistent (HAC) estimators, e.g. newey1987simple, Andrews1991a, to name but two examples, have been followed by the fixed-bandwidth kernel approach to obtain an asymptotically pivotal and mixed-normal test, see e.g. kiefer2000simple and lazarus2018har for a recent review. In the machine learning literature, Lee_Liao_Seo_Shin_2022 also employ the random scaling approach for computationally efficient on-line inference based on the stochastic gradient descent algorithm.
While a simple random scaling can be implemented by the integral of the square of the partial sum process of centered $ q_t $, that is, $ n^{-2} \sum_{t=1}^n \sum_{s=1}^{t} (q_t- n^{-1}\sum_{i=1}^n q_i )^2$, we present a class of more general random scaling methods following the heteroskedasticity and autocorrelation robust (HAR) inference literature. We note that the resulting pivotal distributions differ from the HAR literature, however.
Introduce a kernel function $ k(\cdot) $ that meets the following conditions.
Since $\varepsilon_{t}$ and $\Omega$ are not directly observable in practice, we replace them with the least squares estimates as in Section (ref) and introduce $ q_t = \left( np\right)^{-1/2} x_{t}'\hat{\Omega}^{-1} \hat{e}_{t} \text{\ensuremath{\sum_{s=1}^{t-1}x_{s}\hat{e}_{s}}} $ and its demeaned version, $\bar{q}_{t}= q_t - n^{-1}\sum_{t=2}^{n} q_t$. Then, define a feasible estimate of $\mathcal{V}$ by
Thus, we have a seemingly long-run variance estimate, analogous to traditional HAC/HAR inference, of a nonlinear transformation of the primitive variables.
The choice of bandwidth $ b $ has been a topic of much discussion in the HAC literature. Since $\mathcal{V}$ captures high-order autocovariances in the growing dimensional vector $x_{t}\varepsilon_{t}$, the finite sample variation in the estimate $\hat{\mathcal{V}}$ is generally larger than in more familiar long-run variances, and the moment condition is more expensive. Motivated by this, we follow a fixed bandwidth approach, as in Sun2014.
Our estimator is based on the weighting function $ K_{h}\left(r,s\right) = k\left(h\left(r-s\right)\right)$, where $ h=1/b $. We present numerical results in this paper with $k\left(u\right)=\left(1-\left|u\right|\right)^{h}1\left\{ \left|u\right|<1\right\} $, employing the Bartlett kernel case with $h=1$. Sun2014 terms this the sharp kernel estimator. Other options include the steep quadratic kernel estimator and the orthonormal series estimator with $K$ basis function, of which Sun2014 contains a more detailed discussion. Sun2014 also shows that the centering in $ \bar{q}_t $ can be conveniently represented through a centered version of $ K_h(\cdot) $, that is, $K_{h}^{*}\left(r,s\right) =K_{h}\left(r,s\right)-\int_{0}^{1}K_{h}\left(\tau,s\right)d\tau-\int_{0}^{1}K_{h}\left(r,\tau\right)d\tau +\int_{0}^{1}\int_{0}^{1}K_{h}\left(\tau_{1},\tau_{2}\right)d\tau_{1}d\tau_{2} $.
Building on the representation in Lemma 1 of Sun2014, where the estimate $ \hat{\mathcal{V}} $ is not consistent, we characterize the joint weak limit of $ \hat{\mathcal{V}} $ and $ \mathcal{Q}_n (\gamma) $. For real numbers $a$ and $b$, let $a\vee b$ ($a\wedge b$) denote their maximum (minimum), and introduce a process
where $\left (W\left (r \right ),\bar {W}\left (r \right )\right )^{\prime }$, $r\in[0,1]$, is a bivariate Gaussian process that does not depend on any model parameters including the break point $\gamma$, and has covariance kernel
For any given $\gamma\in\Gamma$, the marginal distribution of $ \mathcal{Q}(\gamma ) $ is standard normal. Thus, the conclusion of Theorem (ref) can be expressed as $\mathcal{Q}_n(\gamma)\overset{d}{\rightarrow}\sqrt{\mathcal{V}}\mathcal{Q}(\gamma)$, pointwise in $\gamma\in\Gamma$. By taking a suitable ratio, we obtain a pivotal variable as in the following theorem, which is the basis of our test statistic.
The asymptotic null distribution is mixed normal and pivotal. The critical values can be tabulated for each $ \gamma $ via Monte Carlo simulation and the replication files provide R code. Note that the same Gaussian process $W(\cdot)$ occurs in both the limiting numerator and denominator, and this process is different from the Brownian motion in Sun2014. In fact, it can be represented by the partial sum of $ \sqrt{t/n}$ times an iid normal sequence. Since the limit also involves another variable $ \bar{W}(\cdot) $, the critical values will be different from those previously tabulated in the literature.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{Bias Correction} The degrees of freedom $p$ provide a correct centering for $ W_n (\gamma) $ in first order asymptotic analysis. However, in the finite sample experiments given in Section (ref), e.g. Figure (ref) and Figure (ref), we find that the bias in $ \mathcal{Q}_n (\gamma) $ gets bigger for typical values of $ p $ in nonparametric regression. Therefore, we propose a bootstrap bias correction of $ \mathcal{Q}_n (\gamma) $. To estimate the bias, we implement the null-imposed wild bootstrap by generating
where $ u_t $ is an iid sequence of centered and normalized variables, e.g. the Rademacher variables, to compute $ \mathcal{Q}_n^{\star} (\gamma) $. It is worthwhile to note that the bootstrap DGP (ref) imposes the null hypothesis $ \delta_2 =0, $ so as not to sacrifice the power of the test. See also gonccalves2007bootstrapping for a thorough discussion on the wild bootstrap for infinite order autoregression. Iterating this $ B $ times, we obtain $ \bar{\mathcal{Q}}_n^{\star} (\gamma) = B^{-1} \sum_j^{B} \mathcal{Q}_n^{\star, j} (\gamma) $, the bootstrap estimate of the bias. In our experiment, $ B=200 $ suffices and thus the bootstrap is not computationally expensive. Therefore, we suggest the following bias corrected test statistic:
The numerical experiments in Section (ref) show that the bootstrap bias corrected test controls the type I error reliably without sacrificing power unduly.\footnote{It is worth noting that the wild bootstrap may not be valid to approximate the quantiles of $ \mathcal{Q}_n (\gamma) $ as it does not capture the high-order dependence embodied in $ \mathcal{V} $.} Now, with the superscript $ \star $ indicating the bootstrap analogue, we have the following result.
Theorem (ref) implies that the bootstrap bias correction is first-order correct but does not imply any higher-order improvement. We demonstrate its merits not analytically but numerically in Section (ref), which is common with the wild bootstrap, see e.g. gonccalves2004bootstrapping and references therein. Details of the components of $W_{n}^{\star}(\gamma )$ are left to Section (ref) of the online supplement.
For a linear regression model $y_t=x_t' \beta+\varepsilon_t$, we consider testing linear restrictions $\mathcal{H}_0^e:R^e\beta=r $, where $R^e$ is a matrix of rank $p\leq dim(\beta)$. For the usual Wald statistic
$\hat{M}=n^{-1}\sum_{t}x_tx_t'$ and $\hat\Omega$ is an estimator of $E\varepsilon_t^2x_tx_t'$, define ${\mathcal{Q}}_{n}^e:=\left({W_{n}^e-p}\right)/{\sqrt{2p}}$. Although the test statistic appears to be very similar to the Chow test, the next theorem shows that the numerator and denominator in our corrected test statistic are related differently, calling for different critical values. Furthermore, the HLV is now obtained by replacing $\Omega^{-1}$ in Assumption (ref) with $L=M^{-1}R^{e\prime}\left(R^e M^{-1}\Omega M^{-1}R^{e\prime}\right)^{-1}R^eM^{-1}$, and the resulting limit denoted $\mathcal{V}^e$.
To estimate $\mathcal{V}^e$ and employ bootstrap bias correction, it is convenient to reformulate the restriction as an exclusion restriction of growing dimension, without loss of generality. Indeed, let $ S $ be the orthogonal complement of $ R^e $ and $ Q=(S, R^e) $. Then, let $ \mathring{x}_t = Q'^{-1}x_t $, $ \delta = Q \beta - (0',r')' $ such that $ \delta = (\delta_1',\delta_2')' $, with $\mathring x_t=\left(\mathring x_{1t}',\mathring x_{2t}'\right)'$ conformably partitioned, $ \delta_2 = R^e \beta - r $ and $ \tilde{y}_t = y_t - \mathring x_{2t} ' r $. We can now test the null hypothesis $ \mathcal{H}_0^e : \delta_2 = 0 $ in the regression of $ \tilde{y}_t $ on $ \mathring x_t $.
This transformation makes it particularly convenient to impose the null in the bootstrap resampling at the bias correction stage. Let $\bar{\mathcal{Q}}_n^\star$ denote the bootstrap bias correction factor for $W_n^e$. This yields the bias-corrected statistic $\mathcal{T}_n^{e,b} = {\left(\mathcal{Q}_n^e - \bar{\mathcal{Q}}_n^{\star}\right) }/{ \sqrt{\hat{\mathcal{V}^e}}}$, where $\hat{\mathcal V}^e$ is defined analogously to $\hat{\mathcal{V}}$, but now with $ q_t = \left( np\right)^{-1/2} \tilde{x}_{2t} '\hat{\Omega}^{e-1} \hat{e}_{t} \text{\ensuremath{\sum_{s=1}^{t-1}\tilde{x}_{2t} \hat{e}_{s}}} $, where $ \tilde{x}_{2t} $ denotes the residuals from the regression of $\mathring x_{2t} $ on $\mathring x_{1t} $ and $\hat\Omega^e=n^{-1}\sum_{t}\tilde{x}_{2t} \tilde{x}_{2t} '\hat{e}_t^2$. Then, with $ W(\cdot) $ defined in (ref), we have the following theorem:
The limiting distribution is mixed normal and pivotal but different from the limit in Theorem (ref). This is because the Chow test considers a quadratic form in $n^{-1/2}\sum_{t=1}^n\psi_t(\gamma)\xi_t$, which differs from this section by introducing a trend into the regressors via the factor $\psi_t(\gamma)$. Due to this difference, the partial sum processes converge to Gaussian processes with different covariance kernels. An R code to compute the critical values is available in the replication files.
This section examines the finite sample properties of our bias corrected HLV-robust test $ \mathcal{T}_n^b $ compared to the standard chi-square test $ W_n $, which does not account for growing $ p $, and the unscaled $\mathcal{Q}_{n}$ statistic with standard normal critical values, which does not account for $ \mathcal{V} $, in terms of bias, size and power.
We will consider the examples below in our Monte Carlo experiments. In Appendix (ref) we check our assumptions for these settings.
\begin{E1}[Multiple Regression of Growing Dimension] koenker1988asymptotic found through his metastudy that it is common practice in econometrics to increase the number of regressors as the sample size $n$ grows, at a rate of roughly $O\left(n^{1/4}\right) $. In this case, the approximation error $r_t$ is not explicitly modeled and may be set as zero. Practitioners thus adopt a flexible approach to modelling, where the assumed model becomes richer with more covariates and with more lagged terms to account for the dynamic effect in the spirit of the distributed lag model, as illustrated in e.g. Stock2015introduction. \end{E1}
\begin{E2}[Infinite-Order Autoregression] This model is one of the most fundamental models in time series analysis, see e.g. brockwell1991time or Hamilton2020time. For the process to be stationary, the coefficients $\left\{ b _{j}\right\} $ in the AR$\left( \infty \right) $ model $ y_{t}=b _{0}+\sum_{j=1}^{\infty }b _{j}y_{t-j}+\varepsilon _{t} $ are assumed to obey a certain decay rate. Specifically, the tail sum of the coefficients satisfies Assumption (ref) if $\sum_{j=p}^{\infty}\left\vert b_{j}\right\vert =o\left( n^{-1/2}\right) $. While we take $p$ as given in our analysis, for practical purposes various methods based on information criteria are available to choose the truncation lag $p$, see e.g. Shibata1980 and references therein. wang2007regression propose a lasso-based autoregressive order selection rule while lee2018asymptotic propose a lag selection rule in an infinite order panel autoregression. For expositional ease, we assume that the observations begin from $ t=1-p $ and $ x_1 = (1,y_0,...,y_{1-p}) $. \end{E2}
\begin{E3}[Nonparametric Series Regression] In case of the nonparametric series least squares estimation of $E\left( y_{t}|z_{t}\right) $, there exists a sequence of transformations of the covariates $z_{t}$ given by $ x_{nt}:=x_{n}\left( z_{t}\right) :{\mathbb{R}}^{k}\mapsto {\mathbb{R}}^{p}, $ and coefficients $\beta _{n}$ such that $ E\left( y_{t}|\mathcal{F}_{t-1}\right) =f\left( z_{t}\right) =x_{nt}^{\prime }\beta _{n}+r_{n}\left( z_{t}\right), $ where $r_{nt}=r_{n}\left( z_{t}\right) $ meets Assumption (ref) for a broad class of functions $ f $, see e.g. Andrews1991, Newey1997, chen2007 and Lee2016. By Lemma 1 of Lee2016, it is met if $ |r_t|_{\infty} = O(p^\alpha) $ for some $ \alpha < 0 $ and $p^{2\alpha }\leq n^{-1}$. Depending on the smoothness of the nonparametric function $f(\cdot )$, the regressor support dimension $k$, and the type of basis functions used, different values of $\alpha $ may be implied, see e.g. Newey1997, chen2007, p. 5573, for examples and further references. Often, the condition ((ref)) holds under the so-called undersmoothing selection of $p$. Another closely related example is the partially linear regression model, e.g. engle1986semiparametric and robinson1988root. Again, while we do not consider data-dependent $p$, for practical purposes the literature proposes methods for the choice of $p$ using cross validation or information criteria, see e.g. p. 5623 of chen2007 for a list of references. \end{E3}
The tests are applied to the setting of the Chow test and testing general linear restrictions. We consider various sample sizes $ n $ and dimensions $ p $ from the three examples, E1-E3, with the error generated from a bounded ARCH process
where $ \phi(x) = x^2 1\{|x|\leq c\} +c^2 1\{|x|> c\} $, $\eta_{t}=\left(u_{t}-Eu_{t}\right)/\sqrt{var\left(u_{t}\right)}$, and $\left\{ u_{t}\right\} $ is an iid sequence from the Marron1992 normal mixture distributions of type 1-3, which we refer to as error 1, 2 and 3. Their error 1 is standard normal. For a standard normal vector $\left(Z_{1},...,Z_{k}\right)$ and multinomial vector $\left(d_{1},...,d_{k}\right)$ with probability $\left(1/5,1/5,3/5\right)$, error 2 skewed unimodal variate is $u_{t}=Z_{1}d_{1}+\left(2Z_{2}/3+1/2\right)d_{2}+d_{3}\left(5Z_{3}/9+13/12\right)$, while error 3 strongly skewed variate is $u_{t}=\sum_{l=0}^{7}d_{l+1}\left(Z_{l+1}\left(2/3\right)^{l}+3\left(\left(2/3\right)^{l}-1\right)\right)$ with equally likely $ d_i $'s. We report results using (ref) with $\alpha\in\left\{ 0.3,0.4,0.5,0.55,0.57\right\} $ and $ c = 2.5 $. Results from $ c=3 $ and $ \infty $ are similar and omitted.
More specifically, for multiple regression, E1, the regressors $ x_t $ consist of independent AR(1) processes with coefficient $ \alpha_x $ and ARCH innovation as in (ref) and their lags of order up to 3. That is, we consider the distributed lag model with a growing number of variables. The first five elements of $\beta$ in (ref) are set as $d_0\left(5^{-1/2},...,5^{-1/2}\right)p^{1/4}n^{-1/2}$ and the others as zeros. When there is a break, all the values become zero after the break so that the value $ d_0 $ controls the magnitude of the change. We vary $ p \in \{5,9,13\} $ to examine the effect of the dimension on our tests.
For the infinite order AR regression, E2, we generate the sample from the MA(1) model $ y_{t}=\varepsilon_{t}+\theta1\left\{ t\leq\mu\right\} \varepsilon_{t-1} $, with $\mu=n$ for the size experiment and $\mu=[n\gamma]$ for the power evaluation, and estimate the AR($ p $) model with $p=9$ for $ n = 250 $ and $ p = 13 $ for $ n = 500 $. For the sieve regression, E3, we consider two variables $ \zeta_{1t} $ and $ \zeta_{2t} $ and their lags $ \zeta_{1,t-1} $ and $ \zeta_{2,t-1} $ as regressors, denoted by $ z_{1t},\cdots,z_{4t} $, after transforming them as $2\arctan\left(\zeta_{it}\right)/\pi$. Each $\zeta_{it}$ follows an AR(1) process with ARCH error. The regression function is set as $f\left(z_{1},\cdots,z_{4}\right) =d_0\left(1,z_{1},\cdots,z_{4},z_{1}^{2},...,z_{4}^{p_{1}}\right)\left(1^{-2},...,p_2^{-2}\right)' +\sqrt{\left|z_{1}\right|/n}$ with $p_{1}=\left\lfloor n^{1/4}\right\rfloor $ and $ p_2=1+4(p_1-1) $. To estimate the regression function, we construct $ x_t $ from polynomial basis functions and its dimension $ p=p_2 $.
We first employ these DGPs to simulate pre-limit values of $ \mathcal{V} $ in (ref) with $ n=m=500, l=0 $ and $ p $ as described above for each case, which are plotted in Figure (ref), reporting averages from 10000 iterations. This serves as a useful illustration to observe visually that pre-limiting $ \mathcal{V} $ deviates from unity for various specifications. A broad observation we make is that the deviation is bigger with larger ARCH coefficients and bigger autocorrelation in $ x_t $, although this feature is not always monotone. To conclude, we observe that the nonlinear serial autocorrelation factor can induce serious distortion in inference without a suitable robustifying treatment, as we provide in section (ref).
Also, Appendix (ref) gives the verification of the high-level conditions in Assumptions 1-4 for these examples.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{Chow Test} We consider three candidate break points as proportions of the sample sizes, $ \gamma \in \{0.2,0.3,0.5\} $. We begin by examining the bias of $ \mathcal{Q}_n (\gamma) $, conventionally centered by the degrees of freedom $ p $, under the null hypothesis. Note that a severe bias in $ \mathcal{Q}_n (\gamma) $ also implies that the size of the Wald test $ W_n (\gamma) $ can be distorted severely. We report the results in Figure (ref), in which the line with dot markers shows the bias in $\mathcal{Q}_n(\gamma)$ for $n=250$ and $n=500$. For E1, (Figures (ref), (ref)), each vertical partition (marked by a dotted vertical line) corresponds to a specific value of $p$. Within each vertical partition the DGP parameters change along the horizontal axes as $(\text{error type}, \alpha)$, in lexicographic order. As $p$ grows, we observe that the $\mathcal{Q}_{n}\left(\gamma\right)$ statistic exhibits severe finite sample bias for all values of the DGP parameters.
A similar visualization of bias in $\mathcal{Q}_n(\gamma)$ for E2 is presented in Figures (ref), (ref). Rather than report values for different $p$, here we focus on the case $p=9$ for $ n = 250 $ and $ p = 13 $ for $ n = 500 $ and allow the values of $\alpha$ and $\theta$ to vary along the horizontal axis lexicographically as $(\text{error type}, \alpha\text{ or }\theta)$, as detailed in the caption. A substantial bias in $\mathcal{Q}_n(\gamma)$ is observed for all cases, regardless of $n=250$ or $n=500$, albeit the biases are generally smaller in the latter case. Finally, Figures (ref), (ref) show the bias in $\mathcal{Q}_n(\gamma)$ for E3, with the same $ p $ as for E2, to mimic the asymptotic regime of a sieve regression, and parameters as in E2. We observe a similar pattern of substantial bias for both sample sizes.
As discussed above, Figures (ref)-(ref) clearly show that the biases present in $\mathcal{Q}_{n}\left(\gamma\right)$ are severe. In these figures we also plot the bias of the bias-corrected HLV-robust statistic $\mathcal{T}^b_n(\gamma)$, shown in black with square markers. The bootstrap bias correction seems to work well for all the cases, substantially alleviating bias. In Figure (ref), we observe that $\mathcal{T}^b_n(\gamma)$ can still exhibit some bias for specific cases but for E1 and E3, unlike the bias of $\mathcal{Q}_n(\gamma)$, this is centered around zero, while for E2 it is generally smaller in absolute value. Thus we recommend the use of the bootstrap bias correction in practice especially when faced with large values of $p$.
We now study the finite sample rejection frequencies of four competing tests: $\mathcal{T}_{n}^{b}(\gamma),\mathcal{T}_{n}(\gamma),\mathcal{Q}_{n}(\gamma),$ and $W_{n}(\gamma)$, with specific parameter values as given in the respective figure captions. As shown earlier, the unknown HLV scaling factor $\mathcal{V}$ varies along different ARCH parameters. This motivates our approach of experimenting with different $\alpha$ values and innovations. The Monte Carlo sizes resulting from the experiment are plotted in Figure (ref), wherein we place a horizontal dotted line to mark the nominal size of 5%. We report results for $\gamma=0.3$. The vertical partitions in each panel of Figure (ref) correspond, as discussed earlier, to increasing values of $p$ from left to right in E1. We cover multiple regression (Figures (ref), (ref)), AR fits (Figures (ref), (ref)) and sieve regression (Figures (ref), (ref)) for $n=250, 500$.
For all DGPs, the usual Wald statistic $W_n(\gamma)$ (diamond markers) over-rejects. Simply standardizing the test statistic $W_n(\gamma)$ to $\mathcal{Q}_n(\gamma)$, hence ignoring the HLV $\mathcal{V}$, does not improve matters. In fact, it usually worsens the problem of over-rejection. This can be seen in the lines with triangle markers. Our HLV-robust statistic $\mathcal{T}_n(\gamma)$ does much better, as the lines with dot markers indicate. While this shows the importance of the correction for $\mathcal{V}$ that we stress in the paper, there is still a tendency to over-reject. On the other hand, applying the bootstrap bias correction and using the bias corrected HLV-robust statistic $\mathcal{T}_n^b(\gamma)$ achieves excellent size control, as can be seen in the line with square markers. The discussion holds regardless of whether $n=250$ or $n=500$. Thus the importance of our proposed testing procedure is clearly visible.
We now analyze the power features of the competing test statistics for the proposed DGPs, allowing for breaks of different magnitudes and setting $\gamma=0.5$. After the break all the coefficients become zero so that the values of $d_0$ govern the size of the breaks in E1 and E3, while the values of $\theta$ do so for E2. The power performance is plotted in Figure (ref), where to conserve space we report results only for $n=400$. Again, we use $p=9$ for E2 and E3, while a range of $p$ is employed for E1. The line marker schemes for each of the competing tests are as described earlier. Examining the figure, the power of our HLV-robust statistics $\mathcal{T}_n(\gamma)$ (dots) and $\mathcal{T}_n^b(\gamma)$ (squares) tracks that of the uncorrected ones as the break size increases for both E2 (center panel) and E3 (right panel). For E1 (left panel), we only report results for $d_0=2$ for clarity. We observe that $W_n(\gamma)$ tends to have the highest power but our statistics still perform reasonably well with power in excess of 80% even for large $p$. Recall that our size experiments earlier indicate that $W_n(\gamma)$ over-rejects, a phenomenon of which high power is likely an artefact. Thus we conclude that our test is able to control size without sacrificing power to an undue extent.
Finally, Figure (ref) reports the size distortions when the sample variance of $ q_t $ multiplied by 2 is employed instead of $\hat{\mathcal{V}} $, noting that $ q_t $ is a martingale difference array. While we discussed the potential reasons for this severe size distortion in Section (ref), the investigation on a more precise approximation to the finite sample distribution of this statistic is an interesting issue but out of the scope of this paper due to the complex nature of the statistic. Nevertheless, Figure (ref) provides numerical evidence that the typical scaling by variance is not sufficient to control size. Specifically, define the test statistic $\mathcal{T}_n^{b,2}(\gamma)$ exactly like $\mathcal{T}_n^{b}(\gamma)$ but with random scaling replaced by twice the variance of $q_t$ and the standard normal approximation. It is clear that the test (triangle markers) is oversized relative to our recommended test $\mathcal{T}_n^{b}(\gamma)$.
\@startsection{subsection}{2} \z@{0.2\linespacing\@plus 0.2\linespacing}{-.5em} {\normalfont}{Testing Linear Restrictions} This section presents the outcomes of bias, size and power experiments for testing general linear restrictions, analogous to those for the Chow test in the preceding discussion. We use the reparameterization of the linear restrictions to the exclusion restrictions ${\delta}_2=0$, as discussed in Section (ref). We focus on E1 with $n=400$, $p=8,12,16$, $d_0=1$, error 1 and 2 disturbances and $\alpha=0.4,0.55$. The results are displayed in Figure (ref), with the same marking scheme as before and three test statistics employed: $W_n^e$, $\mathcal{Q}_n^e$ and $\mathcal{T}_n^{e,b}$. In all three figures, each vertical partition marks a different value of $p$, increasing from left to right.
The left panel of Figure (ref) shows that the bootstrap bias correction indeed improves matters, as was the case for the Chow test. The center panel again demonstrates the importance of our proposed corrections for size control. $W_n^e$ and $\mathcal{Q}_n^e$ tend to over-reject, becoming worse as $p$ increases. $\mathcal{T}_n^{e,b}$ controls size very well for medium to large $p$, while still outperforming $W_n^e$ and $\mathcal{Q}_n^e$ for smaller $p$. The right panel shows that $\mathcal{T}_n^{e,b}$ sacrifices some power relative to $W_n^e$ and $\mathcal{Q}_n^e$, but not unduly so.
We revisit structural stability in the Hamilton2003 study of the effect of oil shocks on economic activities. The autoregressive distributed lag model, ADL$(p,p)$, with quarterly time series of outputs and several oil price measures is employed. For real output, the quarterly growth rate of chain-weighted real GDP is used, while the oil price is the nominal crude oil producer price index, seasonally unadjusted. As in Hamilton, three oil price measures were considered: the growth rate $o_{t}$ from the previous quarter, the rectified linear unit, $o_{t}^{+}=o_{t}1\left\{ o_{t}>0\right\} $, and the net oil price increase, $o_{t}^{n}$, defined as the amount by which log oil prices in quarter $t$ exceed their peak value over the previous 12 months. If it does not exceed the previous peak, then $o_t^n$ is taken to be zero. We extend the original sample using the FRED database at the St. Louis Fed to obtain a sample from January 1949 to October 2019.
First, we re-evaluate structural stability of the GDP dynamics using AR$(p)$ fits, and that of the regression function of GDP growth on oil price change using an ADL$(p,p)$ model with the three alternative measures of oil price change. Following Table 4 in Hamilton2003, we investigate four exogenous disruptions in world petroleum supply. These are: the Arab–Israel War (November 1973), the Iranian Revolution (November 1978), the Iran–Iraq War (October 1980) and the Persian Gulf War (August 1990).
The p-values of the tests are reported in sub-tables (a) and (b) in Table (ref), where for $ W(\gamma)$ these are computed using the ordinary Chow test. We observe that in many cases the usual Chow test supports a structural break in both regressions more strongly than our recommended $\mathcal{T}^b_n(\gamma)$ test. Thus, the evidence for structural instability is often no longer as strong. In fact, the p-values that we calculate using our test exceed those of the standard Chow test in 22 out of 32 cases.
Second, we explore the relevance of the oil price measures and of the nonlinear transformations ($o_t^+$, $o_t^n$) by testing two exclusion restrictions in the ADL$(p,p)$ regression that include all the three oil price measures as covariates. The first exclusion restriction is to set the coefficients of all the measures as zero and the second is to set those of the nonlinear transformations ($o_t^+$, $o_t^n$) to zero. This yields 12 and 8 df, respectively, when $ p=4 $ and 18 and 12 df, respectively, when $ p=6 $. As shown in sub-table (c) of Table (ref), our recommended test $T^{e,b}_n$ produces p-values bigger than 5% for all cases, suggesting the effect of oil price as measured by these transformations is not statistically significant, nor are the nonlinear transformations. The standard Wald test for the exclusion restrictions is more supportive of their inclusion but may lack robustness with large df. Overall, our tests produce larger p-values than the standard Chow or Wald tests in 72% of cases (26 out of 36) cases in Table (ref).
As another measure of economic activity we now consider the industrial production (IP) index. This is available at monthly frequency and thus we consider ADL(12,12) and ADL(18,18) to include lags of one year and one and a half years, respectively. With monthly data, the dimensionality becomes more important: the number of restrictions we test varies from 13 and 25 in the structural break test for the AR(12) and ADL(12,12) regressions to 36 and 48 for the exclusion tests in the ADL(18,18) regression.
The results in Table (ref) illustrate much stronger differences in the conclusions of our test versus the standard Chow test, compared to the GDP study in Table (ref). The rejection of the null of no structural break is now often overturned at reasonable significance levels. For example, for the Arab-Israel War, the ADL$(p,p)$ model fails to reject the null of no structural break for any significance level below 10.6% when using our test $\mathcal{T}^b_n(\gamma)$, in contrast to the standard Chow test ${W}_n(\gamma)$. Conclusions are likewise overturned for the AR$(p)$ model and the Iranian Revolution and Iran-Iraq War, and indeed for both the ADL$(p,p)$ and AR$(p)$ in several cases for the Persian Gulf War.
The results in Tables (ref) and (ref) are not surprising given our simulation evidence. Indeed, our Monte Carlo simulation illustrates the effect of the degrees of freedom (df) on finite sample properties of the two tests, $ W_n(\gamma)$ tends to have larger p values in the AR case ($p+1$ df) than in the ADL case ($2p+1$ df), while $\mathcal{T}^b_n(\gamma)$ would be the opposite. These are exactly the patterns that we also observe. Finally, sub-table (c) of Table (ref) shows even stronger differences than Table (ref)(c), with all conclusions on the exclusion restrictions overturned at significance levels below 11.6%. Overall, our tests produce larger p-values in every single case considered in Table (ref) when compared to the standard Chow or Wald tests.
\setstretch{1}
\newgeometry{left=.25in, right=.25in,bottom=.05in,top=.05in}
\allowdisplaybreaks