Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
36,022 characters · 11 sections · 33 citation commands
Asymptotics of an Explosive Autoregression under Dependence
Keywords: explosive asymptotics, dependence, dependent innovations, asymptotic theory, explosive autoregression, least-squares estimator, unstable autoregression
\begingroup \footnotetext{The author gratefully acknowledges Bent Nielsen, for his many insightful comments and guidance.} \endgroup
\setcounter{tocdepth}{2}
We expand the classic results of And1959 by providing three results. We first demonstrate that the strong convergence result of the centered least-squares estimator to a ratio of a forward and backward average of the innovations holds in settings where the innovations are correlated and non-centered. Second, we demonstrate the requirement of independence in And1959 can be loosened to $\alpha$-mixing. Finally, we provide an autocorrelation-robust feasible test statistic for the explosive parameter $\rho$ of an explosive first-order autoregression with potentially autocorrelated Gaussian ARMA innovations, a natural extension of Anderson's t-statistic originally derived under iid Gaussian innovations. \\ \leavevmode \\ And1959 contains three sets of results for a first-order explosive autoregression. First, provided the innovations are uncorrelated, have a bounded second moment and centering around zero, the least squares estimator converges to a ratio of the forward and backward averages of the innovations. Secondly, in Theorem 2.3, it is shown that if the innovations are additionally assumed to be independent, the limits of the forward and backward averages, if they exist, are independent. Third, under $iid$ Gaussian innovations, in the limit, the forward and backward averages are independent Gaussians, allowing for a t-statistic for the explosive parameter with a limiting Gaussian distribution to be provided. \\ \leavevmode \\ Anderson's original results remain influential in the purely explosive literature. Because the convergence results do not rely on standard invariance principles, Anderson's asymptotic results in the $iid$ Gaussian setting are still frequently utilized when deriving the asymptotic distribution of estimators and tests in an explosive environment. For instance FulHasGoe1981 and Jeg1988 generalize Anderson's Gaussian results into a multivariate setting. Additionally, the Gaussian results have seen more recent use in explosively co-integrated systems, for instance in Theorem 2.2 of PhiMag08 and co-explosive systems, see Corollary 1 of Nie2010. \\ \leavevmode \\ General consistency results, which permit explosivity, but crucially does not require characterization or existence of a limiting distribution of the least squares estimator, have seen advances beyond the $iid$ Gaussian setting. In those cases, the weaker Marcinkiewicz-Zygmund condition, provided in Equation ((ref)), suffices. This condition is generally considered quite strong outside of the explosive setting, as it may exclude stochastic volatility and ARCH models. In this way, LaiWei1983a provide consistency results for an AR(p), later generalized to VAR(p) in LaiWei1985, subsequently generalized to VAR(p) with deterministics in Nie2005. \\ \\ We divide the results into two main sections: Section (ref) generalizes the convergence results of the least-squares estimator to forward and backward averages of the innovations. The section also provides sufficient conditions for the forward average to be non-zero \emph{almost surely}, the backward average to converge in distribution, as well as the generalization of And1959. Section (ref) then derives the exact limiting distribution of the least-squares estimator and t-statistic under Gaussian ARMA innovations and provides a feasible autocorrelation-robust t-statistic. This section also shows how the univariate results can be expanded to higher order autoregressions by providing an example of how the AR(1) results may be used in AR(2) with intercept.
The examined data generating process is
((ref)) is a univariate autoregressive explosive process of order 1, with no intercept or deterministic terms. The autoregressive coefficient $\rho$ is assumed to be greater than one in absolute value. The innovation process $(\epsilon_t)_{t \in \mathbb{N}}$ and initial value $x_0$ are real-valued and defined on a common probability space $(\Omega, \mathcal{F}, \mathbb{P})$. No additional structure is imposed at this stage; intertemporal structure will be introduced explicitly when needed. \\ \leavevmode \\ The main asymptotic objects of interest are
The first is the least-squares estimator of a regression of $x_t$ on $x_{t-1}$. $A_T$ and $B_T$ denote the numerator and denominator of the least squares estimator $\hat{\rho}$, centered around its true value $\rho$. And $A_T/ \sqrt{B_T}$ is a t-statistic, originally introduced in And1959. \\ \leavevmode \\ The notation $A_T, B_T$ is used to match notation defined in And1959. Following that framework, introduce $ \beta := \rho^{-1}$, which by definition lies in $(-1,1)$, allowing for geometric series such as $\sum_{t=1}^T \beta^t$ to converge, a property which will be relied upon heavily in subsequent convergence results. \\ \leavevmode \\ Additionally, introduce
These two objects are geometrically decaying sums of innovations. $z_t$ is a forward average, attributing highest weight to the first innovations, whilst $y_t$ is a backward average. \\ \leavevmode \\ The objects $A_T, B_T, z_T, y_T$ are measurable as they are continuous functions of $(x_0, \epsilon_1, ..., \epsilon_T)$; see Theorem 13.3 of Bil79. Throughout the rest of the paper, unless otherwise stated, all convergence occurs on $(\Omega, \mathcal{F}, \mathbb{P})$. The notation used is as follows: $\overset{a.s.}{\rightarrow}$ denotes almost sure convergence, $\overset{\mathbb{P}}{\rightarrow}$ convergence in probability, $\overset{D}{\rightarrow}$ convergence in distribution. Recall that $\overset{a.s.}{\rightarrow}$ implies $ \overset{\mathbb{P}}{\rightarrow} $, which implies $ \overset{D}{\rightarrow}$. $\mathbb{E}[\cdot]$ denotes the expectations operator with respect to measure $\mathbb{P}$. Theorems, lemmas, propositions and corollaries are proven in the appendix.
In this subsection, we introduce Assumption (ref), which imposes a mild moment condition on the innovations and initial value. We then show that this assumption implies important marginal behavior of $z_T, A_T, B_T$.
Explosive processes have the unique property of converging towards limiting objects at geometric rates, as $x_T$ itself grows at rate $\rho^T$. Terms containing $x_{T-1}$, such as the numerator $A_T$, grow at rate $\rho^T$, whilst terms like the denominator $B_T$, which contain $x_{T-1}^2$, grow at rate $\rho^{2T}$. One therefore expects the OLS estimator to converge at rate $\rho^T$, and subsequent results will confirm this. \\ \\ \leavevmode One of our central results, is realizing that in an explosive setting, previously derived moment bounds retain their order even when allowing for correlation and non-centering of the innovations. The reason for this is the geometric convergence rate in the explosive settings dominates the quadratic growth of cross-terms. This insight is formalized in bounds of $z_T$, which are provided in Lemmas (ref) and (ref) in the Appendix. Once the bounds are established, the same arguments as in And1959 still apply; one simply replaces the old bounds with the new ones. \\ \\ \leavevmode To formalize this insight, begin by defining a candidate probability limit $z$:
With definitions clarified, we are now able to state our first result.
$z$ is of large importance, as the next lemma will demonstrate that the OLS estimator converges to functions of this variable.
Note that $y_T$ and $z$ are random variables with distributions in most contexts. Despite having characterized the marginal behavior of the numerator and denominator, without further assumptions two issues will arise when we try to combine these marginal results to examine the OLS estimator and the t-statistic. Firstly, there is nothing ensuring that the denominator is non-zero with probability 1. Secondly, the convergence of the distribution $y_T$ is not guaranteed. For the moment, we assume two issues away and explore the consequences of their absence. Section (ref) then examines these assumptions, and where useful, provides low-level sufficient conditions that imply their high-level counterparts.\\ \leavevmode \\ We assume the random variable $z$ is non-atomic at $0$, then explore the consequences.
We can now apply the continuous mapping theorem on our previous results without dividing by zero. This allows us to state our first theorem.
We see that Assumption (ref) is sufficient to ensure consistency as well as a preliminary strong convergence result on the scaled OLS estimator and the t-statistic. However, as nothing ensures $y_T$ converges, let alone joint convergence of $(z_T, y_T)$, we are not able to put a limiting object on the right-hand-side of the convergence results. In order to improve the rate in ((ref)) we must have joint convergence of $(y_T, z_T)$. We assume this and explore its consequences.
With this high-level assumption in place, we can establish joint convergence of the numerator terms by combining the $a.s.$ convergence results of Theorem (ref) with the joint convergence of $(y_T, z_T)$ via the Continuous Mapping Theorem.
And finally, combining all three assumptions, we provide the asymptotic behavior of the scaled OLS estimator and the t-statistic $A_T/\sqrt{B_T}$.
If further structure is imposed, the $Sign(z)$ term in ((ref)) can be ignored.
This concludes the generalization of Anderson's first results.
Much of the explosive literature concerned with asymptotic distributions assumes iid Gaussian innovations as this immediately implies Assumptions (ref)-(ref). However, as this paper is concerned with general, potentially dependent innovations, this subsection examines these assumptions and where necessary, provides general low-level sufficient conditions for Assumptions (ref)-(ref) with an emphasis on ARMA- and stochastic volatility models.
Intuitively, if the innovations are continuously distributed, $z$, which is a forward average of innovations, ought to be continuously distributed, and so $\mathbb{P}(z \in \{0\})=0$. For independent innovations this is easy to verify, however, when the innovations/initial value are dependent, verification of this becomes more difficult, as one must rule out edge cases where the innovations exactly cancel each other out. \\ \leavevmode \\ For many latent-state innovations, such as stochastic volatility models, see for instance She05, it may be useful to exploit conditional non-atomicity to verify whether $\{z\in \{0\} \}$ is a null set, we do so in the following lemma.
For Gaussian ARMA $(\epsilon_t)_{t \in \mathbb{N}}$, one may utilize that $z$ is a sum of Gaussians to obtain a similar result. Before that, let us define the exact class of Gaussian ARMA processes we will be examining in the following assumption.
We begin establishing that the moment condition and stationarity are sufficient for marginal convergence in distribution of $y_T$.
However, stationarity does not necessarily imply joint convergence of $(y_T, z_T)$. This is demonstrated in the following counterexample:
We now show that $\alpha$-mixing is sufficient for joint convergence of $(y_T, z_T)$ as well as independence between the limiting objects. We recall the definition of $\alpha$-mixing
We now generalize And1959 by lessening his independence requirement to $\alpha$-mixing in the following theorem.
In this section, we will apply some of the previous general results to demonstrate their use cases. For inference we need to know the distribution of $y$. In general, nothing ensures that this variable is Gaussian or has a standard CDF. We show however, that if the innovations follow a Gaussian ARMA-process, the exact distribution of $y$ can be derived. This shows that feasible inference in the explosive setting goes much beyond the standard iid Gaussian setting originally provided in And1959. In another example, we show that the presented AR(1) results can be extended to an AR(2) with intercept, demonstrating extendability of the AR(1) results to larger models.
Suppose the innovations follow a stationary Gaussian ARMA(p,q) as in Assumption (ref) and recall that $(\gamma_h)_{h \in \mathbb{N}}$ are the covarainces of the process. One may then define $\Gamma = \gamma_0 + 2 \sum_{h=1}^\infty \gamma_h \beta^h$. This leads to the following behavior of $(y_T, z_T)$.
Stable Gaussian ARMAs as in Assumption (ref) satisfy Assumptions (ref)-(ref) by the following arguments: Stable ARMA innovations have a finite second moment, so Assumption (ref) is satisfied. Assumption (ref) is satisfied for ARMA innovations by Lemma (ref). Finally, Lemma (ref) implies Assumptions (ref)-(ref). Moreover, as Gaussian ARMA innovations without an intercept have a symmetric unconditional distribution around zero, the symmetry condition for Corollary (ref) is also satisfied. Therefore, combining Lemma (ref) with Corollary (ref), requiring Assumptions (ref)-(ref) and the previous symmetry condition, yields the following proposition.
As the estimator $\hat{\rho}$ is rate $\rho^T$ consistent, the residuals of the simple regression can be used to consistently recover the autocovariances of the innovations $(\gamma_h)_{h \in \mathbb{N}}$, which may be used to achieve a feasible test statistic for an explosive AR(1) with ARMA innovations. We do so by first establishing some uniform consistency results, beginning with the estimator $\hat{\beta}^h := \hat{\rho}^{-h}$, which also holds beyond the ARMA setting.
Next, let $L_T$ denote a sequence satisfying $L_T \rightarrow \infty$ and $L_T T^{-1} \rightarrow 0$. Define the sample autocovariances for $h \in \mathbb{N}_0$ as
One can then show
This enables us to define a feasible variance estimator of $\Gamma$ as
Which can be combined with Lemmas (ref), (ref) to establish consistency of $\hat{\Gamma}$, enabling inference on $\rho$ in the ARMA setting.
Based on LaiWei1983a and LaiWei1985, Nie2005 provides general consistency results for VAR(p) models with deterministic terms. Those results require the innovations to satisfy the Marcinkiewicz-Zygmund moment condition
where $(\mathcal{F}_{t})$ is an increasing sequence of $\sigma$-fields. This is considered quite strong for most time-series contexts, as it for instance may exclude latent-state processes and ARCH. However, in Section (ref), we demonstrated that Assumptions (ref)-(ref) were sufficient for consistency of $\hat{\rho}$. Moreover, we provided sufficient conditions for latent-state innovations to satisfy Assumption (ref) in Lemma (ref), providing an alternative set of sufficient conditions for consistency that may be used when the Marcinkiewicz-Zygmund condition in ((ref)) is violated or difficult to verify. This alternative set of sufficient conditions for consistency in the AR(1) can be carried over to general autoregressions by re-applying the same diagonalization framework as Nie2005. This is demonstrated in the following example, where we show how the AR(1) results carry over to an AR(2) with intercept.
We generalized Anderon's first result, showing that the assumptions of uncorrelated and independent innovations were not necessary for the numerator and denominator to converge to forward- and backward averages of the innovations. We then demonstrated how the independence assumption of the innovations made in And1959 may be relaxed to $\alpha$-mixing. We provided the exact limiting distributions of the forward- and backward averages in the ARMA setting, demonstrating that inference is feasible outside the $iid$ setting. Finally, we demonstrated how the AR(1) results, in particular the consistency results, may be extended to larger autoregressions via an example.