EconBase
← Back to paper

Asymptotics of an Explosive Autoregression under Dependence

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

36,022 characters

Asymptotics of an Explosive Autoregression under Dependence


\maketitle

\begin{abstract}
    We generalize the convergence results of an explosive autoregression, pioneered in \cite{And1959}, in three ways: First, we demonstrate that the centered least-squares estimator converges geometrically to a ratio of limits, even in settings where the innovations are correlated and not centered around zero. Secondly, we demonstrate that the requirement of independent innovations in \cite[Theorem 2.3,]{And1959} can be relaxed to $\alpha$-mixing. Third, we provide an autocorrelation-robust feasible test statistic for the explosive parameter under Gaussian ARMA innovations.
\end{abstract}

\noindent \textbf{Keywords:} explosive asymptotics, dependence, dependent innovations, asymptotic theory, explosive autoregression, least-squares estimator, unstable autoregression

\begingroup
\footnotetext{The author gratefully acknowledges Bent Nielsen, for his many insightful comments and guidance.}
\endgroup
\newpage


\setcounter{tocdepth}{2}


\section{Introduction}
We expand the classic results of \cite{And1959} by providing three results. We first demonstrate that the strong convergence result of the centered least-squares estimator to a ratio of a forward and backward average of the innovations holds in settings where the innovations are correlated and non-centered. Second, we demonstrate the requirement of independence in \cite[Theorem 2.3,]{And1959} can be loosened to $\alpha$-mixing. Finally, we provide an autocorrelation-robust feasible test statistic for the explosive parameter $\rho$ of an explosive first-order autoregression with potentially autocorrelated Gaussian ARMA innovations, a natural extension of Anderson's t-statistic originally derived under iid Gaussian innovations. \\
\leavevmode \\
\cite{And1959} contains three sets of results for a first-order explosive autoregression. \emph{First}, provided the innovations are uncorrelated, have a bounded second moment and centering around zero, the least squares estimator converges to a ratio of the forward and backward averages of the innovations. \emph{Secondly}, in Theorem 2.3, it is shown that if the innovations are additionally assumed to be independent, the limits of the forward and backward averages, if they exist, are independent. \emph{Third}, under $iid$ Gaussian innovations, in the limit, the forward and backward averages are independent Gaussians, allowing for a t-statistic for the explosive parameter with a limiting Gaussian distribution to be provided.
\\ \leavevmode \\
Anderson's original results remain influential in the purely explosive literature. Because the convergence results do not rely on standard invariance principles, Anderson's asymptotic results in the $iid$ Gaussian setting are still frequently utilized when deriving the asymptotic distribution of estimators and tests in an explosive environment. For instance \cite{FulHasGoe1981} and \cite{Jeg1988} generalize Anderson's Gaussian results into a multivariate setting. Additionally, the Gaussian results have seen more recent use in \emph{explosively co-integrated systems}, for instance in Theorem 2.2 of \cite{PhiMag08} and \emph{co-explosive systems}, see Corollary 1 of \cite{Nie2010}. \\ \leavevmode \\
General consistency results, which permit explosivity, but crucially does not require characterization or existence of a limiting distribution of the least squares estimator, have seen advances beyond the $iid$ Gaussian setting. In those cases, the weaker \emph{Marcinkiewicz-Zygmund} condition, provided in Equation (\ref{eq:mz}), suffices. This condition is generally considered quite strong outside of the explosive setting, as it may exclude stochastic volatility and ARCH models. In this way, \cite{LaiWei1983a} provide consistency results for an AR(p), later generalized to VAR(p) in \cite{LaiWei1985}, subsequently generalized to VAR(p) with deterministics in  \cite{Nie2005}. \\
\\
\noindent We divide the results into two main sections: Section \ref{sec:econometric_theory} generalizes the convergence results of the least-squares estimator to forward and backward averages of the innovations. The section also provides sufficient conditions for the forward average to be non-zero \emph{almost surely}, the backward average to converge in distribution, as well as the generalization of \cite[Theorem 2.3]{And1959}. Section \ref{sec:imp} then derives the exact limiting distribution of the least-squares estimator and t-statistic under Gaussian ARMA innovations and provides a feasible autocorrelation-robust t-statistic. This section also shows how the univariate results can be expanded to higher order autoregressions by providing an example of how the AR(1) results may be used in AR(2) with intercept.

\section{Results} \label{sec:econometric_theory}

\subsection{Definitions} \label{sec:def}
The examined data generating process is
\begin{equation} \label{eq:model}
    x_t = \rho x_{t-1} + \epsilon_t .
\end{equation}
(\ref{eq:model}) is a univariate autoregressive explosive process of order 1, with no intercept or deterministic terms.
The autoregressive coefficient $\rho$ is assumed to be greater than one in absolute value. The innovation process $(\epsilon_t)_{t \in \mathbb{N}}$ and initial value $x_0$ are real-valued and defined on a common probability space $(\Omega, \mathcal{F}, \mathbb{P})$. No additional structure is imposed at this stage; intertemporal structure will be introduced explicitly when needed.
\\ \leavevmode \\
The main asymptotic objects of interest are
\begin{equation*}
    \hat{\rho} := \frac{\sum_{t=1}^T x_t x_{t-1}}{ \sum_{t=1}^T x_{t-1}^2 }, \quad
    \hat{\rho} - \rho = \frac{\sum_{t=1}^T \epsilon_t x_{t-1}}{ \sum_{t=1}^T x_{t-1}^2 } =: \frac{A_T}{B_T} \quad  \text{ and } \quad \frac{A_T}{\sqrt{B_T}}.
\end{equation*}
The first is the least-squares estimator of a regression of $x_t$ on $x_{t-1}$. $A_T$ and $B_T$ denote the numerator and denominator of the least squares estimator $\hat{\rho}$, centered around its true value $\rho$. And $A_T/ \sqrt{B_T}$ is a t-statistic, originally introduced in \cite{And1959}. \\
 \leavevmode \\
The notation $A_T, B_T$ is used to match notation defined in \cite{And1959}. Following that framework, introduce $ \beta := \rho^{-1}$, which by definition lies in $(-1,1)$, allowing for geometric series such as $\sum_{t=1}^T \beta^t$ to converge, a property which will be relied upon heavily in subsequent convergence results. \\ \leavevmode
\\
Additionally, introduce
\begin{align}
    z_t :=& \beta^{t-2} x_{t-1} = \begin{cases}
        \rho x_0 \hspace{82pt} \text{ if }t=1  \\ \sum_{i=1}^{t-1} \beta^{i-1} \epsilon_{i} + \rho x_0 \label{eq:zt_def} \hspace{9 pt} \text{ if } t > 1
    \end{cases},\\
    y_t :=& \sum_{i=1}^t \beta^{t-i} \epsilon_i. \label{eq:y_def}
\end{align}
These two objects are geometrically decaying sums of innovations.  $z_t$ is a \emph{forward average}, attributing highest weight to the \emph{first} innovations, whilst $y_t$ is a \emph{backward average}. \\ \leavevmode
\\
The objects $A_T, B_T, z_T, y_T$ are measurable as they are continuous functions of $(x_0, \epsilon_1, ..., \epsilon_T)$; see Theorem 13.3 of \cite{Bil79}. Throughout the rest of the paper, unless otherwise stated, all convergence occurs on $(\Omega, \mathcal{F}, \mathbb{P})$. The notation used is as follows: $\overset{a.s.}{\rightarrow}$ denotes almost sure convergence, $\overset{\mathbb{P}}{\rightarrow}$ convergence in probability, $\overset{D}{\rightarrow}$ convergence in distribution. Recall that $\overset{a.s.}{\rightarrow}$ implies $ \overset{\mathbb{P}}{\rightarrow} $, which implies $ \overset{D}{\rightarrow}$. $\mathbb{E}[\cdot]$ denotes the expectations operator with respect to measure $\mathbb{P}$.  Theorems, lemmas, propositions and corollaries are proven in the appendix.

\subsection{Convergence Results} \label{sec:found_results}
In this subsection, we introduce Assumption \ref{ass:var}, which imposes a mild moment condition on the innovations and initial value. We then show that this assumption implies important marginal behavior of $z_T, A_T, B_T$.
\begin{assumption} \label{ass:var} The innovations have a \textbf{bounded second moment}, that is: \vspace{-5mm}
\begin{equation}
    \sup_{t \in \mathbb{N}} \mathbb{E} [\epsilon_t^{2}] \leq M^2< \infty, \quad M \in \mathbb{R}_+.
\end{equation}
Additionally, let the initial value $x_0$ satisfy $\mathbb{E} [x_0^2] \leq M^2$.
\end{assumption}
\begin{remark}
    Assumption \ref{ass:var} relaxes the assumptions in \cite{And1959}, where it is assumed that the innovations are: (1) \emph{uncorrelated} (2) have \emph{expectation zero} (3) have a \emph{constant, bounded second moment} (4) The  initial value $x_0$ is a \emph{finite constant}. ARCH/Stochastic volatility innovations generally satisfy these stronger conditions, as they are uncorrelated, whilst ARMA innovations do not, as these are serially correlated. Assumption \ref{ass:var} permits serial correlation, and therefore permits ARMA innovations.
\end{remark}
\noindent Explosive processes have the unique property of converging towards limiting objects at geometric rates, as $x_T$ itself grows at rate $\rho^T$. Terms containing $x_{T-1}$, such as the numerator $A_T$, grow at rate $\rho^T$, whilst terms like the denominator $B_T$, which contain $x_{T-1}^2$, grow at rate $\rho^{2T}$. One therefore expects the OLS estimator to converge at rate $\rho^T$, and subsequent results will confirm this. \\
\\ \leavevmode
One of our central results, is realizing that in an explosive setting, previously derived moment bounds retain their order even when allowing for correlation and non-centering of the innovations. The reason for this is the geometric convergence rate in the explosive settings dominates the quadratic growth of cross-terms. This insight is formalized in bounds of $z_T$, which are provided in Lemmas \ref{lem:z_T_bounds} and \ref{lem:z_bounds} in the Appendix. Once the bounds are established, the same arguments as in \cite{And1959} still apply; one simply replaces the old bounds with the new ones. \\
\\ \leavevmode
\noindent To formalize this insight, begin by defining a candidate probability limit $z$:
\begin{equation}
    z := \begin{cases} \label{eq:zdef}
            \lim_{T \rightarrow \infty} z_T \quad \text{ on }  A := \{ \omega: \lim_{T \rightarrow \infty} z_T \text{ exists in }\mathbb{R} \} \\
            0 \hspace{59pt}  \text{ on } A^c
        \end{cases}
\end{equation}
\begin{remark}
    This definition ensures $z$ is a random variable (taking values in $\mathbb{R}$) by giving it an arbitrary value in $\mathbb{R}$ (here $0$) on sets where $z_T$ diverges. This is a standard trick in probability theory; see for instance the proof of Proposition 2.43 of \cite{Brei68}.
\end{remark}
\noindent With definitions clarified, we are now able to state our first result.

\begin{lemma} \label{lem:measurability}
    Under Assumption \ref{ass:var}, $z_T \overset{a.s.}{\rightarrow} z$, with $\mathbb{E} z^2 <\infty$.
\end{lemma}
\noindent $z$ is of large importance, as the next lemma will demonstrate that the OLS estimator converges to functions of this variable.
\begin{lemma}
\label{lem:B_T}
Under Assumption \ref{ass:var}
\begin{align}
       \beta^{2(T-2)} B_T \hspace{10pt} \overset{a.s.}{\rightarrow}&
        \hspace{3pt} \frac{1}{1-\beta^2} z^2, \label{eq:den_conv} \\
    ( \beta^{T-2} A_T - y_T z ) \overset{a.s.}{\rightarrow}&  \hspace{17pt} \label{eq:num_conv} 0.
\end{align}
\end{lemma}

\noindent  Note that $y_T$ and $z$ are random variables with distributions in most contexts. Despite having characterized the marginal behavior of the numerator and denominator, without further assumptions two issues will arise when we try to combine these marginal results to examine the OLS estimator and the t-statistic. Firstly, there is nothing ensuring that the denominator is non-zero with probability 1. Secondly, the convergence of the distribution $y_T$ is not guaranteed. For the moment, we assume two issues away and explore the consequences of their absence. Section \ref{sec:yz} then examines these assumptions, and where useful, provides low-level sufficient conditions that imply their high-level counterparts.\\
\leavevmode \\
\noindent We assume the random variable $z$ is non-atomic at $0$, then explore the consequences.
\begin{assumption} \label{ass:z_0}
    $\mathbb{P}(z \in \{0\})=0$.
\end{assumption}
\begin{remark}
    In subsection \ref{sec:z_0}, we show that Assumption \ref{ass:z_0} is satisfied if the innovations follow conventional Gaussian ARMA innovations, and provide a sufficient condition that may be used to verify whether this condition is satisfied under other innovations.
\end{remark}

\noindent We can now apply the continuous mapping theorem on our previous results without dividing by zero. This allows us to state our first theorem.
\begin{theorem} \label{thm:consistency}
    Under Assumptions \ref{ass:var}-\ref{ass:z_0}
    \begin{align}
        \left| \frac{\rho^T}{\rho^2-1} (\hat{\rho} - \rho) - \frac{y_T}{z} \right| & \overset{a.s.}{\rightarrow} \, 0 \hspace{65pt} \label{eq:white_dist_noy} \\
        \left| \frac{A_T}{\sqrt{B_T}} - Sign(z)\sqrt{1-\beta^2} y_T  \right| & \overset{a.s.}{\rightarrow}  \, 0. \hspace{65pt} \label{eq:and_est_noy}
    \intertext{Let $(f_T)$ be a sequence converging to 0, then}
        f_T\rho^{T-2} (\hat{\rho}-\rho) &\overset{\mathbb{P}}{\rightarrow} 0. \label{eq:OLS_consistency}
    \end{align}
    Additionally, if $(f_T)$ satisfies $\sum_{t=0}^\infty |f_t| < \infty$, (\ref{eq:OLS_consistency}) converges almost surely.
\end{theorem}
\begin{corollary} \label{cor:consistency}
    Provided Assumptions \ref{ass:var}-\ref{ass:z_0}, the OLS estimator is strongly consistent. That is, $(\hat{\rho}-\rho) \overset{a.s.}{\rightarrow} 0$.
\end{corollary}
\noindent We see that Assumption \ref{ass:z_0} is sufficient to ensure consistency as well as a preliminary strong convergence result on the scaled OLS estimator and the t-statistic. However, as nothing ensures $y_T$ converges, let alone joint convergence of $(z_T, y_T)$, we are not able to put a limiting object on the right-hand-side of the convergence results. In order to improve the rate in (\ref{eq:OLS_consistency}) we must have joint convergence of $(y_T, z_T)$. We assume this and explore its consequences.
\begin{assumption} \label{ass:y_exists}
$(z_T, y_T) \overset{D}{\rightarrow} (z, y)$.
\end{assumption}
\noindent With this high-level assumption in place, we can establish joint convergence of the numerator terms by combining the $a.s.$ convergence results of Theorem \ref{lem:measurability} with the joint convergence of $(y_T, z_T)$ via the Continuous Mapping Theorem.
\begin{lemma} \label{lem:numden_joint}
    Under Assumptions \ref{ass:var}, \ref{ass:y_exists}
    \begin{equation*}
        \left( \beta^{T-2} A_T, (1-\beta^2) \beta^{2(T-2)} B_T \right) \overset{D}{\rightarrow} (yz,z^2).
    \end{equation*}
\end{lemma}

\noindent And finally, combining all three assumptions, we provide the asymptotic behavior of the scaled OLS estimator and the t-statistic $A_T/\sqrt{B_T}$.

\begin{theorem} \label{thm:and_dist}
    Under Assumptions \ref{ass:var}-\ref{ass:y_exists}, jointly
    \begin{align}
        \frac{\rho^T}{\rho^2-1}  (\hat{\rho}-\rho) \quad & \overset{D}{\rightarrow} \quad \frac{y}{z} \label{eq:white_dist} \\
        \frac{A_T}{\sqrt{B_T}} \quad &\overset{D}{\rightarrow} \quad Sign(z) \sqrt{ 1 - \beta^2 } \, y.         \label{eq:and_est}
    \end{align}
\end{theorem}
\begin{remark}
    Note that Lemma \ref{lem:measurability} and Theorems \ref{thm:consistency}-\ref{thm:and_dist} do not require \emph{centering} or \emph{uncorrelatedness}, as is done in \cite{And1959}. Therefore, these theorems apply for Gaussian ARMA innovations, a class of innovations which were previously ruled out  in \cite{And1959}.
\end{remark}
\noindent If further structure is imposed, the $Sign(z)$ term in (\ref{eq:and_est}) can be ignored.
\begin{assumption} \label{ass:indep}
    $y$ and $z$ are independent.
\end{assumption}
\begin{remark}
    In Subsection \ref{sec:indep}, sufficient conditions for Assumptions \ref{ass:y_exists}-\ref{ass:indep} are provided. These are satisfied for many conventional Gaussian ARMA and latent-state processes.
\end{remark}
\begin{corollary} \label{cor:symm}
    Under Assumptions \ref{ass:var}, \ref{ass:y_exists} and \ref{ass:indep} and if the innovations satisfy the symmetry condition
    \begin{equation*}
        (\epsilon_1, \epsilon_2, ..., \epsilon_T) \overset{D}{=} - (\epsilon_1, \epsilon_2, ..., \epsilon_T), \quad \forall T \in \mathbb{N}.
    \end{equation*}
    Then, jointly,
    \begin{equation*}
        \frac{A_T}{\sqrt{B_T}} \overset{D}{\rightarrow} \sqrt{1-\beta^2} y, \quad \frac{\rho^T}{\rho^2-1}  (\hat{\rho}-\rho) \overset{D}{\rightarrow} \frac{y}{z}.
    \end{equation*}
\end{corollary}
\noindent This concludes the generalization of Anderson's first results.

\subsection{Properties of $y$ and $z$} \label{sec:yz}
Much of the explosive literature concerned with asymptotic distributions assumes iid Gaussian innovations as this immediately implies Assumptions \ref{ass:var}-\ref{ass:indep}. However, as this paper is concerned with general, potentially dependent innovations, this subsection examines these assumptions and where necessary, provides general low-level sufficient conditions for Assumptions \ref{ass:z_0}-\ref{ass:indep} with an emphasis on ARMA- and stochastic volatility models.

\subsubsection{Sufficient conditions for Assumption \ref{ass:z_0}} \label{sec:z_0}
Intuitively, if the innovations are continuously distributed, $z$, which is a forward average of innovations, ought to be continuously distributed, and so $\mathbb{P}(z \in \{0\})=0$. For independent innovations this is easy to verify, however, when the innovations/initial value are dependent, verification of this becomes more difficult, as one must rule out edge cases where the innovations exactly cancel each other out. \\
\leavevmode \\
\noindent For many latent-state innovations, such as stochastic volatility models, see for instance \cite{She05}, it may be useful to exploit conditional non-atomicity to verify whether $\{z\in \{0\} \}$ is a null set, we do so in the following lemma.
\begin{lemma} \label{lem:SV}
    Suppose the initial value $x_0 \overset{a.s.}{=} 0$ and the innovations $(\epsilon_t)_{t \in \mathbb{N}}$ take on a form $\epsilon_t = \zeta_t \sigma_t$, where $\zeta_t \overset{iid}{\sim}N(0,1)$ with $\sigma_t^2>0$ $a.s.$ for any fixed $t$ and $\sup_{t>0}\mathbb{E}[\sigma_t^2] < \infty$. Furthermore, let $\sigma(\zeta_t: t>0)$ be independent of $\sigma(\sigma_t^2: t >0)$. Then Assumption  \ref{ass:z_0}, stating $\mathbb{P}(z \in \{0\})=0$ is satisfied.
\end{lemma}
\begin{remark}
    This lemma provides an alternative set of sufficient conditions for consistency in the explosive setting where the assumptions made in \cite{Nie2005} are not satisfied. However, the requirement of $\sigma(\zeta_t: t>0)$ being independent of $\sigma(\sigma_t^2: t >0)$ rules out ARCH.
\end{remark}

\noindent For Gaussian ARMA $(\epsilon_t)_{t \in \mathbb{N}}$, one may utilize that $z$ is a sum of Gaussians to obtain a similar result. Before that, let us define the exact class of  Gaussian ARMA processes we will be examining in the following assumption.

\begin{assumption} \label{ass:arma}
    The innovations $(\epsilon_t)_{t \in \mathbb{N}}$ are a stationary Gaussian ARMA(p,q) process with centering around zero. That is for all $t \in \mathbb{Z}$, letting $L$ denote the lag operator, the innovations take on the form
    \begin{equation*}
        \sum_{i=0}^p (1-\varphi_i L)^{i} \epsilon_t = \sum_{i=0}^q (1-\vartheta_i L)^i \zeta_t, \quad \zeta_t \overset{iid}{\sim} N(0,1).
    \end{equation*}
    Additionally, let the roots of the polynomials $\varphi(\mathsf{z})=\sum_{i=1}^p (1-\varphi_i \mathsf{z})^{i}$ and $\theta(\mathsf{z}) = \sum_{i=1}^q (1-\vartheta_i \mathsf{z})^i $ lie outside the unit circle and let the polynomials have no common roots. Denote the covariances $\gamma_{|t-j|} := cov(\epsilon_t,\epsilon_j)$ and let the initial value of the explosive process be $x_0 = 0$.
\end{assumption}

\begin{lemma} \label{lem:ARMA_znonzero}
    Assumption \ref{ass:arma} implies Assumption \ref{ass:z_0}.
\end{lemma}

\subsubsection{Convergence and Independence of $z$ and $y$} \label{sec:indep}
We begin establishing that the moment condition and stationarity are sufficient for marginal convergence in distribution of $y_T$.

\begin{lemma} \label{lem:y_dist_stat}
    Under Assumption \ref{ass:var} and stationary innovations, there exists a distribution $y$ with a finite second moment, such that $ y_T \overset{D}{\rightarrow} y $.
\end{lemma}

\noindent However, stationarity does not necessarily imply joint convergence of $(y_T, z_T)$. This is demonstrated in the following counterexample:
\begin{example}
    Let $\Omega_1, \Omega_2$ be disjoint sets satisfying $\mathbb{P}(\Omega_1)=\mathbb{P}(\Omega_2)=0.5$ and $\Omega_1 \cup \Omega_2 = \Omega$. Assume the initial value is $x_0=0$. Define the innovation sequence $(\epsilon_t)_{t \in \mathbb{N}}$ via:
    \begin{equation*}
        \big(\epsilon_t(\omega)\big)_{t \in \mathbb{N}} = \begin{cases}
            \big( (-1)^t \big)_{t \in \mathbb{N}} \hspace{11pt} \text{ for } \omega \in \Omega_1 \\
            \big( (-1)^{t+1} \big)_{t \in \mathbb{N}} \text{ for } \omega \in \Omega_2.
        \end{cases}
    \end{equation*}
    That is $(\epsilon_t)_{t \in \mathbb{N}}$ oscillates deterministically between $ -1$ and $1$, with the only source of randomness being whether even or odd indexes take on positive values. Note this process is stationary. \\ \leavevmode
    \\
    \noindent However, in this case $(y_T,z_T)$ is divergent, as convergence along \emph{even} $T$ and \emph{odd} $T$ lead to different CDFs:
    \begin{align*}
        \text{ T odd:} \quad &\, {(1+\beta)}{(y_T, z_T)} \overset{D}{\rightarrow} \begin{cases}
            (-1, -1) \text{ w.p. } 0.5 \\
            (\hspace{10pt} 1, \hspace{10pt}  1)  \text{ w.p. } 0.5
        \end{cases} \\
        \text{ T even:}\quad &  \, {(1+\beta)}(y_T, z_T) \overset{D}{\rightarrow} \begin{cases}
            (-1, \hspace{10pt} 1) \text{ w.p. } 0.5 \\
            ( \hspace{10pt} 1, -1)  \text{ w.p. } 0.5
        \end{cases}
    \end{align*}
    That is along the subsequence of odd integers, $z_T, y_T$
    converge to the same limit, whilst along the subsequence of even integers $z_T, y_T$ converge to limits of opposite signs.
\end{example}


\noindent We now show that $\alpha$-mixing is sufficient for joint convergence of $(y_T, z_T)$ as well as independence between the limiting objects. We recall the definition of $\alpha$-mixing
\begin{assumption}[$\alpha$-mixing]  \label{ass:alpha_mixing}
    Define the $\alpha$-mixing coefficient between two $\sigma$-fields $\mathcal{A}, \mathcal{B}$ as
    \[
    \alpha(\mathcal{A}, \mathcal{B}) := \sup_{A \in \mathcal{A}, B \in\mathcal{B}} \left| \mathbb{P}(A \cap B) - \mathbb{P}(A) \mathbb{P}(B) \right|.\]
    The $\alpha$-mixing coefficient of the process $(\epsilon_t)$ is defined via
    \[ \alpha_\epsilon(h) := \sup_u \alpha \big( \sigma(x_0, \epsilon_t: t \leq u),  \sigma(\epsilon_t: t \geq u+h) \big).\]
    The innovations satisfy $\alpha_\epsilon(h) \rightarrow 0$ as $h \rightarrow \infty$.
\end{assumption}


\begin{remark} \label{rem:mixing}
   By \cite[Theorem 6, pp. 99]{Dou1994}, ARMA innovations as in Assumption \ref{ass:arma} are $\beta$-mixing, which implies $\alpha$-mixing.
\end{remark}

\noindent We now generalize \cite[Theorem 2.3]{And1959} by lessening his independence requirement to $\alpha$-mixing in the following theorem.
\begin{theorem} \label{lem:independence}
    Under Assumptions \ref{ass:var} and \ref{ass:alpha_mixing}, and if $y_T \overset{D}{\rightarrow} y$, then $(y_T, z_T) \overset{D}{\rightarrow} (y,z)$, where $y$ and $z$ are independent and both have a finite second moment.
\end{theorem}
\begin{corollary}
    Suppose the innovations are stationary and satisfy Assumptions \ref{ass:var}, \ref{ass:alpha_mixing}. Then $(y_T, z_T) \overset{D}{\rightarrow} (y,z)$, where $y$ and $z$ are independent and both have a finite second moment.
\end{corollary}
\begin{remark} The corollary is established by noting that \emph{stationarity} of $(\epsilon_t)_{t \in \mathbb{N}}$ is a sufficient condition for $y_T \overset{D}{\rightarrow} y$ by Lemma \ref{lem:y_dist_stat}. Theorem \ref{lem:independence} assumes $y_T \overset{D}{\rightarrow} y$ as this also encompasses situations where $y_T$ converges marginally even when the innovations are nonstationary.
\end{remark}

\section{Implications} \label{sec:imp}
In this section, we will apply some of the previous general results to demonstrate their use cases. For inference we need to know the distribution of $y$. In general, nothing ensures that this variable is Gaussian or has a standard CDF. We show however, that if the innovations follow a Gaussian ARMA-process, the exact distribution of $y$ can be derived. This shows that feasible inference in the explosive setting goes much beyond the standard iid Gaussian setting originally provided in \cite{And1959}. In another example, we show that the presented AR(1) results can be extended to an AR(2) with intercept, demonstrating extendability of the AR(1) results to larger models.

\subsection{Results for Gaussian ARMA Innovations} \label{sec:arma}
Suppose the innovations follow a stationary Gaussian ARMA(p,q) as in Assumption \ref{ass:arma} and recall that $(\gamma_h)_{h \in \mathbb{N}}$ are the covarainces of the process. One may then define $\Gamma = \gamma_0 + 2 \sum_{h=1}^\infty \gamma_h \beta^h$. This leads to the following behavior of $(y_T, z_T)$.
\begin{lemma} \label{lem:arma}
    Under Assumption \ref{ass:arma}, $(y_T, z_T) \overset{D}{\rightarrow} N(0, \sigma^2 I_2)$, with $\sigma^2 = \frac{\Gamma}{1-\beta^2}.$
\end{lemma}
\noindent Stable Gaussian ARMAs as in Assumption \ref{ass:arma} satisfy Assumptions \ref{ass:var}-\ref{ass:indep} by the following arguments: Stable ARMA innovations have a finite second moment, so Assumption \ref{ass:var} is satisfied. Assumption  \ref{ass:z_0} is satisfied for ARMA innovations by Lemma \ref{lem:ARMA_znonzero}. Finally, Lemma \ref{lem:arma} implies Assumptions \ref{ass:y_exists}-\ref{ass:indep}. Moreover, as Gaussian ARMA innovations without an intercept have a symmetric unconditional distribution around zero, the symmetry condition for Corollary \ref{cor:symm} is also satisfied. Therefore, combining Lemma \ref{lem:arma} with Corollary \ref{cor:symm}, requiring Assumptions \ref{ass:var}-\ref{ass:indep} and the previous symmetry condition, yields the following proposition.
\begin{proposition} \label{thm:arma}
    Under Assumption \ref{ass:arma}, jointly,
    \begin{align*}
        \frac{\rho^T}{\rho^2-1} (\hat{\rho} - \rho) &\overset{D}{\rightarrow} Cauchy \\
        \frac{A_T}{\sqrt{B_T}} &\overset{D}{\rightarrow} N \big(0, \Gamma \big).
    \end{align*}
\end{proposition}
\begin{remark}
    \cite{Whi1958} proves that under iid Gaussian errors, the scaled OLS estimator has a limiting Cauchy distribution. Proposition \ref{thm:arma} demonstrates the same limit holds even for potentially autocorrelated stable Gaussian ARMA(p,q) innovations.
\end{remark}

\begin{remark}
    \cite{And1959} demonstrates in the Gaussian iid case, $A_T / \sqrt{B_T}$ has a limiting Gaussian distribution with variance $\gamma_0$. However, in the ARMA(p,q) case, the dependence in the innovations distorts the variance of the t-statistic via the term $2 \sum_{h=1}^\infty \gamma_h \beta^h$.
\end{remark}

 \noindent As the estimator $\hat{\rho}$ is rate $\rho^T$ consistent, the residuals of the simple regression can be used to consistently recover the autocovariances of the innovations $(\gamma_h)_{h \in \mathbb{N}}$, which may be used to achieve a feasible test statistic for an explosive AR(1) with ARMA innovations. We do so by first establishing some uniform consistency results, beginning with the estimator $\hat{\beta}^h := \hat{\rho}^{-h}$, which also holds beyond the ARMA setting.

\begin{lemma} \label{lem:uniformbeta}
    Under Assumptions \ref{ass:var}-\ref{ass:z_0}, one has $\sup_{0<h<T} |\hat{\beta}^h-\beta^h| \overset{a.s.}{\rightarrow}0$. Furthermore, one has $\sum_{h=1}^T |\hat{\beta}^h - \beta^h| \overset{a.s.}{\rightarrow} 0$.
\end{lemma}
\noindent Next, let $L_T$ denote a sequence satisfying $L_T \rightarrow \infty$ and $L_T T^{-1} \rightarrow 0$. Define the sample autocovariances for $h \in \mathbb{N}_0$ as
\begin{equation*}
    \hat{\gamma}_h := (T-h)^{-1} \sum_{t=h+1}^T (x_t - \hat{\rho} x_{t-1})(x_{t-h}-\hat{\rho}x_{t-h-1}).
\end{equation*}
One can then show
\begin{lemma}\label{lem:ACFs}
    Let the innovations satisfy Assumption \ref{ass:arma}. Then
    \begin{equation*} \textstyle
        \sup_{0 \leq h\leq L_T} |\hat{\gamma}_h - \gamma_h| \overset{\mathbb{P}}{\rightarrow} 0.
    \end{equation*}
\end{lemma}
\noindent This enables us to define a feasible variance estimator of $\Gamma$ as
\begin{equation*}
    \hat{\Gamma} := \hat{\gamma}_0 + 2 \sum_{h=1}^{L_T} \hat{\gamma}_h \hat{\beta}^h.
\end{equation*}
Which can be combined with Lemmas \ref{lem:uniformbeta}, \ref{lem:ACFs} to establish consistency of $\hat{\Gamma}$, enabling inference on $\rho$ in the ARMA setting.
\begin{theorem} \label{thm:feasiblearma}
    Under Assumption \ref{ass:arma}, then $\hat{\Gamma} \overset{\mathbb{P}}{\rightarrow} \Gamma$ and $ \hat{\Gamma}^{-1} \frac{A_T}{\sqrt{B_T}} \overset{D}{\rightarrow} N(0,1).$
\end{theorem}


\subsection{Consistency of AR(2)} \label{sec:ar2}
Based on \cite{LaiWei1983a} and \cite{LaiWei1985}, \cite{Nie2005} provides general consistency results for VAR(p) models with deterministic terms. Those results require the innovations to satisfy the \emph{Marcinkiewicz-Zygmund} moment condition
\begin{equation} \label{eq:mz}
    \exists \delta>0 : sup_t \mathbb{E}[\epsilon_t^{2+\delta}|\mathcal{F}_{t-1}] < \infty \quad a.s.
\end{equation}
where $(\mathcal{F}_{t})$ is an increasing sequence of $\sigma$-fields. This is considered quite strong for most time-series contexts, as it for instance may exclude latent-state processes and ARCH. However, in Section \ref{sec:econometric_theory}, we demonstrated that Assumptions \ref{ass:var}-\ref{ass:z_0} were sufficient for consistency of $\hat{\rho}$. Moreover, we provided sufficient conditions for latent-state innovations to satisfy Assumption \ref{ass:z_0} in Lemma \ref{lem:SV}, providing an alternative set of sufficient conditions for consistency that may be used when the Marcinkiewicz-Zygmund condition in (\ref{eq:mz}) is violated or difficult to verify. This alternative set of sufficient conditions for consistency in the AR(1) can be carried over to general autoregressions by re-applying the same diagonalization framework as \cite{Nie2005}. This is demonstrated in the following example, where we show how the AR(1) results carry over to an AR(2) with intercept.
\begin{example}{\label{ex:ar2}}
Consider the explosive AR(2) with intercept $\mu \in \mathbb{R}$ of the form
\begin{equation*}
    (1-\rho L) (1-\alpha L) x_t = \mu + \epsilon_t, \quad 0<\alpha<1<\rho.
\end{equation*}
By stacking terms into a vector $\mathbf{x}_t :=(x_t, x_{t-1}, 1)$ and defining
\begin{equation*}
    \Pi := \begin{pmatrix}
        \alpha+\rho & -\alpha \rho & \mu \\
        1 & 0 & 0 \\ 0 & 0 & 1
    \end{pmatrix}, \quad M := \begin{pmatrix}
        1 & -\rho & \tfrac{\mu}{\alpha-1} \\ 0 & 0 & 1 \\ 1 &-\alpha & \tfrac{\mu}{\rho-1}
    \end{pmatrix} , \quad \Lambda:= \begin{pmatrix}
        \alpha & 0 & 0 \\ 0 & 1 & 0 \\ 0 & 0 & \rho
    \end{pmatrix}
\end{equation*}
one may write the model in companion form as $\mathbf{x}_t = \Pi \mathbf{x}_{t-1} + e_t$, where $e_t = (\epsilon_t,0,0)$. As $M \Pi = \Lambda M$, the companion model  can pre-multiplied by $M$, delivering a diagonalized process $(u_t, 1, w_t) := M\mathbf{x}_t$ of the form
\begin{equation*}
    \begin{pmatrix}
        u_t \\ 1 \\ w_t
    \end{pmatrix} =  \begin{pmatrix}
        \alpha & 0 & 0 \\ 0 & 1 & 0 \\ 0 & 0 & \rho
    \end{pmatrix} \begin{pmatrix}
        u_{t-1} \\ 1 \\ w_{t-1}
    \end{pmatrix} + Me_t.
\end{equation*}
As $M$ is invertible, $x_t$ can be expressed as a linear combination of a stable AR(1), an explosive AR(1) and a deterministic term. \\ \leavevmode \\
This transformation is especially useful if we wish to analyse multivariate least-squares estimators. Define $\hat{\theta}$ as the least-squares estimator arising from a regression of $x_t$ on $x_{t-1}, x_{t-2}$ and an intercept. Defining $Y:=(\mathbf{x}_{T-1}'M', \mathbf{x}_{T-2}'M', ..., \mathbf{x}_{0}'M')$ and $\epsilon = (\epsilon_T, ..., \epsilon_1)$, one may write $\hat{\theta}$, centered around its true value $\theta = (\alpha+\rho, -\alpha \rho, \mu)$ as
\begin{equation*}
    (\hat{\theta} - \theta) = M' (Y'Y)^{-1}Y'\epsilon.
\end{equation*}
Defining
\begin{equation*}
    D_T:= \begin{pmatrix}
        \{\sum_{t=1}^T{u_{t-1}^2}\}^{1/2} & & \\
        & \sqrt{T} & \\
        && \{ \sum_{t=1}^T w_{t-1}^2\}^{1/2}    \end{pmatrix}
\end{equation*}
and
\begin{equation*}
    H_T := \begin{pmatrix}
        0 & \hat{\phi}(u_{t-1},1) & \hat{\phi}(u_{t-1},w_{t-1}) \\
        \hat{\phi}(1, u_{t-1}) & 0 & \hat{\phi}(1,w_{t-1}) \\
        \hat{\phi}(w_{t-1},u_{t-1}) & \hat{\phi}(w_{t-1},1) & 0
    \end{pmatrix}
\end{equation*}
and finally
\begin{equation*}
    \hat{\phi}(a_t,b_t):=\frac{\sum_{t=1}^T a_t b_t}{\sqrt{\sum_{t=1}^T a_t^2} \sqrt{\sum_{t=1}^T b_t^2}},
\end{equation*}
one may decompose $(\hat{\theta}-\theta)$ as
\begin{equation*}
     (\hat{\theta}-\theta) = M' (D_T [I_3 + H_T] D_T)^{-1} \sum_{t=1}^T
     \begin{pmatrix}
        u_{t-1} \epsilon_t \\
        \epsilon_t \\
        w_{t-1} \epsilon_t
    \end{pmatrix}.
\end{equation*}
This form can for instance be used to establish consistency of $\hat{\theta}$. As the focus of this paper is the explosive component, we assume the innovations and stable AR(1) component $u_t$ satisfy the follow high-level behavior:
\begin{assumption} \label{ass:stableAR}
    Let $\sigma_u^2, M^2 \in \mathbb{R}_+$. Moreover, assume $T^{-1} \sum_{t=1}^T \epsilon_t u_{t-1} \overset{\mathbb{P}}{\rightarrow} 0$, $T^{-1} \sum_{t=1}^T \epsilon_t \overset{\mathbb{P}}{\rightarrow} 0$, $T^{-1} \sum_{t=1}^T u_{t-1}^2 \overset{\mathbb{P}}{\rightarrow} \mathbb{E}[u_{t-1}^2] =\sigma_u^2$, $T^{-1} \sum_{t=1}^T u_{t-1} \overset{\mathbb{P}}{\rightarrow} 0$ and assume that $T^{-1} \sum_{t=1}^T \epsilon_t^2 \overset{\mathbb{P}}{\rightarrow} M^2$.
\end{assumption}
\noindent With high-level assumptions in place ensuring stable components are well-behaved, we may use the explosive convergence results provided in this paper to demonstrate that terms containing the explosive component in $(\hat{\theta}-\theta)$
vanish, and $\hat{\theta}$ is therefore consistent.
\begin{proposition} \label{prop:consistency}
     Let the innovations satisfy Assumptions \ref{ass:var}-\ref{ass:z_0} and \ref{ass:stableAR}. Then $\hat{\theta} \overset{\mathbb{P}}{\rightarrow} \theta$.
\end{proposition}

\end{example}

\section{Conclusion}
We generalized Anderon's first result, showing that the assumptions of \emph{uncorrelated} and \emph{independent} innovations were not necessary for the numerator and denominator to converge to forward- and backward averages of the innovations. We then demonstrated how the independence assumption of the innovations made in \cite[Theorem 2.3]{And1959} may be relaxed to $\alpha$-mixing. We provided the exact limiting distributions of the forward- and backward averages in the ARMA setting, demonstrating that inference is feasible outside the $iid$ setting. Finally, we demonstrated how the AR(1) results, in particular the consistency results, may be extended to larger autoregressions via an example.

\newpage