EconBase
← Back to paper

Initial-Condition-Robust Inference in Autoregressive Models

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

35,700 characters

Initial-Condition-Robust Inference in Autoregressive Models


\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long
\global\long
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}}
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle }
\global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle }
\global\long\def\turd#1{\frac{#1}{3}}
\global\long\global\long\def\sand#1{\left\lceil #1\right\vert }
\global\long\def\wich#1{\left\vert #1\right\rfloor }
\global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor }
\global\long\def\abs#1{\left|#1\right|}
\global\long\def\norm#1{\left\Vert #1\right\Vert }
\global\long\def\rest#1{\left.#1\right|}
\global\long\def\inprod#1{\left\langle #1\right\rangle }
\global\long\def\ol#1{\overline{#1}}
\global\long\def\ul#1{\underline{#1}}
\global\long\def\td#1{\widetilde{#1}}
\global\long\def\bs#1{\boldsymbol{#1}}
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long
\title{Initial-Condition-Robust Inference \\in Autoregressive Models}
\author{Donald W. K. Andrews\thanks{Andrews: Department of Economics and Cowles Foundation, Yale University,
\protect[email removed].} \and Ming Li\thanks{Li: Department of Economics and Risk Management Institute, National
University of Singapore, \protect[email removed].} \and Yapeng Zheng\thanks{Zheng: Department of Economics, Chinese University of Hong Kong, \protect\href{http://[email removed]}{[email removed]}.}}
\maketitle
\begin{abstract}
This paper considers confidence intervals (CIs) for the autoregressive
(AR) parameter in an AR model with an AR parameter that may be close
or equal to one. Existing CIs rely on the assumption of a stationary
or fixed initial condition to obtain correct asymptotic coverage and
good finite sample coverage. When this assumption fails, their coverage
can be quite poor. In this paper, we introduce a new CI for the AR
parameter whose coverage probability is completely robust to the initial
condition, both asymptotically and in finite samples. This CI pays
only a small price in terms of its length when the initial condition
is stationary or fixed. The new CI also is robust to conditional heteroskedasticity
of the errors.

\vspace{0.4in}

\emph{Keywords}: Asymptotic size, autoregressive model, confidence
set, initial condition, robustness. \medskip

\emph{JEL Classification Numbers}: C10, C12.
\end{abstract}
\thispagestyle{empty}

\newpage{}

\setcounter{page}{1}

\section{Introduction}

We consider an AR model with an AR parameter that may take values
close or equal to one. Existing CIs in the literature for the AR parameter
assume that the initial condition is stationary, zero, or fixed, e.g.,
see \textcite{stock1991confidence}, \textcite{andrews1993exactly},
\textcite{andrews1994approximately}, \textcite{hansen1999grid},
\textcite{elliott2001confidence}, \textcite{mikusheva2007uniform},
and \textcite[AG14 hereafter]{andrews2014conditional}. If the initial-condition
assumption is not satisfied, then these CIs do not have correct asymptotic
coverage probabilities and their finite sample coverage probabilities
can be poor. For example, simulations reported below show that the
nominal 95\% AG14 CI has finite sample coverage probabilities ranging
from 24.1\% to 93.5\% across 50 cases with highly variable initial
conditions and different types of conditional heteroskedasticity of
the error. A quarter of these CPs are 79.0\% or less.

To circumvent the sensitivity of existing CIs to the initial condition,
this paper introduces a new initial-condition-robust (ICR) CI whose
finite sample CPs do not depend on the initial condition. These CIs
are easy to compute and do not depend on any tuning parameters. We
show that the ICR CI has correct asymptotic size in a uniform sense
under conditions that allow for an arbitrary initial condition and
conditional heteroskedasticity. Simulations show that the finite sample
CPs of the nominal 95\% ICR CI are quite good. They lie between 93.5\%
and 95.0\% in the cases considered.

Like the other CIs referenced above, the ICR CI is constructed by
inverting hypothesis tests concerning the value of the AR parameter.
For all of these CIs, the tests are based on a t statistic that depends
on a least squares (LS) estimator of the AR parameter. In contrast
to existing CIs, the LS estimator used by the ICR tests contains an
additional regressor that eliminates the effect of the initial condition
under the null hypothesis. (For the definition of this regressor,
see Section \ref{sec:Model-and-ICR} below.) In consequence, the ICR
CI has a CP that is invariant to the value of the initial condition.
However, the length of the ICR CI can be effected by the initial condition.
Simulations show that the ICR CI is slightly shorter when the initial
condition is highly variable than when it is stationary or fixed.\footnote{The length of the ICR CI varies mostly with the value of the AR parameter.
It is shortest for AR parameters near one. This is true of existing
CIs as well.}

In scenarios where existing CIs are asymptotically valid (i.e., the
initial condition is stationary or fixed), the ICR CI pays a price
in terms of its length compared to existing CIs. This occurs because
the ICR LS estimator has a higher variance due to the additional regressor
that is included in the LS regression. However, simulations show that
the effect is fairly small. Across 50 scenarios with stationary or
fixed initial conditions and i.i.d., GARCH, or ARCH errors, the ratio
of the expected length of the ICR CI compared to that of the AG14
CI is found to vary between 1.00 and 1.11. On average, the ICR CI
is 3.5\% longer than the AG14 CI in these scenarios.

Based on the ICR CI, we also introduce an asymptotically median-unbiased
interval estimator (MUE) of the AR parameter.

The AR model considered here has been applied in the literature to
exchange rate, commodity and stock prices, and other economic time
series, e.g., see \textcite{kim1993unit}.

This paper is organized as follows. Section \ref{sec:Model-and-ICR}
specifies the AR model and defines the ICR CI. Section \ref{sec:MUE}
defines the median-unbiased interval estimator. Section \ref{sec:Monte-Carlo-Simulations}
provides simulation results concerning CPs and average lengths of
the ICR and AG14 CIs and absolute median biases of the ICR\ MUE.
Section \ref{sec:Asymptotic-Results} provides asymptotic results
for the ICR CI. The Supplemental Material describes how the critical
values were computed, proves Theorems \ref{thm:ICR-CP} and \ref{thm:Jh-convergence},
and provides some additional simulation results.

\section{Model and ICR Confidence Interval}\label{sec:Model-and-ICR}

We consider the AR(1) model with conditional-heteroskedasticity studied
in AG14:
\begin{align}
Y_{i} & =\mu+Y_{i}^{*}\quad\text{for \ensuremath{i=0,1,\dots,n} and}\nonumber \\
Y_{i}^{*} & =\rho Y_{i-1}^{*}+U_{i}\quad\text{for }i=1,2,\dots,n,\label{eq:model}
\end{align}
where the observed data are $\{Y_{0},\dots,Y_{n}\}$, $\rho\in[-1+\varepsilon,1]$
for some $0<\varepsilon<2$, and $\{U_{i}:i=1,2,\dots\}$ are stationary
and ergodic under the distribution $F$, with conditional mean $0$
given a $\sigma$-field $\mathcal{G}_{i-1}$ for which $U_{j}\in\mathcal{G}_{i}$
for all $j\leq i$, conditional variance $\sigma_{i}^{2}=E_{F}(U_{i}^{2}|\mathcal{G}_{i-1})$,
and unconditional variance $\sigma_{U}^{2}\in(0,\infty)$. A detailed
description of the parameter space $\Lambda$ is provided in Section \ref{subsec:Parameter-space}.

By iterative substitution in (\ref{eq:model}), we have
\begin{align}
Y_{i} & =\mu+\rho^{i-1}\cdot\rho\cdot Y_{0}^{*}+\sum_{j=1}^{i}\rho^{i-j}U_{j}\quad\text{for }i=1,\dots,n.\label{eq:representation-ar1}
\end{align}
 Let $Y$, $U$, and $X_{1}$ be $n$-vectors whose $i$-th elements
are $Y_{i}$, $U_{i}$, and $Y_{i-1}$, respectively. Let $X_{2}(\rho)$
be an $n\times2$ matrix whose $i$-th row is $(1,\rho^{i-1})$ when
$-1<\rho<1$.\footnote{Following the convention $0^{0}=1$, we use $\rho^{i-1}$ in the regression
instead of $\rho^{i}$ to avoid perfect multicollinearity when $\rho=0$.} When $\rho=1$, we define $X_{2}(\rho)$ to be an $n\times2$ matrix
whose $i$-th row is $(1,i)$. Let $X(\rho)=[X_{1}:X_{2}(\rho)]$,
$P_{X(\rho)}=X(\rho)(X(\rho)'X(\rho))^{-1}X(\rho)'$, and $M_{X(\rho)}=I_{n}-P_{X(\rho)}$.
Let $\widehat{U}_{i}$ denote the $i$-th element of the residual
vector $M_{X(\rho)}Y$. Let $p_{ii}$ denote the $i$-th diagonal
element of $P_{X(\rho)}$, and define $p_{ii}^{*}=\min\{p_{ii},n^{-1/2}\}$.
Let $\Delta$ be a diagonal $n\times n$ matrix whose $i$-th diagonal
element is $\widehat{U}_{i}/(1-p_{ii}^{*})$. Then, the ICR LS estimator
$\widehat{\rho}_{n}(\rho)$, and the HC5 variance estimator $\widehat{\sigma}_{n}^{2}(\rho)$
are defined as
\begin{align}
\widehat{\rho}_{n}(\rho)= & \ (X_{1}'M_{X_{2}(\rho)}X_{1})^{-1}X_{1}'M_{X_{2}(\rho)}Y\text{ and}\nonumber \\
\widehat{\sigma}_{n}^{2}(\rho)= & \ (n^{-1}X_{1}'M_{X_{2}(\rho)}X_{1})^{-1}(n^{-1}X_{1}'M_{X_{2}(\rho)}\Delta^{2}M_{X_{2}(\rho)}X_{1})(n^{-1}X_{1}'M_{X_{2}(\rho)}X_{1})^{-1},\label{eq:ICR_sigma_hat}
\end{align}
respectively.

Let
\begin{equation}
T_{n}(\rho)=\frac{n^{1/2}(\widehat{\rho}_{n}(\rho)-\rho)}{\widehat{\sigma}_{n}^{2}(\rho)}.\label{eq:t-stat(rho)}
\end{equation}
For suitable sequences $\{\rho_{n}\}_{n\geq1}$ such that $n(1-\rho_{n})\to h\in[0,\infty]$,
we show that
\begin{equation}
T_{n}(\rho_{n})\stackrel{d}{\to} J_{h},\label{eq:asym-distribution}
\end{equation}
where $J_{h}$ is defined in Section \ref{subsec:Asymptotic-Distribution-Jh}
below. Table \ref{tab:ICR-quantiles} provides the quantiles $c_{h}(\alpha/2)$
and $c_{h}(1-\alpha/2)$ of the distribution of $J_{h}$ for $\alpha=.05$
and $.1$, which are the critical values employed by the ICR CI.

The nominal $1-\alpha$ equal-tailed two-sided ICR CI for $\rho$
is
\begin{equation}
CI_{\text{ICR},n}:=\left\{ \rho\in[-1+\epsilon,1]:c_{h}(\alpha/2)\leq T_{n}(\rho)\leq c_{h}(1-\alpha/2)\ \text{for}\ h=n(1-\rho)\right\} ,\label{eq:ICR_CI}
\end{equation}
which can be computed by taking a fine grid of values $\rho\in[-1+\epsilon,1]$
and comparing $T_{n}(\rho)$ to $c_{h}(\alpha/2)$ and $c_{h}(1-\alpha/2)$.
Given these critical values, computation of the CIs is fast.

\begin{table}[t]
\centering
\caption{Quantiles of $J_{h}$ for Use with 90\% and 95\% Equal-Tailed Two-Sided
CIs and MUEs}\label{tab:ICR-quantiles}

\centering
\scriptsize
\begin{tabular}{>{\centering\arraybackslash}m{.98cm}ccccccccccccc}
\toprule
\multicolumn{14}{c}{{Values of $c_{h}\left(\alpha\right)$, the $\alpha^{\text{th}}$ Quantile of $J_{h}$, for Use with 90\% and 95\% Equal-Tailed Two-Sided CI's and MUE's}}\tabularnewline
\midrule
{$h$} & {0} & {.2} & {.4} & {.6} & {.8} & {1} & {1.4} & {1.8} & {2.2} & {2.6} & {3} & {3.4} & {3.8}\tabularnewline
\midrule
{$c_{h}\left(.025\right)$} & {-3.66} & {-3.63} & {-3.60} & {-3.56} & {-3.54} & {-3.52} & {-3.46} & {-3.40} & {-3.36} & {-3.31} & {-3.27} & {-3.23} & {-3.19}\tabularnewline
{$c_{h}\left(.05\right)$} & {-3.41} & {-3.38} & {-3.35} & {-3.31} & {-3.28} & {-3.25} & {-3.20} & {-3.14} & {-3.08} & {-3.04} & {-3.00} & {-2.95} & {-2.91}\tabularnewline
{$c_{h}\left(.5\right)$} & {-2.18} & {-2.13} & {-2.09} & {-2.04} & {-1.99} & {-1.95} & {-1.86} & {-1.78} & {-1.70} & {-1.63} & {-1.57} & {-1.50} & {-1.45}\tabularnewline
{$c_{h}\left(.95\right)$} & {-.94} & {-.87} & {-.80} & {-.74} & {-.68} & {-.62} & {-.50} & {-.39} & {-.29} & {-.19} & {-.11} & {-.03} & {.05}\tabularnewline
{$c_{h}\left(.975\right)$} & {-.65} & {-.59} & {-.52} & {-.45} & {-.38} & {-.32} & {-.21} & {-.08} & {.01} & {.11} & {.19} & {.28} & {.35}\tabularnewline
\midrule
{$h$} & {4.2} & {4.6} & {5} & {6} & {7} & {8} & {9} & {10} & {11} & {12} & {13} & {14} & {15}\tabularnewline
\midrule
{$c_{h}\left(.025\right)$} & {-3.16} & {-3.12} & {-3.09} & {-3.02} & {-2.97} & {-2.90} & {-2.87} & {-2.82} & {-2.79} & {-2.75} & {-2.73} & {-2.71} & {-2.69}\tabularnewline
{$c_{h}\left(.05\right)$} & {-2.87} & {-2.83} & {-2.80} & {-2.72} & {-2.66} & {-2.61} & {-2.56} & {-2.52} & {-2.48} & {-2.45} & {-2.42} & {-2.40} & {-2.38}\tabularnewline
{$c_{h}\left(.5\right)$} & {-1.39} & {-1.34} & {-1.30} & {-1.20} & {-1.11} & {-1.04} & {-.99} & {-.93} & {-.89} & {-.85} & {-.82} & {-.78} & {-.76}\tabularnewline
{$c_{h}\left(.95\right)$} & {.11} & {.18} & {.24} & {.36} & {.46} & {.55} & {.61} & {.68} & {.74} & {.78} & {.81} & {.84} & {.88}\tabularnewline
{$c_{h}\left(.975\right)$} & {.41} & {.48} & {.54} & {.66} & {.77} & {.86} & {.92} & {.99} & {1.05} & {1.08} & {1.12} & {1.15} & {1.20}\tabularnewline
\midrule
{$h$} & {20} & {25} & {30} & {40} & {50} & {60} & {70} & {80} & {90} & {100} & {200} & {300} & {500}\tabularnewline
\midrule
{$c_{h}\left(.025\right)$} & {-2.59} & {-2.53} & {-2.47} & {-2.41} & {-2.36} & {-2.32} & {-2.30} & {-2.27} & {-2.26} & {-2.25} & {-2.16} & {-2.13} & {-2.09}\tabularnewline
{$c_{h}\left(.05\right)$} & {-2.28} & {-2.22} & {-2.15} & {-2.09} & {-2.05} & {-2.01} & {-1.99} & {-1.96} & {-1.94} & {-1.94} & {-1.84} & {-1.81} & {-1.78}\tabularnewline
{$c_{h}\left(.5\right)$} & {-.65} & {-.58} & {-.52} & {-.45} & {-.41} & {-.37} & {-.34} & {-.32} & {-.30} & {-.28} & {-.20} & {-.16} & {-.13}\tabularnewline
{$c_{h}\left(.95\right)$} & {.99} & {1.06} & {1.12} & {1.19} & {1.24} & {1.28} & {1.30} & {1.32} & {1.34} & {1.36} & {1.45} & {1.48} & {1.52}\tabularnewline
{$c_{h}\left(.975\right)$} & {1.30} & {1.38} & {1.43} & {1.50} & {1.55} & {1.59} & {1.62} & {1.64} & {1.66} & {1.68} & {1.76} & {1.79} & {1.83}\tabularnewline
\bottomrule
\end{tabular}

\end{table}


\section{ICR Median-Unbiased Interval Estimator}\label{sec:MUE}

By definition, an estimator $\widehat{\theta}_{n}$ of a parameter
$\theta$ is median unbiased if $P(\widehat{\theta}_{n}\geq\theta)\geq1/2$
and $P(\widehat{\theta}_{n}\leq\theta)\geq1/2$. In this section,
we introduce an ICR MUE of $\rho$ that satisfies an analogous condition.
Also, with probability close to one, this interval estimator is a
point.\footnote{If the median of the asymptotic null distribution of the $t$-statistic
is strictly decreasing in $\rho$, or equivalently, if $c_{h}(0.5)$
is strictly increasing in $h$, then the proposed interval estimator
would be a point estimator with probability one. Because this condition
fails to hold exactly, but almost holds, there is a very small probability
that the estimator is a short interval rather than a point.}

Let $CI_{\text{ICR},n}^{\text{up}}(.5)$ and $CI_{\text{ICR},n}^{\text{low}}(.5)$
denote level $.5$ one-sided upper-bound and lower-bound CIs for $\rho$,
respectively. By definition,
\begin{align}
CI_{\text{ICR},n}^{\text{up}}(.5) & \coloneqq\left\{ \rho\in[-1+\epsilon,1]:c_{h}(0.5)\leq T_{n}(\rho)\ \text{for}\ h=n(1-\rho)\right\} \text{ and}\nonumber \\
CI_{\text{ICR},n}^{\text{low}}(.5) & \coloneqq\left\{ \rho\in[-1+\epsilon,1]:T_{n}(\rho)\leq c_{h}(0.5)\ \text{for}\ h=n(1-\rho)\right\} .\label{eq:CI_low}
\end{align}

The MUE $\widetilde{\rho}_{n}$ of $\rho$ is defined by
\begin{align}
\widetilde{\rho}_{n} & =[\widetilde{\rho}_{n}^{\text{low}},\widetilde{\rho}_{n}^{\text{up}}],\text{ where}\nonumber \\
\widetilde{\rho}_{n}^{\text{up}} & =\max\{\rho:\rho\in CI_{\text{ICR},n}^{\text{up}}(.5)\}\text{ and}\nonumber \\
\widetilde{\rho}_{n}^{\text{low}} & =\min\{\rho:\rho\in CI_{\text{ICR},n}^{\text{low}}(.5)\}.\label{eq:rho_low}
\end{align}
By construction, we have $\widetilde{\rho}_{n}^{\text{low}}\leq\widetilde{\rho}_{n}^{\text{up}}$.\footnote{This holds because $\widetilde{\rho}_{n}^{\text{up}}\geq\sup\{\rho\in[-1+\epsilon,1]:c_{h}(.5)=T_{n}(\rho)\}$
and $\widetilde{\rho}_{n}^{\text{low}}$ is less than or equal to
the infimum of the values in the same set.} In addition, $\widetilde{\rho}_{n}$ is a singleton whenever the
set $\{\rho\in[-1+\epsilon,1]:T_{n}(\rho)=c_{h}(.5)\}$ contains a
single point, in which case $\widetilde{\rho}_{n}$ equals this point.
Table \ref{tab:ICR-quantiles} provides the critical values $c_{h}(.5)$
for a wide range of $h$. Given these critical values, computation
of $\widetilde{\rho}_{n}$ is fast.

\section{Monte Carlo Simulations}\label{sec:Monte-Carlo-Simulations}

We compare the coverage probabilities (CPs) and average lengths (ALs)
of the ICR and AG14 CIs  and report the absolute median biases (ABMs)
of the ICR MUE using Monte Carlo simulations. In Section \hyperref[SM-sec:Additional-Simulation-Results]{\textcolor{black}{\ref*{SM-sec:Additional-Simulation-Results}}}
of the Supplemental Material, we compare the performance of the ICR
and \textcite[Mik07 hereafter]{mikusheva2007uniform} CIs. We focus
on the nominal 95\% equal-tailed two-sided CIs. For the AG14 CI, we
follow the paper's constructions of the test statistic and critical
values.  To calculate the MUE for $\rho$, we follow the procedure
described in Section \ref{sec:MUE} and use $\widetilde{\rho}_{n}$
as the MUE of $\rho$ when $\widetilde{\rho}_{n}$ is a singleton.
When $\widetilde{\rho}_{n}^{\text{up}}\neq\widetilde{\rho}_{n}^{\text{low}}$,
we take $\widetilde{\rho}_{n}^{\text{up}}$ as the MUE for $\rho$
following \textcite[Section 5.1]{andrews2025inference}.

We consider a wide range of $\rho$ values: 0, 0.5, 0.7, 0.9, and
0.99. The innovations are of the form $U_{i}=\sigma_{i}\varepsilon_{i}$,
where $\{\varepsilon_{i}:i=0,\pm1,\dots\}$ are identically and independently
distributed (i.i.d.) standard normal and $\sigma_{i}$ is the multiplicative
conditional heteroskedasticity. Let GARCH-$(ma,ar;\psi)$ denote a
GARCH(1, 1) process with MA, AR, and intercept parameters $(ma,ar;\psi)$,
and let ARCH-$(ar_{1},\dots,ar_{4};\psi)$ denote an ARCH(4) process
with AR parameters $(ar_{1},\dots,ar_{4})$ and intercept $\psi$.
Following AG14, we consider five specifications for the conditional
heteroskedasticity of the innovations: (a) i.i.d. $N(0,1)$ (``i.i.d.''),
(b) GARCH(1,1)-(.05,.9;.001) (``GARCH1''), (c) GARCH(1,1)-(.15,.8;.2)
(``GARCH2''), (d) GARCH(1,1)-(.25,.7;.2) (``GARCH3''), and (e) ARCH(4)-(.3,.2,.2,.2;.2)
(``ARCH''). For the initial conditions, we consider the following
four specifications: (a) Fixed: $Y_{0}^{*}=0$, (b) Stationary: $Y_{0}^{*}=\sum_{i=0}^{\infty}\rho^{i}U_{-i}$,
(c) Scaled $n$: $Y_{0}^{*}\sim\sqrt{n}\cdot\sum_{i=0}^{\infty}\rho^{i}U_{-i}$,
and (d) Explosive: $Y_{0}^{*}\sim n^{3/4}\cdot\sum_{i=0}^{\infty}\rho^{i}U_{-i}$.
We consider a sample size of $n=150$. All simulation results are
based on 30,000 simulation repetitions.

\medskip{}

Tables \ref{tab:CP-ag14} and \ref{tab:CP-ICR} report the CPs of
AG14 and ICR CIs, respectively. The CPs of the AG14 CI in Table \ref{tab:CP-ag14}
are close to the nominal level of 95\% when the initial condition
is fixed at zero or stationary, but fall below 95\% when the initial
condition is more variable, namely in the scaled $n$ and explosive
$Y_{0}^{*}$ cases. In particular, the AG14 CI is robust to non-i.i.d.\
innovations when the initial condition is fixed or stationary.  On
the other hand, for scaled $n$ and explosive $Y_{0}^{*}$ cases,
the AG14 CPs lie in the interval {[}24.1, 93.5{]} with roughly a quarter
of the CPs being less than 79.0. Table \ref{tab:CP-ICR} shows that
the ICR CI achieves CPs close to 95\% in all of the scenarios considered
and for any initial conditions. Specifically, all ICR CI CPs lie in
the interval {[}93.5, 95.0{]}.

Table \ref{tab:length-ratio} reports the ratios of the ALs of the
nominal 95\% ICR CI to the AG14 CI. We include only cases with fixed
or stationary initial conditions because the AG14 CI is not robust
to more variable initial conditions, whereas the ICR CI is. The ratios
are between 1.00 and 1.11. On average, the ICR CI is 3.5\% longer
than the AG14 CI over the fifty cases considered.  Thus, the price
the ICR CI pays in terms of AL to gain robustness against the distribution
of the initial condition is fairly small.

\begin{table}[H]
\centering
\caption{Coverage probabilities ($\times100$) of the nominal 95\% AG14 CI}\label{tab:CP-ag14}

\begin{tabular}{lrrrrrlrrrrr}
\toprule
\multicolumn{1}{r}{\small\textbf{Initial Conditions:}} & \multicolumn{5}{c}{\small\textbf{Fixed}} & & \multicolumn{5}{c}{\small\textbf{Stationary}} \\
\multicolumn{1}{r}{\small\textbf{$\boldsymbol{\rho}$:}} & .00 & .50 & .70 & .90 & .99 & & .00 & .50 & .70 & .90 & .99 \\
\midrule
\small\textbf{i.i.d.} & 94.6 & 94.6 & 94.8 & 94.7 & 95.1 & & 94.6 & 94.5 & 94.8 & 94.7 & 94.5 \\
\small\textbf{GARCH1} & 94.6 & 94.5 & 94.8 & 94.8 & 95.1 & & 94.6 & 94.6 & 94.9 & 94.8 & 94.4 \\
\small\textbf{GARCH2} & 94.4 & 94.6 & 94.5 & 94.9 & 95.2 & & 94.3 & 94.6 & 94.5 & 94.8 & 94.7 \\
\small\textbf{GARCH3} & 94.2 & 94.3 & 94.6 & 94.7 & 95.2 & & 94.1 & 94.3 & 94.6 & 94.6 & 94.6 \\
\small\textbf{ARCH4} & 93.7 & 93.9 & 94.2 & 94.6 & 95.6 & & 93.6 & 94.0 & 94.1 & 94.5 & 94.7 \\
\midrule
\small\textbf{Initial Conditions:} & \multicolumn{5}{c}{\small\textbf{Scaled $\boldsymbol{n}$}} & & \multicolumn{5}{c}{\small\textbf{Explosive}} \\
\midrule
\small\textbf{i.i.d.} & 88.4 & 90.5 & 91.8 & 92.7 & 83.4 & & 63.7 & 80.5 & 86.6 & 90.8 & 77.7 \\
\small\textbf{GARCH1} & 59.1 & 79.0 & 86.4 & 91.7 & 83.2 & & 24.1 & 69.9 & 82.3 & 90.0 & 77.3 \\
\small\textbf{GARCH2} & 90.2 & 91.8 & 92.4 & 93.0 & 83.9 & & 71.0 & 83.3 & 87.8 & 90.9 & 77.4 \\
\small\textbf{GARCH3} & 90.9 & 92.1 & 92.5 & 93.4 & 84.2 & & 74.4 & 84.9 & 88.4 & 91.4 & 77.7 \\
\small\textbf{ARCH4} & 91.1 & 92.3 & 92.8 & 93.5 & 84.4 & & 76.7 & 86.3 & 89.4 & 91.8 & 77.9 \\
\bottomrule
\end{tabular}
\end{table}

\begin{table}[H]
\centering
\caption{Coverage probabilities ($\times100$) of the nominal 95\% ICR CI}\label{tab:CP-ICR}

\begin{tabular}{lrrrrr}
\toprule
\multicolumn{1}{r}{\small\textbf{Initial Conditions:}} & \multicolumn{5}{c}{\small\textbf{Arbitrary}} \\
\multicolumn{1}{r}{\small\textbf{$\boldsymbol{\rho}$:}} & .00 & .50 & .70 & .90 & .99 \\
\midrule
\small\textbf{i.i.d.} & 94.4 & 94.5 & 94.7 & 94.7 & 94.3 \\
\small\textbf{GARCH1} & 94.4 & 94.6 & 94.9 & 95.0 & 94.3 \\
\small\textbf{GARCH2} & 94.1 & 94.6 & 94.4 & 94.9 & 94.2 \\
\small\textbf{GARCH3} & 93.9 & 94.2 & 94.5 & 94.7 & 94.1 \\
\small\textbf{ARCH4} & 93.5 & 93.8 & 93.9 & 94.5 & 94.3 \\
\bottomrule
\end{tabular}
\end{table}

\begin{table}[H]
\centering
\caption{Ratios of the average lengths of the nominal 95\% ICR to AG14 CIs}\label{tab:length-ratio}

\begin{tabular}{lrrrrrlrrrrr}
\toprule
\multicolumn{1}{r}{\small\textbf{Initial Conditions:}} & \multicolumn{5}{c}{\small\textbf{Fixed}} & & \multicolumn{5}{c}{\small\textbf{Stationary}} \\
\multicolumn{1}{r}{\small\textbf{$\boldsymbol{\rho}$:}} & .00 & .50 & .70 & .90 & .99 & & .00 & .50 & .70 & .90 & .99 \\
\midrule
\small\textbf{i.i.d.} & 1.00 & 1.00 & 1.02 & 1.04 & 1.08 & & 1.00 & 1.01 & 1.02 & 1.05 & 1.11 \\
\small\textbf{GARCH1} & 1.04 & 1.05 & 1.06 & 1.07 & 1.08 & & 1.00 & 1.01 & 1.02 & 1.05 & 1.11 \\
\small\textbf{GARCH2} & 1.00 & 1.01 & 1.03 & 1.02 & 1.07 & & 1.00 & 1.01 & 1.03 & 1.04 & 1.11 \\
\small\textbf{GARCH3} & 1.00 & 1.01 & 1.03 & 1.02 & 1.06 & & 1.00 & 1.02 & 1.04 & 1.03 & 1.10 \\
\small\textbf{ARCH4} & 1.00 & 1.02 & 1.03 & 1.01 & 1.06 & & 1.01 & 1.02 & 1.04 & 1.03 & 1.10 \\
\bottomrule
\end{tabular}
\end{table}

\begin{table}[H]
\centering
\caption{Average lengths of the nominal 95\% ICR CI}\label{tab:avg_length_ICR}

\begin{tabular}{lrrrrrlrrrrr}
\toprule
\multicolumn{1}{r}{\small\textbf{Initial Conditions:}} & \multicolumn{5}{c}{\small\textbf{Fixed}} & & \multicolumn{5}{c}{\small\textbf{Stationary}} \\
\multicolumn{1}{r}{\small\textbf{$\boldsymbol{\rho}$:}} & .00 & .50 & .70 & .90 & .99 & & .00 & .50 & .70 & .90 & .99 \\
\midrule
\small\textbf{i.i.d.} & .32 & .28 & .24 & .17 & .08 & & .32 & .28 & .24 & .17 & .07 \\
\small\textbf{GARCH1} & .33 & .30 & .25 & .17 & .08 & & .34 & .30 & .26 & .17 & .08 \\
\small\textbf{GARCH2} & .38 & .34 & .29 & .19 & .08 & & .38 & .34 & .29 & .19 & .08 \\
\small\textbf{GARCH3} & .43 & .38 & .33 & .20 & .09 & & .43 & .38 & .33 & .20 & .08 \\
\small\textbf{ARCH4} & .49 & .44 & .37 & .21 & .09 & & .49 & .44 & .37 & .21 & .09 \\
\midrule
\small\textbf{Initial Conditions:} & \multicolumn{5}{c}{\small\textbf{Scaled $\boldsymbol{n}$}} & & \multicolumn{5}{c}{\small\textbf{Explosive}} \\
\midrule
\small\textbf{i.i.d.} & .31 & .27 & .23 & .14 & .04 & & .28 & .23 & .18 & .09 & .02 \\
\small\textbf{GARCH1} & .28 & .24 & .19 & .12 & .04 & & .18 & .15 & .12 & .07 & .02 \\
\small\textbf{GARCH2} & .37 & .32 & .26 & .15 & .04 & & .32 & .27 & .21 & .10 & .02 \\
\small\textbf{GARCH3} & .41 & .36 & .29 & .16 & .04 & & .36 & .30 & .23 & .11 & .02 \\
\small\textbf{ARCH4} & .46 & .40 & .32 & .17 & .04 & & .39 & .32 & .25 & .12 & .02 \\
\bottomrule
\end{tabular}
\end{table}

Table \ref{tab:avg_length_ICR} presents the ALs of the ICR CI. These
are substantially decreasing in $\rho$, as expected, longer for ARCH
conditional heteroskedasticity than i.i.d., and increasing from GARCH1
to GARCH2 to GARCH3 conditional heteroskedasticity.

\begin{table}[H]
\centering
\caption{Absolute Median Biases of the ICR MUE}\label{tab:mue_absolute_median_bias_ICR}

\begin{tabular}{lrrrrrlrrrrr}
\toprule
\multicolumn{1}{r}{\small\textbf{Initial Conditions:}} & \multicolumn{5}{c}{\small\textbf{Fixed}} & & \multicolumn{5}{c}{\small\textbf{Stationary}} \\
\multicolumn{1}{r}{\small\textbf{$\boldsymbol{\rho}$:}} & .00 & .50 & .70 & .90 & .99 & & .00 & .50 & .70 & .90 & .99 \\
\midrule
\small\textbf{i.i.d.} & .012 & .011 & .006 & .005 & .020 & & .012 & .011 & .006 & .005 & .015 \\
\small\textbf{GARCH1} & .017 & .011 & .006 & .005 & .020 & & .017 & .011 & .006 & .005 & .015 \\
\small\textbf{GARCH2} & .017 & .011 & .006 & .000 & .020 & & .017 & .011 & .006 & .000 & .015 \\
\small\textbf{GARCH3} & .022 & .011 & .011 & .000 & .020 & & .022 & .011 & .011 & .000 & .015 \\
\small\textbf{ARCH4} & .022 & .016 & .011 & .000 & .020 & & .022 & .016 & .011 & .000 & .015 \\
\midrule
\small\textbf{Initial Conditions:} & \multicolumn{5}{c}{\small\textbf{Scaled $\boldsymbol{n}$}} & & \multicolumn{5}{c}{\small\textbf{Explosive}} \\
\midrule
\small\textbf{i.i.d.} & .012 & .011 & .006 & .005 & .010 & & .012 & .011 & .006 & .010 & .010 \\
\small\textbf{GARCH1} & .017 & .011 & .006 & .005 & .010 & & .017 & .011 & .011 & .010 & .010 \\
\small\textbf{GARCH2} & .017 & .011 & .011 & .005 & .010 & & .017 & .011 & .011 & .010 & .010 \\
\small\textbf{GARCH3} & .022 & .011 & .011 & .005 & .010 & & .022 & .011 & .011 & .010 & .010 \\
\small\textbf{ARCH4} & .022 & .016 & .011 & .005 & .010 & & .022 & .016 & .011 & .010 & .010 \\
\bottomrule
\end{tabular}
\end{table}

Table \ref{tab:mue_absolute_median_bias_ICR} reports the AMBs of
the ICR MUE. Across all cases, the AMBs range from 0.000 to 0.022,
indicating that their magnitudes are generally small.

In Section \hyperref[SM-sec:Additional-Simulation-Results]{\textcolor{black}{\ref*{SM-sec:Additional-Simulation-Results}}} of the Supplemental
Material, we compare the performance of the ICR and Mik07 CIs via
simulations. The results are similar to those reported above for the
ICR and AG14 CIs except that the Mik07 CI undercovers both under conditional
heteroskedasticity, and/or scaled $n$ or explosive initial conditions,
but by a lower magnitude in the latter case than the AG14 CI. Again,
we find that the ICR CI is robust to more variable initial conditions
while paying only a small price in terms of CI width in the situations
where the Mik07 CI has correct asymptotic coverage, i.e., when the
innovations are i.i.d.\  and the initial conditions are stationary
(or fixed).


\section{Asymptotic Results}\label{sec:Asymptotic-Results}

This section establishes the correct uniform asymptotic size and asymptotic
similarity of the ICR CI for $\rho$.

\subsection{Parameter Space, Correct Uniform Size, and Asymptotic \\Similarity}\label{subsec:Parameter-space}

Following AG14, the parameter space for $(\rho,F)$ is given by:

\medskip{}

\noindent$\Lambda=\Big\{\lambda=(\rho,F)$: (i) \label{enu:para-space1}$\rho\in[-1+\epsilon,1]$;
\begin{enumerate}[start=2, nosep, leftmargin=4.5em]
\item \label{enu:para-space2}$\{U_{i}:i=1,2,\dots\}$ are stationary and
strong-mixing under $F$ with $E_{F}(U_{i}|\mathcal{G}_{i-1})=0$
a.s., $E_{F}(U_{i}^{2}|\mathcal{G}_{i-1})=\sigma_{i}^{2}$ a.s.,
where $\mathcal{G}_{i}$ is some non-decreasing sequence of $\sigma$-fields
for which $U_{j}\in\mathcal{G}_{i}$ for all $j\leq i$ for $i=1,2,\dots$;
\item \label{enu:para-space3}The strong-mixing numbers $\{\alpha_{F}(m):m\geq1\}$
satisfy $\alpha_{F}(m)\leq Cm^{-3\zeta/(\zeta-3)}$, $\forall m\geq1$;
\item \label{enu:para-space4}$\sup_{i,s,t}E_{F}|\prod_{a\in A}a|^{\zeta}\leq M$,
where $0<i,s,t<\infty$, $i\geq\max(s,t)$, and $A$ is any non-empty
subset of $\{U_{i-s},U_{i-t},U_{i+1}^{2},U_{1}^{2}\}$;
\item \label{enu:para-space5}$E_{F}U_{1}^{2}=1$;
\end{enumerate}
\hspace{2.6em}for some constants $0<\epsilon<2$, $\zeta>3$, and
$C<\infty$$\Big\}$.

\medskip{}

We suppose $E_{F}U_{1}^{2}=1$ for simplicity because the asymptotic
distribution of $T_{n}(\rho)$ does not depend on the unconditional
variance. Parts \ref{enu:para-space2} and \ref{enu:para-space4}
are slightly different from the corresponding conditions in AG14 because
the random variables $\{U_{-j},\ j\geq0\}$ do not enter into the
asymptotic analysis of the ICR CI. The key innovation of this paper
is that we allow arbitrary initial conditions $Y_{0}^{*}$, including
those corresponding to explosive processes.

The main theoretical result of this paper shows that $CI_{\text{ICR},n}$
has correct asymptotic size for the parameter space $\Lambda$ and is asymptotically
similar. Let $P_{\lambda}$ denote probability under $\lambda=\left(\rho,F\right)\in\Lambda$.
\begin{thm}
\label{thm:ICR-CP}Let $\alpha\in(0,1)$. For the parameter space $\Lambda$,
the nominal $1-\alpha$ ICR CI for the parameter $\rho$ satisfies
\[
\underset{n\to\infty}{\lim\inf}\inf_{\lambda\in\Lambda}P_{\lambda}(\rho\in CI_{\text{ICR},n})=\underset{n\to\infty}{\lim\sup}\sup_{\lambda\in\Lambda}P_{\lambda}(\rho\in CI_{\text{ICR},n})=1-\alpha.
\]
\end{thm}

The MUE $\widetilde{\rho}_{n}$ has the following median-unbiasedness
property. This result is a corollary to the one-sided versions of
Theorem \ref{thm:ICR-CP}.\footnote{Corollary \ref{cor:MUE} holds because $CI_{\text{ICR},n}^{\text{up}}(.5)$
and $CI_{\text{ICR},n}^{\text{low}}(.5)$ both have coverage probabilities
of $1/2$ or greater by the proof of Theorem \ref{thm:ICR-CP} applied
to these one-sided CIs and $(\widetilde{\rho}_{n}\geq\rho)\supset CI_{\text{ICR},n}^{\text{up}}(.5)$
and $(\widetilde{\rho}_{n}\leq\rho)\supset CI_{\text{ICR},n}^{\text{low}}(.5)$.}
\begin{cor}
\label{cor:MUE}The MUE estimator $\widetilde{\rho}_{n}$ satisfies
\begin{align*}
\underset{n\to\infty}{\lim\inf}\inf_{\lambda\in\Lambda}P_{\lambda}(\widetilde{\rho}_{n}^{\text{up}}\geq\rho) & \geq1/2\text{ and }\underset{n\to\infty}{\lim\inf}\inf_{\lambda\in\Lambda}P_{\lambda}(\widetilde{\rho}_{n}^{\text{low}}\leq\rho)\geq1/2.
\end{align*}
\end{cor}

\subsection{Asymptotic Distribution}\label{subsec:Asymptotic-Distribution-Jh}

As stated in (\ref{eq:asym-distribution}), $J_{h}$ is the asymptotic
distribution of $T_{n}(\rho_{n})$ under suitable sequences $\{\rho_{n}:n\geq1\}$.
Here, we state this result formally as Theorem \ref{thm:Jh-convergence}
and define $J_{h}$. This result is proved in a similar way to Theorem
1 in \textcite{andrews2012asymptotics}. Using Theorem 2.1 in \textcite{andrews2020generic},
Theorem \ref{thm:Jh-convergence} is sufficient to establish Theorem
\ref{thm:ICR-CP}, which is proved in Section \hyperref[SM-sec:Proof-of-Theorem1]{\textcolor{black}{\ref*{SM-sec:Proof-of-Theorem1}}}
of the Supplemental Material.
\begin{thm}
\label{thm:Jh-convergence}For a sequence $\lambda_{n}=(\rho_{n},F_{n})$
such that $n(1-\rho_{n})\to h\in[0,\infty]$, we have
\[
T_{n}(\rho_{n})=\frac{n^{1/2}(\widehat{\rho}_{n}(\rho_{n})-\rho_{n})}{\widehat{\sigma}_{n}(\rho_{n})}\stackrel{d}{\to} J_{h},
\]
where $J_{h}$ is defined immediately below.
\end{thm}
\begin{rem}
For any subsequence $\{p_{n}\}_{n\geq1}$ of $\{n\}_{n\geq1}$, Theorem
\ref{thm:Jh-convergence} holds with $p_{n}$ in place of $n$ throughout
and $h_{p_{n}}$ in place of $h=h_{n}$.
\end{rem}
For $h=\infty$, $J_{h}$ is the $N(0,1)$ distribution. For $h\in[0,\infty)$,
\begin{equation}
J_{h}:=\int_{0}^{1}I_{f,h}(r)dW(r)\Big/\left(\int_{0}^{1}I_{f,h}(r)^{2}dr\right)^{1/2},\label{eq:J_hintegral}
\end{equation}
where $I_{f,h}(r)$ is defined as follows.

Let $W(\cdot)$ denote a standard Brownian motion on $[0,1]$. For
$h\in(0,\infty)$, define $f_{h}(r)=(1,\exp(-hr))'$, and for $h=0$,
let $f_{h}(r)=(1,r)'$. Define
\begin{align}
I_{h}(r) & :=\begin{cases}
\int_{0}^{r}\exp(-(r-s)h)dW(s) & \text{for }h>0\\
W(r) & \text{for }h=0
\end{cases}\quad\text{ and}\nonumber \\
I_{f,h}(r) & :=I_{h}(r)-\left(\int_{0}^{1}I_{h}(s)f_{h}(s)ds\right)'\left(\int_{0}^{1}f_{h}(s)f_{h}(s)'ds\right)^{-1}f_{h}(r),\label{eq:I_fh}
\end{align}
where the subscript $f$ indicates that we take the residual process
of $I_{h}(\cdot)$ after projecting it onto the space spanned by $f_{h}(\cdot)$.
When $h>0$,
\begin{align}
I_{f,h}(r)=\  & I_{h}(r)-\alpha_{h}(r)\int_{0}^{1}I_{h}(s)ds-\beta_{h}(r)\int_{0}^{1}\frac{1-e^{-hs}}{h}I_{h}(s)ds,
\end{align}
where
\begin{align}
\alpha_{h}(r) & =\ \frac{h[1-e^{-2h}-2(1-e^{-h})e^{-hr}+2he^{-hr}-2(1-e^{-h})]}{h(1-e^{-2h})-2(1-e^{-h})^{2}}\quad\text{and}\nonumber \\
\beta_{h}(r) & =\ \frac{2h^{2}[he^{-hr}-(1-e^{-h})]}{h(1-e^{-2h})-2(1-e^{-h})^{2}}.
\end{align}
When $h=0$, $I_{f,h}(r)$ reduces to detrended Brownian motion:
\begin{equation}
W_{f}(r):=W(r)-(4-6r)\int_{0}^{1}W(s)ds-(12r-6)\int_{0}^{1}sW(s)ds.
\end{equation}
The following lemma establishes the connection between $I_{f,h}(r)$
for $h>0$ and $h=0$.
\begin{lem}
\label{lem:continuous-Ifh} Let $h\downarrow0$, we have
\[
\sup_{r\in[0,1]}|\alpha_{h}(r)-(4-6r)|\to0\quad\text{and}\quad\sup_{r\in[0,1]}|\beta_{h}(r)-(12r-6)|\to0.
\]
\end{lem}
\begin{rem}
This lemma shows that $I_{f,h}(r)$ is uniformly continuous with respect
to $h$. Consequently, $I_{f,h}(r)\to I_{f,0}(r)$ uniformly over
$r\in[0,1]$ as $h\to0$ a.s.\  and the asymptotic distribution is
continuous at $h=0$.
\end{rem}
\newpage{}

\printbibliography[title={\centering\makebox[\textwidth]{References}}]

\clearpage{}

\newpage

\makeatletter
\setcounter{table}{0}
\makeatother

\clearpage
\thispagestyle{empty}

\setcounter{footnote}{0}

\begin{center}
\vspace*{2cm}

{\LARGE Supplemental Material}\\[1.5em]

{\LARGE to}\\[1.5em]

{\LARGE Initial-Condition-Robust Inference}\\[0.5em]
{\LARGE in Autoregressive Models}\\[3em]

\large
Donald W.\ K.\ Andrews\footnote{
Andrews: Department of Economics and Cowles Foundation, Yale University.
[email removed].}
\qquad
Ming Li\footnote{
Li: Department of Economics and Risk Management Institute,
National University of Singapore.
[email removed].}
\qquad
Yapeng Zheng\footnote{
Zheng: Department of Economics,
Chinese University of Hong Kong.
[email removed].}\\[1.5em]

\large
\today

\end{center}

\clearpage
\setcounter{page}{1}


\newpage{}