EconBase
← Back to paper

Initial-Condition-Robust Inference in Autoregressive Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

35,935 characters · 7 sections · 12 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Initial-Condition-Robust Inference in Autoregressive Models

\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long \global\long \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle } \global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle } \global\long\def\turd#1{\frac{#1}{3}} \global\long\global\long\def\sand#1{\left\lceil #1\right\vert } \global\long\def\wich#1{\left\vert #1\right\rfloor } \global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor } \global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\inprod#1{\left\langle #1\right\rangle } \global\long\def\ol#1{\overline{#1}} \global\long\def\ul#1{#1} \global\long\def\td#1{\widetilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long

abstractThis paper considers confidence intervals (CIs) for the autoregressive (AR) parameter in an AR model with an AR parameter that may be close or equal to one. Existing CIs rely on the assumption of a stationary or fixed initial condition to obtain correct asymptotic coverage and good finite sample coverage. When this assumption fails, their coverage can be quite poor. In this paper, we introduce a new CI for the AR parameter whose coverage probability is completely robust to the initial condition, both asymptotically and in finite samples. This CI pays only a small price in terms of its length when the initial condition is stationary or fixed. The new CI also is robust to conditional heteroskedasticity of the errors. Keywords: Asymptotic size, autoregressive model, confidence set, initial condition, robustness. JEL Classification Numbers: C10, C12.

\thispagestyle{empty}

\setcounter{page}{1}

Introduction

We consider an AR model with an AR parameter that may take values close or equal to one. Existing CIs in the literature for the AR parameter assume that the initial condition is stationary, zero, or fixed, e.g., see stock1991confidence, andrews1993exactly, andrews1994approximately, hansen1999grid, elliott2001confidence, mikusheva2007uniform, and andrews2014conditional. If the initial-condition assumption is not satisfied, then these CIs do not have correct asymptotic coverage probabilities and their finite sample coverage probabilities can be poor. For example, simulations reported below show that the nominal 95% AG14 CI has finite sample coverage probabilities ranging from 24.1% to 93.5% across 50 cases with highly variable initial conditions and different types of conditional heteroskedasticity of the error. A quarter of these CPs are 79.0% or less.

To circumvent the sensitivity of existing CIs to the initial condition, this paper introduces a new initial-condition-robust (ICR) CI whose finite sample CPs do not depend on the initial condition. These CIs are easy to compute and do not depend on any tuning parameters. We show that the ICR CI has correct asymptotic size in a uniform sense under conditions that allow for an arbitrary initial condition and conditional heteroskedasticity. Simulations show that the finite sample CPs of the nominal 95% ICR CI are quite good. They lie between 93.5% and 95.0% in the cases considered.

Like the other CIs referenced above, the ICR CI is constructed by inverting hypothesis tests concerning the value of the AR parameter. For all of these CIs, the tests are based on a t statistic that depends on a least squares (LS) estimator of the AR parameter. In contrast to existing CIs, the LS estimator used by the ICR tests contains an additional regressor that eliminates the effect of the initial condition under the null hypothesis. (For the definition of this regressor, see Section (ref) below.) In consequence, the ICR CI has a CP that is invariant to the value of the initial condition. However, the length of the ICR CI can be effected by the initial condition. Simulations show that the ICR CI is slightly shorter when the initial condition is highly variable than when it is stationary or fixed.\footnote{The length of the ICR CI varies mostly with the value of the AR parameter. It is shortest for AR parameters near one. This is true of existing CIs as well.}

In scenarios where existing CIs are asymptotically valid (i.e., the initial condition is stationary or fixed), the ICR CI pays a price in terms of its length compared to existing CIs. This occurs because the ICR LS estimator has a higher variance due to the additional regressor that is included in the LS regression. However, simulations show that the effect is fairly small. Across 50 scenarios with stationary or fixed initial conditions and i.i.d., GARCH, or ARCH errors, the ratio of the expected length of the ICR CI compared to that of the AG14 CI is found to vary between 1.00 and 1.11. On average, the ICR CI is 3.5% longer than the AG14 CI in these scenarios.

Based on the ICR CI, we also introduce an asymptotically median-unbiased interval estimator (MUE) of the AR parameter.

The AR model considered here has been applied in the literature to exchange rate, commodity and stock prices, and other economic time series, e.g., see kim1993unit.

This paper is organized as follows. Section (ref) specifies the AR model and defines the ICR CI. Section (ref) defines the median-unbiased interval estimator. Section (ref) provides simulation results concerning CPs and average lengths of the ICR and AG14 CIs and absolute median biases of the ICR\ MUE. Section (ref) provides asymptotic results for the ICR CI. The Supplemental Material describes how the critical values were computed, proves Theorems (ref) and (ref), and provides some additional simulation results.

Model and ICR Confidence Interval

We consider the AR(1) model with conditional-heteroskedasticity studied in AG14:

align[align omitted — 175 chars of source]

where the observed data are $\{Y_{0},\dots,Y_{n}\}$, $\rho\in[-1+\varepsilon,1]$ for some $0<\varepsilon<2$, and $\{U_{i}:i=1,2,\dots\}$ are stationary and ergodic under the distribution $F$, with conditional mean $0$ given a $\sigma$-field $\mathcal{G}_{i-1}$ for which $U_{j}\in\mathcal{G}_{i}$ for all $j\leq i$, conditional variance $\sigma_{i}^{2}=E_{F}(U_{i}^{2}|\mathcal{G}_{i-1})$, and unconditional variance $\sigma_{U}^{2}\in(0,\infty)$. A detailed description of the parameter space $\Lambda$ is provided in Section (ref).

By iterative substitution in ((ref)), we have

align[align omitted — 147 chars of source]

Let $Y$, $U$, and $X_{1}$ be $n$-vectors whose $i$-th elements are $Y_{i}$, $U_{i}$, and $Y_{i-1}$, respectively. Let $X_{2}(\rho)$ be an $n\times2$ matrix whose $i$-th row is $(1,\rho^{i-1})$ when $-1<\rho<1$.\footnote{Following the convention $0^{0}=1$, we use $\rho^{i-1}$ in the regression instead of $\rho^{i}$ to avoid perfect multicollinearity when $\rho=0$.} When $\rho=1$, we define $X_{2}(\rho)$ to be an $n\times2$ matrix whose $i$-th row is $(1,i)$. Let $X(\rho)=[X_{1}:X_{2}(\rho)]$, $P_{X(\rho)}=X(\rho)(X(\rho)'X(\rho))^{-1}X(\rho)'$, and $M_{X(\rho)}=I_{n}-P_{X(\rho)}$. Let $\widehat{U}_{i}$ denote the $i$-th element of the residual vector $M_{X(\rho)}Y$. Let $p_{ii}$ denote the $i$-th diagonal element of $P_{X(\rho)}$, and define $p_{ii}^{*}=\min\{p_{ii},n^{-1/2}\}$. Let $\Delta$ be a diagonal $n\times n$ matrix whose $i$-th diagonal element is $\widehat{U}_{i}/(1-p_{ii}^{*})$. Then, the ICR LS estimator $\widehat{\rho}_{n}(\rho)$, and the HC5 variance estimator $\widehat{\sigma}_{n}^{2}(\rho)$ are defined as

align[align omitted — 320 chars of source]

respectively.

Let

equation[equation omitted — 128 chars of source]

For suitable sequences $\{\rho_{n}\}_{n\geq1}$ such that $n(1-\rho_{n})\to h\in[0,\infty]$, we show that

equation[equation omitted — 83 chars of source]

where $J_{h}$ is defined in Section (ref) below. Table (ref) provides the quantiles $c_{h}(\alpha/2)$ and $c_{h}(1-\alpha/2)$ of the distribution of $J_{h}$ for $\alpha=.05$ and $.1$, which are the critical values employed by the ICR CI.

The nominal $1-\alpha$ equal-tailed two-sided ICR CI for $\rho$ is

equation[equation omitted — 171 chars of source]

which can be computed by taking a fine grid of values $\rho\in[-1+\epsilon,1]$ and comparing $T_{n}(\rho)$ to $c_{h}(\alpha/2)$ and $c_{h}(1-\alpha/2)$. Given these critical values, computation of the CIs is fast.

table[table omitted — 3,295 chars of source]

ICR Median-Unbiased Interval Estimator

By definition, an estimator $\widehat{\theta}_{n}$ of a parameter $\theta$ is median unbiased if $P(\widehat{\theta}_{n}\geq\theta)\geq1/2$ and $P(\widehat{\theta}_{n}\leq\theta)\geq1/2$. In this section, we introduce an ICR MUE of $\rho$ that satisfies an analogous condition. Also, with probability close to one, this interval estimator is a point.\footnote{If the median of the asymptotic null distribution of the $t$-statistic is strictly decreasing in $\rho$, or equivalently, if $c_{h}(0.5)$ is strictly increasing in $h$, then the proposed interval estimator would be a point estimator with probability one. Because this condition fails to hold exactly, but almost holds, there is a very small probability that the estimator is a short interval rather than a point.}

Let $CI_{\text{ICR},n}^{\text{up}}(.5)$ and $CI_{\text{ICR},n}^{\text{low}}(.5)$ denote level $.5$ one-sided upper-bound and lower-bound CIs for $\rho$, respectively. By definition,

align[align omitted — 328 chars of source]

The MUE $\widetilde{\rho}_{n}$ of $\rho$ is defined by

align[align omitted — 357 chars of source]

By construction, we have $\widetilde{\rho}_{n}^{\text{low}}\leq\widetilde{\rho}_{n}^{\text{up}}$.\footnote{This holds because $\widetilde{\rho}_{n}^{\text{up}}\geq\sup\{\rho\in[-1+\epsilon,1]:c_{h}(.5)=T_{n}(\rho)\}$ and $\widetilde{\rho}_{n}^{\text{low}}$ is less than or equal to the infimum of the values in the same set.} In addition, $\widetilde{\rho}_{n}$ is a singleton whenever the set $\{\rho\in[-1+\epsilon,1]:T_{n}(\rho)=c_{h}(.5)\}$ contains a single point, in which case $\widetilde{\rho}_{n}$ equals this point. Table (ref) provides the critical values $c_{h}(.5)$ for a wide range of $h$. Given these critical values, computation of $\widetilde{\rho}_{n}$ is fast.

Monte Carlo Simulations

We compare the coverage probabilities (CPs) and average lengths (ALs) of the ICR and AG14 CIs and report the absolute median biases (ABMs) of the ICR MUE using Monte Carlo simulations. In Section \hyperref[SM-sec:Additional-Simulation-Results]{\textcolor{black}{\ref*{SM-sec:Additional-Simulation-Results}}} of the Supplemental Material, we compare the performance of the ICR and mikusheva2007uniform CIs. We focus on the nominal 95% equal-tailed two-sided CIs. For the AG14 CI, we follow the paper's constructions of the test statistic and critical values. To calculate the MUE for $\rho$, we follow the procedure described in Section (ref) and use $\widetilde{\rho}_{n}$ as the MUE of $\rho$ when $\widetilde{\rho}_{n}$ is a singleton. When $\widetilde{\rho}_{n}^{\text{up}}\neq\widetilde{\rho}_{n}^{\text{low}}$, we take $\widetilde{\rho}_{n}^{\text{up}}$ as the MUE for $\rho$ following andrews2025inference.

We consider a wide range of $\rho$ values: 0, 0.5, 0.7, 0.9, and 0.99. The innovations are of the form $U_{i}=\sigma_{i}\varepsilon_{i}$, where $\{\varepsilon_{i}:i=0,\pm1,\dots\}$ are identically and independently distributed (i.i.d.) standard normal and $\sigma_{i}$ is the multiplicative conditional heteroskedasticity. Let GARCH-$(ma,ar;\psi)$ denote a GARCH(1, 1) process with MA, AR, and intercept parameters $(ma,ar;\psi)$, and let ARCH-$(ar_{1},\dots,ar_{4};\psi)$ denote an ARCH(4) process with AR parameters $(ar_{1},\dots,ar_{4})$ and intercept $\psi$. Following AG14, we consider five specifications for the conditional heteroskedasticity of the innovations: (a) i.i.d. $N(0,1)$ (“i.i.d.”), (b) GARCH(1,1)-(.05,.9;.001) (“GARCH1”), (c) GARCH(1,1)-(.15,.8;.2) (“GARCH2”), (d) GARCH(1,1)-(.25,.7;.2) (“GARCH3”), and (e) ARCH(4)-(.3,.2,.2,.2;.2) (“ARCH”). For the initial conditions, we consider the following four specifications: (a) Fixed: $Y_{0}^{*}=0$, (b) Stationary: $Y_{0}^{*}=\sum_{i=0}^{\infty}\rho^{i}U_{-i}$, (c) Scaled $n$: $Y_{0}^{*}\sim\sqrt{n}\cdot\sum_{i=0}^{\infty}\rho^{i}U_{-i}$, and (d) Explosive: $Y_{0}^{*}\sim n^{3/4}\cdot\sum_{i=0}^{\infty}\rho^{i}U_{-i}$. We consider a sample size of $n=150$. All simulation results are based on 30,000 simulation repetitions.

Tables (ref) and (ref) report the CPs of AG14 and ICR CIs, respectively. The CPs of the AG14 CI in Table (ref) are close to the nominal level of 95% when the initial condition is fixed at zero or stationary, but fall below 95% when the initial condition is more variable, namely in the scaled $n$ and explosive $Y_{0}^{*}$ cases. In particular, the AG14 CI is robust to non-i.i.d.\ innovations when the initial condition is fixed or stationary. On the other hand, for scaled $n$ and explosive $Y_{0}^{*}$ cases, the AG14 CPs lie in the interval {[}24.1, 93.5{]} with roughly a quarter of the CPs being less than 79.0. Table (ref) shows that the ICR CI achieves CPs close to 95% in all of the scenarios considered and for any initial conditions. Specifically, all ICR CI CPs lie in the interval {[}93.5, 95.0{]}.

Table (ref) reports the ratios of the ALs of the nominal 95% ICR CI to the AG14 CI. We include only cases with fixed or stationary initial conditions because the AG14 CI is not robust to more variable initial conditions, whereas the ICR CI is. The ratios are between 1.00 and 1.11. On average, the ICR CI is 3.5% longer than the AG14 CI over the fifty cases considered. Thus, the price the ICR CI pays in terms of AL to gain robustness against the distribution of the initial condition is fairly small.

table[table omitted — 1,601 chars of source]
table[table omitted — 680 chars of source]
table[table omitted — 957 chars of source]
table[table omitted — 1,486 chars of source]

Table (ref) presents the ALs of the ICR CI. These are substantially decreasing in $\rho$, as expected, longer for ARCH conditional heteroskedasticity than i.i.d., and increasing from GARCH1 to GARCH2 to GARCH3 conditional heteroskedasticity.

table[table omitted — 1,595 chars of source]

Table (ref) reports the AMBs of the ICR MUE. Across all cases, the AMBs range from 0.000 to 0.022, indicating that their magnitudes are generally small.

In Section \hyperref[SM-sec:Additional-Simulation-Results]{\textcolor{black}{\ref*{SM-sec:Additional-Simulation-Results}}} of the Supplemental Material, we compare the performance of the ICR and Mik07 CIs via simulations. The results are similar to those reported above for the ICR and AG14 CIs except that the Mik07 CI undercovers both under conditional heteroskedasticity, and/or scaled $n$ or explosive initial conditions, but by a lower magnitude in the latter case than the AG14 CI. Again, we find that the ICR CI is robust to more variable initial conditions while paying only a small price in terms of CI width in the situations where the Mik07 CI has correct asymptotic coverage, i.e., when the innovations are i.i.d.\ and the initial conditions are stationary (or fixed).

Asymptotic Results

This section establishes the correct uniform asymptotic size and asymptotic similarity of the ICR CI for $\rho$.

Parameter Space, Correct Uniform Size, and Asymptotic \\Similarity

Following AG14, the parameter space for $(\rho,F)$ is given by:

$\Lambda=\Big\{\lambda=(\rho,F)$: (i) $\rho\in[-1+\epsilon,1]$;

enumerate[start=2, nosep, leftmargin=4.5em] • $\{U_{i}:i=1,2,\dots\}$ are stationary and strong-mixing under $F$ with $E_{F}(U_{i}|\mathcal{G}_{i-1})=0$ a.s., $E_{F}(U_{i}^{2}|\mathcal{G}_{i-1})=\sigma_{i}^{2}$ a.s., where $\mathcal{G}_{i}$ is some non-decreasing sequence of $\sigma$-fields for which $U_{j}\in\mathcal{G}_{i}$ for all $j\leq i$ for $i=1,2,\dots$; • The strong-mixing numbers $\{\alpha_{F}(m):m\geq1\}$ satisfy $\alpha_{F}(m)\leq Cm^{-3\zeta/(\zeta-3)}$, $\forall m\geq1$; • $\sup_{i,s,t}E_{F}|\prod_{a\in A}a|^{\zeta}\leq M$, where $0<i,s,t<\infty$, $i\geq\max(s,t)$, and $A$ is any non-empty subset of $\{U_{i-s},U_{i-t},U_{i+1}^{2},U_{1}^{2}\}$; • $E_{F}U_{1}^{2}=1$;

for some constants $0<\epsilon<2$, $\zeta>3$, and $C<\infty$$\Big\}$.

We suppose $E_{F}U_{1}^{2}=1$ for simplicity because the asymptotic distribution of $T_{n}(\rho)$ does not depend on the unconditional variance. Parts (ref) and (ref) are slightly different from the corresponding conditions in AG14 because the random variables $\{U_{-j},\ j\geq0\}$ do not enter into the asymptotic analysis of the ICR CI. The key innovation of this paper is that we allow arbitrary initial conditions $Y_{0}^{*}$, including those corresponding to explosive processes.

The main theoretical result of this paper shows that $CI_{\text{ICR},n}$ has correct asymptotic size for the parameter space $\Lambda$ and is asymptotically similar. Let $P_{\lambda}$ denote probability under $\lambda=\left(\rho,F\right)\in\Lambda$.

thmLet $\alpha\in(0,1)$. For the parameter space $\Lambda$, the nominal $1-\alpha$ ICR CI for the parameter $\rho$ satisfies \[ \underset{n\to\infty}{\lim\inf}\inf_{\lambda\in\Lambda}P_{\lambda}(\rho\in CI_{\text{ICR},n})=\underset{n\to\infty}{\lim\sup}\sup_{\lambda\in\Lambda}P_{\lambda}(\rho\in CI_{\text{ICR},n})=1-\alpha. \]

The MUE $\widetilde{\rho}_{n}$ has the following median-unbiasedness property. This result is a corollary to the one-sided versions of Theorem (ref).\footnote{Corollary (ref) holds because $CI_{\text{ICR},n}^{\text{up}}(.5)$ and $CI_{\text{ICR},n}^{\text{low}}(.5)$ both have coverage probabilities of $1/2$ or greater by the proof of Theorem (ref) applied to these one-sided CIs and $(\widetilde{\rho}_{n}\geq\rho)\supset CI_{\text{ICR},n}^{\text{up}}(.5)$ and $(\widetilde{\rho}_{n}\leq\rho)\supset CI_{\text{ICR},n}^{\text{low}}(.5)$.}

corThe MUE estimator $\widetilde{\rho}_{n}$ satisfies \begin{align*} \underset{n\to\infty}{\lim\inf}\inf_{\lambda\in\Lambda}P_{\lambda}(\widetilde{\rho}_{n}^{up}\geq\rho) & \geq1/2 and \underset{n\to\infty}{\lim\inf}\inf_{\lambda\in\Lambda}P_{\lambda}(\widetilde{\rho}_{n}^{low}\leq\rho)\geq1/2. \end{align*}

Asymptotic Distribution

As stated in ((ref)), $J_{h}$ is the asymptotic distribution of $T_{n}(\rho_{n})$ under suitable sequences $\{\rho_{n}:n\geq1\}$. Here, we state this result formally as Theorem (ref) and define $J_{h}$. This result is proved in a similar way to Theorem 1 in andrews2012asymptotics. Using Theorem 2.1 in andrews2020generic, Theorem (ref) is sufficient to establish Theorem (ref), which is proved in Section \hyperref[SM-sec:Proof-of-Theorem1]{\textcolor{black}{\ref*{SM-sec:Proof-of-Theorem1}}} of the Supplemental Material.

thmFor a sequence $\lambda_{n}=(\rho_{n},F_{n})$ such that $n(1-\rho_{n})\to h\in[0,\infty]$, we have \[ T_{n}(\rho_{n})=\frac{n^{1/2}(\widehat{\rho}_{n}(\rho_{n})-\rho_{n})}{\widehat{\sigma}_{n}(\rho_{n})}\stackrel{d}{\to} J_{h}, \] where $J_{h}$ is defined immediately below.
remFor any subsequence $\{p_{n}\}_{n\geq1}$ of $\{n\}_{n\geq1}$, Theorem (ref) holds with $p_{n}$ in place of $n$ throughout and $h_{p_{n}}$ in place of $h=h_{n}$.

For $h=\infty$, $J_{h}$ is the $N(0,1)$ distribution. For $h\in[0,\infty)$,

equation[equation omitted — 125 chars of source]

where $I_{f,h}(r)$ is defined as follows.

Let $W(\cdot)$ denote a standard Brownian motion on $[0,1]$. For $h\in(0,\infty)$, define $f_{h}(r)=(1,\exp(-hr))'$, and for $h=0$, let $f_{h}(r)=(1,r)'$. Define

align[align omitted — 293 chars of source]

where the subscript $f$ indicates that we take the residual process of $I_{h}(\cdot)$ after projecting it onto the space spanned by $f_{h}(\cdot)$. When $h>0$,

align[align omitted — 128 chars of source]

where

align[align omitted — 233 chars of source]

When $h=0$, $I_{f,h}(r)$ reduces to detrended Brownian motion:

equation[equation omitted — 83 chars of source]

The following lemma establishes the connection between $I_{f,h}(r)$ for $h>0$ and $h=0$.

lemLet $h\downarrow0$, we have \[ \sup_{r\in[0,1]}|\alpha_{h}(r)-(4-6r)|\to0\quad\text{and}\quad\sup_{r\in[0,1]}|\beta_{h}(r)-(12r-6)|\to0. \]
remThis lemma shows that $I_{f,h}(r)$ is uniformly continuous with respect to $h$. Consequently, $I_{f,h}(r)\to I_{f,0}(r)$ uniformly over $r\in[0,1]$ as $h\to0$ a.s.\ and the asymptotic distribution is continuous at $h=0$.

\printbibliography[title={\makebox[\textwidth]{References}}]

\makeatletter \setcounter{table}{0} \makeatother

\thispagestyle{empty}

\setcounter{footnote}{0}

center[center omitted — 591 chars of source]

\setcounter{page}{1}