Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
35,935 characters · 7 sections · 12 citation commands
Initial-Condition-Robust Inference in Autoregressive Models
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long \global\long \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\mc#1{\mathscr{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\bracket#1#2{\left\langle #1\middle\vert#2\right\rangle } \global\long\def\sandvich#1#2#3{\left\langle #1\middle\vert#2\middle\vert#3\right\rangle } \global\long\def\turd#1{\frac{#1}{3}} \global\long\global\long\def\sand#1{\left\lceil #1\right\vert } \global\long\def\wich#1{\left\vert #1\right\rfloor } \global\long\def\sandwich#1#2#3{\left\lceil #1\middle\vert#2\middle\vert#3\right\rfloor } \global\long\def\abs#1{\left|#1\right|} \global\long\def\norm#1{\left\Vert #1\right\Vert } \global\long\def\rest#1{\left.#1\right|} \global\long\def\inprod#1{\left\langle #1\right\rangle } \global\long\def\ol#1{\overline{#1}} \global\long\def\ul#1{#1} \global\long\def\td#1{\widetilde{#1}} \global\long\def\bs#1{\boldsymbol{#1}} \global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long\global\long
\thispagestyle{empty}
\setcounter{page}{1}
We consider an AR model with an AR parameter that may take values close or equal to one. Existing CIs in the literature for the AR parameter assume that the initial condition is stationary, zero, or fixed, e.g., see stock1991confidence, andrews1993exactly, andrews1994approximately, hansen1999grid, elliott2001confidence, mikusheva2007uniform, and andrews2014conditional. If the initial-condition assumption is not satisfied, then these CIs do not have correct asymptotic coverage probabilities and their finite sample coverage probabilities can be poor. For example, simulations reported below show that the nominal 95% AG14 CI has finite sample coverage probabilities ranging from 24.1% to 93.5% across 50 cases with highly variable initial conditions and different types of conditional heteroskedasticity of the error. A quarter of these CPs are 79.0% or less.
To circumvent the sensitivity of existing CIs to the initial condition, this paper introduces a new initial-condition-robust (ICR) CI whose finite sample CPs do not depend on the initial condition. These CIs are easy to compute and do not depend on any tuning parameters. We show that the ICR CI has correct asymptotic size in a uniform sense under conditions that allow for an arbitrary initial condition and conditional heteroskedasticity. Simulations show that the finite sample CPs of the nominal 95% ICR CI are quite good. They lie between 93.5% and 95.0% in the cases considered.
Like the other CIs referenced above, the ICR CI is constructed by inverting hypothesis tests concerning the value of the AR parameter. For all of these CIs, the tests are based on a t statistic that depends on a least squares (LS) estimator of the AR parameter. In contrast to existing CIs, the LS estimator used by the ICR tests contains an additional regressor that eliminates the effect of the initial condition under the null hypothesis. (For the definition of this regressor, see Section (ref) below.) In consequence, the ICR CI has a CP that is invariant to the value of the initial condition. However, the length of the ICR CI can be effected by the initial condition. Simulations show that the ICR CI is slightly shorter when the initial condition is highly variable than when it is stationary or fixed.\footnote{The length of the ICR CI varies mostly with the value of the AR parameter. It is shortest for AR parameters near one. This is true of existing CIs as well.}
In scenarios where existing CIs are asymptotically valid (i.e., the initial condition is stationary or fixed), the ICR CI pays a price in terms of its length compared to existing CIs. This occurs because the ICR LS estimator has a higher variance due to the additional regressor that is included in the LS regression. However, simulations show that the effect is fairly small. Across 50 scenarios with stationary or fixed initial conditions and i.i.d., GARCH, or ARCH errors, the ratio of the expected length of the ICR CI compared to that of the AG14 CI is found to vary between 1.00 and 1.11. On average, the ICR CI is 3.5% longer than the AG14 CI in these scenarios.
Based on the ICR CI, we also introduce an asymptotically median-unbiased interval estimator (MUE) of the AR parameter.
The AR model considered here has been applied in the literature to exchange rate, commodity and stock prices, and other economic time series, e.g., see kim1993unit.
This paper is organized as follows. Section (ref) specifies the AR model and defines the ICR CI. Section (ref) defines the median-unbiased interval estimator. Section (ref) provides simulation results concerning CPs and average lengths of the ICR and AG14 CIs and absolute median biases of the ICR\ MUE. Section (ref) provides asymptotic results for the ICR CI. The Supplemental Material describes how the critical values were computed, proves Theorems (ref) and (ref), and provides some additional simulation results.
We consider the AR(1) model with conditional-heteroskedasticity studied in AG14:
where the observed data are $\{Y_{0},\dots,Y_{n}\}$, $\rho\in[-1+\varepsilon,1]$ for some $0<\varepsilon<2$, and $\{U_{i}:i=1,2,\dots\}$ are stationary and ergodic under the distribution $F$, with conditional mean $0$ given a $\sigma$-field $\mathcal{G}_{i-1}$ for which $U_{j}\in\mathcal{G}_{i}$ for all $j\leq i$, conditional variance $\sigma_{i}^{2}=E_{F}(U_{i}^{2}|\mathcal{G}_{i-1})$, and unconditional variance $\sigma_{U}^{2}\in(0,\infty)$. A detailed description of the parameter space $\Lambda$ is provided in Section (ref).
By iterative substitution in ((ref)), we have
Let $Y$, $U$, and $X_{1}$ be $n$-vectors whose $i$-th elements are $Y_{i}$, $U_{i}$, and $Y_{i-1}$, respectively. Let $X_{2}(\rho)$ be an $n\times2$ matrix whose $i$-th row is $(1,\rho^{i-1})$ when $-1<\rho<1$.\footnote{Following the convention $0^{0}=1$, we use $\rho^{i-1}$ in the regression instead of $\rho^{i}$ to avoid perfect multicollinearity when $\rho=0$.} When $\rho=1$, we define $X_{2}(\rho)$ to be an $n\times2$ matrix whose $i$-th row is $(1,i)$. Let $X(\rho)=[X_{1}:X_{2}(\rho)]$, $P_{X(\rho)}=X(\rho)(X(\rho)'X(\rho))^{-1}X(\rho)'$, and $M_{X(\rho)}=I_{n}-P_{X(\rho)}$. Let $\widehat{U}_{i}$ denote the $i$-th element of the residual vector $M_{X(\rho)}Y$. Let $p_{ii}$ denote the $i$-th diagonal element of $P_{X(\rho)}$, and define $p_{ii}^{*}=\min\{p_{ii},n^{-1/2}\}$. Let $\Delta$ be a diagonal $n\times n$ matrix whose $i$-th diagonal element is $\widehat{U}_{i}/(1-p_{ii}^{*})$. Then, the ICR LS estimator $\widehat{\rho}_{n}(\rho)$, and the HC5 variance estimator $\widehat{\sigma}_{n}^{2}(\rho)$ are defined as
respectively.
Let
For suitable sequences $\{\rho_{n}\}_{n\geq1}$ such that $n(1-\rho_{n})\to h\in[0,\infty]$, we show that
where $J_{h}$ is defined in Section (ref) below. Table (ref) provides the quantiles $c_{h}(\alpha/2)$ and $c_{h}(1-\alpha/2)$ of the distribution of $J_{h}$ for $\alpha=.05$ and $.1$, which are the critical values employed by the ICR CI.
The nominal $1-\alpha$ equal-tailed two-sided ICR CI for $\rho$ is
which can be computed by taking a fine grid of values $\rho\in[-1+\epsilon,1]$ and comparing $T_{n}(\rho)$ to $c_{h}(\alpha/2)$ and $c_{h}(1-\alpha/2)$. Given these critical values, computation of the CIs is fast.
By definition, an estimator $\widehat{\theta}_{n}$ of a parameter $\theta$ is median unbiased if $P(\widehat{\theta}_{n}\geq\theta)\geq1/2$ and $P(\widehat{\theta}_{n}\leq\theta)\geq1/2$. In this section, we introduce an ICR MUE of $\rho$ that satisfies an analogous condition. Also, with probability close to one, this interval estimator is a point.\footnote{If the median of the asymptotic null distribution of the $t$-statistic is strictly decreasing in $\rho$, or equivalently, if $c_{h}(0.5)$ is strictly increasing in $h$, then the proposed interval estimator would be a point estimator with probability one. Because this condition fails to hold exactly, but almost holds, there is a very small probability that the estimator is a short interval rather than a point.}
Let $CI_{\text{ICR},n}^{\text{up}}(.5)$ and $CI_{\text{ICR},n}^{\text{low}}(.5)$ denote level $.5$ one-sided upper-bound and lower-bound CIs for $\rho$, respectively. By definition,
The MUE $\widetilde{\rho}_{n}$ of $\rho$ is defined by
By construction, we have $\widetilde{\rho}_{n}^{\text{low}}\leq\widetilde{\rho}_{n}^{\text{up}}$.\footnote{This holds because $\widetilde{\rho}_{n}^{\text{up}}\geq\sup\{\rho\in[-1+\epsilon,1]:c_{h}(.5)=T_{n}(\rho)\}$ and $\widetilde{\rho}_{n}^{\text{low}}$ is less than or equal to the infimum of the values in the same set.} In addition, $\widetilde{\rho}_{n}$ is a singleton whenever the set $\{\rho\in[-1+\epsilon,1]:T_{n}(\rho)=c_{h}(.5)\}$ contains a single point, in which case $\widetilde{\rho}_{n}$ equals this point. Table (ref) provides the critical values $c_{h}(.5)$ for a wide range of $h$. Given these critical values, computation of $\widetilde{\rho}_{n}$ is fast.
We compare the coverage probabilities (CPs) and average lengths (ALs) of the ICR and AG14 CIs and report the absolute median biases (ABMs) of the ICR MUE using Monte Carlo simulations. In Section \hyperref[SM-sec:Additional-Simulation-Results]{\textcolor{black}{\ref*{SM-sec:Additional-Simulation-Results}}} of the Supplemental Material, we compare the performance of the ICR and mikusheva2007uniform CIs. We focus on the nominal 95% equal-tailed two-sided CIs. For the AG14 CI, we follow the paper's constructions of the test statistic and critical values. To calculate the MUE for $\rho$, we follow the procedure described in Section (ref) and use $\widetilde{\rho}_{n}$ as the MUE of $\rho$ when $\widetilde{\rho}_{n}$ is a singleton. When $\widetilde{\rho}_{n}^{\text{up}}\neq\widetilde{\rho}_{n}^{\text{low}}$, we take $\widetilde{\rho}_{n}^{\text{up}}$ as the MUE for $\rho$ following andrews2025inference.
We consider a wide range of $\rho$ values: 0, 0.5, 0.7, 0.9, and 0.99. The innovations are of the form $U_{i}=\sigma_{i}\varepsilon_{i}$, where $\{\varepsilon_{i}:i=0,\pm1,\dots\}$ are identically and independently distributed (i.i.d.) standard normal and $\sigma_{i}$ is the multiplicative conditional heteroskedasticity. Let GARCH-$(ma,ar;\psi)$ denote a GARCH(1, 1) process with MA, AR, and intercept parameters $(ma,ar;\psi)$, and let ARCH-$(ar_{1},\dots,ar_{4};\psi)$ denote an ARCH(4) process with AR parameters $(ar_{1},\dots,ar_{4})$ and intercept $\psi$. Following AG14, we consider five specifications for the conditional heteroskedasticity of the innovations: (a) i.i.d. $N(0,1)$ (“i.i.d.”), (b) GARCH(1,1)-(.05,.9;.001) (“GARCH1”), (c) GARCH(1,1)-(.15,.8;.2) (“GARCH2”), (d) GARCH(1,1)-(.25,.7;.2) (“GARCH3”), and (e) ARCH(4)-(.3,.2,.2,.2;.2) (“ARCH”). For the initial conditions, we consider the following four specifications: (a) Fixed: $Y_{0}^{*}=0$, (b) Stationary: $Y_{0}^{*}=\sum_{i=0}^{\infty}\rho^{i}U_{-i}$, (c) Scaled $n$: $Y_{0}^{*}\sim\sqrt{n}\cdot\sum_{i=0}^{\infty}\rho^{i}U_{-i}$, and (d) Explosive: $Y_{0}^{*}\sim n^{3/4}\cdot\sum_{i=0}^{\infty}\rho^{i}U_{-i}$. We consider a sample size of $n=150$. All simulation results are based on 30,000 simulation repetitions.
Tables (ref) and (ref) report the CPs of AG14 and ICR CIs, respectively. The CPs of the AG14 CI in Table (ref) are close to the nominal level of 95% when the initial condition is fixed at zero or stationary, but fall below 95% when the initial condition is more variable, namely in the scaled $n$ and explosive $Y_{0}^{*}$ cases. In particular, the AG14 CI is robust to non-i.i.d.\ innovations when the initial condition is fixed or stationary. On the other hand, for scaled $n$ and explosive $Y_{0}^{*}$ cases, the AG14 CPs lie in the interval {[}24.1, 93.5{]} with roughly a quarter of the CPs being less than 79.0. Table (ref) shows that the ICR CI achieves CPs close to 95% in all of the scenarios considered and for any initial conditions. Specifically, all ICR CI CPs lie in the interval {[}93.5, 95.0{]}.
Table (ref) reports the ratios of the ALs of the nominal 95% ICR CI to the AG14 CI. We include only cases with fixed or stationary initial conditions because the AG14 CI is not robust to more variable initial conditions, whereas the ICR CI is. The ratios are between 1.00 and 1.11. On average, the ICR CI is 3.5% longer than the AG14 CI over the fifty cases considered. Thus, the price the ICR CI pays in terms of AL to gain robustness against the distribution of the initial condition is fairly small.
Table (ref) presents the ALs of the ICR CI. These are substantially decreasing in $\rho$, as expected, longer for ARCH conditional heteroskedasticity than i.i.d., and increasing from GARCH1 to GARCH2 to GARCH3 conditional heteroskedasticity.
Table (ref) reports the AMBs of the ICR MUE. Across all cases, the AMBs range from 0.000 to 0.022, indicating that their magnitudes are generally small.
In Section \hyperref[SM-sec:Additional-Simulation-Results]{\textcolor{black}{\ref*{SM-sec:Additional-Simulation-Results}}} of the Supplemental Material, we compare the performance of the ICR and Mik07 CIs via simulations. The results are similar to those reported above for the ICR and AG14 CIs except that the Mik07 CI undercovers both under conditional heteroskedasticity, and/or scaled $n$ or explosive initial conditions, but by a lower magnitude in the latter case than the AG14 CI. Again, we find that the ICR CI is robust to more variable initial conditions while paying only a small price in terms of CI width in the situations where the Mik07 CI has correct asymptotic coverage, i.e., when the innovations are i.i.d.\ and the initial conditions are stationary (or fixed).
This section establishes the correct uniform asymptotic size and asymptotic similarity of the ICR CI for $\rho$.
Following AG14, the parameter space for $(\rho,F)$ is given by:
$\Lambda=\Big\{\lambda=(\rho,F)$: (i) $\rho\in[-1+\epsilon,1]$;
for some constants $0<\epsilon<2$, $\zeta>3$, and $C<\infty$$\Big\}$.
We suppose $E_{F}U_{1}^{2}=1$ for simplicity because the asymptotic distribution of $T_{n}(\rho)$ does not depend on the unconditional variance. Parts (ref) and (ref) are slightly different from the corresponding conditions in AG14 because the random variables $\{U_{-j},\ j\geq0\}$ do not enter into the asymptotic analysis of the ICR CI. The key innovation of this paper is that we allow arbitrary initial conditions $Y_{0}^{*}$, including those corresponding to explosive processes.
The main theoretical result of this paper shows that $CI_{\text{ICR},n}$ has correct asymptotic size for the parameter space $\Lambda$ and is asymptotically similar. Let $P_{\lambda}$ denote probability under $\lambda=\left(\rho,F\right)\in\Lambda$.
The MUE $\widetilde{\rho}_{n}$ has the following median-unbiasedness property. This result is a corollary to the one-sided versions of Theorem (ref).\footnote{Corollary (ref) holds because $CI_{\text{ICR},n}^{\text{up}}(.5)$ and $CI_{\text{ICR},n}^{\text{low}}(.5)$ both have coverage probabilities of $1/2$ or greater by the proof of Theorem (ref) applied to these one-sided CIs and $(\widetilde{\rho}_{n}\geq\rho)\supset CI_{\text{ICR},n}^{\text{up}}(.5)$ and $(\widetilde{\rho}_{n}\leq\rho)\supset CI_{\text{ICR},n}^{\text{low}}(.5)$.}
As stated in ((ref)), $J_{h}$ is the asymptotic distribution of $T_{n}(\rho_{n})$ under suitable sequences $\{\rho_{n}:n\geq1\}$. Here, we state this result formally as Theorem (ref) and define $J_{h}$. This result is proved in a similar way to Theorem 1 in andrews2012asymptotics. Using Theorem 2.1 in andrews2020generic, Theorem (ref) is sufficient to establish Theorem (ref), which is proved in Section \hyperref[SM-sec:Proof-of-Theorem1]{\textcolor{black}{\ref*{SM-sec:Proof-of-Theorem1}}} of the Supplemental Material.
For $h=\infty$, $J_{h}$ is the $N(0,1)$ distribution. For $h\in[0,\infty)$,
where $I_{f,h}(r)$ is defined as follows.
Let $W(\cdot)$ denote a standard Brownian motion on $[0,1]$. For $h\in(0,\infty)$, define $f_{h}(r)=(1,\exp(-hr))'$, and for $h=0$, let $f_{h}(r)=(1,r)'$. Define
where the subscript $f$ indicates that we take the residual process of $I_{h}(\cdot)$ after projecting it onto the space spanned by $f_{h}(\cdot)$. When $h>0$,
where
When $h=0$, $I_{f,h}(r)$ reduces to detrended Brownian motion:
The following lemma establishes the connection between $I_{f,h}(r)$ for $h>0$ and $h=0$.
\printbibliography[title={\makebox[\textwidth]{References}}]
\makeatletter \setcounter{table}{0} \makeatother
\thispagestyle{empty}
\setcounter{footnote}{0}
\setcounter{page}{1}