Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
131,169 characters · 15 sections · 195 citation commands
Theory of Low Frequency Contamination from Nonstationarity and Misspecification: Consequences for HAR Inference
\setcounter{page}{0}
{\bf{JEL Classification}}: C12, C13, C18, C22, C32, C51\\ {\bf{Keywords}}: Edgeworth expansions, Fixed-$b$, HAC standard errors, HAR, Long memory, Long-run variance, Low frequency contamination, Nonstationarity, Outliers, Segmented locally stationary.
\onehalfspacing \thispagestyle{empty} \allowdisplaybreaks
The low frequency contamination can be explained as follows. For a short memory series, the autocorrelation function (ACF) displays exponential decay and vanishes as the lag length $k\rightarrow\infty$, and the periodogram is finite at the origin. Under general forms of nonstationarity involving changes in the mean, we show theoretically that $\widehat{\Gamma}\left(k\right)=\lim_{T\rightarrow\infty}\Gamma_{T}\left(k\right)+d^{*},$ where $\Gamma_{T}\left(k\right)=T^{-1}\sum_{t=k+1}^{T}\mathbb{E}\left(V_{t}V_{t-k}\right)$, $k\geq0$ and $d^{*}>0$ is independent of $k$. Assuming positive dependence for simplicity (i.e., $\lim_{T\rightarrow\infty}\Gamma_{T}\left(k\right)>0$), that means that each sample autocovariance overestimates the true dependence in the data. The bias factor $d^{*}>0$ depends on the type of nonstationarity and in general does not vanish as $T\rightarrow\infty$. In addition, since short memory implies $\Gamma_{T}\left(k\right)\rightarrow0$ as $k\rightarrow\infty$, it follows that $d^{*}$ generates long memory effects since $\widehat{\Gamma}\left(k\right)\thickapprox d^{*}>0$ as $k\rightarrow\infty$. As for the periodogram, $I_{T}\left(\omega\right)$, we show that under nonstationarity $\mathbb{E}\left(I_{T}\left(\omega\right)\right)\rightarrow\infty$ as $\omega\rightarrow0$, a feature also shared by long memory processes.
Recently, casini_hac proposed a new HAC estimator that applies nonparametric smoothing over time in order to account flexibly for nonstationarity. We show theoretically that nonparametric smoothing over time is robust to low frequency contamination and prove that the resulting sample local autocovariance and the local periodogram do not exhibit long memory features. Nonparametric smoothing avoids mixing highly heterogeneous data coming from distinct nonstationary regimes as opposed to what the sample autocovariance and the periodogram do.
Our work is different from the literature on spurious persistence caused by the presence of level shifts or other deterministic trends. perron:90 showed that the presence of breaks in mean often induces spurious non-rejection of the unit root hypothesis, and that the presence of a level shift asymptotically biases the estimate of the AR coefficient towards one. bhattacharya/gupta/waymire:83 demonstrated that certain deterministic trends can induce the spurious presence of long memory. In other contexts, similar issues were discussed by varneskov/christensen:17, diebold/inoue:01, demetrescu/salish:2020, lamoureux/lastrapes:1990, hillebrand:05, granger/hyung:04, mccloskey/hill:2017, mikosh/starica:04, muller/watson:2008 and perron/qu:2010. Our results are different from theirs in that we consider a more general problem and we allow for more general forms of nonstationarity using the segmented locally stationary framework of casini_hac. Importantly, we provide a general solution to these problems and show theoretically its robustness to low frequency contamination. Moreover, we discuss in detail the implications of our theory for HAR inference.
HAR inference relies on estimation of the long-run variance (LRV). The latter, from a time domain perspective, is equivalent to the sum of all autocovariances while from a frequency domain perspective, is equal to $2\pi$ times an integrated time-varying spectral density at the zero frequency. From a time domain perspective, estimation involves a weighted sum of the sample autocovariances, while from a frequency domain perspective estimation is based on a weighted sum of the periodogram ordinates near the zero frequency. Therefore, our results on low frequency contamination for the sample autocovariances and the periodogram can have important implications.
There are two main approaches in HAR inference, one based on traditional asymptotics and the other based on fixed-smoothing asymptotics. The classical approach relies on an LRV estimator using a small bandwidth {[}cf. the HAC estimators of newey/west:87 (newey/west:87, newey/west:94) and andrews:91{]}. Inference is standard because HAR test statistics follow asymptotically standard distributions. It was shown early that HAC standard errors can result in oversized tests when there is substantial temporal dependence. This stimulated a second approach based on an LRV estimator that keeps the bandwidth at a fixed fraction of the sample size and that converges weakly to a random variable {[}cf. Kiefer/vogelsang/bunzel:00{]}. Inference is then based on a nonstandard reference distribution and it is shown that fixed-$b$ achieves high-order refinements {[}e.g., sun/phillips/jin:08{]} and reduces the oversize problem of HAR tests.\footnote{See dou:18, hwang/sun:2017, ibragimov/kattuman/skrobotov:2021, ibragimov/muller:10, jansson:04, Kiefer/vogelsang:02 (Kiefer/vogelsang:02, kiefer/vogelsang:05), lazarus/lewis/stock:17, \textcolor{MyBlue}{Lazarus et al.} lazarus/lewis/stock/watson:18 muller:07 (muller:07, mueller:14), phillips:05, politis:11, potscher/preinerstorfer:18 (preinerstorfer/potscher:16, potscher/preinerstorfer:18, potscher/preinerstorfer:19), robinson:98, sun:14 (sun:13, sun:14a, sun:14) and zhang/shao:13.} However, unlike the classical approach, current fixed-$b$ HAR inference is only valid under stationarity {[}cf. casini_fixed_b_erp{]} as the fixed-$b$ limiting distribution of the $t$/$F$ statistic is non-pivotal under nonstationarity. More recently, a variant of the fixed-$b$ approach {[}see, e.g., sun:14 and lazarus/lewis/stock/watson:18{]} considered the use of small-$b$ asymptotics in conjunction with fixed-$b$ or $t/F$ critical values. These bandwidths are typically larger than the MSE-optimal bandwidths used for the HAC estimators.
Recently, casini_hac questioned the performance of HAR inference under nonstationarity from a theoretical standpoint. Simulation evidence of serious (e.g., non-monotonic) power or related issues in specific HAR inference contexts were documented by altissimo/corradi:2003, casini_CR_Test_Inst_Forecast, casini/perron_Lap_CR_Single_Inf (casini/perron:OUP-Breaks, casini/perron_SC_BP_Lap, casini/perron_Lap_CR_Single_Inf), chan:2020 (chan:2022, chan:2020), crainiceanu/vogelsang:07, deng/perron:06, juhl/xiao:09, kim/perron:09, martins/perron:16, otto/breitung:2021, perron:1991, perron/yamamoto:18, shao/zhang:2010, vogeslang:99 and zhang/lavitas:2018 among others{]}. Our theoretical results show that these issues occur because the unaccounted nonstationarity alters the spectrum at low frequencies. Each sample autocovariance is upward biased ($d^{*}>0$) and the resulting LRV estimators tend to be inflated. When these estimators are used to normalize test statistics, the latter lose power. Interestingly, $d^{*}$ is independent of $k$ so that the more lags are included the more severe is the problem. Further, by virtue of weak dependence, we have that $\Gamma_{T}\left(k\right)\rightarrow0$ as $k\rightarrow\infty$ but $d^{*}>0$ across $k$. We show formally that long bandwidths/fixed-$b$ LRV estimators are expected to suffer most from power losses because they use many/all lagged autocovariances.
To precisely analyze the theoretical properties of the HAR tests under the null hypothesis, we present second-order Edgeworth expansions under nonstationarity for the distribution of the HAC and DK-HAC estimator and for the distribution of the corresponding $t$-test in the linear regression model. Under stationarity the results concerning the HAC estimator were provided by velasco/robinson:01. We show that the order of the approximation error of the expansion is the same as under stationarity from which it follows that the error in rejection probability (ERP) is also the same. The ERP of the $t$-test based on the DK-HAC estimator is slightly larger than that of the $t$-test based on the HAC estimator due to the double smoothing. High-order asymptotic expansions for spectral and other estimates were studied by bhattacharya/ghosh:1978, bentkus/rudzkis:1982, janas:1994, phillips:1977 (phillips:1977, phillips:1980) and taniguchi/puri:1996. The asymptotic expansions of the fixed-$b$ HAR tests under stationarity were developed by jansson:04 and sun/phillips/jin:08. casini_fixed_b_erp showed that under nonstationarity the ERP of the fixed-$b$ HAR tests can be larger than that of HAR tests based on HAC and DK-HAC estimators thereby controverting the conclusion in the literature that the original fixed-$b$ HAR tests have superior null rejection rates relative to HAR tests based on traditional LRV estimators. casini_fixed_b_erp also developed fixed-$b$ methods that are valid under nonstationarity and in fact provide better null rejection rates in finite-sample.
The Monte Carlo results suggest that under the null hypothesis nonstationarity can generate larger size distortions than what one finds under stationarity. In particular, fixed-smoothing methods can exhibit under-rejections whereas HAC and DK-HAC methods can exhibit over-rejections when there is strong persistence. For the latter problem, our second-order Edgeworth expansions could be used to construct corrections to the standard normal critical value. We relegate this opportunity to future research.
The paper is organized as follows. Section (ref) presents the statistical setting and Section (ref) establishes the theoretical results on low frequency contamination. Section (ref) presents the Edgeworth expansions of HAR tests based on the HAC and DK-HAC estimators. The implications of our results for HAR inference are analyzed analytically and computationally through simulations in Section (ref). Section (ref) concludes. The supplemental materials {[}cf. casini/perron_Low_Frequency_Contam_Nonstat:2020_supp{]} contain some additional examples and all mathematical proofs.
Suppose $\{V_{t,T}\}_{t=1}^{T}$ is defined on a probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$, where $\Omega$ is the sample space, $\mathscr{F}$ is the $\sigma$-algebra and $\mathbb{P}$ is a probability measure. In order to analyze time series models that have a time-varying spectrum it is useful to introduce an infill asymptotic setting whereby we rescale the original discrete time horizon $\left[1,\,T\right]$ by dividing each $t$ by $T.$ Letting $u=t/T$ we define a new time scale $u\in\left[0,\,1\right]$ on which as $T\rightarrow\infty$ we observe more and more realizations of $V_{t,T}$ close to time $t$. As a notion of nonstationarity, we use the concept of segmented local stationarity (SLS) introduced in casini_hac. This extends the locally stationary processes {[}cf. dahlhaus:96{]} to allow for structural change and regime switching-type models. SLS processes allow for a finite number of discontinuities in the spectrum over time. We collect the break dates in the set $\mathcal{T}\triangleq\left\{ T_{1}^{0},\,\ldots,\,T_{m}^{0}\right\} $. Let $i\triangleq\sqrt{-1}.$ A function $G\left(\cdot,\,\cdot\right):\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ is said to be left-differentiable at $u_{0}$ if $\partial G\left(u_{0},\omega\right)/\partial_{-}u\triangleq\lim_{u\rightarrow u_{0}^{-}}\left(G\left(u_{0},\,\omega\right)-G\left(u,\,\omega\right)\right)/\left(u_{0}-u\right)$ exists for any $\omega\in\mathbb{R}$. Let $m_{0}\geq0$ be a finite integer.
Definition (ref) states that $V_{t,T}$ has a time-varying spectral representation where both the mean $\mu_{\cdot}\left(\cdot\right)$ and transfer function $A_{\cdot,\cdot,T}^{0}\left(\omega\right)$ are piecewise continuous. Since the transfer function depends on the parameters that enter the second moments of $V_{t,T}$, the smoothness properties of $\mu_{\cdot}\left(\cdot\right)$ and $A$ guarantee that $V_{t,T}$ has a piecewise locally stationary behavior. We require additional smoothness properties for $A$ and an example is presented at the end of this section.
We define the time-varying spectral density as $f_{j}\left(u,\,\omega\right)\triangleq(2\pi)^{-1}|A_{j}\left(u,\,\omega\right)|^{2}$ for $T_{j-1}^{0}/T<u=t/T\leq T_{j}^{0}/T$. Then we can define the local covariance of $V_{t,T}$ at the rescaled time $u$ with $Tu\notin\mathcal{T}$ and lag $k\in\mathbb{Z}$ as $c\left(u,\,k\right)\triangleq\int_{-\pi}^{\pi}e^{i\omega k}f\left(u,\,\omega\right)d\omega$. The same definition is also used when $Tu\in\mathcal{T}$ and $k\geq0$. For $Tu\in\mathcal{T}$ and $k<0$ it is defined as $c\left(u,\,k\right)\triangleq\lim_{T\rightarrow\infty}\int_{-\pi}^{\pi}e^{i\omega k}A\left(u,\,\omega\right)A\left(u-k/T,\,-\omega\right)d\omega$.
Next, we impose conditions on the temporal dependence (we omit the second subscript $T$ when it is clear from the context). Let
where $\left\{ V_{\mathscr{N},t}\right\} $ is a Gaussian sequence with the same mean and covariance structure as $\left\{ V_{t}\right\} $, $\kappa_{V,t}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}\left(u,\,v,\,w\right)$ is the time-$t$ fourth-order cumulant of $(V_{t}^{\left(a_{1}\right)},\,V_{t+u}^{\left(a_{2}\right)},\,V_{t+v}^{\left(a_{3}\right)},$ $\,V_{t+w}^{\left(a_{4}\right)})$ while $\kappa_{\mathscr{N}}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}$ $(t,\,t+u,\,t+v,\,t+w)$ is the time-$t$ centered fourth moment of $V_{t}$ if $V_{t}$ were Gaussian.
If $\left\{ V_{t}\right\} $ is stationary then the cumulant condition of Assumption (ref)-(i) reduces to the standard one used in the time series literature {[}see andrews:91{]}. Note that $\alpha$-mixing and some moment conditions imply that the cumulant condition of Assumption (ref) holds. Part (ii) extends the smoothness conditions on $A\left(u,\,\omega\right)$ in Assumption (ref) to the fourth-order cumulant. These smoothness conditions are not particularly restrictive.
Consider the following time-varying AR(1) process with one break at mid-sample $\lambda_{1}^{0}=0.5$,
where $\rho_{1}\left(\cdot\right)$ and $\rho_{2}\left(\cdot\right)$ are Lipschitz continuous, $\sigma\left(\cdot\right)$ is piecewise Lipschitz continuous and $\left\{ u_{t}\right\} $ are i.i.d. random variables with mean zero and unit variance. Then, $V_{t,T}$ is an SLS process with $A\left(u,\,\omega\right)=\sigma\left(u\right)\left(1+\rho\left(u\right)\exp\left(i\omega\right)\right)$. If $\rho\left(u\right)$ and $\sigma\left(u\right)$ satisfy the same smoothness conditions in $u$ required for $A\left(u,\,\omega\right)$ in Assumption (ref), $\sup_{u\in\left[0,\,1\right]}\left|\rho\left(u\right)\right|<1$ and $\sup_{u\in\left[0,\,1\right]}\sigma\left(u\right)<\infty$, then $V_{t,T}$ fulfills Assumption (ref)-(ref).
In this section we establish theoretical results about the low frequency contamination induced by nonstationarity, misspecification and outliers. We first consider the asymptotic proprieties of two key quantities for inference in time series contexts, i.e., the sample autocovariance and the periodogram. These are defined, respectively, by
where $\overline{V}$ is the sample mean and
which is evaluated at the Fourier frequencies $\omega_{j}=\left(2\pi j\right)/T\in[0,\,\pi]$. In the context of autocorrelated data, hypotheses testing and construction of confidence intervals require estimation of the so-called long-run variance. Traditional HAC estimators are weighted sums of sample autocovariances while frequency domain estimators are weighted sums of the periodograms. casini_hac considered an alternative estimate for the sample autocovariance to be used in the DK-HAC estimators, defined in Section (ref), namely,
where $k\in\mathbb{Z},$ $n_{T}\rightarrow\infty$ satisfying the conditions given below, and
with $\overline{V}{}_{rn_{T},T}=n_{2,T}^{-1}\sum_{s=0}^{n_{2,T}-1}V_{rn_{T}-n_{2,T}/2+s+1}$ and $n_{2,T}\rightarrow\infty$ such that $n_{2,T}/T\rightarrow0$. For notational simplicity we assume that $n_{T}$ and $n_{2,T}$ are even. $\widehat{c}_{T}\left(rn_{T}/T,\,k\right)$ is an estimate of the autocovariance at time $rn_{T}$ and lag $k$, i.e., $\mathrm{cov}(V_{rn_{T}},\,V_{rn_{T}-k})$. One could use a smoothed or tapered version; the estimate $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$ is an integrated local sample autocovariance. It extends $\widehat{\Gamma}\left(k\right)$ to better account for nonstationarity. Similarly, the DK-HAC estimator does not relate to the periodogram but to the local periodogram defined by
where $I_{\mathrm{L},T}\left(u,\,\omega\right)$ is the (untapered) periodogram over a segment of length $n_{T}$ with midpoint $\left\lfloor Tu\right\rfloor $. We also consider the statistical properties of both $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$ and $I_{\mathrm{L},T}\left(u,\,\omega\right)$ under nonstationarity. Define $r_{j}=(\lambda_{j}^{0}-\lambda_{j-1}^{0})$ for $j=1,\ldots,\,m_{0}+1$ with $\lambda_{0}^{0}=0$ and $\lambda_{m_{0}+1}^{0}=1$. Note that $\lambda_{j}^{0}=\sum_{s=0}^{j}r_{s}.$
The low frequency bias is generated by breaks in the mean function. For the sample autocovariance, the bias factor is given by $d^{*}=2^{-1}\sum_{j_{1}\neq j_{2}}r_{j_{1}}r_{j_{2}}(\overline{\mu}_{j_{2}}-\overline{\mu}_{j_{1}})^{2}$ where
with $\mu_{j}\left(\cdot\right)$ defined in (ref) and we use $\sum_{j_{1}\neq j_{2}}$ as a shorthand for $\sum_{\left\{ j_{1},\,j_{2}=1,\ldots,\,m_{0}+1,\,j_{1}\neq j_{2}\right\} }.$ When the mean is constant in each regime $\mu_{j}\left(t/T\right)=\mu_{j}$. Then, $\overline{\mu}_{j}=\mu_{j}$ and $d^{*}=2^{-1}\sum_{j_{1}\neq j_{2}}r_{j_{1}}r_{j_{2}}(\mu_{j_{2}}-\mu_{j_{1}})^{2}.$ If the mean is constant across regimes, then there is no low frequency bias and $d^{*}=0.$
In Section (ref) we generalize the results in the literature on low frequency contamination for the sample autocovariance and the periodogram. In Section (ref) we show that the local sample autocovariance and the local periodogram are in general robust to low frequency contamination.
\citet*{mikosh/starica:04} established some results on the low frequency bias for the sample autocovariance and periodogram under the assumption that $V_{t}$ is stationary in each regime and that the regimes are independent. In Section (ref) in the supplement we extend these results by allowing time-varying mean and autocovariace function in each regime and weak dependence across regimes. Here we present a brief summary of these results. Theorem (ref) shows that for $\left\{ V_{t,T}\right\} $ that satisfies Definition (ref) and Assumption (ref)-(ref), we have
and as $k\rightarrow\infty,$ $\widehat{\Gamma}\left(k\right)\geq d^{*}$ $\mathbb{P}$-a.s. This suggests that $\widehat{\Gamma}\left(k\right)$ is asymptotically the sum of two terms. The first is the autocovariance of $\left\{ V_{t}\right\} $ at lag $k$. The second, $d_{\mathrm{}}^{*}$, is always positive and increases with the difference in the mean across regimes. Thus, the time-varying mean induces a positive bias. The result that $\widehat{\Gamma}\left(k\right)\geq d^{*}$ $\mathbb{P}$-a.s. as $k\rightarrow\infty$ implies that unaccounted nonstationarity generates long memory effects. The intuition is straightforward. A long memory SLS process satisfies $\sum_{k=-\infty}^{\infty}|\Gamma\left(u,\,k\right)|\rightarrow\infty$ for some $u\in\left(0,\,1\right)$, similar to a stationary long memory process.\footnote{In Section (ref) in the supplement we define long memory SLS processes that are characterized by the property $\sum_{k=-\infty}^{\infty}\left|\rho_{V}\left(u,\,k\right)\right|=\infty$ for some $u\in\left[0,\,1\right]$ where $\rho_{V}\left(u,\,k\right)\triangleq\mathrm{Corr}(V_{\left\lfloor Tu\right\rfloor },\,V_{\left\lfloor Tu\right\rfloor +k})$ and $\vartheta\left(u\right)\in\left(0,\,1/2\right)$ is the long memory parameter at time $u$. } The theorem shows that $\widehat{\Gamma}\left(k\right)$ exhibits a similar property and $\widehat{\Gamma}\left(k\right)$ decays more slowly than for a short memory stationary process for small lags and approaches a constant $d^{*}>0$ for large lags.
Theorem (ref) in the supplement analyzes the properties of the periodogram $I_{T}\left(\omega_{l}\right)$ as $\omega\rightarrow0$ when the mean is time-varying. The result states that as $\omega\rightarrow0$ $\mathbb{E}\left(I_{T}\left(\omega\right)\right)$ generally takes unbounded values except for some $\omega$ for which $\mathbb{E}\left(I_{T}\left(\omega\right)\right)$ is bounded below by $2\pi\int_{0}^{1}f\left(u,\,\omega\right)du>0.$ An SLS process with long memory has an unbounded local spectral density $f\left(u,\,\omega\right)$ as $\omega\rightarrow0$ for some $u\in\left[0,\,1\right]$. Since $f\left(\cdot,\,\cdot\right)$ cannot be negative, it follows that $\int_{0}^{1}f\left(u,\,\omega\right)du$ is also unbounded as $\omega\rightarrow0$. Theorem (ref) suggests that nonstationarity consisting of time-varying first moment results in a periodogram sharing features of a long memory series.
This discussion suggests that certain deviations from stationarity can generate a long memory component that leads to overestimation of the true autocovariance. It follows that the LRV is also overestimated. Since the LRV is used to normalize test statistics, this has important consequences for many HAR inference tests characterized by deviations from stationarity under the alternative hypothesis. These include tests for forecast evaluation, tests and inference for structural change models, time-varying parameters models and regime-switching models. In the linear regression model, $V_{t}$ corresponds to the regressors multiplied by the fitted residuals. Unaccounted nonlinearities and outliers can contaminate the mean of $V_{t}$ and therefore contribute to $d^{*}$.
We now consider the behavior of $\widehat{c}_{T}\left(rn_{T}/T,\,k\right)$ defined in (ref) for fixed $k$ as well as for $k\rightarrow\infty$. For notational simplicity we assume that $k$ is even. For $u\in\left(0,\,1\right)$ define $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)=\{\left\lfloor Tu\right\rfloor +k/2-n_{2,T}/2+1,\ldots,\,\left\lfloor Tu\right\rfloor +k/2+n_{2,T}/2\}$, $n_{j,L}\left(u,\,k,\,n_{2,T}\right)=(T_{j}^{0}-(\left\lfloor Tu\right\rfloor +k/2-n_{2,T}/2+1)),$ and $n_{j,R}\left(u,\,k,\,n_{2,T}\right)=((\left\lfloor Tu\right\rfloor +k/2+n_{2,T}/2+1)-T_{j}^{0})$. $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$ denotes a window of length $n_{2,T}$ around $\left\lfloor Tu\right\rfloor $, $n_{j,L}\left(u,\,k,\,n_{2,T}\right)$ (resp. $n_{j,R}\left(u,\,k,\,n_{2,T}\right)$) denotes the distance between the left (resp. right) end point of $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$ and $T_{j}^{0}$.
The theorem shows that the behavior of $\widehat{c}_{T}\left(u,\,k\right)$ depends on whether a change in mean is present, and if so whether it is close enough to $\left\lfloor Tu\right\rfloor $. For a given $u\in\left(0,\,1\right)$ and $k\in\mathbb{Z}$, if the condition of part (i) of the theorem holds, then $\widehat{c}_{T}\left(u,\,k\right)$ is consistent for $\mathrm{cov}(V_{\left\lfloor Tu\right\rfloor }V_{\left\lfloor Tu\right\rfloor -k})=c\left(u,\,k\right)+O\left(T^{-1}\right)$ {[}see casini_hac{]}. If a change-point falls close to either boundary of the window $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$, as specified in case (ii-b), then $\widehat{c}_{T}\left(u,\,k\right)$ remains consistent. The only case in which a non-negligible bias arises is when the change-point falls in a neighborhood around $\left\lfloor Tu\right\rfloor $ sufficiently far from either boundary. This represents case (ii-a), for which a biased estimate results. However, the bias vanishes asymptotically. Since $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$ is an average of $\widehat{c}_{T}\left(rn_{T},\,k\right)$ over blocks $r=1,\ldots,\,\left\lfloor T/n_{T}\right\rfloor $, if case (ii-a) holds then $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)\geq d_{T}^{*}$ as $k\rightarrow\infty$ but $d_{T}^{*}\rightarrow0$ as $T\rightarrow\infty$. Thus, comparing this result with the discussion above on $\widehat{\Gamma}\left(k\right)$ (see also Theorem (ref)), in practice the long memory effects are unlikely to occur when using $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$. Furthermore, one can reduce this problem by appropriately choosing the blocks $r=1,\ldots,\,\left\lfloor T/n_{T}\right\rfloor $. A procedure was proposed in casini_hac using the methods developed in casini/perron:change-point-spectra.
We now study the asymptotic properties of $I_{\mathrm{L},T}\left(u,\,\omega\right)$ as $\omega\rightarrow0$ for $u\in\left[0,\,1\right]$. We consider the Fourier frequencies $\omega_{l}=2\pi l/n_{T}\in(-\pi,\,\pi)$ for an integer $l\neq0$ (mod $n_{T}$). We need the following high-level conditions. Part (i) corresponds to Assumption (ref), part (ii) is satisfied if $\left\{ V_{t}\right\} $ is strong mixing with mixing parameters of size $-2\nu/\left(\nu-1/2\right)$ for some $\nu>1$ such that $\sup_{t\geq1}\mathbb{E}\left|V_{t}\right|^{4\nu}<\infty,$ while part (iii) requires additional smoothness.
It is useful to compare Theorem (ref) with the discussion above about the periodogram (see also Theorem (ref)). Unlike the periodogram, the asymptotic behavior of the local periodogram as $\omega_{l}\rightarrow0$ depends on the vicinity of $u$ to $\lambda_{j}^{0}$ $\left(j=1,\ldots,\,m_{0}\right)$. Since $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$ uses observations in the window $\mathbf{S}\left(u,\,0,\,n_{T}\right)$, if no discontinuity in the mean occurs in this window then $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$ is asymptotically unbiased for the spectral density $f\left(u,\,\omega_{l}\right)$. More complex is its behavior if some $T_{j}^{0}$ falls in $\mathbf{S}\left(u,\,0,\,n_{T}\right)$. The theorem shows that if $T_{j}^{0}$ is close to the boundary, as indicated in case (ii-b), then $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$ is bounded below by $f\left(u,\,\omega_{l}\right)$, similarly to case (i). If instead $T_{j}^{0}$ falls sufficiently close to the mid-point $\left\lfloor Tu\right\rfloor ,$ as indicated in case (ii-a), then $\mathbb{E}\left(I_{\mathrm{L},T}\left(u,\,\omega\right)\right)\rightarrow\infty$ for many values in the sequence $\left\{ \omega_{l}\right\} $ as $\omega_{l}\rightarrow0$ provided it satisfies $n_{T}\omega_{l}^{2}\rightarrow0$ as $T\rightarrow\infty$. Hence, unless $T\lambda_{j}^{0}$ is close to $\left\lfloor Tu\right\rfloor ,$ the local periodogram $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$ behaves very differently from the periodogram $I_{T}\left(\omega_{l}\right)$. Accordingly, nonstationarity is unlikely to generate long memory effects if one uses the local periodogram. As for $\widehat{c}_{T}\left(u,\,k\right)$, if one uses preliminary inference procedures {[}cf. casini:change-point-spectra{]} for the detection and estimation of the discontinuities in the spectrum and for the estimation of their locations, then one can construct the window efficiently and avoid $T_{j}^{0}$ being too close to $\left\lfloor Tu\right\rfloor .$
We now consider Edgeworth expansions for the distribution of the $t$-statistic in the location model based on the HAC and DK-HAC estimator where $\left\{ V_{t}\right\} $ is assumed to have zero-mean and time-varying second moments. This is useful for analyzing the theoretical properties of the null rejection probabilities of the HAR tests under nonstationarity. As in the literature, we make use of the Gaussianity assumption for mathematical convenience.\footnote{This can be relaxed by considering distributions with Gram-Charlier representations at the expense of more complex derivations. } We relax the stationarity assumption used in the literature {[}cf. jansson:04, sun/phillips/jin:08 and velasco/robinson:01{]} which has important consequences for the nature of the results. The results concerning the $t$-test based on the HAC estimator are presented in Section (ref) while those based on the DK-HAC estimator are presented in Section (ref).
Let $\left\{ V_{t}\right\} $ be a zero-mean Gaussian SLS process satisfying Assumption (ref)-(i-iv). Let
which is valid for all $T$ such that $J_{T}>0$ where $J_{T}=T^{-1}\sum_{s=1}^{T}\sum_{t=1}^{T}\mathbb{E}(V_{s}V_{t})$.
The classical HAC estimator is defined as
where $K_{1}\left(\cdot\right)$ is a kernel and $b_{1,T}$ a bandwidth parameter. Under appropriate conditions on $b_{1,T},$ we have $\widehat{J}_{\mathrm{\mathrm{HAC},}T}-J_{T}\overset{\mathbb{P}}{\rightarrow}0$ from which it follows that
Let $\mathbf{V}=(V_{1},\ldots,\,V_{T})'$. Note that $\widehat{J}_{\mathrm{HAC,}T}=\mathbf{V}'W_{b_{1}}\mathbf{V}/T$ where $W_{b_{1}}$ has $\left(r,\,s\right)$th element
such that $\widetilde{K}_{b_{1}}\left(\omega\right)$ is a kernel with smoothing number $b_{1,T}^{-1}$ and $\Pi=(-\pi,\,\pi]$. For an even function $K$ that integrates to one, we define
Note that $\widetilde{K}_{b_{1}}\left(\omega\right)$ is periodic of period $2\pi$, even and satisfies $\smallint_{-\pi}^{\pi}\widetilde{K}_{b_{1}}\left(\omega\right)d\omega=1$. It follows that $w\left(r\right)=\int_{-\infty}^{\infty}e^{irx}K\left(x\right)dx$ and $\widehat{J}_{\mathrm{HAC,}T}=2\pi\int_{\Pi}\widetilde{K}_{b_{1}}\left(\omega\right)I_{T}\left(\omega\right)d\omega$. $\widetilde{K}_{b_{1}}\left(\omega\right)$ is the so-called spectral window generator. We refer to brillinger:75 for a review of these introductory concepts.
We now analyze the joint distribution of $\overline{V}$ and $\widehat{J}_{\mathrm{HAC,}T}$. Let $\mathsf{B}_{T}=\mathbb{E}(\widehat{J}_{\mathrm{HAC,}T})/J_{T}-1$ and $\mathsf{V}_{T}^{2}=\mathrm{Var}(\sqrt{Tb_{1,T}}\widehat{J}_{\mathrm{HAC,}T}/J_{T})$ denote the relative bias and variance, respectively, of $\widehat{J}_{\mathrm{HAC,}T}$. It is convenient to work with standardized statistics with zero mean and unit variance. Write
where $\mathbf{h}=\left(h_{1},\,h_{2}\right)'$. Note that $h_{2}=\mathbf{V}'Q_{T}\mathbf{V}-$$\mathbb{E}\left(\mathbf{V}'Q_{T}\mathbf{V}\right)$ is a centered quadratic form in a Gaussian vector where $Q_{T}=W_{b_{1}}(\sqrt{T/b_{1,T}}\mathsf{V}_{T}J_{T})^{-1}$. The joint characteristic function of $\mathbf{h}$ is
where $\Upsilon_{T}=\mathbb{E}\left(\mathbf{V}'Q_{T}\mathbf{V}\right)=\mathrm{Tr}\left(\Sigma_{V}Q_{T}\right),$ $\Sigma_{V}=\mathbb{E}\left(\mathbf{VV}'\right)$, and $\xi_{T}=\mathbf{1}/\sqrt{TJ_{T}}$ with $\mathbf{1}$ being the $T\times1$ vector $\left(1,1,\ldots,\,1\right)'$. The cumulant generating function of $\mathbf{h}$ is
where $\kappa_{T}\left(r,\,s\right)$ is the cumulant of $\mathbf{h}.$ phillips:1980 considered the distribution of linear and quadratic forms under Gaussianity. From his derivations, the nonzero bivariate cumulants are
We introduce the following assumptions about $\left\{ V_{t}\right\} $ and $f\left(u,\,0\right)$.
Assumptions (ref)-(ref) about the kernel and bandwidth are the same as in velasco/robinson:01 in which a discussion can be found. They are satisfied by most kernels used in practice. The bandwidth condition in Assumption (ref) is sufficient for the consistency of $\widehat{J}_{\mathrm{HAC,}T}$ and is strengthened in Assumption (ref), for some parts of the proofs, which is satisfied by popular MSE-optimal bandwidths {[}cf. andrews:91, casini_comment_andrews91, \textcolor{MyBlue}{Belotti et al.} belotti/casini/catania/grassi/perron_HAC_Sim_Bandws and whilelm:2015{]}.
Assumptions (ref)-(ref) impose conditions on the smoothness and boundedness of the spectral density. Assumption (ref) is implied by $\sum_{k=-\infty}^{\infty}\left|k\right|^{d_{f}+\varrho}\sup_{t}|\mathbb{E}V_{t}V_{t-k}|<\infty$ but it is stronger than necessary because it extends the smoothness restriction to all frequencies. Assumption (ref) does impose some restrictions on $f\left(u,\,\cdot\right)$ beyond the origin, though it is not particularly restrictive since any $p>1$ arbitrarily close to 1 will suffice.
We now analyze the asymptotic distribution of $\widehat{J}_{\mathrm{HAC,}T}$. Under stationarity this was discussed by bentkus/rudzkis:1982 and velasco/robinson:01. From Lemmas (ref)-(ref) in the supplement we obtain
The order of the asymptotic bias $b_{1,T}^{d_{f}}$ depends on the smoothness of the spectral density at $\omega=0$ {[}cf. Assumption (ref){]}. The constant $\overline{c}_{1}$ depends on the moment of order $d_{f}$ of the kernel $K$ and on the smoothness of $f\left(u,\,\omega\right)$ at $\omega=0$. For example, for the time-varying AR(1) in (ref),
If there is positive dependence at time $u$, then $\rho\left(u\right)>0$ and $f^{\left(2\right)}\left(u,\,0\right)<0$. Suppose $K\left(x\right)\geq0$ for all $x$ so that $\mu_{2}\left(K\right)>0$. Then the sign of the bias is determined by the sign of $\int_{0}^{1}f^{\left(2\right)}\left(u,\,0\right)du$. A positive local AR(1) coefficient contributes negative bias which corresponds to the well-known downward bias of the LRV estimator when there is positive dependence. Conversely, with anti-persistence $\rho\left(u\right)<0$ and $f^{\left(2\right)}\left(u,\,0\right)>0$. Since $\rho\left(\cdot\right)$ is time-varying, whether the bias is positive or negative depends on the path of $\rho\left(\cdot\right)$. The smoother the spectral density is at frequency zero, the smoother the kernel and the slower $b_{1,T}$ can be. The factor $\int_{0}^{1}f\left(u,\,0\right)du$ in the denominator follows by definition because $\mathsf{B}_{T}$ is the relative bias.
We present a second-order Edgeworth expansion to approximate the distribution of $\mathbf{h}$, with error $o((Tb_{1,T})^{-1/2})$ and including terms up to order $(Tb_{1,T})^{-1/2}$ to correct the asymptotic normal distribution. This will imply the validity of that expansion for the distribution of $\widehat{J}_{\mathrm{HAC},T}$. For $\mathbf{B}\in\mathscr{B}^{2}$, where $\mathscr{B}^{2}$ is any class of Borel sets in $\mathbb{R}^{2}$, let $\mathbb{Q}_{T}^{\left(2\right)}\left(\mathbf{B}\right)=\int_{\mathbf{B}}\varphi_{2}\left(\mathbf{h}\right)q_{T}^{\left(2\right)}\left(\mathbf{h}\right)d\mathbf{h},$ where $\varphi_{2}\left(\mathbf{h}\right)=\left(2\pi\right)^{-1}\exp\{-\left(1/2\right)\left\Vert \mathbf{h}\right\Vert ^{2}\}$ is the density of the bivariate standard normal distribution,
where $\mathcal{H}_{j}\left(\cdot\right)$ are the univariate Hermite polynomials of order $j$, and $\Xi_{0}\left(0,\,3\right)=\left(4\pi\right)^{1/2}2!\int_{\Pi}K^{3}\left(\omega\right)$ $d\omega\left\Vert K\right\Vert _{2}^{-3}$ and $\Xi_{0}(2,\,1)=\left(4\pi\right)^{1/2}K\left(0\right)\left\Vert K\right\Vert _{2}^{-1}$ (see Lemmas (ref)-(ref)). Let $\left(\partial\mathbf{B}\right)^{\phi}$denote a neighborhood of radius $\phi$ of the boundary of a set $\mathrm{\mathbf{B}}.$ Let $\mathbb{P}_{T}$ denote the probability measure of $\mathbf{h}.$
Theorem (ref) shows that $\mathbb{Q}_{T}^{\left(2\right)}$ is a valid second-order Edgeworth expansion for the measure $\mathbb{P}_{T}$. The method of proof is the same as in velasco/robinson:01. We first approximate the true characteristic function and then apply a smoothing lemma {[}cf. Lemma (ref) in the supplement which is from bhattacharya/rao:1975{]}. The leading term of the approximation error is of order $o((Tb_{1,T})^{-1/2})$ as the second term on the right hand side of (ref) is negligible if $\mathbf{B}$ is convex because $\phi_{T}$ decreases as a power of $T$. This is the same order obtained for the corresponding leading term under stationarity. Since the higher-order correction terms in $q_{T}^{\left(2\right)}$ depend only on $K\left(\cdot\right)$ but not on $f\left(\cdot,\,\cdot\right)$, they are equal to the one obtained under stationarity.
Next, we focus on $Z_{T}$, i.e., a $t$-statistic for the mean. Proceeding as in velasco/robinson:01, we first derive a linear stochastic approximation to $Z_{T}\left(\mathbf{h}\right)$ and show that its distribution is the same as that of $Z_{T}$ up to order $o((Tb_{1,T}){}^{-1/2})$. Then, we show that the asymptotic approximation for the distribution of the linear stochastic approximation is valid also for $Z_{T}$ with the same error $o((Tb_{1,T}){}^{-1/2})$. Using Lemmas (ref)-(ref) in the supplement we can substitute out $\mathsf{B}_{T}$ and $\mathsf{V}_{T}$ in $Z_{T}$ and, by only focusing on the leading terms, we define the following linear stochastic approximation,
The next theorem presents a valid Edgeworth expansion for the distribution of $\widetilde{Z}_{T}$ from that of $\mathbf{h}.$
Theorem (ref) shows the form of the correction term to the standard normal distribution, i.e., $b_{1,T}^{d_{f}}\int_{\mathbf{C}}\varphi\left(x\right)r_{2}\left(x\right)dx$. The error of the approximation is of order $o((Tb_{1,T})^{-1/2})$ which is the same as the one obtained under stationarity by velasco/robinson:01.
Let $\Phi\left(\cdot\right)$ denote the distribution function of the standard normal. Setting $\mathbf{C}=(-\infty,\,z]$, integrating and Taylor expanding $\Phi\left(\cdot\right)$, we obtain, uniformly in $z$,
This shows that under the conditions of Theorem (ref), the standard normal approximation is correct up to order $O((Tb_{1,T})^{-1/2})$. Eq. (ref) has an immediate interpretation. Consider the time-varying AR(1) example in (ref) and suppose $K\left(x\right)\geq0$ for all $x$ so that $\mu_{2}\left(K\right)\geq0$. Given (ref) we know that with local positive persistence (i.e., $\rho\left(u\right)>0$) $f\left(u,\,\omega\right)$ has a peak at $\omega=0$. If the pattern of $\rho\left(u\right)$ is such that $\int_{0}^{1}f^{\left(2\right)}\left(u,\,0\right)du<0$ so that the positive persistence dominates, then $\overline{c}_{1}<0$ and as is well-known the HAC estimator underestimates the true LRV and the corresponding HAC-based test over-rejects. The approximation in (ref) tends to correct this problem as it follows that one uses $\Phi\left(z\left(1+\gamma_{T}\right)\right)$ where $\gamma_{T}\leq0$, so for a given significance level the critical value $z$ is larger in absolute value than the corresponding standard normal critical value. Conversely, if there is anti-persistence, then $\overline{c}_{1}>0$ and the implied critical value is smaller than the corresponding standard normal critical value. For $d_{f}>2$ the reasoning is the same but one has to take into account the sign of $\mu_{d_{f}}\left(K\right)$.
Consider the location model $y_{t}=\beta+V_{t}$ $\left(t=1,\ldots,\,T\right).$ For the null hypothesis $\mathbb{H}_{0}:\,\beta=\beta_{0}$, consider the following $t$-test,
where $\widehat{\beta}$ is the least-squares estimator of $\beta$. Theorem (ref) and (ref) imply that
for any $z\in\mathbb{R},$ where $p\left(z\right)$ is an odd function. When $q=1/\left(1+2d_{f}\right)$ we have $p\left(z\right)=2^{-1}\overline{c}_{1}z\varphi\left(z\right)C^{d_{f}+1/2}$ where $C$ is defined in Assumption (ref). Thus, the error in rejection probability (ERP) of $t_{\mathrm{HAC}}$ is of order $O((Tb_{1,T})^{-1/2})$. If $\left\{ V_{t}\right\} $ is second-order stationary, the results in velasco/robinson:01 imply that the ERP of $t_{\mathrm{HAC}}$ is also of order $O((Tb_{1,T})^{-1/2})$. Below we establish the corresponding ERP when the $t$-statistic is instead normalized by $\widehat{J}_{\mathrm{DK},T}$ and also discuss the ERP of the $t$-test under fixed-$b$ asymptotics.
We now consider the Edgeworth expansion for tests based on the DK-HAC estimator. In order to simplify some parts of the proof here we consider an asymptotically equivalent version of the DK-HAC estimator discussed in Section (ref). Let
where $b_{1,T}$ is a bandwidth sequence and
with $K_{2}$ a kernel and $b_{2,T}$ a bandwidth. Note that $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$ and $\widehat{\Gamma}_{\mathrm{DK}}^{*}\left(k\right)$ are asymptotically equivalent and $\widehat{c}_{T}$ is a special case of $\widehat{c}_{\mathrm{DK,}T}$ with $K_{2}$ being a rectangular kernel and $n_{2,T}=Tb_{2,T}$.
Under Assumptions (ref)-(ref), (ref) and (ref) it holds that $\widehat{J}_{\mathrm{DK},T}^{*}-J_{T}\overset{\mathbb{P}}{\rightarrow}0$ {[}cf. casini_hac{]} and
Note that $\widehat{J}_{\mathrm{DK},T}^{*}=\int_{0}^{1}\mathbf{\widetilde{V}}\left(r\right)'W_{b_{1}}\mathbf{\widetilde{V}}\left(r\right)dr/(Tb_{2,T})$ where $\mathbf{\widetilde{V}}\left(r\right)=(\widetilde{V}_{1}\left(r\right),\,\widetilde{V}_{2}\left(r\right),\ldots,\,\widetilde{V}_{T}\left(r\right))'$ with $\widetilde{V}_{j}\left(r\right)=\sqrt{K_{2}\left(\left(r-j\right)/Tb_{2,T}\right)}V_{j}$ and $W_{b_{1}}$ defined in (ref). Let
$\widetilde{I}_{T}\left(r,\,\omega\right)$ is the local periodogram of $\{\mathbf{\widetilde{V}}\left(r\right)\}$. Then, $\widehat{J}_{\mathrm{DK},T}^{*}=2\pi\int_{0}^{1}\int_{\Pi}\widetilde{K}_{b_{1}}\left(\omega\right)\widetilde{I}_{T}\left(r,\,\omega\right)d\omega dr$.
We begin by analyzing the joint distribution of $\overline{V}$ and $\widehat{J}_{\mathrm{DK},T}^{*}$. Let $\mathsf{B}_{\mathrm{2},T}=\mathbb{E}(\widehat{J}_{\mathrm{DK},T}^{*})/J_{T}-1$ and $\mathsf{V}_{2,T}^{2}=\mathrm{Var}(\sqrt{Tb_{1,T}b_{2,T}}\widehat{J}_{\mathrm{DK},T}^{*}/J_{T})$ denote the relative bias and variance of $\widehat{J}_{\mathrm{DK},T}^{*}$, respectively. It is convenient to work with standardized statistics with zero mean and unit variance. Write
where $\mathbf{v}=\left(v_{1},\,v_{2}\right)'$ with $v_{1}=h_{1}.$ Note that $v_{2}=\int_{0}^{1}(\mathbf{\widetilde{V}}\left(r\right)'Q_{2,T}\mathbf{\widetilde{V}}\left(r\right)-$$\mathbb{E}(\mathbf{\widetilde{V}}\left(r\right)'Q_{2,T}\mathbf{\widetilde{V}}\left(r\right)))dr$ is a centered quadratic form in a Gaussian vector where $Q_{2,T}=W_{b_{1}}(\sqrt{Tb_{2,T}/b_{1,T}}\mathsf{V}_{2,T}J_{T})^{-1}$. The joint characteristic function of $\mathbf{v}$ is
where $\Upsilon_{2,T}=\mathbb{E}(\int_{0}^{1}(\mathbf{\widetilde{V}}\left(r\right)'Q_{2,T}\mathbf{\widetilde{V}}\left(r\right))dr)=\mathrm{Tr}(\Sigma_{\widetilde{V}}Q_{2,T}),$ $\Sigma_{\widetilde{V}}=\mathbb{E}(\int_{0}^{1}(\mathbf{\widetilde{V}}\left(r\right)\mathbf{\widetilde{V}}\left(r\right)')dr)$ and $\xi_{2,T}=\mathbf{1}/\sqrt{Tb_{2,T}J_{T}}$. The cumulant generating function of $\mathbf{v}$ is
where $\kappa_{2,T}\left(r,\,s\right)$ is the cumulant of $\mathbf{v}.$ To obtain more precise bounds in some parts of the proofs we use the following assumption on the cross-partial derivatives of $f\left(u,\,\omega\right)$. Let $\widetilde{\mathbf{C}}$ denote the set of continuity points of $f\left(u,\,\omega\right)$ in $u$, i.e., $\widetilde{\mathbf{C}}=\{\left[0,\,1\right]/\{\lambda_{j}^{0},\,j=1,\ldots,\,m_{0}\}\}$. Define
where
From Lemmas (ref) and (ref), the relative bias of $\widehat{J}_{\mathrm{DK},T}^{*}$ is
where
The factor $\overline{c}_{1}$ in the relative bias $\mathsf{B}_{2,T}$ also enters $\mathsf{B}_{T}$ and we already discussed it. The second factor, $\overline{c}_{2}$, includes two elements. The first depends on the second moment of the kernel $K_{2}$ and on the smoothness over time of the spectral density $f\left(u,\,0\right)$. The second element in $\overline{c}_{2}$ is $\Delta_{f}\left(0\right)$ which depends on the right and left first partial derivatives of $f\left(u,\,0\right)$ with respect to $u$ at the discontinuity points. The more nonstationary is the data the more complex is $\overline{c}_{2}$, and in fact the larger in magnitude are $\partial^{2}f\left(u,\,0\right)/\partial u^{2}$ and $\Delta_{f}\left(0\right)$. For the special case of stationary data, $\overline{c}_{2}=0$. The more nonstationary is the data, the smaller $b_{2,T}$ should be chosen so as to weight more the data locally. The smoothing over sample autocovariances is needed to achieve consistency while the time-smoothing is introduced to more flexibly account for the time-varying properties of the data. The disadvantage of the time-smoothing is that it reduces the effective sample size thereby making accounting for strong dependence more difficult.
We now present a second-order Edgeworth expansion to approximate the distribution of $\mathbf{v}$ with error $o((Tb_{1,T}b_{2,T})^{-1/2})$. The expansion includes terms up to order $(Tb_{1,T}b_{2,T})^{-1/2}$ to correct the asymptotic normal distribution. This implies the validity of that expansion for the distribution of $\widehat{J}_{\mathrm{DK},T}^{*}$. For $\mathbf{B}\in\mathscr{B}^{2}$, let $\mathbb{Q}_{2,T}^{\left(2\right)}(\mathbf{B})=\int_{\mathbf{B}}\varphi_{2}\left(\mathbf{v}\right)q_{2,T}^{\left(2\right)}\left(\mathbf{v}\right)d\mathbf{v},$ where
$\mathcal{H}_{2,j}\left(\cdot\right)$ are the univariate Hermite polynomials of order $j$ and $\Xi_{2,0}(0,\,3)$ and $\Xi_{2,0}(2,\,1)$ are bounded and depend on $K,\,K_{2}$ and on $f\left(u,\,0\right)$ (see Lemmas (ref)-(ref)).
Theorem (ref) shows that $\mathbb{Q}_{2,T}^{\left(2\right)}$ is a valid second-order Edgeworth expansion for the probability measure $\mathbb{P}_{T}$ of $\mathbf{v}.$ The correction $q_{2,T}^{\left(2\right)}\left(\mathbf{v}\right)$ differs from $q_{T}^{\left(2\right)}\left(\mathbf{h}\right)$ in Theorem (ref). This difference depends on the smoothing over time, i.e., on $b_{2,T}$ and $K_{2}\left(\cdot\right)$. The theorem also suggests that the leading term of the error of the approximation is of order $o((Tb_{1,T}b_{2,T})^{-1/2})$.
Next, we focus on $U_{T}$ defined in (ref), i.e., a $t$-statistic based on $\widehat{J}_{\mathrm{DK},T}^{*}$, and present the Edgeworth expansion. We need the following assumption, replacing Assumptions (ref)-(ref), that controls the rate of smoothing over lagged autocovariances and time implied by the bandwidths $b_{1,T}$ and $b_{2,T}$, respectively. It requires that the bias due to smoothing over frequency and over time is of the same order as the correction term obtained in $\mathbb{Q}_{2,T}^{\left(2\right)}\left(\mathbf{B}\right)$ or as the standard deviation of $\widehat{J}_{\mathrm{DK},T}^{*}$. The assumption is satisfied by, for example, the MSE-optimal DK-HAC estimators proposed by \textcolor{MyBlue}{Belotti et al.} belotti/casini/catania/grassi/perron_HAC_Sim_Bandws and casini_hac.
Theorem (ref) shows that the correction term to the standard normal distribution, i.e., $\int_{\mathbf{C}}\varphi\left(x\right)$ $(r_{2}\left(x\right)b_{1,T}^{d_{f}}+r_{3}\left(x\right)b_{2,T}^{2})dx$, depends on both smoothing directions. The error of the approximation is of order $o((Tb_{1,T}$ $b_{2,T})^{-1/2})$ which can be larger than that obtained in Theorem (ref) for the HAC estimators. Similar to (ref), we obtain uniformly in $z$,
where $\mathbf{C}=(-\infty,\,z]$, which suggests that the standard normal approximation is correct up to order $O((Tb_{1,T}b_{2,T})^{-1/2})$. Eq. (ref) has a similar interpretation to (ref). Consider the time-varying AR(1) example in (ref) and suppose $\rho\left(u\right)>0$ for all $u.$ Then, $\overline{c}_{1}<0$. However, the sign of $\overline{c}_{2}$ is not easily determined even for this simple model. For the special case $\rho\left(u\right)=\sin(u\pi/10)$, no break and $\sigma^{2}\left(u\right)=\sigma^{2}$ we have $\overline{c}_{2}<0$. Then, the implied critical value from the approximation is larger than the standard normal critical value. In general, however, the correction to strong persistence might be either attenuated or strengthened by the correction to nonstationarity depending on the true data-generating process.
Returning to the location model, consider the $t$-statistic based on $\widehat{J}_{\mathrm{DK},T}^{*}$,
Theorem (ref) and (ref) imply that
for any $z\in\mathbb{R},$ where $p_{2}\left(z\right)$ is an odd function. Under the conditions of Theorem (ref) $p_{2}\left(z\right)=2^{-1}((C^{d_{f}+1/2}\overline{c}_{1}+C_{2}\overline{c}_{2})z\varphi\left(z\right))$ where $C$ is defined in Assumption (ref), $C_{2}=(\overline{b}C^{d_{f}+1/2})^{1/2}$ and $\overline{b}$ is defined in Assumption (ref). Thus, the ERP of $t_{\mathrm{DK}}$ can be larger than that of $t_{\mathrm{HAC}}$, though the margin is small. This follows from the fact that $\widehat{J}_{\mathrm{DK},T}^{*}$ applies smoothing over two directions. The smoothing over time is useful to flexibly account for nonstationarity. Its benefits appear explicitly under the alternative hypothesis as we show in Section (ref) whereas the ERP refers to the null hypothesis. One can show that the ERP of $t_{\mathrm{DK}}$ and $t_{\mathrm{HAC}}$ remain unchanged if prewhitening is applied, though the proofs are omitted since they are similar.
We can further compare the ERP of $t_{\mathrm{HAC}}$ and $t_{\mathrm{DK}}$ to that of the corresponding $t$-test under the fixed-$b$ asymptotics. casini_fixed_b_erp showed that the limiting distribution of the original fixed-$b$ HAR test statistics under nonstationarity is not pivotal as it depends on the true data-generating process of the errors and regressors. This contrasts to the stationarity case for which the fixed-$b$ limiting distribution is pivotal and the ERP is of order $O(T^{-1})$ {[}see jansson:04 and sun/phillips/jin:08{]}. Based on an ERP of smaller magnitude relative to that of HAR tests based on HAC estimators {[}cf. $O(T^{-1})<O((Tb_{1,T})^{-1/2})${]}, the literature has long suggested that the original fixed-$b$ HAR tests are superior to HAR tests based on HAC estimators. However, this breaks down under nonstationarity as shown by casini_fixed_b_erp who established that (i) the ERP of the original fixed-$b$ HAR tests does not converge to zero because under nonstationarity the fixed-$b$ limiting distribution is different; (ii) for fixed-$b$ HAR tests that use the critical values from the non-pivotal fixed-$b$ limiting distribution the ERP increases by an order of magnitude relative to the stationary case {[}i.e., from $O(T^{-1})$ to $O(T^{-\eta})$ with $\eta\in(0,\,1/2)${]}. Therefore, fixed-$b$ HAR tests can have an ERP larger than that of $t_{\mathrm{HAC}}$ and $t_{\mathrm{DK}}$. Overall, the results based on Edgeworth expansions show that the distortions on the null rejection rates of the HAR tests can arise from time variation in the second moments even when the mean is constant. Thus, these results complement the asymptotic bias results induced by breaks in the mean function.
In this section, we discuss the implications of the theoretical results from Section (ref)-(ref). In Section (ref), we first present a review of HAR inference methods and their connection to the estimates considered in Section (ref). In Section (ref) we present evidence that the HAR inference tests can suffer from larger size distortions under nonstationarity than under stationarity. In Section (ref) we show the consequences of low frequency contamination for the power of the HAR tests and we provide the corresponding theoretical results in Section (ref).
There are two main approaches for HAR inference. Classical HAC standard errors {[}cf. newey/west:87 (newey/west:87, newey/west:94) and andrews:91{]} require estimation of the LRV defined as $J\triangleq\mathrm{lim}_{T\rightarrow\infty}J_{T}$ where $J_{T}$ is defined after (ref). The form of $\left\{ V_{t}\right\} $ depends on the specific problem under study. For example, for a $t$-test on a regression coefficient in the linear model $y_{t}=x{}_{t}\beta_{0}+e_{t}$ $\left(t=1,\ldots,\,T\right)$ we have $V_{t}=x_{t}e_{t}$. Classical HAC estimators take the following form,
where $\widehat{\Gamma}\left(k\right)$ is given in (ref) with $\widehat{V}_{t}=x_{t}\widehat{e}_{t}$ where $\left\{ \widehat{e}_{t}\right\} $ are the least-squares residuals, $K_{1}\left(\cdot\right)$ is a kernel and $b_{1,T}$ is bandwidth. One can use the the Bartlett kernel, advocated by newey/west:87, the quadratic spectral kernel as suggested by andrews:91, or any other kernel suggested in the literature, see e.g. dejong/davidson:00 and ng/perron:1996. Under $b_{1,T}\rightarrow0$ at an appropriate rate, we have $\widehat{J}_{\mathrm{HAC,}T}\overset{\mathbb{P}}{\rightarrow}J.$ Hence, equipped with $\widehat{J}_{\mathrm{HAC,}T}$, HAR inference is standard and simple because HAR test statistics follow asymptotically standard distributions.
HAC standard errors can result in oversized tests when there is substantial temporal dependence {[}e.g., andrews:91{]}. This stimulated a second approach based on LRV estimators that keeps the bandwidth at some fixed fraction of $T$ {[}cf. Kiefer/vogelsang/bunzel:00{]}, e.g., using all autocovariances, so that $\widehat{J}_{\mathrm{\mathrm{KVB},}T}\triangleq T^{-1}\sum_{t=1}^{T}\sum_{s=1}^{T}\left(1-\left|t-s\right|/T\right)$ $\widehat{V}_{t}\widehat{V}_{s}$ which is equivalent to the Newey-West estimator with $b_{1,T}=T^{-1}$. Under fixed-$b$ asymptotics the reference distribution of HAR test statistics is nonstandard. The validity of fixed-$b$ inference rests on stationarity {[}cf. casini_fixed_b_erp{]}. Many authors have considered various versions of $\widehat{J}_{\mathrm{\mathrm{KVB},}T}$. However, the one that leads to HAR inference tests that are least oversized is the original $\widehat{J}_{\mathrm{\mathrm{KVB},}T}$ {[}see casini/perron_PrewhitedHAC for simulation results{]}. For comparison we also report the equally-weighted cosine (EWC) estimator of lazarus/lewis/stock:17. It is an orthogonal series estimators that use long bandwidths,
with $B$ some fixed integer. Assuming $B$ satisfies some conditions, under fixed-$b$ asymptotics a $t$-statistic normalized by $\widehat{J}_{\mathrm{\mathrm{EWC},}T}$ follows a $t_{B}$ distribution where $B$ is the degree of freedom.
Recently, a new HAC estimator was proposed in casini_hac. Motivated by the power impact of low frequency contamination of existing LRV estimators, he proposed a double kernel HAC (DK-HAC) estimator, defined by
where $b_{1,T}$ is a bandwidth sequence and $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$ defined in Section (ref) with $\widehat{c}_{T}\left(\cdot,\,k\right)$ replaced by
with $K_{2}$ a kernel and $b_{2,T}$ a bandwidth. Note that $\widehat{c}_{\mathrm{DK,}T}$ and $\widehat{c}_{T}$ are asymptotically equivalent and the results of Section (ref) continue to hold for $\widehat{c}_{\mathrm{DK,}T}$. More precisely, $\widehat{c}_{T}$ is a special case of $\widehat{c}_{\mathrm{DK,}T}$ with $K_{2}$ being a rectangular kernel and $n_{2,T}=Tb_{2,T}$. This approach falls in the first category of standard inference $\widehat{J}_{\mathrm{DK},T}\overset{\mathbb{P}}{\rightarrow}J$ and HAR test statistics normalized by $\widehat{J}_{\mathrm{DK},T}$ follows standard distribution asymptotically. The DK-HAC estimator involves two kernels: $K_{1}$ smooths the lagged sample autocovariances, akin to the classical HAC estimators, while $K_{2}$ applies smoothing over time. The latter feature is useful to avoid the low frequency contamination. Additionally, casini/perron_PrewhitedHAC proposed prewhitened DK-HAC $(\widehat{J}_{\mathrm{\mathrm{pw},DK},T})$ estimator that improves the size control of HAR tests and enjoys the same asymptotic properties of $\widehat{J}_{\mathrm{DK},T}$. casini_hac and casini/perron_PrewhitedHAC demonstrated via simulations that tests based on $\widehat{J}_{\mathrm{DK},T}$ and $\widehat{J}_{\mathrm{pw,DK},T}$ have superior power properties relative to tests based on the other estimators. In terms of size, the simulation results showed that tests based on $\widehat{J}_{\mathrm{pw,DK},T}$ perform better than those based on $\widehat{J}_{\mathrm{HAC,}T}$ and $\widehat{J}_{\mathrm{DK},T}$, and is competitive with $\widehat{J}_{\mathrm{\mathrm{KVB},}T}$ when the latter works well. We include $\widehat{J}_{\mathrm{DK},T}$ and $\widehat{J}_{\mathrm{pw,DK},T}$ in our simulations below. We report the results only for the DK-HAC estimators that do not use the pre-test for discontinuities in the spectrum {[}cf. casini/perron:change-point-spectra{]} because we do not want the results to be affected by such pre-test.
In order to better understand the effect of nonstationarity on the null rejection rates of HAR tests we first conduct a Monte Carlo analysis where we compare a nonstationary model with a stationary one that has either the same spectral density at frequency zero or the same average dependence. Consider the following four AR(1) data-generating processes (DGPs). DGP 1 is given by
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$ for all $t$. The LRV of DGP 1 is $J=1.826$. DGP 2 is
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$ for all $t$. Its LRV is $J=20.988$. We now introduce two nonstationary DGPs. DGP 3 takes the following form
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$. Note that the spectral density at frequency zero of $V_{t}$ is given by the weighted average of the spectral densities of $V_{t}$ in the two regimes:
Thus, the LRV of $V_{t}$ is $J=2\pi\int_{0}^{1}f\left(u,\,0\right)du=20.988$ which takes the same value as the LRV of DGP 2. Further, DGP 3 has the same average dependence as DGP 1, meaning that the AR(1) coefficient in DGP 1 is equal to the weighted average of the AR(1) coefficients of DGP 3 in the two regimes, i.e., $\overline{\rho}=0.2\cdot0.9+0.8\cdot0.1=0.26$. We also want to verify whether the location of the break in persistence in DGP 3 is important for the bias. Thus, we consider DGP 4:
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$ for all $t$. While in DGP 3 the regime with strong persistence occurs in the first 20% of the sample, in DGP 4 it occurs between the 50% and 70% of the sample. The LRV of DGP 4 is the same as that of DGP 3.
For each DGP we consider three different initial conditions: (a) $V_{0}=0$; (b) $V_{0}\sim\mathscr{N}\left(0,\,1\right)$; (c) $V_{0}\sim\mathscr{N}\left(0,\,4\right)$. This is useful in order to verify whether the initial condition has any effect on the bias generated by changes in the second-order properties. DGP 3(a) should exhibit a smaller bias due to nonstationarity than DGP 3(b,c) and 4. To see this, note that in DGP 3(a) the initial condition is $V_{0}=0.$ Thus, the process starts from zero. Since there is strong persistence in the first 20% of the sample, the process is more likely to stay close to zero in the first regime than when the initial condition is $V_{0}\sim\mathscr{N}\left(0,\,1\right)$ or $V_{0}\sim\mathscr{N}\left(0,\,4\right)$. In DGP 4 the different specifications of the initial condition should not lead to any differences in the bias due to nonstationarity because the regime with strong dependence occurs about mid-sample.
To summarize, we have four DGPs. DGP 1 and 2 are stationary while DGP 3 and 4 are nonstationary. Since DGP 2 has a LRV that takes the same value as that of DGP 3 and 4, this allows us to better separate the effect of persistence from that of nonstationarity in the second moments on the following quantities: $\widehat{J}_{\mathrm{HAC}}$, $-\widehat{c}_{1}b_{1,T}$ and $\widehat{\Gamma}\left(k\right)$ for $k=0,1,\,5,\,10.$ In the simulations below $\widehat{J}_{\mathrm{HAC}}$ is the Newey-West estimator based on a predetermined number of lagged sample autocovariances following the rule $4\left(T/100\right)^{2/9}$ {[}cf. lazarus/lewis/stock/watson:18{]}. We compare $\widehat{\Gamma}\left(k\right)$ to the theoretical value $\Gamma_{T}\left(k\right)$ corresponding to each DGP which can be computed by hand given the simple form of the DGPs. In fact, for the nonstationary DGPs, $\Gamma_{T}\left(k\right)$ is a weighed average of the theoretical autocovariances corresponding to each regime. Here, $\widehat{c}_{1}$ is an estimate of $\overline{c}_{1}$ in (ref) that enters the asymptotic bias of $\widehat{J}_{\mathrm{HAC}}$. In order to compute $\widehat{c}_{1}$ we recall that the asymptotic bias of the LRV estimator based on the Bartlett kernel is given by
where
denotes the index of smoothness of the kernel at zero and $f^{\left(1\right)}\left(u,\,0\right)$ is the index of smoothness of the local spectral density at time $u$ and frequency zero. For the Bartlett kernel $K_{\mathrm{BT},q}=0$ if $q<1$, $K_{\mathrm{BT},q}=1$ if $q=1$ and $K_{\mathrm{BT},q}=\infty$ if $q>1.$ The Parzen characteristic exponent is the largest $q$ such that $K_{\mathrm{BT},q}$ is finite. Thus, the relative bias is
using $K_{\mathrm{BT},1}=1.$ The index of smoothness of $f\left(u,\,\omega\right)$ at $\omega=0$ is defined as
For an AR(1) process with parameters $\rho\left(u\right)$ and $\sigma_{e}^{2}\left(u\right)$, we have $\Gamma\left(u,\,k\right)=\sigma_{e}^{2}\left(u\right)\rho\left(u\right)^{|k|}/(1-\rho\left(u\right)^{2}).$ It follows that
Based on this result we can obtain $\overline{c}_{1}$ for each model. In particular, for model DGP 1, 2, 3 and 4 we have $\overline{c}_{1}=0.55$, 3.92, 9.04 and 9.05, respectively.
We estimate $\overline{c}_{1}$ as follows. For DGP 1, we obtain the OLS residuals $\widehat{V}_{t}$ and estimate $\rho$ and $\sigma_{e}^{2}$ from the autoregression
where $\sigma_{e}^{2}$ is the variance of $e_{t}$. Let these estimates be denoted by $\widehat{\rho}$ and $\widehat{\sigma}_{e}^{2}$, respectively. Then, the estimate of $\overline{c}_{1}$ is defined as
The same applies to DGP 2. For DGP 3, we obtain the estimate of the autoregressive coefficient of $V_{t}$ and of the variance of the innovations by estimating the autoregression in the two regimes separately. That is, we obtain
where we also compute $\widehat{\sigma}_{1,e}^{2}$ and $\widehat{\sigma}_{2,e}^{2}$ which are the sample variances of the residuals $\widehat{e}_{t}$ in the two regimes, respectively. Then, the estimate of $\overline{c}_{1}$ is defined as
The same applies to DGP 4 with the difference that the autoregressive coefficient and the variance of the innovations are estimated separately in each of the three distinct regimes.
We consider the sample size $T=100,\,200$ and 1000, and 50,000 repetitions were used for each DGP. The results are reported in Table (ref). Let us first discuss the finite-sample properties of $\widehat{J}_{\mathrm{HAC}}$. The results clearly suggest that $\widehat{J}_{\mathrm{HAC}}$ deviates substantially from $J$ when the data are nonstationary. $\widehat{J}_{\mathrm{HAC}}$ underestimates $J$ for all DGPs but it does so much more when the DGP is nonstationary. The difference between the values of $\widehat{J}_{\mathrm{HAC}}$ in DGP 2 and those in DGP 3-4 is about one half, e.g., $\widehat{J}_{\mathrm{HAC}}=6.775$ in DGP 2(a) and $\widehat{J}_{\mathrm{HAC}}=3.142$ in DGP 3(a). As the sample size increases the downward bias becomes smaller, though $\widehat{J}_{\mathrm{HAC}}$ still underestimates $J$ for $T=1000$. The downward bias continues to remain larger in DGP 3-4 than in DGP 2 even when $T=1000$. Thus, this evidence based on $\widehat{J}_{\mathrm{HAC}}$ already points out that basic forms of nonstationarity generate bias in the LRV estimator. This bias adds to the well-known bias generated by strong persistence in stationary data documented in the literature.
Let us discuss the relative bias $-\overline{c}_{1}b_{1,T}$ and its estimate $-\widehat{c}_{1}b_{1,T}$. First note that $-\overline{c}_{1}b_{1,T}<0$ and $-\widehat{c}_{1}b_{1,T}<0$ for all DGPs and sample sizes considered. This confirms the downward bias of $\widehat{J}_{\mathrm{HAC}}$ observed above. For a given model, the asymptotic relative bias $-\overline{c}_{1}b_{1,T}$ and its estimate increase with the sample size. The downward bias is much larger for the nonstationary DGP 3-4 than for the stationary DGP 1-2. The estimates $-\widehat{c}_{1}b_{1,T}$ of the relative bias $-\overline{c}_{1}b_{1,T}$ significantly underestimate $-\overline{c}_{1}b_{1,T}$ in DGP 3-4 while in DGP 1-2 the deviations are much smaller. The large deviations of $-\widehat{c}_{1}b_{1,T}$ from $-\overline{c}_{1}b_{1,T}$ continue to hold even for $T=1000.$
We now move to discuss the finite-sample properties of $\widehat{\Gamma}\left(k\right)$. When the data are stationary, $\widehat{\Gamma}\left(k\right)$ is close to $\Gamma_{T}\left(k\right)$ even when $T=100$ and it approaches $\Gamma_{T}\left(k\right)$ when $T=1000$. For nonstationary data, $\widehat{\Gamma}\left(k\right)$ is much farther from $\Gamma_{T}\left(k\right)$. For example, in DGP 2(a) $\widehat{\Gamma}\left(0\right)=2.507$ and $\Gamma_{T}\left(0\right)=2.571$ whereas in DGP 3(a) $\widehat{\Gamma}\left(0\right)=1.589$ and $\Gamma_{T}\left(0\right)=1.861$. Thus, $\widehat{\Gamma}\left(k\right)$ has larger bias (in general downward) when the data are nonstationary. This result is present even when $T=200$. As $T$ increases, $\widehat{\Gamma}\left(k\right)$ approaches $\Gamma_{T}\left(k\right)$ for all DGPs, though the downward bias remains larger in DGP 3-4 than in DGP 1-2.
We repeated this exercise for other DGPs and the conclusions were the same. The results suggest that under nonstationarity the bias in the LRV estimator is affected by multiple factors. In addition to the downward bias arising from strong persistence which is also present under stationarity there is bias generated by the time-varying properties of the process. Under the null hypothesis this time variation occurs in the autocovariance structure of the process. For example, in DGP 3 one has $0.2T$ observations to estimate $2\pi\int_{0}^{0.2}f\left(u,\,0\right)du=0.4\pi f\left(0\right)$ where $f\left(0\right)=1/(2\pi\left(1-2\rho+\rho^{2}\right))$ with $\rho=0.9$, and $0.8T$ observations to estimate $2\pi\int_{0.2}^{1}f\left(u,\,0\right)du=1.6\pi f\left(0\right)$ where $f\left(0\right)=1/(2\pi\left(1-2\rho+\rho^{2}\right))$ with $\rho=0.1$. This is more difficult than estimating $2\pi f\left(0\right)=1/(2\pi\left(1-2\rho+\rho^{2}\right))$ with $\rho=0.7817$ using $T$ observations, which applies to DGP 2. Even if the total sample size is $T$ in both DGP 2 and 3, nonstationarity reduces the effective sample size making the estimation of the LRV in DGP 3 effectively based on a smaller number of observations. For example, $\widehat{\Gamma}\left(k\right)$ involves an average on $\{\widehat{V}_{t}\widehat{V}_{t-k}\}$ for $t=k+1,\ldots,\,T$. Some of these pairs $\{\widehat{V}_{t}\widehat{V}_{t-k}\}$ are such that $\widehat{V}_{t}$ and $\widehat{V}_{t-k}$ belong to two different regimes, and so contribute bias to the estimation of $\Gamma_{T}\left(k\right)$. Under stationarity all the pairs $\{\widehat{V}_{t}\widehat{V}_{t-k}\}$ are such that $\widehat{V}_{t}$ and $\widehat{V}_{t-k}$ belong to the same regime leading to more precise estimates of $\widehat{\Gamma}\left(k\right)$ and LRV. In addition, changes in persistence over short regimes share features similar to shifts in the mean, at least graphically. While the former is consistent with the null hypothesis, the latter is not. This is likely to generate some bias where changes in persistence are confounded with shifts in the mean even when the unconditional mean of the series has not changed. The downward bias due to strong persistence and the bias due to time-varying second-order properties are likely to influence each other making the estimation problem even harder.
We now investigate the consequence of nonstationarity for HAR inference. We obtain the empirical size and power for a two-tailed $t$-test on the intercept normalized by several LRV estimators for the model $y_{t}=\delta+V_{t}$ with $\delta=0$ under the null and $\delta>0$ under the alternative hypothesis. Model M1 involves an SLS process: $V_{t}=0.9V_{t-1}+u_{t}$, $V_{0}\sim\mathscr{N}\left(0,\,1\right)$, $u_{t}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(0,\,1\right)$ for $t=1,\ldots,\,T_{1}^{0}$ with $T_{1}^{0}=T\lambda_{1}^{0}$, and $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$, $\rho\left(t/T\right)=0.3\left(\cos\left(1.5-\cos\left(t/T\right)\right)\right)$, $u_{t}\sim\mathscr{\mathrm{i.i.d.}\,\mathscr{N}}\left(0,\,0.5\right)$ for $t=T_{1}^{0}+1,\ldots,\,T$. Note that $\rho\left(\cdot\right)$ varies between 0.172 and 0.263. We set $\lambda_{1}^{0}=0.1$. In addition to M1, we consider other models: M2 involves a time-varying AR(1) with a break in volatility $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$, $\rho\left(t/T\right)=0.7(\cos\left(1.5t/T\right))$, $u_{t}\sim\mathscr{N}\left(0,\,\sigma_{t}^{2}\right)$, $\sigma_{t}^{2}=5$ for $t\leq4$ and $\sigma_{t}^{2}=0.25$ for $t>4$, $V_{0}\sim\mathscr{N}\left(0,\,5\right)$; M3 involves $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$, $\rho\left(t/T\right)=0.8(\cos\left(1.5t/T\right))$, $u_{t}\sim\mathscr{N}\left(0,\,0.25\right)$, $V_{0}=0$ with outliers $V_{t}\sim\mathrm{Uniform}\left(\underline{c},\,5\underline{c}\right)$ for $t=T/2,\,3T/4$ where $\underline{c}=-1/(\sqrt{2}\mathrm{erfc^{-1}\left(3/2\right))}\mathrm{med}\left(\left|V-\mathrm{med}\left(V\right)\right|\right)$ with $\mathrm{erfc}^{-1}$ the inverse complementary error function, $\mathrm{med}\left(\cdot\right)$ is the median and $V=\left(V_{t}\right)_{t=1}^{T}$;\footnote{In this literature, values smaller than $\underline{c}$ are not classified as outliers. } M4 involves a time varying AR(1) with periods of strong persistence where $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$, $\rho\left(t/T\right)=0.95(\cos\left(1.5t/T\right))$, $u_{t}\sim\mathscr{\mathrm{i.i.d.}\,\mathscr{N}}\left(0,\,0.4\right)$ and $V_{0}\sim\mathscr{N}\left(0,\,4\right)$. $\rho\left(\cdot\right)$ varies between 0.7 and 0.05 in M2, between 0.05 and 0.8 in M3 and between 0.95 and 0.07 in M4.
We consider the DK-HAC estimators with and without prewhitening ($\widehat{J}_{\mathrm{DK},T}$, $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},T}$, $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},\mu,T}$) of casini_hac and casini/perron_PrewhitedHAC, respectively; andrews:91' andrews:91 HAC estimator with and without the prewhitening procedure of andrews/monahan:92; newey/west:87's newey/west:87 HAC estimator with the popular rule to select the number of lags (i.e., $b_{1,T}=(4(T/100)^{2/9})^{-1}$; Newey-West with the fixed-$b$ method of Kiefer/vogelsang/bunzel:00 with $b=1$ (labeled KVB); and the Equally-Weighted Cosine (EWC) of lazarus/lewis/stock/watson:18 with the bandwidth choice recommended by the authors. For the DK-HAC estimators we use the data-dependent methods for the bandwidths, kernels and choice of $n_{T}$ as proposed in casini_hac and casini/perron_PrewhitedHAC, which are optimal under mean-squared error (MSE). Let $\widehat{V}_{t}$ denote the least-squares residual based on $\widehat{\delta}$ where the latter is the least-squares estimate of $\delta$. We set $\widehat{b}_{1,T}=0.6828(\widehat{\phi}\left(2\right)T\widehat{\overline{b}}_{2,T})^{-1/5}$ where
with
and $\widehat{\overline{b}}_{2,T}=\left(n_{T}/T\right)\sum_{r=1}^{\left\lfloor T/n_{T}\right\rfloor -1}$ $\widehat{b}_{2,T}\left(rn_{T}/T\right)$, $\widehat{b}_{2,T}\left(u\right)=1.6786(\widehat{D}_{1}\left(u\right)){}^{-1/5}(\widehat{D}_{2}\left(u\right))^{1/5}T^{-1/5}$ where $\widehat{D}_{2}\left(u\right)\triangleq2\sum_{l=-\left\lfloor T^{4/25}\right\rfloor }^{\left\lfloor T^{4/25}\right\rfloor }\widehat{c}_{\mathrm{DK,}T}\left(u,\,l\right)^{2}$ and
with $\left[S_{\omega}\right]$ being the cardinality of $S_{\omega}$ and $\omega_{s+1}>\omega_{s}$, $\omega_{1}=-\pi,\,\omega_{\left[S_{\omega}\right]}=\pi.$ We set $n_{T}=T^{0.6}$, $S_{\omega}=\{-\pi,\,-3,\,-2,\,-1,\,0,\,1,\,2,\,3,\,\pi\}.$ $K_{1}\left(\cdot\right)$ is the QS kernel and $K_{2}\left(x\right)=6x\left(1-x\right)$ for $x\in\left[0,\,1\right].$
Table (ref) reports the results using 5,000 replications. The $t$-test based on newey/west:87's newey/west:87 and andrews:91' andrews:91 prewhitened HAC estimators are excessively oversized. andrews:91' andrews:91 HAC-based test is slightly undersized while the KVB's fixed-$b$ and EWC-based tests are severely undersized. The fact that the KVB's fixed-$b$ and EWC-based tests have larger size distortions than other tests is consistent with the results in Section (ref) which suggest that they have a larger ERP. For the $t$-test on the intercept, $\widehat{J}_{\mathrm{DK},T}$ can lead to tests that are oversized when there is strong dependence. However, the prewhitened DK-HAC estimators $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},T}$ and $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},\mu,T}$ lead to tests having more accurate rejection rates. Nonstationarity affects the power of the tests based on LRV estimators that rely on $\widehat{\Gamma}\left(k\right)$ or equivalently on $I_{T}\left(\omega\right)$ (e.g., the EWC). The KVB's fixed-$b$ and EWC-based tests suffer from relatively large power losses. The power of tests normalized by newey/west:87's (1987) and andrews:91' andrews:91 prewhitened HAC are not comparable because they are significantly oversized. The DK-HAC-based tests have the best power, the second best being andrews:91' andrews:91 HAC-based test.
Turning to M2, Table (ref) shows some size distortions and power losses for KVB's fixed-$b$ and EWC-based tests. The prewhitened DK-HAC-based tests display accurate size control and good power. newey/west:87's (1987) and andrews:91' andrews:91 prewhitened HAC-based tests are again excessively oversized. andrews:91' andrews:91 HAC-based test and the DK-HAC-based test show a similar performance. For model M3-M4, Table (ref) shows that all methods lead to oversized tests except prewhitened DK-HAC and KVB's fixed-$b$. However, the KVB's fixed-$b$-based tests show substantial unde-rejection that has consequences for power whereas the prewhitened DK-HAC-based-tests show accurate null rejection rates and good power. Finally, the simulations show that the null rejection rates of HAC- and DK-HAC-based tests are not very far from each other, thereby confirming that their respective ERP are close as shown in Section (ref).
We now discuss HAR inference tests for which the low frequency contamination results of Section (ref) hold asymptotically. This means that $d^{*}>0$ for all $T$ and as $T\rightarrow\infty$. This comprises the class of HAR tests that admit a nonstationary alternative hypothesis. This class is very large and includes most HAR tests as discussed in the Introduction. Here we consider the Diebold-Mariano test for the sake of illustration and remark that similar issues apply to other HAR tests.
The Diebold-Mariano test statistic is defined as $t_{\mathrm{DM}}\triangleq T_{n}^{1/2}\overline{d}_{L}/\sqrt{\widehat{J}_{d_{L},T}}$, where $\overline{d}_{L}$ is the average of the loss differentials between two competing forecast models, $\widehat{J}_{d_{L},T}$ is an estimate of the LRV of the loss differential series and $T_{n}$ is the number of observations in the out-of-sample. We use the quadratic loss. We consider an out-of-sample forecasting exercise with a fixed forecasting scheme where, given a sample of $T$ observations, $0.5T$ observations are used for the in-sample and the remaining half is used for prediction {[}see perron/yamamoto:18 for recommendations on using a fixed scheme in the presence of breaks{]}. The DGP under the null hypothesis is given by $y_{t}=1+\beta_{0}x_{t-1}^{(0)}+e_{t}$ where $x_{t-1}^{(0)}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(1,\,1\right)$, $e_{t}=0.3e_{t-1}+u_{t}$ with $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$, and we set $\beta_{0}=1$ and $T=400.$ The two competing models both involve an intercept but differ with respect to the predictor used in place of $x_{t}^{(0)}$. The first forecast model uses $x_{t}^{(1)}$ while the second uses $x_{t}^{(2)}$ where $x_{t}^{(1)}$ and $x_{t}^{(2)}$ are independent $\mathrm{i.i.d.}\,\mathscr{N}\left(1,\,1\right)$ sequences, both independent from $x_{t}^{(0)}$. Each forecast model generates a sequence of $\tau\left(=1\right)$-step ahead out-of-sample losses $L_{t}^{(j)}$ $\left(j=1,\,2\right)$ for $t=T/2+1,\ldots,\,T-\tau.$ Then $d_{t}\triangleq L_{t}^{(2)}-L_{t}^{(1)}$ denotes the loss differential at time $t$. The Diebold-Mariano test rejects the null hypothesis of equal predictive ability when $\overline{d}_{L}$ is sufficiently far from zero. Under the alternative hypothesis, the two competing forecast models are as follows: the first uses $x_{t}^{(1)}=x_{t}^{(0)}+u_{X_{1},t}$ where $u_{X_{1},t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$ while the second uses $x_{t}^{(2)}=x_{t}^{(0)}+0.2z_{t}+2u_{X_{2},t}$ for $t\in\left[1,\ldots,\,3T/4-1,\,3T/4+21,\ldots T\right]$ and $x_{t}^{(2)}=\delta\left(t/T\right)+0.2z_{t}+2u_{X_{2},t}$ for $t=3T/4,\ldots,\,3T/4+20$ with $u_{X_{2},t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$, where $z_{t}$ has the same distribution as $x_{t}^{(0)}.$
We consider four specifications for $\delta\left(\cdot\right).$ In the first $x_{t}^{(2)}$ is subject to an abrupt break in the mean $\delta\left(t/T\right)=\delta>0$; in the second $x_{t}^{(2)}$ is locally stationary with time-varying mean $\delta\left(t/T\right)=\delta\left(\sin\left(t/T-3/4\right)\right)$; in the third specification $x_{t}^{(2)}=x_{t}^{(0)}+0.2z_{t}+2u_{X_{2},t}$ for $t\in[1,\ldots,\,T/2-30,\,T/2$ $+21,\ldots T]$ and $x_{t}^{(2)}=\delta\left(t/T\right)+0.2z_{t}+2u_{X_{2},t}$ for $t=T/2-30,\ldots,\,T/2+20$ with $\delta\left(t/T\right)=\delta(\sin(t/T-1/2$ $-30/T))$; in the fourth $x_{t}^{(2)}$ is the same as in the second with in addition two outliers $x_{t}^{(2)}\sim\mathrm{Uniform}\left(\left|\underline{c}\right|,\,5\left|\underline{c}\right|\right)$ for $t=6T/10,\,8T/10$ where $\underline{c}=-1/(\sqrt{2}\mathrm{erfc^{-1}\left(3/2\right))}\mathrm{med}(|x^{(2)}-\mathrm{med}$ $(x^{(2)})|)$ where $x^{(2)}=(x_{t}^{(2)})_{t=1}^{T}$. That is, in the second model $x_{t}^{(2)}$ is locally stationary only in the out-of-sample, in the third it is locally stationary in both the in-sample and out-of sample and in the fourth model $x_{t}^{(2)}$ has two outliers in the out-of-sample. The location of the outliers is irrelevant for the results; they can also occur in the in-sample.
Table (ref) reports the null rejection rate and the power of the various tests for all models. We begin with the case $\delta\left(t/T\right)=\delta>0$ (top panel). The null rejection rate of the test using the DK-HAC estimators is accurate while the tests using other LRV estimators are oversized with the exception of the KVB's fixed-$b$ method for which the rejection rate is equal to zero. The HAR tests using existing LRV estimators have lower power relative to that obtained with the DK-HAC estimators for small values of $\delta$. When $\delta$ increases the tests standardized by the HAC estimators of andrews:91 and newey/west:87, and by the KVB's fixed-$b$ and EWC LRV estimators display non-monotonic power gradually converging to zero as the alternative gets further away from the null value. In contrast, when using the DK-HAC estimators the test has monotonic power that reaches and maintains unit power. The results for the other models are even stronger. In general, except when using the DK-HAC estimators, all tests display serious power problems. Thus, either form of nonstationarity or outliers leads to similar implications, consistent with our theoretical results.
In order to further assess the theoretical results from Section (ref), Figure (ref) (top panel) reports the plots of $d_{t}$, its sample autocovariances and its periodogram, for $\delta=1$. Figures (ref)-(ref) (top panels) in the supplement report the corresponding plots for $\delta=2,\,5$, respectively. We only consider the case $\delta_{t}=\delta>0$. The other cases lead to the same conclusions. For $\delta=1$, Figure (ref) (top panel) shows that $\widehat{\Gamma}\left(k\right)$ decays slowly. As $\delta$ increases, from Figures (ref) and (ref) (top panels), $\widehat{\Gamma}\left(k\right)$ decays even more slowly at a rate far from the typical exponential decay of short memory processes. This suggests evidence of long memory. However, the data are short memory with small temporal dependence. What is generating the spurious long memory effect is the nonstationarity present under the alternative hypothesis. This is visible in the top panels which present plots of $d_{t}$ for the first specification. The shift in the mean of $d_{t}$ for $t=3T/4,\ldots,\,3T/4+20$ is responsible for the long memory effect. This corresponds to the second term of (ref) in Theorem (ref). The overall behavior of the sample autocovariance is as predicted by Theorem (ref). For small lags, $\widehat{\Gamma}\left(k\right)$ shows a power-like decay and it is positive. As $k$ increases to medium lags, the autocovariances turn negative because the sum of all sample autocovariances has to be equal to zero {[}cf. percival:1992{]}. Next, we move to the bottom panels which plot the periodogram of $\{d_{t}\}$. It is unbounded at frequencies close to $\omega=0$ as predicted by Theorem (ref) and as would occur if long memory was present. It also explains why the Diebold-Mariano test normalized by Newey-West's, Andrews', KVB's fixed-$b$ and EWC's LRV estimators have serious power problems. These LRV estimators are inflated and consequently the tests lose power. The figures show that as we raise $\delta$ the more severe these issues and the power losses so that the power eventually reaches zero. This is consistent with our theory since $d^{*}$ is increasing in $\delta$ (cf. $d^{*}\thickapprox0.1\cdot0.9\delta^{2}$).
We now verify the results about the local sample autocovariance $\widehat{c}_{T}\left(u,\,k\right)$ and the local periodogram from Theorems (ref)-(ref). We set $n_{2,T}=T^{0.6}=36$ following the MSE criterion of casini_hac. We consider (i) $u=236/T$, (ii-a) $u=T_{1}^{0}/T=3/4$ and (ii-b) $u=264/T$. Note that cases (i)-(ii-b) correspond to parts (i)-(ii-b) in Theorems (ref)-(ref). We consider $\delta=1,\,2$ and $5$. According to Theorems (ref)-(ref), we should expect long memory features only for case (ii-a). Figures (ref) and (ref)-(ref) in the supplement confirm this. The results pertaining to case (ii-a) are plotted in the middle panels. They show that the local autocovariance displays slow decay similar to the pattern discussed above for $\widehat{\Gamma}\left(k\right)$ and that this problem becomes more severe as $\delta$ increases. Such long memory features also appear for $I_{\mathrm{L}}\left(3/4,\,\omega\right)$. The bottom panels in Figures (ref) and (ref)-(ref) show that the local periodogram at $u=3/4$ and at a frequency close to $\omega=0$ are extremely large. The latter result is consistent with Theorem (ref)-(ii-a) which suggests that $I_{\mathrm{L,}T}\left(3/4,\,\omega\right)\rightarrow\infty$ as $\omega\rightarrow0$. For case (i) and (ii-b) both figures show that the local autocovariance and the local periodogram do not display long memory features. Indeed, they have forms similar to those of a short memory process, a result consistent with Theorems (ref)-(ref) also for cases (i) and (ii-b).
It is noteworthy to explain why HAR inference based on the DK-HAC estimators does not suffer from the low frequency contamination even for case (ii-a). The DK-HAC estimator computes an average of the local spectral density over time blocks. If one of these blocks contains a discontinuity in the spectrum, then as in case (ii-a) some bias would arise for the local spectral density estimate corresponding to that block. However, by virtue of the time-averaging over blocks that bias becomes negligible. Hence, nonparametric smoothing over time asymptotically cancels the bias, so that inference based on the DK-HAC estimators is robust to nonstationarity.
We present theoretical results about the power of $t_{\mathrm{DM}}$ for the case of general low frequency contamination discussed in Section (ref). In particular, we focus on specification (1) (i.e., $\delta>0$). The same intuition and qualitative theoretical results apply to the other specifications of $\delta\left(\cdot\right)$.
Let $t_{\mathrm{DM},i}=T_{n}^{1/2}\overline{d}_{L}/\sqrt{\widehat{J}_{d_{L},i,T}}$ denote the DM test statistic where $i=\mathrm{DK},\,\mathrm{pwDK},\,\mathrm{KVB},\,\mathrm{EWC},$ $\mathrm{A91},\,\mathrm{pwA}91$, $\mathrm{NW87}$ and $\mathrm{pwNW87}$ with $\widehat{J}_{\mathrm{A91},T}$ and $\widehat{J}_{\mathrm{NW87},T}$ being $\widehat{J}_{\mathrm{HAC},T}$ using the quadratic spectral and Bartlett kernel, respectively. Define the power of $t_{\mathrm{DM},i}$ as $\mathbb{P}_{\delta}(|t_{\mathrm{DM},i}|>z_{1-\alpha/2})$ where $z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal for a two-sided test with significance level $\alpha\in\left(0,\,1\right)$. To avoid repetitions we present the results only for $i=\mathrm{DK},\,\mathrm{KVB}$ and $\mathrm{NW87}$. The results concerning the prewhitening DK-HAC estimator are the same as those corresponding to the DK-HAC estimator while the results concerning the EWC estimator are similar to those corresponding to the KVB's fixed-$b$ estimator, though for the latter the non-monotonic power is more pronounced. The results pertaining to andrews:91' andrews:91 HAC estimator (with and without prewhitening) are the same as those corresponding to newey/west:87's newey/west:87 estimator. Let $n_{\delta}=T-T_{b}-2$ denote the length of the regime in which $x_{t}^{(2)}$ exhibits a shift $\delta$ in the mean. The deviation from the null hypothesis depends on the shift magnitude $\delta$ and on $n_{\delta}$.
Note that Assumption (ref) with $q=1/3$ refers to the MSE-optimal bandwidth for the newey/west:87's newey/west:87 estimator. The conditions $T_{n}^{\zeta}b_{1,T}^{1/2}\rightarrow0$ and $T_{n}^{\zeta}(\widehat{b}_{1,T})^{1/2}\rightarrow0$ mean that the length of the regime in which $x_{t}^{(2)}$ exhibits a shift $\delta$ in the mean increases to infinity at a slower rate than $T$. Theorem (ref) shows that when the HAC estimators or the fixed-$b$ LRV estimators are used, the DM test is not consistent and its power approaches zero. The theorem also implies that the power functions corresponding to tests based on HAC estimators lie above the power functions corresponding to those based on fixed-$b$/EWC LRV estimators. This follows from $|t_{\mathrm{DM},\mathrm{KVB}}|\ll|t_{\mathrm{DM},\mathrm{NW87}}|.$ Another interesting feature is that $|t_{\mathrm{DM},\mathrm{NW87}}|$ and $|t_{\mathrm{DM},\mathrm{KVB}}|$ do not increase in magnitude with $\delta$ because $\delta$ appears in both the numerator and denominator ($\delta$ enters the denominator through the low frequency contamination term $d^{*}$ that accounts for the bias in the HAC and fixed-$b$ estimators (cf. Theorem (ref))). Part (iii) of the theorem suggests that these issues do not occur when the DK-HAC estimator is used since the test is consistent and its power increases with $\delta$ and with the sample size as it should be. These results match the empirical results in Table (ref) discussed above, thereby confirming the relevance of Theorem (ref).
Economic time series often display nonstationary features that are usefully addressed in testing by allowing for some misspecification in standard model formulations. If nonstationarity is not accounted for properly, parameter estimates and, in particular, asymptotic LRV estimates can be largely biased. We establish results on the low frequency contamination induced by nonstationarity and misspecification for the sample autocovariance and the periodogram under general conditions. These estimates can exhibit features akin to long memory when the data are nonstationary short memory. We show, using theoretical arguments, that nonparametric smoothing is robust. Since the autocovariances and the periodogram are basic elements for HAR inference, our results allow a better understanding of LRV estimation. Under the null hypothesis there are larger size distortions than when the data are stationary. Under the alternative hypothesis, existing LRV estimators tend to be inflated and HAR tests can exhibit dramatic power losses. Long bandwidths/fixed-$b$ HAR tests suffer more from low frequency contamination relative to HAR tests based on HAC estimators, whereas the DK-HAC estimators do not suffer from this problem.
Casini, A., T. Deng and P. Perron (2024): Supplement to “Theory of low frequency contamination from nonstationarity and misspecification: consequences for HAR inference\textquotedbl , Econometric Theory Supplementary Material.
\addcontentsline{toc}{section}{References}