The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
131,169 characters
Theory of Low Frequency Contamination from Nonstationarity and Misspecification: Consequences for HAR Inference
\setcounter{page}{0}
\title{\textbf{\Large{}Theory of Low Frequency Contamination from Nonstationarity
and Misspecification: Consequences for HAR Inference}\textbf{}\thanks{We are grateful to Peter C.B. Phillips, Anna Mikusheva and the referees
for useful suggestions. We thank Whitney Newey and Tim Vogelsang for
discussions and Andrew Chesher, Adam McCloskey, Zhongjun Qu and
Daniel Whilem for comments.}\textbf{ }}
\maketitle
\begin{abstract}
We establish theoretical results about the low frequency contamination
(i.e., long memory effects) induced by general nonstationarity for
estimates such as the sample autocovariance and the periodogram, and
deduce consequences for heteroskedasticity and autocorrelation robust
(HAR) inference. We present explicit expressions for the asymptotic
bias of these estimates. We show theoretically that nonparametric
smoothing over time is robust to low frequency contamination.
Nonstationarity can have consequences for both the size and power
of HAR tests. Under the null hypothesis there are larger size distortions
than when data are stationary. Under the alternative hypothesis, existing
LRV estimators tend to be inflated and HAR tests can exhibit dramatic
power losses. Our theory indicates that long bandwidths or fixed-$b$
HAR tests suffer more from low frequency contamination relative to
HAR tests based on HAC estimators, whereas recently introduced double
kernel HAC estimators do not suffer from this problem. We present
second-order Edgeworth expansions under nonstationarity about the
distribution of HAC and DK-HAC estimators and about the corresponding
$t$-test in the regression model. The results show that the distortions
in the rejection rates can be induced by time variation in the second
moments even when there is no break in the mean.
\end{abstract}
\indent {\bf{JEL Classification}}: C12, C13, C18, C22, C32, C51\\
\noindent {\bf{Keywords}}: Edgeworth expansions, Fixed-$b$, HAC standard errors, HAR, Long memory, Long-run variance, Low frequency contamination, Nonstationarity, Outliers, Segmented locally stationary.
\onehalfspacing
\thispagestyle{empty}
\allowdisplaybreaks
\vfill{}
\pagebreak{}
\section{\label{Section Introduction Lap_BP}Introduction}
\begin{onehalfspace}
Many economic and financial time series have nonstationary characteristics
that need to be accounted for in inference {[}see, e.g., \citet{perron:89},
\citet{stock/watson:96}, \citet{ng/wright:13}, and \citet{giacomini/rossi:15}{]}.
We develop theoretical results about the behavior of the sample autocovariance
($\widehat{\Gamma}\left(k\right),\,k\in\mathbb{Z}$) and the periodogram
($I_{T}\left(\omega\right),\,\omega\in\left[-\pi,\,\pi\right]$) for
a short memory nonstationary process. This means processes that have
non-constant moments and whose sum of absolute autocovariances is
finite. The latter rules out processes with unbounded second moments
(e.g., unit root). We show that time-variation in the mean induces
low frequency contamination, meaning that the sample autocovariance
and the periodogram share features that are similar to those of a
long memory series. We present explicit expressions for the asymptotic
bias of these estimates, showing that it is always positive and increases
with the degree of heterogeneity in the data.
\end{onehalfspace}
The low frequency contamination can be explained as follows. For a
short memory series, the autocorrelation function (ACF) displays exponential
decay and vanishes as the lag length $k\rightarrow\infty$, and the
periodogram is finite at the origin. Under general forms of nonstationarity
involving changes in the mean, we show theoretically that $\widehat{\Gamma}\left(k\right)=\lim_{T\rightarrow\infty}\Gamma_{T}\left(k\right)+d^{*},$
where $\Gamma_{T}\left(k\right)=T^{-1}\sum_{t=k+1}^{T}\mathbb{E}\left(V_{t}V_{t-k}\right)$,
$k\geq0$ and $d^{*}>0$ is independent of $k$. Assuming positive
dependence for simplicity (i.e., $\lim_{T\rightarrow\infty}\Gamma_{T}\left(k\right)>0$),
that means that each sample autocovariance overestimates the true
dependence in the data. The bias factor $d^{*}>0$ depends on the
type of nonstationarity and in general does not vanish as $T\rightarrow\infty$.
In addition, since short memory implies $\Gamma_{T}\left(k\right)\rightarrow0$
as $k\rightarrow\infty$, it follows that $d^{*}$ generates long
memory effects since $\widehat{\Gamma}\left(k\right)\thickapprox d^{*}>0$
as $k\rightarrow\infty$. As for the periodogram, $I_{T}\left(\omega\right)$,
we show that under nonstationarity $\mathbb{E}\left(I_{T}\left(\omega\right)\right)\rightarrow\infty$
as $\omega\rightarrow0$, a feature also shared by long memory processes.
\begin{onehalfspace}
Several HAR inference problems in applied work (besides the $t$-
and $F$-test in regression models) are characterized by nonstationary
alternative hypotheses for which $d^{*}>0$ even asymptotically. This
class of tests is very large. Tests for forecast evaluation {[}e.g.,
\citet{casini_CR_Test_Inst_Forecast}, \citet{diebold/mariano:95},
\citeauthor{giacomini/rossi:09} (\citeyear{giacomini/rossi:09},
\citeyear{giacomini/rossi:10}), \citet{giacomini/white:06}, \citet{perron/yamamoto:18}
and \citet{west:96}{]}, tests and inference for structural changes
{[}e.g., \citet{andrews:93}, \citet{bai/perron:98}, \citeauthor{casini/perron_CR_Single_Break}
(\citeyear{casini/perron_SC_BP_Lap}, \citeyear{casini/perron_Lap_CR_Single_Inf},
\citeyear{casini/perron_CR_Single_Break}), \citet{elliott/mueller:07},
and \citet{qu/perron:07}{]}, tests and inference in time-varying
parameters models {[}e.g., \citet{cai:07} and \citet{chen/hong:12}{]},
tests and inference for regime switching models {[}e.g., \citet{hamilton:89}
and \citet{qu/zhuo:2020}{]} and others are part of this class.
\end{onehalfspace}
Recently, \citet{casini_hac} proposed a new HAC estimator that applies
nonparametric smoothing over time in order to account flexibly for
nonstationarity. We show theoretically that nonparametric smoothing
over time is robust to low frequency contamination and prove that
the resulting sample local autocovariance and the local periodogram
do not exhibit long memory features. Nonparametric smoothing avoids
mixing highly heterogeneous data coming from distinct nonstationary
regimes as opposed to what the sample autocovariance and the periodogram
do.
Our work is different from the literature on spurious persistence
caused by the presence of level shifts or other deterministic trends.
\citet{perron:90} showed that the presence of breaks in mean often
induces spurious non-rejection of the unit root hypothesis, and that
the presence of a level shift asymptotically biases the estimate of
the AR coefficient towards one. \citet{bhattacharya/gupta/waymire:83}
demonstrated that certain deterministic trends can induce the spurious
presence of long memory. In other contexts, similar issues were discussed
by \citet{varneskov/christensen:17}, \citet{diebold/inoue:01}, \citet{demetrescu/salish:2020},
\citet{lamoureux/lastrapes:1990}, \citet{hillebrand:05}, \citet{granger/hyung:04},
\citet{mccloskey/hill:2017}, \citet{mikosh/starica:04}, \citet{muller/watson:2008}
and \citet{perron/qu:2010}. Our results are different from theirs
in that we consider a more general problem and we allow for more general
forms of nonstationarity using the segmented locally stationary framework
of \citet{casini_hac}. Importantly, we provide a general solution
to these problems and show theoretically its robustness to low frequency
contamination. Moreover, we discuss in detail the implications of
our theory for HAR inference.
HAR inference relies on estimation of the long-run variance (LRV).
The latter, from a time domain perspective, is equivalent to the sum
of all autocovariances while from a frequency domain perspective,
is equal to $2\pi$ times an integrated time-varying spectral density
at the zero frequency. From a time domain perspective, estimation
involves a weighted sum of the sample autocovariances, while from
a frequency domain perspective estimation is based on a weighted sum
of the periodogram ordinates near the zero frequency. Therefore, our
results on low frequency contamination for the sample autocovariances
and the periodogram can have important implications.
There are two main approaches in HAR inference, one based on traditional
asymptotics and the other based on fixed-smoothing asymptotics. The
classical approach relies on an LRV estimator using a small bandwidth
{[}cf. the HAC estimators of \citeauthor{newey/west:87} (\citeyear{newey/west:87},
\citeyear{newey/west:94}) and \citet{andrews:91}{]}. Inference is
standard because HAR test statistics follow asymptotically standard
distributions. It was shown early that HAC standard errors can result
in oversized tests when there is substantial temporal dependence.
This stimulated a second approach based on an LRV estimator that
keeps the bandwidth at a fixed fraction of the sample size and that
converges weakly to a random variable {[}cf. \citet{Kiefer/vogelsang/bunzel:00}{]}.
Inference is then based on a nonstandard reference distribution
and it is shown that fixed-$b$ achieves high-order refinements {[}e.g.,
\citet{sun/phillips/jin:08}{]} and reduces the oversize problem of
HAR tests.\footnote{See \citet{dou:18}, \citet{hwang/sun:2017}, \citet{ibragimov/kattuman/skrobotov:2021},
\citet{ibragimov/muller:10}, \citet{jansson:04}, \citeauthor{Kiefer/vogelsang:02}
(\citeyear{Kiefer/vogelsang:02}, \citeyear{kiefer/vogelsang:05}),
\citet{lazarus/lewis/stock:17}, \textcolor{MyBlue}{Lazarus et al.}
\citeyearpar{lazarus/lewis/stock/watson:18} \citeauthor{muller:07}
(\citeyear{muller:07}, \citeyear{mueller:14}), \citet{phillips:05},
\citet{politis:11}, \citeauthor{potscher/preinerstorfer:18} (\citeyear{preinerstorfer/potscher:16},
\citeyear{potscher/preinerstorfer:18}, \citeyear{potscher/preinerstorfer:19}),
\citet{robinson:98}, \citeauthor{sun:14} (\citeyear{sun:13}, \citeyear{sun:14a},
\citeyear{sun:14}) and \citet{zhang/shao:13}.} However, unlike the classical approach, current fixed-$b$ HAR inference
is only valid under stationarity {[}cf. \citet{casini_fixed_b_erp}{]}
as the fixed-$b$ limiting distribution of the $t$/$F$ statistic
is non-pivotal under nonstationarity. More recently, a variant of
the fixed-$b$ approach {[}see, e.g., \citet{sun:14} and \citet{lazarus/lewis/stock/watson:18}{]}
considered the use of small-$b$ asymptotics in conjunction with
fixed-$b$ or $t/F$ critical values. These bandwidths are typically
larger than the MSE-optimal bandwidths used for the HAC estimators.
Recently, \citet{casini_hac} questioned the performance of HAR inference
under nonstationarity from a theoretical standpoint. Simulation evidence
of serious (e.g., non-monotonic) power or related issues in specific
HAR inference contexts were documented by \citet{altissimo/corradi:2003},
\citet{casini_CR_Test_Inst_Forecast}, \citeauthor{casini/perron_Lap_CR_Single_Inf}
(\citeyear{casini/perron:OUP-Breaks}, \citeyear{casini/perron_SC_BP_Lap},
\citeyear{casini/perron_Lap_CR_Single_Inf}), \citeauthor{chan:2020}
(\citeyear{chan:2022}, \citeyear{chan:2020}), \citet{crainiceanu/vogelsang:07},
\citet{deng/perron:06}, \citet{juhl/xiao:09}, \citet{kim/perron:09},
\citet{martins/perron:16}, \citet{otto/breitung:2021}, \citet{perron:1991},
\citet{perron/yamamoto:18}, \citet{shao/zhang:2010}, \citet{vogeslang:99}
and \citet{zhang/lavitas:2018} among others{]}. Our theoretical results
show that these issues occur because the unaccounted nonstationarity
alters the spectrum at low frequencies. Each sample autocovariance
is upward biased ($d^{*}>0$) and the resulting LRV estimators tend
to be inflated. When these estimators are used to normalize test statistics,
the latter lose power. Interestingly, $d^{*}$ is independent
of $k$ so that the more lags are included the more severe is the
problem. Further, by virtue of weak dependence, we have that $\Gamma_{T}\left(k\right)\rightarrow0$
as $k\rightarrow\infty$ but $d^{*}>0$ across $k$. We show formally
that long bandwidths/fixed-$b$ LRV estimators are expected to suffer
most from power losses because they use many/all lagged autocovariances.
To precisely analyze the theoretical properties of the HAR tests under
the null hypothesis, we present second-order Edgeworth expansions
under nonstationarity for the distribution of the HAC and DK-HAC estimator
and for the distribution of the corresponding $t$-test in the linear
regression model. Under stationarity the results concerning the HAC
estimator were provided by \citet{velasco/robinson:01}. We show that
the order of the approximation error of the expansion is the same
as under stationarity from which it follows that the error in rejection
probability (ERP) is also the same. The ERP of the $t$-test based
on the DK-HAC estimator is slightly larger than that of the $t$-test
based on the HAC estimator due to the double smoothing. High-order
asymptotic expansions for spectral and other estimates were studied
by \citet{bhattacharya/ghosh:1978}, \citet{bentkus/rudzkis:1982},
\citet{janas:1994}, \citeauthor{phillips:1977} (\citeyear{phillips:1977},
\citeyear{phillips:1980}) and \citet{taniguchi/puri:1996}. The asymptotic
expansions of the fixed-$b$ HAR tests under stationarity were developed
by \citet{jansson:04} and \citet{sun/phillips/jin:08}. \citet{casini_fixed_b_erp}
showed that under nonstationarity the ERP of the fixed-$b$ HAR tests
can be larger than that of HAR tests based on HAC and DK-HAC estimators
thereby controverting the conclusion in the literature that the original
fixed-$b$ HAR tests have superior null rejection rates relative to
HAR tests based on traditional LRV estimators. \citet{casini_fixed_b_erp}
also developed fixed-$b$ methods that are valid under nonstationarity
and in fact provide better null rejection rates in finite-sample.
The Monte Carlo results suggest that under the null hypothesis nonstationarity
can generate larger size distortions than what one finds under stationarity.
In particular, fixed-smoothing methods can exhibit under-rejections
whereas HAC and DK-HAC methods can exhibit over-rejections when there
is strong persistence. For the latter problem, our second-order Edgeworth
expansions could be used to construct corrections to the standard
normal critical value. We relegate this opportunity to future research.
The paper is organized as follows. Section \ref{Section, Statistical Framework for Nonstationarity}
presents the statistical setting and Section \ref{Section Low Freq Cont - Theory}
establishes the theoretical results on low frequency contamination.
Section \ref{Section Edgeworth-Expansions-for} presents the Edgeworth
expansions of HAR tests based on the HAC and DK-HAC estimators. The
implications of our results for HAR inference are analyzed analytically
and computationally through simulations in Section \ref{Section Consequences for HAR}.
Section \ref{Section Conclusions} concludes. The supplemental materials
{[}cf. \citet{casini/perron_Low_Frequency_Contam_Nonstat:2020_supp}{]}
contain some additional examples and all mathematical proofs.
\section{\label{Section, Statistical Framework for Nonstationarity}Statistical
Framework for Nonstationarity}
Suppose $\{V_{t,T}\}_{t=1}^{T}$ is defined on a probability space
$\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$, where $\Omega$
is the sample space, $\mathscr{F}$ is the $\sigma$-algebra and $\mathbb{P}$
is a probability measure. In order to analyze time series models that
have a time-varying spectrum it is useful to introduce an infill asymptotic
setting whereby we rescale the original discrete time horizon $\left[1,\,T\right]$
by dividing each $t$ by $T.$ Letting $u=t/T$ we define a new time
scale $u\in\left[0,\,1\right]$ on which as $T\rightarrow\infty$
we observe more and more realizations of $V_{t,T}$ close to time
$t$. As a notion of nonstationarity, we use the concept of segmented
local stationarity (SLS) introduced in \citet{casini_hac}. This
extends the locally stationary processes {[}cf. \citet{dahlhaus:96}{]}
to allow for structural change and regime switching-type models.
SLS processes allow for a finite number of discontinuities in the
spectrum over time. We collect the break dates in the set $\mathcal{T}\triangleq\left\{ T_{1}^{0},\,\ldots,\,T_{m}^{0}\right\} $.
Let $i\triangleq\sqrt{-1}.$ A function $G\left(\cdot,\,\cdot\right):\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$
is said to be left-differentiable at $u_{0}$ if $\partial G\left(u_{0},\omega\right)/\partial_{-}u\triangleq\lim_{u\rightarrow u_{0}^{-}}\left(G\left(u_{0},\,\omega\right)-G\left(u,\,\omega\right)\right)/\left(u_{0}-u\right)$
exists for any $\omega\in\mathbb{R}$. Let $m_{0}\geq0$ be a finite
integer.
\begin{defn}
\label{Definition Segmented-Locally-Stationary}A sequence of stochastic
processes $\{V_{t,T}\}_{t=1}^{T}$ is called segmented locally stationary
(SLS) with $m_{0}+1$ regimes, transfer function $A^{0}$ and trend
$\mu$ if there exists a representation
\begin{align}
V_{t,T} & =\mu_{j}\left(t/T\right)+\int_{-\pi}^{\pi}\exp\left(i\omega t\right)A_{j,t,T}^{0}\left(\omega\right)d\xi\left(\omega\right),\qquad\qquad\left(t=T_{j-1}^{0}+1,\ldots,\,T_{j}^{0}\right),\label{Eq. Spectral Rep of SLS}
\end{align}
for $j=1,\ldots,\,m_{0}+1$, where by convention $T_{0}^{0}=0$ and
$T_{m_{0}+1}^{0}=T$. The following technical conditions are also
assumed to hold: (i) $\xi\left(\lambda\right)$ is a process on $\left[-\pi,\,\pi\right]$
with $\overline{\xi\left(\omega\right)}=\xi\left(-\omega\right)$
and
\begin{align*}
\mathrm{cum}\left\{ d\xi\left(\omega_{1}\right),\ldots,\,d\xi\left(\omega_{r}\right)\right\} & =\zeta\left(\sum_{j=1}^{r}\omega_{j}\right)g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)d\omega_{1}\ldots d\omega_{r},
\end{align*}
where $\mathrm{cum}\left\{ \cdots\right\} $ denotes the cumulant
spectra of $r$-th order, $g_{1}=0,\,g_{2}\left(\omega\right)=1$,
$\left|g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)\right|\leq M_{r}$
for all $r$ with $M_{r}<\infty$ that may depend on $r$, and $\zeta\left(\omega\right)=\sum_{j=-\infty}^{\infty}\delta\left(\omega+2\pi j\right)$
is the period $2\pi$ extension of the Dirac delta function $\delta\left(\cdot\right)$;
(ii) There exists a $C<\infty$ and a piecewise continuous function
$A:\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ such
that, for each $j=1,\ldots,\,m_{0}+1$, there exists a $2\pi$-periodic
function $A_{j}:\,(\lambda_{j-1}^{0},\,\lambda_{j}^{0}]\times\mathbb{R}\rightarrow\mathbb{C}$
with $A_{j}\left(u,\,-\omega\right)=\overline{A_{j}\left(u,\,\omega\right)}$,
$\lambda_{j}^{0}\triangleq T_{j}^{0}/T$ and for all $T,$
\begin{align}
A\left(u,\,\omega\right) & =A_{j}\left(u,\,\omega\right)\,\mathrm{\,for\,}\,\lambda_{j-1}^{0}<u\leq\lambda_{j}^{0},\label{Eq A(u) =00003D Ai}\\
\sup_{1\leq j\leq m_{0}+1} & \sup_{T_{j-1}^{0}<t\leq T_{j}^{0},\,\omega}\left|A_{j,t,T}^{0}\left(\omega\right)-A_{j}\left(t/T,\,\omega\right)\right|\leq CT^{-1};\label{Eq. 2.4 Smothenss Assumption on A}
\end{align}
(iii) $\mu_{\cdot}\left(\cdot\right)$ is piecewise Lipschitz continuous.
\end{defn}
Definition \ref{Definition Segmented-Locally-Stationary} states that
$V_{t,T}$ has a time-varying spectral representation where both the
mean $\mu_{\cdot}\left(\cdot\right)$ and transfer function $A_{\cdot,\cdot,T}^{0}\left(\omega\right)$
are piecewise continuous. Since the transfer function depends on the
parameters that enter the second moments of $V_{t,T}$, the smoothness
properties of $\mu_{\cdot}\left(\cdot\right)$ and $A$ guarantee
that $V_{t,T}$ has a piecewise locally stationary behavior. We require
additional smoothness properties for $A$ and an example is presented
at the end of this section.
\begin{assumption}
\label{Assumption Smothness of A (for HAC)}(i) $\left\{ V_{t,T}\right\} $
is an SLS process with $m_{0}+1$ regimes; (ii) $A\left(u,\,\omega\right)$
is \textcolor{red}{ }twice continuously differentiable in $u$ at
all $u\neq\lambda_{j}^{0}$, $j=1,\ldots,\,m_{0}+1,$ with bounded
derivatives $\left(\partial/\partial u\right)A\left(u,\,\cdot\right)$
and $\left(\partial^{2}/\partial u^{2}\right)A\left(u,\,\cdot\right)$;
(iii) $\left(\partial^{2}/\partial u^{2}\right)A\left(u,\,\cdot\right)$
is Lipschitz continuous at all $u\neq\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$;
(iv) $A\left(u,\,\omega\right)$ is twice left-differentiable in $u$
at $u=\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with bounded derivatives
$\left(\partial/\partial_{-}u\right)A\left(u,\,\cdot\right)$ and
$\left(\partial^{2}/\partial_{-}u^{2}\right)A\left(u,\,\cdot\right)$
and has piecewise Lipschitz continuous derivative $\left(\partial^{2}/\partial_{-}u^{2}\right)A\left(u,\,\cdot\right)$;
(v) $A\left(u,\,\omega\right)$ is Lipschitz continuous in $\omega.$
\end{assumption}
We define the time-varying spectral density as $f_{j}\left(u,\,\omega\right)\triangleq(2\pi)^{-1}|A_{j}\left(u,\,\omega\right)|^{2}$
for $T_{j-1}^{0}/T<u=t/T\leq T_{j}^{0}/T$. Then we can define the
local covariance of $V_{t,T}$ at the rescaled time $u$ with $Tu\notin\mathcal{T}$
and lag $k\in\mathbb{Z}$ as $c\left(u,\,k\right)\triangleq\int_{-\pi}^{\pi}e^{i\omega k}f\left(u,\,\omega\right)d\omega$.
The same definition is also used when $Tu\in\mathcal{T}$ and $k\geq0$.
For $Tu\in\mathcal{T}$ and $k<0$ it is defined as $c\left(u,\,k\right)\triangleq\lim_{T\rightarrow\infty}\int_{-\pi}^{\pi}e^{i\omega k}A\left(u,\,\omega\right)A\left(u-k/T,\,-\omega\right)d\omega$.
Next, we impose conditions on the temporal dependence (we omit the
second subscript $T$ when it is clear from the context). Let
\begin{align*}
\kappa_{V,t}^{\left(a_{1},a_{2},a_{3},a_{4}\right)} & \left(u,\,v,\,w\right)\\
& \triangleq\kappa^{\left(a_{1},a_{2},a_{3},a_{4}\right)}\left(t,\,t+u,\,t+v,\,t+w\right)-\kappa_{\mathscr{N}}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}\left(t,\,t+u,\,t+v,\,t+w\right)\\
& \triangleq\mathbb{E}\left(V_{t}^{\left(a_{1}\right)}-\mathbb{E}V_{t}^{\left(a_{1}\right)}\right)\left(V_{t+u}^{\left(a_{2}\right)}-\mathbb{E}V_{t+u}^{\left(a_{2}\right)}\right)\left(V_{t+v}^{\left(a_{3}\right)}-\mathbb{E}V_{t+v}^{\left(a_{3}\right)}\right)\left(V_{t+w}^{\left(a_{4}\right)}-\mathbb{E}V_{t+w}^{\left(a_{4}\right)}\right)\\
& \quad-\mathbb{E}\left(V_{\mathscr{N},t}^{\left(a_{1}\right)}-\mathbb{E}V_{\mathscr{N},t}^{\left(a_{1}\right)}\right)\left(V_{\mathscr{N},t+u}^{\left(a_{2}\right)}-\mathbb{E}V_{\mathscr{N},t+u}^{\left(a_{2}\right)}\right)\left(V_{\mathscr{N},t+v}^{\left(a_{3}\right)}-\mathbb{E}V_{\mathscr{N},t+v}^{\left(a_{3}\right)}\right)\left(V_{\mathscr{N},t+w}^{\left(a_{4}\right)}-\mathbb{E}V_{\mathscr{N},t+w}^{\left(a_{4}\right)}\right),
\end{align*}
where $\left\{ V_{\mathscr{N},t}\right\} $ is a Gaussian sequence
with the same mean and covariance structure as $\left\{ V_{t}\right\} $,
$\kappa_{V,t}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}\left(u,\,v,\,w\right)$
is the time-$t$ fourth-order cumulant of $(V_{t}^{\left(a_{1}\right)},\,V_{t+u}^{\left(a_{2}\right)},\,V_{t+v}^{\left(a_{3}\right)},$
$\,V_{t+w}^{\left(a_{4}\right)})$ while $\kappa_{\mathscr{N}}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}$
$(t,\,t+u,\,t+v,\,t+w)$ is the time-$t$ centered fourth moment of
$V_{t}$ if $V_{t}$ were Gaussian.
\begin{assumption}
\label{Assumption A - Dependence}(i) $\sum_{k=-\infty}^{\infty}\sup_{u\in\left[0,\,1\right]}$
$\left\Vert c\left(u,\,k\right)\right\Vert <\infty$ and $\sum_{k=-\infty}^{\infty}\sum_{j=-\infty}^{\infty}\sum_{l=-\infty}^{\infty}\sup_{u\in\left[0,\,1\right]}|\kappa_{V,\left\lfloor Tu\right\rfloor }^{\left(a_{1},a_{2},a_{3},a_{4}\right)}$
$\left(k,\,j,\,l\right)|<\infty$ for all $a_{1},a_{2},a_{3},a_{4}\leq p$.
(ii) For all $a_{1},a_{2},a_{3},a_{4}\leq p$ there exists a function
$\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}:\,\left[0,\,1\right]\times\mathbb{Z}\times\mathbb{Z}\times\mathbb{Z}\rightarrow\mathbb{R}$
such that $\sup_{1\leq j\leq m_{0}+1}\sup_{\lambda_{j-1}^{0}<u\leq\lambda_{j}^{0}}|\kappa_{V,\left\lfloor Tu\right\rfloor }^{\left(a_{1},a_{2},a_{3},a_{4}\right)}\left(k,\,s,\,l\right)-\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}$
$\left(u,\,k,\,s,\,l\right)|\leq LT^{-1}$ for some constant $L$;
the function $\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}\left(u,\,k,\,s,\,l\right)$
is twice differentiable in $u$ at all $u\neq\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$
with bounded derivatives $\left(\partial/\partial u\right)\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}$
$\left(u,\cdot,\cdot,\cdot\right)$ and $\left(\partial^{2}/\partial u^{2}\right)\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}\left(u,\cdot,\cdot,\cdot\right)$,
and twice left-differentiable in $u$ with bounded derivatives $\left(\partial/\partial_{-}u\right)\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}\left(u,\cdot,\cdot,\cdot\right)$
and $\left(\partial^{2}/\partial_{-}u^{2}\right)\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}\left(u,\cdot,\cdot,\cdot\right)$,
and piecewise Lipschitz continuous derivative $\left(\partial^{2}/\partial_{-}u^{2}\right)\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}\left(u,\cdot,\cdot,\cdot\right)$.
\end{assumption}
If $\left\{ V_{t}\right\} $ is stationary then the cumulant condition
of Assumption \ref{Assumption A - Dependence}-(i) reduces to the
standard one used in the time series literature {[}see \citet{andrews:91}{]}.
Note that $\alpha$-mixing and some moment conditions imply that
the cumulant condition of Assumption \ref{Assumption A - Dependence}
holds. Part (ii) extends the smoothness conditions on $A\left(u,\,\omega\right)$
in Assumption \ref{Assumption Smothness of A (for HAC)} to the fourth-order
cumulant. These smoothness conditions are not particularly restrictive.
Consider the following time-varying AR(1) process with one break at
mid-sample $\lambda_{1}^{0}=0.5$,
\begin{align}
V_{t,T} & =\rho\left(t/T\right)V_{t-1,T}+\sigma\left(t/T\right)u_{t},\label{Eq. Example TV AR(1)}\\
\rho\left(u\right) & =\begin{cases}
\rho_{1}\left(u\right), & u\leq0.5\\
\rho_{2}\left(u\right), & u>0.5
\end{cases},\nonumber
\end{align}
where $\rho_{1}\left(\cdot\right)$ and $\rho_{2}\left(\cdot\right)$
are Lipschitz continuous, $\sigma\left(\cdot\right)$ is piecewise
Lipschitz continuous and $\left\{ u_{t}\right\} $ are i.i.d. random
variables with mean zero and unit variance. Then, $V_{t,T}$ is an
SLS process with $A\left(u,\,\omega\right)=\sigma\left(u\right)\left(1+\rho\left(u\right)\exp\left(i\omega\right)\right)$.
If $\rho\left(u\right)$ and $\sigma\left(u\right)$ satisfy the same
smoothness conditions in $u$ required for $A\left(u,\,\omega\right)$
in Assumption \ref{Assumption Smothness of A (for HAC)}, $\sup_{u\in\left[0,\,1\right]}\left|\rho\left(u\right)\right|<1$
and $\sup_{u\in\left[0,\,1\right]}\sigma\left(u\right)<\infty$, then
$V_{t,T}$ fulfills Assumption \ref{Assumption Smothness of A (for HAC)}-\ref{Assumption A - Dependence}.
\section{\label{Section Low Freq Cont - Theory}Theoretical Results on Low
Frequency Contamination}
In this section we establish theoretical results about the low frequency
contamination induced by nonstationarity, misspecification and outliers.
We first consider the asymptotic proprieties of two key quantities
for inference in time series contexts, i.e., the sample autocovariance
and the periodogram. These are defined, respectively, by
\begin{align}
\widehat{\Gamma}\left(k\right) & =T^{-1}\sum_{t=|k|+1}^{T}\left(V_{t}-\overline{V}\right)\left(V_{t-|k|}-\overline{V}\right),\label{Eq. Definition of Gamma(k)}
\end{align}
where $\overline{V}$ is the sample mean and
\begin{align*}
I_{T}\left(\omega\right) & =\left|\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\exp\left(-i\omega t\right)V_{t}\right|^{2},\qquad\qquad\omega\in\left[0,\,\pi\right],
\end{align*}
which is evaluated at the Fourier frequencies $\omega_{j}=\left(2\pi j\right)/T\in[0,\,\pi]$.
In the context of autocorrelated data, hypotheses testing and construction
of confidence intervals require estimation of the so-called long-run
variance. Traditional HAC estimators are weighted sums of sample autocovariances
while frequency domain estimators are weighted sums of the periodograms.
\citet{casini_hac} considered an alternative estimate for the sample
autocovariance to be used in the DK-HAC estimators, defined in Section
\ref{Subsection HAR inference methods}, namely,
\begin{align*}
\widehat{\Gamma}_{\mathrm{DK}}\left(k\right) & \triangleq\frac{n_{T}}{T}\sum_{r=1}^{\left\lfloor T/n_{T}\right\rfloor }\widehat{c}_{T}\left(rn_{T}/T,\,k\right),
\end{align*}
where $k\in\mathbb{Z},$ $n_{T}\rightarrow\infty$ satisfying the
conditions given below, and
\begin{align}
\widehat{c}_{T}\left(rn_{T}/T,\,k\right) & =n_{2,T}^{-1}\sum_{s=0}^{n_{2,T}-1}\left(V_{rn_{T}+\left\lfloor |k/2|\right\rfloor -n_{2,T}/2+s+1}-\overline{V}{}_{rn_{T},T}\right)\left(V_{rn_{T}-\left\lfloor |k/2|\right\rfloor -n_{2,T}/2+s+1}-\overline{V}{}_{rn_{T},T}\right),\label{Eq. chat}
\end{align}
with $\overline{V}{}_{rn_{T},T}=n_{2,T}^{-1}\sum_{s=0}^{n_{2,T}-1}V_{rn_{T}-n_{2,T}/2+s+1}$
and $n_{2,T}\rightarrow\infty$ such that $n_{2,T}/T\rightarrow0$.
For notational simplicity we assume that $n_{T}$ and $n_{2,T}$ are
even. $\widehat{c}_{T}\left(rn_{T}/T,\,k\right)$ is an estimate of
the autocovariance at time $rn_{T}$ and lag $k$, i.e., $\mathrm{cov}(V_{rn_{T}},\,V_{rn_{T}-k})$.
One could use a smoothed or tapered version; the estimate $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$
is an integrated local sample autocovariance. It extends $\widehat{\Gamma}\left(k\right)$
to better account for nonstationarity. Similarly, the DK-HAC estimator
does not relate to the periodogram but to the local periodogram defined
by
\begin{align*}
I_{\mathrm{L},T}\left(u,\,\omega\right) & \triangleq\left|\frac{1}{\sqrt{n_{T}}}\sum_{s=0}^{n_{T}-1}V_{\left\lfloor Tu\right\rfloor -n_{T}/2+s+1,T}\exp\left(-i\omega s\right)\right|^{2},
\end{align*}
where $I_{\mathrm{L},T}\left(u,\,\omega\right)$ is the (untapered)
periodogram over a segment of length $n_{T}$ with midpoint $\left\lfloor Tu\right\rfloor $.
We also consider the statistical properties of both $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$
and $I_{\mathrm{L},T}\left(u,\,\omega\right)$ under nonstationarity.
Define $r_{j}=(\lambda_{j}^{0}-\lambda_{j-1}^{0})$ for $j=1,\ldots,\,m_{0}+1$
with $\lambda_{0}^{0}=0$ and $\lambda_{m_{0}+1}^{0}=1$. Note that
$\lambda_{j}^{0}=\sum_{s=0}^{j}r_{s}.$
The low frequency bias is generated by breaks in the mean function.
For the sample autocovariance, the bias factor is given by $d^{*}=2^{-1}\sum_{j_{1}\neq j_{2}}r_{j_{1}}r_{j_{2}}(\overline{\mu}_{j_{2}}-\overline{\mu}_{j_{1}})^{2}$
where
\begin{align*}
\overline{\mu}_{j} & =r_{j}^{-1}\int_{\lambda_{j-1}^{0}}^{\lambda_{j}^{0}}\mu_{j}\left(u\right)du,\qquad\mathrm{for}\,j=1,\ldots,\,m_{0}+1,
\end{align*}
with $\mu_{j}\left(\cdot\right)$ defined in \eqref{Eq. Spectral Rep of SLS}
and we use $\sum_{j_{1}\neq j_{2}}$ as a shorthand for $\sum_{\left\{ j_{1},\,j_{2}=1,\ldots,\,m_{0}+1,\,j_{1}\neq j_{2}\right\} }.$
When the mean is constant in each regime $\mu_{j}\left(t/T\right)=\mu_{j}$.
Then, $\overline{\mu}_{j}=\mu_{j}$ and $d^{*}=2^{-1}\sum_{j_{1}\neq j_{2}}r_{j_{1}}r_{j_{2}}(\mu_{j_{2}}-\mu_{j_{1}})^{2}.$
If the mean is constant across regimes, then there is no low frequency
bias and $d^{*}=0.$
In Section \ref{Subsection Extension Sample Autocovariance and Periodogram}
we generalize the results in the literature on low frequency contamination
for the sample autocovariance and the periodogram. In Section \ref{Subsection Local Autocovariance and Local Periodogram}
we show that the local sample autocovariance and the local periodogram
are in general robust to low frequency contamination.
\subsection{\label{Subsection Extension Sample Autocovariance and Periodogram}The
Sample Autocovariance and the Periodogram Under Nonstationarity }
\citet*{mikosh/starica:04} established some results on the low frequency
bias for the sample autocovariance and periodogram under the assumption
that $V_{t}$ is stationary in each regime and that the regimes are
independent. In Section \ref{Section Results on Low Frequency Contamination}
in the supplement we extend these results by allowing time-varying
mean and autocovariace function in each regime and weak dependence
across regimes. Here we present a brief summary of these results.
Theorem \ref{Theorem ACF Nonstat} shows that for $\left\{ V_{t,T}\right\} $
that satisfies Definition \ref{Definition Segmented-Locally-Stationary}
and Assumption \ref{Assumption Smothness of A (for HAC)}-\ref{Assumption A - Dependence},
we have
\begin{align}
\widehat{\Gamma}\left(k\right)\geq & \int_{0}^{1}c\left(u,\,k\right)du+d^{*}+o_{\mathrm{a.s}.}\left(1\right),\label{Eq. Gamma_hat(k) Ineq Theorem-1}
\end{align}
and as $k\rightarrow\infty,$ $\widehat{\Gamma}\left(k\right)\geq d^{*}$
$\mathbb{P}$-a.s. This suggests that $\widehat{\Gamma}\left(k\right)$
is asymptotically the sum of two terms. The first is the autocovariance
of $\left\{ V_{t}\right\} $ at lag $k$. The second, $d_{\mathrm{}}^{*}$,
is always positive and increases with the difference in the mean
across regimes. Thus, the time-varying mean induces a positive bias.
The result that $\widehat{\Gamma}\left(k\right)\geq d^{*}$ $\mathbb{P}$-a.s.
as $k\rightarrow\infty$ implies that unaccounted nonstationarity
generates long memory effects. The intuition is straightforward. A
long memory SLS process satisfies $\sum_{k=-\infty}^{\infty}|\Gamma\left(u,\,k\right)|\rightarrow\infty$
for some $u\in\left(0,\,1\right)$, similar to a stationary long memory
process.\footnote{In Section \ref{Subsection Long-Memory-Segmented-Locally} in the
supplement we define long memory SLS processes that are characterized
by the property $\sum_{k=-\infty}^{\infty}\left|\rho_{V}\left(u,\,k\right)\right|=\infty$
for some $u\in\left[0,\,1\right]$ where $\rho_{V}\left(u,\,k\right)\triangleq\mathrm{Corr}(V_{\left\lfloor Tu\right\rfloor },\,V_{\left\lfloor Tu\right\rfloor +k})$
and $\vartheta\left(u\right)\in\left(0,\,1/2\right)$ is the long
memory parameter at time $u$. } The theorem shows that $\widehat{\Gamma}\left(k\right)$ exhibits
a similar property and $\widehat{\Gamma}\left(k\right)$ decays more
slowly than for a short memory stationary process for small lags and
approaches a constant $d^{*}>0$ for large lags.
Theorem \ref{Theorem Periodogram Long Memory Effects} in the supplement
analyzes the properties of the periodogram $I_{T}\left(\omega_{l}\right)$
as $\omega\rightarrow0$ when the mean is time-varying. The result
states that as $\omega\rightarrow0$ $\mathbb{E}\left(I_{T}\left(\omega\right)\right)$
generally takes unbounded values except for some $\omega$ for which
$\mathbb{E}\left(I_{T}\left(\omega\right)\right)$ is bounded below
by $2\pi\int_{0}^{1}f\left(u,\,\omega\right)du>0.$ An SLS process
with long memory has an unbounded local spectral density $f\left(u,\,\omega\right)$
as $\omega\rightarrow0$ for some $u\in\left[0,\,1\right]$. Since
$f\left(\cdot,\,\cdot\right)$ cannot be negative, it follows that
$\int_{0}^{1}f\left(u,\,\omega\right)du$ is also unbounded as $\omega\rightarrow0$.
Theorem \ref{Theorem Periodogram Long Memory Effects} suggests that
nonstationarity consisting of time-varying first moment results in
a periodogram sharing features of a long memory series.
This discussion suggests that certain deviations from stationarity
can generate a long memory component that leads to overestimation
of the true autocovariance. It follows that the LRV is also overestimated.
Since the LRV is used to normalize test statistics, this has important
consequences for many HAR inference tests characterized by deviations
from stationarity under the alternative hypothesis. These include
tests for forecast evaluation, tests and inference for structural
change models, time-varying parameters models and regime-switching
models. In the linear regression model, $V_{t}$ corresponds to the
regressors multiplied by the fitted residuals. Unaccounted nonlinearities
and outliers can contaminate the mean of $V_{t}$ and therefore contribute
to $d^{*}$.
\subsection{\label{Subsection Local Autocovariance and Local Periodogram}The
Sample Local Autocovariance and Local Periodogram Under Nonstationarity }
We now consider the behavior of $\widehat{c}_{T}\left(rn_{T}/T,\,k\right)$
defined in \eqref{Eq. chat} for fixed $k$ as well as for $k\rightarrow\infty$.
For notational simplicity we assume that $k$ is even. For $u\in\left(0,\,1\right)$
define $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)=\{\left\lfloor Tu\right\rfloor +k/2-n_{2,T}/2+1,\ldots,\,\left\lfloor Tu\right\rfloor +k/2+n_{2,T}/2\}$,
$n_{j,L}\left(u,\,k,\,n_{2,T}\right)=(T_{j}^{0}-(\left\lfloor Tu\right\rfloor +k/2-n_{2,T}/2+1)),$
and $n_{j,R}\left(u,\,k,\,n_{2,T}\right)=((\left\lfloor Tu\right\rfloor +k/2+n_{2,T}/2+1)-T_{j}^{0})$.
$\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$ denotes a window of length
$n_{2,T}$ around $\left\lfloor Tu\right\rfloor $, $n_{j,L}\left(u,\,k,\,n_{2,T}\right)$
(resp. $n_{j,R}\left(u,\,k,\,n_{2,T}\right)$) denotes the distance
between the left (resp. right) end point of $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$
and $T_{j}^{0}$.
\begin{thm}
\label{Theorem Local ACF Nonstat}Assume that $\left\{ V_{t,T}\right\} $
satisfies Definition \ref{Definition Segmented-Locally-Stationary},
$n_{T},\,n_{2,T}\rightarrow\infty$ with $n_{T}/T\rightarrow0$, $n_{2,T}/T\rightarrow0$
and $n_{T}/n_{2,T}\rightarrow0.$ Under Assumption \ref{Assumption Smothness of A (for HAC)}-\ref{Assumption A - Dependence},
(i) for $u\in\left(0,\,1\right)$ such that $T_{j}^{0}\notin\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$
for all $j=1,\ldots,\,m_{0}$, $\widehat{c}_{T}\left(u,\,k\right)=c\left(u,\,k\right)+o_{\mathbb{P}}\left(1\right)$;
(ii) for $u\in\left(0,\,1\right)$ such that $T_{j}^{0}\in\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$
for some $j=1,\ldots,\,m_{0}$, we have two sub-cases: (a) if $n_{j,L}\left(u,\,k,\,n_{2,T}\right)/n_{2,T}\rightarrow\gamma$
or $n_{j,R}\left(u,\,k,\,n_{2,T}\right)/n_{2,T}\rightarrow\gamma$
with $\gamma\in\left(0,\,1\right)$, then
\begin{align*}
\widehat{c}_{T}\left(u,\,k\right) & \geq\gamma c\left(\lambda_{j}^{0},\,k\right)+\left(1-\gamma\right)c\left(u,\,k\right)+\gamma\left(1-\gamma\right)\left(\mu_{j}\left(\lambda_{j}^{0}\right)-\mu_{j+1}\left(u\right)\right)^{2}+o_{\mathbb{P}}\left(1\right).
\end{align*}
(b) if $n_{j,L}\left(u,\,k,\,n_{2,T}\right)/n_{2,T}\rightarrow0$
or $n_{j,R}\left(u,\,k,\,n_{2,T}\right)/n_{2,T}\rightarrow0$, then
$\widehat{c}_{T}\left(u,\,k\right)=c\left(u,\,k\right)+o_{\mathbb{P}}\left(1\right)$.
Further, if there exists an $r=1,\ldots,\,\left\lfloor T/n_{T}\right\rfloor $
such that there exists a $j=1,\ldots,\,m_{0}$ with $T_{j}^{0}\in\mathbf{S}\left(rn_{T},\,k,\,n_{2,T}\right)$
satisfying (ii-a), then, as $k\rightarrow\infty$, $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)\geq d_{T}^{*}$
$\mathbb{P}$-a.s., where $d_{T}^{*}=\left(n_{2,T}/T\right)\gamma\left(1-\gamma\right)$
$(\mu_{j}(\lambda_{j}^{0})-\mu_{j+1}\left(u\right))^{2}>0$ and $d_{T}^{*}\rightarrow0$
as $T\rightarrow\infty$.
\end{thm}
The theorem shows that the behavior of $\widehat{c}_{T}\left(u,\,k\right)$
depends on whether a change in mean is present, and if so whether
it is close enough to $\left\lfloor Tu\right\rfloor $. For a given
$u\in\left(0,\,1\right)$ and $k\in\mathbb{Z}$, if the condition
of part (i) of the theorem holds, then $\widehat{c}_{T}\left(u,\,k\right)$
is consistent for $\mathrm{cov}(V_{\left\lfloor Tu\right\rfloor }V_{\left\lfloor Tu\right\rfloor -k})=c\left(u,\,k\right)+O\left(T^{-1}\right)$
{[}see \citet{casini_hac}{]}. If a change-point falls close to either
boundary of the window $\mathbf{S}\left(u,\,k,\,n_{2,T}\right)$,
as specified in case (ii-b), then $\widehat{c}_{T}\left(u,\,k\right)$
remains consistent. The only case in which a non-negligible bias
arises is when the change-point falls in a neighborhood around $\left\lfloor Tu\right\rfloor $
sufficiently far from either boundary. This represents case (ii-a),
for which a biased estimate results. However, the bias vanishes asymptotically.
Since $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$ is an average
of $\widehat{c}_{T}\left(rn_{T},\,k\right)$ over blocks $r=1,\ldots,\,\left\lfloor T/n_{T}\right\rfloor $,
if case (ii-a) holds then $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)\geq d_{T}^{*}$
as $k\rightarrow\infty$ but $d_{T}^{*}\rightarrow0$ as $T\rightarrow\infty$.
Thus, comparing this result with the discussion above on $\widehat{\Gamma}\left(k\right)$
(see also Theorem \ref{Theorem ACF Nonstat}), in practice the long
memory effects are unlikely to occur when using $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$.
Furthermore, one can reduce this problem by appropriately choosing
the blocks $r=1,\ldots,\,\left\lfloor T/n_{T}\right\rfloor $. A
procedure was proposed in \citet{casini_hac} using the methods
developed in \citet{casini/perron:change-point-spectra}.
We now study the asymptotic properties of $I_{\mathrm{L},T}\left(u,\,\omega\right)$
as $\omega\rightarrow0$ for $u\in\left[0,\,1\right]$. We consider
the Fourier frequencies $\omega_{l}=2\pi l/n_{T}\in(-\pi,\,\pi)$
for an integer $l\neq0$ (mod $n_{T}$). We need the following high-level
conditions. Part (i) corresponds to Assumption \ref{Assumption Means for Periodogram},
part (ii) is satisfied if $\left\{ V_{t}\right\} $ is strong mixing
with mixing parameters of size $-2\nu/\left(\nu-1/2\right)$ for
some $\nu>1$ such that $\sup_{t\geq1}\mathbb{E}\left|V_{t}\right|^{4\nu}<\infty,$
while part (iii) requires additional smoothness.
\begin{assumption}
\label{Assumption Means for Local Periodogram}(i) For each $\omega_{l}$
and $u\in\left[0,\,1\right]$ with $T_{j}^{0}\in\mathbf{S}\left(u,\,0,\,n_{T}\right)$
there exist $B_{j}\in\mathbb{R}$, $j=1,\ldots,\,m_{0}$ with $B_{j_{1}}\neq B_{j_{2}}$
for $j_{1}\neq j_{2}$ such that
\begin{align*}
\left|\sum_{s=0}^{n_{T}-1}\mu\left(\left(\left\lfloor Tu\right\rfloor -n_{T}/2+s+1\right)/T\right)\exp\left(-i\omega_{l}s\right)\right|^{2}\geq\\
\quad\left|B_{j}\sum_{s=0}^{T_{j}^{0}-\left(\left\lfloor Tu\right\rfloor -n_{T}/2+1\right)}\exp\left(-i\omega_{l}s\right)+B_{j+1}\sum_{s=T_{j}^{0}-\left(\left\lfloor Tu\right\rfloor -n_{T}/2\right)}^{n_{T}-1}\exp\left(-i\omega_{l}s\right)\right|^{2}.
\end{align*}
(ii) $\left|\Gamma\left(u,\,k\right)\right|=C_{u,k}k^{-m}$ for all
$u\in\left[0,\,1\right]$ and all $k\geq C_{3}T^{\kappa}$ for some
$C_{3}<\infty$ , $C_{u,k}<\infty$ (which depends on $u$ and $k$),
$0<\kappa<1/2$, and $m>2$. (iii) $\sup_{u\in\left[0,\,1\right],\,u\neq\lambda_{0}^{j},\,j=1,\ldots,\,m_{0}}\left(\partial^{2}/\partial u^{2}\right)f\left(u,\,\omega\right)$
is continuous in $\omega.$
\end{assumption}
\begin{thm}
\label{Theorem Local Periodogram Long Memory Effects}Assume that
$\left\{ V_{t,T}\right\} $ satisfies Definition \ref{Definition Segmented-Locally-Stationary}
and that $n_{T}\rightarrow\infty$ with $n_{T}/T\rightarrow0$. Under
Assumption \ref{Assumption Smothness of A (for HAC)}-\ref{Assumption A - Dependence},
and \ref{Assumption Means for Local Periodogram},
(i) for any $u\in\left(0,\,1\right)$ such that $T_{j}^{0}\notin\mathbf{S}\left(u,\,0,\,n_{T}\right)$
for all $j=1,\ldots,\,m_{0}$, $\mathbb{E}\left(I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)\right)\geq f\left(u,\,\omega_{l}\right)$
as $\omega_{l}\rightarrow0$;
(ii) for any $u\in\left(0,\,1\right)$ such that $T_{j}^{0}\in\mathbf{S}\left(u,\,0,\,n_{T}\right)$
for some $j=1,\ldots,\,m_{0}$ we have two sub-cases: (a) if $n_{j,L}\left(u,\,0,\,n_{T}\right)/n_{T}\rightarrow\gamma$
or $n_{j,R}\left(u,\,0,\,n_{T}\right)/n_{T}\rightarrow\gamma$ with
$\gamma\in\left(0,\,1\right),$ and $n_{T}\omega_{l}^{2}\rightarrow0$
as $T\rightarrow\infty$, then $\mathbb{E}\left(I_{\mathrm{L},T}\left(u,\,\omega\right)\right)\rightarrow\infty$
for many values in the sequence $\left\{ \omega_{l}\right\} $ as
$\omega_{l}\rightarrow0$; (b) if $n_{j,L}\left(u,\,0,\,n_{T}\right)/n_{T}\rightarrow0$
or $n_{j,R}\left(u,\,0,\,n_{T}\right)/n_{T}\rightarrow0$, then $\mathbb{E}\left(I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)\right)\geq f\left(u,\,\omega_{l}\right)$
as $\omega_{l}\rightarrow0$.
\end{thm}
It is useful to compare Theorem \ref{Theorem Local Periodogram Long Memory Effects}
with the discussion above about the periodogram (see also Theorem
\ref{Theorem Periodogram Long Memory Effects}). Unlike the periodogram,
the asymptotic behavior of the local periodogram as $\omega_{l}\rightarrow0$
depends on the vicinity of $u$ to $\lambda_{j}^{0}$ $\left(j=1,\ldots,\,m_{0}\right)$.
Since $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$ uses observations
in the window $\mathbf{S}\left(u,\,0,\,n_{T}\right)$, if no discontinuity
in the mean occurs in this window then $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$
is asymptotically unbiased for the spectral density $f\left(u,\,\omega_{l}\right)$.
More complex is its behavior if some $T_{j}^{0}$ falls in $\mathbf{S}\left(u,\,0,\,n_{T}\right)$.
The theorem shows that if $T_{j}^{0}$ is close to the boundary, as
indicated in case (ii-b), then $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$
is bounded below by $f\left(u,\,\omega_{l}\right)$, similarly to
case (i). If instead $T_{j}^{0}$ falls sufficiently close to the
mid-point $\left\lfloor Tu\right\rfloor ,$ as indicated in case (ii-a),
then $\mathbb{E}\left(I_{\mathrm{L},T}\left(u,\,\omega\right)\right)\rightarrow\infty$
for many values in the sequence $\left\{ \omega_{l}\right\} $ as
$\omega_{l}\rightarrow0$ provided it satisfies $n_{T}\omega_{l}^{2}\rightarrow0$
as $T\rightarrow\infty$. Hence, unless $T\lambda_{j}^{0}$ is close
to $\left\lfloor Tu\right\rfloor ,$ the local periodogram $I_{\mathrm{L},T}\left(u,\,\omega_{l}\right)$
behaves very differently from the periodogram $I_{T}\left(\omega_{l}\right)$.
Accordingly, nonstationarity is unlikely to generate long memory effects
if one uses the local periodogram. As for $\widehat{c}_{T}\left(u,\,k\right)$,
if one uses preliminary inference procedures {[}cf. \citet{casini:change-point-spectra}{]}
for the detection and estimation of the discontinuities in the spectrum
and for the estimation of their locations, then one can construct
the window efficiently and avoid $T_{j}^{0}$ being too close to $\left\lfloor Tu\right\rfloor .$
\section{\label{Section Edgeworth-Expansions-for}Edgeworth Expansions for
HAR Tests Under Nonstationarity}
We now consider Edgeworth expansions for the distribution of the
$t$-statistic in the location model based on the HAC and DK-HAC estimator
where $\left\{ V_{t}\right\} $ is assumed to have zero-mean and time-varying
second moments. This is useful for analyzing the theoretical properties
of the null rejection probabilities of the HAR tests under nonstationarity.
As in the literature, we make use of the Gaussianity assumption for
mathematical convenience.\footnote{This can be relaxed by considering distributions with Gram-Charlier
representations at the expense of more complex derivations. } We relax the stationarity assumption used in the literature {[}cf.
\citet{jansson:04}, \citet{sun/phillips/jin:08} and \citet{velasco/robinson:01}{]}
which has important consequences for the nature of the results. The
results concerning the $t$-test based on the HAC estimator are presented
in Section \ref{Subsection: HAC-based-HAR-Tests} while those based
on the DK-HAC estimator are presented in Section \ref{Subsection: DK-HAC-based-HAR-Tests}.
Let $\left\{ V_{t}\right\} $ be a zero-mean Gaussian SLS process
satisfying Assumption \ref{Assumption Smothness of A (for HAC)}-(i-iv).
Let
\begin{align}
h_{1} & \triangleq\frac{\sqrt{T}\,\overline{V}}{\sqrt{J_{T}}}\sim\mathscr{N}\left(0,\,1\right),\label{Eq. (h1)}
\end{align}
which is valid for all $T$ such that $J_{T}>0$ where $J_{T}=T^{-1}\sum_{s=1}^{T}\sum_{t=1}^{T}\mathbb{E}(V_{s}V_{t})$.
\subsection{\label{Subsection: HAC-based-HAR-Tests}HAC-based HAR Tests}
The classical HAC estimator is defined as
\begin{align*}
\widehat{J}_{\mathrm{HAC,}T}\triangleq\sum_{k=-T+1}^{T-1}K_{1}\left(b_{1,T}k\right)\widehat{\Gamma}\left(k\right), & \qquad\widehat{\Gamma}\left(k\right)=T^{-1}\sum_{t=|k|+1}^{T}V_{t}V_{t-|k|},
\end{align*}
where $K_{1}\left(\cdot\right)$ is a kernel and $b_{1,T}$ a bandwidth
parameter. Under appropriate conditions on $b_{1,T},$ we have $\widehat{J}_{\mathrm{\mathrm{HAC},}T}-J_{T}\overset{\mathbb{P}}{\rightarrow}0$
from which it follows that
\begin{align*}
Z_{T} & \triangleq\frac{\sqrt{T}\,\overline{V}}{\sqrt{\widehat{J}_{\mathrm{HAC,}T}}}\overset{d}{\rightarrow}\mathscr{N}\left(0,\,1\right).
\end{align*}
Let $\mathbf{V}=(V_{1},\ldots,\,V_{T})'$. Note that $\widehat{J}_{\mathrm{HAC,}T}=\mathbf{V}'W_{b_{1}}\mathbf{V}/T$
where $W_{b_{1}}$ has $\left(r,\,s\right)$th element
\begin{align}
W_{b_{1}}^{(r,s)} & =w\left(b_{1,T}(r-s)\right)=\int_{\Pi}\widetilde{K}_{b_{1}}\left(\omega\right)e^{i\left(r-s\right)\omega}d\omega,\label{Eq. Definition W_b1}
\end{align}
such that $\widetilde{K}_{b_{1}}\left(\omega\right)$ is a kernel
with smoothing number $b_{1,T}^{-1}$ and $\Pi=(-\pi,\,\pi]$. For
an even function $K$ that integrates to one, we define
\begin{align*}
\widetilde{K}_{b_{1}}\left(\omega\right) & =b_{1,T}^{-1}\sum_{j=-\infty}^{\infty}K\left(b_{1,T}^{-1}(\omega+2\pi j)\right).
\end{align*}
Note that $\widetilde{K}_{b_{1}}\left(\omega\right)$ is periodic
of period $2\pi$, even and satisfies $\smallint_{-\pi}^{\pi}\widetilde{K}_{b_{1}}\left(\omega\right)d\omega=1$.
It follows that $w\left(r\right)=\int_{-\infty}^{\infty}e^{irx}K\left(x\right)dx$
and $\widehat{J}_{\mathrm{HAC,}T}=2\pi\int_{\Pi}\widetilde{K}_{b_{1}}\left(\omega\right)I_{T}\left(\omega\right)d\omega$.
$\widetilde{K}_{b_{1}}\left(\omega\right)$ is the so-called spectral
window generator. We refer to \citet{brillinger:75} for a review
of these introductory concepts.
We now analyze the joint distribution of $\overline{V}$ and $\widehat{J}_{\mathrm{HAC,}T}$.
Let $\mathsf{B}_{T}=\mathbb{E}(\widehat{J}_{\mathrm{HAC,}T})/J_{T}-1$
and $\mathsf{V}_{T}^{2}=\mathrm{Var}(\sqrt{Tb_{1,T}}\widehat{J}_{\mathrm{HAC,}T}/J_{T})$
denote the relative bias and variance, respectively, of $\widehat{J}_{\mathrm{HAC,}T}$.
It is convenient to work with standardized statistics with zero mean
and unit variance. Write
\begin{align*}
Z_{T} & =Z_{T}\left(\mathbf{h}\right)=h_{1}\left(1+\mathrm{\mathsf{B}}_{T}+\mathsf{V}_{T}h_{2}\left(Tb_{1,T}\right)^{-1/2}\right)^{-1/2},\qquad h_{2}=\sqrt{Tb_{1,T}}\left(\frac{\widehat{J}_{\mathrm{HAC,}T}-\mathbb{E}\left(\widehat{J}_{\mathrm{HAC,}T}\right)}{J_{T}\mathsf{V}_{T}}\right),
\end{align*}
where $\mathbf{h}=\left(h_{1},\,h_{2}\right)'$. Note that $h_{2}=\mathbf{V}'Q_{T}\mathbf{V}-$$\mathbb{E}\left(\mathbf{V}'Q_{T}\mathbf{V}\right)$
is a centered quadratic form in a Gaussian vector where $Q_{T}=W_{b_{1}}(\sqrt{T/b_{1,T}}\mathsf{V}_{T}J_{T})^{-1}$.
The joint characteristic function of $\mathbf{h}$ is
\begin{align*}
\psi_{T}\left(\mathbf{t}\right)=\psi_{T}\left(t_{1},\,t_{2}\right) & =\left|I-2it_{2}\Sigma_{V}Q_{T}\right|^{-1/2}\exp\left(-2^{-1}t_{1}^{2}\xi'_{T}\left(I-2it_{2}\Sigma_{V}Q_{T}\right)^{-1}\Sigma_{V}\xi_{T}-it_{2}\Upsilon_{T}\right),
\end{align*}
where $\Upsilon_{T}=\mathbb{E}\left(\mathbf{V}'Q_{T}\mathbf{V}\right)=\mathrm{Tr}\left(\Sigma_{V}Q_{T}\right),$
$\Sigma_{V}=\mathbb{E}\left(\mathbf{VV}'\right)$, and $\xi_{T}=\mathbf{1}/\sqrt{TJ_{T}}$
with $\mathbf{1}$ being the $T\times1$ vector $\left(1,1,\ldots,\,1\right)'$.
The cumulant generating function of $\mathbf{h}$ is
\begin{align*}
\mathrm{K}_{T}\left(t_{1},\,t_{2}\right) & =\log\psi_{T}\left(t_{1},\,t_{2}\right)=\sum_{r=0}^{\infty}\sum_{s=0}^{\infty}\kappa_{T}\left(r,\,s\right)\frac{\left(it_{1}\right)^{r}}{r!}\frac{\left(it_{2}\right)^{r}}{s!},
\end{align*}
where $\kappa_{T}\left(r,\,s\right)$ is the cumulant of $\mathbf{h}.$
\citet{phillips:1980} considered the distribution of linear and quadratic
forms under Gaussianity. From his derivations, the nonzero bivariate
cumulants are
\begin{align*}
\kappa_{T}\left(0,\,s\right) & =2^{s-1}\left(s-1\right)!\mathrm{Tr}\left(\left(\Sigma_{V}Q_{T}\right)^{s}\right),\qquad s>1,\\
\kappa_{T}\left(2,\,s\right) & =2^{s}s!\xi_{T}'\left(\Sigma_{V}Q_{T}\right)^{s}\Sigma_{V}\xi_{T},\qquad\qquad s>0.
\end{align*}
We introduce the following assumptions about $\left\{ V_{t}\right\} $
and $f\left(u,\,0\right)$.
\begin{assumption}
\label{Assumption 1 in VR}For all $u\in\left[0,\,1\right]$, $0<f\left(u,\,0\right)<\infty$
and $f\left(u,\,\omega\right)$ has $d_{f}$ continuous derivatives
$\left(d_{f}\geq2\right)$ $f^{\left(d_{f}\right)}\left(u,\,\omega\right)$
in a neighborhood of $\omega=0$ and the $d_{f}$th derivative satisfies
a Lipschitz condition of order $\varrho$ with $\varrho\in(0,\,1]$.
\end{assumption}
\begin{assumption}
\label{Assumption: Assumption 2 in VR }For all $u$, $f\left(u,\,\omega\right)\in L_{p}$
for some $p>1$, i.e., $\left\Vert f\left(u,\cdot\right)\right\Vert _{p}^{p}=\int_{\Pi}f^{p}\left(u,\,\omega\right)d\omega<\infty.$
\end{assumption}
\begin{assumption}
\label{Assumption 3 VR}$|K\left(x\right)|<\infty$, $K\left(x\right)=K\left(-x\right)$,
$K\left(x\right)=0$ for $x\notin\Pi$ and $\int_{\Pi}K\left(x\right)dx=1$.
\end{assumption}
\begin{assumption}
\label{Assumption 4 in VR}$K\left(x\right)$ satisfies a uniform
Lipschitz condition of order 1 in $\left[-\pi,\,\pi\right]$.
\end{assumption}
\begin{assumption}
\label{Assumption 5 VR}For $j=0,\,1,\ldots,\,d_{f}$, $d_{f}\geq2$
and $r=1,\,2,\ldots$
\begin{align*}
\mu_{j}\left(K^{r}\right) & \triangleq\int_{\Pi}x^{j}\left(K\left(x\right)\right)^{r}dx=\begin{cases}
=0, & j<d_{f},\,r=1;\\
\neq0, & j=d_{f},\,r=1.
\end{cases}
\end{align*}
\end{assumption}
\begin{assumption}
\label{Assumption 6 VR} $b_{1,T}+(Tb_{1,T})^{-1}\rightarrow0$ as
$T\rightarrow\infty$.
\end{assumption}
\begin{assumption}
\label{Assumption 7 VR}$b_{1,T}=CT^{-q}$ where $0<q<1$ and $0<C<\infty$.
\end{assumption}
Assumptions \ref{Assumption 3 VR}-\ref{Assumption 7 VR} about the
kernel and bandwidth are the same as in \citet{velasco/robinson:01}
in which a discussion can be found. They are satisfied by most kernels
used in practice. The bandwidth condition in Assumption \ref{Assumption 6 VR}
is sufficient for the consistency of $\widehat{J}_{\mathrm{HAC,}T}$
and is strengthened in Assumption \ref{Assumption 7 VR}, for some
parts of the proofs, which is satisfied by popular MSE-optimal bandwidths
{[}cf. \citet{andrews:91}, \citet{casini_comment_andrews91}, \textcolor{MyBlue}{Belotti et al.}
\citeyearpar{belotti/casini/catania/grassi/perron_HAC_Sim_Bandws}
and \citet{whilelm:2015}{]}.
Assumptions \ref{Assumption 1 in VR}-\ref{Assumption: Assumption 2 in VR }
impose conditions on the smoothness and boundedness of the spectral
density. Assumption \ref{Assumption 1 in VR} is implied by $\sum_{k=-\infty}^{\infty}\left|k\right|^{d_{f}+\varrho}\sup_{t}|\mathbb{E}V_{t}V_{t-k}|<\infty$
but it is stronger than necessary because it extends the smoothness
restriction to all frequencies. Assumption \ref{Assumption: Assumption 2 in VR }
does impose some restrictions on $f\left(u,\,\cdot\right)$ beyond
the origin, though it is not particularly restrictive since any $p>1$
arbitrarily close to 1 will suffice.
We now analyze the asymptotic distribution of $\widehat{J}_{\mathrm{HAC,}T}$.
Under stationarity this was discussed by \citet{bentkus/rudzkis:1982}
and \citet{velasco/robinson:01}. From Lemmas \ref{Lemma 4.1 in VR}-\ref{Lemma 2 in VR }
in the supplement we obtain
\begin{align}
\mathsf{B}_{T}=\overline{c}_{1}b_{1,T}^{d_{f}}+O\left(b_{1,T}^{d_{f}+\varrho}+T^{-1}\log T\right), & \qquad\mathrm{where}\qquad\overline{c}_{1}=\frac{\mu_{d_{f}}\left(K\right)\int_{0}^{1}f^{\left(d_{f}\right)}\left(u,\,0\right)du}{d_{f}!\int_{0}^{1}f\left(u,\,0\right)du}.\label{Eq. (c1)}
\end{align}
The order of the asymptotic bias $b_{1,T}^{d_{f}}$ depends on the
smoothness of the spectral density at $\omega=0$ {[}cf. Assumption
\ref{Assumption 1 in VR}{]}. The constant $\overline{c}_{1}$ depends
on the moment of order $d_{f}$ of the kernel $K$ and on the smoothness
of $f\left(u,\,\omega\right)$ at $\omega=0$. For example, for the
time-varying AR(1) in \eqref{Eq. Example TV AR(1)},
\begin{align}
f^{\left(2\right)}\left(u,\,0\right) & =-\frac{\sigma^{2}\left(u\right)\rho\left(u\right)}{\pi\left(1+\rho\left(u\right)^{2}-2\rho\left(u\right)\right)^{2}}.\label{Eq.(f(2) (u,0) TV AR(1)}
\end{align}
If there is positive dependence at time $u$, then $\rho\left(u\right)>0$
and $f^{\left(2\right)}\left(u,\,0\right)<0$. Suppose $K\left(x\right)\geq0$
for all $x$ so that $\mu_{2}\left(K\right)>0$. Then the sign of
the bias is determined by the sign of $\int_{0}^{1}f^{\left(2\right)}\left(u,\,0\right)du$.
A positive local AR(1) coefficient contributes negative bias which
corresponds to the well-known downward bias of the LRV estimator when
there is positive dependence. Conversely, with anti-persistence $\rho\left(u\right)<0$
and $f^{\left(2\right)}\left(u,\,0\right)>0$. Since $\rho\left(\cdot\right)$
is time-varying, whether the bias is positive or negative depends
on the path of $\rho\left(\cdot\right)$. The smoother the spectral
density is at frequency zero, the smoother the kernel and the slower
$b_{1,T}$ can be. The factor $\int_{0}^{1}f\left(u,\,0\right)du$
in the denominator follows by definition because $\mathsf{B}_{T}$
is the relative bias.
We present a second-order Edgeworth expansion to approximate the
distribution of $\mathbf{h}$, with error $o((Tb_{1,T})^{-1/2})$
and including terms up to order $(Tb_{1,T})^{-1/2}$ to correct the
asymptotic normal distribution. This will imply the validity of that
expansion for the distribution of $\widehat{J}_{\mathrm{HAC},T}$.
For $\mathbf{B}\in\mathscr{B}^{2}$, where $\mathscr{B}^{2}$ is any
class of Borel sets in $\mathbb{R}^{2}$, let $\mathbb{Q}_{T}^{\left(2\right)}\left(\mathbf{B}\right)=\int_{\mathbf{B}}\varphi_{2}\left(\mathbf{h}\right)q_{T}^{\left(2\right)}\left(\mathbf{h}\right)d\mathbf{h},$
where $\varphi_{2}\left(\mathbf{h}\right)=\left(2\pi\right)^{-1}\exp\{-\left(1/2\right)\left\Vert \mathbf{h}\right\Vert ^{2}\}$
is the density of the bivariate standard normal distribution,
\begin{align*}
q_{T}^{\left(2\right)}\left(\mathbf{h}\right) & =1+(1/3!)\left(Tb_{1,T}\right)^{-1/2}\left(\Xi_{0}(0,\,3)\mathcal{H}_{3}\left(h_{2}\right)+\Xi_{0}(2,\,1)\mathcal{H}_{2}\left(h_{1}\right)\mathcal{H}_{1}\left(h_{2}\right)\right),
\end{align*}
where $\mathcal{H}_{j}\left(\cdot\right)$ are the univariate Hermite
polynomials of order $j$, and $\Xi_{0}\left(0,\,3\right)=\left(4\pi\right)^{1/2}2!\int_{\Pi}K^{3}\left(\omega\right)$
$d\omega\left\Vert K\right\Vert _{2}^{-3}$ and $\Xi_{0}(2,\,1)=\left(4\pi\right)^{1/2}K\left(0\right)\left\Vert K\right\Vert _{2}^{-1}$
(see Lemmas \ref{Lemma 3 in VR}-\ref{Lemma: Lemma 4 in VR}). Let
$\left(\partial\mathbf{B}\right)^{\phi}$denote a neighborhood of
radius $\phi$ of the boundary of a set $\mathrm{\mathbf{B}}.$ Let
$\mathbb{P}_{T}$ denote the probability measure of $\mathbf{h}.$
\begin{thm}
\label{Theorem: Theorem 1 in VR} Let Assumptions \ref{Assumption 1 in VR},
\ref{Assumption: Assumption 2 in VR } $\left(p>1\right)$, \ref{Assumption 3 VR}-\ref{Assumption 4 in VR}
and \ref{Assumption 7 VR} $\left(0<q<1\right)$ hold. For $\phi_{T}=(Tb_{1,T})^{-\varpi}$
with $1/2<\varpi<1$, we have
\begin{align}
\sup_{\mathbf{B}\in\mathscr{B}^{2}}\left|\mathbb{P}_{T}\left(\mathbf{B}\right)-\mathbb{Q}_{T}^{\left(2\right)}\left(\mathbf{B}\right)\right| & =o\left(\left(Tb_{1,T}\right)^{-1/2}\right)+(4/3)\sup_{\mathbf{B}\in\mathscr{B}^{2}}\mathbb{Q}_{T}^{\left(2\right)}\left(\left(\partial\mathbf{B}\right)^{2\phi_{T}}\right).\label{Eq. in Th. 1 in VR}
\end{align}
\end{thm}
Theorem \ref{Theorem: Theorem 1 in VR} shows that $\mathbb{Q}_{T}^{\left(2\right)}$
is a valid second-order Edgeworth expansion for the measure $\mathbb{P}_{T}$.
The method of proof is the same as in \citet{velasco/robinson:01}.
We first approximate the true characteristic function and then apply
a smoothing lemma {[}cf. Lemma \ref{Lemma Bhattacharya and Rao 1975}
in the supplement which is from \citet{bhattacharya/rao:1975}{]}.
The leading term of the approximation error is of order $o((Tb_{1,T})^{-1/2})$
as the second term on the right hand side of \eqref{Eq. in Th. 1 in VR}
is negligible if $\mathbf{B}$ is convex because $\phi_{T}$ decreases
as a power of $T$. This is the same order obtained for the corresponding
leading term under stationarity. Since the higher-order correction
terms in $q_{T}^{\left(2\right)}$ depend only on $K\left(\cdot\right)$
but not on $f\left(\cdot,\,\cdot\right)$, they are equal to the one
obtained under stationarity.
Next, we focus on $Z_{T}$, i.e., a $t$-statistic for the mean.
Proceeding as in \citet{velasco/robinson:01}, we first derive a
linear stochastic approximation to $Z_{T}\left(\mathbf{h}\right)$
and show that its distribution is the same as that of $Z_{T}$ up
to order $o((Tb_{1,T}){}^{-1/2})$. Then, we show that the asymptotic
approximation for the distribution of the linear stochastic approximation
is valid also for $Z_{T}$ with the same error $o((Tb_{1,T}){}^{-1/2})$.
Using Lemmas \ref{Lemma 3 in VR}-\ref{Lemma: Lemma 4 in VR} in the
supplement we can substitute out $\mathsf{B}_{T}$ and $\mathsf{V}_{T}$
in $Z_{T}$ and, by only focusing on the leading terms, we define
the following linear stochastic approximation,
\begin{align*}
\widetilde{Z}_{T} & \triangleq h_{1}\left(1-2^{-1}\overline{c}_{1}b_{1,T}^{d_{f}}-2^{-1}\sqrt{4\pi}\left\Vert K_{2}\right\Vert h_{2}\left(Tb_{1,T}\right)^{-1/2}\right).
\end{align*}
The next theorem presents a valid Edgeworth expansion for the distribution
of $\widetilde{Z}_{T}$ from that of $\mathbf{h}.$
\begin{thm}
\label{Theorem: Theorem 2 in VR}Let Assumptions \ref{Assumption 1 in VR},
\ref{Assumption: Assumption 2 in VR } $\left(p>1\right)$, \ref{Assumption 3 VR}-\ref{Assumption 5 VR}
and \ref{Assumption 7 VR} $(q=1/\left(1+2d_{f}\right))$ hold. For
a convex Borel set $\mathbf{C}$, we have, for $r_{2}\left(x\right)=-\overline{c}_{1}\left(x^{2}-1\right)/2$,
\begin{align}
\sup_{\mathbf{C}}\left|\mathbb{P}\left(Z_{T}\in\mathbf{C}\right)-\int_{\mathbf{C}}\varphi\left(x\right)\left(1+r_{2}\left(x\right)b_{1,T}^{d_{f}}\right)dx\right| & =o\left(\left(Tb_{1,T}\right)^{-1/2}\right).\label{Eq. (3) in VR}
\end{align}
\end{thm}
Theorem \ref{Theorem: Theorem 2 in VR} shows the form of the correction
term to the standard normal distribution, i.e., $b_{1,T}^{d_{f}}\int_{\mathbf{C}}\varphi\left(x\right)r_{2}\left(x\right)dx$.
The error of the approximation is of order $o((Tb_{1,T})^{-1/2})$
which is the same as the one obtained under stationarity by \citet{velasco/robinson:01}.
Let $\Phi\left(\cdot\right)$ denote the distribution function of
the standard normal. Setting $\mathbf{C}=(-\infty,\,z]$, integrating
and Taylor expanding $\Phi\left(\cdot\right)$, we obtain, uniformly
in $z$,
\begin{align}
\mathbb{P}\left(Z_{T}\leq z\right) & =\Phi\left(z\right)+\frac{1}{2}\overline{c}_{1}z\varphi\left(z\right)b_{1,T}^{d_{f}}+o\left(\left(Tb_{1,T}\right)^{-1/2}\right)\label{Eq. (4) in VR}\\
& =\Phi\left(z\left(1+\frac{1}{2}\overline{c}_{1}b_{1,T}^{d_{f}}\right)\right)+o\left(\left(Tb_{1,T}\right)^{-1/2}\right)=\Phi\left(z\right)+O\left(\left(Tb_{1,T}\right)^{-1/2}\right).\nonumber
\end{align}
This shows that under the conditions of Theorem \ref{Theorem: Theorem 2 in VR},
the standard normal approximation is correct up to order $O((Tb_{1,T})^{-1/2})$.
Eq. \eqref{Eq. (4) in VR} has an immediate interpretation. Consider
the time-varying AR(1) example in \eqref{Eq. Example TV AR(1)} and
suppose $K\left(x\right)\geq0$ for all $x$ so that $\mu_{2}\left(K\right)\geq0$.
Given \eqref{Eq.(f(2) (u,0) TV AR(1)} we know that with local positive
persistence (i.e., $\rho\left(u\right)>0$) $f\left(u,\,\omega\right)$
has a peak at $\omega=0$. If the pattern of $\rho\left(u\right)$
is such that $\int_{0}^{1}f^{\left(2\right)}\left(u,\,0\right)du<0$
so that the positive persistence dominates, then $\overline{c}_{1}<0$
and as is well-known the HAC estimator underestimates the true LRV
and the corresponding HAC-based test over-rejects. The approximation
in \eqref{Eq. (4) in VR} tends to correct this problem as it follows
that one uses $\Phi\left(z\left(1+\gamma_{T}\right)\right)$ where
$\gamma_{T}\leq0$, so for a given significance level the critical
value $z$ is larger in absolute value than the corresponding standard
normal critical value. Conversely, if there is anti-persistence, then
$\overline{c}_{1}>0$ and the implied critical value is smaller than
the corresponding standard normal critical value. For $d_{f}>2$ the
reasoning is the same but one has to take into account the sign of
$\mu_{d_{f}}\left(K\right)$.
Consider the location model $y_{t}=\beta+V_{t}$ $\left(t=1,\ldots,\,T\right).$
For the null hypothesis $\mathbb{H}_{0}:\,\beta=\beta_{0}$, consider
the following $t$-test,
\begin{align*}
t_{\mathrm{HAC}} & =\frac{\sqrt{T}\left(\widehat{\beta}-\beta_{0}\right)}{\sqrt{\widehat{J}_{\mathrm{HAC},T}}},
\end{align*}
where $\widehat{\beta}$ is the least-squares estimator of $\beta$.
Theorem \ref{Theorem: Theorem 2 in VR} and \eqref{Eq. (4) in VR}
imply that
\begin{align}
\mathbb{P}\left(t_{\mathrm{HAC}}\leq z\right) & =\Phi\left(z\right)+p\left(z\right)\left(Tb_{1,T}\right)^{-1/2}+o\left(\left(Tb_{1,T}\right)^{-1/2}\right),\label{Eq. ERP t-hac}
\end{align}
for any $z\in\mathbb{R},$ where $p\left(z\right)$ is an odd function.
When $q=1/\left(1+2d_{f}\right)$ we have $p\left(z\right)=2^{-1}\overline{c}_{1}z\varphi\left(z\right)C^{d_{f}+1/2}$
where $C$ is defined in Assumption \ref{Assumption 7 VR}. Thus,
the error in rejection probability (ERP) of $t_{\mathrm{HAC}}$ is
of order $O((Tb_{1,T})^{-1/2})$. If $\left\{ V_{t}\right\} $ is
second-order stationary, the results in \citet{velasco/robinson:01}
imply that the ERP of $t_{\mathrm{HAC}}$ is also of order $O((Tb_{1,T})^{-1/2})$.
Below we establish the corresponding ERP when the $t$-statistic is
instead normalized by $\widehat{J}_{\mathrm{DK},T}$ and also discuss
the ERP of the $t$-test under fixed-$b$ asymptotics.
\subsection{\label{Subsection: DK-HAC-based-HAR-Tests}DK-HAC-based HAR Tests }
We now consider the Edgeworth expansion for tests based on the DK-HAC
estimator. In order to simplify some parts of the proof here we
consider an asymptotically equivalent version of the DK-HAC estimator
discussed in Section \ref{Section Consequences for HAR}. Let
\begin{align*}
\widehat{J}_{\mathrm{DK},T}^{*}=\sum_{k=-T+1}^{T-1}K_{1}\left(b_{1,T}k\right)\widehat{\Gamma}_{\mathrm{DK}}^{*}\left(k\right), & \qquad\widehat{\Gamma}_{\mathrm{DK}}^{*}\left(k\right)\triangleq\int_{0}^{1}\widehat{c}_{\mathrm{DK,}T}\left(r,\,k\right)dr,
\end{align*}
where $b_{1,T}$ is a bandwidth sequence and
\begin{align*}
\widehat{c}_{\mathrm{DK,}T}\left(r,\,k\right) & =\left(Tb_{2,T}\right)^{-1}\sum_{s=|k|+1}^{T}K_{2}\left(\frac{\left(Tr-\left(s-|k|/2\right)\right)/T}{b_{2,T}}\right)V_{s}V{}_{s-|k|},
\end{align*}
with $K_{2}$ a kernel and $b_{2,T}$ a bandwidth. Note that $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$
and $\widehat{\Gamma}_{\mathrm{DK}}^{*}\left(k\right)$ are asymptotically
equivalent and $\widehat{c}_{T}$ is a special case of $\widehat{c}_{\mathrm{DK,}T}$
with $K_{2}$ being a rectangular kernel and $n_{2,T}=Tb_{2,T}$.
\begin{assumption}
\label{Assumption K2 and b2}$K_{2}\left(\cdot\right):\,\mathbb{R}\rightarrow\left[0,\,\infty\right]$,
$K_{2}\left(x\right)=K_{2}\left(1-x\right)$, ${\textstyle \int_{0}^{1}}K_{2}\left(x\right)dx=1$,
$K_{2}\left(x\right)=0$ for $x\notin\left[0,\,1\right]$ and $K_{2}\left(\cdot\right)$
is continuous. The bandwidth sequence $\{b_{2,T}\}$ satisfies $b_{2,T}\rightarrow0$,
$b_{2,T}^{2}/b_{1,T}^{q_{2}}\rightarrow\overline{b}\in[0,\,\infty)$
and $1/Tb_{1,T}b_{2,T}\rightarrow0$ where $q_{2}$ is the index of
smoothness of $K_{1}\left(\cdot\right)$ at 0.
\end{assumption}
Under Assumptions \ref{Assumption 3 VR}-\ref{Assumption 4 in VR},
\ref{Assumption 6 VR} and \ref{Assumption K2 and b2} it holds that
$\widehat{J}_{\mathrm{DK},T}^{*}-J_{T}\overset{\mathbb{P}}{\rightarrow}0$
{[}cf. \citet{casini_hac}{]} and
\begin{align}
U_{T} & \triangleq\frac{\sqrt{T}\,\overline{V}}{\sqrt{\widehat{J}_{\mathrm{DK},T}^{*}}}\overset{d}{\rightarrow}\mathscr{N}\left(0,\,1\right).\label{Eq. U_T}
\end{align}
Note that $\widehat{J}_{\mathrm{DK},T}^{*}=\int_{0}^{1}\mathbf{\widetilde{V}}\left(r\right)'W_{b_{1}}\mathbf{\widetilde{V}}\left(r\right)dr/(Tb_{2,T})$
where $\mathbf{\widetilde{V}}\left(r\right)=(\widetilde{V}_{1}\left(r\right),\,\widetilde{V}_{2}\left(r\right),\ldots,\,\widetilde{V}_{T}\left(r\right))'$
with $\widetilde{V}_{j}\left(r\right)=\sqrt{K_{2}\left(\left(r-j\right)/Tb_{2,T}\right)}V_{j}$
and $W_{b_{1}}$ defined in \eqref{Eq. Definition W_b1}. Let
\begin{align*}
\widetilde{I}_{T}\left(r,\,\omega\right) & =\frac{1}{2\pi Tb_{2,T}}\left|\sum_{t=1}^{T}\exp\left(-i\omega t\right)\widetilde{V}_{t}\left(r\right)\right|^{2}.
\end{align*}
$\widetilde{I}_{T}\left(r,\,\omega\right)$ is the local periodogram
of $\{\mathbf{\widetilde{V}}\left(r\right)\}$. Then, $\widehat{J}_{\mathrm{DK},T}^{*}=2\pi\int_{0}^{1}\int_{\Pi}\widetilde{K}_{b_{1}}\left(\omega\right)\widetilde{I}_{T}\left(r,\,\omega\right)d\omega dr$.
We begin by analyzing the joint distribution of $\overline{V}$ and
$\widehat{J}_{\mathrm{DK},T}^{*}$. Let $\mathsf{B}_{\mathrm{2},T}=\mathbb{E}(\widehat{J}_{\mathrm{DK},T}^{*})/J_{T}-1$
and $\mathsf{V}_{2,T}^{2}=\mathrm{Var}(\sqrt{Tb_{1,T}b_{2,T}}\widehat{J}_{\mathrm{DK},T}^{*}/J_{T})$
denote the relative bias and variance of $\widehat{J}_{\mathrm{DK},T}^{*}$,
respectively. It is convenient to work with standardized statistics
with zero mean and unit variance. Write
\begin{align*}
U_{T} & =U_{T}\left(\mathbf{v}\right)=v_{1}\left(1+\mathrm{\mathsf{B}}_{2,T}+\mathsf{V}_{2,T}v_{2}\left(Tb_{1,T}b_{2,T}\right)^{-1/2}\right)^{-1/2},\quad v_{2}=\sqrt{Tb_{1,T}b_{2,T}}\left(\frac{\widehat{J}_{\mathrm{DK},T}^{*}-\mathbb{E}\left(\widehat{J}_{\mathrm{DK},T}^{*}\right)}{J_{T}\mathsf{V}_{2,T}}\right),
\end{align*}
where $\mathbf{v}=\left(v_{1},\,v_{2}\right)'$ with $v_{1}=h_{1}.$
Note that $v_{2}=\int_{0}^{1}(\mathbf{\widetilde{V}}\left(r\right)'Q_{2,T}\mathbf{\widetilde{V}}\left(r\right)-$$\mathbb{E}(\mathbf{\widetilde{V}}\left(r\right)'Q_{2,T}\mathbf{\widetilde{V}}\left(r\right)))dr$
is a centered quadratic form in a Gaussian vector where $Q_{2,T}=W_{b_{1}}(\sqrt{Tb_{2,T}/b_{1,T}}\mathsf{V}_{2,T}J_{T})^{-1}$.
The joint characteristic function of $\mathbf{v}$ is
\begin{align*}
\psi_{2,T}\left(t_{1},\,t_{2}\right) & =\left|I-2it_{2}\Sigma_{\widetilde{V}}Q_{2,T}\right|^{-1/2}\exp\left\{ -2^{-1}t_{1}^{2}\xi'_{2,T}\left(I-2it_{2}\Sigma_{\widetilde{V}}Q_{2,T}\right)^{-1}\Sigma_{\widetilde{V}}\xi_{2,T}-it_{2}\Upsilon_{2,T}\right\} ,
\end{align*}
where $\Upsilon_{2,T}=\mathbb{E}(\int_{0}^{1}(\mathbf{\widetilde{V}}\left(r\right)'Q_{2,T}\mathbf{\widetilde{V}}\left(r\right))dr)=\mathrm{Tr}(\Sigma_{\widetilde{V}}Q_{2,T}),$
$\Sigma_{\widetilde{V}}=\mathbb{E}(\int_{0}^{1}(\mathbf{\widetilde{V}}\left(r\right)\mathbf{\widetilde{V}}\left(r\right)')dr)$
and $\xi_{2,T}=\mathbf{1}/\sqrt{Tb_{2,T}J_{T}}$. The cumulant generating
function of $\mathbf{v}$ is
\begin{align*}
\mathrm{K}_{2,T}\left(t_{1},\,t_{2}\right) & =\log\psi_{2,T}\left(t_{1},\,t_{2}\right)=\sum_{r=0}^{\infty}\sum_{s=0}^{\infty}\kappa_{2,T}\left(r,\,s\right)\frac{\left(it_{1}\right)^{r}}{r!}\frac{\left(it_{2}\right)^{r}}{s!},
\end{align*}
where $\kappa_{2,T}\left(r,\,s\right)$ is the cumulant of $\mathbf{v}.$
To obtain more precise bounds in some parts of the proofs we use
the following assumption on the cross-partial derivatives of $f\left(u,\,\omega\right)$.
Let $\widetilde{\mathbf{C}}$ denote the set of continuity points
of $f\left(u,\,\omega\right)$ in $u$, i.e., $\widetilde{\mathbf{C}}=\{\left[0,\,1\right]/\{\lambda_{j}^{0},\,j=1,\ldots,\,m_{0}\}\}$.
Define
\begin{align*}
\Delta_{f}\left(\omega\right) & =\sum_{j=1}^{m_{0}}\int_{0}^{1}\left(\frac{\partial}{\partial u_{-}}f\left(\lambda_{j}^{0},\,\omega\right)\int_{0}^{1-s}xK_{2}\left(x\right)dx+\frac{\partial}{\partial u_{+}}f\left(\lambda_{j}^{0},\,\omega\right)\int_{1-s}^{1}xK_{2}\left(x\right)dx\right)ds,
\end{align*}
where
\begin{align*}
\frac{\partial}{\partial u_{-}}f\left(\lambda_{j}^{0},\,\omega\right)=\underset{h\uparrow0}{\lim}\frac{f\left(\lambda_{j}^{0}+h,\,\omega\right)-f\left(\lambda_{j}^{0},\,\omega\right)}{h}, & \qquad\frac{\partial}{\partial u_{+}}f\left(\lambda_{j}^{0},\,\omega\right)=\underset{h\downarrow0}{\lim}\frac{f\left(\lambda_{j}^{0}+h,\,\omega\right)-f\left(\lambda_{j}^{0},\,\omega\right)}{h}.
\end{align*}
\begin{assumption}
\label{Assumption Lip of d2 f(u,w)}For $u\in\widetilde{\mathbf{C}},$
$\left(\partial^{2}/\partial u^{2}\right)f\left(u,\,\omega\right)$
has $d_{f}$ continuous derivatives in $\omega$ in a neighborhood
of $\omega=0,$ the $d_{f}$ derivative satisfying a Lipschitz condition
of order $\varrho_{2}\in(0,\,1]$. \\
For $u\notin\widetilde{\mathbf{C}},$ $\left(\partial/\partial u_{-}\right)f\left(u,\,\omega\right)$
and $\left(\partial/\partial u_{+}\right)f\left(u,\,\omega\right)$
have $d_{f}$ continuous derivatives in $\omega$ in a neighborhood
of $\omega=0,$ the $d_{f}$ derivative satisfying a Lipschitz condition
of order $\varrho_{2}\in(0,\,1]$.
\end{assumption}
From Lemmas \ref{Lemma 4.1 in VR} and \ref{Lemma 2 in VR DK-HAC},
the relative bias of $\widehat{J}_{\mathrm{DK},T}^{*}$ is
\begin{align*}
\mathsf{B}_{2,T}=\overline{c}_{1}b_{1,T}^{d_{f}}+\overline{c}_{2}b_{2,T}^{2}+O\left(b_{1,T}^{d_{f}+\varrho}+T^{-1}\log T+\left(Tb_{2,T}\right)^{-1}\right)+o\left(b_{2,T}^{2}\right) & ,
\end{align*}
where
\begin{align*}
\overline{c}_{1}=\frac{\mu_{d_{f}}\left(K\right)\int_{0}^{1}f^{\left(d_{f}\right)}\left(u,\,0\right)du}{d_{f}!\int_{0}^{1}f\left(u,\,0\right)du}, & \qquad\overline{c}_{2}=\frac{2^{-1}\int_{0}^{1}x^{2}K_{2}\left(x\right)dx\int_{\widetilde{\mathbf{C}}}\frac{\partial^{2}}{\partial u^{2}}f\left(u,\,0\right)du+\Delta_{f}\left(0\right)}{\int_{0}^{1}f\left(u,\,0\right)du}.
\end{align*}
The factor $\overline{c}_{1}$ in the relative bias $\mathsf{B}_{2,T}$
also enters $\mathsf{B}_{T}$ and we already discussed it. The second
factor, $\overline{c}_{2}$, includes two elements. The first depends
on the second moment of the kernel $K_{2}$ and on the smoothness
over time of the spectral density $f\left(u,\,0\right)$. The second
element in $\overline{c}_{2}$ is $\Delta_{f}\left(0\right)$ which
depends on the right and left first partial derivatives of $f\left(u,\,0\right)$
with respect to $u$ at the discontinuity points. The more nonstationary
is the data the more complex is $\overline{c}_{2}$, and in fact the
larger in magnitude are $\partial^{2}f\left(u,\,0\right)/\partial u^{2}$
and $\Delta_{f}\left(0\right)$. For the special case of stationary
data, $\overline{c}_{2}=0$. The more nonstationary is the data,
the smaller $b_{2,T}$ should be chosen so as to weight more the data
locally. The smoothing over sample autocovariances is needed to achieve
consistency while the time-smoothing is introduced to more flexibly
account for the time-varying properties of the data. The disadvantage
of the time-smoothing is that it reduces the effective sample size
thereby making accounting for strong dependence more difficult.
We now present a second-order Edgeworth expansion to approximate
the distribution of $\mathbf{v}$ with error $o((Tb_{1,T}b_{2,T})^{-1/2})$.
The expansion includes terms up to order $(Tb_{1,T}b_{2,T})^{-1/2}$
to correct the asymptotic normal distribution. This implies the validity
of that expansion for the distribution of $\widehat{J}_{\mathrm{DK},T}^{*}$.
For $\mathbf{B}\in\mathscr{B}^{2}$, let $\mathbb{Q}_{2,T}^{\left(2\right)}(\mathbf{B})=\int_{\mathbf{B}}\varphi_{2}\left(\mathbf{v}\right)q_{2,T}^{\left(2\right)}\left(\mathbf{v}\right)d\mathbf{v},$
where
\begin{align*}
q_{2,T}^{\left(2\right)}\left(\mathbf{v}\right) & =1+(1/3!)\left(Tb_{1,T}b_{2,T}\right)^{-1/2}\left\{ \Xi_{2,0}(0,\,3)\mathcal{H}_{2,3}\left(v_{2}\right)+\Xi_{2,0}(2,\,1)\mathcal{H}_{2,2}\left(v_{1}\right)\mathcal{H}_{2,1}\left(v_{1}\right)\right\} ,
\end{align*}
$\mathcal{H}_{2,j}\left(\cdot\right)$ are the univariate Hermite
polynomials of order $j$ and $\Xi_{2,0}(0,\,3)$ and $\Xi_{2,0}(2,\,1)$
are bounded and depend on $K,\,K_{2}$ and on $f\left(u,\,0\right)$
(see Lemmas \ref{Lemma: Proposition 1 in VR DK-HAC}-\ref{Lemma: Proposition 2 in VR DK-HAC}).
\begin{thm}
\label{Theorem: Theorem 1 in VR DK-HAC}Let Assumptions \ref{Assumption 1 in VR},
\ref{Assumption: Assumption 2 in VR } $\left(p>1\right)$, \ref{Assumption 3 VR}-\ref{Assumption 4 in VR},
\ref{Assumption 7 VR} $\left(0<q<1\right)$, \ref{Assumption K2 and b2}-\ref{Assumption Lip of d2 f(u,w)}
hold. For $\phi_{T}=(Tb_{1,T}b_{2,T})^{-\varpi}$ with $1/2<\varpi<1$,
and every class $\mathscr{B}^{2}$ of Borel sets in $\mathbb{R}^{2}$,
we have
\begin{align}
\sup_{\mathbf{B}\in\mathscr{B}^{2}}\left|\mathbb{P}_{T}\left(\mathbf{B}\right)-\mathbb{Q}_{2,T}^{\left(2\right)}\left(\mathbf{B}\right)\right| & =o\left(\left(Tb_{1,T}b_{2,T}\right)^{-1/2}\right)+(4/3)\sup_{\mathbf{B}\in\mathscr{B}^{2}}\mathbb{Q}_{2,T}^{\left(2\right)}\left(\left(\partial\mathbf{B}\right)^{2\phi_{T}}\right).\label{Eq. in Th. 1 in VR DK-HAC}
\end{align}
\end{thm}
Theorem \ref{Theorem: Theorem 1 in VR DK-HAC} shows that $\mathbb{Q}_{2,T}^{\left(2\right)}$
is a valid second-order Edgeworth expansion for the probability measure
$\mathbb{P}_{T}$ of $\mathbf{v}.$ The correction $q_{2,T}^{\left(2\right)}\left(\mathbf{v}\right)$
differs from $q_{T}^{\left(2\right)}\left(\mathbf{h}\right)$ in
Theorem \ref{Theorem: Theorem 1 in VR}. This difference depends on
the smoothing over time, i.e., on $b_{2,T}$ and $K_{2}\left(\cdot\right)$.
The theorem also suggests that the leading term of the error of the
approximation is of order $o((Tb_{1,T}b_{2,T})^{-1/2})$.
Next, we focus on $U_{T}$ defined in \eqref{Eq. U_T}, i.e., a
$t$-statistic based on $\widehat{J}_{\mathrm{DK},T}^{*}$, and
present the Edgeworth expansion. We need the following assumption,
replacing Assumptions \ref{Assumption 6 VR}-\ref{Assumption 7 VR},
that controls the rate of smoothing over lagged autocovariances and
time implied by the bandwidths $b_{1,T}$ and $b_{2,T}$, respectively.
It requires that the bias due to smoothing over frequency and over
time is of the same order as the correction term obtained in $\mathbb{Q}_{2,T}^{\left(2\right)}\left(\mathbf{B}\right)$
or as the standard deviation of $\widehat{J}_{\mathrm{DK},T}^{*}$.
The assumption is satisfied by, for example, the MSE-optimal DK-HAC
estimators proposed by \textcolor{MyBlue}{Belotti et al.} \citeyearpar{belotti/casini/catania/grassi/perron_HAC_Sim_Bandws}
and \citet{casini_hac}.
\begin{assumption}
\label{Assumption condition Repalce Assumption 7}The bandwidths $b_{1,T}\rightarrow0$
and $b_{2,T}\rightarrow0$ satisfy $0<b_{1,T}^{d_{f}}\left(Tb_{1,T}b_{2,T}\right)^{-1/2}<\infty$
and $0<b_{2,T}^{2}\left(Tb_{1,T}b_{2,T}\right)^{-1/2}<\infty$.
\end{assumption}
\begin{thm}
\label{Theorem: Theorem 2 in VR DK-HAC}Let Assumptions \ref{Assumption 1 in VR},
\ref{Assumption: Assumption 2 in VR } $\left(p>1\right)$, \ref{Assumption 3 VR}-\ref{Assumption 5 VR},
and \ref{Assumption K2 and b2}-\ref{Assumption condition Repalce Assumption 7}
hold. For convex Borel sets $\mathbf{C}$, we have, for $r_{2}\left(x\right)=-\overline{c}_{1}\left(x^{2}-1\right)/2$
and $r_{3}\left(x\right)=-\overline{c}_{2}\left(x^{2}-1\right)/2$,
\begin{align}
\sup_{\mathbf{C}}\left|\mathbb{P}\left(U_{T}\in\mathbf{C}\right)-\int_{\mathbf{C}}\varphi\left(x\right)\left(1+r_{2}\left(x\right)b_{1,T}^{d_{f}}+r_{3}\left(x\right)b_{2,T}^{2}\right)dx\right| & =o\left(\left(Tb_{1,T}b_{2,T}\right)^{-1/2}\right).\label{Eq. (3) in VR-1}
\end{align}
\end{thm}
Theorem \ref{Theorem: Theorem 2 in VR DK-HAC} shows that the correction
term to the standard normal distribution, i.e., $\int_{\mathbf{C}}\varphi\left(x\right)$
$(r_{2}\left(x\right)b_{1,T}^{d_{f}}+r_{3}\left(x\right)b_{2,T}^{2})dx$,
depends on both smoothing directions. The error of the approximation
is of order $o((Tb_{1,T}$ $b_{2,T})^{-1/2})$ which can be larger
than that obtained in Theorem \ref{Theorem: Theorem 2 in VR} for
the HAC estimators. Similar to \eqref{Eq. (4) in VR}, we obtain uniformly
in $z$,
\begin{align}
\mathbb{P}\left(U_{T}\leq z\right) & =\Phi\left(z\left(1+\frac{1}{2}\overline{c}_{1}b_{1,T}^{d_{f}}+\frac{1}{2}\overline{c}_{2}b_{2,T}^{2}\right)\right)+O\left(\left(Tb_{1,T}b_{2,T}\right)^{-1/2}\right),\label{Eq. (4) in VR DK-HAC}
\end{align}
where $\mathbf{C}=(-\infty,\,z]$, which suggests that the standard
normal approximation is correct up to order $O((Tb_{1,T}b_{2,T})^{-1/2})$.
Eq. \eqref{Eq. (4) in VR DK-HAC} has a similar interpretation to
\eqref{Eq. (4) in VR}. Consider the time-varying AR(1) example in
\eqref{Eq. Example TV AR(1)} and suppose $\rho\left(u\right)>0$
for all $u.$ Then, $\overline{c}_{1}<0$. However, the sign of $\overline{c}_{2}$
is not easily determined even for this simple model. For the special
case $\rho\left(u\right)=\sin(u\pi/10)$, no break and $\sigma^{2}\left(u\right)=\sigma^{2}$
we have $\overline{c}_{2}<0$. Then, the implied critical value from
the approximation is larger than the standard normal critical value.
In general, however, the correction to strong persistence might be
either attenuated or strengthened by the correction to nonstationarity
depending on the true data-generating process.
Returning to the location model, consider the $t$-statistic based
on $\widehat{J}_{\mathrm{DK},T}^{*}$,
\begin{align*}
t_{\mathrm{DK}} & =\frac{\sqrt{T}\left(\widehat{\beta}-\beta_{0}\right)}{\sqrt{\widehat{J}_{\mathrm{DK},T}^{*}}}.
\end{align*}
Theorem \ref{Theorem: Theorem 2 in VR DK-HAC} and \eqref{Eq. (4) in VR DK-HAC}
imply that
\begin{align}
\mathbb{P}\left(t_{\mathrm{DK}}\leq z\right) & =\Phi\left(z\right)+p_{2}\left(z\right)\left(Tb_{1,T}b_{2,T}\right)^{-1/2}+o\left(\left(Tb_{1,T}b_{2,T}\right)^{-1/2}\right),\label{Eq. ERP t-DK}
\end{align}
for any $z\in\mathbb{R},$ where $p_{2}\left(z\right)$ is an odd
function. Under the conditions of Theorem \ref{Theorem: Theorem 2 in VR DK-HAC}
$p_{2}\left(z\right)=2^{-1}((C^{d_{f}+1/2}\overline{c}_{1}+C_{2}\overline{c}_{2})z\varphi\left(z\right))$
where $C$ is defined in Assumption \ref{Assumption 7 VR}, $C_{2}=(\overline{b}C^{d_{f}+1/2})^{1/2}$
and $\overline{b}$ is defined in Assumption \ref{Assumption K2 and b2}.
Thus, the ERP of $t_{\mathrm{DK}}$ can be larger than that of $t_{\mathrm{HAC}}$,
though the margin is small. This follows from the fact that $\widehat{J}_{\mathrm{DK},T}^{*}$
applies smoothing over two directions. The smoothing over time is
useful to flexibly account for nonstationarity. Its benefits appear
explicitly under the alternative hypothesis as we show in Section
\ref{Section Consequences for HAR} whereas the ERP refers to the
null hypothesis. One can show that the ERP of $t_{\mathrm{DK}}$ and
$t_{\mathrm{HAC}}$ remain unchanged if prewhitening is applied, though
the proofs are omitted since they are similar.
We can further compare the ERP of $t_{\mathrm{HAC}}$ and $t_{\mathrm{DK}}$
to that of the corresponding $t$-test under the fixed-$b$ asymptotics.
\citet{casini_fixed_b_erp} showed that the limiting distribution
of the original fixed-$b$ HAR test statistics under nonstationarity
is not pivotal as it depends on the true data-generating process of
the errors and regressors. This contrasts to the stationarity case
for which the fixed-$b$ limiting distribution is pivotal and the
ERP is of order $O(T^{-1})$ {[}see \citet{jansson:04} and \citet{sun/phillips/jin:08}{]}.
Based on an ERP of smaller magnitude relative to that of HAR tests
based on HAC estimators {[}cf. $O(T^{-1})<O((Tb_{1,T})^{-1/2})${]},
the literature has long suggested that the original fixed-$b$ HAR
tests are superior to HAR tests based on HAC estimators. However,
this breaks down under nonstationarity as shown by \citet{casini_fixed_b_erp}
who established that (i) the ERP of the original fixed-$b$ HAR tests
does not converge to zero because under nonstationarity the fixed-$b$
limiting distribution is different; (ii) for fixed-$b$ HAR tests
that use the critical values from the non-pivotal fixed-$b$ limiting
distribution the ERP increases by an order of magnitude relative to
the stationary case {[}i.e., from $O(T^{-1})$ to $O(T^{-\eta})$
with $\eta\in(0,\,1/2)${]}. Therefore, fixed-$b$ HAR tests can have
an ERP larger than that of $t_{\mathrm{HAC}}$ and $t_{\mathrm{DK}}$.
Overall, the results based on Edgeworth expansions show that the distortions
on the null rejection rates of the HAR tests can arise from time variation
in the second moments even when the mean is constant. Thus, these
results complement the asymptotic bias results induced by breaks in
the mean function.
\section{\label{Section Consequences for HAR}Consequences for HAR Inference}
In this section, we discuss the implications of the theoretical results
from Section \ref{Section Low Freq Cont - Theory}-\ref{Section Edgeworth-Expansions-for}.
In Section \ref{Subsection HAR inference methods}, we first present
a review of HAR inference methods and their connection to the estimates
considered in Section \ref{Section Low Freq Cont - Theory}. In Section
\ref{Subsec Finite-Sample-Low-Frequency} we present evidence that
the HAR inference tests can suffer from larger size distortions under
nonstationarity than under stationarity. In Section \ref{Subsec General-Low-Frequency}
we show the consequences of low frequency contamination for the power
of the HAR tests and we provide the corresponding theoretical results
in Section \ref{Subsec: Theoretical-Results-on Power}.
\subsection{\label{Subsection HAR inference methods}HAR Inference Methods}
There are two main approaches for HAR inference. Classical HAC standard
errors {[}cf. \citeauthor{newey/west:87} (\citeyear{newey/west:87},
\citeyear{newey/west:94}) and \citet{andrews:91}{]} require estimation
of the LRV defined as $J\triangleq\mathrm{lim}_{T\rightarrow\infty}J_{T}$
where $J_{T}$ is defined after \eqref{Eq. (h1)}. The form of $\left\{ V_{t}\right\} $
depends on the specific problem under study. For example, for a $t$-test
on a regression coefficient in the linear model $y_{t}=x{}_{t}\beta_{0}+e_{t}$
$\left(t=1,\ldots,\,T\right)$ we have $V_{t}=x_{t}e_{t}$. Classical
HAC estimators take the following form,
\begin{align*}
\widehat{J}_{\mathrm{HAC,}T}=\sum_{k=-T+1}^{T-1}K_{1}\left(b_{1,T}k\right)\widehat{\Gamma}\left(k\right) & ,
\end{align*}
where $\widehat{\Gamma}\left(k\right)$ is given in \eqref{Eq. Definition of Gamma(k)}
with $\widehat{V}_{t}=x_{t}\widehat{e}_{t}$ where $\left\{ \widehat{e}_{t}\right\} $
are the least-squares residuals, $K_{1}\left(\cdot\right)$ is a kernel
and $b_{1,T}$ is bandwidth. One can use the the Bartlett kernel,
advocated by \citet{newey/west:87}, the quadratic spectral kernel
as suggested by \citet{andrews:91}, or any other kernel suggested
in the literature, see e.g. \citet{dejong/davidson:00} and \citet{ng/perron:1996}.
Under $b_{1,T}\rightarrow0$ at an appropriate rate, we have $\widehat{J}_{\mathrm{HAC,}T}\overset{\mathbb{P}}{\rightarrow}J.$
Hence, equipped with $\widehat{J}_{\mathrm{HAC,}T}$, HAR inference
is standard and simple because HAR test statistics follow asymptotically
standard distributions.
HAC standard errors can result in oversized tests when there is
substantial temporal dependence {[}e.g., \citet{andrews:91}{]}. This
stimulated a second approach based on LRV estimators that keeps the
bandwidth at some fixed fraction of $T$ {[}cf. \citet{Kiefer/vogelsang/bunzel:00}{]},
e.g., using all autocovariances, so that $\widehat{J}_{\mathrm{\mathrm{KVB},}T}\triangleq T^{-1}\sum_{t=1}^{T}\sum_{s=1}^{T}\left(1-\left|t-s\right|/T\right)$
$\widehat{V}_{t}\widehat{V}_{s}$ which is equivalent to the Newey-West
estimator with $b_{1,T}=T^{-1}$. Under fixed-$b$ asymptotics
the reference distribution of HAR test statistics is nonstandard.
The validity of fixed-$b$ inference rests on stationarity {[}cf.
\citet{casini_fixed_b_erp}{]}. Many authors have considered various
versions of $\widehat{J}_{\mathrm{\mathrm{KVB},}T}$. However, the
one that leads to HAR inference tests that are least oversized is
the original $\widehat{J}_{\mathrm{\mathrm{KVB},}T}$ {[}see \citet{casini/perron_PrewhitedHAC}
for simulation results{]}. For comparison we also report the equally-weighted
cosine (EWC) estimator of \citet{lazarus/lewis/stock:17}. It is an
orthogonal series estimators that use long bandwidths,
\begin{align*}
\widehat{J}_{\mathrm{\mathrm{EWC},}T} & \triangleq B^{-1}\sum_{j=1}^{B}\Lambda_{j}^{2},\qquad\mathrm{where}\quad\Lambda_{j}=\sqrt{\frac{2}{T}}\sum_{t=1}^{T}\widehat{V}_{t}\cos\left(\pi j\left(\frac{t-1/2}{T}\right)\right)
\end{align*}
with $B$ some fixed integer. Assuming $B$ satisfies some conditions,
under fixed-$b$ asymptotics a $t$-statistic normalized by $\widehat{J}_{\mathrm{\mathrm{EWC},}T}$
follows a $t_{B}$ distribution where $B$ is the degree of freedom.
Recently, a new HAC estimator was proposed in \citet{casini_hac}.
Motivated by the power impact of low frequency contamination of existing
LRV estimators, he proposed a double kernel HAC (DK-HAC) estimator,
defined by
\begin{align*}
\widehat{J}_{\mathrm{DK},T}\triangleq\sum_{k=-T+1}^{T-1}K_{1}\left(b_{1,T}k\right)\widehat{\Gamma}_{\mathrm{DK}}\left(k\right) & ,
\end{align*}
where $b_{1,T}$ is a bandwidth sequence and $\widehat{\Gamma}_{\mathrm{DK}}\left(k\right)$
defined in Section \ref{Section Low Freq Cont - Theory} with $\widehat{c}_{T}\left(\cdot,\,k\right)$
replaced by
\begin{align*}
\widehat{c}_{\mathrm{DK,}T}\left(rn_{T}/T,\,k\right) & =\left(Tb_{2,T}\right)^{-1}\sum_{s=|k|+1}^{T}K_{2}\left(\frac{\left(rn_{T}-\left(s-|k|/2\right)\right)/T}{b_{2,T}}\right)\widehat{V}_{s}\widehat{V}{}_{s-|k|},
\end{align*}
with $K_{2}$ a kernel and $b_{2,T}$ a bandwidth. Note that $\widehat{c}_{\mathrm{DK,}T}$
and $\widehat{c}_{T}$ are asymptotically equivalent and the results
of Section \ref{Section Low Freq Cont - Theory} continue to hold
for $\widehat{c}_{\mathrm{DK,}T}$. More precisely, $\widehat{c}_{T}$
is a special case of $\widehat{c}_{\mathrm{DK,}T}$ with $K_{2}$
being a rectangular kernel and $n_{2,T}=Tb_{2,T}$. This approach
falls in the first category of standard inference $\widehat{J}_{\mathrm{DK},T}\overset{\mathbb{P}}{\rightarrow}J$
and HAR test statistics normalized by $\widehat{J}_{\mathrm{DK},T}$
follows standard distribution asymptotically. The DK-HAC estimator
involves two kernels: $K_{1}$ smooths the lagged sample autocovariances,
akin to the classical HAC estimators, while $K_{2}$ applies smoothing
over time. The latter feature is useful to avoid the low frequency
contamination. Additionally, \citet{casini/perron_PrewhitedHAC} proposed
prewhitened DK-HAC $(\widehat{J}_{\mathrm{\mathrm{pw},DK},T})$ estimator
that improves the size control of HAR tests and enjoys the same
asymptotic properties of $\widehat{J}_{\mathrm{DK},T}$. \citet{casini_hac}
and \citet{casini/perron_PrewhitedHAC} demonstrated via simulations
that tests based on $\widehat{J}_{\mathrm{DK},T}$ and $\widehat{J}_{\mathrm{pw,DK},T}$
have superior power properties relative to tests based on the other
estimators. In terms of size, the simulation results showed that
tests based on $\widehat{J}_{\mathrm{pw,DK},T}$ perform better than
those based on $\widehat{J}_{\mathrm{HAC,}T}$ and $\widehat{J}_{\mathrm{DK},T}$,
and is competitive with $\widehat{J}_{\mathrm{\mathrm{KVB},}T}$ when
the latter works well. We include $\widehat{J}_{\mathrm{DK},T}$
and $\widehat{J}_{\mathrm{pw,DK},T}$ in our simulations below. We
report the results only for the DK-HAC estimators that do not use
the pre-test for discontinuities in the spectrum {[}cf. \citet{casini/perron:change-point-spectra}{]}
because we do not want the results to be affected by such pre-test.
\subsection{\label{Subsec Finite-Sample-Low-Frequency}Null Rejection Rates and
Power in Finite-Sample}
In order to better understand the effect of nonstationarity on the
null rejection rates of HAR tests we first conduct a Monte Carlo analysis
where we compare a nonstationary model with a stationary one that
has either the same spectral density at frequency zero or the same
average dependence. Consider the following four AR(1) data-generating
processes (DGPs). DGP 1 is given by
\begin{align*}
V_{t} & =0.26V_{t-1}+e_{t},\qquad t=1,\ldots,\,T,
\end{align*}
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$ for all $t$. The
LRV of DGP 1 is $J=1.826$. DGP 2 is
\begin{align*}
V_{t} & =0.7817V_{t-1}+e_{t},\qquad t=1,\ldots,\,T,
\end{align*}
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$ for all $t$. Its
LRV is $J=20.988$. We now introduce two nonstationary DGPs. DGP 3
takes the following form
\begin{align*}
V_{t} & =\begin{cases}
0.9V_{t-1}+e_{t}, & 1\leq t\leq0.2T\\
0.1V_{t-1}+e_{t}, & 0.2T<t\leq T,
\end{cases}
\end{align*}
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$. Note that the spectral
density at frequency zero of $V_{t}$ is given by the weighted average
of the spectral densities of $V_{t}$ in the two regimes:
\begin{align*}
f\left(0\right)=\int_{0}^{1}f\left(u,\,0\right)du & =0.2\frac{1}{2\pi\left(1-2\cdot0.9+0.9^{2}\right)}+0.8\frac{1}{2\pi\left(1-2\cdot0.1+0.1^{2}\right)}=3.342.
\end{align*}
Thus, the LRV of $V_{t}$ is $J=2\pi\int_{0}^{1}f\left(u,\,0\right)du=20.988$
which takes the same value as the LRV of DGP 2. Further, DGP 3 has
the same average dependence as DGP 1, meaning that the AR(1) coefficient
in DGP 1 is equal to the weighted average of the AR(1) coefficients
of DGP 3 in the two regimes, i.e., $\overline{\rho}=0.2\cdot0.9+0.8\cdot0.1=0.26$.
We also want to verify whether the location of the break in persistence
in DGP 3 is important for the bias. Thus, we consider DGP 4:
\begin{align*}
V_{t} & =\begin{cases}
0.1V_{t-1}+e_{t}, & 1\leq t\leq0.5T\\
0.9V_{t-1}+e_{t}, & 0.5T<t\leq0.5T+0.2T\\
0.1V_{t-1}+e_{t}, & 0.5T+0.2T<t\leq T,
\end{cases}
\end{align*}
where $e_{t}\sim\mathscr{N}\left(0,\,1\right)$ for all $t$. While
in DGP 3 the regime with strong persistence occurs in the first 20\%
of the sample, in DGP 4 it occurs between the 50\% and 70\% of the
sample. The LRV of DGP 4 is the same as that of DGP 3.
For each DGP we consider three different initial conditions: (a) $V_{0}=0$;
(b) $V_{0}\sim\mathscr{N}\left(0,\,1\right)$; (c) $V_{0}\sim\mathscr{N}\left(0,\,4\right)$.
This is useful in order to verify whether the initial condition has
any effect on the bias generated by changes in the second-order properties.
DGP 3(a) should exhibit a smaller bias due to nonstationarity than
DGP 3(b,c) and 4. To see this, note that in DGP 3(a) the initial condition
is $V_{0}=0.$ Thus, the process starts from zero. Since there is
strong persistence in the first 20\% of the sample, the process is
more likely to stay close to zero in the first regime than when the
initial condition is $V_{0}\sim\mathscr{N}\left(0,\,1\right)$ or
$V_{0}\sim\mathscr{N}\left(0,\,4\right)$. In DGP 4 the different
specifications of the initial condition should not lead to any differences
in the bias due to nonstationarity because the regime with strong
dependence occurs about mid-sample.
To summarize, we have four DGPs. DGP 1 and 2 are stationary while
DGP 3 and 4 are nonstationary. Since DGP 2 has a LRV that takes the
same value as that of DGP 3 and 4, this allows us to better separate
the effect of persistence from that of nonstationarity in the second
moments on the following quantities: $\widehat{J}_{\mathrm{HAC}}$,
$-\widehat{c}_{1}b_{1,T}$ and $\widehat{\Gamma}\left(k\right)$ for
$k=0,1,\,5,\,10.$ In the simulations below $\widehat{J}_{\mathrm{HAC}}$
is the Newey-West estimator based on a predetermined number of lagged
sample autocovariances following the rule $4\left(T/100\right)^{2/9}$
{[}cf. \citet{lazarus/lewis/stock/watson:18}{]}. We compare $\widehat{\Gamma}\left(k\right)$
to the theoretical value $\Gamma_{T}\left(k\right)$ corresponding
to each DGP which can be computed by hand given the simple form of
the DGPs. In fact, for the nonstationary DGPs, $\Gamma_{T}\left(k\right)$
is a weighed average of the theoretical autocovariances corresponding
to each regime. Here, $\widehat{c}_{1}$ is an estimate of $\overline{c}_{1}$
in \eqref{Eq. (c1)} that enters the asymptotic bias of $\widehat{J}_{\mathrm{HAC}}$.
In order to compute $\widehat{c}_{1}$ we recall that the asymptotic
bias of the LRV estimator based on the Bartlett kernel is given by
\begin{align*}
\lim_{T\rightarrow\infty}b_{1,T}^{-1}\mathbb{E}\left(\widehat{J}_{\mathrm{HAC}}-J_{T}\right) & =-2\pi K_{\mathrm{BT},1}\int_{0}^{1}f^{\left(1\right)}\left(u,\,0\right)du,
\end{align*}
where
\begin{align*}
K_{\mathrm{BT},q} & =\lim_{x\rightarrow0}\frac{1-K_{\mathrm{BT}}\left(x\right)}{\left|x\right|^{q}}
\end{align*}
denotes the index of smoothness of the kernel at zero and $f^{\left(1\right)}\left(u,\,0\right)$
is the index of smoothness of the local spectral density at time $u$
and frequency zero. For the Bartlett kernel $K_{\mathrm{BT},q}=0$
if $q<1$, $K_{\mathrm{BT},q}=1$ if $q=1$ and $K_{\mathrm{BT},q}=\infty$
if $q>1.$ The Parzen characteristic exponent is the largest $q$
such that $K_{\mathrm{BT},q}$ is finite. Thus, the relative bias
is
\begin{align*}
\lim_{T\rightarrow\infty}b_{1,T}^{-1}\mathbb{E}\left(\widehat{J}_{\mathrm{HAC}}/J_{T}-1\right) & =-K_{\mathrm{BT},1}\frac{\int_{0}^{1}f^{\left(1\right)}\left(u,\,0\right)du}{\int_{0}^{1}f\left(u,\,0\right)du}=-\overline{c}_{1},
\end{align*}
using $K_{\mathrm{BT},1}=1.$ The index of smoothness of $f\left(u,\,\omega\right)$
at $\omega=0$ is defined as
\begin{align*}
f^{\left(1\right)}\left(u,\,0\right) & =\frac{1}{2\pi}\sum_{k=-\infty}^{\infty}|k|\Gamma\left(u,\,k\right).
\end{align*}
For an AR(1) process with parameters $\rho\left(u\right)$ and $\sigma_{e}^{2}\left(u\right)$,
we have $\Gamma\left(u,\,k\right)=\sigma_{e}^{2}\left(u\right)\rho\left(u\right)^{|k|}/(1-\rho\left(u\right)^{2}).$
It follows that
\begin{align*}
f^{\left(1\right)}\left(u,\,0\right) & =-\frac{1}{2\pi}\frac{2\rho\left(u\right)\sigma_{e}^{2}\left(u\right)}{\left(\rho\left(u\right)-1\right)^{3}\left(1+\rho\left(u\right)\right)}.
\end{align*}
Based on this result we can obtain $\overline{c}_{1}$ for each model.
In particular, for model DGP 1, 2, 3 and 4 we have $\overline{c}_{1}=0.55$,
3.92, 9.04 and 9.05, respectively.
We estimate $\overline{c}_{1}$ as follows. For DGP 1, we obtain the
OLS residuals $\widehat{V}_{t}$ and estimate $\rho$ and $\sigma_{e}^{2}$
from the autoregression
\begin{align*}
\widehat{V}_{t} & =\rho\widehat{V}_{t-1}+e_{t},\qquad\qquad t=1,\ldots,\,T,
\end{align*}
where $\sigma_{e}^{2}$ is the variance of $e_{t}$. Let these estimates
be denoted by $\widehat{\rho}$ and $\widehat{\sigma}_{e}^{2}$, respectively.
Then, the estimate of $\overline{c}_{1}$ is defined as
\begin{align*}
\widehat{c}_{1} & =-\frac{2\widehat{\rho}\widehat{\sigma}_{e}^{2}}{\widehat{J}_{\mathrm{HAC}}\left(\widehat{\rho}-1\right)^{3}\left(1+\widehat{\rho}\right)}.
\end{align*}
The same applies to DGP 2. For DGP 3, we obtain the estimate of the
autoregressive coefficient of $V_{t}$ and of the variance of the
innovations by estimating the autoregression in the two regimes separately.
That is, we obtain
\begin{align*}
\widehat{V}_{t} & =\begin{cases}
\widehat{\rho}_{1}\widehat{V}_{t-1}+\widehat{e}_{t}, & 1\leq t\leq0.2T\\
\widehat{\rho}_{2}\widehat{V}_{t-1}+\widehat{e}_{t}, & 0.2T<t\leq T,
\end{cases}
\end{align*}
where we also compute $\widehat{\sigma}_{1,e}^{2}$ and $\widehat{\sigma}_{2,e}^{2}$
which are the sample variances of the residuals $\widehat{e}_{t}$
in the two regimes, respectively. Then, the estimate of $\overline{c}_{1}$
is defined as
\begin{align*}
\widehat{c}_{1} & =-0.2\frac{2\widehat{\rho}_{1}\widehat{\sigma}_{1,e}^{2}}{\widehat{J}_{\mathrm{HAC}}\left(\widehat{\rho}_{1}-1\right)^{3}\left(1+\widehat{\rho}_{1}\right)}-0.8\frac{2\widehat{\rho}_{2}\widehat{\sigma}_{2,e}^{2}}{\widehat{J}_{\mathrm{HAC}}\left(\widehat{\rho}_{2}-1\right)^{3}\left(1+\widehat{\rho}_{2}\right)}.
\end{align*}
The same applies to DGP 4 with the difference that the autoregressive
coefficient and the variance of the innovations are estimated separately
in each of the three distinct regimes.
We consider the sample size $T=100,\,200$ and 1000, and 50,000 repetitions
were used for each DGP. The results are reported in Table \ref{Table T=00003D100, 200, 1000}.
Let us first discuss the finite-sample properties of $\widehat{J}_{\mathrm{HAC}}$.
The results clearly suggest that $\widehat{J}_{\mathrm{HAC}}$ deviates
substantially from $J$ when the data are nonstationary. $\widehat{J}_{\mathrm{HAC}}$
underestimates $J$ for all DGPs but it does so much more when the
DGP is nonstationary. The difference between the values of $\widehat{J}_{\mathrm{HAC}}$
in DGP 2 and those in DGP 3-4 is about one half, e.g., $\widehat{J}_{\mathrm{HAC}}=6.775$
in DGP 2(a) and $\widehat{J}_{\mathrm{HAC}}=3.142$ in DGP 3(a). As
the sample size increases the downward bias becomes smaller, though
$\widehat{J}_{\mathrm{HAC}}$ still underestimates $J$ for $T=1000$.
The downward bias continues to remain larger in DGP 3-4 than in DGP
2 even when $T=1000$. Thus, this evidence based on $\widehat{J}_{\mathrm{HAC}}$
already points out that basic forms of nonstationarity generate bias
in the LRV estimator. This bias adds to the well-known bias generated
by strong persistence in stationary data documented in the literature.
Let us discuss the relative bias $-\overline{c}_{1}b_{1,T}$ and its
estimate $-\widehat{c}_{1}b_{1,T}$. First note that $-\overline{c}_{1}b_{1,T}<0$
and $-\widehat{c}_{1}b_{1,T}<0$ for all DGPs and sample sizes considered.
This confirms the downward bias of $\widehat{J}_{\mathrm{HAC}}$ observed
above. For a given model, the asymptotic relative bias $-\overline{c}_{1}b_{1,T}$
and its estimate increase with the sample size. The downward bias
is much larger for the nonstationary DGP 3-4 than for the stationary
DGP 1-2. The estimates $-\widehat{c}_{1}b_{1,T}$ of the relative
bias $-\overline{c}_{1}b_{1,T}$ significantly underestimate $-\overline{c}_{1}b_{1,T}$
in DGP 3-4 while in DGP 1-2 the deviations are much smaller. The large
deviations of $-\widehat{c}_{1}b_{1,T}$ from $-\overline{c}_{1}b_{1,T}$
continue to hold even for $T=1000.$
We now move to discuss the finite-sample properties of $\widehat{\Gamma}\left(k\right)$.
When the data are stationary, $\widehat{\Gamma}\left(k\right)$ is
close to $\Gamma_{T}\left(k\right)$ even when $T=100$ and it approaches
$\Gamma_{T}\left(k\right)$ when $T=1000$. For nonstationary data,
$\widehat{\Gamma}\left(k\right)$ is much farther from $\Gamma_{T}\left(k\right)$.
For example, in DGP 2(a) $\widehat{\Gamma}\left(0\right)=2.507$ and
$\Gamma_{T}\left(0\right)=2.571$ whereas in DGP 3(a) $\widehat{\Gamma}\left(0\right)=1.589$
and $\Gamma_{T}\left(0\right)=1.861$. Thus, $\widehat{\Gamma}\left(k\right)$
has larger bias (in general downward) when the data are nonstationary.
This result is present even when $T=200$. As $T$ increases, $\widehat{\Gamma}\left(k\right)$
approaches $\Gamma_{T}\left(k\right)$ for all DGPs, though the downward
bias remains larger in DGP 3-4 than in DGP 1-2.
We repeated this exercise for other DGPs and the conclusions were
the same. The results suggest that under nonstationarity the bias
in the LRV estimator is affected by multiple factors. In addition
to the downward bias arising from strong persistence which is also
present under stationarity there is bias generated by the time-varying
properties of the process. Under the null hypothesis this time variation
occurs in the autocovariance structure of the process. For example,
in DGP 3 one has $0.2T$ observations to estimate $2\pi\int_{0}^{0.2}f\left(u,\,0\right)du=0.4\pi f\left(0\right)$
where $f\left(0\right)=1/(2\pi\left(1-2\rho+\rho^{2}\right))$ with
$\rho=0.9$, and $0.8T$ observations to estimate $2\pi\int_{0.2}^{1}f\left(u,\,0\right)du=1.6\pi f\left(0\right)$
where $f\left(0\right)=1/(2\pi\left(1-2\rho+\rho^{2}\right))$ with
$\rho=0.1$. This is more difficult than estimating $2\pi f\left(0\right)=1/(2\pi\left(1-2\rho+\rho^{2}\right))$
with $\rho=0.7817$ using $T$ observations, which applies to DGP
2. Even if the total sample size is $T$ in both DGP 2 and 3, nonstationarity
reduces the effective sample size making the estimation of the LRV
in DGP 3 effectively based on a smaller number of observations. For
example, $\widehat{\Gamma}\left(k\right)$ involves an average on
$\{\widehat{V}_{t}\widehat{V}_{t-k}\}$ for $t=k+1,\ldots,\,T$. Some
of these pairs $\{\widehat{V}_{t}\widehat{V}_{t-k}\}$ are such that
$\widehat{V}_{t}$ and $\widehat{V}_{t-k}$ belong to two different
regimes, and so contribute bias to the estimation of $\Gamma_{T}\left(k\right)$.
Under stationarity all the pairs $\{\widehat{V}_{t}\widehat{V}_{t-k}\}$
are such that $\widehat{V}_{t}$ and $\widehat{V}_{t-k}$ belong to
the same regime leading to more precise estimates of $\widehat{\Gamma}\left(k\right)$
and LRV. In addition, changes in persistence over short regimes share
features similar to shifts in the mean, at least graphically. While
the former is consistent with the null hypothesis, the latter is not.
This is likely to generate some bias where changes in persistence
are confounded with shifts in the mean even when the unconditional
mean of the series has not changed. The downward bias due to strong
persistence and the bias due to time-varying second-order properties
are likely to influence each other making the estimation problem even
harder.
We now investigate the consequence of nonstationarity for HAR inference.
We obtain the empirical size and power for a two-tailed $t$-test
on the intercept normalized by several LRV estimators for the model
$y_{t}=\delta+V_{t}$ with $\delta=0$ under the null and $\delta>0$
under the alternative hypothesis. Model M1 involves an SLS process:
$V_{t}=0.9V_{t-1}+u_{t}$, $V_{0}\sim\mathscr{N}\left(0,\,1\right)$,
$u_{t}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(0,\,1\right)$ for $t=1,\ldots,\,T_{1}^{0}$
with $T_{1}^{0}=T\lambda_{1}^{0}$, and $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$,
$\rho\left(t/T\right)=0.3\left(\cos\left(1.5-\cos\left(t/T\right)\right)\right)$,
$u_{t}\sim\mathscr{\mathrm{i.i.d.}\,\mathscr{N}}\left(0,\,0.5\right)$
for $t=T_{1}^{0}+1,\ldots,\,T$. Note that $\rho\left(\cdot\right)$
varies between 0.172 and 0.263. We set $\lambda_{1}^{0}=0.1$. In
addition to M1, we consider other models: M2 involves a time-varying
AR(1) with a break in volatility $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$,
$\rho\left(t/T\right)=0.7(\cos\left(1.5t/T\right))$, $u_{t}\sim\mathscr{N}\left(0,\,\sigma_{t}^{2}\right)$,
$\sigma_{t}^{2}=5$ for $t\leq4$ and $\sigma_{t}^{2}=0.25$ for $t>4$,
$V_{0}\sim\mathscr{N}\left(0,\,5\right)$; M3 involves $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$,
$\rho\left(t/T\right)=0.8(\cos\left(1.5t/T\right))$, $u_{t}\sim\mathscr{N}\left(0,\,0.25\right)$,
$V_{0}=0$ with outliers $V_{t}\sim\mathrm{Uniform}\left(\underline{c},\,5\underline{c}\right)$
for $t=T/2,\,3T/4$ where $\underline{c}=-1/(\sqrt{2}\mathrm{erfc^{-1}\left(3/2\right))}\mathrm{med}\left(\left|V-\mathrm{med}\left(V\right)\right|\right)$
with $\mathrm{erfc}^{-1}$ the inverse complementary error function,
$\mathrm{med}\left(\cdot\right)$ is the median and $V=\left(V_{t}\right)_{t=1}^{T}$;\footnote{In this literature, values smaller than $\underline{c}$ are not classified
as outliers. } M4 involves a time varying AR(1) with periods of strong persistence
where $V_{t}=\rho\left(t/T\right)V_{t-1}+u_{t}$, $\rho\left(t/T\right)=0.95(\cos\left(1.5t/T\right))$,
$u_{t}\sim\mathscr{\mathrm{i.i.d.}\,\mathscr{N}}\left(0,\,0.4\right)$
and $V_{0}\sim\mathscr{N}\left(0,\,4\right)$. $\rho\left(\cdot\right)$
varies between 0.7 and 0.05 in M2, between 0.05 and 0.8 in M3 and
between 0.95 and 0.07 in M4.
We consider the DK-HAC estimators with and without prewhitening ($\widehat{J}_{\mathrm{DK},T}$,
$\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},T}$, $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},\mu,T}$)
of \citet{casini_hac} and \citet{casini/perron_PrewhitedHAC}, respectively;
\citeauthor{andrews:91}' \citeyearpar{andrews:91} HAC estimator
with and without the prewhitening procedure of \citet{andrews/monahan:92};
\citeauthor{newey/west:87}'s \citeyearpar{newey/west:87} HAC estimator
with the popular rule to select the number of lags (i.e., $b_{1,T}=(4(T/100)^{2/9})^{-1}$;
Newey-West with the fixed-$b$ method of \citet{Kiefer/vogelsang/bunzel:00}
with $b=1$ (labeled KVB); and the Equally-Weighted Cosine (EWC) of
\citet{lazarus/lewis/stock/watson:18} with the bandwidth choice recommended
by the authors. For the DK-HAC estimators we use the data-dependent
methods for the bandwidths, kernels and choice of $n_{T}$ as proposed
in \citet{casini_hac} and \citet{casini/perron_PrewhitedHAC}, which
are optimal under mean-squared error (MSE). Let $\widehat{V}_{t}$
denote the least-squares residual based on $\widehat{\delta}$ where
the latter is the least-squares estimate of $\delta$. We set $\widehat{b}_{1,T}=0.6828(\widehat{\phi}\left(2\right)T\widehat{\overline{b}}_{2,T})^{-1/5}$
where
\begin{align*}
\widehat{\phi}\left(2\right) & =\left(18\left(\frac{n_{T}}{T}\sum_{j=0}^{\left\lfloor T/n_{3,T}\right\rfloor -1}\frac{\left(\widehat{\sigma}\left(\left(jn_{T}+1\right)/T\right)\widehat{a}_{1}\left(\left(jn_{T}+1\right)/T\right)\right)^{2}}{\left(1-\widehat{a}_{1}\left(\left(jn_{T}+1\right)/T\right)\right)^{4}}\right)^{2}\right)/\\
& \quad\left(\frac{n_{T}}{T}\sum_{j=0}^{\left\lfloor T/n_{3,T}\right\rfloor -1}\frac{\left(\widehat{\sigma}\left(\left(jn_{T}+1\right)/T\right)\right)^{2}}{\left(1-\widehat{a}_{1}\left(\left(jn_{T}+1\right)/T\right)\right)^{2}}\right)^{2},
\end{align*}
with
\begin{align*}
\widehat{a}_{1}\left(u\right)=\frac{\sum_{j=t-n_{T}+1}^{t}\widehat{V}_{j}\widehat{V}_{j-1}}{\sum_{j=t-n_{T}+1}^{t}(\widehat{V}_{j-1})^{2}}, & \qquad\mathrm{and}\qquad\widehat{\sigma}\left(u\right)=(\sum_{j=t-n_{T}+1}^{t}(\widehat{V}_{j}-\widehat{a}_{1}\left(u\right)\widehat{V}_{j-1})^{2})^{1/2},
\end{align*}
and $\widehat{\overline{b}}_{2,T}=\left(n_{T}/T\right)\sum_{r=1}^{\left\lfloor T/n_{T}\right\rfloor -1}$
$\widehat{b}_{2,T}\left(rn_{T}/T\right)$, $\widehat{b}_{2,T}\left(u\right)=1.6786(\widehat{D}_{1}\left(u\right)){}^{-1/5}(\widehat{D}_{2}\left(u\right))^{1/5}T^{-1/5}$
where $\widehat{D}_{2}\left(u\right)\triangleq2\sum_{l=-\left\lfloor T^{4/25}\right\rfloor }^{\left\lfloor T^{4/25}\right\rfloor }\widehat{c}_{\mathrm{DK,}T}\left(u,\,l\right)^{2}$
and
\begin{align*}
\widehat{D}_{1}\left(u\right) & \triangleq(\left[S_{\omega}\right]^{-1}\sum_{s\in S_{\omega}}[3\pi^{-1}(1+0.8(\cos1.5+\cos4\pi u)\exp(-i\omega_{s}))^{-4}(0.8(-4\pi\sin(4\pi u)))\exp(-i\omega_{s})\\
& \quad-\pi^{-1}\left|1+0.8(\cos1.5+\cos4\pi u)\exp(-i\omega_{s})\right|^{-3}(0.8(-16\pi^{2}\cos(4\pi u)))\exp(-i\omega_{s})])^{2},
\end{align*}
with $\left[S_{\omega}\right]$ being the cardinality of $S_{\omega}$
and $\omega_{s+1}>\omega_{s}$, $\omega_{1}=-\pi,\,\omega_{\left[S_{\omega}\right]}=\pi.$
We set $n_{T}=T^{0.6}$, $S_{\omega}=\{-\pi,\,-3,\,-2,\,-1,\,0,\,1,\,2,\,3,\,\pi\}.$
$K_{1}\left(\cdot\right)$ is the QS kernel and $K_{2}\left(x\right)=6x\left(1-x\right)$
for $x\in\left[0,\,1\right].$
Table \ref{Table S1 - Figure 1} reports the results using 5,000 replications.
The $t$-test based on \citeauthor{newey/west:87}'s \citeyearpar{newey/west:87}
and \citeauthor{andrews:91}' \citeyearpar{andrews:91} prewhitened
HAC estimators are excessively oversized. \citeauthor{andrews:91}'
\citeyearpar{andrews:91} HAC-based test is slightly undersized while
the KVB's fixed-$b$ and EWC-based tests are severely undersized.
The fact that the KVB's fixed-$b$ and EWC-based tests have larger
size distortions than other tests is consistent with the results in
Section \ref{Section Edgeworth-Expansions-for} which suggest that
they have a larger ERP. For the $t$-test on the intercept, $\widehat{J}_{\mathrm{DK},T}$
can lead to tests that are oversized when there is strong dependence.
However, the prewhitened DK-HAC estimators $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},T}$
and $\widehat{J}_{\mathrm{DK,pw},\mathrm{SLS},\mu,T}$ lead to tests
having more accurate rejection rates. Nonstationarity affects the
power of the tests based on LRV estimators that rely on $\widehat{\Gamma}\left(k\right)$
or equivalently on $I_{T}\left(\omega\right)$ (e.g., the EWC). The
KVB's fixed-$b$ and EWC-based tests suffer from relatively large
power losses. The power of tests normalized by \citeauthor{newey/west:87}'s
(1987) and \citeauthor{andrews:91}' \citeyearpar{andrews:91} prewhitened
HAC are not comparable because they are significantly oversized. The
DK-HAC-based tests have the best power, the second best being \citeauthor{andrews:91}'
\citeyearpar{andrews:91} HAC-based test.
Turning to M2, Table \ref{Table S1 - Figure 1} shows some size distortions
and power losses for KVB's fixed-$b$ and EWC-based tests. The prewhitened
DK-HAC-based tests display accurate size control and good power. \citeauthor{newey/west:87}'s
(1987) and \citeauthor{andrews:91}' \citeyearpar{andrews:91} prewhitened
HAC-based tests are again excessively oversized. \citeauthor{andrews:91}'
\citeyearpar{andrews:91} HAC-based test and the DK-HAC-based test
show a similar performance. For model M3-M4, Table \ref{Table S1 - Figure 1}
shows that all methods lead to oversized tests except prewhitened
DK-HAC and KVB's fixed-$b$. However, the KVB's fixed-$b$-based tests
show substantial unde-rejection that has consequences for power whereas
the prewhitened DK-HAC-based-tests show accurate null rejection rates
and good power. Finally, the simulations show that the null rejection
rates of HAC- and DK-HAC-based tests are not very far from each other,
thereby confirming that their respective ERP are close as shown in
Section \ref{Section Edgeworth-Expansions-for}.
\subsection{\label{Subsec General-Low-Frequency}General Low Frequency Contamination}
We now discuss HAR inference tests for which the low frequency contamination
results of Section \ref{Section Low Freq Cont - Theory} hold asymptotically.
This means that $d^{*}>0$ for all $T$ and as $T\rightarrow\infty$.
This comprises the class of HAR tests that admit a nonstationary alternative
hypothesis. This class is very large and includes most HAR tests
as discussed in the Introduction. Here we consider the Diebold-Mariano
test for the sake of illustration and remark that similar issues apply
to other HAR tests.
The Diebold-Mariano test statistic is defined as $t_{\mathrm{DM}}\triangleq T_{n}^{1/2}\overline{d}_{L}/\sqrt{\widehat{J}_{d_{L},T}}$,
where $\overline{d}_{L}$ is the average of the loss differentials
between two competing forecast models, $\widehat{J}_{d_{L},T}$ is
an estimate of the LRV of the loss differential series and $T_{n}$
is the number of observations in the out-of-sample. We use the quadratic
loss. We consider an out-of-sample forecasting exercise with a fixed
forecasting scheme where, given a sample of $T$ observations, $0.5T$
observations are used for the in-sample and the remaining half is
used for prediction {[}see \citet{perron/yamamoto:18} for recommendations
on using a fixed scheme in the presence of breaks{]}. The DGP under
the null hypothesis is given by $y_{t}=1+\beta_{0}x_{t-1}^{(0)}+e_{t}$
where $x_{t-1}^{(0)}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(1,\,1\right)$,
$e_{t}=0.3e_{t-1}+u_{t}$ with $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$,
and we set $\beta_{0}=1$ and $T=400.$ The two competing models both
involve an intercept but differ with respect to the predictor used
in place of $x_{t}^{(0)}$. The first forecast model uses $x_{t}^{(1)}$
while the second uses $x_{t}^{(2)}$ where $x_{t}^{(1)}$ and $x_{t}^{(2)}$
are independent $\mathrm{i.i.d.}\,\mathscr{N}\left(1,\,1\right)$
sequences, both independent from $x_{t}^{(0)}$. Each forecast model
generates a sequence of $\tau\left(=1\right)$-step ahead out-of-sample
losses $L_{t}^{(j)}$ $\left(j=1,\,2\right)$ for $t=T/2+1,\ldots,\,T-\tau.$
Then $d_{t}\triangleq L_{t}^{(2)}-L_{t}^{(1)}$ denotes the loss differential
at time $t$. The Diebold-Mariano test rejects the null hypothesis
of equal predictive ability when $\overline{d}_{L}$ is sufficiently
far from zero. Under the alternative hypothesis, the two competing
forecast models are as follows: the first uses $x_{t}^{(1)}=x_{t}^{(0)}+u_{X_{1},t}$
where $u_{X_{1},t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$
while the second uses $x_{t}^{(2)}=x_{t}^{(0)}+0.2z_{t}+2u_{X_{2},t}$
for $t\in\left[1,\ldots,\,3T/4-1,\,3T/4+21,\ldots T\right]$ and $x_{t}^{(2)}=\delta\left(t/T\right)+0.2z_{t}+2u_{X_{2},t}$
for $t=3T/4,\ldots,\,3T/4+20$ with $u_{X_{2},t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$,
where $z_{t}$ has the same distribution as $x_{t}^{(0)}.$
We consider four specifications for $\delta\left(\cdot\right).$ In
the first $x_{t}^{(2)}$ is subject to an abrupt break in the mean
$\delta\left(t/T\right)=\delta>0$; in the second $x_{t}^{(2)}$ is
locally stationary with time-varying mean $\delta\left(t/T\right)=\delta\left(\sin\left(t/T-3/4\right)\right)$;
in the third specification $x_{t}^{(2)}=x_{t}^{(0)}+0.2z_{t}+2u_{X_{2},t}$
for $t\in[1,\ldots,\,T/2-30,\,T/2$ $+21,\ldots T]$ and $x_{t}^{(2)}=\delta\left(t/T\right)+0.2z_{t}+2u_{X_{2},t}$
for $t=T/2-30,\ldots,\,T/2+20$ with $\delta\left(t/T\right)=\delta(\sin(t/T-1/2$
$-30/T))$; in the fourth $x_{t}^{(2)}$ is the same as in the second
with in addition two outliers $x_{t}^{(2)}\sim\mathrm{Uniform}\left(\left|\underline{c}\right|,\,5\left|\underline{c}\right|\right)$
for $t=6T/10,\,8T/10$ where $\underline{c}=-1/(\sqrt{2}\mathrm{erfc^{-1}\left(3/2\right))}\mathrm{med}(|x^{(2)}-\mathrm{med}$
$(x^{(2)})|)$ where $x^{(2)}=(x_{t}^{(2)})_{t=1}^{T}$. That is,
in the second model $x_{t}^{(2)}$ is locally stationary only in the
out-of-sample, in the third it is locally stationary in both the in-sample
and out-of sample and in the fourth model $x_{t}^{(2)}$ has two outliers
in the out-of-sample. The location of the outliers is irrelevant for
the results; they can also occur in the in-sample.
Table \ref{Table Power DM Test} reports the null rejection rate and
the power of the various tests for all models. We begin with the case
$\delta\left(t/T\right)=\delta>0$ (top panel). The null rejection
rate of the test using the DK-HAC estimators is accurate while the
tests using other LRV estimators are oversized with the exception
of the KVB's fixed-$b$ method for which the rejection rate is equal
to zero. The HAR tests using existing LRV estimators have lower power
relative to that obtained with the DK-HAC estimators for small values
of $\delta$. When $\delta$ increases the tests standardized by the
HAC estimators of \citet{andrews:91} and \citet{newey/west:87},
and by the KVB's fixed-$b$ and EWC LRV estimators display non-monotonic
power gradually converging to zero as the alternative gets further
away from the null value. In contrast, when using the DK-HAC estimators
the test has monotonic power that reaches and maintains unit power.
The results for the other models are even stronger. In general, except
when using the DK-HAC estimators, all tests display serious power
problems. Thus, either form of nonstationarity or outliers leads to
similar implications, consistent with our theoretical results.
In order to further assess the theoretical results from Section \ref{Section Low Freq Cont - Theory},
Figure \ref{Fig_1_ET} (top panel) reports the plots of $d_{t}$,
its sample autocovariances and its periodogram, for $\delta=1$.
Figures \ref{Fig_2_ET}-\ref{Fig_3_ET} (top panels) in the supplement
report the corresponding plots for $\delta=2,\,5$, respectively.
We only consider the case $\delta_{t}=\delta>0$. The other cases
lead to the same conclusions. For $\delta=1$, Figure \ref{Fig_1_ET}
(top panel) shows that $\widehat{\Gamma}\left(k\right)$ decays slowly.
As $\delta$ increases, from Figures \ref{Fig_2_ET} and \ref{Fig_3_ET}
(top panels), $\widehat{\Gamma}\left(k\right)$ decays even more slowly
at a rate far from the typical exponential decay of short memory processes.
This suggests evidence of long memory. However, the data are short
memory with small temporal dependence. What is generating the spurious
long memory effect is the nonstationarity present under the alternative
hypothesis. This is visible in the top panels which present plots
of $d_{t}$ for the first specification. The shift in the mean of
$d_{t}$ for $t=3T/4,\ldots,\,3T/4+20$ is responsible for the long
memory effect. This corresponds to the second term of \eqref{Eq. Gamma_hat(k) Ineq Theorem}
in Theorem \ref{Theorem ACF Nonstat}. The overall behavior of the
sample autocovariance is as predicted by Theorem \ref{Theorem ACF Nonstat}.
For small lags, $\widehat{\Gamma}\left(k\right)$ shows a power-like
decay and it is positive. As $k$ increases to medium lags, the autocovariances
turn negative because the sum of all sample autocovariances has to
be equal to zero {[}cf. \citet{percival:1992}{]}. Next, we
move to the bottom panels which plot the periodogram of $\{d_{t}\}$.
It is unbounded at frequencies close to $\omega=0$ as predicted by
Theorem \ref{Theorem Periodogram Long Memory Effects} and as would
occur if long memory was present. It also explains why the Diebold-Mariano
test normalized by Newey-West's, Andrews', KVB's fixed-$b$ and EWC's
LRV estimators have serious power problems. These LRV estimators are
inflated and consequently the tests lose power. The figures show
that as we raise $\delta$ the more severe these issues and the
power losses so that the power eventually reaches zero. This is consistent
with our theory since $d^{*}$ is increasing in $\delta$ (cf. $d^{*}\thickapprox0.1\cdot0.9\delta^{2}$).
We now verify the results about the local sample autocovariance $\widehat{c}_{T}\left(u,\,k\right)$
and the local periodogram from Theorems \ref{Theorem Local ACF Nonstat}-\ref{Theorem Local Periodogram Long Memory Effects}.
We set $n_{2,T}=T^{0.6}=36$ following the MSE criterion of \citet{casini_hac}.
We consider (i) $u=236/T$, (ii-a) $u=T_{1}^{0}/T=3/4$ and (ii-b)
$u=264/T$. Note that cases (i)-(ii-b) correspond to parts (i)-(ii-b)
in Theorems \ref{Theorem Local ACF Nonstat}-\ref{Theorem Local Periodogram Long Memory Effects}.
We consider $\delta=1,\,2$ and $5$. According to Theorems \ref{Theorem Local ACF Nonstat}-\ref{Theorem Local Periodogram Long Memory Effects},
we should expect long memory features only for case (ii-a). Figures
\ref{Fig_1_ET} and \ref{Fig_2_ET}-\ref{Fig_3_ET} in the supplement
confirm this. The results pertaining to case (ii-a) are plotted in
the middle panels. They show that the local autocovariance displays
slow decay similar to the pattern discussed above for $\widehat{\Gamma}\left(k\right)$
and that this problem becomes more severe as $\delta$ increases.
Such long memory features also appear for $I_{\mathrm{L}}\left(3/4,\,\omega\right)$.
The bottom panels in Figures \ref{Fig_1_ET} and \ref{Fig_2_ET}-\ref{Fig_3_ET}
show that the local periodogram at $u=3/4$ and at a frequency close
to $\omega=0$ are extremely large. The latter result is consistent
with Theorem \ref{Theorem Local Periodogram Long Memory Effects}-(ii-a)
which suggests that $I_{\mathrm{L,}T}\left(3/4,\,\omega\right)\rightarrow\infty$
as $\omega\rightarrow0$. For case (i) and (ii-b) both figures show
that the local autocovariance and the local periodogram do not display
long memory features. Indeed, they have forms similar to those of
a short memory process, a result consistent with Theorems \ref{Theorem Local ACF Nonstat}-\ref{Theorem Local Periodogram Long Memory Effects}
also for cases (i) and (ii-b).
It is noteworthy to explain why HAR inference based on the DK-HAC
estimators does not suffer from the low frequency contamination even
for case (ii-a). The DK-HAC estimator computes an average of the local
spectral density over time blocks. If one of these blocks contains
a discontinuity in the spectrum, then as in case (ii-a) some bias
would arise for the local spectral density estimate corresponding
to that block. However, by virtue of the time-averaging over blocks
that bias becomes negligible. Hence, nonparametric smoothing over
time asymptotically cancels the bias, so that inference based on the
DK-HAC estimators is robust to nonstationarity.
\subsection{\label{Subsec: Theoretical-Results-on Power}Theoretical Results
about the Power}
We present theoretical results about the power of $t_{\mathrm{DM}}$
for the case of general low frequency contamination discussed in Section
\ref{Subsec General-Low-Frequency}. In particular, we focus on specification
(1) (i.e., $\delta>0$). The same intuition and qualitative theoretical
results apply to the other specifications of $\delta\left(\cdot\right)$.
Let $t_{\mathrm{DM},i}=T_{n}^{1/2}\overline{d}_{L}/\sqrt{\widehat{J}_{d_{L},i,T}}$
denote the DM test statistic where $i=\mathrm{DK},\,\mathrm{pwDK},\,\mathrm{KVB},\,\mathrm{EWC},$
$\mathrm{A91},\,\mathrm{pwA}91$, $\mathrm{NW87}$ and $\mathrm{pwNW87}$
with $\widehat{J}_{\mathrm{A91},T}$ and $\widehat{J}_{\mathrm{NW87},T}$
being $\widehat{J}_{\mathrm{HAC},T}$ using the quadratic spectral
and Bartlett kernel, respectively. Define the power of $t_{\mathrm{DM},i}$
as $\mathbb{P}_{\delta}(|t_{\mathrm{DM},i}|>z_{1-\alpha/2})$ where
$z_{1-\alpha/2}$ is the $1-\alpha/2$ quantile of the standard normal
for a two-sided test with significance level $\alpha\in\left(0,\,1\right)$.
To avoid repetitions we present the results only for $i=\mathrm{DK},\,\mathrm{KVB}$
and $\mathrm{NW87}$. The results concerning the prewhitening DK-HAC
estimator are the same as those corresponding to the DK-HAC estimator
while the results concerning the EWC estimator are similar to those
corresponding to the KVB's fixed-$b$ estimator, though for the latter
the non-monotonic power is more pronounced. The results pertaining
to \citeauthor{andrews:91}' \citeyearpar{andrews:91} HAC estimator
(with and without prewhitening) are the same as those corresponding
to \citeauthor{newey/west:87}'s \citeyearpar{newey/west:87} estimator.
Let $n_{\delta}=T-T_{b}-2$ denote the length of the regime in which
$x_{t}^{(2)}$ exhibits a shift $\delta$ in the mean. The deviation
from the null hypothesis depends on the shift magnitude $\delta$
and on $n_{\delta}$.
\begin{thm}
\label{Theorem Power DM HAR Tests}Let $\left\{ d_{t}-\mathbb{E}(d_{t})\right\} _{t=1}^{T_{n}}$
be an SLS process satisfying Assumption \ref{Assumption Smothness of A (for HAC)}-(i-iv)
and \ref{Assumption A - Dependence}. Let Assumptions \ref{Assumption 3 VR}-\ref{Assumption 4 in VR}
hold and $n_{\delta}=O(T_{n}^{1/2+\zeta})$ where $\zeta\in\left(0,\,1/2\right)$
such that $T_{n}^{\zeta}b_{1,T}^{1/2}\rightarrow0$ and $T_{n}^{\zeta}(\widehat{b}_{1,T})^{1/2}\rightarrow0$.
Then, we have:
(i) Under Assumption \ref{Assumption 6 VR}, $\mathbb{P}_{\delta}(|t_{\mathrm{DM},\mathrm{NW87}}|>z_{\alpha})\rightarrow0$$.$
If Assumption \ref{Assumption 6 VR} is replaced by Assumption \ref{Assumption 7 VR}
with $q=1/3$, then $|t_{\mathrm{DM},\mathrm{NW87}}|=O_{\mathbb{P}}(T_{n}^{\zeta-1/6})$
and $\mathbb{P}_{\delta}(|t_{\mathrm{DM},\mathrm{NW87}}|>z_{\alpha})$$\rightarrow0.$
(ii) If $b_{1,T}=T^{-1}$, then $|t_{\mathrm{DM},\mathrm{KVB}}|=O_{\mathbb{P}}(T_{n}^{\zeta-1/2})$
and $\mathbb{P}_{\delta}(|t_{\mathrm{DM},\mathrm{KVB}}|>z_{\alpha})$$\rightarrow0.$
(iii) Under Assumption \ref{Assumption K2 and b2}, $|t_{\mathrm{DM},\mathrm{DK}}|=\delta^{2}O_{\mathbb{P}}(T_{n}^{\zeta})$
and $\mathbb{P}_{\delta}(|t_{\mathrm{DM},\mathrm{DK}}|>z_{\alpha})$$\rightarrow1$.
\end{thm}
Note that Assumption \ref{Assumption 7 VR} with $q=1/3$ refers
to the MSE-optimal bandwidth for the \citeauthor{newey/west:87}'s
\citeyearpar{newey/west:87} estimator. The conditions $T_{n}^{\zeta}b_{1,T}^{1/2}\rightarrow0$
and $T_{n}^{\zeta}(\widehat{b}_{1,T})^{1/2}\rightarrow0$ mean that
the length of the regime in which $x_{t}^{(2)}$ exhibits a shift
$\delta$ in the mean increases to infinity at a slower rate than
$T$. Theorem \ref{Theorem Power DM HAR Tests} shows that when the
HAC estimators or the fixed-$b$ LRV estimators are used, the DM test
is not consistent and its power approaches zero. The theorem also
implies that the power functions corresponding to tests based on HAC
estimators lie above the power functions corresponding to those based
on fixed-$b$/EWC LRV estimators. This follows from $|t_{\mathrm{DM},\mathrm{KVB}}|\ll|t_{\mathrm{DM},\mathrm{NW87}}|.$
Another interesting feature is that $|t_{\mathrm{DM},\mathrm{NW87}}|$
and $|t_{\mathrm{DM},\mathrm{KVB}}|$ do not increase in magnitude
with $\delta$ because $\delta$ appears in both the numerator and
denominator ($\delta$ enters the denominator through the low frequency
contamination term $d^{*}$ that accounts for the bias in the HAC
and fixed-$b$ estimators (cf. Theorem \ref{Theorem ACF Nonstat})).
Part (iii) of the theorem suggests that these issues do not occur
when the DK-HAC estimator is used since the test is consistent and
its power increases with $\delta$ and with the sample size as it
should be. These results match the empirical results in Table \ref{Table Power DM Test}
discussed above, thereby confirming the relevance of Theorem \ref{Theorem Power DM HAR Tests}.
\section{\label{Section Conclusions}Conclusions}
Economic time series often display nonstationary features that are
usefully addressed in testing by allowing for some misspecification
in standard model formulations. If nonstationarity is not accounted
for properly, parameter estimates and, in particular, asymptotic LRV
estimates can be largely biased. We establish results on the low
frequency contamination induced by nonstationarity and misspecification
for the sample autocovariance and the periodogram under general conditions.
These estimates can exhibit features akin to long memory when the
data are nonstationary short memory. We show, using theoretical arguments,
that nonparametric smoothing is robust. Since the autocovariances
and the periodogram are basic elements for HAR inference, our results
allow a better understanding of LRV estimation. Under the
null hypothesis there are larger size distortions than when the data
are stationary. Under the alternative hypothesis, existing LRV estimators
tend to be inflated and HAR tests can exhibit dramatic power losses.
Long bandwidths/fixed-$b$ HAR tests suffer more from low frequency
contamination relative to HAR tests based on HAC estimators, whereas
the DK-HAC estimators do not suffer from this problem.
\section*{Supplemental Materials}
Casini, A., T. Deng and P. Perron (2024): Supplement to ``Theory
of low frequency contamination from nonstationarity and misspecification:
consequences for HAR inference\textquotedbl , Econometric Theory
Supplementary Material.
\newpage{}
\bibliographystyle{elsarticle-harv}
\bibliography{References_JoE}
\addcontentsline{toc}{section}{References}
\newpage{}
\clearpage