Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
133,904 characters · 13 sections · 158 citation commands
Prewhitened Long-Run Variance Estimation Robust to Nonstationarity
\setcounter{page}{0}
\raggedbottom
{\bf{JEL Classification}}: C12, C13, C18, C22, C32, C51\\ {\bf{Keywords}}: Asymptotic Minimax MSE, Data-dependent bandwidths, HAC, HAR, Long-run variance, Nonstationarity, Prewhitening, Spectral density.
\onehalfspacing \thispagestyle{empty}
Heteroskedasticity and autocorrelation robust (HAR) inference requires estimation of the relevant asymptotic variance or simply the long-run variance (LRV). A large literature has considered this problem. In econometrics, andrews:91 and newey/west:87 (newey/west:87; newey/west:94) extended the scope of kernel-based autocorrelation and heteroskedastic consistent (HAC) estimators of the LRV {[}see also dejong/davidson:00 and hansen:92ecma{]}. Test statistics normalized by HAC estimators follow standard asymptotic distributions under the null hypothesis under mild conditions.
It was early noted that classical HAC estimators lead to test statistics that do not correctly control the rejection rates under the null hypothesis when there is strong serial dependence in the data. A vast literature has considered this issue. Kiefer/vogelsang/bunzel:00 and Kiefer/vogelsang:02 (Kiefer/vogelsang:02; kiefer/vogelsang:05) introduced the fixed-$b$ LRV estimators for stationary sequences which are characterized by using a fixed bandwidth {[}e.g., the Newey-West/Bartlett estimator including all lags{]}. The crucial difference relative to HAC estimators is that the LRV estimator is not consistent under fixed-$b$ asymptotics and inference is nonstandard. Test statistics under the null hypotheses asymptotically follow nonstandard distributions whose critical values are obtained numerically. This has limited the use of fixed-$b$ in practice. The advantage of the fixed-$b$ framework is that it yields test statistics with more accurate null rejection rates when there is strong dependence.\footnote{See jansson:04 and sun/phillips/jin:08 for theoretical results based on asymptotic expansions. }
There is widespread evidence that the processes governing economic data are nonstationary. By nonstationary we mean non-constant moments. As in the literature, we consider processes whose sum of absolute autocovariances is finite. That is, we rule out processes with unbounded second moments (e.g., unit root). The latter can be handled by taking first-differences or applying some de-trending technique. Nonstationarity can occur for several reasons: changes in the moments induced by changes in the model parameters that govern the data (e.g., the Great Moderation with the decline in variance for many macroeconomic variables or the effects of the COVID-19 pandemic); smooth changes in the distributions of the processes that arise from transitory dynamics; and so on. HAR inference requires the estimation of the LRV of some relevant process, $V_{t}$ say.\footnote{For example, in the linear regression model $V_{t}=x_{t}e_{t}$ where $x_{t}$ is a vector of regressors and $e_{t}$ is a disturbance. } We first analyze the case with $\mathbb{E}(V_{t})=0$ for all $t$, since it is the leading case that applies under the null hypothesis. This will allow us to derive useful properties to construct bandwidths (and so on) to have tests with the correct null rejection rates. Thus, under the null hypothesis, nonstationary occurs through time-varying autocovariances $\mathbb{E}(V_{t}V_{t-k})$. We recognize that in some cases, the null hypothesis may involve a non-constant mean (e.g., when the model is misspecified). As in the literature, we do not address this case since the results can only be obtained on a case by case basis. Under the alternative hypothesis, $\mathbb{E}(V_{t})\neq0$, and $\mathbb{E}(V_{t})$ as well as $\mathbb{E}(V_{t}V_{t-k})$ can be time-varying. In most HAR inference problems the leading case is with a non-zero mean. The literature has so far not properly addressed this leading case. Our aim is to devise a method for this leading case that delivers useful estimates such that the tests have good power. Hence, we shall also consider the properties of our estimator when the mean of $V_{t}$ is non-zero and show that it leads to test having good monotonic power, unlike what is available in the literature.
The objective of this paper is to propose an estimator of the LRV that has the following properties: (i) it can be used for any hypothesis testing problem both within and outside the linear regression model and is valid under both stationarity and nonstationarity; (ii) it can be used without the need to develop further asymptotic analyses to determine the null limiting distribution of the test statistics; (iii) it leads to tests that have accurate null rejection rates even with strong dependence; such tests are consistent in any hypothesis testing problem, and in particular, in testing problems characterized by a nonstationary alternative hypothesis.\footnote{By nonstationary alternative hypothesis we mean alternative hypothesis such that $\mathbb{E}(V_{t})$ is time-varying.} None of the existing procedures satisfies all three properties. Fixed-$b$ methods rely on nonstandard limit theory and require one to derive the null limiting distribution on a case-by-case basis.\footnote{lazarus/lewis/stock:17 pointed out the usefulness for empirical work of having test statistics that follow asymptotically standard distributions rather than nonstandard distributions whose critical value has to be obtained by simulations.} casini_fixed_b_erp showed that the original fixed-$b$ methods are not theoretically valid under nonstationarity since the null limiting distribution of the test statistics is then not pivotal. More recently, a variant of the fixed-$b$ approach {[}see, e.g., sun:14 and lazarus/lewis/stock/watson:18{]} considered the use of small-$b$ asymptotics (i.e., small-bandwidths) in conjunction with fixed-$b$ critical values. In general, the latter methods do not satisfy (i)-(ii) since they use fixed-$b$ critical values, and we show below that they may not lead to consistent tests under nonstationary alternative hypothesis. Traditional HAC estimators satisfy (i)-(ii) since they are consistent for the LRV so that a test statistic studentized by an HAC estimator follows asymptotically a standard distribution. A long-lasting problem with HAC estimators is that they lead to HAR tests that can be oversized when there is strong dependence. To address this issue, andrews/monahan:92 proposed the prewhitened HAC estimators which substantially reduce the oversized problem under stationarity with HAR tests having null rejection rates similar to those of recent methods based on fixed-$b$ {[}e.g., the EWP and EWC methods of lazarus/lewis/stock:17 and \textcolor{MyBlue}{Lazarus et al.} lazarus/lewis/stock/watson:18, respectively{]}. However, we show theoretically that existing prewhitened and non-prewhitened LRV estimators lead to HAR tests that are not consistent in contexts characterized by nonstationary alternative hypotheses. This has been a recurrent problem in the time series econometrics literature.\footnote{Simulation evidence of serious (e.g., non-monotonic) power problems was documented by altissimo/corradi:2003, casini_CR_Test_Inst_Forecast, casini/perron_Lap_CR_Single_Inf (casini/perron_Oxford_Survey, casini/perron_Lap_CR_Single_Inf, casini/perron_SC_BP_Lap), chan:2020, chang/perron:18, crainiceanu/vogelsang:07, demetrescu/salish:2020, deng/perron:06, juhl/xiao:09, kim/perron:09, martins/perron:16, otto/breitung:2021, perron:1991, perron/yamamoto:18, shao/zhang:2010, vogeslang:99 and zhang/lavitas:2018 among others.} It occurs, for instance, when using tests involving structural breaks based on estimating the model under the null hypothesis; e.g., tests for forecast evaluation {[}e.g., diebold/mariano:95, giacomini/white:06 and west:96{]}, tests for forecast instability {[}cf. casini_CR_Test_Inst_Forecast, giacomini/rossi:09 (giacomini/rossi:09, giacomini/rossi:10) and perron/yamamoto:18{]}, CUSUM tests for structural change {[}see, e.g., brown/durbin/evans:1975 and ploberger/kramer:1992{]} tests and inference in time-varying parameters models {[}e.g., cai:07 and chen/hong:12{]}, tests and inference for regime switching models {[}e.g., hamilton:89 and qu/zhuo:2020{]}.
To improve the power properties of HAR tests based on HAC estimators, casini_hac proposed to modify the HAC estimators by adding a second kernel which applies smoothing over time. Such double kernel HAC estimators (DK-HAC) satisfy properties (i)-(iii) except that they can be oversized when there is high serial correlation. We introduce a novel nonparametric nonlinear VAR prewhitening procedure to apply prior to constructing the DK-HAC estimators. The key property is that our prewhitening procedure is applied locally in time through nonparametric time smoothing. This allows us to account flexibly for the time-varying second-order properties of the data and to reduce the asymptotic bias arising from nonparametric estimation. Our prewhitening is robust to nonstationarity unlike previous prewhitened procedures {[}e.g., andrews/monahan:92, preinerstorfer:17, rho/shao:13 and xiao/linton:02{]}. The latter are sensitive to estimation errors in the whitening step when there is nonstationarity in the autoregressive dynamics. For example, with AR(1) prewhitening the resulting LRV estimator is given by $\widehat{J}_{\mathrm{HAC},\mathrm{pw}}=\widehat{J}_{\mathrm{HAC},V^{*}}/(1-\widehat{a}_{1})^{2}$ where $\widehat{a}_{1}$ is the estimated parameter in the regression $V_{t}=a_{1}V_{t-1}+V_{t}^{*}$ involving the process of interest $\{V_{t}\}$ and $\widehat{J}_{\mathrm{HAC},V^{*}}$ is a classical HAC estimator applied to the prewhitened residuals $\{V_{t}^{*}\}$. Under nonstationarity in $\{V_{t}\}$, $\widehat{a}_{1}$ is biased toward one, {[}cf. perron:89{]}. This makes the recoloring step unstable as $(1-\widehat{a}_{1})^{2}$ approaches zero and more so as the nonstationarity increases. Hence, $\widehat{J}_{\mathrm{HAC},\mathrm{pw}}$ is inflated and test statistics lose power.
The consistency, rate of convergence and MSE of the new prewhitening procedure are established under nonstationarity using the segmented locally stationary framework. We then establish the consistency, rate of convergence and minimax MSE bounds for the DK-HAC estimator under general nonstationarity (i.e., unconditionally heteroskedastic processes) and discuss how these results can be used to show that the prewhitened DK-HAC estimators are valid under general nonstationarity. The new minimax MSE bounds generalizes the MSE bounds in andrews:91 as follows. andrews:91 expressed the bounds in terms of the distributions of two different second-order stationary processes. The two distributions provide upper and lower bounds, respectively, to the autocovariances of the nonstationary processes in some class. We show that this class can be enlarged substantially if the two distributions are taken to be those of some nonstationary processes that satisfy segmented locally stationarity. This allows for more variability of $\mathbb{E}(V_{t}V_{t-k})$ and serial dependence of $\left\{ V_{t}\right\} $. Thus, our bounds apply to a richer class of processes. The new bounds also provide information on how nonstationarity influences the estimation bias.
The paper makes several theoretical contributions to the HAR inference literature. First, it establishes the consistency and MSE-optimality under nonstationarity of a local prewhitening procedure applied to the double-kernel HAC estimator. Most of the existing literature focused on the stationary case, (i.e., $\mathbb{E}\left(V_{t}V_{t-k}\right)$ depends on $k$ but not on $t$), and considered a typical LRV estimator that applies smoothing only over lagged sample autocovariances. We allow $\mathbb{E}\left(V_{t}V_{t-k}\right)$ to depend on $t$ as well as $k$ and consider a prewhitened LRV estimator that applies non-parametric smoothing over lagged sample autocovariances and time. We establish the theoretical properties of the prewhitened LRV estimator using data-dependent bandwidths that flexibly account for nonstationarity unlike previously proposed data-dependent bandwidths. Second, we show that imposing restrictions on nonstationarity allows one to obtain superior minimax MSE bounds relative to those obtained under stationarity. The usefulness of these bounds is twofold. On the one hand, they allow the construction of data-dependent bandwidths that lead to a more efficient LRV estimator. On the other hand, they are used to show the validity of the (prewhitened) LRV estimator under general forms of nonstationarity (e.g., more general than segmented locally stationarity).
The prewhitened DK-HAC estimators lead to HAR tests with null rejection rates close to the nominal even with strong dependence. Furthermore, we show theoretically that the prewhitened DK-HAC estimators lead to HAR tests that are consistent even under nonstationary alternative hypotheses whereas existing HAC-based and fixed-$b$-based HAR tests are not consistent with their power converging to zero as nonstationarity increases. The simulations demonstrate that these theoretical results provide accurate predictions about the finite-sample behavior of the tests.
The paper is organized as follows. Section (ref) introduces the nonlinear VAR prewhitening procedure and its asymptotic results are established in Section (ref). Section (ref) establishes the theoretical validity of the DK-HAC estimators under general nonstationarity and presents new minimax MSE bounds. Section (ref) presents some theoretical results about the power of HAR tests under nonstationary alternative hypotheses. Section (ref) presents the simulation results. Section (ref) concludes. The supplemental materials {[}cf. casini/perron_PrewhitedHAC_Supp{]} contain all mathematical proofs.
Suppose $\left\{ V_{t}\right\} _{t=1}^{T}$ is defined on an abstract probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$, where $\Omega$ is the sample space, $\mathscr{F}$ is the $\sigma$-algebra and $\mathbb{P}$ is a probability measure. HAR inference requires the estimation of asymptotic variances of the form $J\triangleq\mathrm{lim}_{T\rightarrow\infty}J_{T}$ where
with $V_{t}(\beta)$ a random $p$-vector for each $\beta\in\Theta\subset\mathbb{R}^{p_{\beta}}$ and $\mathbb{E}(V_{t}(\beta_{0}))=0$ for all $t$ under the null hypothesis provided that the underlying model is correctly specified. We allow for $\mathbb{E}(V_{t})\neq0$ in Section (ref) when we analyze the theoretical properties of the power of the tests. For the linear regression model $y_{t}=x'_{t}\beta_{0}+e_{t}$, we have $V_{t}(\beta_{0})=x_{t}e_{t}.$ More generally, in nonlinear dynamic models, we have under mild conditions, \[ (B_{T}J_{T}B_{T})^{-1/2}\sqrt{T}(\widehat{\beta}-\beta_{0})\overset{d}{\rightarrow}\mathscr{N}(0,\,I_{p_{\beta}}), \] where $B_{T}$ is a nonrandom $p_{\beta}\times p$ matrix. Often it is easy to construct estimators $\widehat{B}_{T}$ such that $\widehat{B}_{T}-B_{T}\overset{\mathbb{P}}{\rightarrow}0$. Thus, one needs a consistent estimator of $J=\lim_{T\rightarrow\infty}J_{T}$ to construct a consistent estimator of $\lim_{T\rightarrow\infty}B_{T}J_{T}B'_{T}.$ Our goal is to consider the estimation of $J$ under nonstationarity.
Under nonstationarity the autocovariance of $V_{t}$ depends on the calendar time at which it is computed in addition to the lag. That is, $\Gamma_{u}\left(k\right)\triangleq\mathbb{E}(V_{Tu}V'_{Tu-k})$ where $u=t/T$ for some lag $k\in\mathbb{Z}$. The rescaled time index $u\in\left[0,\,1\right]$ is introduced because under nonstationarity we use the infill asymptotics. We now define the local spectral density of $V_{t}$ at time $u$ and frequency $\omega,$ $f\left(u,\,\omega\right)$. It is an important quantity because it summarizes the second-order properties of $V_{t}$. It is defined as the squared modulus of the transfer function $A\left(u,\,\omega\right)$ where the latter appears in the spectral representation of $V_{t}$ {[}see eq. (ref) in the supplement{]}. That is, $f\left(u,\,\omega\right)=|A\left(u,\,\omega\right)|^{2}$. The local spectral density can also be defined implicitly from the definition of $c\left(u,\,k\right)$ which is the approximation to the local autocovariance $\Gamma_{u}\left(k\right)$ where
and $i=\sqrt{-1}.$ In fact, Lemma S.A.1 in casini_hac showed that, under the assumptions we introduce below, $\Gamma_{u}\left(k\right)=c\left(u,\,k\right)+O\left(T^{-1}\right)$ where $O\left(T^{-1}\right)$ is the error due to the infill asymptotic approximation. Eq. (ref) relates the local autocovariance of $\{V_{t}\}$ at rescaled time $u$ and lag $k$ to its local spectral density at $u$. Thus, the nonstationary properties of $\{V_{t}\}$, which are reflected in the time-varying behavior of the autocovariance function $\Gamma_{u}\left(k\right)$, depend on the smoothness properties of $f\left(u,\,\omega\right)$ in $u$. For example, if $\{V_{t}\}$ is stationary, then $\Gamma_{u}\left(k\right)=\Gamma\left(k\right)$ for all $u$, $c\left(u,\,k\right)$ is constant in $u$, $f\left(u,\,\omega\right)=f\left(\omega\right)$ and (ref) reduces to $\Gamma\left(k\right)=\int_{-\pi}^{\pi}e^{i\omega k}f\left(\omega\right)d\omega$. These coincide with textbook definitions used under stationarity {[}see, e.g., brillinger:75{]}. If $f\left(u,\,\omega\right)$ is continuous in $u$ then $\{V_{t}\}$ is locally stationary {[}cf. dahlhaus:96{]}.\footnote{In econometrics, locally stationary processes are often referred to as time-varying parameter processes.} For example, consider a time-varying AR(1) $V_{t}=a\left(t/T\right)V_{t-1}+u_{t}$ where $u_{t}$ is a zero-mean i.i.d. process with unit variance and $a\left(\cdot\right)$ is continuous with $a\left(t/T\right)\in(-1,\,1)$ for all $t.$ Then $V_{t}$ is a locally stationary AR(1) with a local spectral density $f\left(u,\,\omega\right)$ that is continuous in $u$. We impose restrictions on the smoothness of $f\left(u,\,\omega\right)$ in $u$ which allow for considerable forms of nonstationarity in $\left\{ V_{t}\right\} $ including most of the nonstationary models used in econometrics.\footnote{A function $g\left(\cdot\right):\,\left[0,\,1\right]\mapsto\mathbb{R}$ is said to be piecewise (Lipschitz) continuous if there exists a finite subdivision $\left\{ x_{0},\,x_{1},\ldots,\,x_{n}\right\} $ of $\left[0,\,1\right]$ where $x_{0}=0$ and $x_{n}=1$, such that for all $i\in\left\{ 1,\,2,\ldots,\,n\right\} $ $g$ is (Lipschitz) continuous on $\left(x_{i-1},\,x_{i}\right)$. }
Assumption (ref) implies that $\left\{ V_{t}\right\} $ is segmented locally stationary (SLS) (see Definition (ref) in the supplement). It is similar to Assumption 3.1 in casini_hac where the latter imposes smoothness conditions on the transfer function $A\left(u,\,\omega\right)$ whereas here we directly make assumptions on the local spectral density $f\left(u,\,\omega\right)$. The class of SLS processes allows for relevant features such as structural change, regime switching-type and threshold model and includes general time-varying parameter processes, locally stationary processes and stationary processes.\footnote{For general time-varying parameter processes we mean linear and nonlinear processes whose parameters can change smoothly as well as abruptly. See Example 2.1 in casini_hac for some examples.} Assumption (ref) requires $f\left(u,\,\cdot\right)$ to be twice differentiable at the continuity points and left-differentiable at the discontinuity points. The zero-mean assumption holds under the null hypothesis. To focus on the main intuition, we first consider the case of SLS processes and then extend the results to general nonstationarity processes in Section (ref).\footnote{For general nonstationarity we mean a process with a time-varying spectral density that does not satisfy piecewise Lipschitz continuity.} The latter require more technical notations and assumptions. In Section (ref) we present the prewhitening DK-HAC estimator while in Section (ref) we discuss its data-dependent bandwidths.
Under Assumption (ref), the argument at the beginning of Section 2.1 in casini_hac suggests that $J=2\pi\int_{0}^{1}f\left(u,\,0\right)du$. The right-hand side can be seen as a function, say $\widetilde{f}\left(\omega\right)$, evaluated at the zero frequency $\omega=0.$ The intuition behind prewhitening is simple, though the mechanics under nonstationarity are quite different. Suppose one is estimating $\widetilde{f}\left(0\right)$ nonparametrically by averaging asymptotically unbiased estimators of $\widetilde{f}\left(\omega\right)$ at a number of points $\omega$ in a neighborhood of $0$. The flatter is the function $\widetilde{f}\left(\omega\right)$ around 0, the smaller the estimation bias. The idea is to transform the data such that the function of the transformed data, say $\widetilde{f}^{*}(\omega)$, is flatter in the neighborhood of $\omega=0$. Then, using the transformed data one can estimate $\widetilde{f}^{*}(0)$ by averaging asymptotically unbiased estimators of $\widetilde{f}^{*}(\omega)$ at points $\omega$ in the neighborhood of 0. The resulting bias should be less than that incurred by estimating $\widetilde{f}\left(0\right)$ since $\widetilde{f}^{*}(\omega)$ is flatter than $\widetilde{f}\left(\omega\right)$. Finally, one can apply the inverse of the transformation from $\widetilde{f}\left(\omega\right)$ to $\widetilde{f}^{*}(\omega)$ to obtain an estimator of $\widetilde{f}\left(\omega\right)$ from the estimator of $\widetilde{f}^{*}(\omega)$. This is how it works under stationarity. However, under nonstationarity one applies both the transformation and the inverse transformation locally in time, otherwise the prewhitening procedure may be unreliable as nonstationarity induces an additional source of bias in both the transformation and its inverse.
The proposed prewhitening procedure is based on the following three steps:
Step 1 (whitening step): Divide the sample in $\left\lfloor T/n_{T}\right\rfloor $ blocks, each with $n_{T}$ observations. Let $\widehat{V}_{t}=V_{t}(\widehat{\beta})$, where $\widehat{\beta}$ is a $\sqrt{T}$-consistent estimator of $\beta_{0}$. For each block $r=0,\ldots,\,\left\lfloor T/n_{T}\right\rfloor $, run the following VAR$(p_{A})$,
where $\widehat{A}_{r,j}$ for $j=1,\ldots,\,p_{A}$ are $p\times p$ least-squares estimators and $\widehat{V}_{t}^{*}=V_{t}^{*}(\widehat{\beta})$ are the prewhitened residuals. The VAR in (ref) is used to “soak up” some of the serial dependence in $\widehat{V}_{t}$ and to leave one with residuals $\{\widehat{V}_{t}^{*}\}$ that are closer to white noise.\footnote{Since the residuals $\{\widehat{V}_{t}^{*}\}$ are closer to a white noise process, they have a flatter spectral density at $\omega=0$ than $\{\widehat{V}_{t}\}$ because a white noise process has a flat spectral density.} That is why it is called “whitening step”.
Step 2 (recoloring step): Take the prewhitened residuals $\widehat{V}_{t}^{*}$, transform them by applying an inverse transformation $\widehat{V}_{t}^{*}\mapsto\widehat{V}_{D,t}^{*}=\widehat{D}_{t}\widehat{V}_{t}^{*}$ where $\widehat{D}_{t}=(I_{p}-\sum_{j=1}^{p_{A}}\widehat{A}_{D,t,j})^{-1}$ with $\widehat{A}_{D,t,j}=\widehat{A}_{r,j}$ for $t=rn_{T}+1,\ldots,\,\left(r+1\right)n_{T}$. This implies that the transformed residuals $\widehat{V}_{D,t}^{*}$ have been “recolored” (i.e., the dependence has been added back). Note that the matrix $\widehat{D}_{t}$ is the same for all $t$ in a given block. In this way the appropriate amount of dependence is added back, i.e., no contamination from possibly different strengths of dependence occurring in other blocks.
Step 3 (prewhitened DK-HAC estimation): Construct the prewhitened DK-HAC estimator $\widehat{J}_{\mathrm{pw},T}$ using $\widehat{V}_{D,t}^{*}$:
with $K_{1}\left(\cdot\right)$ a real-valued kernel in the class $\boldsymbol{K}_{3}$ defined below, $\widehat{b}_{1,T}^{*}$ is a data-dependent bandwidth sequence to be discussed below, $n_{T}\rightarrow\infty$, and
$K_{2}^{*}$ being a kernel, $\widehat{b}_{2,T}^{*}$ a data-dependent bandwidth sequence to be defined below.
In order to guarantee positive semi-definiteness, one needs to use a data taper or, e.g., for $k\geq0$,
see casini_hac.
In Step 1 the last block is $t=\left\lfloor T/n_{T}\right\rfloor n_{T}+1,\ldots,\,T$. The order of the VAR, $p_{A},$ can potentially change across blocks but, for notational ease, we assume it is the same for each $r$. The choices of $n_{T}$ and how to optimally split the sample depend on the property of the spectrum of $\{\widehat{V}_{t}\}$. A test for breaks versus smooth changes in the spectrum of $\{\widehat{V}_{t}\}$ is introduced in casini:change-point-spectra. The latter could be employed here to efficiently determine the sample-splitting. This would result in the sample being split in blocks with the property that within each block $\{\widehat{V}_{t}\}$ is locally stationary. However, this is not required for the theoretical validity. The least-squares estimation within blocks yields consistent estimators $\widehat{A}_{r,j}$ for some $A_{r,j}$ even when the fitted VAR is not the true model. The fitted VAR is used only to yield residuals $\{\widehat{V}_{t}^{*}\}$ that are closer to white noise so that their spectral density at zero is flatter, implying less asymptotic bias when estimating it nonparametrically.
Below we assume that $\widehat{A}_{r,j}\overset{\mathbb{P}}{\rightarrow}A_{r,j}$ for some $A_{r,j}\in\mathbb{R}^{p\times p}$ for all $r$ and $j$ which follow from standard arguments. For $K_{1}$ we suggest using the Quadratic Spectral (QS) kernel \[ K_{\mathrm{1}}^{\mathrm{QS}}\left(x\right)=\left(25/\left(12\pi^{2}x^{2}\right)\right)\left[\frac{\sin\left(6\pi x/5\right)}{6\pi x/5}-\cos\left(6\pi x/5\right)\right], \] and for $K_{2}$ a quadratic-type kernel {[}cf. epanechnikov:69{]} given by $K_{2}\left(x\right)=6x\left(1-x\right),\,0\leq x\leq1$. These kernels are optimal under an MSE criterion {[}see casini_hac{]}.
There has been some recent works on LRV estimation in statistics that relate to ours. kawka:2020 studied the asymptotic properties of classical spectral estimators for a linear time-varying AR process where the AR coefficients can have a finite number of discontinuities. Since classical spectral estimators do not involve any local smoothing over time, and since he focused on linear processes and did not consider data-dependent bandwidths, his framework required simpler assumptions. He also considered an estimate of the spectrum profile which is defined similarly to the variance profile of cavaliere/taylor:2007. That is based on a recursive estimate of the spectral density which is, however, different from applying local smoothing. The local smoothing is important to better account for nonstationarity as shown in casini/perron_Low_Frequency_Contam_Nonstat:2020. potiron/mykland:2020 showed that in the context of estimation of higher powers of volatility for high frequency data the local smoothing can lead to substantial efficiency gains. Although our setting is complicated by serial dependence and the fact that the class of estimators has a slower rate of convergence than the parametric $\sqrt{T}$-rate, the theoretical results on the power of the HAR tests below suggest that the local smoothing yields more powerful tests. In addition, casini/perron_Low_Frequency_Contam_Nonstat:2020 showed that under nonstationarity the sample autocovariance can be upward biased asymptotically relative to the integrated local sample autocovariance, both for fixed lag $k$ and for $k\rightarrow\infty.$ An alternative way to deal with a time-varying mean has been considered by chan:2020 (chan:2022, chan:2020) who proposed a LRV estimator which uses difference-based statistics that combine local smoothing and lagged differences of the series. His results confirmed that the local smoothing is important to enhance efficiency. However, he required covariance stationarity and did not study the theoretical properties of HAR tests normalized by the proposed LRV estimator.
For data-dependent bandwidths, we use plug-in estimates of the optimal value that minimizes some MSE criterion, see Section (ref) and casini_hac. Let $\Gamma_{D,u}\left(k\right)=\mathrm{Cov}(V_{D,Tu}^{*},\,V_{D,Tu-k}^{*})$ and $C_{pp}=\sum_{j=1}^{p}\sum_{l=1}^{p}\iota_{j}\iota_{l}'\otimes\iota_{l}\iota_{j}'$, where $\iota_{i}$ is the $i$-th elementary $p$-vector. The notation $W$ and $\widetilde{W}$ are used for some $p^{2}\times p^{2}$ weight matrices. Let $F(K_{2})\triangleq\int_{0}^{1}K_{2}^{2}\left(x\right)dx,$ $H\left(K_{2}\right)\triangleq(\int_{0}^{1}x^{2}K_{2}\left(x\right)dx)^{2}$,
where $c_{D}^{*}\left(u,\,l\right)=\mathrm{Cov}(V_{D,Tu}^{*},\,V_{D,Tu-l}^{*})$, $V_{D,t}^{*}=D_{t}V_{t}^{*}$,
The optimal $b_{2,T}$ is given by {[}see casini_hac{]} \[ b_{2,T}^{\mathrm{opt,*}}\left(u\right)=[H\left(K_{2}^{\mathrm{}}\right)D_{1,D}\left(u\right)]^{-1/5}\left(F\left(K_{2}\right)\left(D_{2,D}\left(u\right)\right)\right)^{1/5}T^{-1/5}. \] Let
$K_{1,q}<\infty$ if and only if $K_{1}\left(x\right)$ is $q$ times differentiable at zero. Let $f_{D}^{*}\left(u,\,\omega\right)=\sum_{k=-\infty}^{\infty}c_{D}^{*}\left(u,\,k\right)e^{-i\omega k}$ and define the index of smoothness of $f_{D}^{*}\left(u,\,\omega\right)$ at $\omega=0$ by $f_{D}^{*\left(q\right)}\left(u,\,0\right)\triangleq\left(2\pi\right)^{-1}\sum_{k=-\infty}^{\infty}\left|k\right|^{q}c_{D}^{*}\left(u,\,k\right)$. Let
The optimal $b_{1,T}$ given the optimal value $b_{2,T}^{\mathrm{opt,*}}$ is given by {[}see casini_hac{]}, \[ b_{1,T}^{\mathrm{opt,*}}=(2qK_{1,q}^{2}\phi_{D}\left(q\right)T\overline{b}_{2,T}^{\mathrm{opt}}/\left(\smallint K_{1}^{2}\left(y\right)dy\smallint K_{2}^{2}\left(x\right)dx\right))^{-1/\left(2q+1\right)}, \] with $\overline{b}_{2,T}^{\mathrm{opt,*}}=\int_{0}^{1}b_{2,T}^{\mathrm{opt},*}\left(u\right)du$. For the QS kernel, $q=2$, $K_{1,2}=1.421223$, and $\int K_{1}^{2}\left(x\right)dx=1$. For the optimal $K_{2}^{\mathrm{}}$ we have $H(K_{2}^{\mathrm{opt}})=0.09$ and $F(K_{2}^{\mathrm{opt}})=1.2$.
The bandwidths $(b_{1,T}^{\mathrm{opt,*}},\,\overline{b}_{2,T}^{\mathrm{opt,*}})$ are optimal under a sequential MSE criterion that determines the optimal $b_{1}$ as a function of the optimal $b_{2}\left(u\right)$. Thus, the latter influences the former but not vice-versa. However, this has the advantage that the optimal $b_{2}\left(u\right)$ is allowed to change over time. \textcolor{MyBlue}{Belotti et al.} belotti/casini/catania/grassi/perron_HAC_Sim_Bandws proposed an alternative criterion that determines the optimal $b_{1}$ and $b_{2}$ that jointly minimize the global MSE. The latter yields an optimal $b_{2}$ that does not depend on $u$ and so it does not perform as well as the sequential method when the data is far from stationary.
In order to construct a data-dependent bandwidth for $b_{2,T}\left(u\right)$, we need consistent estimators of $D_{1,D}\left(u\right)$ and $D_{2,D}\left(u\right)$. We set $\widetilde{W}^{\left(r,r\right)}=p^{-1}$ for all $r$ which corresponds to the normalization used below for $W$. In order to replace $D_{1,D}\left(u\right)$ we make a parametric assumption and estimate $D_{1,D}\left(u\right)$ under this assumption. Following casini_hac, the approximating parametric assumption is that $V_{D,t}^{*}$ belongs to the class of class of locally stationary first-order autoregressive (AR(l)) models with certain restrictions on the smoothness of the parameters. Under this approximating parametric assumption, the estimator of $D_{1,D}\left(u\right)$ is
where $\left[S_{\omega}\right]$ is the cardinality of $S_{\omega}$ and $\omega_{s+1}>\omega_{s}$ with $\omega_{1}=-\pi,\,\omega_{\left[S_{\omega}\right]}=\pi.$ We set $S_{\omega}=\{-\pi,\,-3,\,-2,\,-1,\,0,\,1,\,2,\,3,\,\pi\}$. The estimator of $D_{2,D}\left(u\right)$ is given by \[ \widehat{D}_{2,D}\left(u_{0}\right)\triangleq2p^{-1}\sum_{r=1}^{p}\sum_{l=-\left\lfloor T^{4/25}\right\rfloor }^{\left\lfloor T^{4/25}\right\rfloor }\left(\widehat{c}_{D,T}^{*,\left(r,r\right)}\left(u_{0},\,l\right)\right)^{2}, \] where the number of summands grows at the same rate as the inverse of the optimal bandwidth $b_{1,T}^{\mathrm{opt,*}}$. Hence, the estimator of the optimal bandwidth $b_{2,T}^{\mathrm{opt},*}$ is given by
The data-dependent bandwidth parameter $\widehat{b}_{1,T}^{*}$ is then defined as follows. First, one specifies $p$ univariate approximating parametric models given by $\{V_{D,t}^{*\left(r\right)}\}$ for $r=1,\ldots,\,p$. Second, one estimates the parameters of the approximating parametric model by least-squares. Third, one substitutes these estimates into $\phi_{D}\left(q\right)$ with the estimate denoted by $\widehat{\phi}_{D}\left(q\right)$. This yields the data-dependent bandwidth parameter
For the QS kernel, we have $\widehat{b}_{1,T}^{*}=0.6828(\widehat{\phi}_{D}\left(2\right)T\widehat{\overline{b}}_{2,T}^{*})^{-1/5}$. As mentioned above, the suggested approximating parametric models are the locally stationary AR(l) models given by $V_{D,t}^{*\left(r\right)}=a_{1}^{\left(r\right)}\left(t/T\right)$ $V_{D,t-1}^{*\left(r\right)}+u_{t}^{\left(r\right)}$, $r=1,\ldots,\,p$. Let $\widehat{a}_{1}^{\left(r\right)}\left(u\right)$ and $(\widehat{\sigma}^{\left(r\right)}\left(u\right))^{2}$ be the least-squares estimators of the autoregressive and innovation variance parameters computed using data close to $u=t/T$:
where $n_{2,T}\rightarrow\infty$.\footnote{See, for example, Dahlhaus/Giraitis:98 for a discussion about nonparametric local parameter estimates in the context of locally stationary time series.} These are simply least-squares estimators based on rolling windows. Then, for $q=2$, we have
where $W^{\left(r,r\right)},\,r=1,\ldots,\,p$ are pre-specified weights and $n_{3,T}\rightarrow\infty$. The usual choice for $W^{\left(r,r\right)}$ is one for all $r$ except that which corresponds to an intercept in which case it is zero. Let $\widehat{\theta}=(\int_{0}^{1}\widehat{a}_{1}^{\left(1\right)}\left(u\right)du,\,\int_{0}^{1}(\widehat{\sigma}^{\left(1\right)}\left(u\right))^{2}du,\ldots,\,\int_{0}^{1}\widehat{a}_{1}^{\left(p\right)}\left(u\right)du,\,\int_{0}^{1}(\widehat{\sigma}^{\left(p\right)}\left(u\right))^{2}du)'$ and let $\theta^{*}$ denote the probability limit of $\widehat{\theta}$. If the locally stationary AR(1) parametric model is not correctly specified for $V_{D,t}^{*}$, then the probability limit of $\widehat{\phi}_{D}\left(q\right)$ need not be equal to $\phi_{D}\left(q\right)$. Let $\phi_{\theta^{*}}\in\mathbb{R}$ be the probability limit of $\widehat{\phi}_{D}\left(q\right)$ (i.e., $\widehat{\phi}_{D}\left(q\right)-\phi_{\theta^{*}}=o_{\mathbb{P}}\left(1\right))$. When the locally stationary AR(1) parametric model is correctly specified we have $\phi_{\theta^{*}}=\phi_{D}\left(q\right)$.
In this section, we analyze the asymptotic properties of $\widehat{J}_{\mathrm{pw},T}$ for the case with $\mathbb{E}\left(V_{t}\right)=0$ for all $t$, which is relevant under the null hypothesis provided that the model is correctly specified. Let $K$ denote a generic kernel and $K^{(q)}$ be defined as $K_{1,q}$ in (ref) with $K_{1}$ replaced by $K$. Let
Note that $\boldsymbol{K}_{3}$ depends on $T$, though we omit this dependence. $\boldsymbol{K}_{3}$ contains commonly used kernels, e.g., QS, Bartlett, Parzen, and Tukey-Hanning, with the exception of the truncated kernel. For the QS, Parzen, and Tukey-Hanning kernels, $q=2$. For the Bartlett kernel, $q=1$. The condition $q<17/2$ in part (iv) is a technical condition needed to control the deviation $|\widehat{b}_{1,T}^{*}-b_{\theta_{1},T}|$, where $b_{\theta_{1},T}$ is defined as $\widehat{b}_{1,T}^{*}$ {[}cf. (ref) below{]} but with $\widehat{\phi}_{D}\left(q\right)$ replaced by $\phi_{\theta^{*}}$.
For $K_{2}$ we consider the same class of kernels $\boldsymbol{K}_{2}$ as considered by casini_hac:
We define \[ \mathrm{MSE}(Tb_{1,T}b_{2,T},\,\widetilde{J}_{T},\,J_{T},\,W)=Tb_{1,T}b_{2,T}\mathbb{E}[\mathrm{vec}(\widetilde{J}_{T}-J_{T})'W\mathrm{vec}(\widetilde{J}_{T}-J_{T})]. \] We need to impose conditions on the temporal dependence of $\{V_{t}\}$. Let
where $\{V_{\mathscr{N},t}\}$ is a Gaussian sequence with the same mean and covariance structure as $\left\{ V_{t}\right\} $. $\kappa_{V,t}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}\left(u,\,v,\,w\right)$ is the time-$t$ fourth-order cumulant of $(V_{t}^{\left(a_{1}\right)},\,V_{t+u}^{\left(a_{2}\right)},\,V_{t+v}^{\left(a_{3}\right)},$ $\,V_{t+w}^{\left(a_{4}\right)})$ while $\kappa_{\mathscr{N}}^{\left(a_{1},a_{2},a_{3},a_{4}\right)}$ $(t,\,t+u,\,t+v,\,t+w)$ is the time-$t$ centered fourth moment of $V_{t}$ if $V_{t}$ were Gaussian. Let $\lambda_{\max}\left(A\right)$ denote the largest eigenvalue of the matrix $A$.
If $\left\{ V_{t,T}\right\} $ is stationary then the cumulant condition of Assumption (ref)-(i) reduces to the standard one used in the time series literature {[}see, e.g., Assumption A in andrews:91{]}. We do not require fourth-order stationarity but only that the time-$t=Tu$ fourth order cumulant is locally constant in a neighborhood of a continuity point $u$. As explained in casini_hac, using an argument similar to that used in Lemma 1 in andrews:91, one can show that $\alpha$-mixing and moment conditions imply that the cumulant condition of Assumption (ref)-(i) holds. Part (ii) essentially requires that the approximating cumulant function $\widetilde{\kappa}_{a_{1},a_{2},a_{3},a_{4}}(u,\,k,$ $\,s,\,l)$ satisfies similar smoothness restrictions as $f\left(u,\,\cdot\right)$ (i.e., twice differentiability at the continuity points and twice left-differentiable at the discontinuity points).
Assumption (ref)-(i,iii) is an extension of Assumption B in andrews:91 to a nonstationary setting. Part (i) follows from asymptotic normality of $\sqrt{T}(\widehat{\beta}-\beta_{0})$. Part (ii)-(iii) are common conditions used to obtain the asymptotic normality of $\sqrt{T}(\widehat{\beta}-\beta_{0})$ under nonstationarity. In order to obtain rate of convergence results we shall replace Assumption (ref) with the following assumption.
Assumption (ref) is needed to show that the effect of using $\widehat{\beta}$ rather than $\beta_{0}$ when constructing $\widehat{J}_{\mathrm{pw},T}$ is at most $o_{\mathbb{P}}\left(1\right)$; it is an extension of Assumption C in andrews:91. Parts (i)-(ii) of Assumption (ref) are the nonparametric analogue to Assumption E-F in andrews:91. Part (iii) is satisfied if $\left\{ V_{t}\right\} $ is strong mixing with mixing numbers that are less stringent than those sufficient for the cumulant condition in Assumption (ref)-(i). Part (iv) and (vi) extend (i)-(ii) to $\widehat{D}_{1}$ and $\widehat{D}_{2}$. Part (v) is needed to apply the convergence of Riemann sums. Under Assumption (ref) the effect of using the bandwidths $\widehat{b}_{1,T}^{*}$ and $\widehat{b}_{2,T}^{*}$ rather than $b_{\theta_{1},T}$ and $\overline{b}_{\theta_{2},T}$ (defined below in (ref)) when constructing $\widehat{J}_{\mathrm{pw},T}$ is at most $o_{\mathbb{P}}\left(1\right)$.
Given the restrictions below on $n_{T}$, Assumption (ref) is satisfied by standard nonparametric estimators. For the consistency of $\widehat{J}_{T,\mathrm{pw}},$ Assumption (ref), (ref)-(ref), (ref)-(i,iv) and (ref) are sufficient. For the rate of convergence and asymptotic MSE results additional conditions are needed. Let
where $\overline{b}_{\theta_{2},T}\triangleq\int_{0}^{1}[H\left(K_{2}\right)$ $D_{1,D}\left(u\right)]^{-1/5}\left(F\left(K_{2}\right)D_{2,D}\left(u\right)\right)^{1/5}T^{-1/5}du$. Recall that the bandwidths $\widehat{\overline{b}}_{2,T}^{*},$ $\widehat{b}_{2,T}^{*}$ and $\widehat{b}_{1,T}^{*}$ are defined by (ref), (ref) and (ref), respectively.
A result corresponding to Theorem (ref) for non-prewhitened DK-HAC estimators is established in Theorem 5.1 in casini_hac under the same assumptions with the exception of Assumption (ref). Note that for $u$ a continuity point, $f_{D}^{*}\left(u,\,\omega\right)=D\left(u,\,\omega\right)f^{*}\left(u,\,\omega\right)D\left(u,\,\omega\right)',$ where $D\left(u,\,\omega\right)=(I_{p}-\sum_{j=1}^{p_{A}}A_{D,j}\left(u\right)e^{-ij\omega})^{-1}$ with $A_{D,j}\left(u\right)=A_{D,Tu,j}+O\left(T^{-1}\right)$ and $f^{*}\left(u,\,\omega\right)$ is the local spectral density of $\left\{ V_{t}^{*}\right\} $. Since $D\left(u-k/T,\,\omega\right)=D\left(u,\,\omega\right)+O\left(T^{-1}\right)$ by local stationarity, we have
A meaningful comparison between prewhitened and non-prewhitened DK-HAC estimators $\widehat{J}_{T}$ can be made only if reasonable choices of the bandwdiths $b_{1,T}$ and $b_{2,T}$ are made. When the optimal bandwidths for $\widehat{J}_{\mathrm{pw},T}$ and $\widehat{J}_{T}$ are used we find that $\widehat{J}_{\mathrm{pw},T}$ has smaller asymptotic MSE than $\widehat{J}_{T}$ if and only if (assuming $p=1$, i.e., the scalar case, with $w_{1,1}=1$)
A numerical comparison would be tedious since the condition depends on the true data-generating process of $\{V_{t}\}$ and the VAR approximation for $\widehat{V}_{t}=V_{t}(\widehat{\beta})$. Under stationarity, grenander/rosenblatt:57 and andrews/monahan:92 considered a few examples. We can make a few observations on the difference between the condition (ref) and an analogous condition for the case with $\{V_{t}\}$ second-order stationary and $D_{s}=D=(1-\sum_{j=1}^{p_{A}}A_{j})^{-1}$ for all $s$ {[}cf. andrews/monahan:92{]}. The condition in andrews/monahan:92 is then
where the quantities $f^{q}\left(0\right)$ and $f^{*\left(q\right)}\left(0\right)$ do not depend on $u$ by stationarity. The main difference between the two conditions (ref)-(ref) is that the part involving the asymptotic variance is missing in (ref). The quantities $|f^{*\left(q\right)}\left(0\right)|D^{2}$ and $|f^{q}\left(0\right)|$ are from the asymptotic squared bias. This is a consequence of the fact that prewhitened and non-prewhitened HAC estimators have the same asymptotic variance under stationarity when the optimal bandwidths are used. This property does not hold when $\{V_{t}\}$ is nonstationary. The condition (ref) suggests instead that, in general, both the asymptotic squared bias and asymptotic variance of prewhitened and non-prewhitened HAC estimators can be different. Simulations in andrews/monahan:92 showed that this is indeed the case even under stationarity: the variance of the prewhitened HAC estimators is larger than that of the non-prewhitened HAC estimators\textemdash this feature is consistent with our theoretical results but not with theirs.
Both the smoothing over lagged autocovariances and over time influence the bias of $\widehat{J}_{\mathrm{pw},T}$. The contribution to the bias due to smoothing over lagged autocovariances is $O(b_{1,T}^{q})$ while the contribution due to smoothing over time is $O(b_{2,T}^{2})$. Note that the continuity points and the discontinuity points here induce a bias of the same order $b_{2,T}^{2}$. For the continuity points, $O(b_{2,T}^{2})$ follows from the usual argument. In the neighborhood of a discontinuity point $[\lambda_{j}^{0}-b_{2,T},\,\lambda_{j}^{0}+b_{2,T}]$, the bias of the local smoothing is $O(b_{2,T})$. However, when averaging over blocks or equivalently integrating over $u\in\left[0,\,1\right]$, this bias becomes $O(b_{2,T}^{2})$ since there are only a finite number of discontinuity points and so each discontinuity point contributes $O(b_{2,T}^{2})$ to the integrated bias. For $\widehat{J}_{\mathrm{pw},T}$ we have $(\widehat{\overline{b}}_{2,T}^{*})^{2}/(\widehat{b}_{1,T}^{*})^{q}\rightarrow0$ since $q=2.$ Thus, the bias due to smoothing over lagged autocovariances dominates the bias due to smoothing over time.
In this section we discuss the case where $\left\{ V_{t}\right\} $ is unconditionally heteroskedastic and establish new MSE bounds which we compare to existing ones. To focus on the main intuition and for comparison purposes, we consider the non-prewhitened DK-HAC estimator \[ \widehat{J}_{T}(b_{1,T},\,b_{2,T})=\sum_{k=-T+1}^{T-1}K_{1}(b_{1,T}k)\widehat{\Gamma}\left(k\right), \] where $\widehat{\Gamma}\left(k\right)$ is defined analogously to $\widehat{\Gamma}_{D}^{*}\left(k\right)$ but with $\widehat{V}_{t}$ in place of $\widehat{V}_{D,t}^{*}$. We use the new MSE bounds to show that the data-dependent bandwidths for the DK-HAC estimator are minimax MSE-optimal also under general nonstationarity. Corresponding results for the prewhitened estimator $\widehat{J}_{\mathrm{pw},T}$ can be obtained by using the results of Section (ref), though the proofs are more lengthy with no special gain in intuition.
We provide theoretical results under the assumption that $\{V_{t}\}$ is generated by some distribution $\mathscr{P}$ and so defined on the probability space $(\Omega,\,\mathscr{F},\,\widetilde{\mathbb{P}})$ where $\mathscr{P}=\widetilde{\mathbb{P}}\circ V^{-1}$, $\widetilde{\mathbb{P}}$ is different from $\mathbb{P}$ used in Section (ref)-(ref) and $V$ is a random variable that is a measurable function $V:\,\Omega\mapsto\mathbb{R}$. $\mathbb{E}_{\mathscr{P}}$ denotes the expectation taken under $\mathscr{P}$. We establish lower and upper bounds on the MSE under $\mathscr{P}$ and use a minimax MSE criterion for optimality. Define the sample size dependent spectral density of $\{V_{t}\}$ as \[ f_{\mathscr{P},T}\left(\omega\right)\triangleq\left(2\pi\right)^{-1}\sum_{k=-T+1}^{T-1}\Gamma_{\mathscr{P},T}\left(k\right)\exp\left(-i\omega k\right),\,\,\,\mathrm{for\,\,\,}\omega\in\left[-\pi,\,\pi\right], \] where
The estimand is then given by
The theoretical bounds are derived in terms of two distributions $\mathscr{P}_{w}$, $w=L,\,U$, under which $\{V_{t}\}$ is zero-mean SLS with $m_{0}+1$ regimes and satisfies Assumption (ref) and (ref) with autocovariance function $\{\Gamma_{\mathscr{P}_{w},t/T}\left(k\right)\}$. Then, $\{a'V_{t}\}$ has spectral density $f_{\mathscr{P}_{w},a}\left(\omega\right)\triangleq\int_{0}^{1}f_{\mathscr{P}_{w},a}\left(u,\,\omega\right)du,$ where
Let $\kappa_{\mathscr{P},aV,t}\left(k,\,j,\,m\right)$ denote the time-$t$ fourth-order cumulant of $(a'V_{t},\,a'V_{t+k},\,a'V_{t+j},\,a'V_{t+m})$ under $\mathscr{P}$. For two matrices $A$ and $B$, $A\leq B$ if and only if $A_{ij}\leq B_{ij}$ for all $i$ and $j$. Define
To derive the MSE bounds for a given class of general nonstationary processes one needs to impose restrictions on the autocovariance function of the processes in the class relative to the autocovariance function of some process whose second-order properties are known. This approach was also used by andrews:91 who, however, relied on stationarity. $\boldsymbol{P}_{U}$ includes all distributions such that the autocovariances of $\left\{ V_{t}\right\} $ are bounded above by those of some SLS process with distribution $\mathscr{P}_{U}$, thereby allowing considerable variability of $\Gamma_{\mathscr{P},t/T}\left(k\right)$ for given $t$ and $k.$ The set $\boldsymbol{P}_{L}$ requires the autocovariances of $\left\{ V_{t}\right\} $ to be bounded below by positive semidefinite autocovariances of some SLS process with distribution $\mathscr{P}_{L}$. Let $c_{\mathscr{P}_{w}}\left(u,\,k\right)=\int e^{i\omega k}\Gamma_{\mathscr{P}_{w},u}\left(k\right)d\omega$ denote the local autocovariance associated to the distribution $\mathscr{P}_{w},$ $w=L,\,U.$ Let
Note that $\boldsymbol{K}_{3}\subset\boldsymbol{K}_{1}$. In particular, $\boldsymbol{K}_{1}$ includes also the truncated kernel.
Consider the following generalization of Assumption (ref) and (ref):
Let $\mathrm{MSE}_{\mathscr{P}}\left(\cdot\right)$ denote the MSE of $\cdot$ under $\mathscr{P}$ and let $\boldsymbol{K}_{1,+}=\left\{ K_{1}\left(\cdot\right)\in\boldsymbol{K}_{1}:\,K_{1}\left(x\right)\geq0\,\forall x\right\} $. $\boldsymbol{K}_{1,+}$ is a subset of $\boldsymbol{K}_{1}$ that contains all kernels that are non-negative and is used for some results below. The QS kernel is not in $\boldsymbol{K}_{1,+}$. The smoothness of $f_{\mathscr{P}_{w},a}\left(u,\,\omega\right)$ at $\omega=0$ is indexed by \[ f_{\mathscr{P}_{w},a}^{\left(q\right)}\left(u,\,0\right)=\left(2\pi\right)^{-1}\sum_{k=-\infty}^{\infty}\left|k\right|^{q}a'\Gamma_{\mathscr{P}_{w},u}\left(k\right)a,\,\,\,\,\mathrm{for}\,q\in[0,\,\infty),\,w=L,\,U. \] We first consider the MSE bounds for $\widetilde{J}_{T}$ which is constructed using $V_{t}(\beta_{0})$ rather than $\widehat{V}_{t}$.
The theoretical bounds in Theorem (ref) are sharper than the ones in \textcolor{MyBlue}{Andrews (1988; 1991)} which are based on stationarity (i.e., the autocovariances that dominate the autocovariances of any $\mathscr{P}\in\boldsymbol{P}_{U}$ are assumed by andrews:91 to be those of a stationary process).\footnote{There are a couple of technical issues in Section 8 in andrews:91. In particular, the MSE bound is not correct. See casini_comment_andrews91 for details.} Given that stationarity is a special case of SLS, our bounds apply to a wider class of processes. Furthermore, they are more informative because they change with the specific type of nonstationarity unlike \textcolor{MyBlue}{Andrews'} andrews:91 bounds that depend on the spectral density of a stationary process.
The theorem is derived under $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow0$ (i.e., the bias due to smoothing over time is of smaller order than that due to smoothing over lagged autocovariances). When instead $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow\nu\in\left(0,\,\infty\right)$, there is an additional term in the bound. For example, in part (i) this term is \[ \left(\pi\nu\int_{0}^{1}x^{2}K_{2}\left(x\right)dx\int_{\mathbf{\mathbf{\widetilde{\mathbf{C}}}_{\mathscr{P}_{U}}}}\left(\partial^{2}/\partial u^{2}\right)f_{\mathscr{P}_{U},a}\left(u,\,0\right)du+2\pi\nu\Delta_{f_{\mathscr{P}_{U},a}}\left(0\right)\right)^{2}+\Xi, \] where $\mathbf{\widetilde{\mathbf{C}}}_{\mathscr{P}_{U}}$ is the set of continuity points under $\mathscr{P}_{U}$,
with $\left\{ \lambda_{j}^{0}\right\} _{j=1}^{m_{0}}$ being the discontinuity points, $m_{0}$ being a finite integer,
and $\Xi$ depends on the cross-products of the bias terms due to smoothing over time and lagged autocovariances. Some of the results of this paper are extended to the case $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow\nu\in\left(0,\,\infty\right)$ in \textcolor{MyBlue}{Belotti et al. (2021)}. \nocite{belotti/casini/catania/grassi/perron_HAC_Sim_Bandws} Thus, our bounds show how nonstationarity influences the bias-variance trade-off. They also highlight how it is affected by the smoothing over the time direction versus the autocovariance lags direction. These are important elements in order to understand the properties of HAR tests normalized by LRV estimators.
We now extend the results in Theorem (ref) to the estimator $\widehat{J}_{T}$ that uses $V_{t}(\widehat{\beta})$. The following assumptions extend Assumption (ref)-(ref) to the distribution $\mathscr{P}.$
To show the asymptotic equivalence of the MSE of $a'\widehat{J}_{T}a$ to that of $a'\widetilde{J}_{T}a$ we need an additional assumption which was also used by andrews:91. Let $|A|$ denote the vector or matrix of absolute values of the elements of $A.$ Define
Let $H_{1,T}^{\left(r\right)}$, $\widehat{\beta}^{\left(r\right)}$ and $\beta_{0}^{\left(r\right)}$ denote the $r$-th elements of $H_{1,T}$, $\widehat{\beta}$ and $\beta_{0}$, respectively, for $r=1,\ldots,\,p$.
Theorem (ref) extends the consistency, rate of convergence, MSE results of Theorem 3.2 in casini_hac. The asymptotic equivalence of the MSE implies that the bounds in Theorem (ref) apply to $\widehat{J}_{T}$ as well. The MSE equivalence is used to show that the optimal kernels and bandwidths results below apply to $\widehat{J}_{T}$ as well as to $\widetilde{J}_{T}$. Similar results can be shown for the prewhitened estimator $\widehat{J}_{T,\mathrm{pw}}$. For this case, the sets $\boldsymbol{P}_{U}$ and $\boldsymbol{P}_{L}$ would need to be defined in terms of the autocovariance function of $V_{D,t}^{*}=D_{t}V_{t}^{*}.$ The distributions $\mathscr{P}_{U}$ and $\mathscr{P}_{L}$ that form an envelope for the autocovariances of $V_{D,t}^{*}$ may depend on different prewhitening models.
We use the sequential MSE procedure that first determines the optimal $b_{2,T}\left(u\right)$ and then determines the optimal $b_{1,T}$ as function of the integrated optimal $\overline{b}_{2,T}$, see casini_hac. The results for the global MSE criterion can easily be extended using similar arguments as those used in this section.
We consider distributions $\mathscr{P}\in\boldsymbol{P}_{U,2}$ where $\boldsymbol{P}_{U,2}\subseteq\boldsymbol{P}_{U}$ is defined below. We need to restrict attention to a subset $\boldsymbol{P}_{U,2}$ of $\boldsymbol{P}_{U}$ for technical reasons related to the derivation of the optimal bandwidth $b_{2,T}^{\mathrm{opt}}\left(u\right)$. The distributions in $\boldsymbol{P}_{U,2}$ restrict the degree of nonstationarity by requiring some smoothness of the local autocovariance. This is intuitive since the optimality of $b_{2,T}^{\mathrm{opt}}\left(u\right)$ is justified under smoothness locally in time. We remark, however, that the optimality of $b_{1,T}$ and $K_{1}$ determined below holds over all distributions $\mathscr{P}\in\boldsymbol{P}_{U}.$ We show that the resulting optimal kernels are $K_{1}^{\mathrm{opt}}\left(\cdot\right)$ and $K_{2}^{\mathrm{opt}}\left(\cdot\right)$ from Section (ref).
Let $\mathbf{\widetilde{C}}_{\mathscr{\mathscr{P}}_{U}}$ denote the set of continuity points $u\in\left(0,\,1\right)$ under $\mathscr{\mathscr{P}}_{U}$. For any $a\in\mathbb{R}^{p}$ and $u_{0}\in\mathbf{\widetilde{C}}_{\mathscr{\mathscr{P}}_{U}}$ consider the following inequality,
which essentially requires that the distribution $\mathscr{\mathscr{P}}_{U}$ has locally a larger degree of nonstationarity than that of the distribution $\mathscr{P}$. We consider the following class of distributions,
Let
We now obtain the optimal $K_{1}\left(\cdot\right)$ and $b_{1,T}$ as a function of $\overline{b}_{2,T}^{\mathrm{opt}}=\int_{0}^{1}b_{2,T}^{\mathrm{opt}}\left(u\right)du$ and $K_{2}^{\mathrm{opt}}\left(\cdot\right)$. For some results below, we consider a subset of $\boldsymbol{K}_{1}$ defined by $\boldsymbol{\widetilde{K}}_{1}=\{K_{1}\left(\cdot\right)\in\boldsymbol{K}_{1}|\,\widetilde{K}\left(\omega\right)\geq0\,\forall\,\omega\in\mathbb{R}\}$ where $\widetilde{K}\left(\omega\right)=\left(2\pi\right)^{-1}\int_{-\infty}^{\infty}K_{1}\left(x\right)e^{-ix\omega}dx.$ The function $\widetilde{K}\left(\omega\right)$ is referred to as the spectral window generator corresponding to the kernel $K_{1}\left(\cdot\right)$. The set $\widetilde{\boldsymbol{K}}_{1}$ contains all kernels $K_{1}$ that generate positive semidefinite estimators in finite samples. $\boldsymbol{\widetilde{K}}_{1}$ contains the Bartlett, Parzen, and QS kernels, but not the truncated or Tukey-Hanning kernels.
We adopt the notation $\widehat{J}_{T}(b_{1,T})=\widehat{J}_{T}(b_{1,T},\,b_{2,T},\,K_{2,0})$ for the estimator $\widehat{J}_{T}$ that uses $K_{2,0}\left(\cdot\right)\in\boldsymbol{K}_{2}$, $b_{1,T}$ and $b_{2,T}=\overline{b}_{2,T}^{\mathrm{opt}}+o(T^{-1/5})$ where $\overline{b}_{2,T}^{\mathrm{opt}}=\int_{0}^{1}b_{2,T}^{\mathrm{opt}}\left(u\right)du$. Let $\widehat{J}_{T}^{\mathrm{QS}}(b_{1,T})$ denote the estimator based on the QS kernel $K_{1}^{\mathrm{QS}}\left(\cdot\right)$. We then compare two kernels $K_{1}$ using comparable\textcolor{red}{ }bandwidths $b_{1,T}$ which are defined as follows. Given $K_{1}\left(\cdot\right)\in\widetilde{\boldsymbol{K}}_{1}$, the QS kernel $K_{1}^{\mathrm{QS}}\left(\cdot\right)$, and a bandwidth $\left\{ b_{1,T}\right\} $ to be used with the QS kernel, define a comparable bandwidth $\{b_{1,T,K_{1}}\}$ for use with $K_{1}\left(\cdot\right)$ such that both kernel/bandwidth combinations have the same maximum asymptotic variance over $\mathscr{P}\in\boldsymbol{P}_{U}$ when scaled by the same factor $Tb_{1,T}b_{2,T}$. This means that $b_{1,T,K_{1}}$ is such that
This definition yields $b_{1,T,K_{1}}=b_{1,T}/(\int K_{1}^{2}\left(x\right)dx).$ Note that for the QS kernel, $K_{1}^{\mathrm{QS}}\left(x\right)$, we have $b_{1,T,\mathrm{QS}}=b_{1,T}$ since $\int(K_{1}^{\mathrm{QS}}(x))^{2}dx=1$.
We now consider the asymptotically optimal choice of $b_{1,T}$ for a given kernel $K_{1}\left(\cdot\right)$ for which $K_{1,q}\in\left(0,\,\infty\right)$ for some $q$, and given $K_{2}^{\mathrm{opt}}$ and $\overline{b}_{2,T}^{\mathrm{opt}}$. We continue to use a minimax optimality criterion. However, unlike the results of Proposition (ref) and Theorem (ref), in which an optimal kernel was found that was the same for any dominating distribution $\boldsymbol{P}_{U,2}$ and $\boldsymbol{P}_{U}$, respectively, the optimal bandwidth $b_{1,T}$ depends on a scalar parameter $\phi\left(q\right)$ that is a function of $\mathscr{P}_{U}$ and $q$.
Let $w_{r},\,\,r=1,\ldots,\,p$, be a set of non-negative weights summing to one. We consider a weighted squared error loss function \[ \mathrm{L}(\widehat{J}_{T},\,J_{\mathscr{P},T})=\sum_{r=1}^{p}w_{r}(\widehat{J}_{T}^{\left(r,r\right)}(b_{1,T})-J_{\mathscr{P},T}^{\left(r,r\right)})^{2}. \] A common choice is $w_{r}=1/p$ for $r=1,\ldots,\,p$. For a given dominating distribution $\mathscr{P}_{U}$, define
where $a^{\left(r\right)}$ is a $p$-vector with the $r$-th element one and all other elements zero. For any given $\phi\left(q\right)\in\left(0,\,\infty\right)$, let $\mathscr{\boldsymbol{P}}_{U}\left(\phi\right)$ denote some set $\mathscr{\boldsymbol{P}}_{U}$ whose dominating distribution $\mathscr{P}_{U}$ satisfies (ref).
We now show that the DK-HAC estimators based on data-dependent bandwidths with similar form as $\widehat{b}_{1,T}^{*}$ and $\widehat{\overline{b}}_{2,T}^{*}$ (cf. Section (ref)) have the same asymptotic MSE properties as the estimators based on optimal fixed bandwidth sequences $b_{1,T}^{\mathrm{opt}}$ and $\overline{b}_{2,T}^{\mathrm{opt}}$ that depend on the unknown distribution $\mathscr{P}$.
We consider the data-dependent bandwidths $\widehat{b}_{1,T}$ and $\widehat{\overline{b}}_{2,T}$ from casini_hac which are defined as $\widehat{b}_{1,T}^{*}$ and $\widehat{\overline{b}}_{2,T}^{*}$, repetitively, with $\widehat{V}_{t}$ in place of $\widehat{V}_{D,t}^{*}$. We choose a parametric model for $\{a^{\left(r\right)\prime}V_{t}\}$, $r=1,\,\ldots,\,p$, where $a^{\left(r\right)}$ is a $p$-vector with the $r$-th element one and all other elements zero. We use the same locally stationary AR(1) models as in Section (ref), i.e., \[ V_{t}^{\left(r\right)}=a_{1}^{\left(r\right)}\left(t/T\right)V_{t-1}^{\left(r\right)}+u_{t}^{\left(r\right)}, \] with estimated parameters $\widehat{a}_{1}^{\left(r\right)}\left(\cdot\right)$ and $\widehat{\sigma}^{\left(r\right)}\left(\cdot\right).$ Let \[ \widehat{\theta}=\left(\int_{0}^{1}\widehat{a}_{1}^{\left(1\right)}\left(u\right)du,\,\int_{0}^{1}\left(\widehat{\sigma}^{\left(1\right)}\left(u\right)\right){}^{2}du,\ldots,\,\int_{0}^{1}\widehat{a}_{1}^{\left(p\right)}\left(u\right)du,\,\int_{0}^{1}\left(\widehat{\sigma}^{\left(p\right)}\left(u\right)\right){}^{2}du\right)', \] and $\theta_{\mathscr{P}}^{*}$ denote the probability limit of $\widehat{\theta}$. We only consider distributions $\mathscr{P}$ for which $\theta_{\mathscr{P}}^{*}$ exists. We construct $\widehat{\phi}\left(q\right)=\widehat{\phi}_{D}\left(q\right)$ as in Section (ref) but using the estimate $\widehat{\theta}$. The probability limit of $\widehat{\phi}\left(q\right)$ is denoted by $\phi_{\theta^{*}}\left(q\right)$. Let $\phi_{\mathscr{P}}\left(\cdot\right)$ be the value of $\phi\left(\cdot\right)$ from (ref) obtained when $\mathscr{P}_{U}$ is given by the approximating distribution with parameter $\theta_{\mathscr{P}}^{*}$. For some $\underline{\phi}$, $\overline{\phi}$ such that $0<\underline{\phi}\leq\overline{\phi}<\infty$, define
where
with $b>1+1/q$. The class of distributions $\boldsymbol{P}_{U,3}$ corresponds to the class $P_{1,1}$ used by andrews:88hac. The lower bound $0<\underline{\phi}\leq\phi_{\mathscr{P}}\left(q\right)$ in part (i) eliminates any distribution for which $\phi_{\mathscr{P}}\left(\cdot\right)=0$. For example, white noise sequences do not belong to $\boldsymbol{P}_{U,3}$ since then $\phi\left(q\right)=0$. We discuss these cases at the end of the section. Part (ii) imposes a condition on the temporal dependence of the distribution $\mathscr{P}_{U}$ and is similar to Assumption (ref)-(iii). Part (iii) is satisfied by a wide class of SLS processes as shown by casini_hac. Part (iv) was also used by andrews:88hac, though the interval $\boldsymbol{S}\left(q,\,b,\,l\right)$ is tighter as it takes into account of the time smoothing.
Let
denote the optimal bandwidth for the case in which $\mathscr{P}_{U}$ equals the approximating parametric model with parameter $\theta_{\mathscr{P}}^{*}$. Let \[ \widehat{D}_{2,a}\left(u\right)\triangleq2\sum_{l=-\left\lfloor T^{4/25}\right\rfloor }^{\left\lfloor T^{4/25}\right\rfloor }a'\widehat{c}_{T}\left(u_{0},\,l\right)\widehat{c}_{T}\left(u_{0},\,l\right)'a, \] where $\widehat{c}_{T}$ is defined as $\widehat{c}_{D,T}^{*}$ with $\widehat{V}_{t}$ and $\widehat{b}_{2,T}$ in place of $\widehat{V}_{D,t}^{*}$ and $\widehat{b}_{2,T}^{*}$, respectively,
Any estimator $\widehat{\phi}$ based on kernel nonparametric estimators of $\widehat{a}_{1}^{\left(r\right)}\left(\cdot\right)$ and $\widehat{\sigma}^{\left(r\right)}\left(\cdot\right)$ satisfies Assumption (ref)-(i). Assumption (ref)-(ii) extends Assumption (ref)-(vi) to the distribution $\mathscr{P}$ and is are useful to show that the effect of using $\widehat{b}_{1,T}$ and $\widehat{\overline{b}}_{2,T}$ rather than $b_{1,\theta_{\mathscr{P}},T}$ and $\overline{b}_{2,T}^{\mathrm{opt}}$ when constructing $\widehat{J}_{T}$ is at most $o_{\mathbb{P}}\left(1\right)$. The following result shows that $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{\overline{b}}_{2,T})$ has the same asymptotic MSE properties under $\mathscr{P}$ as the estimator $\widehat{J}_{T}(b_{1,\theta_{\mathscr{P}},T},\,\overline{b}_{2,T}^{\mathrm{opt}})$. Since the asymptotic MSE properties of the estimators with fixed bandwidth parameters have been determined in Section (ref), from this result follows the consistency of $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{\overline{b}}_{2,T})$ and its asymptotic optimality properties.
Theorem (ref) combined with Theorem (ref) and Theorem (ref)-(iii) establish upper and lower bounds on the asymptotic MSE. Results on asymptotic minimax optimality for data-dependent bandwidths parameters can be obtained using Theorem (ref), Theorem (ref)-(iii) and Theorem (ref)-(ref).
It remains to consider the case $\phi_{\mathscr{P}}\left(\cdot\right)=0$. When this occurs, $\widehat{\phi}^{-1}\left(\cdot\right)$ is $O_{\mathscr{P}}((T/n_{3,T})^{2}+n_{2,T})$. Under the additional condition $((T/n_{3,T})^{2}+n_{2,T})/T^{4/5}\rightarrow c\in[0,\,\infty)$ in Assumption (ref)-(i) we have $\widehat{b}_{1,T}=O_{\mathscr{P}}\left(1\right)$. Thus, $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{\overline{b}}_{2,T})-J_{\mathscr{P},T}\overset{\mathscr{P}}{\rightarrow}0$ also when the series is white noise. This is important in applied work because often researchers use robust standard errors even when they are not aware of whether any dependence is present at all.
A long-lasting problem in time series econometrics is the low/non-monotonic power of HAR inference tests under nonstationary alternative hypotheses. The problem involves HAR tests outside the regression model that can be characterized by an alternative hypothesis involving $\mathbb{E}\left(V_{t}\right)=\mu_{t}$ with $\mu_{t}\neq0$ for at least one $t$. The process $\mu_{t}$ can be any piecewise continuous function of $t$. For example, tests for structural breaks, tests for regime switching and tests for time-varying parameters can be framed in this way. To see this, consider a linear regression model,
The null hypothesis of no break in the regression coefficient of $x_{t}$ is written as $H_{\beta,0}:\,\beta_{t}=\beta_{0}$ for all $t$ for some $\beta_{0}\in\mathbb{R}^{p}$ {[}see, e.g., andrews:93{]}. The alternative hypothesis may be of several forms. Let $H_{\beta,1}:\,\beta_{t}=\beta\left(t/T\right)$ $\left(t=1,\ldots,\,T\right)$ for some piecewise continuous function $\beta\left(\cdot\right)$. Estimating (ref) by least-squares yields $y_{t}=x'_{t}\widehat{\beta}+\widehat{e}_{t}$ for all $t$ where $\widehat{\beta}$ is the least-squares estimate and $\left\{ \widehat{e}_{t}\right\} $ are the least-squares residuals. Letting $V_{t}=x_{t}\widehat{e}_{t}$, the null hypothesis $H_{\beta,0}$ can be rewritten as $H_{0}:\,\mathbb{E}\left(V_{t}\right)=0$ for all $t$ while the alternative hypothesis $H_{\beta,1}$ can be rewritten as $H_{1}:\,\mathbb{E}\left(V_{t}\right)=\mu_{t}$ where $\mu_{t}\neq0$ for at least one $t.$ Structural break tests are based on an estimate of the LRV of $V_{t}=x_{t}\widehat{e}_{t}$. While under $H_{0}$ $V_{t}$ is zero-mean, under $H_{1}$ the mean of $V_{t}$ is time-varying. Under the alternative hypothesis, it is sufficient for consistency of the test that the LRV estimator converges to some positive semidefinite matrix since the numerator of the test statistics diverges to infinity. However, time-variation in the mean of $V_{t}$ severely biases upward traditional LRV estimators which then lead to tests with non-monotonic power. casini/perron_Low_Frequency_Contam_Nonstat:2020 established analytical results for this phenomenon, which they referred to as low frequency contamination. We show that the proposed nonlinear prewhitened DK-HAC estimator accounts for nonstationarity also under the alternative hypothesis and leads to consistent tests with good monotonic power. Although the main theoretical result of this section is presented for a particular HAR test and a particular form of $H_{1}$, this result is general enough to provide guidance for most cases discussed in the literature.
We present theoretical results about the power of a popular forecast evaluation test, namely the Diebold-Mariano test {[}cf. diebold/mariano:95{]}, which can be also framed as above. We focus on the Diebold-Mariano test for ease of the exposition. Similar results hold for the other HAR inference tests that can be framed as above, though the proofs change slightly depending on the specific test statistic. Suppose the goal is to forecast some variable $y_{t}$. Two forecast models are used: $y_{t}=\beta^{\left(1\right)}+\beta^{\left(2\right)}x_{t-1}^{(i)}+e_{t}$ where $x_{t-1}^{(i)}$ is some predictor and $i=1,\,2$. That is, each forecast model uses an intercept and a predictor. The parameters $\beta^{\left(1\right)}$ and $\beta^{\left(2\right)}$ are estimated using least-squares in the in-sample $t=1,\ldots,\,T_{m}$ with a fixed forecasting scheme. Each forecast model generates a sequence of $\tau\left(=1\right)$-step ahead out-of-sample losses $L_{t}^{(i)}$ $\left(i=1,\,2\right)$ for $t=T_{m}+1,\ldots,\,T-\tau.$ Then $d_{t}\triangleq L_{t}^{(2)}-L_{t}^{(1)}$ denotes the loss differential at time $t$. Let $\overline{d}_{L}$ denote the average of the loss differentials. The Diebold-Mariano test statistic is defined as $t_{\mathrm{DM}}\triangleq T_{n}^{1/2}\overline{d}_{L}/\widehat{J}_{d_{L},T}^{1/2}$, where $\widehat{J}_{d_{L},T}$ is an estimate of the LRV of the loss differentials and $T_{n}$ is the number of observations in the out-of-sample. Throughout, we use the quadratic loss. The true model is $y_{t}=\beta_{0}^{\left(1\right)}+\beta_{0}^{\left(2\right)}x_{t-1}^{(0)}+e_{t}$ where $x_{t-1}^{(0)}$ is a predictor and $e_{t}$ is a zero-mean error. We assume that the conditions for consistency and asymptotic normality of the least-squares estimates of $\beta_{0}^{\left(1\right)}$ and $\beta_{0}^{\left(2\right)}$ are satisfied.
In this setting, $\widehat{V}_{t}=d_{t}.$ The hypothesis testing problem is given by
$H_{0}$ corresponds to equal predictive ability between the two forecast models while $H_{1}$ corresponds to the two forecast models performing differently.
Since we want to study the power of $t_{\mathrm{DM}}$, we need to work under the alternative hypothesis. The two competing forecast models are as follows: the first model uses the actual true predictor (i.e., $x_{t-1}^{(1)}=x_{t-1}^{(0)}$ for all $t$) while the second model differs in that in place of $x_{t-1}^{(0)}$ it uses $x_{t-1}^{(2)}=x_{t-1}^{(0)}+u_{X_{2},t}$ for $t\leq T_{b}$ and $x_{t-1}^{(2)}=\delta+x_{t-1}^{(0)}+u_{X_{2},t}$ for $t>T_{b}$ with $T_{b}>T_{m}$, and $u_{X_{2},t}$ is a zero-mean error term. Evidently, the null hypotheses of equal predictive ability should be rejected whenever $\delta>0$. We consider $t_{\mathrm{DM}}$ normalized by different LRV estimators. The HAC estimator is defined as
where $K_{1}\left(\cdot\right)$ is a kernel (e.g., the Bartlett and QS) and $b_{T}$ a bandwidth. Kiefer/vogelsang/bunzel:00 proposed to use a LRV estimator that keeps $b_{T}$ at a fixed fraction of $T$, i.e., $\widehat{J}_{\mathrm{\mathrm{KVB},}T}\triangleq T^{-1}\sum_{t=1}^{T}\sum_{s=1}^{T}$ $\left(1-\left|t-s\right|/T\right)\widehat{V}_{t}\widehat{V}_{s}$ which is equivalent to the Newey-West estimator with $b_{T}=T^{-1}$.
We present theoretical results about the power of $t_{\mathrm{DM}}$. Let $t_{\mathrm{DM},i}=T_{n}^{1/2}\overline{d}_{L}/\sqrt{\widehat{J}_{d_{L},i,T}}$ denote the DM test statistic where $i=\mathrm{DK},\,\mathrm{pwDK},\,\mathrm{KVB},\,\mathrm{EWC},$ $\mathrm{A91},\,\mathrm{pwA}91$, $\mathrm{NW87}$ and $\mathrm{pwNW87}$. $\widehat{J}_{d_{L},\mathrm{A91},T}$ and $\widehat{J}_{d_{L},\mathrm{NW87},T}$ are $\widehat{J}_{d_{L},\mathrm{HAC,}T}$ where $K_{1}\left(\cdot\right)$ is the Bartlett and QS kernel, respectively.\footnote{Since $\{\widehat{V}_{t}\}$ is only observed in the out-of-sample, the LRV estimators use a sample of $T_{n}$ observations.} $\widehat{J}_{d_{L},\mathrm{pwA91},T}$ and $\widehat{J}_{d_{L},\mathrm{pwNW87},T}$ are the prewhitened HAC estimators using the QS and Bartlett kernel, respectively, and the prewhitening procedure of andrews/monahan:92. “DK” refers to the DK-HAC estimator from casini_hac with the MSE-optimal kernels and bandwidths whereas “pwDK” refers to the prewhitened DK-HAC estimator $\widehat{J}_{\mathrm{pw},T}$ in (ref). Define the power of $t_{\mathrm{DM},i}$ as $\mathbb{P}_{\delta}(|t_{\mathrm{DM},i}|>z_{\alpha})$ where $z_{\alpha}$ is the two-sided standard normal critical value and $\alpha\in\left(0,\,1\right)$ is the significance level. To avoid repetitions we present the results only for $i=\mathrm{DK},\,\mathrm{pwDK},\,\mathrm{KVB},\,\mathrm{NW87}$ and $\mathrm{pwNW87}$. The results concerning the EWC estimator are the same as those for the KVB's fixed-$b$ estimator. The results pertaining to andrews:91' andrews:91 HAC estimator (with and without prewhitening) are the same as those corresponding to newey/west:87's newey/west:87 estimator (with and without prewhitening, respectively). For the HAC and DK-HAC estimators we report the results for the MSE-optimal bandwidth {[}see andrews:91, casini_hac and whilelm:2015{]}.\footnote{For the HAC estimators we also report the result for any bandwidth choice $b_{T}\rightarrow0$ such that $Tb_{T}\rightarrow\infty$, which is sufficient for the consistency of the estimator. } We set $n_{T}=n_{2,T}=n_{3,T}=T^{2/3}$ which satisfy the growth rate bounds {[}see casini_hac for details{]}. Let $n_{\delta}=T-T_{b}-2$ denote the length of the regime in which $x_{t}^{(2)}$ exhibits a shift $\delta$ in the mean. The alternative hypothesis depends on the shift magnitude $\delta$ and on how long the shift lasts for. Here the latter is $n_{\delta}$. More generally, this is the set of time points such that $\mathbb{E}(\widehat{V}_{t})=\mu_{t}\neq0$ holds.
Note that $b_{T}=O(T^{-1/3})$ in parts (i)-(ii) refers to the MSE-optimal bandwidth for the newey/west:87's newey/west:87 estimator. The conditions $T_{n}^{\zeta}b_{T}^{1/2}\rightarrow0$ and $T_{n}^{\zeta}(\widehat{b}_{1,T}^{*})^{1/2}\rightarrow0$ mean that the length of the regime in which $x_{t}^{(2)}$ exhibits a shift $\delta$ in the mean increases to infinity at a slower rate than $T$. Theorem (ref) implies that when the (prewhitened or non-prewhitened) HAC estimators or the fixed-$b$ LRV estimators are used, the DM test is not consistent and its power converges to zero. The theorem suggests that prewhitened and non-prewhitened HAC estimators suffer from this problem in a similar way. The theorem also implies that the power functions corresponding to tests based on HAC estimators lie above the power functions corresponding to those based on fixed-$b$/EWC LRV estimators. An additional feature is that $|t_{\mathrm{DM},\mathrm{NW87}}|$, $|t_{\mathrm{DM},\mathrm{pwNW87}}|$ and $|t_{\mathrm{DM},\mathrm{KVB}}|$ do not increase in magnitude with $\delta$ because $\delta$ appears in both the numerator and denominator. The results concerning the DK-HAC estimator and the prewhitened DK-HAC estimator $\widehat{J}_{\mathrm{pw},T}$ show that these issues do not occur when these estimators are used. In fact, the test is consistent and its power increases with $\delta$ and with the sample size. We provide finite-sample evidence in support of these theoretical results in Section (ref).
We now show that the prewhitened DK-HAC estimators lead to HAR inference tests that have accurate null rejection rates when there is strong dependence and have superior power properties relative to those based on traditional LRV estimators. We consider HAR tests in the linear regression model as well as applied to the forecast evaluation literature, namely the Diebold-Mariano test and the forecast breakdown test of giacomini/rossi:09.
The linear regression models have an intercept and a stochastic regressor. We focus on the $t$-statistics $t_{r}=\sqrt{T}(\widehat{\beta}^{\left(r\right)}-\beta_{0}^{\left(r\right)})/\sqrt{\widehat{J}_{X,T}^{\left(r,r\right)}}$ where $\widehat{J}_{X,T}$ is a consistent estimator of the limit of $\mathrm{Var}(\sqrt{T}(\widehat{\beta}-\beta_{0}))$ and $r=1,\,2$. $t_{1}$ is the $t$-statistic for the parameter associated to the intercept while $t_{2}$ is associated to the stochastic regressor. Two regression models are considered. We run a $t$-test on the intercept in model M1 whereas a $t$-test on the coefficient of the stochastic regressor is run in model M2. The models are,
for the $t$-test on the intercept and
for the $t$-test on $\beta_{0}^{\left(2\right)}$ where $\delta=0$ under the null hypotheses. In model M1 we set $\beta_{0}^{\left(1\right)}=0$, $\beta_{0}^{\left(2\right)}=1$, $x_{t}\sim\mathrm{i.i.d}.\,\mathscr{N}\left(1,\,1\right)$ and $e_{t}=\rho e_{t-1}+u_{t},\,\rho=0.4,\,0.9,\,u_{t}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(0,\,0.7\right)$. Model M2 involves segmented locally stationary errors: $\beta_{0}^{\left(1\right)}=\beta_{0}^{\left(2\right)}=0$, $x_{t}=0.6+0.8x_{t-1}+u_{x,t},\,u_{x,t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$ and $e_{t}=\rho_{t}e_{t-1}+u_{t},\,\rho_{t}=\max\left\{ 0,\,0.8\left(\cos\left(1.5-\cos\left(5t/T\right)\right)\right)\right\} $ for $t<4T/5$ and $e_{t}=0.5e_{t-1}+u_{t},\,u_{t}\sim\mathrm{\mathrm{i.i.d.\,}}\mathscr{N}\left(0,\,1\right)$ for $t\geq4T/5$. Note that $\rho_{t}$ varies smoothly between 0 and 0.7021. Then, $\widehat{J}_{X,T}=(X'X/T)^{-1}\widehat{J}_{T}(X'X/T)^{-1}$ where $X=\left[X_{1},\ldots,\,X_{T}\right]'$ and $X_{t}=[1,\,x_{t}]'$.
Next, we move to the forecast evaluation tests. The Diebold-Mariano test statistic is defined as in Section (ref), $t_{\mathrm{DM}}\triangleq T_{n}^{1/2}\overline{d}_{L}/\widehat{J}_{d_{L},T}^{1/2}$. In model M3 we consider an out-of-sample forecasting exercise with a fixed scheme where, given a sample of $T$ observations, $0.5T$ observations are used for the in-sample and the remaining half is used for prediction. To evaluate the empirical size, we specify the following data-generating process and the two forecasting models that have equal predictive ability. The true model for $y_{t}$ is given by $y_{t}=\beta_{0}^{\left(1\right)}+\beta_{0}^{\left(2\right)}x_{t-1}^{(0)}+e_{t}$ where $x_{t-1}^{(0)}\sim\mathrm{i.i.d.}\,\mathscr{N}\left(1,\,1\right)$, $e_{t}=0.8e_{t-1}+u_{t}$ with $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$ and we set $\beta_{0}^{\left(1\right)}=0,\,\beta_{0}^{\left(2\right)}=1.$ The two competing models differ on the predictor used in place of $x_{t}^{(0)}$. The first forecast model uses $x_{t}^{(1)}$ while the second uses $x_{t}^{(2)}$ where $x_{t}^{(1)}$ and $x_{t}^{(2)}$ are $\mathrm{i.i.d.}\,\mathscr{N}\left(1,\,1\right)$ sequences, both independent from $x_{t}^{(0)}$. Each forecast model generates a sequence of $\tau\left(=1\right)$-step ahead out-of-sample losses $L_{t}^{(i)}$ $\left(i=1,\,2\right)$ for $t=T/2+1,\ldots,\,T-\tau.$ Then $d_{t}\triangleq L_{t}^{(2)}-L_{t}^{(1)}$ denotes the loss differential at time $t$. The test rejects the null of equal predictive ability when (after normalization) $\overline{d}_{L}$ is sufficiently far from zero.
Next, we specify the alternative hypotheses for the Diebold-Mariano test. The two competing forecast models are as follows: the first model uses the actual true data-generating process while the second model differs in that in place of $x_{t-1}^{(0)}$ it uses $x_{t-1}^{(2)}=x_{t-1}^{(0)}+u_{X_{2},t}$ for $t\leq3T/4$ and $x_{t-1}^{(2)}=\delta+x_{t-1}^{(0)}+u_{X_{2},t}$ for $t>3T/4$, with $u_{X_{2},t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$. The null hypotheses of equal predictive ability should be rejected whenever $\delta>0$.
Finally, we consider model M4 which we use for investigating the performance the $t$-test for forecast breakdown of giacomini/rossi:09. Suppose we want to forecast a variable $y_{t}$ which follows $y_{t}=\beta_{0}^{\left(1\right)}+\beta_{0}^{\left(2\right)}x_{t-1}+\delta x_{t-1}\mathbf{1}\{t>T_{1}^{0}\}+e_{t}$ where $x_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(1.5,\,1.5\right)$ and $e_{t}=0.3e_{t-1}+u_{t}$ with $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,0.7\right)$, $\beta_{0}^{\left(1\right)}=\beta_{0}^{\left(2\right)}=1$ and $T_{1}^{0}=T\lambda_{1}^{0}$ with $\lambda_{1}^{0}=0.85$. The test detects a forecast breakdown when the average of the out-of-sample losses differs significantly from the average of the in-sample losses. The in-sample is used to obtain estimates of $\beta_{0}^{\left(1\right)}$ and $\beta_{0}^{\left(2\right)}$ which are in turn used to construct out-of-sample forecasts $\widehat{y}_{t}=\widehat{\beta}_{0}^{\left(1\right)}+\widehat{\beta}_{0}^{\left(2\right)}x_{t-1}$. The test is defined as $t^{\mathrm{GR}}\triangleq T_{n}^{1/2}\overline{SL}/\widehat{J}_{SL}^{1/2}$ where $\overline{SL}\triangleq T_{n}^{-1}\sum_{t=T_{m}+1}^{T-\tau}SL_{t+\tau}$, $SL_{t+\tau}$ is the surprise loss at time $t+\tau$, i.e., the difference between the time $t+\tau$ out-of-sample loss and in-sample-average loss, $SL_{t+\tau}=L_{t+\tau}-\overline{L}_{t+\tau}$. Here $T_{n}$ is the sample size in the out-of-sample, $T_{m}$ is the sample size in the in-sample and $\widehat{J}_{SL}$ is a LRV estimator. We consider a fixed forecasting scheme and $\tau=1.$
We consider the following DK-HAC estimators: $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS}}=\widehat{J}_{T,\mathrm{pw}}$ as discussed in Section (ref), $\widehat{J}_{T,\mathrm{pw},1}$ which uses prewhitening with a single block {[}$n_{T}=T$ in (ref){]} (i.e., stationary prewhitening), $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS,}\mu}$ which uses prewhitening involving a VAR(1) with time-varying intercept {[}i.e., with $\widehat{\mu}_{t}$ in (ref){]}. The asymptotic properties of $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS,}\mu}$ are the same as those of $\widehat{J}_{T,\mathrm{pw,\mathrm{SLS}}}$ since $\widehat{\mu}_{t}$ plays no role in the theory given the zero-mean assumption on $\left\{ V_{t}\right\} $. However, it leads to power enhancement under nonstationary alternative hypotheses. The asymptotic properties of $\widehat{J}_{T,\mathrm{pw},1}$ follow as a special case from the properties of $\widehat{J}_{T,\mathrm{pw,\mathrm{SLS}}}$. We set $n_{T}=n_{2,T}=n_{3,T}=T^{2/3}.$ For the test of giacomini/rossi:09 we do not report the results for $\widehat{J}_{T,\mathrm{pw},1}$ because the stationarity assumption is clearly violated under the alternative. We compare tests using these estimators to those using the following estimates: andrews:91' andrews:91 HAC estimator with automatic bandwidth; andrews:91' andrews:91 HAC estimator with automatic bandwidth and the prewhitening procedure of andrews/monahan:92; newey/west:87's newey/west:87 HAC estimator with the automatic bandwidth as proposed in newey/west:94; newey/west:87's newey/west:87 HAC estimator with the automatic bandwidth as proposed in newey/west:94 and the prewhitening procedure; Newey-West with the fixed-$b$ method of Kiefer/vogelsang/bunzel:00; the Empirical Weighted Cosine (EWC) of lazarus/lewis/stock/watson:18. We consider the following sample sizes: $T=200,\,400$ for M1-M2 and $T=400,\,800$ for model M3-M4. We set $T_{m}=200,\,400$ for M3 and $T_{m}=240,\,480$ for M4. The nominal size is $\alpha=0.05$ throughout.
Table (ref)-(ref) report the rejection rates under the null hypothesis for model M1-M4. We begin with model M1 with medium dependence ($\rho=0.4$). The prewhitened DK-HAC estimators lead to tests with accurate rejection rates that are slightly better than those obtained with Newey-West with fixed-$b$ and to EWC. In contrast, the classical HAC estimators of andrews:91 and newey/west:87 are less accurate with rejection rates higher than the nominal level. The prewhitening of andrews/monahan:92 helps to reduce the size distortions but they still persist for the Newey-West estimator even for $T=400.$ For higher dependence (i.e., $\rho=0.9$), using EWC and $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS,}\mu}$ yield oversized tests, though by a small margin. The best size control is achieved using the Newey-West with fixed-$b$ (KVB), $\widehat{J}_{T,\mathrm{pw},1}$ and $\widehat{J}_{T,\mathrm{pw,\mathrm{SLS}}}$.
For model M2, Newey-West with fixed-$b$ and the prewhitened DK-HAC $(\widehat{J}_{T,\mathrm{pw},1},\,\widehat{J}_{T,\mathrm{pw},\mathrm{SLS}},$ $\widehat{J}_{T,\mathrm{pw},\mu})$ allow accurate rejection rates. In some cases, tests based on the prewhitened DK-HAC are superior to those based on fixed-$b$ (KVB). The tests with EWC are slightly oversized when $T=200$ but close to the nominal level when $T=400$. The classical HAC of andrews:91 and newey/west:87, either prewhitened or not, imply oversized tests with $T=200.$
Turning to the HAR tests for forecast evaluation, Table (ref) reports some striking results. First, tests based on the Newey-West with fixed-$b$ (KVB) have size essentially equal to zero, while those based on the EWC and prewhitened or non-prewhitened classical HAC estimators are oversized. The prewhitened DK-HAC allows more accurate tests. For model M4, many of the tests have size equal to or close to zero. This occurs using the classical HAC, either prewhitened or not and EWC. The prewhitened DK-HAC estimators and Newey-West with fixed-$b$ (KVB) allow controlling the size reasonably well. Overall, Table (ref)-(ref) in part confirm previous evidence and in part suggest new facts. Newey-West with fixed-$b$ (KVB) leads to better size control than using the classical HAC estimators of andrews:91 and newey/west:87 even when the latter are used in conjunction with the prewhitening device of andrews/monahan:92. The new result is that several of the LRV estimators proposed in the literature can lead to tests having null rejection rates equal to or close to zero. This occurs because the null hypotheses involves nonstationary data generating mechanisms. These LRV estimators are inflated and the associated test statistics are undersized. This is expected to have negative consequences for the power of the tests, as we will see below. The estimators proposed in this paper perform well in leading to tests that control the null rejection rates for all cases. They are in general competitive with using the Newey-West with fixed-$b$ (KVB) when the latter does not fail and in some cases can also outperform it.
Table (ref)-(ref) report the empirical power of the tests for model M1-M4. For model M1 with $\rho=0.9$ and M2 we see that all tests have good and monotonic power. It is fair to compare tests based on the DK-HAC estimators relative to using Newey-West with fixed-$b$ (KVB) since they have similar well-controlled null rejection rates. Tests based on the Newey-West with fixed-$b$ (KVB) sacrifice power more than using the DK-HAC estimators and the difference is substantial. The classical HAC estimators have higher power but it is unfair to compare them since they are often oversized. A similar argument applies to using the EWC.
We now move to the forecast evaluation tests. For both models M3 and M4 we observe several features of interests. Essentially all tests proposed previously experience severe power issues. The power is either non-monotonic, very low or equal zero. This holds when using the classical HAC estimators of andrews:91 as well as newey/west:87 irrespective of whether prewhitening is used, with the EWC and the Newey-West with fixed-$b$ (KVB). The only exceptions are tests based on the newey/west:87's newey/west:87 and andrews:91' andrews:91 HAC estimator with prewhitening in model M4 that display some power but much lower compared to using the prewhitened DK-HAC estimators. The latter have excellent power. The reason for the severe power problems for the previous LRV-based tests is that models M3 and M4 involve nonstationary alternative hypotheses. The sample autocovariances become inflated and overestimate the true autocovariances. The theoretical results about the power in Theorem (ref) suggest that this issue becomes more severe as $\delta$ increases, which explains the non-monotonic power for some of the tests, with tests based on fixed-$b$ methods that include many lags suffering most. The double smoothing in the DK-HAC estimators allows to avoid this problem because it flexibly accounts for nonstationarity. The key idea is not to mix observations belonging to different regimes. Simulation results for additional data-generating processes involving ARMA, ARCH and heteroskedastic errors are not discussed here because the results are qualitatively equivalent.
We introduce a nonparametric nonlinear VAR prewhitened long-run variance (LRV) estimator for the construction of standard errors robust to autocorrelation and heteroskedasticity that can be used for hypothesis testing both within and outside the linear regression model. HAR tests normalized by the proposed estimator exhibit accurate null rejection rates even when there is strong dependence. We show theoretically that existing estimators lead to HAR tests that have low/non-monotonic power under nonstationary alternative hypotheses while the proposed estimator has good monotonic power thereby addressing a long-lasting problem in time series econometrics. The proposed method is theoretically valid under general nonstationary random variables. We also establish mean-squared error bounds for LRV estimation that are sharper than previously established and use them to determine the data-dependent bandwidths.
The supplement for online publication {[}cf. casini/perron_PrewhitedHAC_Supp{]} presents the proofs of the results in the paper.
\addcontentsline{toc}{section}{References}
\pagenumbering{arabic}