EconBase
← Back to paper

Simultaneous Bandwidths Determination for DK-HAC Estimators and Long-Run Variance Estimation in Nonparametric Settings

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

90,142 characters · 15 sections · 118 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Simultaneous Bandwidths Determination for DK-HAC Estimators and Long-Run Variance Estimation in Nonparametric Settings

\thispagestyle{empty} \setcounter{page}{0} \raggedbottom

abstract{We consider the derivation of data-dependent simultaneous bandwidths for double kernel heteroskedasticity and autocorrelation consistent (DK-HAC) estimators. In addition to the usual smoothing over lagged autocovariances for classical HAC estimators, the DK-HAC estimator also applies smoothing over the time direction. We obtain the optimal bandwidths that jointly minimize the global asymptotic MSE criterion and discuss the trade-off between bias and variance with respect to smoothing over lagged autocovariances and over time. Unlike the MSE results of andrews:91, we establish how nonstationarity affects the bias-variance trade-off. We use the plug-in approach to construct data-dependent bandwidths for the DK-HAC estimators and compare them with the DK-HAC estimators from casini_hac that use data-dependent bandwidths obtained from a sequential MSE criterion. The former performs better in terms of size control, especially with stationary and close to stationary data. Finally, we consider long-run variance estimation under the assumption that the series is a function of a nonparametric estimator rather than of a semiparametric estimator that enjoys the usual $\sqrt{T}$ rate of convergence. Thus, we also establish the validity of consistent long-run variance estimation in nonparametric parameter estimation settings.}

{{ {\bf{JEL Classification}}: C12, C13, C18, C22, C32, C51\\ {\bf{Keywords}}: Fixed-$b$, HAC standard errors, HAR, Long-run variance, Nonstationarity, Misspecification, Outliers, Segmented locally stationary.}}

\onehalfspacing \thispagestyle{empty}

Introduction

Long-run variance (LRV) estimation has a long history in econometrics and statistics since it plays a key role for heteroskedasticity and autocorrelation robust (HAR) inference. The classical approach in HAR inference relies on consistent estimation of the LRV. newey/west:87 and andrews:91 proposed kernel heteroskedasticity and autocorrelation consistent (HAC) estimators and showed their consistency. However, recent work by casini_hac showed that, both in the linear regression model and other contexts, their results do not provide accurate approximations in that test statistics normalized by classical HAC estimators may exhibit size distortions and substantial power losses. Issues with the power have been shown for a variety of HAR testing problems outside the regression model {[}e.g., altissimo/corradi:2003, Casini (casini_CR_Test_Inst_Forecast), Casini and Perron (casini/perron_Oxford_Survey, casini/perron_Lap_CR_Single_Inf, casini/perron_SC_BP_Lap), chan:2020, chang/perron:18, crainiceanu/vogelsang:07, deng/perron:06, juhl/xiao:09, kim/perron:09, martins/perron:16, perron/yamamoto:18 and vogeslang:99{]}. casini/perron_Low_Frequency_Contam_Nonstat:2020 showed theoretically that such power issues are generated by low frequency contamination induced by nonstationarity. More specifically, nonstationarity biases upward each sample autocovariance. Thus, LRV estimators are inflated and HAR test statistics lose power. These issues can also be provoked by misspecification, nonstationary alternative hypotheses and outliers. They also showed that LRV estimators that rely on fixed-$b$ or versions thereof suffer more from these problems than classical HAC estimators since the former use a larger number of sample autocovariances.\footnote{The fixed-$b$ literature is extensive. Pioneering contribution of Kiefer/vogelsang/bunzel:00 and Kiefer and Vogelsang (Kiefer/vogelsang:02; kiefer/vogelsang:05) introduced the fixed-$b$ LRV estimators. Additional contributions can be found in dou:18, lazarus/lewis/stock:17, lazarus/lewis/stock/watson:18, Goncalves/vogelsang:11, dejong/davidson:00, ibragimov/muller:10, jansson:04, muller:07 (2007, 2014)\nocite{mueller:14}, phillips:05, politis:11, preinerstorfer/potscher:16, potscher/preinerstorfer:18 (2018; 2019)\nocite{potscher/preinerstorfer:19}, robinson:98, Sun sun:13,sun:14,sun:14a, velasco/robinson:01 and zhang/shao:13. }

In order to flexibly account for nonstationarity, casini_hac introduced a double kernel HAC (DK-HAC) estimator that applies kernel smoothing over two directions. In addition to the usual smoothing over lagged autocovariances used in classical HAC estimators, the DK-HAC estimator uses a second kernel that applies smoothing over time. The latter accounts for time variation in the covariance structure of time series which is a relevant feature in economics and finance. Since the DK-HAC uses two kernels and bandwidths, one cannot rely on the theory of andrews:91 or newey/west:94 for selecting the bandwidths. casini_hac considered a sequential MSE criterion that determines the optimal bandwidth controlling the number of lags as a function of the optimal bandwidth controlling the smoothing over time. Thus, the latter influences the former but not viceversa. However, each smoothing affects the bias-variance trade-off so that the two bandwidths should affect each others optimal value. Consequently, it is useful to consider an alternative criterion to select the bandwidths. In this paper, we consider simultaneous bandwidths determination obtained by jointly minimizing the asymptotic MSE of the DK-HAC estimator. We obtain the asymptotic optimal formula for the two bandwidths and use the plug-in approach to replace unknown quantities by consistent estimates. Our results are established under the nonstationary framework characterized by segmented locally stationary processes {[}cf. casini_hac{]}. The latter extends the locally stationary framework of dahlhaus:96 to allow for discontinuities in the spectrum. Thus, the class of segmented locally stationary processes includes structural break models {[}see e.g., bai/perron:98 and casini/perron_CR_Single_Break{]}, time-varying parameter models {[}see e.g., cai:07{]} and regime switching {[}cf. hamilton:89{]}.

We establish the consistency, rate of convergence and asymptotic MSE results for the DK-HAC estimators with data-dependent simultaneous bandwidths. The optimal bandwidths have the same order $O(T^{-1/6})$ whereas under the sequential criterion the optimal bandwidths smoothing over time has an order $O(T^{-1/5})$ and the optimal bandwidth smoothing the lagged autocovariances has an order $O(T^{-4/25})$. Thus, asymptotically, the joint MSE criterion implies the use of (marginally) more lagged autocovariances and a longer segment length for the smoothing over time relative to the sequential criterion. Hence, the former should control more accurately the variance due to nonstationarity while the latter should control better the bias. If the degree of nonstationarity is high then the theory suggests that one should expect the sequential criterion to perform marginally better. The difference in the smoothing over lags is very minor between the order of the corresponding bandwidths implied by the two criteria. Our simulation analysis supports this view as we show that the joint MSE criterion performs better especially when the degree of nonstationarity is not too high.

Overall, we find that HAR tests normalized by DK-HAC estimators strike the best balance between size and power among the existing LRV estimators and we also find that using the bandwidths selected from the joint MSE criterion yields tests that perform better than the sequential criterion in terms of size control. The optimal rate $O(T^{-1/6})$ is also found by neumann/von_sachs:1997 and dahlhaus:12 in the context of local spectral density estimates under local stationarity. Under both sequential and joint MSE criterion the optimal kernels are found to be the same, i.e., the quadratic spectral kernel for smoothing over autocovariance lags {[}similar to andrews:91{]} and a parabolic kernel {[}cf. epanechnikov:69{]} for smoothing over time.

Another contribution of the paper is develop asymptotic results for consistent LRV estimation in nonparametric parameter estimation settings. newey/west:87 and andrews:91 established the consistency of HAC estimators for the long-run variance of some series $\{V_{t}(\widehat{\beta})\}$ where $\widehat{\beta}$ is a semiparametric estimator of $\beta_{0}$ having the usual parametric rate of convergence $\sqrt{T}$ {[}i.e., they assumed that $\sqrt{T}(\widehat{\beta}-\beta_{0})=O_{\mathbb{P}}\left(1\right)${]}. For example, in the linear regression model estimated by least-squares, $V_{t}(\widehat{\beta})=\widehat{e}_{t}x_{t}$ where $\{\widehat{e}_{t}\}$ are the least-squares residuals and $\{x_{t}\}$ is a vector of regressors. Unfortunately, the condition $\sqrt{T}(\widehat{\beta}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ does not hold for nonparametric estimators $\widehat{\beta}_{\mathrm{np}}$ since they satisfy $T^{\vartheta}(\widehat{\beta}_{\mathrm{np}}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ for some $\vartheta\in\left(0,\,1/2\right)$. For example, for tests for forecast evaluation often forecasters use nonparametric kernel methods to obtain the forecasts {[}i.e., $\{V_{t}(\widehat{\beta})\}=L(e_{t}(\widehat{\beta}_{\mathrm{np}}))$ where $L\left(\cdot\right)$ is a forecast loss, $e_{t}(\cdot)$ is a forecast error and $\widehat{\beta}_{\mathrm{np}}$ is, e.g., a rolling window estimate of a parameter that is used to construct the forecasts{]}. Given the widespread use of nonparametric methods in applied work, it is useful to extend the theoretical results of HAC and DK-HAC estimators for these settings. We establish the validity of HAC and DK-HAC estimators including the validity of the corresponding estimators based on data-dependent bandwidths.

The remainder of the paper is organized as follows. Section (ref) introduces the statistical setting and the joint MSE criterion. Section (ref) presents consistency, rates of convergence, asymptotic MSE results, and optimal kernels and bandwidths for the DK-HAC estimators using the joint MSE criterion. Section (ref) develops a data-dependent method for simultaneous bandwidth parameters selection and its asymptotic properties are then discussed. Section (ref) presents theoretical results for LRV estimation in nonparametric parameter estimation. Section (ref) presents Monte Carlo results about the small-sample size and power of HAR tests based on the DK-HAC estimators using the proposed automatic simultaneous bandwidths. We also provide comparisons with a variety of other approaches. Section (ref) concludes the paper. The supplemental material {[}belotti/casini/catania/grassi/perron_HAC_Sim_Bandws_Supp{]} contains the mathematical proofs. The code to implement the proposed methods is available online in $\mathsf{\mathrm{\mathtt{Matlab}}}$, $\mathtt{R}$ and $\mathtt{Stata}$ languages.

The Statistical Environment

We consider the estimation of the LRV $J\triangleq\mathrm{lim}_{T\rightarrow\infty}J_{T}$ where $J_{T}=T^{-1}\sum_{s=1}^{T}\sum_{t=1}^{T}\mathbb{E}(V_{s}\left(\beta_{0}\right)V_{t}\left(\beta_{0}\right)')$ with $V_{t}\left(\beta\right)$ being a random $p$-vector for each $\beta\in\Theta$. For example, for the linear model $V_{t}\left(\beta\right)=\left(y_{t}-x'_{t}\beta\right)x_{t}$. The classical approach for inference in the context of serially correlated data is based on consistent estimation of $J$. newey/west:87 and andrews:91 considered the class of kernel HAC estimators, where the subscript Cla stands for classical,

align*[align* omitted — 446 chars of source]

$\widehat{V}_{t}=V_{t}(\widehat{\beta})$, $K_{1}\left(\cdot\right)$ is a real-valued kernel in the class $\boldsymbol{K}_{1}$ defined below and $b_{1,T}$ is a bandwidth sequence. The factor $T/\left(T-p\right)$ is an optional small-sample degrees of freedom adjustment. For the Newey-West estimator $K_{1}$ corresponds to the Bartlett kernel while for Andrews' andrews:91 $K_{1}$ corresponds to the quadratic spectral (QS) kernel. Data-dependent methods for the selection of $b_{1,T}$ were proposed by newey/west:94 and andrews:91, respectively. Under appropriate conditions on $b_{1,T}\rightarrow0$ they showed that $\widehat{J}_{\mathrm{Cla},T}\overset{\mathbb{P}}{\rightarrow}J$. When $\left\{ V_{t}\right\} $ is second-order stationary, $J=2\pi f\left(0\right)$ where $f\left(0\right)$ is the spectral density of $\left\{ V_{t}\right\} $ at frequency zero. Most of the LRV estimation literature has focused on the stationarity assumption for $\left\{ V_{t}\right\} $ {[}e.g., Kiefer/vogelsang/bunzel:00, muller:07 and lazarus/lewis/stock:17{]}. Unlike the HAC estimators, fixed-$b$ (and versions thereof) LRV estimators require stationarity of $\left\{ V_{t}\right\} $. The latter assumption is restrictive for economic and financial time series. The properties of $J$ under nonstationarity were studied recently by casini_hac who showed that if $\left\{ V_{t}\right\} $ is either locally stationary or segmented locally stationary (SLS), then $J=2\pi\int_{0}^{1}f\left(u,\,0\right)du$ where $f\left(u,\,0\right)$ is the time-varying spectral density at rescaled time $u=t/T$ and frequency zero. For locally stationary processes, $f\left(u,\,0\right)$ is smooth in $u$ while for SLS processes $f\left(u,\,0\right)$ can in addition contain a finite-number of discontinuities. The number of discontinuities can actually grow to infinity with unchanged results though at the expense of slightly more complex derivations. Since the assumption of a finite number of discontinuities capture well the idea that a finite number of regimes or structural breaks is enough to account for structural changes (or big events) in economic time series we maintain this assumption here. The latter is relaxed by casini/perron_PrewhitedHAC.

Under nonstationarity casini_hac argued that an extension of the classical HAC estimators can actually account flexibly for the time-varying properties of the data. He proposed the class of double kernel HAC (DK-HAC) estimators,

align*[align* omitted — 373 chars of source]

where $n_{T}\rightarrow\infty$ satisfies the conditions given below, and

align[align omitted — 471 chars of source]

with $K_{2}^{*}$ being a real-valued kernel and $b_{2,T}$ is a bandwidth sequence. $\widehat{c}_{T}\left(u,\,k\right)$ is an estimate of the local autocovariance $c\left(u,\,k\right)=\mathbb{E}(V_{\left\lfloor Tu\right\rfloor },\,V'_{\left\lfloor Tu\right\rfloor -k})+O\left(T^{-1}\right)$ {[}under regularity conditions; see casini_hac{]} at lag $k$ and time $u=rn_{T}/T$. $\widehat{\Gamma}\left(k\right)$ estimates the local autocovariance across blocks of length $n_{T}$ and then takes an average over the blocks. The estimator $\widehat{J}_{T}$ involves two kernels: $K_{1}$ smooths the lagged autocovariances\textemdash akin to the classical HAC estimators\textemdash while $K_{2}$ applies smoothing over time. The smoothing over time better account for nonstationarity and makes $\widehat{J}_{\mathrm{DK,}T}$ robust to low frequency contamination. See casini/perron_Low_Frequency_Contam_Nonstat:2020 who showed theoretically that existing LRV estimators are contaminated by nonstationarity so that they become inflated with consequent large power losses when the estimators are used to normalize HAR test statistics.

casini_hac considered adaptive estimators $\widehat{J}_{\mathrm{DK,}T}$ for which $b_{1,T}$ and $b_{2,T}$ are data-dependent. Observe that the optimal $b_{2,T}$ actually depends on the properties of $\left\{ V_{t,T}\right\} $ in any given block {[}i.e., $b_{2,T}=b_{2,T}\left(t/T\right)${]}. Let

align*[align* omitted — 373 chars of source]

where $\widetilde{W}_{T}$ is some $p\times p$ positive semidefinite matrix. He considered a sequential MSE criterion to determine the optimal kernels and bandwidths. For $K_{1},$ the result states that the QS kernel minimizes the asymptotic MSE for any $K_{2}\left(\cdot\right)$. The optimal $b_{1,T}^{\mathrm{opt}}$ and $b_{2,T}^{\mathrm{opt}}$ satisfy the following,

align[align omitted — 729 chars of source]

$\widehat{J}_{T}(b_{1,T},\,\overline{b}_{2,T}^{\mathrm{opt}})$ indicates the estimator $\widehat{J}_{T}$ that uses $b_{1,T}^{\mathrm{}}$ and $\overline{b}_{2,T}^{\mathrm{opt}}$. Eq. (ref) holds as $T\rightarrow\infty$. The above criterion determines the globally optimal $b_{1,T}^{\mathrm{opt}}$ given the integrated locally optimal $b_{2,T}^{\mathrm{opt}}\left(u\right)$. Under (ref), only $b_{2,T}$ affects $b_{1,T}$ but not vice-versa. Intuitively, this is a limitation because it is likely that in order to minimize the global MSE the bandwidths $b_{1,T}$ and $b_{2,T}$ affect each other.

In this paper, we consider a more theoretically appealing criterion to determine the optimal bandwidths. That is, we consider bandwidths $(\widetilde{b}_{1,T}^{\mathrm{opt}},\,\widetilde{b}_{2,T}^{\mathrm{opt}})$ that jointly minimize the global asymptotic relative MSE, denoted by ReMSE,

align[align omitted — 369 chars of source]

where $W_{T}$ is $p^{2}\times p^{2}$ weight matrix. Under (ref), $\widetilde{b}_{1,T}^{\mathrm{opt}}$ and $\widetilde{b}_{2,T}^{\mathrm{opt}}$ affect each other simultaneously. This is a more reasonable property. In Section (ref) we solve for the sequences $(\widetilde{b}_{1,T}^{\mathrm{opt}},\,\widetilde{b}_{2,T}^{\mathrm{opt}})$ that minimize (ref). We propose a data-dependent method for $(\widetilde{b}_{1,T}^{\mathrm{opt}},\,\widetilde{b}_{2,T}^{\mathrm{opt}})$ in Section (ref).

The literature on LRV estimation has routinely focused on the case where $\widehat{V}_{t}$ is a function of a parameter estimate $\widehat{\beta}$ that enjoys a standard $\sqrt{T}$ parametric rate of convergence. While this is an important case, the recent increasing use of nonparametric methods suggests that the case where $\widehat{\beta}$ enjoys a nonparametric rate of convergence slower than $\sqrt{T}$ is of potential interest. Hence, in Section (ref) we consider consistent LRV estimation under the latter framework and develop corresponding results for the classical HAC as well as the DK-HAC estimators.

We consider the following standard classes of kernels {[}cf. andrews:91{]},

align[align omitted — 799 chars of source]

The class $\boldsymbol{K}_{1}$ was also considered by andrews:91. Examples of kernels in $\boldsymbol{K}_{1}$ include the Truncated, Bartlett, Parzen, Quadratic Spectral (QS) and Tukey-Hanning kernels. The QS kernel was shown to be optimal for $\widehat{J}_{\mathrm{Cla},T}$ under the MSE criterion by andrews:91 and for $\widehat{J}_{T}$ under a sequential MSE criterion by casini_hac,

align*[align* omitted — 158 chars of source]

The class $\boldsymbol{K}_{2}$ was also considered by, for example, Dahlhaus/Giraitis:98.

Throughout we adopt the following notational conventions. The $j$th element of a vector $x$ is indicated by $x^{\left(j\right)}$ while the $\left(j,\,l\right)$th element of a matrix $X$ is indicated as $X^{\left(j,\,l\right)}$. $\mathrm{tr}\left(\cdot\right)$ denotes the trace function and $\otimes$ denotes the tensor (or Kronecker) product operator. The $p^{2}\times p^{2}$ matrix $C_{pp}$ is a commutation matrix that transforms $\mathrm{vec}\left(A\right)$ into $\mathrm{vec}\left(A'\right)$, i.e., $C_{pp}=\sum_{j=1}^{p}\sum_{l=1}^{p}\iota_{j}\iota_{l}'\otimes\iota_{l}\iota_{j}'$, where $\iota_{j}$ is the $j$th elementary $p$-vector. $\lambda_{\max}\left(A\right)$ denotes the largest eigenvalue of the matrix $A$. $W$ and $\widetilde{W}$ are used for $p^{2}\times p^{2}$ weight matrices. $\mathbb{C}$ is used for the set of complex numbers and $\overline{A}$ for the complex conjugate of $A\in\mathbb{C}$. Let $0=\lambda_{0}<\lambda_{1}<\ldots<\lambda_{m}<\lambda_{m+1}=1$. A function $G\left(\cdot,\,\cdot\right):\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ is said to be piecewise (Lipschitz) continuous with $m+1$ segments if it is (Lipschitz) continuous within each segment. For example, it is piecewise Lipschitz continuous if for each segment $j=1,\ldots,\,m+1$ it satisfies $\sup_{u\neq v}\left|G\left(u,\,\omega\right)-G\left(v,\,\omega\right)\right|\leq K\left|u-v\right|$ for any $\omega\in\mathbb{R}$ with $\lambda_{j-1}<u,\,v\leq\lambda_{j}$ for some $K<\infty.$ We define $G_{j}\left(u,\,\omega\right)=G\left(u,\,\omega\right)$ for $\lambda_{j-1}<u\leq\lambda_{j}$, so $G_{j}\left(u,\,\omega\right)$ is Lipschitz continuous for each $j.$ If we say piecewise Lipschitz continuous with index $\vartheta>0$, then the above inequality is replaced by $\sup_{u\neq v}\left|G\left(u,\,\omega\right)-G\left(v,\,\omega\right)\right|\leq K\left|u-v\right|^{\vartheta}$. A function $G\left(\cdot,\,\cdot\right):\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ is said to be left-differentiable at $u_{0}$ if $\partial G\left(u_{0},\omega\right)/\partial_{-}u\triangleq\lim_{u\rightarrow u_{0}^{-}}\left(G\left(u_{0},\,\omega\right)-G\left(u,\,\omega\right)\right)/\left(u_{0}-u\right)$ exists for any $\omega\in\mathbb{R}$. We use $\left\lfloor \cdot\right\rfloor $ to denote the largest smaller integer function. The symbol “$\triangleq$” is for definitional equivalence.

Simultaneous Bandwidths Determination for DK-HAC Estimators

In Section (ref) we present the consistency, rate of convergence and asymptotic MSE properties of predetermined bandwidths for the DK-HAC estimators. We use the MSE results to determine the optimal bandwidths and kernels in Section (ref). We use the framework for nonstationarity introduced in casini_hac. That is, we assume that $\left\{ V_{t,T}\right\} $ is segmented locally stationary (SLS). Suppose $\left\{ V_{t}\right\} _{t=1}^{T}$ is defined on an abstract probability space $\left(\Omega,\,\mathscr{F},\,\mathbb{P}\right)$, where $\Omega$ is the sample space, $\mathscr{F}$ is the $\sigma$-algebra and $\mathbb{P}$ is a probability measure. We use an infill asymptotic setting and rescale the original discrete time horizon $\left[1,\,T\right]$ by dividing each $t$ by $T.$ Letting $u=t/T$ and $T\rightarrow\infty,$ this defines a new time scale $u\in\left[0,\,1\right]$. Let $i\triangleq\sqrt{-1}$.

defnA sequence of stochastic processes $\{V_{t,T}\}_{t=1}^{T}$ is called Segmented Locally Stationarity (SLS) with $m_{0}+1$ regimes, transfer function $A^{0}$ and trend $\mu_{\cdot}$ if there exists a representation \begin{align} V_{t,T} & =\mu_{j}\left(t/T\right)+\int_{-\pi}^{\pi}\exp\left(i\omega t\right)A_{j,t,T}^{0}\left(\omega\right)d\xi\left(\omega\right),\qquad\qquad\left(t=T_{j-1}^{0}+1,\ldots,\,T_{j}^{0}\right), \end{align} for $j=1,\ldots,\,m_{0}+1$, where by convention $T_{0}^{0}=0$ and $T_{m_{0}+1}^{0}=T$ and the following holds: (i) $\xi\left(\omega\right)$ is a stochastic process on $\left[-\pi,\,\pi\right]$ with $\overline{\xi\left(\omega\right)}=\xi\left(-\omega\right)$ and \begin{align*} \mathrm{cum}\left\{ d\xi\left(\omega_{1}\right),\ldots,\,d\xi\left(\omega_{r}\right)\right\} & =\varphi\left(\sum_{j=1}^{r}\omega_{j}\right)g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)d\omega_{1}\ldots d\omega_{r}, \end{align*} where $\mathrm{cum}\left\{ \cdot\right\} $ is the cumulant of $r$th order, $g_{1}=0,\,g_{2}\left(\omega\right)=1$, $\left|g_{r}\left(\omega_{1},\ldots,\,\omega_{r-1}\right)\right|\leq M_{r}<\infty$ and $\varphi\left(\omega\right)=\sum_{j=-\infty}^{\infty}\delta\left(\omega+2\pi j\right)$ is the period $2\pi$ extension of the Dirac delta function $\delta\left(\cdot\right)$. (ii) There exists a constant $K>0$ and a piecewise continuous function $A:\,\left[0,\,1\right]\times\mathbb{R}\rightarrow\mathbb{C}$ such that, for each $j=1,\ldots,\,m_{0}+1$, there exists a $2\pi$-periodic function $A_{j}:\,(\lambda_{j-1},\,\lambda_{j}]\times\mathbb{R}\rightarrow\mathbb{C}$ with $A_{j}\left(u,\,-\omega\right)=\overline{A_{j}\left(u,\,\omega\right)}$, $\lambda_{j}^{0}\triangleq T_{j}^{0}/T$ and for all $T,$ \begin{align} A\left(u,\,\omega\right) & =A_{j}\left(u,\,\omega\right)\,\mathrm{\,for\,}\,\lambda_{j-1}^{0}<u\leq\lambda_{j}^{0},\\ \sup_{1\leq j\leq m_{0}+1} & \sup_{T_{j-1}^{0}<t\leq T_{j}^{0},\,\omega}\left|A_{j,t,T}^{0}\left(\omega\right)-A_{j}\left(t/T,\,\omega\right)\right|\leq KT^{-1}. \end{align} (iii) $\mu_{j}\left(t/T\right)$ is piecewise continuous.

Observe that this representation is similar to the spectral representation of stationary processes {[}see anderson:71, brillinger:75, hannan:70 and priestley:85 for introductory concepts{]}. The main difference is that $A\left(t/T,\,\omega\right)$ and $\mu\left(t/T\right)$ are not constant in $t$. dahlhaus:96 used the time-varying spectral representation to define the so-called locally stationary processes which are characterized, broadly speaking, by smoothness conditions on $\mu\left(\cdot\right)$ and $A\left(\cdot,\,\cdot\right)$. Locally stationary processes are often referred to as time-varying parameter processes {[}see e.g., cai:07 and chen/hong:12{]}. However, the smoothness restrictions exclude many prominent models that account for time variation in the parameters. For example, structural change and regime switching-type models do not belong to this class because parameter changes occur suddenly at a particular time. Thus, the class of SLS processes is more general and likely to be more useful. Stationarity and local stationarity are recovered as special cases of the SLS definition.

Let $\mathcal{T}\triangleq\{T_{1}^{0},\,\ldots,\,T_{m_{0}}^{0}\}$. The spectrum of $V_{t,T}$ is defined (for fixed $T$) as

align*[align* omitted — 569 chars of source]

with $A_{1,t,T}^{0}\left(\omega\right)=A_{1}\left(0,\,\omega\right)$ for $t<1$ and $A_{m+1,t,T}^{0}\left(\omega\right)=A_{m+1}\left(1,\,\omega\right)$ for $t>T$. casini_hac showed that $f_{j,T}\left(u,\,\omega\right)$ tends in mean-squared to $f_{j}\left(u,\,\omega\right)\triangleq\left|A_{j}\left(u,\,\omega\right)\right|^{2}$ for $T_{j-1}^{0}/T<u=t/T\leq T_{j}^{0}/T$, which is the spectrum that corresponds to the spectral representation. Therefore, we call $f_{j}\left(u,\,\omega\right)$ the time-varying spectral density matrix of the process. Given $f\left(u,\,\omega\right),$ we can define the local covariance of $V_{t,T}$ at rescaled time $u$ with $Tu\notin\mathcal{T}$ and lag $k\in\mathbb{Z}$ as $c\left(u,\,k\right)\triangleq\int_{-\pi}^{\pi}e^{i\omega k}f\left(u,\,\omega\right)d\omega$. The same definition is also used when $Tu\in\mathcal{T}$ and $k\geq0$. For $Tu\in\mathcal{T}$ and $k<0$ it is defined as $c\left(u,\,k\right)\triangleq\int_{-\pi}^{\pi}e^{i\omega k}A\left(u,\,\omega\right)A\left(u-k/T,\,-\omega\right)d\omega$.

Asymptotic MSE Properties of DK-HAC estimators

Let $\widetilde{J}_{T}$ denote the pseudo-estimator identical to $\widehat{J}_{T}$ but based on $\{V_{t,T}\}=\{V_{t,T}\left(\beta_{0}\right)\}$ rather than on $\{\widehat{V}_{t,T}\}=\{V_{t,T}(\widehat{\beta})\}$.

assumption(i) $\{V_{t,T}\}$ is a mean-zero SLS process with $m_{0}+1$ regimes; (ii) $A\left(u,\,\omega\right)$ is \textcolor{red}twice continuously differentiable in $u$ at all $u\neq\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly bounded derivatives $\left(\partial/\partial u\right)A\left(u,\,\cdot\right)$ and $\left(\partial^{2}/\partial u^{2}\right)A\left(u,\,\cdot\right)$, and Lipschitz continuous in the second component with index $\vartheta=1$; (iii) $\left(\partial^{2}/\partial u^{2}\right)A\left(u,\,\cdot\right)$ is Lipschitz continuous at all $u\neq\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$; (iv) $A\left(u,\,\omega\right)$ is twice left-differentiable in $u$ at $u=\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly bounded derivatives $\left(\partial/\partial_{-}u\right)A\left(u,\,\cdot\right)$ and $\left(\partial^{2}/\partial_{-}u^{2}\right)A\left(u,\,\cdot\right)$ and has piecewise Lipschitz continuous derivative $\left(\partial^{2}/\partial_{-}u^{2}\right)A\left(u,\,\cdot\right)$.

We also need to impose conditions on the temporal dependence of $V_{t}=V_{t,T}$. Let

align*[align* omitted — 521 chars of source]

where $\left\{ V_{\mathscr{N},t}\right\} $ is a Gaussian sequence with the same mean and covariance structure as $\left\{ V_{t}\right\} $. $\kappa_{V,t}^{\left(a,b,c,d\right)}\left(u,\,v,\,w\right)$ is the time-$t$ fourth-order cumulant of $(V_{t}^{\left(a\right)},\,V_{t+u}^{\left(b\right)},\,V_{t+v}^{\left(c\right)},$ $\,V_{t+w}^{\left(d\right)})$ while $\kappa_{\mathscr{N}}^{\left(a,b,c,d\right)}(t,\,t+u,$ $\,t+v,\,t+w)$ is the time-$t$ centered fourth moment of $V_{t}$ if $V_{t}$ were Gaussian.

assumption(i) $\sum_{k=-\infty}^{\infty}\sup_{u\in\left[0,\,1\right]}$ $\left\Vert c\left(u,\,k\right)\right\Vert <\infty$, $\sum_{k=-\infty}^{\infty}\sup_{u\in\left[0,\,1\right]}\left\Vert \left(\partial^{2}/\partial u^{2}\right)c\left(u,\,k\right)\right\Vert <\infty$ and $\sum_{k=-\infty}^{\infty}\sum_{j=-\infty}^{\infty}\sum_{l=-\infty}^{\infty}\sup_{u\in\left[0,\,1\right]}|\kappa_{V,\left\lfloor Tu\right\rfloor }^{\left(a,b,c,d\right)}$ $\left(k,\,j,\,l\right)|<\infty$ for all $a,\,b,\,c,\,d\leq p$. (ii) For all $a,\,b,\,c,\,d\leq p$ there exists a function $\widetilde{\kappa}_{a,b,c,d}:\,\left[0,\,1\right]\times\mathbb{Z}\times\mathbb{Z}\times\mathbb{Z}\rightarrow\mathbb{R}$ such that $\sup_{u\in\left(0,\,1\right)}|\kappa_{V,\left\lfloor Tu\right\rfloor }^{\left(a,b,c,d\right)}\left(k,\,s,\,l\right)$ $-\widetilde{\kappa}_{a,b,c,d}\left(u,\,k,\,s,\,l\right)|\leq KT^{-1}$ for some constant $K$; the function $\widetilde{\kappa}_{a,b,c,d}\left(u,\,k,\,s,\,l\right)$ is twice differentiable in $u$ at all $u\neq\lambda_{j}^{0}$, $(j=1,\ldots,\,m_{0}+1)$, with uniformly bounded derivatives $\left(\partial/\partial u\right)\widetilde{\kappa}_{a,b,c,d}\left(u,\cdot,\cdot,\cdot\right)$ and $\left(\partial^{2}/\partial u^{2}\right)\widetilde{\kappa}_{a,b,c,d}\left(u,\cdot,\cdot,\cdot\right)$, and twice left-differentiable in $u$ at $u=\lambda_{j}^{0}$ $(j=1,\ldots,\,m_{0}+1)$ with uniformly bounded derivatives $\left(\partial/\partial_{-}u\right)\widetilde{\kappa}_{a,b,c,d}\left(u,\cdot,\cdot,\cdot\right)$ and $\left(\partial^{2}/\partial_{-}u^{2}\right)\widetilde{\kappa}_{a,b,c,d}$ $\left(u,\cdot,\cdot,\cdot\right)$ and piecewise Lipschitz continuous derivative $\left(\partial^{2}/\partial_{-}u^{2}\right)\widetilde{\kappa}_{a,b,c,d}\left(u,\cdot,\cdot,\cdot\right)$.

We do not require fourth-order stationarity but only that the time-$t=Tu$ fourth order cumulant is locally constant in a neighborhood of $u$.

Following parzen:57, we define $K_{1,q}\triangleq\lim_{x\downarrow0}\left(1-K_{1}\left(x\right)\right)/\left|x\right|^{q}$ for $q\in[0,\,\infty);$ $q$ increases with the smoothness of $K_{1}\left(\cdot\right)$ with the largest value being such that $K_{1,q}<\infty$. When $q$ is an even integer, $K_{1,q}=-\left(d^{q}K_{1}\left(x\right)/dx^{q}\right)|_{x=0}/q!$ and $K_{1,q}<\infty$ if and only if $K_{1}\left(x\right)$ is $q$ times differentiable at zero. We define the index of smoothness of $f\left(u,\,\omega\right)$ at $\omega=0$ by $f^{\left(q\right)}\left(u,\,0\right)\triangleq\left(2\pi\right)^{-1}\sum_{k=-\infty}^{\infty}\left|k\right|^{q}c\left(u,\,k\right)$, for $q\in[0,\,\infty)$. If $q$ is even, then $f^{\left(q\right)}\left(u,\,0\right)=\left(-1\right)^{q/2}\left(d^{q}f\left(u,\,\omega\right)/d\omega^{q}\right)|_{\omega=0}$. Further, $||f^{\left(q\right)}\left(u,\,0\right)||<\infty$ if and only if $f\left(u,\,\omega\right)$ is $q$ times differentiable at $\omega=0$. We define

align[align omitted — 237 chars of source]
thmSuppose $K_{1}\left(\cdot\right)\in\boldsymbol{K}_{1}$, $K_{2}\left(\cdot\right)\in\boldsymbol{K}_{2}$, Assumption (ref)-(ref) hold, $b_{1,T},\,b_{2,T}\rightarrow0$, $n_{T}\rightarrow\infty,\,n_{T}/T\rightarrow0$ and $1/Tb_{1,T}b_{2,T}\rightarrow0$. We have: (i) \begin{align*} \lim_{T\rightarrow\infty} & Tb_{1,T}b_{2,T}\mathrm{Var}\left[\mathrm{vec}\left(\widetilde{J}_{T}\right)\right]\\ & =4\pi^{2}\int K_{1}^{2}\left(y\right)dy\int_{0}^{1}K_{2}^{2}\left(x\right)dx\left(I+C_{pp}\right)\left(\int_{0}^{1}f\left(u,\,0\right)du\right)\otimes\left(\int_{0}^{1}f\left(v,\,0\right)dv\right). \end{align*} (ii) If $1/Tb_{1,T}^{q}b_{2,T}\rightarrow0$, $n_{T}/Tb_{1,T}^{q}\rightarrow0$ and $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow\nu\in\left(0,\,\infty\right)$ for some $q\in[0,\,\infty)$ for which $K_{1,q},$ $||\int_{0}^{1}f^{\left(q\right)}\left(u,\,0\right)du||\in[0,\,\infty)$ then $\lim_{T\rightarrow\infty}b_{1,T}^{-q}\mathbb{E}(\widetilde{J}_{T}-J_{T})=\mathsf{B}_{1}+\mathsf{B}_{2}$ where $\mathsf{B}_{1}=-2\pi K_{1,q}\int_{0}^{1}f^{\left(q\right)}\left(u,\,0\right)du$ and $\mathrm{\mathsf{B}}_{2}=2^{-1}\nu\int_{0}^{1}x^{2}K_{2}\left(x\right)\sum_{k=-\infty}^{\infty}\int_{0}^{1}\left(\partial^{2}/\partial u^{2}\right)c\left(u,\,k\right)du.$ (iii) If $n_{T}/Tb_{1,T}^{q}\rightarrow0$, $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow\nu$ and $Tb_{1,T}^{2q+1}b_{2,T}\rightarrow\gamma\in\left(0,\,\infty\right)$ for some $q\in[0,\,\infty)$ for which $K_{1,q},\,||\int_{0}^{1}f^{\left(q\right)}\left(u,\,0\right)du||\in[0,\,\infty)$ , then \begin{align*} \lim_{T\rightarrow\infty} & \mathrm{MSE}\left(Tb_{1,T}b_{2,T},\,\widetilde{J}_{T},\,W\right)=4\pi^{2}\left[\gamma\left(4\pi^{2}\right)^{-1}\mathrm{vec}\left(\mathsf{B}_{1}+\mathsf{B}_{2}\right)'W\mathrm{vec}\left(\mathsf{B}_{1}+\mathsf{B}_{2}\right)\right.\\ & \quad\left.+\int K_{1}^{2}\left(y\right)dy\int K_{2}^{2}\left(x\right)dx\,\mathrm{tr}W\left(I_{p^{2}}+C_{pp}\right)\left(\int_{0}^{1}f\left(u,\,0\right)du\right)\otimes\left(\int_{0}^{1}f\left(v,\,0\right)dv\right)\right]. \end{align*}

The bias expression in part (ii) of Theorem (ref) is different from the corresponding one in casini_hac because $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow\nu\in\left(0,\,\infty\right)$ replaces $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow0$ there. The extra term is $\mathsf{B}_{2}$. This means that both $b_{1,T}$ and $b_{2,T}$ affect the bias as well as the variance. It is therefore possible to consider a joint minimization of the asymptotic MSE with respect to $b_{1,T}$ and $b_{2,T}$. Note that $\mathsf{B}_{2}=0$ when $\int_{0}^{1}\left(\partial^{2}/\partial^{2}u\right)c\left(u,\,k\right)du=0$. The latter occurs when the process is stationary. We now move to the results concerning $\widehat{J}_{T}$.

assumption(i) $\sqrt{T}(\widehat{\beta}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$; (ii) $\sup_{u\in\left[0,\,1\right]}\mathbb{E}||V_{\left\lfloor Tu\right\rfloor }||^{2}<\infty$; (iii) $\sup_{u\in\left[0,\,1\right]}\mathbb{E}\sup_{\beta\in\Theta}$ $||\left(\partial/\partial\beta'\right)V_{\left\lfloor Tu\right\rfloor }\left(\beta\right)||^{2}<\infty$; (iv) $\int_{-\infty}^{\infty}\left|K_{1}\left(y\right)\right|dy,$ $\int_{0}^{1}\left|K_{2}\left(x\right)\right|dx<\infty.$

Assumption (ref)(i)-(iii) is the same as Assumption B in andrews:91. Part (i) is satisfied by standard (semi)parametric estimators. In Section (ref) we relax this assumption and consider nonparametric estimators that satisfy $T^{\vartheta}(\widehat{\beta}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ where $\vartheta\in\left(0,\,1/2\right)$. In order to obtain rate of convergence results we replace Assumption (ref) with the following assumptions.

assumption(i) Assumption (ref) holds with $V_{t,T}$ replaced by \begin{align*} \left(V'_{\left\lfloor Tu\right\rfloor },\,\mathrm{vec}\left(\left(\frac{\partial}{\partial\beta'}V_{\left\lfloor Tu\right\rfloor }\left(\beta_{0}\right)\right)-\mathbb{E}\left(\frac{\partial}{\partial\beta'}V_{\left\lfloor Tu\right\rfloor }\left(\beta_{0}\right)\right)\right)'\right)' & . \end{align*} (ii) $\sup_{u\in\left[0,\,1\right]}\mathbb{E}(\sup_{\beta\in\Theta}||\left(\partial^{2}/\partial\beta\partial\beta'\right)V_{\left\lfloor Tu\right\rfloor }^{\left(a\right)}\left(\beta\right)||^{2})<\infty$ for all $a=1,\ldots,\,p$.
assumptionLet $W_{T}$ denote a $p^{2}\times p^{2}$ weight matrix such that $W_{T}\overset{\mathbb{P}}{\rightarrow}W$.
thmSuppose $K_{1}\left(\cdot\right)\in\boldsymbol{K}_{1}$, $K_{2}\left(\cdot\right)\in\boldsymbol{K}_{2}$, $b_{1,T},\,b_{2,T}\rightarrow0$, $n_{T}\rightarrow\infty,\,n_{T}/Tb_{1,T}\rightarrow0,$ and $1/Tb_{1,T}b_{2,T}\rightarrow0$. We have: (i) If Assumption (ref)-(ref) hold, $\sqrt{T}b_{1,T}\rightarrow\infty$, $b_{2,T}/b_{1,T}\rightarrow\nu\in[0,\,\infty)$ then $\widehat{J}_{T}-J_{T}\overset{\mathbb{P}}{\rightarrow}0$ and $\widehat{J}_{T}-\widetilde{J}_{T}\overset{\mathbb{P}}{\rightarrow}0$. (ii) If Assumption (ref), (ref)-(ref) hold, $n_{T}/Tb_{1,T}^{q}\rightarrow0$, $1/Tb_{1,T}^{q}b_{2,T}\rightarrow0$, $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow\nu\in[0,\,\infty)$ and $Tb_{1,T}^{2q+1}b_{2,T}\rightarrow\gamma\in\left(0,\,\infty\right)$ for some $q\in[0,\,\infty)$ for which $K_{1,q},\,||\int_{0}^{1}f^{\left(q\right)}\left(u,\,0\right)du||\in[0,\,\infty)$, then $\sqrt{Tb_{1,T}b_{2,T}}(\widehat{J}_{T}-J_{T})=O_{\mathbb{P}}\left(1\right)$ and $\sqrt{Tb_{1,T}}(\widehat{J}_{T}-\widetilde{J}_{T})=o_{\mathbb{P}}\left(1\right).$ (iii) Under the conditions of part (ii) with $\nu\in\left(0,\,\infty\right)$ and Assumption (ref), \begin{align*} \lim_{T\rightarrow\infty}\mathrm{MSE}\left(Tb_{1,T}b_{2,T},\,\widehat{J}_{T},\,W_{T}\right)=\lim_{T\rightarrow\infty}\mathrm{MSE}\left(Tb_{1,T}b_{2,T},\,\widetilde{J}_{T},\,W\right) & . \end{align*}

Part (ii) yields the consistency of $\widehat{J}_{T}$ with $b_{1,T}$ only required to be $o\left(Tb_{2,T}\right)$. This rate is slower than the corresponding rate $o\left(T\right)$ of the classical kernel HAC estimators as shown by andrews:91 in his Theorem 1-(b). However, this property is of little practical import because optimal growth rates typically are less than $T^{1/2}$\textemdash for the QS kernel the optimal growth rate is $T^{1/5}$ while it is $T^{1/3}$ for the Barteltt. Part (ii) of the theorem presents the rate of convergence of $\widehat{J}_{T}$ which is $\sqrt{Tb_{2,T}b_{1,T}}$, the same rate shown by casini_hac when $b_{2,T}^{2}/b_{1,T}^{q}\rightarrow0$. Thus, the presence of the bias term $\mathsf{B}_{2}$ does not alter the rate of convergence. In Section (ref), we compare the rate of convergence of $\widehat{J}_{T}$ with optimal bandwidths $(\widetilde{b}_{1,T}^{\mathrm{opt}},\,\widetilde{b}_{2,T}^{\mathrm{opt}})$ from the joint MSE criterion (ref) with that using $(b_{1,T}^{\mathrm{opt}},\,b_{2,T}^{\mathrm{opt}})$ from the sequential MSE criterion (ref), and with that of the classical HAC estimators when the corresponding optimal bandwidths are used.

Optimal Bandwidths and Kernels

We consider the optimal bandwidths $\left(\widetilde{b}_{1,T}^{\mathrm{opt}},\,\widetilde{b}_{2,T}^{\mathrm{opt}}\right)$ and kernels $\widetilde{K}_{1}^{\mathrm{opt}}$ and $\widetilde{K}_{2}^{\mathrm{opt}}$ that minimize the global asymptotic relative MSE (ref) given by

align*[align* omitted — 270 chars of source]

Let $\Xi_{1,1}=-K_{1,q}$, $\Xi_{1,2}=\left(4\pi\right)^{-1}\int_{0}^{1}x^{2}K_{2}\left(x\right)dx,$ $\Xi_{2}=\int K_{1}^{2}\left(y\right)dy\int_{0}^{1}K_{2}^{2}\left(x\right)dx,$

align*[align* omitted — 331 chars of source]
thmSuppose Assumption (ref), (ref)-(ref) hold, $\int_{0}^{1}||f^{\left(2\right)}\left(u,\,0\right)||du<\infty$, $\mathrm{vec}\left(\Delta_{1,1,0}\right)'W$ $\mathrm{vec}\left(\Delta_{1,1,0}\right)>0$, $\mathrm{vec}\left(\Delta_{1,2}\right)'W\mathrm{vec}\left(\Delta_{1,2}\right)>0$ and $W$ is positive definite. Then, $\lim_{T\rightarrow\infty}\mathrm{ReMSE}$ $(\widehat{J}_{T}\left(b_{1,T},\,b_{2,T}\right)J^{-1},\,W_{T})$ is jointly minimized by \begin{align*} \widetilde{b}_{1,T}^{\mathrm{opt}} & =0.46\left(\frac{\mathrm{vec}\left(\Delta_{1,2}\right)'W\mathrm{vec}\left(\Delta_{1,2}\right)}{\left(\mathrm{vec}\left(\Delta_{1,1,0}\right)'W\mathrm{vec}\left(\Delta_{1,1,0}\right)\right)^{5}}\right)^{1/24}T^{-1/6},\\ \widetilde{b}_{2,T}^{\mathrm{opt}} & =3.56\left(\frac{\mathrm{vec}\left(\Delta_{1,1,0}\right)'W\mathrm{vec}\left(\Delta_{1,1,0}\right)}{\left(\mathrm{vec}\left(\Delta_{1,2}\right)'W\mathrm{vec}\left(\Delta_{1,2}\right)\right)^{5}}\right)^{1/24}T^{-1/6}. \end{align*} Furthermore, the optimal kernels are given by $K_{1}^{\mathrm{opt}}=K_{1}^{\mathrm{QS}}$ and $K_{2}^{\mathrm{opt}}\left(x\right)=6x\left(1-x\right)$ for $x\in\left[0,\,1\right]$.

The requirement $\int_{0}^{1}||f^{\left(2\right)}\left(u,\,0\right)||du<\infty$ is not stringent and reduces to the one used by andrews:91 when $\left\{ V_{t,T}\right\} $ is stationary. Note that $\Delta_{1,1,0}$ accounts for the relative variation of $\int_{0}^{1}f\left(u,\,\omega\right)$ around $\omega=0$ whereas $\Delta_{1,2}$ accounts for the relative time variation (i.e., nonstationarity). The theorem states that as $\Delta_{1,1,0}$ increases $\widetilde{b}_{1,T}^{\mathrm{opt}}$ becomes smaller while $\widetilde{b}_{2,T}^{\mathrm{opt}}$ becomes larger. This is intuitive. With more variation around the zero frequency, more smoothing is required over the frequency direction and less over the time direction. Conversely, the more nonstationary is the data the more smoothing is required over the time direction (i.e., $\widetilde{b}_{2,T}^{\mathrm{opt}}$ is smaller and the optimal block length $T\widetilde{b}_{2,T}^{\mathrm{opt}}$ smaller) relative to the frequency direction. Both optimal bandwidths $(\widetilde{b}_{1,T}^{\mathrm{opt}},\,\widetilde{b}_{2,T}^{\mathrm{opt}})$ have the same order $O(T^{-1/6}).$ We can compare it with $b_{1,T}^{\mathrm{opt}}=O(T^{-4/25})$ and $\overline{b}_{2,T}^{\mathrm{opt}}=O(T^{-1/5})$ resulting from the sequential MSE criterion in casini_hac. The latter leads to a slightly smaller block length relative to the global criterion (ref) {[}i.e., $O(T\overline{b}_{2,T}^{\mathrm{opt}})<O(T\widetilde{b}_{2,T}^{\mathrm{opt}})${]}. Since $K_{2}$ applies overlapping smoothing, a smaller block length is beneficial if there is substantial nonstationarity. On the same note, a smaller block length is less exposed to low frequency contamination since it allows to better account for nonstationarity. The rate of convergence when the optimal bandwidths are used is $O(T^{1/3})$ which is sightly faster than the corresponding rate of convergence when $(b_{1,T}^{\mathrm{opt}},\,b_{2,T}^{\mathrm{opt}})$. The latter is $O(T^{0.32})$, so the difference is small.

Data-Dependent Bandwidths

In this section we consider estimators $\widehat{J}_{T}$ that use bandwidths $b_{1,T}$ and $b_{2,T}$ whose values are determined via data-dependent methods. We use the “plug-in” method which is characterized by plugging-in estimates of unknown quantities into a formula for an optimal bandwidth parameter (i.e., the expressions for $\widetilde{b}_{1,T}^{\mathrm{opt}}$ and $\widetilde{b}_{2,T}^{\mathrm{opt}}$). Section (ref) discusses the implementation of the automatic bandwidths, while Section (ref) presents the corresponding theoretical results.

Implementation

The first step for the construction of data-dependent bandwidth parameters is to specify $p$ univariate parametric models for $V_{t}=(V_{t}^{\left(1\right)},\ldots,\,V_{t}^{\left(p\right)})$. The second step involves the estimation of the parameters of the parametric models. Here standard estimation methods are local least-squares (LS) (i.e., LS method applied to rolling windows) and nonparametric kernel methods. Let

align*[align* omitted — 425 chars of source]

In a third step, we replace the unknown parameters in $\phi_{1}$ and $\phi_{2}$ with corresponding estimates. Such estimates $\widehat{\phi}_{1}$ and $\widehat{\phi}_{2}$ are then substituted into the expression for $\widetilde{b}_{1,T}^{\mathrm{opt}}$ and $\widetilde{b}_{2,T}^{\mathrm{opt}}$ to yield

align[align omitted — 142 chars of source]

In practice, a reasonable candidate to be used as an approximating parametric model is the first order autoregressive {[}AR(l){]} model for $\{V_{t}^{\left(r\right)}\},\,r=1,\ldots,\,p$ (with different parameters for each $r$) or a first order vector autoregressive {[}VAR(l){]} model for $\{V_{t}\}$ {[}see andrews:91{]}. However, in our context it is reasonable to allow the parameters to be time-varying. For parsimony, we consider time-varying AR(1) models with no break points in the spectrum (i.e., $V_{t}^{\left(r\right)}=a_{1}\left(t/T\right)V_{t-1}^{\left(r\right)}+u_{t}^{\left(r\right)}$).

The use of $p$ univariate parametric models requires $W$ to be a diagonal matrix. This leads to $\phi_{1}=\phi_{1,1}/\phi_{1,2}^{5}$ and $\phi_{2}=\phi_{1,2}/\phi_{1,1}^{5}$ where

align*[align* omitted — 467 chars of source]

The usual choice is $W^{\left(r,r\right)}=1$ for all $r$. An estimate of $f^{\left(r,r\right)}\left(u,\,0\right)$ $\left(r=1,\ldots,\,p\right)$ is $\widehat{f}^{\left(r,r\right)}\left(u,\,0\right)=\left(2\pi\right)^{-1}(\widehat{\sigma}^{\left(r\right)}\left(u\right))^{2}(1-\widehat{a}_{1}^{\left(r\right)}\left(u\right))^{-2}$ while $f^{\left(2\right)\left(r,r\right)}\left(u,\,0\right)$ can be estimated by $\widehat{f}^{\left(2\right)\left(r,r\right)}\left(u,\,0\right)=3\pi^{-1}$ $((\widehat{\sigma}^{\left(r\right)}\left(u\right))^{2}\widehat{a}_{1}^{\left(r\right)}\left(u\right))(1-\widehat{a}_{1}^{\left(r\right)}\left(u\right))^{-4}$ where $\widehat{a}_{1}^{\left(r\right)}\left(u\right)$ and $\widehat{\sigma}^{\left(r\right)}\left(u\right)$ are the LS estimates computed using local data to the left of $u=t/T$:

align[align omitted — 493 chars of source]

where $n_{2,T}\rightarrow\infty$. More complex is the estimation of $\overline{\Delta}_{1,2,1}\triangleq\sum_{k=-\infty}^{\infty}\int_{0}^{1}\left(\partial^{2}/\partial u^{2}\right)c\left(u,\,k\right)$ because it involves the second partial derivative of $c\left(u,\,k\right).$ We need a further parametric assumption. We assume that the parameters of the approximating time-varying AR(1) models change slowly such that the smoothness of $f\left(\cdot,\,\omega\right)$ and thus of $c\left(\cdot,\,\cdot\right)$ is the same to the one that would arise if $a_{1}\left(u\right)=0.8\left(\cos1.5+\cos4\pi u\right)$ and $\sigma\left(u\right)=\sigma=1$ for all $u\in\left[0,\,1\right]$ {[}cf. dahlhaus:12{]}. Then, $\Delta_{1,2,1}^{\left(r,r\right)}\left(u,\,k\right)\triangleq\left(\partial^{2}/\partial u^{2}\right)c^{\left(r,r\right)}\left(u,\,k\right)$ can be computed analytically:

align*[align* omitted — 498 chars of source]

An estimate of $\Delta_{1,2,1}^{\left(r,r\right)}\left(u,\,k\right)$ is given by

align*[align* omitted — 563 chars of source]

where $\left[S_{\omega}\right]$ is the cardinality of $S_{\omega}$ and $\omega_{s+1}>\omega_{s}$ with $\omega_{1}=-\pi,\,\omega_{\left[S_{\omega}\right]}=\pi.$ In our simulations, we use $S_{\omega}=\left\{ -\pi,\,-3,\,-2,\,-1,\,0,\,1,\,2,\,3,\,\pi\right\} $. We can average over $u$ and sum over $k$ to obtain an estimate of $\overline{\Delta}_{1,2,1}^{\left(r,r\right)}:$ $\widehat{\overline{\Delta}}_{1,2,1}^{\left(r,r\right)}=\sum_{k=-\left\lfloor T^{1/6}\right\rfloor }^{\left\lfloor T^{1/6}\right\rfloor }\frac{n_{3,T}}{T}\sum_{j=0}^{\left\lfloor T/n_{3,T}\right\rfloor }\widehat{\Delta}_{1,2,1}^{\left(r,r\right)}\left(jn_{T}/T,\,k\right)$ where the number of summands over $k$ grows at the same rate as $1/\widetilde{b}_{1,T}^{\mathrm{opt}}$; a different choice is allowed as long as it grows at a slower rate than $T^{2/5}$ but our sensitivity analysis does not indicate significant changes.

Then, $\widehat{b}_{1,T}=0.46\widehat{\phi}_{1}^{1/24}T^{-1/6}$ and $\widehat{b}_{2,T}=3.56\widehat{\phi}_{2}^{1/24}T^{-1/6}$ where $\widehat{\phi}_{1}=\widehat{\phi}_{1,1}/\widehat{\phi}_{1,2}^{5}$, $\widehat{\phi}_{2}=\widehat{\phi}_{1,2}/\widehat{\phi}_{1,1}^{5},$

align*[align* omitted — 913 chars of source]

For most of the results below we can take $n_{3,T}=n_{2,T}=n_{T}.$

Theoretical Results

We establish results corresponding to Theorem (ref) for the estimator $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ that uses $\widehat{b}_{1,T}$ and $\widehat{b}_{2,T}$. We restrict the class of admissible kernels to the following,

align*[align* omitted — 948 chars of source]

Let $\widehat{\theta}$ denote the estimator of the parameter of the approximate (time-varying) parametric model(s) introduced above {[}i.e., $\widehat{\theta}=(\int_{0}^{1}\widehat{a}_{1}\left(u\right)du,\,\int_{0}^{1}\widehat{\sigma}_{1}^{2}\left(u\right)du,\ldots,\,\int_{0}^{1}\widehat{a}_{p}^{2}\left(u\right)du,\,\int_{0}^{1}\widehat{\sigma}_{p}^{2}\left(u\right)du)'${]}. Let $\theta^{*}$ denote the probability limit of $\widehat{\theta}$. $\widehat{\phi}_{1}$ and $\widehat{\phi}_{2}$ are the values of $\phi_{1}$ and $\phi_{2}$, receptively, with $\widehat{\theta}$ instead of $\theta$. The probability limits of $\widehat{\phi}_{1}$ and $\widehat{\phi}_{2}$ are denoted by $\phi_{1,\theta^{*}}$ and $\phi_{2,\theta^{*}},$ respectively.

assumption(i) $\widehat{\phi}_{1}=O\mathbb{_{P}}\left(1\right),$ $1/\widehat{\phi}_{1}=O\mathbb{_{P}}\left(1\right),$ $\widehat{\phi}_{2}=O\mathbb{_{P}}\left(1\right),$ and $1/\widehat{\phi}_{2}=O\mathbb{_{P}}\left(1\right)$; (ii) $\inf\bigl\{ T/n_{3,T},\,\sqrt{n_{2,T}}\bigr\}((\widehat{\phi}_{1}-\phi_{1,\theta^{*}}),\,(\widehat{\phi}_{2}-\phi_{2,\theta^{*}}))'=O_{\mathbb{P}}\left(1\right)$ for some $\phi_{1,\theta^{*}},\,\phi_{2,\theta^{*}}\in\left(0,\,\infty\right)$ where $n_{2,T}/T+n_{3,T}/T\rightarrow0,$ $n_{2,T}^{10/6}/T\rightarrow[c_{2},\,\infty),$ $n_{3,T}^{10/6}/T\rightarrow[c_{3},\,\infty)$ with $0<c_{2},\,c_{3}<\infty$; (iii) $\sup_{u\in\left[0,\,1\right]}\lambda_{\max}(\Gamma_{u}\left(k\right))$ $\leq C_{3}k^{-l}$ for all $k\geq0$ for some $C_{3}<\infty$ and some $l>3$, where $q$ is as in $\boldsymbol{K}_{3}$; (iv) $|\omega_{s+1}-\omega_{s}|=O\left(T^{-1}\right)$ and $\left[S_{\omega}\right]=O\left(T\right)$; (v) $\boldsymbol{K}_{2}$ includes kernels that satisfy $|K_{2}\left(x\right)-K_{2}\left(y\right)|\leq C_{4}\left|x-y\right|$ for all $x,\,y\in\mathbb{R}$ and some constant $C_{4}<\infty$.

Parts (i) and (v) are sufficient for the consistency of $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T}).$ Parts (ii)-(iii) and (iv)-(v) are required for the rate of convergence and MSE results. Note that $\phi_{1,\theta^{*}}$ and $\phi_{2,\theta^{*}}$ coincide with the optimal values $\phi_{1}$ and $\phi_{2}$, respectively, only when the approximate parametric model indexed by $\theta^{*}$ corresponds to the true data-generating mechanism.

Let $b_{\theta_{1},T}=0.46\phi_{1,\theta^{*}}^{1/24}T^{-1/6}$ and $b_{\theta_{2},T}=3.56\phi_{2,\theta^{*}}^{1/24}T^{-1/6}$. The asymptotic properties of $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ are shown to be equivalent to those of $\widehat{J}_{T}(b_{\theta_{1},T},\,b_{\theta_{2},T})$ where the theoretical properties of the latter follow from Theorem (ref).

thmSuppose $K_{1}\left(\cdot\right)\in\boldsymbol{K}_{3}$, $q$ is as in $\boldsymbol{K}_{3}$, $K_{2}\left(\cdot\right)\in\boldsymbol{K}_{2}$, $n_{T}\rightarrow\infty,\,n_{T}/Tb_{\theta_{1},T}\rightarrow0,$ and $||\int_{0}^{1}f^{\left(q\right)}\left(u,\,0\right)du||<\infty$. Then, we have: (i) If Assumption (ref)-(ref) and (ref)-(i,v) hold and $n_{3,T}=n_{2,T}=n_{T},$ then $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})-J_{T}\overset{\mathbb{P}}{\rightarrow}0$. (ii) If Assumption (ref), (ref)-(ref) and (ref)-(ii,iii,iv,v) hold $q\leq2$ and $n_{T}/Tb_{\theta_{1},T}^{2}\rightarrow0,$ then $\sqrt{Tb_{\theta_{1},T}b_{\theta_{2},T}}$ $(\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})-J_{T})=O_{\mathbb{P}}\left(1\right)$ and $\sqrt{Tb_{\theta_{1},T}b_{\theta_{2},T}}(\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})-\widehat{J}_{T}(b_{\theta_{1},T},\,b_{\theta_{2},T}))=o_{\mathbb{P}}\left(1\right)$. (iii) If Assumption (ref), (ref)-(ref) and (ref)-(ii,iii,iv,v) hold, then \begin{align*} \lim_{T\rightarrow\infty} & \mathrm{MSE}\left(T^{2/3},\,\widehat{J}_{T}\left(\widehat{b}_{1,T},\,\widehat{b}_{2,T}\right),\,W_{T}\right)\\ & =\lim_{T\rightarrow\infty}\mathrm{MSE}\left(Tb_{\theta_{1},T}b_{\theta_{2},T},\,\widehat{J}_{T}\left(b_{\theta_{1},T},\,b_{\theta_{2},T}\right),\,W_{T}\right). \end{align*}

When the chosen parametric model indexed by $\theta$ is correct, it follows that $\phi_{1,\theta^{*}}=\phi_{1}$, $\phi_{2,\theta^{*}}=\phi_{2}$, $\widehat{\phi}_{1}\overset{\mathbb{P}}{\rightarrow}\phi_{1}$ and $\widehat{\phi}_{2}\overset{\mathbb{P}}{\rightarrow}\phi_{2}$. The theorem then implies that $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ exhibits the same optimality properties presented in Theorem (ref).

Consistent LRV in the Context of Nonparametric Parameter Estimates

We relax the assumption that $\{V_{t}(\widehat{\beta})\}$ is a function of a semiparametric estimator $\widehat{\beta}$ satisfying $\sqrt{T}(\widehat{\beta}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$. This holds, for example, in the linear regression model estimated by least-squares where $V_{t}(\widehat{\beta})=\widehat{e}_{t}x_{t}$ with $\{\widehat{e}_{t}\}$ being the fitted residuals and $\{x_{t}\}$ being a vector of regressors. However, there are many HAR inference contexts where one needs an estimate of the LRV based on a sequence of observations $\{V_{t}(\widehat{\beta}_{\mathrm{np}})\}$ where $\widehat{\beta}_{\mathrm{np}}$ is a nonparametric estimator that satisfies $T^{\vartheta}(\widehat{\beta}_{\mathrm{np}}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ for some $\vartheta\in\left(0,\,1/2\right)$. For example, in forecasting one needs an estimate of the LRV to obtain a pivotal asymptotic distribution for forecast evaluation tests while one has access to a sequence $\{V_{t}(\widehat{\beta}_{\mathrm{np}})\}$ obtained from nonparametric estimation using some in-sample. Given that nonparametric methods have received a great deal of attention in applied work lately, it is useful to extend the theory of HAC and DK-HAC estimators to these settings. We consider the HAC estimators in Section (ref) and the DK-HAC estimators in Section (ref).

Classical HAC Estimators

We show that the classical HAC estimators that use the data-dependent bandwidths suggested in andrews:91 remain valid when $\sqrt{T}(\widehat{\beta}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ is replaced by $T^{\vartheta}(\widehat{\beta}_{\mathrm{np}}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ for some $\vartheta\in\left(0,\,1/2\right)$. We work under the same assumptions as in andrews:91. Under stationarity we have $\Gamma_{u}\left(k\right)=\Gamma\left(k\right)$ and $\kappa_{V,\left\lfloor Tu\right\rfloor }^{\left(a,b,c,d\right)}\left(k,\,s,\,l\right)=\kappa_{V,0}^{\left(a,b,c,d\right)}\left(k,\,s,\,l\right)$ for any $u\in\left[0,\,1\right]$.

assumption$\left\{ V_{t}\right\} $ is a mean-zero, fourth-order stationary sequence with $\sum_{k=-\infty}^{\infty}\left\Vert \Gamma\left(k\right)\right\Vert <\infty$ and $\sum_{k=-\infty}^{\infty}\sum_{s=-\infty}^{\infty}\sum_{l=-\infty}^{\infty}|\kappa_{V,0}^{\left(a,b,c,d\right)}\left(k,\,s,\,l\right)|<\infty$ $\forall a,\,b,\,c,\,d\leq p$.
assumption(i) $T^{\vartheta}(\widehat{\beta}_{\mathrm{np}}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ for some $\vartheta\in\left(0,\,1/2\right)$; (ii) $\sup_{t\geq1}\mathbb{E}\left\Vert V_{t}\right\Vert ^{2}<\infty$; (iii) $\sup_{t\geq1}\mathbb{E}\sup_{\beta\in\Theta}\left\Vert \left(\partial/\partial\beta\right)V_{t}\left(\beta\right)\right\Vert ^{2}<\infty$; (iv) $\int\left|K_{1}\left(y\right)\right|dy<\infty$.
assumption(i) Assumption (ref) holds with $V_{t}$ replaced by \begin{align*} \left(V'_{t},\,\mathrm{vec}\left(\left(\frac{\partial}{\partial\beta'}V_{t}\left(\beta_{0}\right)\right)-\mathbb{E}\left(\frac{\partial}{\partial\beta'}V_{t}\left(\beta_{0}\right)\right)\right)'\right)' & . \end{align*} (ii) $\sup_{t\geq1}\mathbb{E}(\sup_{\beta\in\Theta}||\left(\partial^{2}/\partial\beta\partial\beta'\right)V_{t}^{\left(a\right)}\left(\beta\right)||^{2})<\infty$ for all $a=1,\ldots,\,p$.

Let $\widehat{b}_{\mathrm{Cla,}1,T}=\left(qK_{1,q}^{2}\widehat{\alpha}\left(q\right)T/\int K_{1}^{2}\left(x\right)dx\right)^{-1/\left(2q+1\right)}$. The form of $\widehat{\alpha}\left(q\right)$ depends on the approximating parametric model for $\{V_{t}^{\left(r\right)}\}$. andrews:91 considered stationarity AR(1) models for $\{V_{t}^{\left(r\right)}\}$, which result in

align[align omitted — 819 chars of source]

Let

align[align omitted — 636 chars of source]
assumption$\widehat{\alpha}\left(q\right)=O_{\mathbb{P}}\left(1\right)$ and $1/\widehat{\alpha}\left(q\right)=O_{\mathbb{P}}\left(1\right)$.
thmSuppose $K_{1}\left(\cdot\right)\in\boldsymbol{K}_{3,\mathrm{Cla}}$, $||f^{\left(q\right)}||<\infty$, $q>\left(1/\vartheta-1\right)/2$, and Assumption (ref)-(ref) hold, then $\widehat{J}_{\mathrm{Cla,}T}(\widehat{b}_{\mathrm{Cla,1,}T})-J_{T}\overset{\mathbb{P}}{\rightarrow}0$.

DK-HAC Estimators

We extend the consistency result in Theorem (ref)-(i) assuming that $\widehat{V}_{t}=V_{t}(\widehat{\beta}_{\mathrm{np}}).$ Thus, we replace Assumption (ref) by the following.

assumption(i) $T^{\vartheta}(\widehat{\beta}_{\mathrm{np}}-\beta_{0})=O_{\mathbb{P}}\left(1\right)$ for some $\vartheta\in\left(0,\,1/2\right)$; (ii)-(iv) from Assumption (ref) continue to hold.
thmSuppose $K_{1}\left(\cdot\right)\in\boldsymbol{K}_{3}$, $K_{2}\left(\cdot\right)\in\boldsymbol{K}_{2}$, $T^{\vartheta}b_{\theta_{1},T}b_{\theta_{2},T}\rightarrow\infty$, $n_{T}\rightarrow\infty,\,n_{T}/Tb_{\theta_{1},T}\rightarrow0,$ and $||\int_{0}^{1}f^{\left(q\right)}\left(u,\,0\right)du||<\infty$. If Assumption (ref)-(ref), (ref)-(i,v), (ref) hold and $n_{3,T}=n_{2,T}=n_{T},$ then $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})-J_{T}\overset{\mathbb{P}}{\rightarrow}0$.

Theorem (ref)-(ref) require different conditions on the parameter $\vartheta$ that controls the rate of convergence of the nonparametric estimator. In Theorem (ref) this conditions depends on $q$ while in Theorem (ref) it depends on $q$ through $b_{\theta_{1},T}$ and also on the smoothing over time through $b_{\theta_{2},T}$. For both HAC and DK-HAC estimators, the condition allows for standard nonparametric estimators with optimal nonparametric convergence rate.

Small-Sample Evaluations

In this section, we conduct a Monte Carlo analysis to evaluate the performance of the DK-HAC estimator based on the data-dependent bandwidths determined via the joint MSE criterion (ref). We consider HAR tests in the linear regression model as well as HAR tests for forecast breakdown, i.e., the test of giacomini/rossi:09. The linear regression models have an intercept and a stochastic regressor. We focus on the $t$-statistics $t_{r}=\sqrt{T}(\widehat{\beta}^{\left(r\right)}-\beta_{0}^{\left(r\right)})/\sqrt{\widehat{J}_{T}^{\left(r,r\right)}}$ where $\widehat{J}_{T}$ is an estimate of the limit of $\mathrm{Var}(\sqrt{T}(\widehat{\beta}-\beta_{0}))$ and $r=1,\,2$. $t_{1}$ is the $t$-statistic for the parameter associated to the intercept while $t_{2}$ is associated to the stochastic regressor $x_{t}$. We omit the discussion of the results concerning to the $F$-test since they are qualitatively similar. Three basic regression models are considered. We run a $t$-test on the intercept in model M1 and a $t$-test on the coefficient of the stochastic regressor in model M2 and M3. The models are based on,

align[align omitted — 142 chars of source]

for the $t$-test on the intercept (i.e., $t_{1}$) and

align[align omitted — 160 chars of source]

for the $t$-test on $\beta_{0}^{\left(2\right)}$ (i.e., $t_{2}$) where $\delta=0$ under the null. We consider the following models:

itemize• M1: $e_{t}=0.4e_{t-1}+u_{t},\,u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,0.5\right),$ $x_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(1,\,1\right)$, $\beta_{0}^{\left(1\right)}=0$ and $\beta_{0}^{\left(2\right)}=1.$ • M2: $e_{t}=0.4e_{t-1}+u_{t},\,u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right),$ $x_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(1,\,1\right)$, and $\beta_{0}^{\left(1\right)}=\beta_{0}^{\left(2\right)}=0.$ • M3: segmented locally stationary errors $e_{t}=\rho_{t}e_{t-1}+u_{t},\,u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right),\,\rho_{t}=\max\{0,\,-1$ $\left(\cos\left(1.5-\cos\left(5t/T\right)\right)\right)\}$\footnote{That is, $\rho_{t}$ varies smoothly between 0 and 0.8071.} for $t\notin\left(4T/5+1,\,4T/5+h\right)$ and $e_{t}=0.99e_{t-1}+u_{t},\,u_{t}\sim\mathrm{\mathrm{i.i.d.}}$ $\mathscr{N}\left(0,\,1\right)$ for $t\in\left(4T/5+1,\,4T/5+h\right)$ where $h=10$ for $T=200$ and $h=30$ for $T=400$, and $x_{t}=1+0.6x_{t-1}+u_{X,t},\,u_{X,t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$.

Finally, we consider model M4 which we use to investigate the performance of giacomini/rossi:09's (2009) test for forecast breakdown. Suppose we want to forecast a variable $y_{t}$ generated by $y_{t}=\beta_{0}^{\left(1\right)}+\beta_{0}^{\left(2\right)}x_{t-1}+e_{t}$ where $x_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(1,\,1.2\right)$ and $e_{t}=0.3e_{t-1}+u_{t}$ with $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$. For a given forecast model and forecasting scheme, the test of giacomini/rossi:09 detects a forecast breakdown when the average of the out-of-sample losses differs significantly from the average of the in-sample losses. The in-sample is used to obtain estimates of $\beta_{0}^{\left(1\right)}$ and $\beta_{0}^{\left(2\right)}$ which are in turn used to construct out-of-sample forecasts $\widehat{y}_{t}=\widehat{\beta}_{0}^{\left(1\right)}+\widehat{\beta}_{0}^{\left(2\right)}x_{t-1}$. We set $\beta_{0}^{\left(1\right)}=\beta_{0}^{\left(2\right)}=1.$ We consider a fixed forecasting scheme. GR's (2009) test statistic is defined as $t^{\mathrm{GR}}\triangleq\sqrt{T_{n}}\overline{SL}/\sqrt{\widehat{J}_{SL}}$ where $\overline{SL}\triangleq T_{n}^{-1}\sum_{t=T_{m}}^{T-\tau}SL_{t+\tau}$, $SL_{t+\tau}$ is the surprise loss at time $t+\tau$ (i.e., the difference between the time $t+\tau$ out-of-sample loss and in-sample loss, $SL_{t+\tau}=L_{t+\tau}-\overline{L}_{t+\tau}$), $T_{n}$ is the sample size in the out-of-sample, $T_{m}$ is the sample size in the in-sample and $\widehat{J}_{SL}$ is a LRV estimator. We restrict attention to one-step ahead forecasts (i.e., $\tau=1$). Under $H_{1}:\,\mathbb{E}(\overline{SL})\neq0$, we have $y_{t}=1+x_{t-1}+\delta x_{t-1}\mathbf{1}\left\{ t>T_{1}^{0}\right\} +e_{t}$, where $T_{1}^{0}=T\lambda_{1}^{0}$ with $\lambda_{1}^{0}=0.7$. Under this specification there is a break in the coefficient associated with $x_{t-1}$. Thus, there is a forecast instability or failure and the test of giacomini/rossi:09 should reject $H_{0}$. We set $T_{m}=0.4T$ and $T_{n}=0.6T$.

Throughout our study we consider the following LRV estimators: $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ with $K_{1}^{\mathrm{opt}},\,K_{2}^{\mathrm{opt}}$ and automatic bandwidths; the same $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ with in addition the prewhitening of casini/perron_PrewhitedHAC; casini_hac's (2021) DK-HAC with $K_{1}^{\mathrm{opt}},\,K_{2}^{\mathrm{opt}}$ and automatic bandwidths from the sequential method; DK-HAC with prewhitening; Newey and West's (1987) HAC estimator with the automatic bandwidth as proposed in newey/west:94; the same with the prewhitening procedure of andrews/monahan:92; Newey-West estimator with the fixed-$b$ method of Kiefer/vogelsang/bunzel:00; the Empirical Weighted Cosine (EWC) of lazarus/lewis/stock:17.\footnote{We have excluded Andrews's (1991) HAC estimator since its performance is similar to that of the Newey-West estimator.} casini/perron_PrewhitedHAC proposed three methods related to prewhitening: (1) $\widehat{J}_{T,\mathrm{pw},1}$ uses a stationary model to whiten the data; (2) $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS}}$ uses a nonstationary model to whiten the data; (3) $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS},\mu}$ is the same as $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS}}$ but it adds a time-varying intercept in the VAR to whiten the data.

We set $n_{T}=T^{0.66}$ as explained in casini_hac and $n_{2,T}=n_{3,T}=n_{T}.$ Simulation results for models involving ARMA, ARCH and heteroskedastic errors are not discussed here because the results are qualitatively equivalent. The significance level is $\alpha=0.05$ throughout.

Empirical Sizes of HAR Inference Tests

Table (ref)-(ref) report the rejection rates for model M1-M4. As a general pattern, we confirm previous evidence that Newey-West's (1987) HAC estimator leads to $t$-tests that are oversized when the data are stationary and there is substantial dependence {[}cf. model M1-M2{]}. This is a long-discussed issue in the literature. Newey-West with prewhitening is often effective in reducing the oversize problem under stationarity. However, the simulation results below and in the literature show that the prewhitened Newey-West-based tests can be oversized when there is high serial dependence. Among the existing methods, the rejection rates of the Newey-West-based tests with fixed-$b$ are accurate in model M1-M2. Overall, the results in the literature along with those in casini_hac and casini/perron_PrewhitedHAC showed that under stationarity the original fixed-$b$ method of KVB is the method which is in general the least oversized across different degrees of dependence among all existing methods. EWC performs similarly to KVB's fixed-$b$. Among the recently introduced DK-HAC estimators, Table (ref) reports evidence that the non-prewhitened DK-HAC from casini_hac leads to HAR tests that are a bit oversized whereas the tests based on the new DK-HAC with simultaneous data-dependent bandwidths, $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$, are more accurate. The results also show that the tests based on the prewhitened DK-HAC estimators are competitive with those based on KVB's fixed-$b$ in controlling the size. In particular, tests based on $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ with prewhitening are more accurate than those using the prewhitened DK-HAC with sequential data-dependent bandwidths. Since $\widehat{J}_{T,\mathrm{pw},1}$ uses a stationarity VAR model to whiten the data, it works as well as $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS}}$ and $\widehat{J}_{T,\mathrm{pw},\mathrm{SLS},\mu}$ when stationarity actually holds, as documented in Table (ref).

Turning to nonstationary data and to the GR test, Table (ref) casts concerns about the finite-sample performance of existing methods in this context. For both model M3 and M4, existing long-run variance estimators lead to HAR tests that have either size equal or close to zero. The methods that use long bandwidths (i.e., many lagged autocovariances) such as KVB's fixed-$b$ and EWC suffer most from this problem relative to using the Newey-West estimator. This is demonstrated in casini/perron_Low_Frequency_Contam_Nonstat:2020 who showed theoretically that nonstationarity induces positive bias for each sample autocovariance. That bias is constant across lag orders. Since existing LRV estimators are weighted sum of sample autocovariances (or weighted sum of periodogram ordinates), the more lags are included the larger is the positive bias. Thus, LRV estimators are inflated and HAR tests have lower rejection rates than the significance level. As we show below, this mechanism has consequences for power as well. In model M3-M4, tests based on the non-prewhitened DK-HAC $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ performs well although tests based on the prewhitened DK-HAC are more accurate. $\widehat{J}_{T,\mathrm{pw},1}$ leads to tests that are slightly less accurate because it uses stationarity and when the latter is violated its performance is affected. In model M4, KVB's fixed-$b$ and prewhitened DK-HAC are associated to rejection rates relatively close to the significance level.

In summary, the prewhitened DK-HAC estimators yield $t$-test in regression models with rejection rates that are relatively close to the exact size. The DK-HAC with simultaneous bandwidths developed in Section (ref) performs better (i.e., the associated null rejection rates are closer to the significance level and approach it from below) than the corresponding DK-HAC estimators with sequential bandwidths when the data are stationary. This is in accordance with our theoretical results. Also for nonstationary data the simultaneous bandwidths perform in general better than the sequential bandwidths, though the margin is smaller. The non-prewhitened DK-HAC can lead to oversized $t$-test on the intercept if there is high dependence. Our results confirm the oversize problem induced by the use of the Newey-West estimators documented in the literature under stationarity. Fixed-$b$ HAR tests control the size well when the data are stationary but can be severely undersized under nonstationarity, a problem that also affects tests based on the Newey-West. Thus, prewhitened DK-HAC estimators are competitive to fixed-$b$ methods under stationarity and they perform well also when the data are nonstationary.

Empirical Power of HAR Inference Tests

For model M1-M4 we report the power results in Table (ref)-(ref). The sample size is $T=200$. For model M1, tests based on the Newey and West's (1987) HAC and on the non-prewhitened DK-HAC estimators have the highest power but they were more oversized than the tests based on other methods. KVB's fixed-$b$ leads to $t$-tests that sacrifices some power relative to using the prewhitened DK-HAC estimators while EWC-based tests have lower power locally to $\delta=0$ (i.e., $\delta=0.1$ and $0.2$). In model M2, a similar pattern holds. HAR tests normalized by either classical HAC or DK-HAC estimators have similarly good power while HAR tests based on KVB's fixed-$b$ have relatively less power. In model M3, the best power is achieved with Newey-West's (1987) HAC estimator followed by $\widehat{J}_{T}(\widehat{b}_{1,T},\,\widehat{b}_{2,T})$ and EWC. Using KVB's fixed-$b$ leads to large power losses. In model M4, it appears that all versions of the classical HAC estimators of newey/west:87, the KVB's fixed-$b$ and EWC lead to $t$-tests that have, essentially, zero power for all $\delta$. In contrast, the $t$-test standardized by the DK-HAC estimators have good power. Among the latter DK-HAC estimators, the ones that use the sequential bandwidths have slightly higher power but they margin is very small. This follows from the usual size-power trade-off since the simultaneous bandwidths led to tests that have more accurate size control.

The severe power problems of tests based on classical HAC estimators, KVB's fixed-$b$ and EWC can be simply reconciled with the fact that under the alternative hypotheses the spectrum of $V_{t}$ is not constant. Existing estimators estimate an average of a time-varying spectrum. Because of this instability in the spectrum, they overestimate the dependence in $V_{t}$. casini/perron_Low_Frequency_Contam_Nonstat:2020 showed that nonstationarity/misspecification alters the low frequency components of a time series making the latter appear as more persistent. Since classical HAC estimators are a weighted sum of an infinite number of low frequency periodogram ordinates, these estimates tend to be inflated. Similarly, LRV estimators using long bandwidths are weighted sum of a large number of sample autocovariances. Each sample autocovariance is biased upward so that the latter estimates are even more inflated than the classical HAC estimators. This explains why KVB's fixed-$b$ and EWC HAR tests have large power problems, even though classical HAC estimators are also affected.

casini/perron_Low_Frequency_Contam_Nonstat:2020 showed that the introduction of the smoothing over time in the DK-HAC estimators avoids such low frequency contamination. This follows because observations belonging to different regimes do not overlap when computing sample autocovariances. This guarantees excellent power properties also under nonstationarity/misspecification or under nonstationary alternative hypotheses (e.g., GR test discussed above). Simulation evidence suggests that tests based on the DK-HAC with simultaneous bandwidths are robust to low frequency contamination and overall performs better than tests based on the DK-HAC with sequential bandwidths especially with respect to size control.

Conclusions

We considered the derivation of data-dependent simultaneous bandwidths for double kernel heteroskedasticity and autocorrelation consistent (DK-HAC) estimators. We obtained the optimal bandwidths that jointly minimize the global asymptotic MSE criterion and discussed the trade-off between bias and variance with respect to smoothing over lagged autocovariances and over time. We highlighted how the derived MSE bounds are influenced by nonstationarity unlike the MSE bounds in andrews:91. We compared the DK-HAC estimators with simultaneous bandwidths to the DK-HAC estimators with bandwidths from the sequential MSE criterion. The new method leads to HAR tests that performs better in terms of size control, especially with stationary and close to stationary data. Finally, we considered long-run variance estimation where the relevant observations are a function of a nonparametric estimator and established the validity of the HAC and DK-HAC estimators in this setting.{ }\nocite{casini/perron_Oxford_Survey,casini/perron:change-point-spectra,casini/perron_CR_Single_Break,casini/perron_Lap_CR_Single_Inf,casini/perron_Low_Frequency_Contam_Nonstat:2020,casini/perron_PrewhitedHAC,casini/perron_SC_BP_Lap,casini_CR_Test_Inst_Forecast,casini_hac,casini_diss}Hence, we also extended the consistency results in andrews:91 and newey/west:87 to nonparametric estimation settings.

\addcontentsline{toc}{section}{References}

\pagenumbering{arabic}