Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,796 characters · 10 sections · 96 citation commands
The Fixed-$b$ Limiting Distribution and the ERP of HAR Tests Under Nonstationarity
\setcounter{page}{0}
\raggedbottom
{\bf{JEL Classification}}: C12, C13, C18, C22, C32, C51\\ {\bf{Keywords}}: Asymptotic expansion, Fixed-$b$, HAC standard errors, HAR inference, Long-run variance, Nonstationarity.
\onehalfspacing \thispagestyle{empty}
The construction of standard errors robust to autocorrelation and heteroskedasticity is important for empirical work because economic and financial time series exhibit temporal dependence. The early literature focused on heteroskedasticity and autocorrelation consistent (HAC) estimators of the asymptotic variance of test statistics (or simply the long-run variance (LRV) of the relevant series) {[}see, e.g., newey/west:87 (newey/west:87; newey/west:94), andrews:91, andrews/monahan:92, hansen:92ecma, dejong/davidson:00{]}. This approach aims at devising a good estimate of the LRV. Over the last twenty years, the literature has focused on methods based on fixed-$b$ asymptotics. These involve an inconsistent estimate of the LRV that keeps the bandwidth at a fixed fraction of the sample size. This approach was initiated by Kiefer/vogelsang/bunzel:00 and Kiefer/vogelsang:02 (Kiefer/vogelsang:02; vogelsang/kiefer:2002ET). They developed the analysis assuming stationarity and showed that valid heteroskedasticity and autocorrelation robust (HAR) inference is feasible even without a consistent estimator of the LRV. Inconsistency results in a pivotal nonstandard limiting distribution whose critical value can be obtained by simulations (e.g., a $t$-statistic on a coefficient in the linear regression model will not follow asymptotically a standard normal distribution but a distribution involving a ratio of Gaussian processes). Theoretical results based on asymptotic expansions suggested that fixed-$b$ HAR test statistics exhibit an error in rejection probability (ERP) that is smaller than that associated to test statistics based on HAC estimators {[}see jansson:04 and sun/phillips/jin:08{]}. This supported extensive finite-sample evidence in the literature documenting that the fixed-$b$ approach leads to HAR test statistics with more accurate null rejection rates when the data are stationary with strong temporal dependence than those associated to test statistics based on HAC estimators. Since then the literature has mostly concentrated on various refinements of fixed-$b$ HAR inference while maintaining the stationarity assumption, mostly to have tests having null rejection rates closer to the nominal level.
Although stationarity rarely holds in economic and financial time series, the literature has surprisingly ignored investigating the theoretical and empirical properties of existing fixed-$b$ HAR inference when stationarity does not hold. By nonstationary we mean non-constant moments. As in the literature, we consider processes whose sum of absolute autocovariances is finite. This rules out processes with unbounded second moments (e.g., unit root). Nonstationarity can occur for several reasons: changes in the moments of the relevant time series induced by changes in the model parameters that govern the data (e.g., the Great Moderation with the decline in variance for many macroeconomic variables, the effects of the financial crisis of 2007\textendash 2008 or of the COVID-19 pandemic); smooth changes in the distributions governing the data that arise from transitory dynamics from one regime to another. Unfortunately, the theoretical properties of fixed-$b$ HAR inference change substantially when stationarity does not hold. The contribution of the paper is to establish such theoretical results and discuss their relevance for inference in empirical work.
We show that the limiting distribution of HAR test statistics under fixed-$b$ asymptotics is not pivotal when the data are nonstationary. It takes the form of a complicated function of Gaussian processes and depends on the second moments of the relevant series. For example, in the case of the linear regression model, it depends on the second moments of the regressors and errors. Hence, fixed-$b$ inference methods based on stationarity are not theoretically valid in general. The nuisance parameters entering the fixed-$b$ limiting distribution can be consistently estimated under small-$b$ asymptotics but only with nonparametric rate of convergence. We develop asymptotic expansions under nonstationarity and we show that the ERP is an order of magnitude larger than that obtained under stationarity by jansson:04 and sun/phillips/jin:08 {[}cf. $O(T^{-\gamma})$ with $\gamma<1/2$ versus $O(T^{-1})$ where $T$ is the sample size{]}. Further, we show that the ERP of fixed-$b$ HAR tests is also larger than that of HAR tests based on HAC estimators. It follows from our results that if one uses fixed-$b$ methods based on the pivotal fixed-$b$ limiting distribution obtained under stationarity but the data are nonstationary, then the ERP does not even converge to zero as the sample size increases because that is not the correct limiting distribution. Hence, our results provide formal support to the claims in ibragimov/muller:10 and mueller:14 who mentioned that the stationarity assumption required by existing fixed-$b$ methods can be a limitation.
The pivotal property breaks down because under nonstationarity the LRV estimator that uses a fixed-bandwidth is not asymptotically proportional to the LRV. A non-pivotal limiting distribution results in a much more complex type of inference in practice. The increase in the ERP from the stationary case arises from fact that the nuisance parameters have to be estimated. It is the discrepancy between these estimates and their probability limits that is reflected in the leading term of the asymptotic expansion.
Our theoretical results reconcile with recent finite-sample evidence that showed that fixed-$b$ HAR tests can perform poorly when the data are nonstationary. These issues have been documented extensively by \textcolor{MyBlue}{Belotti et al.} belotti/casini/catania/grassi/perron_HAC_Sim_Bandws, casini_hac, casini/perron_PrewhitedHAC and casini/perron_Low_Frequency_Contam_Nonstat:2020 who considered $t$-tests in the linear regression models as well as HAR tests outside the linear regression model, and a variety of data-generating processes. They provided evidence that existing fixed-$b$ HAR tests can be severely undersized and can exhibit non-monotonic power. The more nonstationary the data are, the stronger the distortions. This is especially visible in HAR inference contexts characterized by a stationary null hypothesis and a nonstationary alternative hypothesis {[}e.g., tests for structural breaks, tests for regime-switching, tests for time-varying parameters and threshold effects, and tests for forecast evaluation{]}. In such cases, the power of fixed-$b$ HAR tests can be zero irrespective of how large the sample size is and how far the alternative is from the null value.
In our simulation analysis we focus on HAR inference in the linear regression model. The empirical results corroborate the predictions of our ERP results as the existing fixed-$b$ method yields substantial under-rejections. On the other hand, LRV estimators using small-$b$ bandwidths and standard asymptotic distributions avoid these under-rejection issues but likely exhibit over-rejections in the context of strongly persistent (stationary or nonstationary) processes. To address these problems, we propose a feasible inference method based on the non-pivotal nonstationary fixed-$b$ limiting distribution that involves replacing the nuisance parameters by nonparametric estimates and obtaining the critical values by simulating the limiting distribution. For some representative data-generating processes in a simple location model the new method leads to HAR test statistics with accurate null rejection rates irrespective of whether the data are stationary or not and of the strength of the serial dependence.
Recent works in HAR inference {[}see, e.g., sun:14 and lazarus/lewis/stock/watson:18{]} considered the use of small-$b$ asymptotics (i.e., small-bandwidths) in conjunction with fixed-$b$ critical values. These bandwidths are typically larger than the MSE-optimal bandwidths used for the HAC estimators. The idea is that as $b_{T}\rightarrow0$ the fixed-$b$ limiting distribution approximates the standard asymptotic distribution based on small-$b$ asymptotics. Although our results are obtained for fixed-bandwidths, they might suggest that using the critical values from the new fixed-$b$ limiting distribution would improve the finite-sample performance under nonstationarity. This is an interesting research question which, however, deserves its own research.
The remainder of the paper is organized as follows. Section (ref) introduces the statistical problem in the well-known setting of the linear regression model. In Section (ref) we study the limiting distribution of $t$- and $F$-type test statistics. Section (ref) develops the asymptotic expansions and presents the results on the ERP. Section (ref) presents Monte Carlo simulations. Section (ref) concludes the paper. The supplemental material {[}cf. casini_fixed_b_erp_supp{]} contains all mathematical proofs.
We consider the linear regression model
where $\beta_{0}\in\Theta\subset\mathbb{R}^{p}$, $y_{t}$ is an observation on the dependent variable, $x_{t}$ is a $p$-vector of regressors and $e_{t}$ is an unobserved disturbance that is autocorrelated and possibly conditionally heteroskedastic, and $\mathbb{E}(e_{t}|\,x_{t})=0$. The problem addressed is testing linear hypotheses about $\beta_{0}$. We consider the ordinary least squares (OLS) estimator $\widehat{\beta}=(\sum_{t=1}^{T}x_{t}x'_{t})^{-1}\sum_{t=1}^{T}x_{t}y_{t}$. Let $V_{t}=x_{t}e_{t}$. Define $S_{\left\lfloor Tr\right\rfloor }=\sum_{t=1}^{\left\lfloor Tr\right\rfloor }V_{t}$ where $\left\lfloor Tr\right\rfloor $ denotes the integer part of $Tr$. Using ordinary manipulations,
The variance of $T^{-1/2}S_{T}$ plays an important role for constructing tests about $\beta_{0}$. Its exact formula depends on the assumptions about $\left\{ V_{t}\right\} $. We begin with the following notational conventions.
A function $g\left(\cdot\right):\,\left[0,\,1\right]\mapsto\mathbb{R}$ is said to be piecewise (Lipschitz) continuous if it is (Lipschitz) continuous except on a set of discontinuity points that has zero Lebesgue measure. A matrix is said to be piecewise (Lipschitz) continuous if each of its element is piecewise (Lipschitz) continuous. Let $W_{p}\left(r\right)$ denote a $p$-vector of independent standard Wiener processes where $r\in\left[0,\,1\right]$. We use $\overset{\mathbb{P}}{\rightarrow},\,\Rightarrow$ and $\overset{d}{\rightarrow}$ to denote convergence in probability, weak convergence and convergence in distribution, respectively. The following assumptions are sufficient to establish the asymptotic distribution of the test statistics. Let $\Omega\left(u\right)$ denote some $p\times p$ positive semidefinite matrix.
Assumption (ref) states a functional law for nonstationary processes {[}see, e.g., aldous:1978 and merlevede/peligra/utev:2019{]}. If $\left\{ V_{t}\right\} $ is second-order stationary, then $\Sigma\left(u\right)=\Sigma$ for all $u$ and Assumption (ref) reduces to $T^{-1/2}S_{\left\lfloor Tr\right\rfloor }\Rightarrow\Sigma W_{p}\left(r\right)$. The fixed-$b$ literature has routinely used the assumption of second-order stationarity {[}see, e.g., Kiefer/vogelsang/bunzel:00, jansson:04, sun/phillips/jin:08 and lazarus/lewis/stock:17{]}. We relax this assumption substantially as we allow for general time-variation in the second moments of the regressors and errors which encompasses most of the nonstationary processes used in econometrics and statistics. For example, it allows for structural breaks, regime-switching, time-varying parameters and segmented local stationarity in the second moments of $\{V_{t}\}$. With regards to the temporal dependence, Assumption (ref) holds under a variety of regularity conditions. For example, standard mixing conditions and (time-varying) invertible ARMA processes are allowed.
Assumption (ref) allows for structural breaks as well as smooth variation in the second moments of the regressors.\footnote{Assumption (ref) also allows for polynomial trending regressors as long as they are written in the form $(t/T)^{l}$ ($l\geq0$), or more generally, written as a piecewise continuous function of the time trend, say $g(t/T)$.} The fixed-$b$ literature required $Q\left(u\right)=Q$ for all $u$ in which case Assumption (ref) reduces to $T^{-1}\sum_{t=1}^{\left\lfloor Tr\right\rfloor }x_{t}x'_{t}\overset{\mathbb{P}}{\rightarrow}rQ$. The latter is quite restrictive in practice. The uniform convergence, boundness and positive definiteness of $Q\left(\cdot\right)$ are satisfied for a fairly general class of processes. As in previous works, Assumption (ref)-(ref) rule out unit roots and long memory. Let
Under Assumption (ref)-(ref), the limit of $\mathrm{Var}(T^{-1/2}S_{T})$ is given by {[}cf. casini_hac{]}
where $c\left(u,\,k\right)=\mathbb{E}(V_{\left\lfloor Tu\right\rfloor }V_{\left\lfloor Tu\right\rfloor -k})+O(T^{-1})$. By the Cholesky decomposition $\Omega\left(u\right)=\Sigma\left(u\right)\Sigma\left(u\right)'$ and so $\Omega=\int_{0}^{1}\Omega\left(u\right)du$. Note that $\Omega=2\pi\int_{0}^{1}f\left(u,\,0\right)du$ where $f\left(u,\,0\right)$ is the local spectral density matrix of $\left\{ V_{t}\right\} $ at rescaled time $u$ and frequency $0$. For $u$ a continuity point, $f\left(u,\,\omega\right)$ is defined implicitly by the relation $\mathbb{E}(V_{\left\lfloor Tu\right\rfloor }V_{\left\lfloor Tu\right\rfloor -k})=\int_{-\pi}^{\pi}e^{i\omega k}f\left(u,\,\omega\right)d\omega$; see casini_hac for more details. If $\left\{ V_{t}\right\} $ is second-order stationary, then $\Omega=\Sigma\Sigma'=2\pi f\left(0\right)$ since $f(u,\,0)=f(0)$.
Under Assumption (ref)-(ref), it directly follows, using standard arguments, that
where $\Omega^{1/2}$ is the matrix square-root of $\Omega$ and $\overline{Q}\triangleq\int_{0}^{1}Q\left(u\right)du$. Under second-order stationarity $Q\left(u\right)=Q$, $\Sigma\left(u\right)=\Sigma$ and (ref) reduces to
The classical approach to testing hypotheses about $\beta_{0}$ is based on studentization. Provided that a consistent estimator of $\overline{Q}^{-1}\Omega\overline{Q}^{-1}$ can be constructed, it is possible to construct a test statistic whose asymptotic distribution is free of nuisance parameters. The term $\overline{Q}$ can be consistently estimated straightforwardly using $\widehat{Q}=T^{-1}\sum_{t=1}^{T}x_{t}x'_{t}$. Consistent estimators of $\Omega$ are known as HAC estimators {[}see, e.g., newey/west:87, andrews:91, dejong/davidson:00 and casini_hac{]}. HAC estimators take the following general form,
where
$\widehat{V}_{t}=x_{t}\widehat{e}_{t}$ and $\left\{ \widehat{e}_{t}\right\} $ are the OLS residuals, $K\left(\cdot\right)$ is a kernel and $b_{T}$ is a bandwidth sequence. Under $b_{T}\rightarrow0$ at an appropriate rate, we have $\widehat{\Omega}_{\mathrm{HAC}}\overset{\mathbb{P}}{\rightarrow}\Omega.$ An alternative to $\widehat{\Omega}_{\mathrm{HAC}}$ is the double-kernel HAC (DK-HAC) estimator, say $\widehat{\Omega}_{\mathrm{DK-HAC}}$, proposed by casini_hac to flexibly account for nonstationarity. $\widehat{\Omega}_{\mathrm{DK-HAC}}$ uses an additional kernel for smoothing over time; see casini_hac for details. Under appropriate conditions on the bandwidths, we have $\widehat{\Omega}_{\mathrm{DK-HAC}}\overset{\mathbb{P}}{\rightarrow}\Omega$. Hence, equipped with either $\widehat{\Omega}_{\mathrm{HAC}}$ or $\widehat{\Omega}_{\mathrm{DK-HAC}}$, HAR inference is standard because test statistics follow asymptotically standard distributions.
An alternative approach to HAR inference relies on inconsistent estimation of $\Omega.$ Kiefer/vogelsang:02 (Kiefer/vogelsang:02, vogelsang/kiefer:2002ET, kiefer/vogelsang:05) proposed to use the following “estimator”,\footnote{As a notational matter, it is useful to remark that the more recent fixed-$b$ literature does not refer to $\widehat{\Omega}_{\mathrm{fixed}-b}$ as an estimator. This recent literature rather uses the terminology “fixed-$b$” to refer to an asymptotic embedding. We may sometime refer to $\widehat{\Omega}_{\mathrm{fixed}-b}$ as an estimator. This should not create any confusion since our results are provided for the case of $b$ fixed which corresponds to the early fixed-$b$ literature.}
where $b\in(0,\,1]$ is fixed. Note that $\widehat{\Omega}_{\mathrm{fixed}-b}$ is equivalent to $\widehat{\Omega}_{\mathrm{HAC}}$ with $b_{T}=\left(bT\right)^{-1}$. $\widehat{\Omega}_{\mathrm{fixed}-b}$ is inconsistent for $\Omega$. Kiefer/vogelsang/bunzel:00 showed that an asymptotic distribution theory for HAR tests is possible even with an inconsistent estimate of $\Omega.$ One first has to derive the limiting distribution of $\widehat{\Omega}_{\mathrm{fixed}-b}$ under the null hypothesis. Then, one can use it to obtain the limiting distribution of the test statistic of interest which typically involves a ratio of Gaussian processes. Thus, from the inconsistency of $\widehat{\Omega}_{\mathrm{fixed}-b}$, HAR test statistics will not follow asymptotically standard distributions.
In this section we study the limiting distribution of the HAR tests for linear hypothesis in the linear regression model under fixed-$b$ asymptotics when the data are nonstationary. Consider testing the null hypothesis $H_{0}:\,R\beta_{0}=r$ against the alternative hypothesis $H_{1}:\,R\beta_{0}\neq r$ where $R$ is a $q\times p$ matrix with rank $q$ and $r$ is a $q\times1$ vector. Using $\widehat{\Omega}_{\mathrm{fixed}-b}$ an $F$-test can be constructed as follows:
For testing one restriction, $q=1$, one can use the following $t$-statistic:
Let $B_{p}\left(r\right)=W_{p}\left(r\right)-rW_{p}\left(1\right)$ denote the $p\times1$ vector of Brownian bridges. Consider the following class of kernels,
Examples of kernels in $\boldsymbol{K}$ include the Truncated, Bartlett, Parzen, Quadratic Spectral (QS) and Tukey-Hanning kernels. kiefer/vogelsang:05 showed under stationarity that
for $K\in\boldsymbol{K}$ with $K''\left(x\right)$ assumed to exist for $x\in\left[-1,\,1\right]$ and to be continuous.\footnote{Note that $K''\left(x\right)$ does not exist for some popular kernels. This is the case for the Bartlett kernel for which $K''\left(0\right)$ does not exist. However, kiefer/vogelsang:05 showed that for the Bartlett kernel it holds that
Recall that the Bartlett kernel is defined as $K_{\mathrm{BT}}\left(x\right)=1-\left|x\right|$ for $\left|x\right|\leq1$ and $K_{\mathrm{BT}}\left(x\right)=0$ otherwise. } A key feature of the result in (ref) is that $\widehat{\Omega}_{\mathrm{fixed}-b}$ is asymptotically proportional to $\Omega$ through $\Sigma\Sigma'$. The null asymptotic distributions of $F_{\mathrm{fixed}-b}$ and $t_{\mathrm{fixed}-b}$ under stationarity are given, respectively, by
and
Both null distributions are pivotal. Thus, valid testing is possible without consistent estimation of $\Omega.$ This result crucially hinges on stationarity. To see this, consider $t_{\mathrm{fixed}-b}$ for the single-regressor case ($p=1)$ and for the null hypothesis $H_{0}:\,\beta_{0}=0$. Its numerator and denominator are asymptotically equivalent to, respectively, $Q^{-1}\Sigma W_{1}\left(1\right)$ and
Since $W_{1}\left(1\right)$ and $B_{1}\left(r\right)$ are independent, $t_{\mathrm{fixed}-b}$ is a ratio of two independent random variables. The factor $Q^{-1}\Sigma$ cancels because it appears in both numerator and denominator. It follows that the asymptotic null distribution is pivotal. We show that this argument break downs when stationarity does not hold. Under nonstationarity the factor in the denominator corresponding to $Q^{-1}\Sigma$ will depend on the rescaled time $s$ and $r$, and enter the integrand. Thus, it will not cancel out.
We now present the results about the fixed-$b$ limiting distribution of the HAR tests. For $r\in\left[0,\,1\right],$ let
We begin with the following theorem which provides the limiting distribution of $\widehat{\Omega}_{\mathrm{fixed}-b}$.
Theorem (ref) show that, unlike in the stationary case, $\widehat{\Omega}_{\mathrm{fixed}-b}$ is not asymptotically proportional to $\Omega$. This anticipates that asymptotically pivotal tests for null hypotheses involving $\beta_{0}$ cannot be constructed. The limiting distribution depends on $K''\left(\cdot\right),$ $b$ and most importantly on $\Sigma\left(\cdot\right)$ and $Q\left(\cdot\right)$ so that it is not free of nuisance parameters.
We now present the limiting distribution of $F_{\mathrm{fixed}-b}$ and $t_{\mathrm{fixed}-b}$ under $H_{0}.$
Theorem (ref) shows that the asymptotic distribution of the $F$ and $t$ test statistics under fixed-$b$ asymptotics under nonstationarity are not pivotal. This contrasts with the stationary case where the asymptotic distributions depend only on the kernel and bandwidth. Consequently, fixed-$b$ inference based on stationarity is not theoretically valid under nonstationarity. The limiting distributions of $F_{\mathrm{fixed}-b}$ and $t_{\mathrm{fixed}-b}$ depend on nuisance parameters such as the time-varying autocovariance function of $\left\{ V_{t}\right\} $ through $\Sigma\left(\cdot\right)$ and the second moments of the regressors through $Q\left(\cdot\right)$. An inspection of the proof shows that it is practically impossible to make $F_{\mathrm{fixed}-b}$ and $t_{\mathrm{fixed}-b}$ pivotal by studentization based on any sequence of inconsistent covariance matrix estimates. This follows because the LRV is time-varying and, as noted above, this break downs the property that both numerator and denominator are asymptotically proportional to $\Omega$ so that the nuisance parameters cancel out. On the other hand, the property that under nonstationarity the fixed-$b$ asymptotic framework yields non-pivotal asymptotic distributions that depend on the underlying second-order properties of $\left\{ V_{t}\right\} $ may suggest that reliable HAR inference is more challenging.
Theorem (ref) suggests that valid inference under fixed-$b$ asymptotics is going to be more complex in terms of practical implementation relative to when the data are stationary. In the literature, complexity in the implementation has been recognized as a strong disadvantage for the success of a given method in empirical work {[}see, e.g., lazarus/lewis/stock:17{]}. The simplest way to use Theorem (ref) for conducting inference is to replace the nuisance parameters by consistent estimates. This means constructing estimates of $\Sigma\left(u\right)$, $Q\left(u\right)$ and $\overline{Q}.$ For $\overline{Q}$ the argument is straightforward. As under stationarity, one can use $\widehat{Q}=T^{-1}\sum_{t=1}^{T}x_{t}x'_{t}$ since $\widehat{Q}-\overline{Q}\overset{\mathbb{P}}{\rightarrow}0$ also under nonstationarity. More complex is the case for $\Sigma\left(u\right)$ and $Q\left(u\right)$. Nonparametric estimators for $\Sigma\left(u\right)$ and $Q\left(u\right)$ can be constructed. This requires introducing bandwidths and kernels as well as a criterion for their choice. Then, one plugs-in these estimates into the limit distribution and the critical value can be obtained by simulations. However, since $\Sigma\left(u\right)$, $Q\left(u\right)$ and $\overline{Q}$ depend on the data, the critical values need to be obtained on a case-by-case basis. We consider this approach in our simulation analysis in Section (ref) below.
It is interesting to briefly discuss the properties of fixed-$b$ HAR inference when $\left\{ V_{t}\right\} $ follows more general forms of nonstationary, i.e., $\left\{ V_{t}\right\} $ does not satisfy Assumption (ref)-(ref). Assumption (ref)-(ref) are satisfied if $\left\{ V_{t}\right\} $ is, e.g., segmented locally stationary, locally stationary and, of course, stationary. However, if $\left\{ V_{t}\right\} $ is a sequence of unconditionally heteroskedastic random variables such that $Q\left(s\right)$ and $\Sigma\left(s\right)$ do not satisfy the smoothness restrictions in Assumption (ref)-(ref) then Theorem (ref)-(ref) do not hold. For example, consider $V_{t}=\rho_{t}V_{t-1}+u_{t}$ where $u_{t}\sim i.i.d.\,\mathscr{N}\left(0,\,1\right)$ and $\rho_{t}\in\left(-1,\,1\right)$ for all $t$. Segmented local stationarity corresponds to $\rho_{t}$ being piecewise continuous, local stationarity corresponds to $\rho_{t}$ being continuous and stationarity corresponds to $\rho_{t}$ being constant. If $\rho_{t}$ does not satisfy any of these restrictions, Assumption (ref)-(ref) do not hold. For unconditionally heteroskedastic random variables the asymptotic distributions of $F_{\mathrm{fixed}-b}$ and $t_{\mathrm{fixed}-b}$ remain unknown since they cannot be characterized. Thus, for general nonstationary random variables fixed-$b$ inference based on the asymptotic distribution is infeasible. This highlights one major difference from HAR inference based on consistent estimation of $\Omega.$ HAC and DK-HAC estimators are consistent for $\Omega$ also for general nonstationary random variables so that HAR test statistics follow asymptotically the usual standard distributions. For example, a $t$-statistic studentized by a HAC or DK-HAC estimator will follow asymptotically a standard normal.
Before concluding this section, we note that a few papers explored fixed-$b$ asymptotics in settings that involve other forms of nonstationarity that do not fall within the standard HAR inference problem with $\sqrt{T}$-asymptotically normal OLS estimator. bunzel/vogelsang:2005 allowed for deterministic trends and integrated of order one (I(1)) errors. vogelsang/wagner:2013 considered fixed-$b$ asymptotics for unit root tests while vogelsang/wagner:2014 focused on cointegrating regression. xu:2012 and demetrescu/hanck/kruse-becher:2023 considered multivariate trend tests and hypothesis tests in a GMM framework, respectively, allowing for time-varying volatility and no serial correlation.
We develop high-order asymptotic expansions and obtain the ERP of fixed-$b$ HAR tests. The results in velasco/robinson:01, jansson:04 and sun/phillips/jin:08 suggest that under stationarity the ERP of $F_{\mathrm{fixed}-b}$ and of $t_{\mathrm{fixed}-b}$ are smaller than those of the conventional HAC-based HAR tests. We show that the opposite is true when stationarity does not hold.
Consider the location model $y_{t}=\beta_{0}+e_{t}$ $(t=1,\ldots,\,T)$. We have $V_{t}=e_{t}$. Under the assumption that $V_{t}$ is stationary and Gaussian, velasco/robinson:01 developed second-order Edgeworth expansions and showed that
for any $z\in\mathbb{R}$ where
$b_{T}\rightarrow0$, $\varPhi\left(\cdot\right)$ is the distribution function of the standard normal and $d\left(\cdot\right)$ is an odd function. The ERP is the leading term of the right-hand side of (ref). Since $b_{T}=O(T^{-\eta})$ with $0<\eta<1$, the ERP of $t_{\mathrm{HAC}}$ is $O(T^{-\gamma})$ with $\gamma<1/2$. It follows that the leading term of $\mathbb{P}(F_{\mathrm{HAC}}\leq c)$ where $F_{\mathrm{HAC}}=T(\widehat{\beta}-\beta_{0})^{2}/\widehat{\Omega}_{\mathrm{HAC}}$ is of the form $2d(\sqrt{c})T^{-\gamma}=O(T^{-\gamma})$ for any $c>0.$
jansson:04 and sun/phillips/jin:08 showed that\footnote{Actually jansson:04 showed that the bound was $O(T^{-1}\log T)$. Using a different proof strategy, sun/phillips/jin:08 showed that the $\log T$ term can be dropped. }
Thus, $F_{\mathrm{fixed}-b}$ has a smaller ERP than $F_{\mathrm{HAC}}$ {[}cf. $O(T^{-1})$ versus $O(T^{-\gamma})${]}. This implies that the rate of convergence of $F_{\mathrm{fixed}-b}$ to its (nonstandard) limiting null distribution is faster than the rate of convergence of $F_{\mathrm{HAC}}$ to a $\chi_{1}^{2}$. These results reconciled with finite-sample evidence in the literature showing that the null rejection rates of $F_{\mathrm{fixed}-b}$ and $t_{\mathrm{fixed}-b}$ are more accurate than those of $F_{\mathrm{HAC}}$ and $t_{\mathrm{HAC}}$, respectively, when that data are stationary.
We now address the question of whether these results extend to nonstationarity. It turns out that the answer is negative. This provides an analytical explanation for the Monte Carlo experiments that have appeared recently in casini_hac, casini/perron_Low_Frequency_Contam_Nonstat:2020 and casini/perron_PrewhitedHAC who found serious distortions in the rejection rates of fixed-$b$ HAR tests under the null and alternative hypotheses when the data are nonstationary. These distortions being often much larger than those corresponding to the conventional HAC-based HAR tests.
Theorem (ref) showed that the fixed-$b$ HAR tests are not pivotal. Thus, a natural way to conduct inference based on the fixed-$b$ asymptotic distribution is to construct consistent estimates of its nuisance parameters. We introduce a general nonparametric estimator of $\Sigma\left(u\right)$. Let
and
where $K_{h_{1}}\left(\cdot\right)\in\boldsymbol{K}$, $h_{1}$ is a bandwidth sequence satisfying $h_{1}\rightarrow0$,
with $K_{h_{2}}\left(\cdot\right)\in\boldsymbol{K}_{2}$ and $h_{2}$ is a bandwidth sequence satisfying $h_{2}\rightarrow0$. Since $x_{t}=1$ for all $t$, we have $Q\left(u\right)=1$ for all $u\in\left[0,\,1\right]$ in (ref)-(ref). Thus, we set $\widehat{Q}\left(u\right)=1$ for all $u\in\left[0,\,1\right]$. For arbitrary $x_{t},$ one can take
As in the literature, we focus on the simple location model with Gaussian errors. The Gaussianity assumption can be relaxed by considering distributions with, for example, Gram-Charlier representations at the expenses of more complex derivations {[}see, e.g., phillips:1980{]}. The following assumption on $V_{t}$ facilitates the development of the higher order expansions and is weaker than the one used by sun/phillips/jin:08 since they also imposed second-order stationarity.
In order to develop the asymptotic expansions we use the following conditions on the kernel which were also used by andrews:91 and sun/phillips/jin:08.
All of the commonly used kernels with the exception of the truncated kernel satisfy Assumption (ref)-(i, ii). sun/phillips/jin:08 required piecewise smoothness on $K\left(\cdot\right)$ instead of the Lipschitz condition. Part (iii) ensures that the associated LRV estimator is positive semidefinite. The commonly used kernels that satisfy part (i, iii) are the Bartlett, Parzen and quadratic spectral (QS) kernels. For the Bartlett kernel, $q_{0}=1$, while for the Parzen and QS kernels, $q_{0}=2.$
As in sun/phillips/jin:08, we present the asymptotic expansion for the test statistic studentized by $\widehat{\Omega}_{\mathrm{fixed}-b}$ defined in (ref). Let $K_{b}=K\left(\cdot/b\right)$. Lemma (ref) in the supplement extends Theorem (ref) to $\widehat{\Omega}_{\mathrm{fixed}-b}$ using the kernels that satisfy Assumption (ref). Under Assumption (ref) and (ref)-(ref), Lemma (ref) shows that $\widehat{\Omega}_{\mathrm{fixed}-b}\Rightarrow\mathscr{G}_{b}$ where
and
Let $C_{\Omega}=\sup_{s\in\left[0,\,1\right]}\Omega\left(s\right)$, $C_{2,\Omega}=\max\{C_{\Omega},\,1\}$,
and
Theorem (ref) shows that the ERP associated to $t_{\mathrm{fixed}-b}$ is $O((Th_{1}h_{2})^{-1/2})$. This is an order of magnitude larger than the ERP associated to $t_{\mathrm{fixed}-b}$ under stationarity, $O(T^{-1})$, where the latter was established by jansson:04 and sun/phillips/jin:08. The increase in the ERP of $t_{\mathrm{fixed}-b}$ is the price one has to pay for not having a pivotal distribution under nonstationarity. This is intuitive. Without a pivotal distribution, one has to obtain estimates of the nuisance parameters. However, the nuisance parameters can be consistently estimated only under small-$b$ asymptotics. The latter estimates enjoy a nonparametric rate of convergence which then results in a larger ERP since it is the discrepancy between these estimates and their probability limits that is reflected in the leading term of the asymptotic expansion.
The conventional fixed-$b$ methods use a fixed bandwidth and the critical value from the pivotal fixed-$b$ limiting distribution obtained under the assumption of stationarity. Our results suggest that the ERP associated to such fixed-$b$ HAR tests is $O\left(1\right)$. This follows because that critical value is not theoretically valid, i.e., it is from the pivotal fixed-$b$ limiting distribution which, however, is different from the non-pivotal fixed-$b$ limiting distribution under nonstationarity. Thus, as $T\rightarrow\infty$ the ERP does not converge to zero, implying large distortions in the null rejection rates even for unbounded sample sizes.
Theorem (ref) implies that the theoretical properties of fixed-$b$ inference changes substantially depending on whether the data are stationarity or not. In particular, it suggests that the approximations based on fixed-$b$ asymptotics obtained under stationarity in the literature are not valid and do not provide a good approximation when stationarity does not hold. This contrasts to HAR inference tests based on consistent long-run variance estimators which are valid also under nonstationarity and have the same asymptotic distribution regardless of whether the data are stationary or not.\footnote{It is useful to remind that even though they are generally valid their finite-sample performance can be poor if there is strong dependence under either stationarity or nonstationarity as documented in the literature.} Theorem (ref) also provides formal support to the arguments in ibragimov/muller:10 and mueller:14 who mentioned that the stationarity assumption used by fixed-$b$ methods is a disadvantage relative to conventional methods including the $t$-statistic approach.
Additional comments: 1. It is useful to compare Theorem (ref) with the results for the ERP associated to $t_{\mathrm{HAC}}$. Under stationarity velasco/robinson:01 showed that the ERP associated to $t_{\mathrm{HAC}}$ is $O((Tb_{T})^{-1/2})$ where $b_{T}\rightarrow0.$ casini/perron_Low_Frequency_Contam_Nonstat:2020 showed that under nonstationarity the ERP associated to $t_{\mathrm{HAC}}$ has the same order as under stationarity, i.e., $O((Tb_{T})^{-1/2})$. Thus, it is sufficient to compare $O((Th_{1}h_{2})^{-1/2})$ and $O((Tb_{T})^{-1/2})$. Since $h_{1}$ and $b_{T}$ are the bandwidths used for smoothing over lagged autocovariances, they may have a similar order. It follows that the ERP associated to $t_{\mathrm{fixed}-b}$ may be larger than that associated to $t_{\mathrm{HAC}}$. In addition, the ERP associated to $t_{\mathrm{fixed}-b}$ based on conventional fixed-$b$ methods that rely on stationarity is much larger than the ERP associated to $t_{\mathrm{HAC}}$ since the former is $O\left(1\right)$.
2. A more recent development in the literature {[}see, e.g., sun:14 and lazarus/lewis/stock/watson:18{]} considered the use of small-$b$ asymptotics (i.e., small-bandwidths) and fixed-$b$ critical values. These bandwidths are typically larger than the MSE-optimal bandwidths used for $t_{\mathrm{HAC}}$ {[}see eq. (3) in lazarus/lewis/stock/watson:18, and the equation for $b^{*}$ on p. 666 of sun:14 and the related discussion there{]}. As $b_{T}\rightarrow0$ the fixed-$b$ limiting distribution approximates the standard asymptotic distribution based on small-$b$ asymptotics. Thus, the fixed-$b$ critical values converge to the standard normal critical values for the case of a $t$-test. In the limit the ERP of these HAR tests should be the same as that of $t_{\mathrm{HAC}}$. However, recent simulation results in the literature show that these HAR tests have different finite-sample rejection probabilities from those of $t_{\mathrm{HAC}}$. Hence, although the result in Theorem (ref) only speaks for fixed-bandwidths, it might suggest that using the fixed-$b$ critical values from the new fixed-$b$ limiting distribution may improve the finite-sample performance of these recent fixed-$b$ methods under nonstationarity.
3. Overall, the theoretical results contrast with what the early fixed-$b$ literature showed under stationarity {[}see jansson:04{]}, namely that the original fixed-$b$ HAR inference is theoretically superior to HAR inference based on consistent long-run variance estimators.
4. Our theoretical results complement the recent finite-sample evidence in \textcolor{MyBlue}{Belotti et al.} belotti/casini/catania/grassi/perron_HAC_Sim_Bandws, casini_hac, casini/perron_PrewhitedHAC and casini/perron_Low_Frequency_Contam_Nonstat:2020. Their simulation results showed that existing fixed-$b$ HAR inference tests perform poorly in terms of the accuracy of the null rejection rates and of power when stationarity does not hold. They considered $t$-tests in the linear regression models and HAR tests outside the linear regression model, and a variety of data-generating processes. They provided evidence that fixed-$b$ HAR tests can be severely undersized and can exhibit non-monotonic power. Some of these issues are generated by the low frequency contamination induced by nonstationarity which biases upward each sample autocovariance $\widehat{\Gamma}\left(k\right)$. Since $\widehat{\Omega}_{\mathrm{fixed}-b}$ uses many lagged autocovariances as $b$ is fixed, it is inflated which then results in size distortions and lower power.
5. The fixed-$b$ limiting distribution under nonstationarity is complex to use in practice as it depends in a complicated way on nuisance parameters. The procedure in Theorem (ref) replaces the nuisance parameter $\Sigma\left(\cdot\right)$ by a consistent estimate. This procedure represents a natural starting point to study the properties of fixed-$b$ inference under nonstationarity and so the corresponding ERP results may provide general guidance. There are certainly other procedures that could be used. It is beyond the scope of the paper to investigate how best to use the non-pivotal fixed-$b$ limiting distribution. If one wants to consider other procedures such as finding a conservative upper bound for the critical value that holds under all possible values of the nuisance parameter, bootstrap-based autocorrelation robust tests, modification of the test statistic\footnote{hwang/sun:2017 proposed a modification to the trinity of test statistics in the two-step GMM setting and showed that the modified test statistics are asymptotically $F$ distributed under fixed-$b$ asymptotics.}, etc., then one has to face the challenge that these methods are not as simple as HAC-based inference which can be a disadvantage as recently argued by lazarus/lewis/stock:17. Future work should investigate on possible alternative fixed-$b$ procedures that exploit the results in Section (ref).
6. The requirement $Th_{1}h_{2}\rightarrow\infty$ is a standard condition for consistency of nonparametric estimators such as $\widehat{\Omega}\left(u\right)$. The requirement $b<1/(16C_{2,\Omega}\int_{-\infty}^{\infty}|K\left(x\right)|dx)$ is similar to the one used by sun/phillips/jin:08. It can be relaxed at the expenses of more complex derivations.
7. As remarked at the end of Section (ref), for unconditionally heteroskedastic random variables, standard fixed-$b$ HAR inference is infeasible. Thus, the associated ERP does not convergence to zero. In contrast, HAR inference based on consistent long-run variance estimator is valid and the associated ERP is again $O((Tb_{T})^{-1/2})$ with $b_{T}\rightarrow0.$
In this section we conduct a Monte Carlo analysis to evaluate the effectiveness of the theoretical results of Section (ref)-(ref). We consider the empirical null rejection rates and local power of the $t$-statistic in a simple location model:
We consider the following data-generating processes for $e_{t}$. In model M1 we specify $e_{t}$ as an AR(1) with a break in the autoregressive coefficient,
where $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$ and $e_{0}\sim\mathscr{N}\left(0,\,1\right)$. Further, the initial condition of $e_{t}$ in the second regime is not the realized $e_{0.2T}$ but we set $e_{0.2T}\sim\mathscr{N}\left(0,\,1\right)$ so that the two regimes are independent.\footnote{The results are unchanged when we use the realized $e_{0.2T}$ as the initial condition for the second regime.} Model M2 involves a locally stationary AR(1):
where $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$ and $e_{0}\sim\mathscr{N}\left(0,\,1\right)$. Note that $\rho_{t}$ varies between 0.055 to 0.850. In model M3 we consider a stationary AR(1) $e_{t}=0.9e_{t-1}+u_{t}$ where $u_{t}\sim\mathrm{i.i.d.\,}\mathscr{N}\left(0,\,1\right)$ and $e_{0}\sim\mathscr{N}\left(0,\,1\right)$. We consider the following test statistic:
where $\widehat{\beta}=\overline{y}=T^{-1}\sum_{t=1}^{T}y_{t}$ and $\widehat{\Omega}_{b}=\widehat{\Omega}_{\mathrm{fixed-}b}$ as defined in (ref) with the Bartlett kernel and $\widehat{V}_{t}=y_{t}-\overline{y}$. We set $\beta_{0}=0$ and consider the null hypothesis that $\beta_{0}\leq0$ against the alternative that $\beta_{0}>0$. The results for a two-sided version of the test are qualitatively similar. We report the results for the sample sizes $T=250,\,500$ and 5,000 replications were used in all cases. To illustrate how the reference distribution works as the bandwidth $b$ varies in finite-sample, we compute the rejection probabilities for $t_{b}$ implemented using $b=0.02,\,0.04,\,0.06,\ldots,\,0.96,\,0.98,\,1$. We set the asymptotic significance level to 0.05 and consider the following critical values. The usual standard normal critical value of 1.645 for all values of $b$; the critical value from the stationary fixed-$b$ distribution in kiefer/vogelsang:05 for the corresponding $b$; the critical values from the infeasible and feasible nonstationary fixed-$b$ distribution for the corresponding $b$. The infeasible version is obtained by simulating the distribution in Theorem (ref) with the true value of $\Sigma\left(u\right)$ and $Q\left(u\right)$. Note that in our setting $Q\left(u\right)=1$ for all $u.$ The feasible version is obtained by simulating the distribution in Theorem (ref) where $\Sigma\left(u\right)$ is replaced by $\widehat{\Sigma}\left(u\right)$ defined in (ref) with $h_{1}=(Th_{2})^{-4/5}$ and $h_{2}=T{}^{-1/3}$, $K_{h_{1}}$ is Bartlett kernel and $K_{h_{2}}$ is the rectangular kernel. Other choices for $h_{1}$ and $h_{2}$ that satisfy the conditions of Theorem (ref) are possible. However, we verified in unreported simulations that the empirical rejection probabilities are only very marginally sensitive to the choice of $h_{1}$ and $h_{2}$ (i.e., within the $\pm1\%$ range about the ones reported below). In particular, any combination of $\left(h_{1},\,h_{2}\right)$ leads to empirical rejection rates that are more accurate than those from the stationary fixed-$b$ distribution of kiefer/vogelsang:05. Similarly, different choices for the kernels $K_{h_{1}}$ and $K_{h_{2}}$ yield quantitative similar results. We do not report them for brevity.
The results about the null rejection rates are depicted in Figures (ref)-(ref). To facilitate the reading of the figures, we also report the null rejection rates for each method for $b=0.25,\,0.5,\,0.75,\,1$ in Table (ref). In each figure, the line with the label, $N\left(0,\,1\right)$, plots the rejection probabilities when the critical value 1.645 is used. The figures also depict plots of the rejection probabilities using the stationary fixed-$b$ asymptotic critical values and the infeasible and feasible nonstationary fixed-$b$ asymptotic critical values. The results are striking. In the nonstationary models M1 and M2 the stationary fixed-$b$ method yields rejection rates that are substantially below the significance level for all values of $b$ except for $b$ very small.\footnote{This is consistent with the empirical results in casini/perron_Low_Frequency_Contam_Nonstat:2020.} The latter feature is expected since for small $b$ the fixed-$b$ asymptotic theory reduces to the small-$b$ asymptotics, or the fixed-$b$ distribution reduces to the standard normal distribution. Given that a small $b$ means that a small number of sample autocovariances are used in $\widehat{\Omega}_{b}$, we expect the test statistic to over-reject.\footnote{This feature also appeared in the simulations of kiefer/vogelsang:05.} In contrast, both the infeasible and feasible nonstationary fixed-$b$ distributions lead to empirical rejection rates that are very close to 0.05 for all values of $b$ except for $b$ very small. Interestingly, for $T=500$ the stationary fixed-$b$ critical values yield rejection rates that sometimes are worse than for $T=250$, i.e., the under-rejection becomes more pronounced with a larger the sample size. This does not happen when the nonstationary fixed-$b$ critical values are used, in fact they yield more accurate rejection rates as the sample size increases. Hence, these results corroborate the relevance of the nonstationary fixed-$b$ distribution theory relative to the stationary fixed-$b$ distribution theory when the data are nonstationary.
The pattern of the rejection rates when the standard normal critical value is used is similar to that found by kiefer/vogelsang:05. When $b$ is small, there are substantial over-rejections. The rejections fall as $b$ increases but then rise again as $b$ increases further. For all values of $b$ the rejection rates are beyond 0.05. This confirms the size distortions documented in the literature and that the accuracy of the small-$b$ asymptotics depends on $b$.
It is useful to analyze the performance of the nonstationary fixed-$b$ distribution when the data are indeed stationary with strong dependence. The results are reported in Figure (ref)-(ref). First note that Theorem (ref) implies that when the data are stationary the nonstationary fixed-$b$ distribution reduces to the stationary fixed-$b$ distribution obtained by kiefer/vogelsang:05. In fact, the figures show that the rejection rates corresponding to the infeasible nonstationary fixed-$b$ critical values coincide with those corresponding to the stationary fixed-$b$ critical values. As it is well-known, the standard normal critical value leads to large over-rejections for all $b.$ In contrast, the empirical rejection rates corresponding to either fixed-$b$ critical values are much more accurate. The feasible nonstationary fixed-$b$ critical values yield rejection rates that are essentially the same as the ones from the stationary fixed-$b$ distribution. For $b\in\left[0.2,\,1\right]$ the rejection rates of either fixed-$b$ method are stable and therefore equally accurate. For small values of $b$ also the fixed-$b$ critical values yield over-rejections. This is obvious since a small $b$ does not correspond to a fixed-$b$ asymptotic theory with $b$ required to be fixed. Rather, the fixed-$b$ distribution approximates the standard normal distribution as $b\rightarrow0$ and so no gain is expected from using the fixed-$b$ critical values for values of $b$ that are too small.
Let us comment on the difference between the infeasible and feasible nonstationary fixed-$b$ inference. The proof of Theorem (ref) suggests that the infeasible fixed-$b$ distribution enjoys a smaller ERP than its feasible counterpart. The feasible fixed-$b$ method depends on the nonparametric estimators of $\Sigma\left(u\right)$ and $Q\left(u\right)$ which in turn depend on the choice of $h_{1}$ and $h_{2}$. However, as noted above, the rejection rates are not much sensitive to the choice of $h_{1}$ and $h_{2}$. For the reported results, the performance of the feasible inference is not very different from that of the infeasible inference. The feasible inference is slightly more accurate in model M1-M2 and slightly less in model M3. Our unreported simulations involving other choices of $h_{1}$ and $h_{2}$ showed that the feasible inference can often be slightly worse than the infeasible inference as the theory would suggest.
We now move to discuss the local power results. We consider the following null and alternative hypotheses
where $c=\delta\sqrt{\Omega}>0$ is a constant. The local power is computed for $\delta=0,\,0.2,\,0.4,\ldots,\,4.8,\,5$ using 5% asymptotic null critical values. What we report is the size-unadjusted power. We only report the results for model M1 since the results for model M2-M3 are qualitative similar. Figure (ref) and (ref) plot the power across methods for a given value of $b$. Figure (ref) plots the power across values of $b$ for a given method. The power corresponding to the standard normal critical value is much higher than that associated to the fixed-$b$ methods. This is intuitive given that the standard normal critical value leads to oversized tests. Hence, it is fair to focus on the power comparison between the stationary fixed-$b$ method and the infeasible and feasible nonstationary fixed-$b$ methods since they are not oversized. The figures show that the power gain from using the nonstationary fixed-$b$ critical values is roughly 10% for $T=250$ and roughly 15% for $T=500.$ All tests have monotonic power and as $\delta$ increases the power differences become smaller and smaller. These features continue to hold for model M2. Thus, the under-rejections of the stationary fixed-$b$ inference lead to power losses relative to the nonstationary fixed-$b$ inference. In model M3 the power functions associated to the stationary and nonstationary fixed-$b$ methods are essentially the same. Thus, there is no loss in using the nonstationary fixed-$b$ inference when the data are stationary.
To sum up, the theoretical results of Section (ref)-(ref) provide useful predictions about the finite-sample accuracy of the stationary and nonstationary fixed-$b$ distributions. The nonstationary fixed-$b$ distribution yields empirical null rejection rates that are accurate for both stationary and nonstationary data. The stationary fixed-$b$ distribution yields null rejection rates that are substantially below the significance level when the data are nonstationary. It follows that its corresponding power is lower than that associated to the nonstationary fixed-$b$ distribution.
Finally, recent works by casini_hac and casini/perron_Low_Frequency_Contam_Nonstat:2020 showed that in the context of nonstationary alternative hypotheses the stationary fixed-$b$ method as well as the traditional small-$b$ HAC methods exhibit non-monotonic power. This involves testing problems outside the linear regression (e.g., tests for structural breaks, time-varying parameters and regime switching, and tests for forecast evaluation). By construction the ERP results in Section (ref) are only relevant for the size properties of the tests and thus are not suitable for nonstationary alternative hypotheses. We verified via simulations that also the nonstationary fixed-$b$ method can suffer from non-monotonic power in those contexts, though by a smaller extent. The only methods that have monotonic power are those based on the DK-HAC estimators of casini_hac. Hence, it would be interesting to combine the DK-HAC estimation with the nonstationary fixed-$b$ asymptotics in future research.
This paper has shown that the theoretical properties of fixed-$b$ HAR inference change depending on whether the data are stationary or not. Under nonstationarity we established that fixed-$b$ HAR test statistics have a limiting distribution that is not pivotal and that their ERP is an order of magnitude larger than that under stationarity and can be larger than that of HAR tests based on traditional HAC estimators. These theoretical results reconcile with recent finite-sample evidence showing that fixed-$b$ HAR test statistics can perform poorly when the data are nonstationary, both in terms of distortions in the null rejection rates and of non-monotonic power. Overall, the results highlight a new facet of the trade-off between size and power in HAR inference, i.e., the methods that achieve better null rejection rates under stationarity are the ones that suffer more from under-rejection under nonstationarity (i.e., time-varying autocovariance structure), and vice-versa. A new inference method based on the nonstationary fixed-$b$ distribution is proposed and it is shown to provide accurate null rejection rates for hypothesis testing in the linear regression model irrespective of whether the data are stationary or not and of the strength of the dependence as verified for some representative data-generating processes in a simple location model.
The supplement for online publication {[}cf. casini_fixed_b_erp_supp{]} contains the proofs of the results.
\addcontentsline{toc}{section}{References}
\pagenumbering{arabic}