EconBase
← Back to paper

On Robust Inference in Time Series Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

94,631 characters · 23 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On Robust Inference in Time Series Regression

center[center omitted — 109 chars of source]

\thispagestyle{empty}

{

spacing{1} Abstract: Least squares regression with heteroskedasticity consistent standard errors (“OLS-HC regression") has proved very useful in cross section environments. However, several major difficulties, which are generally overlooked, must be confronted when transferring the HC technology to time series environments via heteroskedasticity and autocorrelation consistent standard errors (“OLS-HAC regression"). First, in plausible time-series environments, OLS parameter estimates can be inconsistent, so that OLS-HAC inference fails even asymptotically. Second, most economic time series have autocorrelation, which renders OLS parameter estimates inefficient. Third, autocorrelation similarly renders conditional predictions based on OLS parameter estimates inefficient. Finally, the structure of popular HAC covariance matrix estimators is ill-suited for capturing the autoregressive autocorrelation typically present in economic time series, which produces large size distortions and reduced power in HAC-based hypothesis testing, in all but the largest samples. We show that all four problems are largely avoided by the use of a simple and easily-implemented dynamic regression procedure, which we call DURBIN. We demonstrate the advantages of DURBIN with detailed simulations covering a range of practical issues. \scriptsize Acknowledgments: For detailed comments we are greatly indebted to the editor, co-editor, and two referees. In addition we gratefully acknowledge useful discussions and/or comments from Rob Engle, Domenico Giannone, Jim Hamilton, Daniel Lewis, Nour Meddahi, Ulrich M{\"u}ller, Serena Ng, Lasse Pedersen, Pierre Perron, Peter Phillips, Mikkel Plagborg-Moller, Peter Schmidt, George Tauchen, Tim Volgelsang, Mark Watson, Ken West, and Jeff Wooldridge. We are also grateful to seminar participants at Michigan State University and the University of Pennsylvania, and conference participants at the 2022 NBER Summer Institute, the 2023 joint meetings of the Royal Economic Society and Scottish Economic Society, and the 2023 Copenhagen Conference on Advances in Financial Econometrics. All remaining errors or misunderstandings are ours alone. Key Words: Serial correlation, heteroskedasticity and autocorrelation consistent (HAC) regression, Durbin regression, dynamic regression Contact: {[email removed]}, {[email removed]}, {[email removed]}, {[email removed]}, \\[email removed] JEL Codes: C13, C22, C31

}

{ \setcounter{page}{1} \thispagestyle{empty} }

Introduction

For nearly a century, regression with heteroskedastic and/or autocorrelated disturbances has featured prominently in empirical economics research. For many decades, attention centered on modeling the heteroskedasticity or autocorrelation in the context of feasible generalized least squares (FGLS) estimation.

The dominant estimation approach in recent decades, however, is ordinary least squares (OLS) with standard errors adjusted to achieve valid asymptotic inference without taking a stand on the form of heteroskedasticity or autocorrelation. The idea traces to the classic contribution of White1980, who considered OLS regression with heteroskedasticity consistent (HC) standard errors (“OLS-HC regression") in cross-sectional environments, where sample sizes are typically very large, little or no information is available regarding the form of any possible heteroskedasticity, and serial correlation is irrelevant. In such environments HC standard errors are appropriate and justly emphasized (e.g. angrist2008mostly).

In an elegant extension, newey1987simple generalize White's estimator from cross sections to time series, with possible heteroskedasticity and serial correlation, by replacing White's covariance matrix estimator with an appropriate time-series analog based on an estimator of a spectral density at frequency zero.\footnote{The Newey-West estimator collapses to the White1980 estimator if serial correlation is absent, but appropriately incorporates serial correlation in the calculation of robust standard errors when serial correlation is present.} Such OLS regression with heteroskedasticity and autocorrelation consistent (HAC) standard errors (“OLS-HAC regression") has become extremely popular in time series environments.

In this paper we argue, however, that, in contrast to cross section OLS-HC regression, time series OLS-HAC regression as typically implemented is likely to be problematic, for a variety of reasons:

enumerate• In plausible time-series environments, OLS parameter estimates can be inconsistent, so that OLS-HAC inference fails even asymptotically. And moreover, even when OLS parameter estimates are consistent: • OLS parameter estimates can be highly inefficient in the presence of serial correlation, compared to estimators that account for the serial correlation. • OLS-HAC regression discards valuable predictive information in serially-correlated disturbances and hence produces sub-optimal (inefficient) forecasts, whereas accurate out-of-sample prediction is often a central concern in time series econometrics. • Newey-West-style HAC covariance matrix estimators are ill-suited for capturing the autoregressive autocorrelation typically present in economic time series, which can produce large size distortions, and large power reductions even when the size is not distorted.

Claim (ref) is not widely appreciated, with the exception of glsp, whose results and approach complement ours.\footnote{By now parts of our paper and theirs are entangled. A preliminary version of our paper was presented at the 2016 NBER-NSF Time Series Conference at Columbia University. Our first-draft working paper was released in March 2022, with no knowledge of their work-in-progress. Their first-draft working paper was released in September 2022, with knowledge of ours. Our second-draft working paper was released in June 2022, with knowledge of theirs. This third draft of our paper was released on \today.} Claim (ref) is well known, but its importance in finite samples is ignored when using OLS-HAC regression. Claim (ref) is obvious, but again ignored when using OLS-HAC regression. Claim (ref) is appreciated and has motivated several important refinements of the Newey-West HAC covariance matrix estimator (e.g.,andrews1991heteroskedasticity, kiefer2002heteroskedasticity, lazarus2018har), as well as use of spectral density estimators that differ from the Newey-West lag-window estimator (e.g., muller2014hac). However, those refinements have been only partially successful.

Against the background of the above claims (ref)-(ref), which we will substantiate in detail, we proceed to make a constructive contribution. We propose an alternative to OLS-HAC regression based on so-called “Durbin regressions" (durbin1970testing). Working in a very general environment that includes most dynamic specifications of interest as special cases, we show that the new procedure simultaneously addresses claims (ref)-(ref) above. Indeed, the Durbin regression procedure performs well in all situations, dominating the traditional OLS-HAC and FGLS procedures.

Our paper proceeds as follows. In section (ref) we introduce the basic data-generating process and estimators, including not only traditional OLS-HAC regression and our Durbin regression, but also traditional FGLS and a recently-proposed modified FGLS procedure. In section (ref) we present a generalized modeling framework. In section (ref) we present extensive simulation evidence. We conclude in section (ref), and we present supplementary results in three Appendices.

Data Generating Process and Estimators

Traditional OLS-HAC regression focuses exclusively on OLS parameter estimation, assuming consistency and surrendering on efficiency. But, as we emphasize in this section, even OLS consistency cannot be assumed without significant loss of generality. Moreover, aspects of the consistency and efficiency of OLS and various competitors, under various conditions, are nuanced and not widely appreciated. Hence in this section we begin by reviewing aspects of OLS consistency and efficiency in comparison to competitors -- in particular, a new procedure that we propose based on durbin1970testing regressions, a new modified FGLS procedure, and traditional FGLS -- in a sequence of progressively-richer dynamic environments.

Data-Generating Process

We start with the standard data-generating process (DGP) in the OLS-HAC regression literature,

equation[equation omitted — 63 chars of source]

where $t=1,2,...,T$, $\beta$ \ is a $k$-vector of parameters, $x_{t}$ is a $k$-vector of covariance-stationary covariates and $u_{t}$ is a scalar covariance-stationary disturbance with $E(u_{t}u_{t}')=\sigma^{2}\Omega$.\footnote{Because $u_{t}$ is covariance-stationary, it can be serially correlated and/or conditionally heteroskedastic. In this paper we emphasize serial correlation exclusively, because serial correlation is the unique feature of time-series data relative to cross-section data. Cross sections do of course sometimes have a spatial dimension and therefore a natural ordering in space if not in time, and spatial correlation has recently begun to receive attention from a HAC estimation perspective, as in muller2022spatial. Spatial HAC estimation is, however, beyond the scope of this paper.} DGP ((ref)) is usually augmented with conditions such that OLS is consistent. Then the econometrician generally aims to provide standard error corrections that enable asymptotically valid inference. Note that such OLS-HAC regression involves just a static regression of $y_t$ on $x_{t}$, basically imported directly from cross-sectional micro-econometrics, with dynamics allowed only through $u_{t}$. We will later argue that such a framework is uncompelling in time-series environments, but it is the industry standard in OLS-HAC regression, so we maintain it for now.

Crucial insights will flow from adopting a starting point that allows for significant generality regarding possible relationships between $x_{t}$ and $u_{t}$. In particular, consider the Wold representation of the Gaussian vector process $z_{t}=(x_{t}^{\prime},u_{t})^{\prime}$,

equation[equation omitted — 79 chars of source]

The coefficient matrices are $\Xi_{0}=I$ and \[ \Xi_{i}=

pmatrix[pmatrix omitted — 71 chars of source]

, \] and ${\varepsilon}_{t}=({\varepsilon}_{x,t}^{\prime},{\varepsilon}_{u,t})^{\prime}$ is a vector white noise innovation process with $E({\varepsilon}_{t})=0$ and $\ E({\varepsilon}_{t}{\varepsilon}_{s}^{\prime})=0$ for $s \ne t$, and contemporaneous covariance matrix $E(\varepsilon_{t}\varepsilon_{t}^{\prime})=\Sigma$, where

\[ \Sigma=

pmatrix[pmatrix omitted — 75 chars of source]

. \] Under mild regularity conditions, the infinite vector moving-average representation (ref) is equivalent to the infinite vector-autoregressive (VAR) representation\footnote{Such regularity conditions include assumptions on the rate of decline of $\left\Vert \Xi_{i}\right\Vert $ towards zero as $i\rightarrow\infty$, for suitable norm $\left\Vert .\right\Vert $, to control the persistence of the process and to avoid phenomena such as long memory that complicate the analysis. For details see, e.g., dav2002 and references cited therein.}

equation[equation omitted — 84 chars of source]

where \[ \Psi_{i}=

pmatrix[pmatrix omitted — 72 chars of source]

. \] This setting encompasses a variety of DGPs, and we will consider the consistency and efficiency properties of different estimators under various restrictions imposed on ((ref)).

We now proceed to consider various estimation strategies that may be appropriate in the environment given by ((ref)) and ((ref)).

OLS Parameter Estimation and HAC Covariance Matrix Estimation

The OLS estimator of the regression parameter is of course \[ \hat{\beta}=(X^{\prime}X)^{-1}X^{\prime}Y. \] If $\Omega=I$, the limiting distribution of the OLS estimator is \[ T^{1/2}\left( \hat{\beta}_{OLS}-\beta\right) \rightarrow N\left( 0,\sigma^{2}Q^{-1}\right), \] where $Q=p\lim_{T\rightarrow\infty}\left( T^{-1}X^{\prime}X\right)$.

Based on the $VAR$ representation ((ref)), we define “block diagonality" ($BD$) as holding when ${\Psi}_{ux,i}={\Psi}_{xu,i}=0$, for all $i$, and ${\Sigma}_{xu}=0$. The BD condition implies strong exogeneity, namely that $E(u_s|x_t)=0$ for all $s$ and $t$.\footnote{Strong exogeneity is sometimes called strict exogeneity.} In the $BD$ environment OLS is consistent but asymptotically inefficient, with limiting distribution \[ T^{1/2}(\hat{\beta}_{OLS}-\beta)\rightarrow N(0,V), \] where $V=Q^{-1}\Omega Q^{-1}$. The key object in $V$ \ is $\Omega$, which is the spectrum of $\ x_{t}u_{t}$ at frequency zero. HAC inference estimates $V$ using \[ \widehat{V}=Q^{-1}\widehat{\Omega}Q^{-1}, \] where $\widehat{\Omega}$ is a consistent estimator of $\Omega$, so that $\widehat{V}$ is consistent for $V$. Different choices for $\widehat{\Omega}$ therefore define different HAC covariance matrix estimators and are the main issue in implementing OLS-HAC regression, as we discuss subsequently in section (ref).

FGLS Estimation

If condition BD holds, and if the matrix $\Omega$ is known, then GLS is a consistent and asymptotically efficient estimator of $\beta$. However, $\Omega$ is almost always unknown, in which case attention turns to FGLS as defined by amemiya1973generalized, which is again both consistent and asymptotically efficient provided that condition BD holds.\footnote{Recent contributions to the FGLS literature include RW2017 for heteroskedastic environments, and kapetanios2016semiparametric for dynamic environments.}

The OLS-HAC regression literature was historically motivated by environments where OLS is consistent for $\beta$, but where condition $BD$ simultaneously fails in such a way that FGLS is inconsistent. Such situations are possible, and we will discuss a classic such situation hansen1980forward at some length in section (ref) below, but they are by no means the only or the most important possibility. Indeed there is much more to investigate when $BD$ fails, as emphasized in the insightful work of glsp.

We now consider an alternative estimation procedure that avoids the above discussed OLS-HAC and FGLS complications and always delivers consistent (and sometimes fully efficient) estimates of $\beta$, together with reliable asymptotic inference.

Durbin Estimation and its Relatives

A natural third approach to estimation and inference, which we will argue is generally preferable to both OLS-HAC and FGLS, is based on the “durbin1970testing regression", given by

equation[equation omitted — 168 chars of source]

where ${\varepsilon}_{y,t}$ is serially uncorrelated, and uncorrelated with $y_{t-j}$ and ${x}_{t-j}$ for all $j$. The Durbin regression “cleans out" disturbance dynamics by its direct inclusion of $y_{t-j}$ and ${x}_{t-j}$, so that standard OLS estimation and inference are trustworthy. We refer to the Durbin regression, and the associated estimator of $\beta$, as DURBIN. Crucially, note well that the DGP remains (ref) and (ref); DURBIN is simply a certain procedure (regression) that can be implemented on data from that DGP, just as OLS and FGLS are certain procedures that can be implemented on data from that DGP.

Operationally, it is of course necessary to use a finite order approximation to the infinite order DURBIN regression (ref),

equation[equation omitted — 160 chars of source]

with finite lag order $p$ selected using a data based procedure, typically an information criterion, and increasing at a suitable rate. The theoretical validity of such a procedure for producing valid asymptotic estimation and inference is well known (see, e.g., lewis1985prediction or hannan2012statistical), and we shall have more to say about it when we later implement DURBIN in the simulations of section (ref).

We can also write the finite order DURBIN approximation as

equation[equation omitted — 164 chars of source]

which emphasizes the extent of the parameterization and lag structure. We will later explore in greater detail the relationship between the DGP given by ((ref)) and ((ref)) and the DURBIN regression (ref), which is effectively one equation of a VAR and appears to be the originator of the autoregressive distributed lag (ADL) model model, which is widely used in empirical econometric work.

An estimator closely related to DURBIN, recently proposed by glsp, is a variation on FGLS. We refer to it as FGLS-D (short for “FGLS-Durbin"). While FGLS uses a first-stage OLS regression, FGLS-D uses a first-stage DURBIN regression ((ref)). Under $BD$, it follows that FGLS-D is also efficient. However, when $BD$ does not hold, FGLS-D may not be efficient or even consistent, while DURBIN remains consistent.

Given that condition $BD$ may not hold, it is important to consider the implications of its violation for the various methods of estimation and inference. To see the effects of the various sub-conditions embedded in condition $BD$, we will relax it in sequential stages. First, we impose only that ${\Psi}_{ux,i}=0$ for all $i$ and that ${\Sigma}_{xu}=0$, so that $x$ is weakly exogenous (that is, $E(u_s|x_t)=0$, $\forall s, t {\le}s$) but not strongly exogenous.\footnote{Weak exogeneity is sometimes called predeterminedness. See anna.} ${x}_{t}$ now depends on lags of $u_{t}$, but not vice versa. We refer to this restriction as $GEXOG$ (“GLS exogeneity"). Clearly, OLS is now inconsistent, as is FGLS, which uses OLS residuals, while FGLS-D remains consistent and efficient. Importantly, DURBIN remains consistent, even if not fully efficient, throughout.

Second, we impose only ${\Sigma}_{xu}=0$, so that $x$ is neither strongly nor weakly exogenous. We denote this condition by $EBD$ (“error variance block diagonal"). $u_{t}$ now depends on lags of ${x}_{t}$, and the finite-ordered FGLS autoregression for $u_{t}$ is no longer valid. Therefore, neither FGLS nor FGLS-D is consistent. DURBIN, however, remains consistent under $EBD$, and moreover it is also efficient.

To see the consistency and efficiency of DURBIN under EBD, note that, using ((ref)) and $u_{t}=y_{t}-{x}_{t}^{\prime}{\beta}$, we can write ((ref)) as

align[align omitted — 580 chars of source]

Noting that the relationship ${\gamma}_{j}={\Psi}_{xu,j}-\psi_{u,j}{\beta}$ gives a one-to-one mapping between ${\gamma}_{j}$ and ${\Psi}_{xu,j}$, given values for $\psi_{u,j}$ and ${\beta}$, we immediately obtain efficiency for DURBIN.\footnote{Of course, if ${\Psi}_{xu,j}=0$, then DURBIN, which estimates ${\gamma}_{j}$, is over-parameterized, providing a simple argument showing that DURBIN is inefficient under $GEXOG$.}

Finally, we impose no restrictions at all, in which case all methods become inconsistent and the use of instrumentation appears to be the only way forward.

In summary, OLS requires stronger conditions for consistency than the FGLS variants. The FGLS variants, in turn, require stronger conditions for consistency than DURBIN. Hence, overall, DURBIN has attractive consistency features in comparison with OLS and the FGLS variants. On the other hand, when the FGLS variants are consistent, they are also fully efficient. We shall see how such trade-offs resolve themselves in the simulations of section (ref) below.

A Generalized Data-Generating Process

We now move from the basic DGP (ref) to a generalized version that subsumes all cases of interest.

Data-Generating Process

Henceforth we work with the data-generating process given by

equation[equation omitted — 153 chars of source]

We emphasize that ${u}_{t}$ may also be a dynamic process, to allow, for example, for missing covariates. In particular, we continue to allow $(x_{t}^{\prime},u_{t})^{\prime}$ to follow the the vector moving average (ref), or equivalently, the vector autoregression ((ref)). This generalized DGP covers most linear dynamic relationships of conceivable interest. We use $NDY$ (“no dynamics in $y$") to refer to the restriction imposed on the generalized DGP ((ref)) to get the basic DGP ((ref)), namely $\phi_{j}={\gamma}_{i,j}=0$ $\, \forall i,j$.

We also emphasize that (ref) is now the data generating process, and various regressions could be fit to its data realizations in various attempts at estimation and inference for $\beta$. One such regression, for example, is FGLS. Clearly the use of FGLS in environments characterized by the generalized DGP (ref) accounts only for ${x}_{t}^{\prime}{\beta}$ and therefore ignores all terms involving lags, resulting in misspecification of the conditional mean part of the fitted regression. That is, the only way lagged information is used in FGLS is through estimation of the error covariance matrix, which neglects the problem of misspecification of the conditional mean.

DURBIN is another such regression that can be fit to the generalized DGP (ref). Indeed DURBIN can perfectly accommodate the generalized DGP, because, in precise parallel to ((ref)), we have

align*[align* omitted — 978 chars of source]

which is a DURBIN regression with

align*[align* omitted — 188 chars of source]

The above relationships between the parameters of the generalized DGP ((ref)) and the DURBIN regression ((ref)) also show that the generalized DGP is so richly parameterized that not all parameters are identified through estimation of ((ref)) alone. $\beta$ is always identified and consistently estimable via DURBIN, however, even in cases where OLS, FGLS, and FGLS-D are inconsistent.

table[table omitted — 1,168 chars of source]

Estimator Comparisons

In Table (ref) we summarize the consistency and efficiency properties of all estimators in the leading environments that we have considered, all of which are specializations of the generalized DGP given by (ref) and (ref). Table 1 makes clear the important trade-off between the occasional efficiency of FGLS/FGLS-D and the robust consistency of DURBIN. That is, although FGLS is sometimes efficient when DURBIN is not (under $NDY+BD$ and $NDY+GEXOG$), DURBIN is always at least consistent, and FGLS is not.

Indeed the $EBD$ row of Table (ref) is starkly revealing, as for example it includes simple and natural DGPs like \[ y_{t}=x_{t}\beta+\phi y_{t-1}+x_{t-1}\gamma+u_{t}. \] The conventional FGLS procedure would be to regress $y_{t}$ on $x_{t}$, and then to regress the residuals on lagged residuals, thereby obtaining the Cochrane-Orcutt filter to apply to the $y_{t}$ and $x_{t}$ series. One strongly suspects, and our subsequent simulations in section (ref) show clearly, that FGLS will perform poorly in this environment unless $ \gamma \approx \beta \phi$, in which case the DGP reduces (approximately) to just a static regression of $y_{t}$ on $x_{t}$ with $AR(1)$ disturbances.\footnote{The restriction $ \gamma \approx \beta \phi$ is known as the common factor restriction.}

Hausman Tests

Table (ref) also highlights the potential usefulness of tests for validity of the various restrictions. If for example, one “knew" that $NDY +GEXOG $ held, then FGLS or FGLS-D would be fully appealing estimators (consistent and efficient) whereas DURBIN would be less appealing (consistent but not efficient). Alternatively, if one knew that instead $NDY + EBD$ held, then FGLS or FGLS-D would be highly unappealing (inconsistent) whereas DURBIN would be fully appealing (consistent and efficient).

Hausman tests are available, as follows. Clearly, restrictions on the parameters of ((ref)) determine the comparative desirability of alternative methods of estimation and inference for $\beta$. The key restriction is $BD$. Under the null hypothesis that $BD$ holds with $u$ serially correlated, OLS is consistent but not efficient, while FGLS is both consistent and efficient. Under the alternative hypothesis that $BD$ fails, OLS and FGLS are generally both inconsistent but have different limits, which depend on the parameters of ((ref)). As a result, Hausman tests can be used.

In particular, one may wish to query whether $\beta_1=\beta_2$, where \[ \label{sim}E(y_{t}|x_{t})=x_{t}^{\prime}\beta_1 \] and \[ E(y_{t}|x_{t},x_{t-1},y_{t-1}, x_{t-2},y_{t-2},...)=x_{t}^{\prime}{\beta_2} + \sum_{j=1}^{\infty}\phi_{j}y_{t-j}+\sum_{j=1}^{\infty}x_{t-j}^{\prime} \gamma_{j}. \] Under the null hypothesis, FGLS should be used. Otherwise one should consider using FGLS-D or DURBIN if one is interested in ${\beta_2}$ as would typically be the case, or consider using OLS if for some reason $\beta_1$ is of interest.

Overall, however, we find it preferable simply to use DURBIN under all circumstances, unless there is some compelling reason to do otherwise. There are three reasons:

enumerate• An acceptable HAC estimator of the variance of the OLS estimator may not be available when implementing a Hausman test. Indeed the poor performance of OLS-HAC is the theme of this paper. • As regards consistent/efficient estimation, it will be clear from the simulation results in section (ref) below that the MSE cost of using DURBIN when a more efficient estimator is available (i.e., when $BD$ or at least $GEXOG$ holds) is generally small, whereas the MSE cost of not using DURBIN can be very large when neither $BD$ nor $GEXOG$ holds. • As regards consistent inference, it will also be clear from the simulation results in section (ref) below that DURBIN-based inference performs well in all circumstances that we investigate, both in terms of test size and power, in contrast to all other methods that we consider, where inference often fails.

We will shortly turn to the extensive simulation results alluded to in points (ref) and (ref) above, but first we briefly consider DURBIN vs other estimation approaches in the important context of predictive inference.

Predictive Inference

As is clear from Table (ref), OLS is rarely consistent in time-series situations of interest. One case where OLS is consistent and simultaneously FGLS is inconsistent involves multi-step forecast evaluation, where one tests whether a forecast $x_{t}$ is unbiased for $y_{t+k}$. That is, one tests whether \[ E(y_{t+k}|x_{t})=x_{t}, \] for $k \ge 1$.

One of the earliest analyses of this problem was by hansen1980forward, where $y_{t+k}$ represented the $k$-period-ahead spot exchange rate and $x_{t}$ represented the current $k$-period forward rate. The null hypothesis of $\beta=1$ implies moving-average disturbances, producing a violation of strong exogeneity while nevertheless satisfying weak exogeneity. hansen1980forward recognized that FGLS can be inconsistent in such a situation, whereas OLS remains consistent, albeit inefficient. They recognized, moreover, that the OLS standard error was inconsistent and therefore required a “correction" -- and OLS-HAC was born.

Note however, that DURBIN is also perfectly applicable in the Hansen-Hodrick environment, delivering not only consistent standard errors, but also efficient as opposed to merely consistent parameter estimates.\footnote{For a full empirical analysis, see baillie2023new.} In particular, under the null of unbiasedness, the error term, \[ u_{t+k}=y_{t+k}-x_{t}, \] satisfies $Cov(u_{t+j}u_{t})=0$ for $j>k$, which implies that $u_{t+k}$ can be represented by an $MA(k-1)$ process. Hence we can write

equation[equation omitted — 72 chars of source]

where $\varepsilon_{t}$ is a white noise process and $\theta(L)$ is a polynomial in the lag operator of order $k- 1$.

Conceptually, equation ((ref)) is merely a restricted DURBIN model, because on using the filter $\theta(L)^{-1}$ we obtain

equation[equation omitted — 134 chars of source]

The filtered explanatory variable is uncorrelated with current and future innovations, $\varepsilon_{t+k}$, so that estimation of equation (ref) by OLS will produce consistent and asymptotically efficient estimates of the regression parameters. In practice it is convenient to use the approximation $\theta(L)^{-1}\approx\pi(L)$, where $\pi(L)=\left( 1-\pi_{1}L-...-\pi_{p}L^{p}\right)$ is a $p$th-order lag-operator polynomial with all roots outside the unit circle. DURBIN will then be

equation[equation omitted — 77 chars of source]

which is a restricted version of the generalized DGP ((ref)) and can also be estimated by restricted OLS.

Simulation Evidence on Estimation and Testing

In this section we examine, via simulation, the sampling properties of the various estimators, the properties of forecasts that use those estimated parameters, and crucially, the size and power of associated hypothesis tests.

Simulation Design

The main simulation results will comprise four data generation processes that impose different assumptions on the generalized DGP given by ((ref)) and ((ref)):

enumerate• Autoregressive Disturbances, AR(1) ($NDY + BD$) \begin{align} y_{t} & =\beta x_{t}+u_{t}\nonumber\% \begin{pmatrix} x_{t}\\ u_{t} \end{pmatrix} & = \begin{pmatrix} 0.7 & 0\\ 0 & \rho \end{pmatrix} \begin{pmatrix} x_{t-1}\\ u_{t-1} \end{pmatrix} + \begin{pmatrix} {\varepsilon}_{x,t}\\ {\varepsilon}_{u,t} \end{pmatrix} \end{align} • Triangular vector autoregression (VAR) on ((ref)) ($NDY + GEXOG$) \begin{align} y_{t} & =\beta x_{t}+u_{t}\nonumber\% \begin{pmatrix} x_{t}\\ u_{t} \end{pmatrix} & = \begin{pmatrix} \psi_{11} & \psi_{12}\\ 0 & \psi_{22} \end{pmatrix} \begin{pmatrix} x_{t-1}\\ u_{t-1} \end{pmatrix} + \begin{pmatrix} {\varepsilon}_{x,t}\\ {\varepsilon}_{u,t} \end{pmatrix} \end{align} • Unrestricted VAR on ((ref)) ($NDY + EBD$) \begin{align} y_{t} & =\beta x_{t}+u_{t}\nonumber\% \begin{pmatrix} x_{t}\\ u_{t} \end{pmatrix} & = \begin{pmatrix} \psi_{11} & \psi_{12}\\ \psi_{21} & \psi_{22} \end{pmatrix} \begin{pmatrix} x_{t-1}\\ u_{t-1} \end{pmatrix} + \begin{pmatrix} {\varepsilon}_{x,t}\\ {\varepsilon}_{u,t} \end{pmatrix} \end{align} • Dynamic Regression ($EBD$) \begin{align} y_{t} & =\beta x_{t}+\rho y_{t-1}-0.5 x_{t-1} +u_{t}\nonumber\% \begin{pmatrix} x_{t}\\ u_{t} \end{pmatrix} & = \begin{pmatrix} 0.7 & 0\\ 0 & 0 \end{pmatrix} \begin{pmatrix} x_{t-1}\\ u_{t-1} \end{pmatrix} + \begin{pmatrix} {\varepsilon}_{x,t}\\ {\varepsilon}_{u,t} \end{pmatrix}. \end{align}

In all cases, $(\varepsilon_{x,t} , \varepsilon_{u,t})^{\prime}\sim iid N(0,I)$ with $t=1,...,T$. We explore $T \in\{ 50, 200, 600, 2500\}$, which also spans the relevant range for macroeconomics, where structural change and other considerations tend to keep sample spans to roughly “the most recent fifty years"; that is, sample sizes of 50 years, 200 quarters, 600 months, or approximately 2500 weeks. Including $T=2500$ also lets us check our Monte Carlo results against known large-sample results.

The autoregressive DGP in ((ref)) matches the design in lazarus2018har. We explore $\rho\in\{ 0, .3, .5, .7, .9, .95, .99\}$, which spans the relevant range for economics. All $\rho$ values are positive, as economic time series are generally positively serially correlated, and they range from white noise to the very strong serial correlation often of relevance in macroeconomic series. Including the white noise case ($\rho=0$) allows us to check our Monte Carlo results against known results for the iid case.

In the simulations for the triangular VAR DGP in ((ref)), we consider the following values for the matrix $\Psi$: \[ \Psi_{1}=

pmatrix[pmatrix omitted — 36 chars of source]

\quad\quad\Psi_{1}^{*}=

pmatrix[pmatrix omitted — 36 chars of source]

. \] $\Psi_{1}^*$ has a larger leading eigenvalue than $\Psi_{1}$ (0.6 versus 0.5) and hence exhibits stronger autoregressive features.

For the unrestricted VAR DGP in ((ref)), we consider the following values: \[ \Psi_{2}=

pmatrix[pmatrix omitted — 38 chars of source]

\quad\quad\Psi_{2}^{*}=

pmatrix[pmatrix omitted — 38 chars of source]

\] As before, $\Psi_{2}^{*}$ was selected to be similar to $\Psi_{2}$ but with a larger leading eigenvalue (0.97 for $\Psi_{2}^{*}$ and 0.91 for $\Psi_{2}$).

For the dynamic regression DGP in ((ref)), we consider various parameter values for the the coefficient on $y_{t-1}$, namely $\rho\in\{0, 0.5, 0.7, 0.95\}$. When $\rho=0.5$, the common factor restriction introduced in footnote (ref) holds, in which case we expect FGLS and FGLS-D to perform well.

For all DGPs in our simulations we perform 10,000 Monte Carlo replications. We simulate exact realizations of $x$ and $u$ by drawing $x_{0}$ and $u_{0}$ from their stationary distribution at each Monte Carlo replication, and we use common random numbers whenever appropriate.

Operational Considerations

Next, we detail operational matters relating to our implementation of the various estimators we use in our simulations.

OLS-HAC

OLS-HAC estimation proceeds from the approach previously outlined in section (ref); namely \[ T^{1/2}(\hat{\beta}_{OLS}-\beta)\rightarrow N(0,V), \] where $V=Q^{-1}\Omega Q^{-1}$ and \[ \Omega=\sum_{\tau=-\infty}^{\infty}\Gamma(\tau) \] where $\Gamma(\tau)=cov(x_{t}u_{t},x_{t-\tau}u_{t-\tau}),$ and $\tau=0,\pm1,...$

The key object in $V$ is $\Omega$, the spectrum of $xu$ at frequency zero. The OLS-HAC approach uses \[ \widehat{V}=Q^{-1}\widehat{\Omega}Q^{-1}, \] where $\widehat{\Omega}$ is a consistent estimator of $\Omega$ and hence $\widehat{V}$ delivers a consistent estimator of $V$.

A large literature on consistent estimation of $\Omega$ can be traced back to at least hansen1980forward. The most popular approach is due to newey1987simple, who propose lag-window estimation with linearly-decreasing (Bartlett) lag window:

equation[equation omitted — 262 chars of source]

where \[ \boldsymbol{\widehat{\Gamma}}_{\tau}=\frac{1}{T}\sum_{t=1}^{T}\hat{u}_{t} x_{t}x_{t-\tau}^{\prime}\hat{u}_{t-\tau}, \] the $\hat{u}_{t}$ are OLS regression residuals, and $T$ is sample size. Indeed, many leading HAC estimators are of the form ((ref)), distinguished only by their choice of truncation lag $h$.

We will explore several leading truncation lag choices, including:

enumerate• NW: Newey-West ((ref)) with $h=\lceil(T/100)^{2/9}\rceil$. This $h$ choice is a standard textbook recommendation (e.g.,wooldridge2015introductory). • NW-A: Newey-West ((ref)) with $h = \lceil0.75T^{1/3}\rceil$. This $h$ choice is also standard, arising when a formula in andrews1991heteroskedasticity is specialized to the case of a first-order autoregression with coefficient 0.25. • NW-LLSW: Newey-West ((ref)) with $h=\lceil1.3T^{1/2}\rceil$, as proposed by lazarus2018har. Its use of $T^{1/2}$ rather than $T^{2/9}$ or $T^{1/3}$ as in NW or NW-A, respectively, produces higher truncation lags. For example, if $T=200$, then NW selects $h=5$ but NW-LLSW selects $h=19$. • NW-KV: Newey-West ((ref)) with $h=T$, as proposed by kiefer2002heteroskedasticity, which builds on kiefer2000simple. Setting $h=T$ is of course the maximum possible truncation lag.

We will also explore the muller2014hac HAC estimator (we denote it by M), which is not in the Newey-West family. Instead, it is an orthogonal series estimator, that uses a type-II discrete cosine transform to produce an equally-weighted average of projections on cosines. The M estimator is: \[ \widehat{\Omega}=\frac{1}{\nu} \sum_{j=1}^{\nu}\widehat{\Lambda} _{j}\widehat{\Lambda}_{j}^{\prime}, \] where \[ \widehat{\Lambda}_{j}= \sqrt{\frac{2}{T}}\sum_{t=1}^{T} (x_{t}\hat{u}_{t} )\cos\left( \pi j \left( \frac{t-1/2}{T}\right) \right) \] The M truncation parameter, $\nu$, is the total number of cosines included in the average projection. lazarus2018har suggest setting $\nu=\lfloor0.4T^{2/3} \rfloor$, producing the M-LLSW estimator.

FGLS and FGLS-D

If the data follow the DGP in ((ref)), namely $y_{t}=x_{t}\beta+u_{t}$, and there exists a known lag operator polynomial (filter) $\Phi(L)$ that reduces $u_{t}$ to white noise $\varepsilon_{t}$ (i.e., $\Phi(L)u_{t}=\varepsilon_{t}$), then GLS estimation of $\beta$ is appropriate, and it amounts to running an OLS regression on transformed data. Specifically, one regresses $\tilde{y}_{t}$ on $\tilde{x}_{t}$, where $\tilde{y}_{t}=\Phi(L)y_{t}$ and $\tilde{x}_{t}=\Phi(L)x_{t}$.

In practice, however, $\Phi(L)$ is unknown and needs to be approximated. The FGLS estimator uses $\phi(L)=1-\phi_{1}L-\phi_{2}L^{2}-...-\phi_{p}L^{p}$ and proceeds as follows:

enumerate• Run an OLS regression of $y_{t}$ on $x_{t}$, and obtain the residuals $\hat{u}_{t}$. • Fit an AR$(p)$ model to $\hat{u}_{t}$ (in particular, run an OLS regression of $\hat{u}_{t}$ on $\hat{u}_{t-1},...,\hat{u}_{t-p}$, with $p$ selected by AIC or BIC), and obtain the coefficients $\hat{\phi}_{1}, ...,\hat{\phi}_{p}$. • Construct the transformed data, \begin{align*} \tilde{x}_{t} & =x_{t}-\hat{\phi}_{1} x_{t-1}- ...- \hat{\phi}_{p} x_{t-p}\\ \tilde{y}_{t} & =y_{t}-\hat{\phi}_{1} y_{t-1}- ...- \hat{\phi}_{p} y_{t-p}. \end{align*} • Run an OLS regression of $\tilde{y}_{t}$ on $\tilde{x}_{t}$ to obtain the FGLS estimator of $\beta$.

The FGLS-D estimator relies on a different first-stage procedure, replacing the regressions in steps (ref) and (ref) above with a single DURBIN regression, proceeding as follows:

enumerate• Run the OLS DURBIN regression (with $p$ selected by AIC or BIC), \begin{align*} y_{t} & =\sum_{j=1}^{p}\varphi_{j}y_{t-j}+ \sum_{i=1}^{k} \beta_{i} x_{i,t} +\sum_{j=1}^{p} \sum_{i=1}^{k} \gamma_{i,j}x_{i,t-j}+\varepsilon_{t}. \end{align*} • Use the estimated coefficients on the lags of $y_{t}$, $\hat{\varphi}_{1},...,\hat{\varphi}_{p}$, to construct the transformed data, \begin{align*} \tilde{x}_{t} & =x_{t}-\hat{\varphi}_{1} x_{t-1}- ...- \hat{\varphi}_{p} x_{t-p}\\ \tilde{y}_{t} & =y_{t}-\hat{\varphi}_{1} y_{t-1}- ...- \hat{\varphi}_{p} y_{t-p}. \end{align*} • Run the OLS regression of $\tilde{y}_{t}$ on $\tilde{x}_{t}$ to obtain the FGLS-D estimator of $\beta$.

DURBIN

As previously noted, the DURBIN regression augments regression ((ref)) with lags of $y$ and $x$ to capture dynamics, very much in the spirit of an arbitrary equation in a vector autoregression, as suggested by durbin1970testing.\footnote{Also related is the important recent work of Plagborg, who study lag-augmented local projection estimators of impulse-response functions in vector autoregressions.} The $p^{\mathrm{th}}$-order DURBIN regression is

equation[equation omitted — 167 chars of source]

which has $p+k+kp$ parameters.

If $u_{t}$ in equation (ref) is a finite-ordered AR$(p)$ process with $p$ known, then DURBIN holds exactly. In particular, we have\footnote{Note that DURBIN does not impose the common factor restriction embedded in ((ref)), namely that $\gamma_{i,j}=\beta_{i}\phi_{j} ~ \forall i,j$, in which case DURBIN coincides with FGLS. See Sargan1964 and HendryMizon1978.}

align[align omitted — 320 chars of source]

Hence the usual asymptotic inference is immediately available:

equation[equation omitted — 98 chars of source]

where $\vartheta_{OLS}$ is the vector of DURBIN parameters,

equation[equation omitted — 93 chars of source]

and $z_{t}^{\prime}=\left( y_{t-1},...,y_{t-p},~x_{1,t},...,x_{k,t} ,~x_{1,t-1},...,x_{k,t-1},~...,~x_{1,t-p},...,x_{k,t-p}\right) $.

In the more compelling case where $p$ is unknown and must be selected (implemented in our Monte Carlo below), the DURBIN regression ((ref)) is approximate rather than exact. However, the limiting distribution ((ref)) remains valid if $p$ is selected suitably grenander1981abstract, hannan2012statistical, as achieved by standard criteria with well-known optimality properties.\footnote{OLS-HAC regression, in contrast, typically relies on one or another of various \textquotedblleft rules of thumb" for bandwidth (truncation, $h$ or $\nu$) selection. “Automatic" bandwidth selection has, however, been considered in Andrews-Newey-West environments by andrews1991heteroskedasticity, AndrewsMonahan1992, and NeweyWest1994, among others.} In particular, if a $p_{max}$ is known such that $p\leq p_{max}$, then a consistent selection criterion (in the model selection sense) like BIC is a natural choice. Alternatively, in the absence of a $p_{max}$, an efficient selection criterion (in the model selection sense) like AIC is a natural choice.\footnote{In the Gaussian case, we have $\text{BIC}=T\log(\mathrm{SSE})+\log(T)(p+k+kp)$ and $\mathrm{AIC}=T\log(\mathrm{SSE})+ 2(p + k + kp)$, where SSE is the DURBIN regression sum of squared errors.}

Estimation Accuracy

We first examine the accuracy of our four estimators (OLS, FGLS, FGLS-D, and DURBIN) under our four DGPs ($NDY + BD$, $NDY + GEXOG$, $NDY + EBD$, $BD$). The key object of interest is $\mathrm{RE_{est}}$, the efficiency of DURBIN relative to OLS, FGLS or FGLS-D. For example: \[ \mathrm{RE_{est}(OLS) = \frac{MSE(\text{OLS})}{MSE(DURBIN)}. } \] We also show MSE and bias.\footnote{Note that {all} OLS-HAC estimators simply use the OLS estimator of $\beta$. Particular HAC estimators will have particular effects on the standard errors of $\hat{\beta}$, but not on $\hat{\beta}$ itself, which always remains just $\hat{\beta}_{OLS}$.}

table[table omitted — 5,363 chars of source]
figure[figure omitted — 683 chars of source]

\paragraph{Autoregressive Disturbances DGP ($NDY + BD$).}

Results appear in Table (ref). Let us begin directly with the $\mathrm{RE_{est}}$ results for DURBIN relative to OLS. For any fixed sample size $T$, $\mathrm{RE_{est}}$ is increasing in serial correlation strength $\rho$. Consider, for example, a leading case like $T=200$ corresponding, to fifty years of quarterly data. For $\rho=0$, $\mathrm{RE_{est}}$ is close to 1, as it should be since there is no serial correlation. $\mathrm{RE_{est}}$ grows quickly as $\rho$ increases, however, reaching 2.9 when $\rho=0.7$ and 36.3 when $\rho= 0.95$.

In contrast, for any fixed serial correlation strength $\rho$, $\mathrm{RE_{est}}$ stabilizes quickly in sample size $T$ and remains approximately constant. Consider, for example, a realistic case like $\rho=0.9$. $\mathrm{RE_{est}}$ remains at approximately $\mathrm{RE_{est} =12}$ for all sample sizes $T\in\{50, 200, 600, 2500\}$. Hence $\mathrm{RE_{est}}$ is clearly driven by serial correlation strength and not by sample size.

In Figure (ref) we provide a visual representation of the $\mathrm{RE_{est}}$ of DURBIN relative to OLS presented in Table (ref). It reveals clearly that $\mathrm{RE_{est}}$ is driven entirely by the degree of serial correlation and not by sample size.

Now consider separately the MSEs for OLS and DURBIN that underlie $\mathrm{RE_{est}}$. For any fixed sample size $T$, the MSE of OLS is strongly increasing in serial correlation strength $\rho$ (because the OLS estimator ignores serial correlation), whereas the MSE from DURBIN is invariant to serial correlation strength (because the DURBIN estimator controls for serial correlation). That is why the $\mathrm{RE_{est}}$ ratio is also strongly increasing in $\rho$, as documented earlier. In contrast, for any fixed serial correlation strength $\rho$, the MSEs for both OLS and DURBIN decrease with sample size $T$ (as they must, since both OLS and DURBIN are consistent), but they decrease proportionately, so that the $\mathrm{RE_{est}}$ ratio is invariant to $T$, as documented earlier.

Next, let us examine the bias and variance components that underlie the MSEs. First consider bias. Both the OLS and DURBIN estimators are theoretically unbiased for any serial correlation strength and sample size, and the Monte Carlo confirms the theory: the estimated biases are always negligible and invariant to $\rho$.\footnote{Moreover the estimated biases decrease with $T$, as expected, by consistency.} Moreover, given the scale of the bias, the patterns mentioned above for MSE will correspond to patterns in variance: OLS variance increases sharply with serial correlation strength (because OLS ignores serial correlation), whereas DURBIN variance does not (because DURBIN controls for serial correlation), and both variances decrease with sample size (by consistency), but they do so proportionately. That is, the MSE patterns between OLS and DURBIN, and hence the corresponding $\mathrm{RE_{est}}$ patterns, are driven entirely by variance.

\paragraph{Triangular and Unrestricted VAR DGPs ($NDY + GEXOG$, $NDY + EBD$).}

Results appear in Table (ref). $\Psi_{1}$ and $\Psi_{1}^{*}$ correspond to different parameterizations of the $NDY + GEXOG$ DGP, and $\Psi_{2}$ and $\Psi_{2}^{*}$ correspond to different parameterizations of the $NDY + EBD$ DGP.\footnote{Recall that $\Psi_{1}^{*}$ has a larger leading eigenvalue than does $\Psi_{1}$, and $\Psi_{2}^{*}$ has a larger leading eigenvalue than $\Psi_{2}$.} For all sample sizes, OLS and FGLS exhibit large bias and MSE, which is expected since they are indeed inconsistent under both $NDY + GEXOG$ and $NDY + EBD$. As a result, the RE$_{\text{est}}$'s for DURBIN relative to OLS and FGLS in Table (ref) are very large: DURBIN dominates both.

\paragraph{Dynamic Regression DGP ($EBD$).}

Results appear in Table (ref). In the $EBD$ case, OLS, FGLS and FGLS-D are in general inconsistent, whereas DURBIN remains consistent. This is reflected in the large biases and MSEs of the other estimators compared to DURBIN, and hence the high efficiency of DURBIN relative to OLS and FGLS.

A notable exception is when $\rho=0.5$, in which case the common factor restriction holds, so that it is possible to write the dynamic regression as a single-regressor equation (with just $x_{t}$) and a disturbance with AR(1) serial correlation. Put differently, in this case the DGP in ((ref)) can be rewritten in the form of ((ref)), so that FGLS and FGLS-D are consistent and efficient and should have lower MSE than Durbin. Table (ref) shows that this is the case for all sample sizes. This result highlights the role that the common factor restriction plays; if it holds, it guarantees that all dynamics enter through the disturbance term, so that FGLS and FGLS-D dominate DURBIN, but if it does not hold (and there is no reason why it should hold), DURBIN dominates.

table[table omitted — 3,936 chars of source]
table[table omitted — 4,361 chars of source]

Prediction Accuracy

One of the primary uses of regression and dynamic regression is for ex ante prediction. There is substantial previous literature related to the task of prediction. In particular, Baillie1979 has considered the situation of predictions from the regression model with $AR(p)$ errors and the properties of prediction from static regressions and also with optimal multi-step predictions in the sense of minimum MSE\ predictions. Baillie1979 also derived results on the efficiency of these predictors with and without estimated parameters. One conclusion concerns the importance of including the full effects of dynamics from the $AR(p)$ regression model in the predictor. In this case, the complete structural dynamic predictor generally has substantial asymptotic and small sample efficiency gains over predictors from static regressions. Similar effects and properties are found in more complicated dynamic models such as the DGP considered in section (ref) of this paper.

We now consider one-step-ahead predictions relying on the OLS and DURBIN estimation strategies. The results reflect that an explicit modeling of autocorrelation can be used for improved prediction. OLS estimators neglect this and therefore produce suboptimal predictions. To see this, first consider the case of a DGP with autoregressive disturbances and known parameter $\beta=1$.\footnote{We start with the case of known parameter $\beta$, as it can easily be solved analytically.} Specifically consider the DGP given by

align*[align* omitted — 122 chars of source]

with all shocks $N(0,1)$ and orthogonal at all leads and lags. For this DGP, the optimal prediction accounting for serial correlation in $u$ is

align[align omitted — 131 chars of source]

and the corresponding prediction error is $e_{t+1}^{opt}=\varepsilon_{x,t+1}{+}\varepsilon_{u,t+1}$, with variance $\sigma_{opt}^{2}{=}2$.

The suboptimal prediction, failing to account for serial correlation in $u$, is just the first term in ((ref)), \[ y^{subopt}_{t+1,t} = \rho x_{t}, \] with corresponding prediction error $e^{subopt}_{t+1} = \varepsilon_{x,t+1}{+}u_{t+1}$, and variance $\sigma^{2}_{subopt}{=}1 {+}\frac{1}{1{-}\rho^{2}}$.

Both predictions are unbiased, so the prediction efficiency of DURBIN relative to OLS ($\mathrm{RE_{pred}}$) is just the relative variance, which is

equation[equation omitted — 142 chars of source]

$\mathrm{RE_{pred}}$ is bounded below by 1, which occurs when $\rho{=}0$, and $\mathrm{RE_{pred} {\rightarrow} \infty}$ monotonically as $\rho{\rightarrow}1$.

Now we consider the case of estimated parameters, which is more complicated. In Table (ref) we show $\mathrm{RE_{pred}}$ estimated by Monte Carlo, accounting for parameter estimation uncertainty. For all but the most extreme cases (e.g., $T=50$ with $\rho=0.99$) the Monte Carlo results are almost identical to the analytic result ((ref)) that ignores parameter estimation uncertainty.\footnote{This is because the effects of parameter estimation uncertainty on MSPE vanish quickly (like $1/T$ rather than $1/\sqrt{T}$), as is well known. Hence the earlier-documented poor estimation efficiency of OLS relative to DURBIN, although a large problem for some purposes, is not an important problem for prediction.} Hence $\mathrm{RE_{pred}}$ depends strongly on $\rho$ but not on $T$. More precisely, for any $T$ we of course obtain $\mathrm{RE_{pred}=1}$ in the white noise case ($\rho=0$), but then $\mathrm{RE_{pred}}$ grows quickly in $\rho$, and for any $\rho$, $\mathrm{RE_{pred}}$ stabilizes extremely quickly in $T$ and is basically constant.

table[table omitted — 1,243 chars of source]

Inference

Now we consider the finite-sample properties of hypothesis tests associated with the various estimation procedures. We first consider test sizes, after which we consider rejection frequencies. In all tables in this section we consider the following estimators: OLS with unadjusted standard errors, five OLS-HAC estimators (NW, NW-A, NW-LLSW, NW-KV, and M-LLSW), FGLS, FGLS-D, and two implementations of DURBIN, one using BIC for lag order selection and the other using AIC. Additionally, we have included two Hausman tests; the first null hypothesis is that FGLS is efficient relative to OLS, and the second is that FGLS-D is efficient relative to DURBIN.

table[table omitted — 5,450 chars of source]
table[table omitted — 4,169 chars of source]
table[table omitted — 4,581 chars of source]

Size

Table (ref) contains results for the autoregressive disturbances DGP, $NDY + BD$. First, tests based on OLS are incorrectly sized for all $(\rho,T)$ combinations, except when $\rho=0$, and the size distortions become huge as $\rho$ grows. Second, the various NW HAC corrections reduce but do not eliminate the size distortion. In particular, distortion generally remains in the economically crucial region of $\rho\in[0.5, 0.99]$, depending on the sample size and the precise NW version used. NW and NW-A are worst, NW-LLSW are better, and NW-KV is the best. The M-LLSW HAC correction is different in that it exhibits an approximately correct size across $(\rho,T)$ combinations. Finally, tests based on FGLS, FGLS-D and DURBIN, in contrast, are correctly sized for all $(\rho,T)$ combinations, even with extremely strong autocorrelation. This holds regardless of whether DURBIN lag order selection is done with BIC or AIC.

Table (ref) contains results for the two VAR DGPs, $NDY + GEXOG$ and $NDY + EBD$. In the {NDY+ GEXOG} environment, OLS and FGLS are inconsistent, which produces large size distortions. In contrast, DURBIN and FGLS-D are consistent; they should outperform OLS and FGLS, and they do. DURBIN and FLGS-D should perform similarly, and they do. In the {NDY +EBD} environment, OLS, FGLS, and FGLS-D are inconsistent, and all have large size distortions. DURBIN, however, remains consistent and performs admirably.

Finally, Table (ref) contains results for the dynamic regression DGP, $EBD$. In this environment DURBIN should perform well, and it does, whereas all other test sizes are distorted, except at or near the very special common-factor case of $\rho=0.5$.

figure[figure omitted — 650 chars of source]

Power

Only tests that are correctly sized are of real interest, because only correctly-sized tests produce trustworthy and interpretable rejections. As we have shown, DURBIN satisfies that requirement, whereas OLS-HAC regression does not. One could simply stop there, but it is of interest to compare rejection frequencies in a few laboratory environments where the DGP is known. We do so in Figure (ref) for three of our DGPs with $T=200$ and various persistence parameters, comparing OLS-HAC (Kiefer-Vogelsang, LLSW), FGLS, and FGLS-D.

In the top row Figure (ref) we show rejection frequencies for the autoregressive disturbances environment, $NDY + BD$. All estimators are consistent, and all tests have correct size when $\beta=1$, i.e., when the true parameter equals its value under the null hypothesis. Moving away from the null however, it is clear that OLS-HAC power is inferior to that of DURBIN, because OLS is inefficient relative to DURBIN. Moreover, the inferior power performance of OLS-HAC increases with disturbance persistence ($\rho$), precisely because the relative inefficiency of OLS increases with persistence. Finally, DURBIN, FGLS and FGLS-D have virtually identical power curves.

In the middle row Figure (ref) we show rejection frequencies for the triangular VAR case, $NDY + GEXOG$. OLS-HAC and FGLS are so badly mis-sized that it is not worth discussing them, whereas FGLS-D is asymptotically correctly sized but is still over-sized for $T=200$. Only DURBIN is trustworthy. Moving from the middle-left to middle-right panel (higher persistence), the superiority of Durbin is amplified.

In the bottom row Figure (ref) we show rejection frequencies for the unrestricted VAR case, $NDY + EBD$. FGLS-D fails even asymptotically, so it is not surprising that its finite-sample performance is much worse than in the middle-row triangular VAR $NDY + GEXOG$ case. DURBIN, however, remains trustworthy. Moving from the bottom-left to bottom-right panel (higher persistence), the superiority of DURBIN is amplified, just as in the triangular case.

Concluding Remarks and Directions for Future Research

We have considered issues surrounding the time-series application of OLS regression with HAC standard errors. Although the OLS-HAC methodology is often sensible in cross-section regression situations, we argued that it is not generally an effective procedure in time-series regressions. Such regressions usually possess persistent autocorrelation, which causes OLS-HAC regressions to be highly sub-optimal for parameter estimation (in terms of efficiency), inference (in terms of both test size and power), and prediction.

We showed that the OLS-HAC problems are largely avoided by the use of a simple dynamic regression procedure, DURBIN. We demonstrated the significant advantages of DURBIN with detailed simulations covering a range of practical environments and issues. Effectively, DURBIN is a powerful tool for pre-whitened HAC estimation, in the tradition of AndrewsMonahan1992 -- indeed such a good pre-whitening tool that there's rarely any need for subsequent HAC estimation.

On the other hand, DURBIN is of course not a panacea. For example, DURBIN may struggle in small samples when dynamics have a strong moving-average component. Our Monte Carlo makes clear, however, that for all but the most extreme environments, DURBIN with lag order selected using standard information criteria performs consistently well. Indeed that is the key message of our paper.

In future work, one could generalize the DURBIN regression in various ways. For example, one could allow different lag lengths for $y$ and the $x_{i}$'s. One could also allow for heteroskedasticity, which we suppressed in this paper so as to focus exclusively on autocorrelation, for example by allowing for $GARCH$ disturbances in the DURBIN regression.