Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
104,830 characters · 24 sections · 92 citation commands
Local Projections or VARs? A Primer for Macroeconomists
}
Keywords: dynamic causal effect, impulse response, local projection, misspecification, vector autoregression. JEL codes: C22, C32.
Applied macroeconomists routinely seek to estimate the dynamic causal effects---or impulse response functions---of aggregate shocks to policies or fundamentals. The two dominant empirical methods for doing so are structural vector autoregressions (henceforth interchangeably referred to as VARs or SVARs), going back to Sims1980, and local projections (LPs), as introduced by Jorda2005. The primary objectives of this paper are, first, to review how to think conceptually about the choice between these two estimation methods, and second, to offer concrete recommendations for applied practice.
We begin with the observation that the choice between LPs and VARs has nothing to do with questions of identification: for any given LP and the economic identifying assumptions that it implements, there exists an equivalent VAR, and vice versa. At their core, typical identification schemes in empirical macroeconomics propose to recover the causal effects of aggregate shocks or policies as---sometimes very simple, and sometimes rather complicated---functions of the autocovariances of time series data. Conceptually, LPs and VARs are simply two ways of estimating these autocovariances: they share a common large-sample estimand, and only differ in how they exploit a given finite data set.
In the short samples typical of applied work, the choice between LPs and VARs is one of navigating a bias-variance trade-off. At one end of this spectrum, LPs (or equivalently, VARs with many lags) robustly have low bias, though at the cost of materially elevated variance. At the other end, VARs with few lags tend to deliver sizable precision gains, but at the cost of potentially substantial biases. The intuition is that short-lag VARs extrapolate based only on the first few autocovariances of the data, while LPs flexibly estimate autocovariances at all horizons, without extrapolation. In practice, the variance cost of LPs is so substantial that VARs are typically preferable in terms of mean squared error. The VAR's bias, however, severely threatens the accuracy of its uncertainty assessments, while LP confidence intervals instead accurately reflect statistical uncertainty, by virtue of the LP's small bias. There is “no free lunch”: if VARs afford any precision gains relative to LPs, then their uncertainty assessments are necessarily fragile. In fact, because plain LP is semiparametrically efficient, there is actually no impulse response estimator that is more efficient than LP in large samples without sacrificing robustness.
Our overall recommendation is that researchers who wish to understand what can be learned about dynamic causal effects from the data---and therefore require confidence intervals with accurate coverage probability---should report results based on LPs, or equivalently based on VARs with many more lags than is currently used in applied practice. Confidence intervals produced by short-lag VARs (or shrinkage techniques such as Bayesian VARs and penalized local projections) are simply too fragile to be trusted. However, if the researcher's goal is merely to forecast or to produce a point estimate of a causal effect for policy analysis, then short-lag VARs or shrinkage techniques are attractive due to their favorable mean squared error in many settings Li2024.
We conclude with a “how-to” list of practical recommendations for LP estimation and inference. We cover the importance of lag augmentation, selection of control variables and lag length, bias correction, and confidence interval construction.
Due to our focus on simple, concrete take-aways for applied practice, our review of the literature leaves out several specialized or new topics. These include panel data Almuzara2024, nonlinear specifications Caravello2024,Gonccalves2024, simultaneous confidence bands MontielOlea2019, variance decompositions Plagborg2022, and certain more technical structural shock identification schemes Uhlig2005,Baumeister2015. For brevity, we touch only briefly on proxy or instrumental variable identification, and we abstract from weak instrument issues, even though these are likely to be important in practice Montiel2021. Excellent reviews of LPs, VARs, and the relationship between them include Kilian2017, Stock2018, Baumeister2024, and Jorda2025.
\paragraph{Outline.} The organization of the paper follows the outline above: we first review questions of identification, then discuss theoretically and quantify empirically the bias-variance trade-off between LPs and VARs, next ask how to navigate that trade-off in practice, and finally close with a list of concrete practical recommendations. Throughout, our analysis will be deliberately simple, emphasizing clarity over comprehensiveness; for the interested reader we give references to the original literature. To structure our discussion, we summarize our main insights as “lessons” that should be viewed as guidelines rather than as formal propositions. Supplementary details---in particular for our empirically calibrated simulations---are provided in the online appendix, and replication code is available online.\footnote{The replication code is available at \url{https://github.com/ckwolf92/lp_var_nberma}.}
In this section we review the basic definition of the LP and VAR estimators, and we establish that they share the exact same estimand (i.e., large-sample limit) when the estimation lag length is large. This equivalence result is nonparametric, and in particular does not require the data generating process to be linear or finite-dimensional. The purpose of economic identifying assumptions is then to establish that this common estimand is actually interesting. The message of this section is that the choice between LPs and VARs is completely orthogonal to structural questions of identification.
We begin with a discussion of the estimand of LPs. LP practitioners seek to estimate the dynamic causal effects of macroeconomic “shocks” on aggregate outcomes. To this end, they run linear regressions of the following form, estimated separately by ordinary least squares (OLS) for each horizon $h = 0, 1, 2, \dots$:
They then report as the impulse responses of interest the regression coefficients $\{ \theta_h^\text{LP} \}_h$. In the above regression, $y_t$ is a (scalar) outcome of interest, $x_t$ is a (scalar) impulse variable, $r_t$ is a vector of time series that is included as contemporaneous controls, $w_t = (r_t', x_t, y_t, q_t')'$ collects all variables included as lagged controls, with $q_t$ a potential additional control vector, and $\xi_{h,t}$ is the multi-step forecast error in the regression. Here and in the rest of this section, we consider the limit where the empirical sample size is infinitely large, to abstract from finite-sample issues that will be the focus of the following sections.
By the Frisch-Waugh-Lovell theorem, the regression coefficient $\theta_h^\text{LP}$ has a simple interpretation: it equals the coefficient from a projection of the outcome $y_{t+h}$ on a “shock” $\tilde{x}_t$ that is given by the residual from a projection of $x_t$ on the control variables in the regression (ref):
This LP estimand is economically interesting because, under some assumptions, it can be given a structural interpretation. The argument is most transparent when a researcher directly observes an unpredictable shock (or a valid proxy/instrument for that shock), e.g., by observing the high-frequency responses of an asset price in narrow time windows around policy announcements Kaenzig2021. In this case, she can straightforwardly use the LP (ref) to estimate the causal effects of that shock: the impulse variable $x_t$ is the observed shock, the contemporaneous control vector $r_t$ is empty, and finally further controls $q_t$ are not necessary for consistent estimation, but should be included for efficiency and robustness reasons, as discussed in (ref).\footnote{If the researcher is willing to assume an underlying linear Structural Vector Moving Average (SVMA) model, then the LP regression coefficients equal the shock's true impulse responses (up to scale). If instead she assumes a more general non-linear causal model, then the estimand is a particular weighted average of marginal effects Plagborg2021,Kolesar2024.} Note that in this case the residualized shock $\tilde{x}_t$ simply equals the shock $x_t$ itself, by virtue of being unpredictable.
In another common class of applications, the impulse variable $x_t$ equals a policy instrument (such as the federal funds rate), the contemporaneous controls $r_t$ equal endogenous variables in the policymaker's reaction function (such as output and inflation), and finally some further lagged controls $q_t$ may be included as well. The residualized shock $\tilde{x}_t$ then equals the disturbance in the policy rule (such as a monetary policy shock) under the classic timing assumption that the endogenous variables $r_t$ do not respond within the period to this disturbance Christiano1999.\footnote{By defining the shock as the residual in a policy rule where all variables are observed by the econometrician, we are implicitly imposing the assumption of invertibility, i.e., that the policy shock is spanned by current and lagged observed variables. We will comment further on this structural assumption in (ref).} Intuitively, conditional on the included lagged and contemporaneous controls, the leftover movements in the policy instrument $x_t$ purely reflect shocks to policy, orthogonal to all other disturbances---i.e., exogeneity conditional on controls. We will elaborate on other, more involved ways of identifying macroeconomic shocks in (ref).
We now consider the second estimation method: VARs. The first step in VAR analysis is to estimate a reduced-form VAR in the vector of all observed time series $w_t$ (using the same definition as above):
VAR practitioners translate this reduced-form model into structural impulse response functions using identification conditions. Identifying assumptions are used to map the reduced-form residuals $u_t$ into interpretable structural shocks, and those shocks are then propagated forward through the VAR (ref). We begin with one popular way of doing so: orthogonalizing the reduced-form residuals $u_t$ using what is known as a recursive ordering. Let $BB' = \Sigma$ be the Cholesky decomposition of the covariance matrix $\Sigma \equiv \operatorname*{Var}(u_t)$ of the forecast errors $u_t$, where $B$ is a lower-triangular matrix with positive diagonal entries. As in (ref), we focus on the response of $y_t$ to the orthogonalized shock to $x_t$. Letting $e_x$ and $e_y$ denote the two unit vectors such that $x_t=e_x'w_t$ and $y_t=e_y'w_t$, the structural VAR impulse response estimate equals
where the reduced-form impulse responses satisfy the recursion $C_h = \sum_{\ell=1}^{\min \lbrace h,p \rbrace} A_\ell C_{h-\ell}$, with $C_0$ equal to the identity matrix. Evidently, the VAR impulse responses are obtained by extrapolation: the parameters of the reduced-form VAR (ref) are estimated from observed autocovariances out to lag $p$, and then the parametric structure of the model is exploited to compute responses at all horizons $h$, including $h>p$.
By standard properties of Cholesky decompositions, this procedure isolates as the “shock” the (rescaled) residual in a projection of $u_{x,t} \equiv e_x'u_t$ on $u_{r,t} \equiv e_r'u_t$, where $e_r$ is the unit vector such that $r_t=e_r'w_t$. By definition of the residuals $u_t$, this shock is the same as the residual in a projection of $x_t$ on $r_t$ and $p$ lags of the observables $w_t$---which in turn is precisely the shock $\tilde{x}_t$ in the LP (ref).\footnote{In the case of observed shock or proxy variable identification, the VAR specification that we discuss here includes the shock $x_t$ as an “internal instrument”. The alternative “external instrument” VAR procedure Stock2008,Mertens2013 is conceptually distinct and requires the additional assumption of invertibility Stock2018,Plagborg2021.} Differently from the LP, however, the impulse response coefficients $\{ \theta_h^\text{VAR} \}_h$ are now not direct projection coefficients of future outcomes $\lbrace y_{t+h} \rbrace_h$ on this common “shock;” instead, the dynamic effects of the shock are iteratively propagated forward through the estimated VAR model (ref).
The simple linear regression logic that we just used to interpret the LP and VAR estimands reveals that the methods are intimately connected: they deliver the dynamic causal effects of the exact same implicit “shock” $\tilde{x}_t$, just in one case through direct projection (LP), in the other through iterative propagation (VAR). On impact ($h = 0$), they are thus necessarily the same. Furthermore, as the lag length $p$ becomes arbitrarily large, the extrapolative reduced-form VAR model (ref) is sufficiently flexible to perfectly capture all autocovariance properties of the data, making iterative forecasts equivalent to direct projections; thus, the LP and VAR estimands are identical at all horizons $h$ Plagborg2021,Xu2023.\footnote{Here we abstract from an inessential issue: the implied LP and VAR shocks may have different variances. But once rescaled to have the same units (e.g., standard deviation equal to 1), equivalence follows.} Though this result pertains to linear estimation methods, we stress that the equivalence is extremely general because it is nonparametric, in the sense that it does not impose any particular parametric restrictions (linear or otherwise) on the underlying data generating process. Finally, for general but finite estimation lag length $p$, the extrapolative VAR model still captures the relevant autocovariance properties well out to lag $p$, so that the LP and VAR estimands are typically very close at horizons $h \leq p$, but not necessarily for $h>p$.\footnote{See Plagborg2021 for further discussion of the finite-$p$ case. As shown there, the LP and VAR estimands at $h \leq p$ for finite $p$ in general differ slightly because the computed impulse responses under LP depend on autocovariances up to lag $p + h$, not just up to $p$. Ludwig2024 shows that the LP impulse response can be written exactly as a function of VAR impulse responses with varying lag lengths.}
We provide a visual illustration of the second lesson in (ref), which plots LP and VAR impulse response estimands for a recursively identified monetary policy shock, following the classical identification approach of Christiano1999. The underlying data generating process (DGP) is the dynamic factor model that we study later in (ref); importantly, that DGP does not satisfy a finite-order VAR model. The three panels then plot the VAR and LP estimands for estimation lag lengths $p = \{ 2, 6, 12 \}$ (in red dashed and blue dashed-dotted, respectively), as well as the (common) population estimand ($p = \infty$, in grey). Consistent with our discussion above, LP and VAR estimands are close to each other at horizons $h \leq p$, before then deviating; in particular, for $p$ sufficiently large, they are identical throughout.
The large-sample equivalence between LPs and VARs extends beyond the observed shock and recursive identification schemes discussed above. A common approach to identifying macroeconomic shocks begins with the assumption of “invertibility”, i.e., that the true shocks are spanned by current and lagged time series observables $w_t$ Fernandez2007. Identification schemes like the recursive ordering of Christiano1999---which, as we showed above, can equivalently be implemented using either LPs or VARs---combine the assumption of invertibility with short-run timing restrictions on the shocks. Invertibility may however also be combined with other restrictions to identify shocks, including for example long-run restrictions Blanchard1989 or sign restrictions Uhlig2005. In all these identification approaches, the structural shock of interest ultimately equals $\beta'u_t$, where $u_t$ is the reduced-form forecast error from (ref), and the weight vector $\beta$ is a (potentially complicated) function of the researcher's identifying assumptions. A structural VAR practitioner would then report the impulse response estimate $\theta_h^\text{VAR}=e_y'C_h \Sigma \beta$. But by our earlier Frisch-Waugh-Lovell logic, we can alternatively compute impulse responses with respect to the same shock $\beta'u_t$ from the more general LP
Intuitively, even such richer approaches to identification ultimately amount to the researcher giving a causal interpretation to some (perhaps complicated) function of the autocovariance function of the observed data. Either VARs or LPs can then be used to compute this shared estimand of interest, with the two methods necessarily agreeing in large samples when the estimation lag length $p$ is large.\footnote{See Plagborg2021 for a detailed discussion of how exactly identification schemes like long-run and sign restrictions map into $\beta$.}
In conclusion, the choice between LPs and VARs has absolutely nothing to do with the identifying assumptions necessary to isolate a given shock of interest. Instead, the question is simply which method is better at (i) recovering this common estimand in finite samples and (ii) quantifying the associated statistical uncertainty. This is why estimation and inference---and not identification---will be the focus of the remainder of the paper.
We now dig deeper into the econometric properties of LPs and VARs. First, we motivate our analysis by demonstrating that, despite the population equivalence between LPs and VARs when the estimation lag length is large, the choice between the two estimators does matter in practice when using a small or moderate lag length, as is typical in applied work. Second, we present simple illustrative simulations of the properties of LP and VAR estimators. Finally, we review the available econometric theory. The overarching takeaway of the analysis will be that, in finite data sets, there is a clear and inevitable bias-variance trade-off between LPs and (short-lag) VARs: small bias and large variance for LPs, and vice versa for VARs. In fact, because plain LP is semiparametrically efficient, no estimator can outperform it in terms of variance without sacrificing robustness.
We first establish that, in finite samples and with a moderate estimation lag length, LPs and VARs can provide meaningfully different estimates of their common estimand (that would, as we saw, obtain in large samples if we were able to control for very many lags). Our analysis is based on the literature synthesis of Ramey2016. We replicate four of the headline applications of that paper, for shocks to monetary policy, taxes, government purchases, and technology, respectively. We then use LPs and VARs to estimate impulse responses for several variables and at several horizons, throughout staying as close as possible to the specifications considered by Ramey2016. Across all response variables and horizons, we then compare standard errors and compute the differences in point estimates. Implementation details are provided in (ref).
(ref) shows that there are meaningful differences in both precision as well as location of the two estimators: VAR impulse responses often have substantially lower standard errors than LPs (left panel), and the two sets of point estimates can be quite far apart (right panel). In the left panel we display a box plot of the ratio of VAR and LP standard errors across all different shocks and outcome variables, separately for short horizons (blue, $\leq$ one year) as well as long horizons (grey, $>$ one year). At short horizons (i.e., for $h$ close to or below the lag length), VARs are only somewhat more precise than LPs, with the median standard error ratio only slightly below one, consistent with our theoretical discussion in (ref).\footnote{Asymptotically, VAR standard errors are weakly smaller than LP standard errors, so the standard error ratio is bounded above by $1$, as we will explain later. This need not be the case in small samples, however.} At long horizons, VARs are instead materially more precise, yet again as expected given the preceding discussion. The right panel then shows the difference in point estimates between the two methods, normalized by the VAR standard error. We see that differences can often be of much greater magnitude than the VAR standard error. The gaps are instead smaller at short impulse response horizons, as expected.
These empirical results can of course not directly tell us whether the LP or the VAR estimates tend to be closer to the truth. For this, we will next turn to a simulation exercise in a simple data generating process, before presenting the general econometric theory.
As a first step toward gaining intuition for the econometric properties of LPs and VARs, we present some illustrative simulations. For pedagogical reasons, we here assume a simple univariate ARMA$(1,1)$ DGP, while leaving empirically realistic simulations to later sections:
where $y_t$ is an observed scalar outcome variable, $\varepsilon_t$ is an unobserved scalar shock, and $\rho$ and $\alpha$ are parameters. We are interested in the impulse response of $y_t$ with respect to $\varepsilon_t$ at horizon $h$, which has the formula $\theta_h \equiv \rho^h + \alpha \rho^{h-1}$ for $h \geq 1$ (and $\theta_0=1$).
(ref) shows that, in this DGP, there is a stark bias-variance trade-off between LPs and (short-lag) VARs. The figure shows histograms of the LP (blue) as well as VAR (red) estimators of the impulse response at horizon $h=2$ for $\rho=0.85$, $\alpha=0.1$, and with a sample size of $T=240$. Both estimators control for one lag of the data, with larger lag lengths to be considered below.\footnote{The VAR estimator uses an AR(1) specification. The LP estimator regresses $y_{t+h}$ on $y_t$, while controlling for $y_{t-1}$. Both estimators include an intercept and omit the bias corrections discussed in (ref).} We can see that the LP estimator is centered close to the true value of the impulse response, $\theta_h$. Intuitively, because LP just directly projects the future outcome of interest $y_{t+h}$ on $y_t$ (controlling for $y_{t-1}$), it is able to pick up on the full ARMA dynamics of the DGP (ref). The VAR estimator, in contrast, suffers from extrapolation bias: its impulse response estimate $\hat{\rho}^h$ is obtained by first estimating the autoregressive parameter $\hat{\rho}$ through a one-period-ahead forecast regression of $y_t$ on $y_{t-1}$, and then iterating forward $h$ steps using the parametric AR(1) model. The VAR estimator therefore is not directly informed by the sample autocovariances at horizons 2 and longer, causing it to miss out on the more intricate moving average dynamics present in the DGP. Yet, the figure also shows that the parametric extrapolation has a clear benefit: because the first-order autoregressive coefficient $\hat{\rho}$ is more precisely estimated than the longer-horizon autocovariances, the sampling distribution of the VAR impulse response estimator is less dispersed than the LP estimator. In summary, we see a trade-off: the sampling distribution of the LP estimator is well-centered but dispersed, while that of the VAR estimator is more tightly concentrated but centered incorrectly.
Panel (a) of (ref) exhibits how the nature of the bias-variance trade-off differs across DGPs and impulse response horizons. On impact, $h=0$, the two estimators are numerically identical (they both equal 1). However, at intermediate horizons the bias-variance trade-off is stark, for the reasons discussed earlier. The larger the moving average coefficient $\alpha$, the more misspecified is the one-lag VAR specification, and consequently the larger is the VAR bias. The VAR bias then decreases at long horizons, since in stationary DGPs the impulse responses must converge to zero as $h \to \infty$, a feature that is mechanically enforced by the VAR estimator $\hat{\rho}^h$ (but not by the LP estimator). However, the figure also reveals that the speed of convergence of the impulse responses to zero depends on the persistence parameter $\rho$, so it is not clear a priori at what horizons we should expect the impulse responses to be small---and thus when we can stop worrying about the VAR bias.
Finally, panel (b) of (ref) shows how changes in the estimation lag length $p$ shape the bias-variance trade-off, demonstrating that the bias of the VAR estimator at shorter horizons is reduced when the lag length is increased. In this panel, the LP and VAR estimators now control for $p$ lags of the data $y_{t-1},\dots,y_{t-p}$ rather than just one. When we control for $p=4$ lags, the VAR impulse response estimator is now nearly unbiased out to horizon $h=4$. This comes at a cost, however: the VAR variance increases substantially, to essentially equal that of LP at those horizons. When we control for even more lags, the VAR bias is now reduced at even longer horizons, but again at the cost of higher variance. These results hark back to our second lesson: an LP is essentially a VAR that controls for a lot of lags. Consistent with this interpretation, the bias and standard deviation of the VAR estimator depend dramatically on the choice of lag length, while this choice matters much less for the LP estimator.
In the next section we will establish that these simulation-based conclusions in fact hold up theoretically in a much wider class of multivariate models. Subsequent sections will then use large-scale, empirically calibrated simulation studies to demonstrate that these takeaways are practically relevant, in addition to being theoretically well-founded.
To gain a deeper understanding of the bias-variance trade-off, we review some of the main theoretical results from Montiel2024.
\paragraph{Univariate model.} We begin with the univariate ARMA$(1,1)$ model (ref). Montiel2024 show that the VAR and LP estimators of the true impulse response $\theta_h$ are both approximately normally distributed in large samples, though with differing bias and variance:
where $\overset{\cdot}{\sim}$ denotes “approximate distribution in large samples” (in the usual formal sense of asymptotic normality), $b_h(p)$ is a VAR bias term that depends on the estimation lag length $p$, and the large-sample variances satisfy $\tau_{h,\text{LP}}^2 \geq \tau_{h,\text{VAR}}^2(p)$; we will return later to the observation that the LP large-sample variance does not depend on the estimation lag length.
The formal result that yields (ref) requires the moving average coefficient $\alpha$ in the DGP (ref) to be relatively small. Thus, our analysis gives the VAR estimator the benefit of the doubt by viewing the true DGP as close to---but not exactly equal to---an AR(1), in a way that yields a non-trivial trade-off between model misspecification and statistical sampling error.\footnote{Formally, $\alpha$ needs to be proportional to the magnitude of the standard deviation of the estimators (i.e., $\alpha=\alpha_T \propto T^{-1/2}$). If we instead analyzed the asymptotic properties of fixed-lag VAR estimators under the DGP with fixed parameter $\alpha$ (as in Braun1993), the conclusions would be stark but empirically uninteresting: for any moving average coefficient $\alpha \neq 0$, the VAR estimator would be inconsistent for the true impulse response due to misspecification, so for large sample sizes the ratio of the bias to the standard deviation of the estimator diverges to infinity, yielding a trivial bias/variance trade-off in the limit. The “local-to-zero” modeling device of setting $\alpha \propto T^{-1/2}$ Schorfheide2005 should not be viewed as a literal description of reality, but rather as a technical device intended to tractably and accurately capture key finite-sample phenomena, similar to the econometric literatures on weak instruments or near-unit roots. Our later simulations will show that the lessons learned from these local-to-zero asymptotics are borne out in a large class of realistic DGPs.} The popularity and empirical success of autoregressive estimators in the forecasting literature suggest that carefully implemented autoregressive specifications are often not grossly misspecified. However, it seems unlikely that such simple parametric models can successfully capture every aspect of the real world. These observations make our setting a natural one to evaluate the properties of VAR and LP estimators.
We now study in greater detail the bias and variance expressions in (ref).
To summarize, in the simple ARMA$(1,1)$ model (ref) considered so far, LPs and VARs indeed lie on opposite ends of a bias-variance trade-off, with this trade-off vanishing as the VAR estimation lag length is increased, consistent with our earlier simulations.
\paragraph{Multivariate generalization.} Montiel2024 show that all the above-mentioned qualitative lessons go through in a much wider class of multivariate VARMA$(p_0,q_0)$ models, where $p_0$ is some finite true autoregressive order, while the moving average lag length $q_0$ could be infinite. As before, this set-up gives the benefit of the doubt to VAR practitioners by modeling the moving average coefficients as relatively “small.” Importantly, this class is general enough to be empirically relevant, for two main reasons. First, the class is consistent with typical linearized structural macroeconomic models.\footnote{Linearized DSGE models in macroeconomics almost always have a VARMA representation, though they typically do not have an exact finite-order VAR representation.} Second, it captures several important kinds of dynamic misspecification: even if the true DGP were a finite-order VAR, if the econometrician accidentally omits some lags or relevant control variables from the empirical specification, then those omissions will show up as moving average terms Granger1976.
It turns out that, even in this much larger class of models, the large-sample distributions (ref) continue to apply, and thus so does the bias-variance trade-off, as long as we consider LP and VAR estimators that control for at least $p \geq p_0$ lags of the data; if we fail to control for at least $p_0$ lags, either estimator could be badly biased in large samples. The way to interpret this result in practice is that it is important to control for those variables and lags that are strongly predictive of either the outcome variable of interest or the impulse variable; in the present theoretical setting, this amounts to controlling for $p_0$ lags of the data vector, since the moving average coefficients are small relative to the autoregressive coefficients. But once we control for all of these strong predictors, the LP estimator has zero asymptotic bias and constant variance as a function of $p$, even though further lags of the data do have some modest remaining predictive power. In other words, LP is relatively insensitive to the omission of moderately important control variables and lags (modeled through the moving average terms). By contrast, the VAR estimator has low variance yet generally nonzero bias even once we control for the most important predictors; intuitively, the issue is that the VAR does not project on the most important predictors horizon-by-horizon (like LP), but instead only does so on impact, and then iterates forward using its particular parametric structure.
There is of course a simple way to “bias-correct” the VAR estimator, as already discussed: use a very large estimation lag length. In particular, in the present model set-up, it can be shown that the large-sample bias $b_h(p)$ of the VAR estimator at horizon $h$ is zero when the estimation lag length $p$ exceeds $p_0+h$. However, when the lag length is chosen this large, then the variance is necessarily inflated to that of LP---in other words, the VAR bias is zero only because long-lag VAR and LP estimators are equivalent in large samples. In practice, we do not know exactly how large the true autoregressive lag length $p_0$ is, so to be certain that we included enough lags to eliminate meaningful biases, we would need to compare the VAR results with comparable LP results and verify that they are close to each other.
\paragraph{No free lunch for VARs.} As the final step in our theoretical analysis we show that even relatively minor amounts of VAR misspecification can yield economically large biases. This discussion will rationalize much of what we find in our later simulations.
Our approach here is to ask how large the VAR bias $b_h(p)$ can possibly be, as a function of the magnitude of the dynamic misspecification. Montiel2024 show that the ratio of the absolute bias to the standard deviation of the VAR estimator is (asymptotically) bounded by
where $T$ is the sample size and $\mathcal{M}$ is a measure of the overall magnitude of the misspecification of the VAR, namely the fraction of the variance of the moving average residual that is explained by the lagged shocks in the VARMA model. Equation (ref) states that the VAR's bias is bounded above by a simple function of just two numbers: the potential amount of misspecification, and the ratio of LP-to-VAR standard errors. In a well-specified finite-order VAR model there are no lagged shocks that enter into the residual, and so in this case there is no misspecification, i.e., $\mathcal{M}=0$, and the bias is zero. In the simple ARMA$(1,1)$ model (ref), $\mathcal{M} = \alpha^2$. For a concrete example, suppose that $\mathcal{M}=0.01$, so that the lagged shocks account for a mere 1% of the variance of the residual, and furthermore assume that $T=100$ and the VAR standard error equals 0.4 times the LP standard error, $\tau_{h,\text{VAR}}(p)/\tau_{h,\text{LP}}=0.4$, roughly consistent with the median standard error ratio at longer horizons in (ref). Then the upper bound (ref) equals $\sqrt{5.25} \approx 2.29$, so that the bias of the VAR estimator can be more than twice as large as its standard error, despite the minor degree of misspecification in this example. Biases of this magnitude are obviously very worrisome when drawing statistical inferences from conventional VAR regression output. Though the formula (ref) represents an upper bound on the extent of the bias, it can be shown that there always exists a residual moving average process (with small coefficients satisfying the imposed misspecification magnitude $\mathcal{M}$) that achieves this bound. Moreover, the particular form of this “least favorable” residual process does not appear to be unreasonable ex ante based on economic theory, and conventional statistical tests have low power to detect it ex post in the data.
To summarize, the worst-case bias formula (ref) implies that the only way for VAR practitioners to guarantee that the bias is negligible is to include so many lags $p$ in the specification that the standard error becomes equal to that of LP. Conversely, for any VAR specification that delivers material efficiency gains relative to LP, there is reason to worry that the VAR bias could be large relative to the standard error.
The absence of a free lunch has particular bite at long horizons. If the DGP is stationary, impulse responses must be close to zero at very long horizons, but it is rare to possess precise prior knowledge about exactly at what rate the impulse responses decay. Unless the lag length is very large, the VAR estimator simply extrapolates the long-run responses based on the short-run empirical autocovariances, eventually enforcing an exponential rate of decay of the impulse responses in stationary environments. The VAR-estimated impulse responses at long horizons will therefore tend to have very small standard errors relative to LP (which, as we have discussed, does not enforce exponential decay). But this, then, is precisely when the potential for VAR bias will also be high: if the estimated rate of decay is inaccurate, this can dramatically affect the magnitude of estimated long-horizon responses, and their bias can be several times larger than their standard errors, as indicated by the bias bound (ref). In turn, this can lead to large errors in estimating such key features as the half-life or quarter-life of the impulse response function, or its cumulative value (as would be relevant for estimating hysteresis effects, say). LPs remain approximately unbiased at long horizons, and their large standard errors accurately reflect the fundamental issue that any finite data set only has limited information about what happens in the long run, in the absence of prior information. In particular, LPs simply do not allow the estimation of impulse responses at ultra-long horizons $h$ beyond the observed sampled size $T$ (or rather, the effective sample size $T-p$); VARs instead produce such estimates by pure extrapolation, with potential for severe bias. The fact that an LP at horizon $h$ uses only $T-p-h$ data points (as opposed to $T-p$ for a VAR) is not a drawback of the procedure; it is a necessary consequence of the desire to avoid the extrapolation that the VAR estimator relies on.
In addition to VARs, the literature has proposed several variants of and alternatives to local projection estimators intended to increase efficiency, including for example Generalized Least Squares (GLS) variants of LP, Autoregressive Distributed Lag (ADL) regressions, and parametrized LPs. This goal, however, faces one key theoretical obstacle: LPs are semiparametrically efficient and therefore cannot be improved upon without sacrificing robustness or imposing substantive further restrictions. In this sense, the bias-variance trade-off between LPs and VARs extends also to comparisons with other estimation procedures.
The plain OLS LP estimator is semiparametrically efficient in linear models. This was initially conjectured by Plagborg2021 based on the equivalence of the LP and VAR estimands (recall (ref)), with Xu2023 offering the formal proof. Specifically, in a general dynamic linear model with unrestricted lag structure, the efficient quasi-maximum likelihood estimator is a VAR estimator with an asymptotically increasing lag length.\footnote{More precisely, a VAR estimator with asymptotically increasing lag length reaches the semiparametric efficiency bound associated with the conditional moment restrictions implied by a VAR($\infty$) model with shocks that are martingale difference sequences. Xu2023 requires the model to be stationary with homoskedastic shocks, but it seems likely that these assumptions can be relaxed.} The OLS LP estimator (also with an increasing lag length) is asymptotically equivalent with this estimator. Thus, in this setting, any well-behaved impulse response estimator must either have weakly higher variance than plain LP in large samples, or it must be inconsistent in some DGPs where plain LP is consistent. In other words, any efficiency gains can only come at the expense of imposing substantive restrictions on the transmission mechanism or the identification of the shock. Of course, such additional restrictions may very well be plausible in some applications, but they should always be acknowledged.
This discussion has immediate implications for several of the recently proposed alternatives to the plain OLS LP estimator. First, GLS variants of LP cannot uniformly improve on Jorda2005's (Jorda2005) original OLS version. In particular, Xu2023 shows that the variance gains of the GLS approach of Lusompa2023 only obtain when the true DGP is a VAR with short lag length.\footnote{In fact, some implementations of GLS LP Breitung2023 are equivalent in large samples with a VAR estimator and therefore subject to the same bias issues discussed earlier, unless the estimation lag length is large.} Second, ADL regressions featuring the shock $x_t$---effectively an LP that controls additionally for future values $x_{t+1},\dots,x_{t+h}$ of the shock $x_t$---can provide efficiency gains, as they exploit the information that the shock is directly observed Choi2019,Baek2022. Intuitively, because the shock is serially uncorrelated, future shocks are valid controls and help increase efficiency, as discussed further in (ref). However, a crucial caveat is that, if the serially uncorrelated shock is not in fact directly observed but needs to be pre-estimated, then the ADL standard errors must be adjusted to reflect the generated regressors. Not only is this complicated to do in practice, it may also negate any efficiency gains relative to plain LP. Indeed, no ADL procedure with pre-estimated shocks can outperform plain LP for general shock identification procedures. Third, parametrized LPs impose shape restrictions on the impulse response functions Barnichon2018. Such restrictions will increase precision but in general also lead to bias.
In (ref) we have made some broad conceptual points on identification as well as on the bias-variance trade-off between LPs, VARs, and related techniques. In the remainder of the paper we give recommendations on how to navigate this trade-off in empirical practice. (ref) begins with specification choices and implementation details for a range of LP and VAR estimators. (ref) then provides a quantitative, empirically calibrated assessment of the bias-variance properties of these estimators, while (ref) discusses implications for inference. Finally (ref) concludes with our recommendations for applied practice.
Here we discuss three of the main practical considerations involved in estimating impulse responses through either LPs or VARs: selection of control variables and lags, bias correction, and whether to shrink the estimates towards a priori plausible values.
In applications of LP or VAR estimators, a ubiquitous question is which variables and how many lags we should control for. There are four main factors to consider.
The remaining two factors are specific to LP estimation. LPs that control for lagged data are referred to as lag-augmented.
\paragraph{Practical takeaways.} VAR practitioners tend to select control variables and lags with an eye towards their identifying assumptions as well as predictive power, consistent with the first two factors reviewed above. The precise lag length is then often selected using the AIC or by using conventional fixed lengths such as 4 for quarterly data and 12 for monthly data, as recommended by Kilian2017. We will review the performance of these VAR specification choices in our later simulations.
For LPs, there is instead relatively little consensus on how to select controls and lags, and so we suggest the following procedure, guided by the above considerations. In general, researchers should always control for (i) variables that are central to their identification scheme, (ii) the outcome and impulse variables, and (iii) any additional variables that strongly predict either the outcome or the impulse variables (or both), as suggested by past experience or economic theory. Among this set, the precise choice of lags and controls can be guided by the AIC or by other model selection criteria, proceeding as follows: first, estimate an auxiliary VAR that includes the outcome and impulse variables and other potential controls; second, use the AIC to select the set of controls and lag length as in conventional VAR practice Kilian2017; third, use the selected variables and lags as controls in subsequent LPs.\footnote{Alternatively, we could do two separate model selection exercises for a single-equation one-step-ahead forecast of the outcome and a single-equation one-step-ahead forecast of the impulse variable. We should then use the union of the two sets of selected control variables and lags in subsequent LPs. } There is no internal inconsistency in using an auxiliary VAR for selecting an LP specification: though the data-driven model selection is inevitably subject to small errors that threaten the validity of VAR-based impulse response inference, LP-based inference is robust to such minor model misspecification, as discussed earlier.
As a final comment on applied practice, while some LP analysts run regressions with long differences $y_{t+h}-y_{t-1}$ of the outcome variable on the left-hand side, this is actually redundant when using lagged controls. By standard OLS algebra, this outcome transformation has no impact whatsoever on the estimated LP coefficient once we control for at least one lag of the outcome, which we have argued is advisable anyway.
The theoretical analysis reviewed in (ref) focused on one key source of bias: dynamic misspecification. In the small data sets that are typical in applied macroeconomics, however, LPs and VARs are subject to another important source of bias: persistence of the data. In particular, impulse response estimates are typically biased towards displaying less persistent effects than the actual truth. For VARs, Kilian2017 recommend applying the Pope1990 bias correction to the estimated VAR coefficients to partially remove this source of bias. For LPs, Herbst2024 find that a qualitatively similar bias is present and propose a simple bias correction. Though Li2024 find that the LP bias correction increases the variance of the estimator without entirely removing bias, it is advisable to apply the correction since the entire justification for using LP over VAR prioritizes bias over variance. We demonstrate the advantages of the bias correction in practice in simulations below.
Piger2025 propose to ameliorate the small-sample LP bias by running the regression on first-differenced data and then cumulating the impulse responses to recover the level response. Whether or not to difference the data is a familiar conundrum from VAR analysis, with most textbooks preferring the levels specification (Hamilton1994; Kilian2017). The worry is that differencing the data can reduce the predictive power of the lagged controls, thus causing potentially large efficiency losses---even in large samples. In our simulations in (ref), bootstrap confidence intervals based on LPs that are estimated in levels and with the Herbst2024 correction robustly deliver reliable inference.\footnote{Note that, in contrast to differencing, the Herbst2024 bias correction becomes negligible (and as a result does not impair efficiency) when the sample size is large.} Hence, our tentative view is that the potential efficiency costs of the Piger2025 proposal may well outweigh the marginal benefits, though more research in this area would be welcome.
We also stress that none of the bias correction procedures discussed in this subsection deal with the bias caused by dynamic misspecification. Hence, all of our earlier points about the trade-off between LPs and VARs continue to apply to bias-corrected estimators.
Disappointed by the often jagged-looking impulse response functions that are produced by least-squares LPs and VARs, some researchers have turned to variants of these methods that employ shrinkage---i.e., nudging the least-squares estimate in a direction that is regarded as a priori plausible. The penalized LP estimator of Barnichon2019 smooths out the estimated impulse responses across horizons by adding a term to the OLS objective function that penalizes deviations from a quadratic impulse response function, with the overall degree of shrinkage determined by cross-validation. Similarly, Bayesian VARs shrink unrestricted estimates towards simpler dynamics, such as independent random walks or white noise processes Doan1984,Todd1984,Litterman1986,Giannone2015. Finally, model averaging or selection techniques use data-dependent rules to partially or fully shade the LP estimator toward a short-lag VAR estimate Ferreira2023,Gonzalez2025.
While the attraction of these shrinkage estimators is that they reduce the variance of the jagged least-squares estimators, they will do so explicitly by introducing additional bias. Whether the added bias is actually worthwhile depends both on the DGP and on the researcher's objective function, as we discuss in the next two sections.
This section uses empirically calibrated simulations to quantify the bias-variance trade-off between various implementations of LPs and VARs. The simulation evidence complements our theoretical treatment, revealing what shape the bias-variance trade-off is likely to take in practice, and so laying the groundwork for our later practical recommendations.
Prior simulation studies comparing the performance of LP and VAR methods include Kilian2011, Bruns2021, and Li2024; we here build on and extend the latter study. Just like those authors, we consider a large menu of possible DGPs, all estimated to mimic as closely as possible the properties of the universe of U.S. macroeconomic data. We provide a high-level overview of the DGPs here, with implementation details relegated to the original work of Li2024 as well as to (ref).
\paragraph{Data generating processes.} Our simulations are based on a dynamic factor model (DFM) fitted to a large number of standard U.S. macroeconomic time series. We use the data set of Stock2016, which consists of 207 quarterly U.S. time series spanning a range of variable categories, from quantities to prices. We then fit two separate DFMs to this data set: a non-stationary variant that allows for cointegrating relationships among the factors, as well as a stationary one for which all data series have been pre-transformed to ensure stationarity. Differently from Li2024, we allow for conditional heteroskedasticity of the shocks in the equations for both the factors as well as the idiosyncratic disturbances. The heteroskedasticity is modeled through ARCH processes with empirically estimated parameters. Overall, the resulting DFMs imply complex and varied dynamics for macroeconomic time series of the sort encountered in applied practice.
Given these DFMs, we proceed to construct a diverse array of lower-dimensional DGPs by considering hundreds of different subsets of time series, with variable selection closely emulating applied practice. Specifically, each DGP consists of a set of five observable time series, selected at random from those of the DFM's series that are most commonly used in applied practice. We then consider a researcher who observes data of those particular selected series, and identifies monetary and fiscal policy shocks through either a recursive ordering of these observable time series, or by also additionally measuring the policy shock, as in our discussion of identification in (ref). For the monetary policy DGPs we restrict the vector of observables to always contain the federal funds rate, while for fiscal policy DGPs we always include government spending; for recursive identification these two policy variables are ordered last and first, respectively. The researcher is then interested in the response of one of the other four variables to the identified shock. She estimates this response using several different estimation strategies, to be discussed in detail below. Given that the true DFM is known to us, we can through simulations compute estimator biases, variances, and mean squared errors (for this section) as well as confidence interval properties (for the uncertainty assessments in the next section).
We stress that the DFMs and the resulting DGPs that we construct do not admit finite-order VAR representations, yielding a non-trivial bias-variance trade-off between LP and VAR estimators. The latent nature of the macroeconomic factors and the idiosyncratic disturbances induce moving average dynamics in the processes for the observed time series, consistent with our discussion of dynamic misspecification in (ref).
Our overall objective is to provide an empirically grounded assessment of which estimators are likely to perform well on average in settings typically encountered in applied work, and thus can serve as attractive default procedures. While we cannot claim that our empirically calibrated DFMs capture all aspects of the real world, we contend that a minimum condition for an econometric procedure to be useful for applied work is that it should at least perform well on the kinds of DGPs that we consider. The empirically grounded DGPs will furthermore showcase the quantitative bite of the general theoretical considerations reviewed in (ref).
\paragraph{Estimators.} We consider several variants of the LP and VAR estimators introduced in (ref)---variants that differ in the selection of the lag length $p$, the choice of control variables $w_t$, and whether we apply the shrinkage techniques discussed in (ref). All estimators include an intercept. We only provide a brief overview here, with further technical details in (ref).
In this section we present simulation results for 200 stationary and 200 non-stationary DGPs, for observed shock identification and averaging across monetary and fiscal shocks, with 100 DGPs for each. Results separately by the type of shock and for recursive identification are broadly similar, and relegated to (ref). We approximate population biases and variances by averaging across 1,000 Monte Carlo simulations per DGP.
\paragraph{Bias-variance trade-off.} (ref) reveals that, in applied practice, researchers do indeed face a stark bias-variance trade-off between LPs, VARs, and intermediate shrinkage techniques, consistent with our earlier theoretical discussion. The left panels of the figure show the median absolute bias (across all the monetary and fiscal policy shock DGPs) for the stationary (top) as well as the non-stationary (bottom) encompassing DFM, while the right panels show the median standard deviation, in both cases normalized by the overall scale of the true impulse response. Red lines correspond to VAR estimators, and blue lines indicate LP estimators, with the various line styles corresponding to several different options for the selection of lag length, controls, and shrinkage.\footnote{We only report results for bias-corrected LPs and VARs, consistent with our discussion in (ref). Bias correction barely matters in the stationary DGPs, but LP biases would be materially larger in non-stationary DGPs in the absence of bias correction, particularly at medium and long horizons.} While all least-squares LP estimators have relatively low bias across horizons, the VAR estimators tend to suffer from substantial bias at intermediate horizons, as well as at long horizons in the more persistent DGPs. Finally, and as expected, shrinkage---here in the form of either penalized LPs or Bayesian VARs---reduces variance at the cost of higher bias.
While LPs are relatively insensitive to dynamic specification, the bias-variance properties of VAR estimators are shaped decisively by the estimation lag length, the horizon of interest, and the choice of controls. The AIC typically selects a very short lag length in our DGPs (the mean lag length selected equals $p = 1.88$), and so we see a sharp bias-variance trade-off already at short impulse horizons. Manually increasing the VAR lag length to $p = 4$ (dark red, dashed) aligns LPs and VARs up to horizon $h = 4$, again consistent with the theory reviewed earlier. The performance of LPs is, on the other hand, virtually unaffected by the precise number of estimation lags (in solid blue vs.\ dashed dark blue).\footnote{(ref) shows that the number of lags plays a more important role under recursive identification, since then the lags also matter for spanning the shock of interest (recall (ref)).} Furthermore, in the stationary DGPs, both the VAR bias and standard deviation are small at long horizons, since all impulse responses converge to zero relatively fast. In practice, however, the speed of this convergence is uncertain; in particular, in the non-stationary DGPs (displayed in the bottom panel), the bias-variance trade-off remains stark even at much longer horizons. Finally, comparing the solid and light-dashed lines, we see that the effect of including either few or many controls is negligible for LPs, but important for VARs.
\paragraph{Mean squared error.} We now further assess the quantitative nature of the bias-variance trade-off using the familiar mean squared error (MSE) criterion. For a given estimator $\hat{\theta}_h$ of the true impulse response $\theta_h$, the MSE is defined as
In words, MSE weights equally an estimator's (squared) bias and variance. (ref) plots the MSE of the estimators, again separately for stationary (left) and non-stationary (right) DGPs. We see that MSE is typically minimized by a VAR method, either Bayesian or least-squares. Thus, while we know theoretically that there exist some DGPs close to finite-order VARs for which LPs outperform VARs in terms of MSE, VARs are favored in the typical DGP.\footnote{According to the worst-case bias result (ref) in (ref), even a minor amount of misspecification $\mathcal{M} \geq 1/T$ could suffice for VAR MSE to exceed LP MSE Montiel2024. In the DGPs that we consider, the magnitude of misspecification is indeed quite material (e.g., for the stationary DFM, the average $\sqrt{T \times \mathcal{M}}$ across DGPs is $2.64$ for $p = 4$ VAR lags). However, it turns out that the type of misspecification is not close to the worst case on average, causing the VAR MSE to typically be smaller. We thank Isaiah Andrews for raising this point. } This finding is a consequence of the fact that, across our DGPs, the bias reduction of LP typically comes at more than a one-to-one cost in terms of the standard deviation.
This lesson is the main takeaway of Li2024, and it is also consistent with the conclusions of Marcellino2006 in a forecasting context. Some researchers may, however, be more concerned with minimizing bias than minimizing variance, in which case plain LPs could still be preferable over VARs (and the shrinkage estimators). In the next section we will argue that if the goal is not merely to construct a point estimate but also to accurately convey the uncertainty surrounding the estimate, then we are endogenously forced to heavily prioritize bias.
Applied macroeconomists not only report point estimates of dynamic causal effects, they also want to quantify the statistical uncertainty of those estimates. This is an important task, since the standard errors are often of roughly the same magnitude as the estimates themselves. The accuracy of uncertainty assessments is conventionally evaluated by the coverage of the implied confidence interval: the probability that the reported interval covers the true impulse response should be at least 90% (say), at every horizon and regardless of the shape of the true impulse response function.
This section argues that LPs---or VARs with very long lag lengths---are the only known procedures that can robustly achieve satisfactory coverage of confidence intervals in practice. The intuition is that even quite small amounts of bias (relative to the standard error) can cause severe coverage distortions, and as we have seen, only LPs (or equivalently VARs with very long lags) robustly achieve low bias across a range of empirically relevant DGPs. In other words, while short-lag VARs deliver point estimates that are accurate most of the time, only LPs deliver confidence intervals that are valid (almost) all the time.
We begin by reviewing methods for computing standard errors for LP estimation of impulse responses. We will not review VAR-based inference methods, since these are already covered in textbooks such as Kilian2017.
For LP impulse responses, heteroskedasticity-robust standard errors suffice for accurate uncertainty assessments. In the baseline LP regression (ref), the residual $\xi_{h,t}$ is typically serially correlated because it is a multi-step forecast error, which would suggest the use of Heteroskedasticity and Autocorrelation Consistent (HAC) standard errors, such as Newey-West. Fortunately, however, MontielOlea2021 show that HAC corrections are unnecessary under a weak assumption on the shocks, and conventional heteroskedasticity-robust standard errors for OLS suffice. The reason is that, while the LP regression residual $\xi_{h,t}$ is indeed serially correlated, what matters for the distribution of the LP estimator is the product of the residual and the residualized shock $\tilde{x}_t$ defined in (ref). This product is serially uncorrelated (though typically heteroskedastic) under the natural assumption that the shocks are unpredictable (i.e., conditionally mean independent) from their past and future values. Even if this assumption is slightly violated, it is still likely that the well-documented practical challenges of HAC estimation Lazarus2018,Herbst2024 will make the conventional heteroskedasticity-robust standard errors a better choice for applied work.
Given a standard error $\hat{\tau}_{h,\text{LP}}$ for the LP impulse response estimate $\hat{\theta}_h^\text{LP}$, a level-$(1-a)$ confidence interval can be obtained with the usual formula:
where $z_{1-a/2}$ is the $(1-a/2)$ quantile of the standard normal distribution (e.g., $z_{1-a/2} \approx 1.64$ for a $1-a=90\%$ confidence interval). MontielOlea2021 find that a bootstrap confidence interval can further improve coverage in small samples. However, the bootstrap they consider is only valid for reduced-form impulse responses, not structural impulse responses. Our simulation study in (ref) suggests that confidence intervals obtained by bootstrapping LP estimates using the residual block bootstrap of Brueggemann2016 have accurate coverage for structural impulse responses.
The next sections compare---first theoretically, then through simulations---the quality of uncertainty assessments based on either LP or VAR inference.
In this section we establish theoretically that conventional VAR confidence intervals can exhibit severe coverage distortions even under relatively small amounts of dynamic misspecification, while LP intervals are instead much more robust. Our analysis here builds on and extends our review of the bias-variance trade-off in (ref).
\paragraph{How bias affects coverage.} The approximate VAR sampling distribution in (ref) together with a straightforward calculation reveals that the coverage probability of the conventional VAR confidence interval is a decreasing function of the bias/standard-error ratio $|b_h(p)|/\tau_{h,\text{VAR}}(p)$; specifically, it equals
The LP confidence interval, on the other hand, robustly achieves the nominal coverage level, by virtue of LP's zero (asymptotic) bias.
(ref) plots the coverage probability of the VAR confidence interval as a function of the (scaled) VAR bias and for a range of target coverage levels. We see that even a moderate ratio of bias to standard error yields large coverage distortions. In (ref) we gave a numerical example in which a small amount of VAR misspecification caused the bias to be around 2.29 times the standard error; this ratio would cause a putative 90% confidence interval to cover the true impulse response with probability less than 30%! The right panel of (ref) indeed suggests that the bias/standard-error ratio of VARs exceeds 2 at long horizons in many applications.\footnote{The scaled difference $(\hat{\theta}_h^\text{LP}-\hat{\theta}_h^\text{VAR})/\tau_{h,\text{VAR}}(p)$ of LP and VAR estimates is an unbiased estimate of $b_h(p)/\tau_{h,\text{VAR}}(p)$. An important caveat is that (ref) reports the absolute value of the scaled difference, which is an overestimate of the absolute scaled bias.} (ref) thus implies that a researcher who is interested in guaranteeing that the coverage of a reported confidence interval is not too low relative to the target coverage must necessarily prioritize bias over variance.
Because bias is so costly for coverage, LP confidence intervals---or equivalently VAR confidence intervals with a very large estimation lag length---are necessarily preferred over short-lag VARs when it comes to uncertainty assessments. Recall that, in the face of minor amounts of misspecification, we can only guarantee a low bias/standard-error ratio for VARs when we use very many lags (usually many more lags than indicated by conventional model selection or evaluation procedures). In practice, it thus only appears to be reasonable to trust VAR confidence intervals if they are approximately as wide as corresponding LP intervals.
We furthermore stress that one does not need to believe that the (short-lag) VAR bias is large in most applications in order to prefer LP over VAR confidence intervals. In conventional statistical and econometric practice, we seek to control coverage uniformly across a wide range of empirically relevant DGPs, and not merely for the “typical” DGP. In other words, our uncertainty assessment should reliably indicate that uncertainty is high whenever that is in fact the case. Hence, the mere fact that VARs can sometimes be badly biased, as shown theoretically in (ref) and then practically in (ref), would militate against short-lag VAR confidence intervals. Our simulations in (ref) will illustrate these theoretical observations.
\paragraph{MSE vs.\ coverage.} To reiterate, the optimal procedure when it comes to producing confidence intervals with correct coverage is not the same as the MSE-minimizing procedure. Based on the approximate distributions (ref) of LP and VAR, a simple calculation shows that VARs are preferred to LP on MSE grounds if and only if
Once again making reference to the right panel of (ref), and focusing on long horizons, we may expect the left-hand side of the above inequality to be close to 2 in many applications. The left panel of the same figure shows that the median value of $\tau_{h,\text{LP}}/\tau_{h,\text{VAR}}(p)$ can be close to 0.4 in practice, corresponding to a value of $\sqrt{5.25} \approx 2.29$ for the right-hand side of the above inequality. Thus, in “typical” applications, VAR point estimators are preferable from the MSE perspective, yet the associated VAR confidence interval can have poor coverage.
\paragraph{Shrinkage, model selection, and Bayesian inference.} If we seek confidence intervals that have correct coverage regardless of what the true impulse response function looks like, then it is likely impossible to improve much upon the plain LP confidence interval, at least in large samples. One might hope that shrinkage, model averaging, or model selection techniques could be used to develop shorter confidence intervals that “adapt” to the nature of the DGP; that is, decide in a data-dependent manner to what extent we should report wide LP intervals or some narrower interval that imposes smoothness or parametric structure. However, when applied to the theoretical framework of (ref), mathematical impossibility results from microeconometrics Pratt1961,Armstrong2018 strongly suggest that the potential for such adaptation is very modest. In other words, we can only get away with reporting (meaningfully) narrower confidence intervals than LP if we give up on the usual notion of coverage.\footnote{Freyaldenhoven2024 propose one such way to change the researcher objective.} Indeed, to our knowledge, none of the existing confidence intervals based on shrinkage procedures for LPs or VARs have frequentist coverage guarantees comparable to that of the plain LP interval, and we will see below that some of them indeed have poor coverage in simulations.
A Bayesian who is completely confident in their VAR model and prior can of course ignore LP evidence; however, any Bayesian who is worried about slight model misspecification is forced to use LP as a diagnostic in the same way as a frequentist. This is because, in large samples, the Bayesian VAR posterior distribution behaves just like the sampling distribution of a frequentist VAR estimator, and we have already seen that the latter is highly sensitive to small amounts of dynamic misspecification.
While the previous subsection discussed the fragility of VAR inference under mild misspecification, LP confidence intervals can in fact also have more accurate coverage than conventional VAR confidence intervals even if the VAR model is correctly specified. We review the intuition here and refer to MontielOlea2021 for technical details.
MontielOlea2021 show that, assuming a correctly specified VAR model, LP confidence intervals control coverage more robustly at long horizons and across a wider range of persistence of the underlying DGP than VAR intervals. Intuitively, at long horizons $h$, the VAR impulse response is a highly nonlinear transformation of the estimated parameters (e.g., recall the exponential formula $\hat{\rho}^h$ in the AR(1) case). This spells trouble for traditional VAR inference procedures because they ultimately rely on a linearization argument for asymptotic validity. By contrast, LP simply relies on linear regressions. Moreover, in DGPs with (near-)unit roots, it is well known that autoregressive coefficient estimates can have non-normal distributions, which dramatically complicates the calculation of appropriate critical values for impulse responses. LP, however, is equivalent to a projection of the outcome on the residualized shock $\tilde{x}_t$ in (ref), and the latter is robustly stationary (provided we employ lag augmentation as recommended in (ref)), so that the LP coefficient will typically still have a normal distribution even if the data has stochastic trends.
We now complement the theoretical insights of the previous two subsections with simulation results. Our simulations are again based on the large menu of DGPs described in (ref).
\paragraph{Confidence intervals.} We report the coverage of confidence intervals constructed from several variants of LP and VAR estimators, with a target confidence level of 90%. Further implementation details are provided in (ref).
\paragraph{Results.} As in (ref) we present simulation results for 200 stationary and 200 non-stationary DGPs, for observed shock identification and averaging across monetary and fiscal policy shocks, with 100 DGPs for each. Coverage results separately by type of shock and for recursive identification are similar, and relegated to (ref). Coverage is approximated by averaging over 1,000 Monte Carlo simulations per DGP.
As expected from the theoretical discussions in (ref), LPs tend to provide robust uncertainty assessments, while VARs and shrinkage techniques do not. To establish this, (ref) shows the fraction of DGPs at which any given inference method delivers confidence intervals with coverage probability above 80% (and thus close to the target level of 90%), with VARs in red and LPs in blue.\footnote{(ref) furthermore reports average coverage and interval width.} Both analytical and bootstrap LP confidence intervals attain accurate coverage levels for a large fraction of DGPs at all horizons, with the important exception that the bootstrap confidence interval is the only one that works well for non-stationary DGPs at long horizons. By contrast, for the clear majority of DGPs that we consider here, VARs with AIC-selected lag length have a coverage probability below 80% at all horizons $h>2$. The coverage of longer-lag, bootstrapped VARs is somewhat better, but we see that the coverage of even the best-performing VAR confidence interval is substantially worse than that of the LP bootstrap confidence intervals. The same conclusions extend to shrinkage methods: both Bayesian VARs and penalized LPs undercover severely. That is, even though most of our DGPs feature quite smooth impulse response functions, the shrinkage procedures nevertheless suffer from non-negligible bias, as documented in (ref), which has an outsize effect on coverage.
To link back to the discussion in (ref), recall that Li2024 found in their simulations that a preference for LP over VAR requires an overwhelming concern for bias over variance. We have seen that a desire for robust assessments of statistical uncertainty endogenously creates such a concern.
Based on the 12 lessons we have drawn from the available theory and simulation evidence on the econometric properties of LPs and VARs, we make the following summary recommendations for applied researchers who seek to perform inference on the dynamic causal effects of macroeconomic shocks to policies or fundamentals.
\phantomsection \addcontentsline{toc}{section}{References}