EconBase
← Back to paper

Inference for Local Projections

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

77,235 characters · 18 sections · 61 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Inference for Local Projections

\makeatletter \makeatother

\thispagestyle{empty}

abstract\singlespacing Inference for impulse responses estimated with local projections presents interesting challenges and opportunities. Analysts typically want to assess the precision of individual estimates, explore the dynamic evolution of the response over particular regions, and generally determine whether the impulse generates a response that is any different from the null of no effect. Each of these goals requires a different approach to inference. In this article, we provide an overview of results that have appeared in the literature in the past 20 years along with some new procedures that we introduce here. JEL classification codes: C11, C12, C22, C32, C44, E17. Keywords: local projections, impulse response, instrumental variables, confidence bands, simultaneous bands, significance bands, wild block bootstrap.

\allowbreak

\quad

\pagenumbering{arabic}

Introduction

Impulse responses are often used to characterize dynamic systems. When the analyst wants to remain agnostic about the data generating process (DGP), impulse responses are often estimated using the method of local projections proposed by Jorda2005 or LPs for short. LPs consist of projecting future outcome variables on current and past information on the intervention (or impulse), the outcome and other exogenous variables. In other words, they are simple regressions with a particular dynamic error structure.

This article discusses how to conduct inference for LPs with an emphasis on general principles. Relative to vector autoregressions (VARs), LPs generally have lower bias but higher variance though in infinite samples, when the lag structure grows with the sample, both methods generate the same impulse response and similar parameter estimation uncertainty PMWolf2021, Xu2023. We organize our discussion around three main topics: (1) point-wise inference; (2) simultaneous inference; and (3) significance.

Point-wise inference, by far the most popular, refers to the uncertainty with which individual parameters of the impulse response are estimated. It is presented graphically by means of error bands. Simultaneous inference refers to the uncertainty about sets of impulse response coefficients. It can be presented graphically as error bounds that accommodate families of hypotheses. Because these bounds are generally wider than error bands, practitioners do not generally show them. Both point-wise and simultaneous inference are based on the Wald principle.

In this paper, we introduce a new concept: significance bands. These bands refer to the null hypothesis that the impulse does not generate a response, or in the parlance of experiments, that the policy intervention has no effect. The Lagrange-Multiplier (LM) principle considerably simplifies the construction of significance bands, which are a useful complement to the practice of presenting error bands since many researchers are often concerned about the absence of a response to an impulse.

We set the stage with a brief introduction to LPs to highlight where the properties of the estimated response come from and how they affect inference. Next, we discuss point-wise inference. We highlight several recent developments in the literature, including feasible generalized least-squares methods (LP-FGLS), lag-augmentation (LP-LA), and with a brief mention of Bayesian inference (BLP).

Thinking about simultaneous inference requires system estimation of LPs. Thus we begin by framing such system estimation within the class of generalized method of moments (GMM) estimators. This allows us to cover instrumental variable estimation. With this scaffolding in place, we review two important papers in this arena, Jorda2009 and MontielOleaPM2019.

The review of the literature up to this stage sets us up to introduce the novel concept of significance bands. We will show that the LM principle provides a much simpler approach to think about the absence of a response and to present the information in a convenient graphical way. In addition, we also introduce a straightforward wild block-bootstrap procedure that is simple to implement.

By now, there is a sprawling literature that extends LPs into a variety of new areas. Though we cannot cover all these new developments, we will briefly touch on methods to smooth LPs in the context of producing more efficient responses. We will also discuss applications with panel data and the inferential challenges that these data introduce before concluding.

Local projections: an introduction

There are many situations in the study of dynamic systems where the analyst is interested in the following statistic:

align*[align* omitted — 180 chars of source]

where $s_t$ is the variable that generates the impulse from $s_t = s_0$ to $s_t = s_0 + \delta$ for some initial value $s_0$. Denote $y_t$ as the outcome variable of interest, whose response to an impulse in $s_t$ we want to characterize. Finally, the vector of variables $\bm x_t$ contains lags of the outcome, lags of the impulse, and lags of any other exogenous or predetermined variables, including the constant and deterministic time trends.

We will refer to $s_t$ interchangeably as the treatment, intervention, or shock. At this point, we assume that the shift from $s_0$ to $s_0 + \delta$ is exogenously determined. In practice, identification assumptions and methods will be required. However, since the emphasis here is on inference, we will often prefer to keep things simple. The conditional mean function, $\mathbb E (\cdot | \cdot)$ can in principle be nonlinear, a case considered by AngristK2011 and \citet*{AJK2016}. We will mostly focus on linear settings, but it is important to maintain the distinction for now.

We use the notation $\mathcal{R}_{s \rightarrow y|\bm x}(h, s_0, \delta)$ instead of $\mathcal{R}_{sy|\bm x}(h, s_0, \delta)$ when additional identifying assumptions, such as exogeneity of $s_t$ justify a causal interpretation of the difference in conditional means. The notation $\mathcal{R}_{s \rightarrow y|\bm x}(h, s_0, \delta)$ is meant to convey the direction of the intervention $s$ onto the outcome $y$, conditional on $\bm x$. When it is clear, we will simplify the notation to $\mathcal{R}_{s y}(h)$ to denote the response of $y$ to an impulse in $s$, $h$ periods in the future.

For example, when $s_t \in \{0,1\}$ is randomly assigned, the conditional expectation $ E(y_{t+h}|s_t , \bm x_t)$ does not depend on $\bm x_t$ and only takes two values. Hence, a natural estimate of the response is simply:

align*[align* omitted — 216 chars of source]

\newline \newline

For $h = 0$, one can think of this difference in means as a measure of the average treatment effect in a randomized controlled trial or RCT. The same statistic can be estimated in regression form:

align[align omitted — 75 chars of source]

where $\mu_h = \mathbb{E}(y_{t+h}|s_t = 0)$ and $\beta_h = \mathbb{E}(y_{t+h}|s_t = 1) -\mathbb{E}(y_{t+h}|s_t = 0) = \mathcal{R}_{s y}(h)$.

(ref) can be easily generalized to allow for $s_t \in \mathbb{R}$ and to include a vector of controls $\bm x_t$ which, even if $s_t$ is randomly assigned, will improve efficiency. Using a linear framework, we arrive at the typical expression of a local projection:

align[align omitted — 89 chars of source]

If $s_t$ is not purely randomly assigned but depends on $\bm x_t$, then including $\bm x_t$ may reduce bias while the effect on the standard errors is ambiguous. The reason is that $\bm x_t$ could be correlated with $s_t$, but have no direct effect on $y_{t+h}$. In that case, including $\bm x_t$ in the regression reduces the variance of $s_t$ without reducing the variance of the regression residual. However, if one is willing to assume that the assignment of $s_t$ is as good as random only given $\bm x_t$, then one can appeal to the conditional independence assumptions carefully spelled out in, e.g., AngristK2011 and AJK2016. In this scenario, $\bm x_t$ needs to be included in the regression.

If $s_t$ is not randomly assigned even conditional on $\bm x_t$, but there is an instrument or instruments for it, call them $\bm z_t$, then clearly (ref) can be estimated using instrumental variables and we will be more specific later on about the relevance and exogeneity assumptions needed in this case.

Another obvious extension is to allow for the relationship between $y_{t+h}$ and the right-hand side variables of this expression to be possibly nonlinear, in which case the values of $s_0$ and $\bm x_t$ will be determinative of the actual response estimated. Hence, consider the following general additive conditional mean specification:

align[align omitted — 118 chars of source]

where $s_0$ is the initial state for $s_t$, $\delta$ is the impulse from that initial state, and $\bm \theta$ is the parameter vector. All other components have been previously defined.

Based on (ref), then:

align*[align* omitted — 176 chars of source]

where the notation makes explicit the dependence of the response on $s_0$, $\delta$, and $\bm x_t$. For example, consider a simple nonlinear LP (though linear in parameters):

align[align omitted — 106 chars of source]

then clearly $\mu_h(s_t = s_0 + \delta, x_t; \bm \theta_h) = \alpha_h + \beta_h (s_0 + \delta) + \gamma_h x_t + \phi_h (s_0 + \delta)^2 x_t$ whereas $\mu_h(s_t = s_0, x_t; \bm \theta_h) = \alpha_h + \beta_h s_0 + \gamma_h x_t + \phi s_0^2 x_t$ and hence $\mathcal{R}_{sy}(s_0, \delta, x_t, h) = \beta_h \delta + \phi_h(\delta^2 + 2\delta s_0)x_t$, which will depend on $s_0$, the initial value of the intervention variable, $\delta$ the size of the intervention, and $x_t$, the value of the current state variable. This example highlights the importance of being careful in how nonlinear responses are calculated as it has direct bearing on how inference needs to be constructed.

Assuming instruments are available and they meet relevance and exogeneity conditions that we will make more precise below, then (ref) can be estimated by the Generalized Method of Moments (GMM) based on the moment conditions:

align[align omitted — 147 chars of source]

Furthermore, as we shall explain below, (ref) for each $h$ can be stacked and estimated jointly as a system. Stacking as a system will prove useful when estimates of the covariance matrix of responses across all horizons is required for simultaneous inference or when it is required for nonlinear models such as the one in (ref).

Properties of the residual

Inference based on (ref) obviously depends on the properties of the residual, $v_{t+h}$. We present the main ideas of what differentiates local projections from traditional regression results in as simple a setup as possible. Hence suppose the data are generated by the following covariance-stationary AR(1) model. Using the companion form, the AR(1) encompasses more general AR(p) models and of course, in vector form, general VAR(p) models so the example, while simple, is helpful in thinking of more complex settings:

align[align omitted — 74 chars of source]

where $u_t$ is a zero-mean white noise process with $\mathbb{E} u_t^2 = \sigma^2_u$. All of these assumptions can be relaxed and made much more general but they suffice to make the point.

The goal is to calculate the response of $y_{t+h}$ to a shock $u_t$. Iterating forward $h$ periods into the future on the previous expression, we arrive at a common expression for a local projection, along the lines of that in (ref) where the shock of interest is now $u_t$ instead of some other exogenous variable. Hence we have:

align[align omitted — 75 chars of source]

where $\mu_h = m (1 + \rho + \hdots + \rho^{h-1})$; $\beta_{h} = \rho^{h}$, and $v_{t+h} = u_{t+h} + \rho u_{t+h-1} + \hdots + \rho^{h-1} u_{t+1}$. Given our assumptions, (ref) can be estimated by ordinary least-squares (OLS) and thus, a consistent estimate of $\mathcal{R}_{uy}(h)$ is $\hat \beta_h$.

However, notice that the residuals have an MA(h) structure that affects the computation of the standard error $\hat \beta_h$ since:

align*[align* omitted — 124 chars of source]

from where

align*[align* omitted — 146 chars of source]

Given our assumptions on $u_t$ it is easy to see that $\frac{1}{T-h}\sum_{t=2}^{T-h} y_t^2 \overset{p}{\to} \frac{1}{(1 - \rho^2)} \sigma^2_u$, whereas the numerator will converge in distribution to a Normal random variable whose variance, $\omega^2$ is:

align*[align* omitted — 167 chars of source]

where $v_{t+h} = u_{t+h} + \rho u_{t+h-1} + \hdots + \rho^{h-1} u_{t+1}$ and $y_t = u_t + \rho u_{t-1} + \hdots $. Putting the pieces back together we arrive at:

align[align omitted — 127 chars of source]

These derivations show that approximating $\omega$ in finite samples will require a heteroscedasticity and autocorrelation consistent estimator. Jorda2005 originally proposed using the Newey-West estimator as a simple solution. Since then, several developments that we discuss next have provided more elegant solutions.

LP Inference: The issues

In thinking about inference, it will be important to clearly outline the objective of the inferential procedure to design the best approach. In this regard at least three obvious objectives come to mind:

enumerate• Pointwise inference: How should one assess the precision of individual estimated response coefficients and the value that they attain? • Simultaneous inference: How should one assess subsets of response coefficients or the trajectory of the response as a whole? • Significance: What is the best way to test the null that the intervention has no effect on the outcome?

Layered on top of these three possible objectives, the analyst should consider the properties of alternative procedures available. PMWolf2021 and Xu2023 show that VARs and LPs are asymptotically equivalent when the data are generated by a VAR($\infty$) as long as the lag length is allowed to grow to infinity with the sample size. In finite samples and under the same VAR($\infty$) assumption, JST2024 show that, whereas the truncation lag used to estimate the VAR, say $p$, ensures consistency of the first $p$ lags of the VAR, it does not ensure consistency of the impulse response beyond the $p^{th}$ horizon whereas this is not the case for LPs: LPs remains consistent. Moreover, recent work by PMMOQW2024 shows that LPs are even robust to misspecification of the truncation lag.

Further, the bias-efficiency trade-off between VARs and LPs, determine the probability coverage properties of each approach. Going back to PMMOQW2024, these authors show that LP inference is robust to relatively large amounts of misspecification even when compared to VARs that are only mildly misspecified.

How did they arrive at this conclusion? PMMOQW2024 use a setting where they assume that the data are generated by a VAR where the residuals follow a moving-average structure that vanishes as the sample grows. In large samples, the VAR is clearly correctly specified. However, VARs undercover even when misspecification is so small, that it would be difficult or impossible to detect in small samples. On the other hand, they show that LPs have the correct coverage even when misspecification is so large that it can be easily detected in a finite sample.

Finally, a common concern is not just to derive procedures that are valid in large samples, but to construct methods that account for small sample properties of the data that may be far from the large sample ideal. These will usually include simulation-based inference, such as the bootstrap, as we shall discuss. Bayesian inference is also possible. Although LPs are not a generative model with which to construct the likelihood, Tanaka2020 and MirandaRicco2023, for example, provide Bayesian estimation methods of LPs. LPs can also be used to shrink large dimensional Bayesian VARs toward the responses generated with LPs as AgrippinoRicco2017 show.

Pointwise Inference

(ref) and (ref) more specifically, provide intuition for why, generally speaking, the residuals of the LP will have a moving average structure and how this might affect inference as a result. Jorda2005 proposed using a heteroscedasticity and autocorrelation consistent (HAC) estimator such as NeweyWest1987. Much of the literature appears to follow a similar strategy, even when it comes to panel data, where the Driscoll-Kraay DriscollKraay1998 covariance estimator is used instead.\footnote{The Driscoll-Kraay estimator can be seen as the extension to panel data of the Newey-West estimator.} However, whether the Driscoll-Kraay estimator is the right approach depends on the dimensions of the panel. Implicitly, the assumption is that $T >> N$, where $T$ is the time dimension of the panel and $N$ is the cross-section dimension. We reserve a more thorough discussion of panel data inference to a later section. Instead, we begin by discussing the recent LP inferential procedures introduced for time series data in the literature.

LP-FGLS

Lusompa2023 proposes a parametric feasible GLS procedure that directly accounts for the specific MA structure of the residuals. This LP-FGLS procedure can be best explained with the simple AR(1) example in (ref). The idea is to use the residuals from the first local projection, $\hat u_t$ as follows.

\paragraph{LP-GLS algorithm}

itemize• For $h = 1$ estimate: \begin{align*} y_t = m_1 + \beta_1 y_{t-1} + u_t \quad \rightarrow \quad \hat \beta_1, \, \hat u_t \end{align*} • For $h = 2$, construct $\tilde y_{t+1} = y_{t+1} - \hat \beta_1 \hat u_t$ and hence estimate: \begin{align*} \tilde y_{t+1} = m_2 + \beta_2 y_t + v_{t+1} \quad \rightarrow \quad \hat \beta_2 \end{align*} • For $h = 3$, construct $\tilde y_{t+2} = y_{t+2} - \hat \beta_1 \hat u_{t+1} - \hat \beta_2 \hat u_t$ and hence estimate: \begin{align*} \tilde y_{t+2} = m_3 + \beta_3 y_t + v_{t+2} \quad \rightarrow \quad \hat \beta_3 \end{align*} • For $h > 3$ sequentially generate $\tilde y_{t+h} = y_{t+h} - \hat \beta_1 \hat u_{t+h-1} - \hat \beta_2 \hat u_{t+h-2} - \hdots - \hat \beta_h \hat u_t$ and regress: \begin{align*} \tilde y_{t+h} = m_h + \beta_{h+1} y_t + v_{t+h} \end{align*}

The residuals $\hat v_{t+h}$ will be approximately white noise and hence heteroscedasticity robust inference is sufficient.

In a related paper, BreitungBrugemann2023 propose a similar approach that consists of using the transformation $\tilde y_{t+h} = y_{t+h} - \hat u_{t+h-1}$ and use $\{\hat u_t, \hdots, \hat u_{t+h-2} \}$ as additional regressors. BreitungBrugemann2023 show that this correction is as efficient as responses estimated with the correct vector autoregression in finite samples. Lusompa2023 also provides bootstrap versions of the LP-FGLS procedure, which have good efficiency gains relative to standard procedures.

figure[figure omitted — 517 chars of source]

As an illustration, (ref) compares estimates of an impulse response using the usual Newey-West correction with estimates based on this LP-GLS procedure. The responses are estimated using simulated data from a simple bivariate model given by:

align[align omitted — 335 chars of source]

with a sample of 150 observations (after disregarding 1,000 initial observations). (ref) shows the response of $x$ to a shock to $u^x$ and illustrates that both methods generate almost identical response estimates and very similar error bands with a slight efficiency edge for LP-FGLS. This is expected since the DGP fits quite well with the theoretical background.

Lag-augmentation of LPs

MOPM2021 introduce the idea of using lag-augmentation to conduct robust inference with LPs. They show that this procedure is uniformly valid over both stationary and non-stationary data and over a wide range of response horizons. Going back to the stylized model AR(1) in equation (ref) with $m = 0$, assume $u_t$ is strictly stationary and further assume $E(u_t| \{u_s\}_{s\ne t}) = 0$ almost surely. We make these assumptions to follow the setup in MOPM2021. Using similar notation to that in their paper, let $\beta(\rho,h)$ denote the LP parameter used to estimate the impulse response $\rho^h$, that is

align[align omitted — 143 chars of source]

Next, MOPM2021 suggest adding $y_{t-1}$ as an additional regressor to (ref). The purpose of this lag augmentation is to make the effective regressor of interest stationary even if the data $y_t$ has a unit root. MOPM2021 show that, rearranging, the lag-augmented local projection can be written as

align[align omitted — 109 chars of source]

Although $u_t$ is stationary and therefore would sidestep distortions to the normal distribution caused by near-to-unity asymptotics, it is not directly observed. However, due to the linear relationship between $y_t$ and $u_t$, the feasible local projection onto $(y_t, y_{t-1})$ provides an estimate of $\beta(\rho, h)$ precisely equal to the one that would be obtained from the projection onto $(u_t, y_{t-1})$.

Thus, the actual regression to be estimated is:

align[align omitted — 132 chars of source]

Lag-augmentation has two benefits. As MOPM2021 show, the distribution of $\hat \beta(h)$ of this feasible lag-augmented local projection is uniformly normal in $\rho \in [-1,1]$ using similar arguments to those used of lag-augmentation in AR inference SimsStockWatson1990, TodaYamamoto1995, DoladoLutkepohl1996, InoueKilian2002, InoueKilian2020. The second benefit is that it simplifies the computation of standard errors (though now the convergence rate will be $T^{1/2}$ instead of $T$).

In particular, it is sufficient to use a heteroskedasticity-robust routine to estimate standard errors for $\hat \beta(h)$, like the usual White correction (in STATA, {\tt{reg}} with the option {\tt{robust}} or even better, {\tt{hc3}}). How can we magically dispense with the moving average structure of the residuals evident in (ref)? From (ref), note that $u_t$ was assumed to be uncorrelated with past and future values of itself, and therefore the regression score $\xi_t(\rho, h) u_t$ is serially uncorrelated. To see this, note that the standard error formula in the ideal regression of (ref) would be

align*[align* omitted — 133 chars of source]

But by similar linearity arguments used to justify the feasible augmented local projection, it can be calculated directly from (ref) using White corrected standard errors as indicated.

Several remarks deserve mention. First, MOPM2021 show that lag-augmented LP inference is relatively robust to persistent data and provides appropriate coverage even at relatively long horizons (as long as $h_T/T \to 0$). Second, lag-augmentation is shown to work more generally when the DGP is assumed to be a VAR($p$) or a vector error correction model (VECM), though we are not aware that similar results have been derived for panel data in settings where the time dimension is larger than the cross section dimension, i.e., $T \gg N$. Of course, when $N \gg T$, asymptotic results are driven by the cross-sectional dimension of the panel and then the asymptotic distribution is normal even when the data are persistent. Third, lag-augmentation can also be applied to identified LPs PMWolf2021, MOPM2021.

As an illustration of how Newey-West and lag-augmented error bands compare, we revert back to data simulated from the process shown in (ref) but this time use LP-OLS estimates with error bands calculated using Newey-West vs lag-augmentation. This is shown in (ref). The figure shows that both methods generate similar error bands. In fact, several experiments (not reported here) suggest that, for stationary data, the coverage is very similar between methods. Lag-augmented bands tend to be somewhat more conservative the more persistent the data.

figure[figure omitted — 494 chars of source]

Finally, MOPM2021 provide bootstrap procedures that we briefly sketch here though the reader should go to the original source for details. Suppose that one wants to provide inference for an impulse response estimated with lag-augmented LPs for which one can also obtain the standard error as described earlier (i.e., using White corrected standard errors). MOPM2021 then suggest estimating the corresponding VAR($p$).\footnote{One can also bias-adjust the VAR coefficients using the correction by Pope1990.} This VAR will serve two purposes. One is to construct the equivalent response to that estimated with LPs, whose difference is then used to construct the $t$-ratio using the LP standard error. The second is to generate bootstrap replicates of the data using a parametric wild bootstrap GoncalvesKilian2004 based on the VAR($p$). Using these bootstrap replicates, then one estimates the lag-augmented LP responses and their standard errors. These are the ingredients necessary to then construct a percentile-$t$ confidence interval as usual.

Joint inference

Error bands, as they are overwhelmingly presented in the literature, are the result of inverting the t-statistics of individual hypotheses, as we have seen. When the interest turns to assessing the overall trajectory of the impulse response, the question becomes one of simultaneous inference. Testing joint hypotheses requires the covariance matrix of the response coefficients. Thus, we begin this section by providing a system Generalized Method of Moments (GMM) setup that will allow us to obtain this covariance matrix. Moreover, it will also allow us to consider instrumental variable estimation methods.

Although joint hypotheses tests are straightforward to construct (under rather general assumptions), presenting the evidence graphically, as we did with error bands, is not. After introducing the GMM estimator for LPs, we discuss two approaches to achieve this, one based on Scheff\'e's S-method, and another based on a sup-$t$ test. Both approaches generate more conservative error bounds designed to accommodate hypotheses tests on subsets of response coefficients. In practice, these methods have found less echo in the literature precisely because the bounds are wider. However, some of the insights from these procedures are valuable.

LP-GMM

GMM provides a convenient framework to estimate LPs jointly, as we advanced when discussing (ref). Let $\bm y_t(H) = (y_t, \hdots, y_{t+H})'$ be an $(H+1) \times 1$ vector that collects the outcome variable observed at increasingly distant horizons into the future. Let $S_t = I_{H+1} \otimes s_t$ where $I_{H+1}$ is the identity matrix of order $H+1$ and $s_t$ is the intervention variable. We collect the error terms in $\bm v_t(H) = (v_t, \hdots, v_{t+h})'$. The response coefficients are $\bm \beta = (\beta_0, \hdots, \beta_H)'$. Exogenous and predetermined variables collected in $\bm x_t$ could be easily included by defining $X_t = I_{H+1} \otimes \bm x_t$. However, for simplicity, we will mostly set them aside for the presentation. In linear settings, we could simply invoke the Frisch-Waugh-Lovell theorem and orthogonalize the outcome and the intervention with respect to the controls before proceeding. Finally, suppose an $1 \times l$ vector $\bm z_t$ of instrumental variables is available with $l \ge 1$. Hence we construct $Z_t = I_{H+1} \otimes \bm z_t.$ If no instruments are available, but $s_t$ is exogenously determined (perhaps conditional on controls), then one can simply set $Z_t = S_t$.

Using these definitions, the population moment conditions of the system of local projections can be expressed as:

align*[align* omitted — 54 chars of source]

Thus, the corresponding finite sample GMM problem can be written as:

align*[align* omitted — 203 chars of source]

with $N = T^* - t_0$, where $t_0$ denotes the first observation available after accounting for possible lags in the control set and $T^* = T-H-1$. One could choose to set $\hat \Lambda = I_{l\times (H+1)}$. The estimator, referred to as the equally weighted estimator, yields consistent estimates of $\bm \beta$, but the covariance matrix is not valid for inference, as is well-known. \newline \newline \newline

The optimal weighting matrix, correcting for heteroscedasticity and autocorrelation with, for example, a Barlett correction, is:

align*[align* omitted — 241 chars of source]

where $\tilde{\bm v_t}(H)$ refers to the residuals based on the equally weighted estimator, which we know to be consistent.

StockWatson2018 and PMWolf2022 provide appropriate relevance and exogeneity conditions (under relatively general assumptions for the underlying DGP) for the analysis of LPs. In our model, these can be stated as:

itemize• Relevance: $E(\bm z_t's_t| \bm x_t) \neq \bm 0$ • Exogenenity: $E(Z_{t}' \bm v_t(H)|X_t) = \bm 0$

where we explicitly include the controls in the conditioning set to make clearer their influence (despite having proceeded by projecting their influence away in a first stage or simply when they are subsumed in the vector of instruments, $\bm z_t$). Our exogeneity condition is satisfied under the contemporaneous and lead-lag exogeneity conditions of StockWatson2018.

Based on the relevance and exogeneity assumptions just stated and standard results from the algebra of GMM, estimates of the impulse response can be obtained from:

align*[align* omitted — 223 chars of source]

which will be consistent and asymptotically normally distributed with approximate covariance matrix given by:

align[align omitted — 191 chars of source]

Joint estimation of the impulse response using this GMM set-up is useful as it provides an estimate of the covariance matrix, $\hat \Omega_\beta$, a key ingredient to construct any joint hypothesis test of interest. It will also be an important ingredient in the construction of error bands that accommodate simultaneous inference, as the next section briefly explains.

Simultaneous inference

Analysts interested in testing features of the impulse response involving more than one parameter can set up the appropriate hypothesis test as usual using the F-test, for example, and report the results of the test in a table. However, how should one represent simultaneous inference graphically if one were interested in representing bounds with appropriate probability coverage that accommodate a variety of hypothesis tests an analyst may conduct? Impulse response coefficients are correlated. Thus, the usual practice of inverting the t-ratio to display error bands will not provide the correct probability coverage. The correct approach requires that we construct error bands that account for the simultaneous nature of the family of hypotheses of interest.

This problem was highlighted by Jorda2009. His solution relied on Scheff\'e's multiple comparison approximation or S-method Scheffe1953. Asymptotically, the sum of the squares of the t-ratios of the LP is approximately $\chi^2_H$ distributed. Thus, the critical value with which to construct simultaneous error bands is $(c_\alpha^2/H)^{1/2}$ where $c_\alpha^2$ refers to the critical value of a $\chi^2_H$. The advantage of this method is that it does not require simulation methods to obtain the critical value.

More recently, MontielOleaPM2019 proposed a more efficient approximation based on the sup-$t$ procedure. Although this method requires simulation methods, it is shown to produce the narrowest bands for a given probability coverage. In particular, let $\bm \beta = (\beta_1, \hdots, \beta_H),$ and assume $\bm \beta \overset{d}{\to} N(0, \Omega_\beta).$ The goal is to find $c_\alpha$ such that the error bands defined by:

align*[align* omitted — 241 chars of source]

such that:

align*[align* omitted — 66 chars of source]

MontielOleaPM2019 show that:

align*[align* omitted — 212 chars of source]

Although the distribution of $ \max_{h=0,1, \hdots, H} \left| \Omega_{h,h}^{-1} V_h \right|$ is unknown, it is relatively easy to obtain any desired quantile of this distribution. Thus, error bands for simultaneous inference can be constructed according to the following algorithm: \newline \newline

Plug-in sup-$t$ algorithm

itemize• Step 1: Draw $M$ i.i.d. normal vectors $\hat V^{(m)} \sim N_H(\bm 0_H, \hat \Omega_\beta), \, m = 1, \hdots, M.$ • Step 2: Define $\hat q_{1-\alpha},$ the empirical $1 - \alpha$ quantile of $ \max_{h=0,1, \hdots, H} \left| \Omega_{h,h}^{-1/2} V_h \right|$ across $m = 1, \hdots, M.$ • Step 3: Construct the error bands for each $h$ as: $[\hat \beta_h - \hat q_{1-\alpha} \hat \Omega_{h,h}, \hat \beta_h + \hat q_{1-\alpha} \hat \Omega_{h,h}]; \, h = 0,1, \hdots, H$

MontielOleaPM2019 further provide bootstrap and Bayesian versions of this procedure. The interested reader should consult their paper for more details.

To get a sense of how much wider the sup-$t$ bands are, we revisit the example of (ref) and (ref). We simulate data from the same model and then construct 95% sup-$t$ bands using joint estimation of all responses simultaneously to obtain the resulting covariance matrix. In (ref), the sup-$t$ bands are shown as dashed blue lines. Relative to the usual Newey-West bands, shown as a shaded region in red, it is easy to see that they are wider but not by a very large amount.

figure[figure omitted — 484 chars of source]

Significance

In a randomized controlled trial (RCT), a common hypothesis of interest is to test the null that the treatment is ineffective. Similarly, an impulse response can be thought of as the response to a treatment observed over time and therefore, a common hypothesis of interest will be to assess whether the response is different from zero, just as in the RCT example. It turns out that under the null, the distribution of the appropriate hypothesis test can simplify the construction of significance bands considerably.

To understand the basic issues, we use a simple example where the controls are set aside to keep the notation simple. For example, we may presume that the outcome $y_{t+h}$, the intervention, $s_t$, and the instrument $z_t$ have been previously orthogonalized with respect to the controls $\bm x_t$ based on the Frisch-Waugh-Lovell theorem. Thus, consider the following local projections estimated by instrumental variables regression:

align*[align* omitted — 109 chars of source]

We assume that $z_t$, meets the conditions for relevance and exogeneity, as well as an exclusion restriction. Then we postulate that:

itemize• Relevance: $E(s_t \, z_t) \ne 0$. • Exogeneity: $E(v_{t+h} \, z_{t}) = 0 \,\,\, \forall h=0, 1, \hdots, H - 1$.

Depending on the setting, $z_t$ may include $s_t$ itself, such as when $s_t$ is an observable shock, and then the discussion returns to a more traditional OLS setting. Or if $s_t$ is, conditional on $x_t$, sequentially exogenous. This would be the case in a recursive identification scheme. We further assume that $y_t$, $s_t$, and $z_t$ are covariance stationary. This assumption is not necessary to ensure consistency of the local projection, but will make deriving our inferential procedures and the presentation in this section straightforward.

Based on this simple set up, the instrumental variable estimator for $\beta_h$ can be written as:

align[align omitted — 181 chars of source]

where we note that we will evaluate the statistic under the null $H_0:\beta_h = 0$. Under standard regularity conditions, and the instrumental variable assumptions for local projections, it is easy to see that:

align*[align* omitted — 97 chars of source]

Next, consider the numerator in (ref) evaluated at the null $H_0: \beta_h = 0$:

align*[align* omitted — 95 chars of source]

where $s_{\eta,h}^2$ is given by

align*[align* omitted — 204 chars of source]

and where the RHS is the typical HAC type variance formula. The equality follows from the null hypothesis that $\beta_h = 0$ for $h = 0,1, \hdots, H-1$. It then follows that the limiting distribution of $\hat\beta_h$ is given by

align[align omitted — 171 chars of source]

From (ref) it is easy to derive a $1-\alpha$ percent band around the zero null so that:

align*[align* omitted — 154 chars of source]

where $\zeta_{\alpha/2}$ is the critical value of a standard normal variable at $\alpha/2$ and for a standard normal, $\zeta_{1 - \alpha/2} = -\zeta_{\alpha/2}$, as is well known.

To construct feasible confidence intervals we need to replace $\sigma_h$ with an estimate. The LM principle requires that $\sigma_h$ be estimated using the conventional formula for HAC robust standard errors for the just identified two-stage least squares estimator, but evaluated at $\beta_h = 0$. This is accomplished by estimating $s_{\eta,h}^2$ with the long-run variance of $\eta_{t,h} = z_t y_{t+h}$. The estimator for $\sigma_h^2$ then is based on

align[align omitted — 79 chars of source]

When plotting a significance band of an impulse response up to $H$ periods, we are essentially conducting a joint hypothesis test. Intuitively, the more horizons considered, the more likely it is to spuriously reject the null when the null is true in a finite sample. A simple way to address this issue is with a Bonferroni adjustment as proposed in Dunn1961 so that the significance bands for each $\hat \beta_h$ become:

align*[align* omitted — 138 chars of source]

The joint probability that the estimated impulse response lies within the confidence band is given by:

align*[align* omitted — 206 chars of source]

where the inequality holds in large samples and when the null hypothesis of a zero response is true. Similarly, the test of the joint hypothesis that all response coefficients are zero rejects when:

align*[align* omitted — 158 chars of source]

for at least one $h$. By the same argument, it follows that the size of such a test is not more than $\alpha$ in large samples.

Under stronger assumptions such as independence between $z_t$ and $v_s$ for all $t$ and $s$, or homoskedasticity of $v_{t+h}$ such that $E(v_{t+h - j} v_{t+h}|z_t, z_{t-j} \> )=E(v_{t+h - j}v_{t+h})$ a further simplification of the expression for $s_{\eta,h}^2$ is possible. We obtain

align[align omitted — 256 chars of source]

where the equality follows from the independence between $z_t$ and $v_{t+h}$. We define $\gamma_{z,j}$ and $\gamma_{y,j}$ as the $j^{th}$ autocovariances of $z$ and $y$ respectively. Importantly, note that $\omega$ is no longer a function of the horizon $h$ under these additional restrictions. The implication is that under the additional restrictions of homoskedasticity or independence, the significance bands will be constant as a function of the horizon h.

Under these stronger conditions we can write (ref) under the null hypothesis as:

align[align omitted — 234 chars of source]

A simple example provides further intuition and a connection to well-known results. In the special case where $z = s$, and $y$ and $s$ are serially uncorrelated, (ref) simplifies even further to:

align*[align* omitted — 58 chars of source]

Thus, when $y = s = z$ and $y$ is a white noise and hence $\gamma_{y,0} = \gamma_{s,0}$ so that $\sigma^2 = 1$, the local projection estimator is simply an estimator of the autocorrelation function. Hence, applying the same derivations as in (ref), it is easy to see that one recovers the well known\footnote{Not Barlett corrected.} bands for the autocorrelogram of $y$. Specifically, focus on $h = 1$ in the special case that $y$ is a white noise but one estimates an AR(1) model:

align[align omitted — 81 chars of source]

This is the well known case where the 95% asymptotic significance bands in a correlogram are calculated as $\pm 1.96 \times 1/\sqrt{n}$ and provides a nice window into our proposed procedures. Importantly, notice that the bands do not depend on the horizon (in fact, they also do not depend on the variance in this special case). Whenever an autocorrelation coefficient exceeds the band, the interpretation is that said coefficient can be deemed to be different from zero. This, of course, means that the hypothesis that the impulse/treatment has no effect on the outcome can be rejected.

Practical implementation

Constructing significance bands in practice based on the results from the previous section is straightforward and can be implemented using standard statistical software. We provide a STATA example to illustrate this point and that corresponds to the figures displayed in the paper. The basic steps can be summarized as follows:

\hrule Significance bands using asymptotic approximations \hrule

enumerate• Calculate the sample average of the product $s_t \, z_t$. Call this $\hat \gamma_{sz}$. • Construct the auxiliary variable $\eta_{t,h} = y_{t+h} \,z_t$ and regress $\eta_{t,h}$ on a constant. The Newey-West estimate of the standard error of the intercept coefficient is an estimate of $s_{\hat{\eta},h}$. • An estimate of $\sigma/\sqrt{T-h}\,$, call it $\hat s_{\beta_h}\,$, is therefore: \begin{align*} \hat s_{{\beta}_h} = \frac{\hat s_{{\hat \eta},h}}{\hat{\gamma}_{sz}} \end{align*} • Construct the significance bands as: \begin{align*} \left[\zeta_{\alpha/2(H+1)} \hat s_{{\beta}_h}, \,\, \zeta_{1 - \alpha/2(H+1)} \hat s_{\beta_h} \right] \end{align*}

\hrule

A bootstrap procedure is equally easy to construct. Note that we do not take a position on the data generating process (DGP). Therefore, we apply the bootstrap directly to step 2 of the previous construction of the significance band. Because of the time series dependence and the possible existence of heteroscedasticity, we will use a wild-block bootstrap GoncalvesKilian2007. The STATA implementation only requires a few lines of code. Thus, the entire procedure can be described as follows:

\hrule Significance bands using the Wild-Block Bootstrap \hrule

enumerate• Calculate the sample average of $s_t \, z_t$. Call this $\hat \gamma_{sz}$. • Construct the auxiliary variable $\eta_{t,h} = y_{t+h} \,z_t$ and regress $\eta_{t+h}$ on a constant. The Wild Block bootstrap estimate of the standard error of the intercept coefficient is an estimate of $s_{\hat{\eta},h}$. • An estimate of $\sigma/\sqrt{T-h}\,$, call it $\hat s^b_{\beta_h}\,$, is therefore: \begin{align*} \hat s^b_{{\beta}_h} = \frac{\hat s^b_{{\hat \eta},h}}{\hat{\gamma}_{sz}} \end{align*} • Construct the significance bands as: \begin{align*} \left[\zeta_{\alpha/2(H+1)} \hat s^b_{{\beta}_h}, \,\, \zeta_{1 - \alpha/2(H+1)} \hat s^b_{\beta_h} \right] \end{align*}

Application: the response of shelter inflation to monetary policy

We showcase these procedures with a simple application to shelter inflation. As the COVID-19 pandemic was winding down, inflation measured by the personal consumption expenditures (PCE) index, excluding food and energy (core PCE inflation), peaked at around 5.5% on March 2022. In response, the Federal Reserve raised the federal funds rate and subsequently inflation declined to 2.6% by June 2024 (the last data point as of the writing of this paper). However, inflation has been slow to travel back to the 2% target, in large part, because shelter inflation (which mainly measures rents that tenants face and owners de facto pay themselves) has been slow to decline.

Hence, we evaluate how responsive shelter inflation is to interest rates. We use the series of monetary shocks recently provided by BauerSwanson2023. These shocks are obtained from high frequency financial data after removing information effects. The sample available is January 1988 to December 2019 so as to avoid polluting our estimates with the COVID-19 pandemic. We use as controls 12 lags of shelter Personal Consumption Expenditures (PCE) inflation, the unemployment rate (to control for aggregate conditions in the economy), and lags of the federal funds rate. The left-hand side variable is the long difference of 100 times the log of shelter PCE price index (PCEPI), that is $100(y_{t+h} - y_{t-1})$, where $y_t$ is the log of shelter PCEPI. This means that the response reflects the cumulative percentage change in the level of shelter prices up to $h$ periods since the shock.

(ref) displays estimates of this cumulative response along with Newey-West confidence bands (for 1 and 2 standard deviations in width) along with significance bands estimated with both analytic and bootstrap methods. Based on the conventional error bands displayed in the figure, one might be tempted to conclude that shelter inflation does not respond to monetary shocks. The error bands contain zero throughout the 48 periods displayed.

However, note that the response of shelter PCEPI is almost uniformly negative each month over the 4 years displayed. Indeed, an F-test of the null that the response coefficients are jointly zero is estrenuosly rejected (the p-value is 6.18e-78). And in fact, the significance bands displayed show that, except for about two and one half years after the shock, the response is clearly different from zero. Analytic bands are narrower than bootstrap bands for about 30 periods, after which they are slightly wider.

figure[figure omitted — 816 chars of source]

Monte Carlo evidence

This section presents a couple of simple experiments in graphical form to assess the calculation of significance bands using both the asymptotic approximation and the wild-block bootstrap procedures discussed in the previous section. The data are generated as follows:

align*[align* omitted — 243 chars of source]

This simple system encapsulates several features. First, the treatment variable, $s_t$, affects the outcome, $y_t$, contemporaneously. The outcome is itself serially correlated with a coefficient $0.75$. The idea is to have internal propagation dynamics. Next, the intervention responds to feedback from the value of the outcome in the previous period, but also has some internal propagation dynamics. In addition, movements in the intervention are caused by the exogenous variable $z_t$, which will act as our instrumental variable. Finally, the coefficient $\beta$, which captures the effect of the treatment on the outcome, has values between 0 and 0.75. When $\beta = 0$ we have the null model with which to assess the size of the test. Increasing the value of $\beta$ allows us to assess the power of the significance bands.

We generate samples of 100, and 500 observations with 500 burn-in observations that are discarded to avoid initialization problems. For each sample size and for the different values of $\beta$ we generate 1,000 Monte Carlo replications. The implementation of the wild block-bootstrap is based on 1,000 bootstrap replications as well. For the Newey-West step as well as for the block size in the bootstrap, we use 8 lags. (ref) displays the results for sample sizes of 100 and 500 observations.

The figure summarizes quite a bit of information. The shaded bands around the mean estimate of the impulse response showcase the $25^{th}$ and the $975^{th}$ largest values for each coefficient estimate in the Monte Carlo simulation. The dashed lines correspond to the significance bands. Both Newey-West and the bootstrap procedures (using 8 lags) generate nearly indistinguishable values so the differences cannot be seen with the naked eye. For each Monte Carlo exercise, we construct rejection rates for each type of band constructed. The rate is calculated as the share of replications where one or more impulse response coefficients exceed the significance bands.

Several results deserve comment. First, consider the size of the test. We have chosen a rather conservative strategy with a window of size 8 both for Newey-West and for the block-size in the implementation of the bootstrap. As a result, with a small sample of 100 observations, the size is about 10% instead of the nominal 5%, though with 500 observations the size is close to 4%.

However, even with this conservative choice, the power of the test is respectable with a sample size of 100, improving from about 25% when $\beta = 0.25$ to about 95% when $\beta = 0.75$. These numbers jump with 500 observations with about 95% for $\beta = 0.25$ and 100% even for $\beta = 0.5$.

Other LP settings

Work on extensions to LPs is ongoing. Two settings in particular, have direct implications for how we think about LP inference. LPs can be seen as a semiparametric estimate of the true impulse response and the reason why they have lower bias. Thus, some authors have proposed methods to either smooth the LPs via some low dimensional series approximation BarnichonBrownlees2018, BarnichonMatthes2018, or by pairing LPs with Bayesian methods to shrink high-dimensional VARs toward the LP response Tanaka2020, MirandaRicco2023.

The other setting that is popular in empirical research is panel data. Because LPs are a single-equation method, they lend themselves well to panel data settings. However, although there is an extensive literature using LPs with panels and there is an extensive literature on inference in panel data settings, little is known about LP inference in panel data settings. In the next sections, we give a brief overview on these less traditional settings.

Smoothing local projections

Smoothing can be used to reduce the uncertainty about the impulse response. Consider the GMM moment condition that we introduced in (ref), repeated here for convenience:

align*[align* omitted — 53 chars of source]
figure[figure omitted — 980 chars of source]

\FloatBarrier

where $\bm \beta$ is a $(H+1) \times 1$ dimensional vector with covariance matrix $\Omega_\beta$, which can be estimated as shown in (ref). Smoothing can be thought of as replacing $\bm \beta$ with a lower-dimensional function, say $\phi(h, \bm \theta)$ with $dim(\bm \theta) << dim(\bm \beta)$. In BarnichonBrownlees2018, the authors propose using the B-spline method of EilersMarx1996, whereas BarnichonMatthes2018 propose using Gaussian basis functions.

In general, assuming that one can obtain LP estimates that are asymptotically normal (such as when estimating LPs using the GMM setup in (ref)) so that $\hat{\bm \beta} \overset{d}{\to} N(\bm \beta, \Omega_\beta)$ and $\hat \Omega_\beta \overset{p}{\to} \Omega_\beta$, then one can setup the minimum distance problem:

align[align omitted — 155 chars of source]

to estimate $\hat \phi(h, \hat{\bm \theta})$ efficiently. Moreover, under standard regularity conditions, the quality of the approximation can be judged with an overidentification restrictions test since $\hat Q(\hat{\bm \theta}) \to \chi^2_q$ with $q = dim(\bm \beta) - dim(\bm \theta)$. The variance of $\bm \theta$ is $\hat \Omega_\theta = (\Phi_0' \hat \Omega_\beta^{-1} \Phi_0)^{-1}$ where $\Phi_0 = \partial \phi(h, \bm \theta)/\partial \bm \theta$ evaluated at the true value $\bm \theta_0$. An example of how smoothing can make response estimates more efficient is LPMW2022.

Panel data

Panel data offers more opportunities to explore data using LPs and more opportunities to conduct inference. As we anticipated in (ref), when the $T \to \infty$ with $N$ fixed or when $N$ grows at a slower rate than $T$, the asymptotic distribution of the LP estimates will be dominated by the time-dimension of the panel and in that case, the counterpart to Newey-West HAC standard errors is to use the DriscollKraay1998 covariance estimator.\footnote{The command {\tt xtscc} in STATA implements this type of covariance matrix estimation.}

Cluster robust inference can be used in situations where $N \to \infty$ with $T$ fixed. In such a setting, autocovariances are relatively efficiently estimated and there is no need to specify the residual autocorrelation structure. When $T$ is relatively small, a recommended correction for heteroscedasticity is the wild bootstrap CameronGM2008, CanaySS2021, MacKinnonOR2019\footnote{In STATA, this type of boostrap can be implemented with the user supplied command {\tt boottest}.}. If $N$ is relatively large, the asymptotic distribution will be dominated by the cross-sectional dimension of the panel. In that case, stationarity or lack thereof plays no role in computing standard errors.

Relatedly, in a recent paper, MeiShengShi2023 show that incidental parameter biases Nickell1981 crop up when the dimensions of the panel $N,T \to \infty$ as $N/T \to c,\, c \in (0, \infty)$. To avoid this bias, these authors suggest using the split panel jacknife estimator of Dhaene2016 and Chudik2018. Denote $\hat \beta_h$ the full sample estimate with fixed-effects and $\hat \beta_h^a$ and $\beta_h^b$ estimates based on splitting the sample along the time series dimension into two halves, $T = T_a + T_b$. Then, the bias corrected estimate of the impulse response is $\tilde \beta_h = 2 \hat \beta_h - 0.5(\hat \beta_h^a + \hat \beta_h^b).$ However, note that a recent paper by HHKW2024 shows that the split panel jackknife bias correction is generally higher order inefficient and may significantly increase estimator variance in finite samples compared to higher order efficient bias correction methods.

Conclusion

The error structure of local projections and the kind of hypotheses implicit in the way inference is obtained and communicated requires some care. Residual serial correlation in local projections can be dealt with and more recent developments show that it can be obtained rather easily using lag augmentation and heteroscedasticity robust methods. Bounds for simultaneous inference are also relatively easy to construct, though in most situations we expect that practitioners will simply report formal hypothesis tests.

The significance bands introduced in this review are simple to construct and can be easily displayed alongside the usual confidence bands. While confidence bands inform the reader about the estimation uncertainty of each coefficient, significance bands inform the reader about the significance of the impulse response itself.

Panel data settings present their own challenges and opportunities. The time series and cross-sectional dimensions of the panel play a critical role in choosing the best inferential procedures. Cluster robust inference generally offers an attractive approach, but cannot always be directly used. Inference in panel data settings is an ever growing field and new developments are constantly arriving to improve existing methods.

Applications and extensions of local projections continue to grow. In this review we are unable to cover every single scenario. However, we hope to have provided the reader with general principles that can then be tailored to each specific extension.

\singlespacing

\setcounter{secnumdepth}{0}

thebibliography\bibitem[\citename{Angrist & Kuersteiner, }2011]{AngristK2011} Angrist, Joshua D., & Kuersteiner, Guido M. 2011. \newblock Causal Effects of Monetary Shocks: Semiparametric Conditional Independence Tests with a Multinomial Propensity Score. \newblock {\em The Review of Economics and Statistics}, {\bf 93}(3), 725--747. \bibitem[\citename{Angrist {\em et al.}, }2016]{AJK2016} Angrist, Joshua D., Jord{\`a}, {\`O}scar, & Kuersteiner, Guido M. 2016. \newblock Semiparametric Estimates of Monetary Policy Effects: String Theory Revisited. \newblock {\em Journal of Business and Economic Statistics}, {\bf http://dx.doi.org/10.1080/07350015.2016.1204919}. \bibitem[\citename{Barnichon & Brownlees, }2019]{BarnichonBrownlees2018} Barnichon, Regis, & Brownlees, Christian. 2019. \newblock {Impulse Response Estimation by Smooth Local Projections}. \newblock {\em Review of Economics and Statistics}, {\bf 101}(3), 522--530. \bibitem[\citename{Barnichon & Matthes, }2018]{BarnichonMatthes2018} Barnichon, Regis, & Matthes, Christian. 2018. \newblock Functional Approximation of Impulse Responses. \newblock {\em Journal of Monetary Economics}, {\bf 99}, 41--55. \bibitem[\citename{Bauer & Swanson, }2023]{BauerSwanson2023} Bauer, Michael D, & Swanson, Eric T. 2023. \newblock A reassessment of monetary policy surprises and high-frequency identification. \newblock {\em NBER Macroeconomics Annual}, {\bf 37}(1), 87--155. \bibitem[\citename{Breitung & Br\"uggemann, }2023]{BreitungBrugemann2023} Breitung, J\"org, & Br\"uggemann, Ralf. 2023. \newblock Projection estimators for structural impulse responses. \newblock {\em Oxford Bulletin of Economics and Statistics}, {\bf 85}(6), 1320--1340. \bibitem[\citename{Cameron {\em et al.}, }2008]{CameronGM2008} Cameron, A. Colin, Gelbach, Jonah B., & Miller, Douglas L. 2008. \newblock Bootstrap-based improvements for inference with clustered errors. \newblock {\em Review of Economics and Statistics}, {\bf 90}(3), 414--427. \bibitem[\citename{Canay {\em et al.}, }2021]{CanaySS2021} Canay, Ivan A., Santos, Andres, & Shaikh, Azeem M. 2021. \newblock The wild bootstrap with a “small” number of “large” clusters. \newblock {\em Review of Economics and Statistics}, {\bf 103}(2), 346--363. \bibitem[\citename{Chudik {\em et al.}, }2018]{Chudik2018} Chudik, Alexander, Pesaran, M. Hashem, & Yang, Jui-Chung. 2018. \newblock Half-panel jackknife fixed-effects estimation of linear panels with weakly exogenous regressors. \newblock {\em Journal of Applied Econometrics}, {\bf 33}(6), 816--836. \bibitem[\citename{Dhaene & Jochmans, }2016]{Dhaene2016} Dhaene, Geert, & Jochmans, Koen. 2016. \newblock Bias-corrected estimation of panel vector autoregressions. \newblock {\em Economics Letters}, {\bf 145}, 98--103. \bibitem[\citename{Dolado & L{\"u}tkepohl, }1996]{DoladoLutkepohl1996} Dolado, Juan J., & L{\"u}tkepohl, Helmut. 1996. \newblock Making Wald tests work for cointegrated VAR systems. \newblock {\em Econometric Reviews}, {\bf 15}(4), 369--386. \bibitem[\citename{Driscoll & Kraay, }1998]{DriscollKraay1998} Driscoll, John C., & Kraay, Aart C. 1998. \newblock Consistent Covariance Matrix Estimation with Spatially Dependent Panel Data. \newblock {\em Review of Economics and Statistics}, {\bf 80}(4), 549--560. \bibitem[\citename{Dunn, }1961]{Dunn1961} Dunn, Olive Jean. 1961. \newblock Multiple comparisons among means. \newblock {\em Journal of the American statistical association}, {\bf 56}(293), 52--64. \bibitem[\citename{Eilers & Marx, }1996]{EilersMarx1996} Eilers, Paul H. C., & Marx, Brian D. 1996. \newblock Flexible smoothing with B-splines and penalties. \newblock {\em Statistical Science}, {\bf 11}(2), 89--121. \bibitem[\citename{Ferreira {\em et al.}, }2023]{MirandaRicco2023} Ferreira, Leonardo N., Miranda-Agrippino, Silvia, & Ricco, Giovanni. 2023. \newblock {Bayesian Local Projections}. \newblock {\em Review of Economics and Statistics}, 05, 1--45. \bibitem[\citename{Gon{\c{c}}alves & Kilian, }2004]{GoncalvesKilian2004} Gon{\c{c}}alves, S{\'\i}lvia, & Kilian, Lutz. 2004. \newblock Bootstrapping autoregressions with conditional heteroskedasticity of unknown form. \newblock {\em Journal of Econometrics}, {\bf 123}(1), 89--120. \bibitem[\citename{Gon{\c{c}}alves & Kilian, }2007]{GoncalvesKilian2007} Gon{\c{c}}alves, S{\'\i}lvia, & Kilian, Lutz. 2007. \newblock Asymptotic and bootstrap inference for AR ($\infty$) processes with conditional heteroskedasticity. \newblock {\em Econometric Reviews}, {\bf 26}(6), 609--641. \bibitem[\citename{Hahn {\em et al.}, }2024]{HHKW2024} Hahn, Jinyong, Hughes, David W., Kuersteiner, Guido, & Newey, Whitney K. 2024. \newblock Efficient bias correction for cross section and panel data. \newblock {\em Quantitative Economics}, {\bf 15}, 783--816. \bibitem[\citename{Inoue & Kilian, }2002]{InoueKilian2002} Inoue, Atsushi, & Kilian, Lutz. 2002. \newblock Bootstrapping autoregressive processes with possible unit roots. \newblock {\em Econometrica}, {\bf 70}(1), 377--391. \bibitem[\citename{Inoue & Kilian, }2020]{InoueKilian2020} Inoue, Atsushi, & Kilian, Lutz. 2020. \newblock The uniform validity of impulse response inference in autoregressions. \newblock {\em Journal of Econometrics}, {\bf 215}(2), 450--472. \bibitem[\citename{Jord{\`a}, }2005]{Jorda2005} Jord{\`a}, {\`O}scar. 2005. \newblock {Estimation and Inference of Impulse Responses by Local Projections}. \newblock {\em American Economic Review}, {\bf 95}(1), 161--182. \bibitem[\citename{Jord\`a, }2009]{Jorda2009} Jord\`a, \`Oscar. 2009. \newblock Simultaneous confidence regions for impulse responses. \newblock {\em Review of Economics and Statistics}, {\bf 91}(3), 629--647. \bibitem[\citename{Jord{\`a} {\em et al.}, }2024]{JST2024} Jord{\`a}, {\`O}scar, Singh, Sanjay R., & Taylor, Alan M. 2024. \newblock The long-run effects of monetary policy. \newblock {\em Review of Economics and Statistics}, {\bf forthcoming}. \bibitem[\citename{Li {\em et al.}, }2024]{LPMW2022} Li, Dake, Plagborg-M{\o}ller, Mikkel, & Wolf, Christian K. 2024. \newblock Local projections vs. vars: Lessons from thousands of dgps. \newblock {\em Journal of Econometrics,}, {\bf forthcoming}. \bibitem[\citename{Lusompa, }2023]{Lusompa2023} Lusompa, Amaze. 2023. \newblock Local projections, autocorrelation, and efficiency. \newblock {\em Quantitative Economics}, {\bf 14}(4), 1199--1220. \bibitem[\citename{Mei {\em et al.}, }2023]{MeiShengShi2023} Mei, Ziwei, Sheng, Liugang, & Shi, Zhentao. 2023. \newblock {\em {Nickell Bias in Panel Local Projection: Financial Crises Are Worse Than You Think}}. \newblock Unpublished. \url{https://arxiv.org/pdf/2302.13455.pdf}. \bibitem[\citename{Miranda-Agrippino & Ricco, }2017]{AgrippinoRicco2017} Miranda-Agrippino, Silvia, & Ricco, Giovanni. 2017. \newblock The Transmission of Monetary Policy Shocks. \newblock {\em Center for Macroeconomics working paper}, {\bf DP}(11). \bibitem[\citename{Montiel-Olea & Plagborg-M{\o}ller, }2019]{MontielOleaPM2019} Montiel-Olea, Jos\'e Luis, & Plagborg-M{\o}ller, Mikkel. 2019. \newblock Simultaneous Confidence Bands: Theory, Implementation, and an Application to SVARs. \newblock {\em Journal of Applied Econometrics}, {\bf 34}(1), 1--17. \bibitem[\citename{Montiel Olea & Plagborg-M{\o}ller, }2021]{MOPM2021} Montiel Olea, Jos{\'e} Luis, & Plagborg-M{\o}ller, Mikkel. 2021. \newblock Local projection inference is simpler and more robust than you think. \newblock {\em Econometrica}, {\bf 89}(4), 1789--1823. \bibitem[\citename{Newey & West, }1987]{NeweyWest1987} Newey, Whitney K., & West, Kenneth D. 1987. \newblock {A Simple, Positive Semi-definite, Heteroskedasticity and Autocorrelation Consistent Covariance Matrix}. \newblock {\em Econometrica}, {\bf 55}(3), 703--708. \bibitem[\citename{Nickell, }1981]{Nickell1981} Nickell, Stephen. 1981. \newblock Biases in dynamic models with fixed effects. \newblock {\em Econometrica: Journal of the econometric society}, 1417--1426. \bibitem[\citename{Plagborg-M{\o}ller & Wolf, }2021]{PMWolf2021} Plagborg-M{\o}ller, Mikkel, & Wolf, Christian K. 2021. \newblock Local projections and VARs estimate the same impulse responses. \newblock {\em Econometrica}, {\bf 89}(2), 955--980. \bibitem[\citename{Plagborg-M{\o}ller & Wolf, }2022]{PMWolf2022} Plagborg-M{\o}ller, Mikkel, & Wolf, Christian K. 2022. \newblock Instrumental Variable Identification of Dynamic Variance Decompositions. \newblock {\em Journal of Political Economy}, {\bf 130}(8), 2164--2202. \bibitem[\citename{Plagborg-M{\o}ller {\em et al.}, }2024]{PMMOQW2024} Plagborg-M{\o}ller, Mikkel, Montiel-Olea, Jos\'e Luis, Qian, Eric, & Wolf, Christian K. 2024. \newblock {\em Double Robustness of Local Projections and Some Unpleasant VARithmetic}. \newblock Tech. rept. 32495. NBER, http://www.nber.org/papers/w32495. \bibitem[\citename{Pope, }1990]{Pope1990} Pope, Alun Lloyd. 1990. \newblock Biases of estimators in multivariate non-Gaussian autoregressions. \newblock {\em Journal of Time Series Analysis}, {\bf 11}(3), 249--258. \bibitem[\citename{Roodman {\em et al.}, }2019]{MacKinnonOR2019} Roodman, David, Nielsen, Morten {\O}rregaard, MacKinnon, James G., & Webb, Matthew D. 2019. \newblock Fast and wild: Bootstrap inference in Stata using boottest. \newblock {\em The Stata Journal}, {\bf 19}(1), 4--60. \bibitem[\citename{Scheff\'e, }1953]{Scheffe1953} Scheff\'e, Henry. 1953. \newblock A method for judging all contrasts in the analysis of variance*. \newblock {\em Biometrika}, {\bf 40}(1-2), 87--110. \bibitem[\citename{Sims {\em et al.}, }1990]{SimsStockWatson1990} Sims, Christopher A., Stock, James H., & Watson, Mark W. 1990. \newblock Inference in linear time series models with some unit roots. \newblock {\em Econometrica}, {\bf 58}(1), 113--144. \bibitem[\citename{Stock & Watson, }2018]{StockWatson2018} Stock, James H., & Watson, Mark W. 2018. \newblock Identification and Estimation of Dynamic Causal Effects in Macroeconomics Using External Instruments. \newblock {\em Economic Journal}, {\bf 128}(610), 917--948. \bibitem[\citename{Tanaka, }2020]{Tanaka2020} Tanaka, Masahiro. 2020. \newblock Bayesian inference of local projections with roughness penalty priors. \newblock {\em Computational Economics}, {\bf 55}(2), 629--651. \bibitem[\citename{Toda & Yamamoto, }1995]{TodaYamamoto1995} Toda, Hiro Y., & Yamamoto, Taku. 1995. \newblock Statistical inference in vector autoregressions with possibly integrated processes. \newblock {\em Journal of Econometrics}, {\bf 66}(1-2), 225--250. \bibitem[\citename{Xu, }2023]{Xu2023} Xu, Ke-Li. 2023. \newblock {\em Local Projection Based Inference under General Conditions}. \newblock Unpublished. \url{https://ssrn.com/abstract=4372388}.