EconBase
← Back to paper

Quasi-Bayesian Local Projections: Simultaneous Inference and Extension to the Instrumental Variable Method

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

22,924 characters · 10 sections · 36 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Quasi-Bayesian Local Projections: Simultaneous Inference and Extension to the Instrumental Variable Method

abstractLocal projections (LPs) are widely used for impulse response analysis, but Bayesian methods face challenges due to the absence of a likelihood function. Existing approaches rely on pseudo-likelihoods, which often result in poorly calibrated posteriors. We propose a quasi-Bayesian method based on the Laplace-type estimator, where a quasi-likelihood is constructed using a generalized method of moments criterion. This approach avoids strict distributional assumptions, ensures well-calibrated inferences, and supports simultaneous credible bands. Additionally, it can be naturally extended to the instrumental variable method. We validate our approach through Monte Carlo simulations.

\paragraph*{Keywords:}

local projections; Bayesian inference; Laplace-type estimator; instrumental variable method; generalized method of moments

Introduction

Local projections (LPs) Jorda2005 relate an outcome variable across horizons to exogenous variables and are primarily used for impulse response analysis. For reviews, see Jorda2023 and Jorda2025.

Although frequentist inference for LPs is well studied, Bayesian approaches remain limited Tanaka2020a,Ferreiraforthcoming. A key challenge is that LPs do not define a likelihood, forcing existing studies to rely on pseudo-likelihoods. Tanaka2020a treats LPs as seemingly unrelated regressions with multivariate normal errors, estimated using Gibbs sampling. While Tanaka2020b demonstrates consistency under a specific data-generating process, the posterior is misspecified and credible intervals can have poor coverage. Ferreiraforthcoming propose a quasi-Bayesian method using Mueller2013's (2013) sandwich estimator. However, their approach relies on asymptotic arguments and equation-by-equation variance estimates, which prevents proper belief updating and constructing simultaneous credible bands for impulse response functions (IRFs).

This study introduces a quasi-Bayesian framework using the Laplace-type estimator (LTE) Kim2002,Chernozhukov2003. The LTE constructs a quasi-likelihood based on the generalized method of moments Hansen1982, avoiding strong distributional assumptions and extending naturally to LPs with instrumental variables (LP-IV) Ramey2018,Stock2018.\footnote{Goh2022 construct the quasi-likelihood for an IV regression in a similar manner.}

The approach offers three advantages. First, the LTE-based quasi-posterior is “well calibrated” with credible intervals aligning with asymptotic theory. Second, it enables simultaneous credible bands MontielOlea2019. Third, it accommodates IV estimation, providing what appears to be the first feasible Bayesian method for LP-IV.\footnote{Huber2024 extend Tanaka2020a's (2020a) pseudo-likelihood approach, inheriting its misspecification and ignoring the correlations between first- and second-stage errors.}

The remainder of the paper is structured as follows. Section 2 presents the LTE framework and posterior analysis and extends the framework to IV estimation; Section 3 conducts a simulation study; and Section 4 concludes.

Quasi-Bayesian Inference of LPs

LPs and the pseudo-posterior

LPs estimate the relationship between a response variable observed at different horizons, $y_{t},y_{t+1},...,y_{t+H}$, and $J$ regressors, $\boldsymbol{x}_{t}$. Specifically, the model is given by \[ y_{t+h}=\boldsymbol{\theta}_{\left(h\right)}^{\top}\boldsymbol{x}_{t}+u_{\left(h\right),t+h},\quad h=0,1,...,H;t=1,...,T, \] where $\boldsymbol{x}_{t}$ includes an exogenous shock, an intercept, lags of $y_{t}$, and controls. The first coefficient, $\theta_{1,\left(h\right)}$, represents the response to the structural shock and its sequence, $\theta_{1,\left(0\right)},...,\theta_{1,\left(H\right)}$, forms the IRF. An alternative long-differenced (LD) specification replaces $y_{t+h}$ with $y_{t+h}-y_{t-1}$, which can reduce bias and improve coverage Piger2025.

Bayesian LP studies Tanaka2020a,Ferreiraforthcoming typically stack LPs into a system of seemingly unrelated regressions with multivariate normal errors: \[ \boldsymbol{y}_{t}=\boldsymbol{\Theta}^{\top}\boldsymbol{x}_{t}+\boldsymbol{u}_{t},\quad\boldsymbol{u}_{t}\sim\mathcal{N}\left(\boldsymbol{0}_{H+1},\boldsymbol{\Sigma}\right), \] with $\boldsymbol{\Theta}=\left(\boldsymbol{\theta}_{\left(0\right)},...,\boldsymbol{\theta}_{\left(H\right)}\right)$. This leads to a pseudo-likelihood in standard form. While the point estimates of $\boldsymbol{\Theta}$ may be consistent, posterior uncertainty is misspecified; therefore, credible intervals lack proper coverage.

Ferreiraforthcoming address this by applying Mueller2013's (2013) sandwich estimator equation by equation. However, their approach faces two key issues: (i) the posterior does not represent a proper belief update, as the contribution of priors to posteriors is not appropriately handled, leading to invalid point and interval estimates; and (ii) they cannot generate simultaneous credible bands because uncertainty is quantified equation by equation. Adjusting the pseudo-likelihood with a learning rate, as in generalized/Gibbs posteriors Martin2022,Wu2023, is a potential remedy but existing selection methods are unsuitable for this.

The quasi-posterior

We propose inferring LPs using the LTE Kim2002,Chernozhukov2003, which is based on moment conditions, $\mathbb{E}\left[\boldsymbol{m}_{t}\left(\boldsymbol{\theta}\right)\right]=\boldsymbol{0}$ with $\boldsymbol{\theta}=\textrm{vec}\left(\boldsymbol{\Theta}\right)$. For LPs, the moment function is defined as \[ \boldsymbol{m}_{t}\left(\boldsymbol{\theta}\right)=\left(

array[array omitted — 329 chars of source]

\right). \] The sample mean is denoted by $\bar{\boldsymbol{m}}_{T}\left(\boldsymbol{\theta}\right)=T^{-1}\sum_{t=1}^{T}\boldsymbol{m}_{t}\left(\boldsymbol{\theta}\right)$. Under standard conditions Hansen1982, $\sqrt{T}\bar{\boldsymbol{m}}_{T}\left(\boldsymbol{\theta}^{\text{true}}\right)$ is asymptotically normal: \[ \sqrt{T}\bar{\boldsymbol{m}}_{T}\left(\boldsymbol{\theta}^{\text{true}}\right)\overset{d}{\rightarrow}\mathcal{N}\left(\boldsymbol{0},\;\boldsymbol{V}\left(\boldsymbol{\theta}^{\text{true}}\right)\right),\quad\text{as }T\rightarrow\infty, \] with covariance $\boldsymbol{V}\left(\boldsymbol{\theta}^{\text{true}}\right)=\mathbb{E}\left[\boldsymbol{m}_{t}\left(\boldsymbol{\theta}\right)\boldsymbol{m}_{t}\left(\boldsymbol{\theta}\right)^{\top}\right]$. A consistent estimator $\hat{\boldsymbol{V}}$ is obtained using the ordinary least squares (OLS) estimator $\hat{\boldsymbol{\theta}}^{\text{OLS}}$. The quasi-likelihood is then \[ \tilde{\mathcal{L}}\left(\mathcal{D}|\boldsymbol{\theta}\right)\propto\exp\left\{ -\frac{T}{2}\bar{\boldsymbol{m}}_{T}\left(\boldsymbol{\theta}\right)^{\top}\hat{\boldsymbol{W}}\bar{\boldsymbol{m}}_{T}\left(\boldsymbol{\theta}\right)\right\} ,\quad\hat{\boldsymbol{W}}=\hat{\boldsymbol{V}}\left(\hat{\boldsymbol{\theta}}^{\text{OLS}}\right)^{-1}. \]

With the prior $p\left(\boldsymbol{\theta}\right)$, posterior draws are generated from $\pi\left(\boldsymbol{\theta}|\mathcal{D}\right)\propto\tilde{\mathcal{L}}\left(\mathcal{D}|\boldsymbol{\theta}\right)p\left(\boldsymbol{\theta}\right)$. Fixing $\hat{\boldsymbol{W}}$ avoids repeated matrix inversions, unlike adaptive weighting schemes Yin2009,Frazier2024, which are computationally demanding and unstable. The posterior mean $\hat{\boldsymbol{\theta}}$ has an asymptotic normal distribution: \[ \sqrt{T}\left(\hat{\boldsymbol{\theta}}-\boldsymbol{\theta}^{\text{true}}\right)\overset{d}{\rightarrow}\mathcal{N}\left(\boldsymbol{0},\boldsymbol{\Omega}^{*}\right),\quad\text{as }T\rightarrow\infty, \] \[ \boldsymbol{\Omega}^{*}=\left(\boldsymbol{G}^{\top}\hat{\boldsymbol{W}}\boldsymbol{G}\right)^{-1}\boldsymbol{G}^{\top}\hat{\boldsymbol{W}}\boldsymbol{V}\left(\boldsymbol{\theta}^{\text{true}}\right)\hat{\boldsymbol{W}}\boldsymbol{G}\left(\boldsymbol{G}^{\top}\hat{\boldsymbol{W}}\boldsymbol{G}\right)^{-1}, \] \[ \boldsymbol{G}=\boldsymbol{I}_{H+1}\otimes\left(-\frac{1}{T}\boldsymbol{X}^{\top}\boldsymbol{X}\right). \] Because of the consistency of the OLS estimator, the covariance of $\hat{\boldsymbol{\theta}}$ is well approximated by $T^{-1}\left(\boldsymbol{G}^{\top}\hat{\boldsymbol{W}}\boldsymbol{G}\right)^{-1}$; therefore, the quasi-likelihood behaves like a proper likelihood when $T$ is large. The covariance $\boldsymbol{\Omega}=T^{-1}\boldsymbol{\Omega}^{*}$ is estimated by replacing $\boldsymbol{V}\left(\boldsymbol{\theta}^{\text{true}}\right)$ with $\boldsymbol{V}\left(\hat{\boldsymbol{\theta}}\right)$.

Compared with Ferreiraforthcoming, our uncertainty quantification also relies on asymptotics but the quasi-posterior is comparatively “well calibrated.” Simulations show that unlike pseudo-posteriors, our approach avoids discrepancies between Bayesian and frequentist measures when priors are weak and the gap diminishes---even with informative priors---as $T$ grows.

For posterior simulation, we reframe the quasi-likelihood as

eqnarray*[eqnarray* omitted — 704 chars of source]

This representation facilitates efficient posterior simulation (see Appendix A.2).

Simultaneous credible band

In our framework, a simultaneous credible band can be computed using the method proposed by MontielOlea2019. A simultaneous $1-\alpha$ credible band for the coefficients of a structural shock, $\boldsymbol{\theta}_{1}=\left(\theta_{1,\left(0\right)},\theta_{1,\left(1\right)},...,\theta_{1,\left(H\right)}\right)$, is given by the Cartesian product of intervals: $\hat{\mathcal{C}}=\prod_{h=0}^{H}\hat{\mathcal{C}}_{\left(h\right)}$, where $\hat{\mathcal{C}}_{h}$ denotes an interval such that the true value of $\boldsymbol{\theta}_{1}$, denoted by $\boldsymbol{\theta}_{1}^{\textrm{true}}$, is asymptotically covered with a probability of at least $1-\alpha$: \[ \liminf_{T\rightarrow\infty}P\left(\boldsymbol{\theta}_{1}^{\textrm{true}}\in\hat{\mathcal{C}}\right)=\liminf_{T\rightarrow\infty}P\left(\theta_{1,\left(h\right)}^{\textrm{true}}\in\hat{\mathcal{C}}_{\left(h\right)}\;\textrm{for }h=0,1,...,H\right)\geq1-\alpha. \] Let $\hat{\boldsymbol{\Omega}}_{1}$ denote the estimated posterior covariance matrix for $\boldsymbol{\theta}_{1}$, obtained by deleting the rows and columns irrelevant to $\boldsymbol{\theta}_{1}$ from the full matrix $\hat{\boldsymbol{\Omega}}$. Let $\hat{\varsigma}_{\left(h\right)}$ denote the point-wise standard error for $\theta_{1,\left(h\right)}$, which is computed as $\hat{\varsigma}_{\left(h\right)}=T^{-\frac{1}{2}}\sqrt{\hat{\boldsymbol{\Omega}}_{1,\left(h,h\right)}}$, where $\hat{\boldsymbol{\Omega}}_{1,\left(h,h\right)}$ denotes the $\left(h+1\right)$th diagonal element of $\hat{\boldsymbol{\Omega}}_{1}$. Given a critical value $c$, a credible band is defined as \[ \hat{\mathcal{B}}\left(c\right)=\hat{B}\left(\hat{Q}_{1-\alpha}\right)=\bigtimes{}_{h=0}^{H}\left[\hat{\theta}_{1,\left(h\right)}-\hat{\varsigma}_{\left(h\right)}c,\;\hat{\theta}_{1,\left(h\right)}+\hat{\varsigma}_{\left(h\right)}c\right]. \] Random vectors $\boldsymbol{e}_{1},...,\boldsymbol{e}_{N}$ are generated from a multivariate normal distribution with mean zero and covariance matrix $\hat{\boldsymbol{\Omega}}_{1}$, \[ \boldsymbol{e}_{i}=\left(e_{i,\left(0\right)},e_{i,\left(1\right)},...,e_{i,\left(H\right)}\right)^{\top}\sim\mathcal{N}\left(\boldsymbol{0}_{H+1},\hat{\boldsymbol{\Omega}}_{1}\right), \] and then $c$ is chosen using the draws: \[ c=q_{\xi}\left(\max_{h=0,1,...,H}\left|\hat{\Omega}_{1,\left(h,h\right)}^{-\frac{1}{2}}e_{i,\left(h\right)}\right|\right), \] where $q_{\xi}\left(\cdot\right)$ denotes the $\xi$-quantile function. The pseudo-code of this procedure is shown in Algorithm A.1 in Appendix A.3.

Extension to the IV method

Our framework can be naturally extended to the IV method Jorda2015,Ramey2016,Stock2018. Let $\boldsymbol{z}_{t}$ denote a vector consisting of an IV and exogenous variables, which may overlap with the covariates $\boldsymbol{x}_{t}$. As we assume that the dimensions of $\boldsymbol{z}_{t}$ and $\boldsymbol{x}_{t}$ are the same, the model is just-identified. An LP-IV is specified as

eqnarray*[eqnarray* omitted — 199 chars of source]

where $u_{t}^{\prime}$ and $u_{\left(h\right),t+h}$ are error terms. We infer $\boldsymbol{\theta}$ using the following moment conditions: \[ \boldsymbol{m}_{t}\left(\boldsymbol{\theta}\right)=\left(

array[array omitted — 329 chars of source]

\right). \] See Stock2018,Rambachan2025 for discussions on the point identification of dynamic causal effects. The covariance of the moment function $\boldsymbol{V}\left(\boldsymbol{\theta}^{\text{true}}\right)$ is estimated using the two-stage least squares estimator.

This approach differs from standard Bayesian IV methods Kleibergen2003,Conley2008,Lopes2014 in two respects. First, it does not require assumptions about the joint distribution of $u_{t}^{\prime}$ and $u_{\left(0\right),t},....,u_{\left(H\right),t+H}$. Second, it avoids explicitly specifying or estimating the first-stage equation.

In this study, we focus solely on the use of external instruments Stock2012,Mertens2013. LPs can also be applied with internal instruments, obtained by estimating another model, such as structural vector autoregressions Ramey2011,Plagborg-Moller2021. However, this two-step approach is not fully compatible with Bayesian analysis because it ignores the uncertainty associated with the first-stage estimation.

Simulation study

LPs

To assess finite-sample properties, we conducted a simulation comparing four approaches. Pseudo-raw uses the pseudo-likelihood of Tanaka2020a with Gibbs sampling and computes credible intervals from the percentiles of the draws. Pseudo-asymp, which corresponds to Ferreiraforthcoming, applies Mueller2013's (2013) sandwich estimator to the same draws. LTE-raw uses draws from the LTE quasi-posterior with Gibbs sampling, while LTE-asymp evaluates credible intervals with the asymptotic covariance estimator. We employed the non-informative (NI) prior and roughness-penalty (RP) prior Tanaka2020a, with shrinkage parameters inferred using standard half-Cauchy priors. Where necessary, a scale-invariant Jeffreys prior is assigned to $\boldsymbol{\Sigma}$, $p\left(\boldsymbol{\Sigma}\right)\propto\left|\boldsymbol{\Sigma}\right|^{-\left(H+2\right)/2}$. Both the standard estimator and the heteroskedasticity- and autocorrelation-robust (HAR) estimator Newey1987 were used to estimate the covariance. We considered $T\in\left\{ 200,500,1000\right\} $, with $H=L=7$, generating 1,000 datasets. A total of 50,000 draws were simulated and the last 40,000 draws were used for the analysis. See Appendices A.1 and A.2 for further details.

Tables A.1 and A.2 in the Appendix report the complete results for the coverage of the point-wise 90% credible interval (P-Coverage). Three findings stand out. First, the HAR covariance estimator led to under-coverage, while the standard estimator performed better, consistent with MontielOlea2021. Second, the LD and level specifications yielded similar results. Third, Pseudo-raw coverage was far from nominal, while Pseudo-asymp, LTE-raw, and LTE-asymp delivered well-calibrated intervals.

Figure 1 displays the results for P-Coverage. We focus on the results for the LD specification and standard covariance estimator. With the NI prior, Pseudo-raw and Pseudo-asymp diverged substantially, while LTE-raw and LTE-asymp were nearly identical and had close to nominal coverage. With the RP prior, Pseudo-raw and Pseudo-asymp exhibited significantly different coverage, implying that the contribution of the prior was not properly accounted for in the posterior simulation. By contrast, the discrepancies between LTE-raw and LTE-asymp were smaller and diminished with larger samples.

figure[figure omitted — 177 chars of source]

We also compare the point and interval estimates across the methods. With the NI prior, the posterior means for Pseudo-raw and LTE-raw were nearly identical (Figure A.2), whereas they diverged with the RP prior, reflecting different prior contributions (Figure A.3). The interval lengths for Pseudo-raw and Pseudo-asymp differed markedly under both priors (Figures A.4 and A.5). By contrast, LTE-raw and LTE-asymp produced nearly identical intervals with the NI prior (Figure A.6), and their differences were smaller and diminished as sample size grew under the RP prior (Figure A.7).

Table 1 summarizes the results for the coverage of the simultaneous 90% credible interval (S-Coverage). The results were similar for LTE-raw and LTE-asymp. Although coverage tended to fall below 90%, it generally approached the nominal level with NI priors or larger samples, consistent with the theory.

table[table omitted — 626 chars of source]

LP-IV

We investigated the finite-sample properties of the proposed LP-IV approach. The simulation setting was essentially identical to that in Section 3.1. The vector $\boldsymbol{z}_{t}$ was the same as $\boldsymbol{x}_{t}$, except that the structural shock (the first entry of $\boldsymbol{x}_{t}$) was replaced with an IV.

Tables A.5-A.8 in the Appendix present the results. As in the case without IVs, the LD and level specifications performed similarly, with no clear dominance, while the standard covariance estimator outperformed the HAR estimator, which tended to underestimate uncertainty. Accordingly, we focus on the results for the LD specification and standard covariance estimator.

Figure 2 shows the results for P-Coverage. Under the NI prior (first row), LTE-raw and LTE-asymp exhibited nearly identical coverage, reaching the nominal level. With the RP prior, their coverage differed but converged as the sample size increased. Table 2 reports the S-Coverage results. With the NI prior, LTE-raw and LTE-asymp again performed almost identically. With the RP prior, their coverage diverged but gradually converged as $T$ increased. Overall, these results suggest that the proposed approach performs well for LP-IV.

figure[figure omitted — 179 chars of source]
table[table omitted — 628 chars of source]

Conclusion

In this study, we introduced a novel quasi-Bayesian approach for inferring LPs. The proposed method offers three main advantages over existing approaches. First, the quasi-posterior based on a generalized method of moments criterion is “well calibrated”: its credible intervals closely match their asymptotic counterparts, ensuring a proper balance between likelihood and prior in posterior estimation. Second, the method enables the estimation of simultaneous credible bands. Third, it naturally extends to IV estimation. While the frequentist literature has made methodological and empirical contributions, this is the first study to infer LP-IV within a Bayesian framework. These advances have broad implications for applied macroeconomics and econometrics, where LPs are increasingly used to study dynamic causal effects. We hope this research will stimulate further methodological development and empirical application.

\includepdf[pages=-]{appendix.pdf}