Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
66,771 characters · 16 sections · 88 citation commands
Projection Inference for Set-Identified SVARs.
A Structural Vector Autoregression (SVAR) (Sims:1980) is a time series model that brings theoretical restrictions into a linear, multivariate autoregression. The theoretical restrictions are used to transform the reduced-form parameters of the multivariate autoregression (regression coefficients and the covariance matrix of residuals) into structural parameters that are more amenable to policy interpretation. Depending on the restrictions imposed, the map between reduced-form and structural parameters can be one-to-one (a point-identified SVAR) or one-to-many (a set-identified SVAR).
It is now customary for empirical macroeconomic studies to impose sign and/or equality restrictions on structural dynamic responses in SVARs in order to set-identify the model, as in the pioneering work of Faust:1998 and uhlig:2005. The vast majority of these studies use numerical methods to construct Bayesian posterior credible sets for the coefficients of the structural impulse-response function.
Despite the popularity of the Bayesian approach, a practical concern is the fact that posterior inference for the structural parameters continues to be influenced by prior beliefs even if the sample size is infinite. This point has been documented---in detail and generality---in the work of giacomini_kitagawa:2014.\footnote{See also Poirier:1998, Gustafson:2009, and Moon_Schorfheide:2012} Baumeister_Hamilton:2014 also provide an explicit characterization of the influence of prior beliefs on posterior distributions for structural parameters in set-identified SVARs.
This paper studies the properties of the classical and well-known projection method to conduct simultaneous inference about the coefficients of the structural impulse-response function (and their identified set). The projection method does not rely on the specification of prior beliefs for set-identified parameters. The concrete proposal is to `project' a typical Wald ellipsoid for the reduced-form parameters of a VAR. As we explain below, the suggested nominal $1-\alpha$ projection region consists of all the structural parameters of interest compatible with the reduced-form parameters in a nominal $1-\alpha$ Wald ellipsoid.
Our main result shows that a nominal $1-\alpha$ projection region has---asymptotically and under mild assumptions---both frequentist coverage and robust Bayesian credibility of at least $1-\alpha$. Moreover, building on Kaido_Molinari_Stoye:2014, we show that our baseline projection can be `calibrated\textquoteright\:to eliminate excessive robust Bayesian credibility.
The remainder of the paper is organized as follows. Section (ref) presents an overview of the projection approach. Section (ref) presents the SVAR model and establishes the frequentist coverage of projection. Section (ref) establishes the asymptotic robust Bayesian credibility of the projection region. Section (ref) presents the `calibration\textquoteright\:algorithm designed to eliminate the excess of robust Bayesian credibility. Section (ref) applies projection in the context of the demand/supply SVAR for the U.S. labor market. Section (ref) concludes.
Let $\mu$ denote the parameters of a reduced-form vector autoregression; i.e., the slope coefficients in the regression model and the covariance matrix of residuals. Let $\lambda$ denote the structural parameter of interest; i.e., the response of some variable $i$ to a structural shock $j$ at horizon $k$ (or a vector of responses). In set-identified SVARs there is a known map between $\mu$ and the lower and upper bounds for the components of $\lambda$; see uhlig:2005. Consequently, the smallest and largest value of a particular structural coefficient of interest can be written, simply and succinctly, as $\underline{v}(\mu)$ and $\overline{v}(\mu)$.
Our projection region for $\lambda$ (and for its identified set) is based on a straightforward application of the classical idea of projection inference; see scheffe1953method, dufour1990exact, and Dufour:2005, Dufuour:2007. Let $\widehat{\mu}_{T}$ denote the sample least squares estimator for $\mu$ and let CS$_{T}(1-\alpha; \mu)$ denote its nominal $1-\alpha$ Wald confidence ellipsoid. If, asymptotically, CS$_{T}(1-\alpha; \mu)$ covers the parameter $\mu$ with probability $1-\alpha$, then, asymptotically, the interval
covers the set-identified parameter $\lambda$ (and its identified set) with probability of at least $1-\alpha$ (uniformly over a large class of data generating processes).\footnote{The application of projection inference to SVARs was first suggested by Moon_Schorfheide:2012 (p. 11, NBER working paper 14882). The projection approach is also briefly mentioned in the work Kline_Tamer:2015 (Remark 8) in the context of set-identified models. None of these papers established the properties for projection inference discussed in our work.}
In many applications there is interest in conducting simultaneous inference on $h$ structural parameters; for example, if one wants to analyze the response of variable $i$ to a structural shock $j$ for all horizons ranging from period 1 to $h$ as in Jorda:2009, Kilian_Inoue:2013, Inoue:2014, and Lutkepohl:2015. In this case, one can show that the projection region is given by:
which covers the structural coefficients $(\lambda_1, \ldots \lambda_h )$ and their identified set with probability at least equal to $1-\alpha$ as the sample size grows large. The only assumption required to guarantee the frequentist coverage of our projection region is the asymptotic validity of the confidence set for the reduced-form parameters, $\mu$.
{ General Applicability:} The validity of our projection method requires no regularity assumptions (like continuity or differentiability) on the bounds of the identified set $\underline{v}(\cdot)$ and $\overline{v}(\cdot)$. This means we can handle the typical application of set-identified SVARs in the empirical macroeconomics literature (exclusion restrictions on contemporaneous coefficients, long-run restrictions, elasticity bounds, and of course sign/zero restrictions on the responses of different variables at different horizons for different shocks).
{ Computational Feasibility:} The implementation of our projection approach requires neither numerical inversion of hypothesis tests nor sampling from the space of rotation matrices. Instead, we use state-of-the-art optimization algorithms to solve for the maximum and minimum value of a mathematical program to compute the two end points of the confidence interval in ($\ref{equation:ProjectionCS}$).
{ Robust Bayesian Credibility:} In the spirit of making our results appealing to Bayesian decision makers---and following the seminal work of giacomini_kitagawa:2014---we show that our suggested nominal $1-\alpha$ projection region will have---as the sample size grows large---robust Bayesian credibility of at least $1-\alpha$. This means that the asymptotic posterior probability that the vector of structural parameters of interest belongs to the projection region will be at least $1-\alpha$; for a fixed prior on the reduced-form parameters, $\mu$, and for any prior on the set-identified parameters. A sufficient condition to establish the robust Bayesian credibility of projection is that the prior for $\mu$ used to compute credibility satisfies the Bernstein-von Mises theorem.
{ `Calibrated\textquoteright\:Projection:} Despite the features highlighted above, projection inference is conservative both for a frequentist and a robust Bayesian. Both the asymptotic confidence level and the asymptotic robust credibility of projection can be strictly above $1-\alpha$. Kaido_Molinari_Stoye:2014 [henceforth, KMS] refer to the excess frequentist coverage as projection conservatism and develop an innovative calibration approach to eliminate it.\footnote{Another paper proposing a procedure to eliminate the frequentist excess coverage in moment-inequality models is Bugni_Canay_Shi:2014. Adapting their profiling idea to our set-up could be of theoretical interest and of practical relevance. We leave this question open for future research. } The calibration exercise in KMS requires, in the SVAR context, the computation of Monte-Carlo coverage probabilities for the projection region over an exhaustive grid of values for the reduced-form parameters, $\mu$. In several SVAR applications, the dimension of $\mu$ compromises the construction of an exhaustive grid. Instead of insisting on removing excessive frequentist coverage, we suggest practitioners to calibrate projection to attain a robust Bayesian credibility of exactly $1-\alpha$. Broadly speaking, the calibration consists of drawing $\mu$ from its posterior distribution (or a suitable large-sample Gaussian approximation); evaluating the functions $\underline{v}(\mu), \overline{v}(\mu)$ for each draw of $\mu$; and decreasing the radius defining the projection region until it contains exactly $(1-\alpha)\%$ of the values of $\underline{v}(\mu), \overline{v}(\mu)$ (for different horizons and different shocks if desired).
{ Illustrative Example:} The illustrative example in this paper is a simple demand and supply model of the U.S. labor market. We estimate standard Bayesian credible sets for the dynamic responses of wages and employment using the Normal-Wishart-Haar prior specification in uhlig:2005 and also the alternative prior specification recently proposed by Baumeister_Hamilton:2014. The main set-identifying assumptions are sign restrictions on contemporaneous responses: an expansionary structural demand shock increases wages and employment upon impact; an expansionary structural supply shock decreases wages but increases employment, also upon impact.\footnote{Following Baumeister_Hamilton:2014 we further consider bounds on the wage elasticity of both labor demand and labor supply, and bounds on the long-run impact of a demand shock on employment. }
The Bayesian credible sets for this application illustrate the attractiveness of set-identified SVARs. The data, combined with prior beliefs, and with the (set)-identifying assumptions imply that the initial responses to demand and supply shocks persist in the medium-run, which was not restricted ex-ante. The Bayesian credible sets for this application also illustrate how the quantitative results in set-identified SVARs can be affected by the prior specification. For example, under the prior in Baumeister_Hamilton:2014 the 5-year ahead response of employment to a demand shock could be as large as $4\%$; whereas under the priors in uhlig:2005 the same effect is at most $2\%$.
Our baseline projection approach allows us to get a prior-free assessment about the magnitude (and direction) of the structural responses of interest. For example, the largest value in our projection region for the 5-year response of employment to a demand shock is around 2.5%. This effect is larger than the one implied by the prior in uhlig:2005, but smaller than the one implied by the priors in Baumeister_Hamilton:2014.
Our baseline projection approach---though informative about the effects of demand shocks---is not conclusive about the medium-run effects of structural supply shocks on wages and employment (the projection region allows for both positive and negative responses). This could be a consequence of either the robustness of projection or its conservativeness. To disentangle these effects, we calibrate projection to guarantee that it has exact robust Bayesian credibility. The calibrated projection shows that an expansionary supply shock will decrease wages in each quarter over a 5 year horizon. However, the qualitative effects of supply shocks on employment remain undetermined. The simple SVAR for the labor market illustrates the usefulness of both the baseline and the calibrated projection to analyze the robustness of quantitative and qualitative results in SVARs to prior beliefs.
There continues to be interest in departing from the standard Bayesian analysis of set-identified SVARs in an attempt to provide robustness to the choice of priors. Below we provide a short description of the similarities and differences between our projection approach and three alternative methods available in the literature.
a) In a pioneering paper, Moon-Granziera-Schorfheide:2013 [MSG] proposed a Bonferroni frequentist inference using a moment-inequality, minimum distance framework based on Andrews_Soares:2010.\footnote{An earlier working paper version of MSG also considered a projection of the Andrews_Soares:2010 statistic for moment inequalities for each hypothetical value of the structural matrix $B$. Instead, our paper projects Wald ellipsoid for reduced form parameters $\mu$. Our approach has a computational advantage, since we avoid computationally costly grid search step over values of $B$.} In terms of applicability, their procedures are designed for set-identified SVARs that impose restrictions on the dynamic responses of only one structural shock. It is possible to extend their approach to the same class of models that we consider; there is, however, a serious issue regarding computational feasibility. Specifically, the Bonferroni approach in MSG requires the researcher to compute---by simulation---a critical value for each single orthogonal matrix of dimension $n \times n$, where $n$ is the dimension of the SVAR. Our baseline implementation of the projection method does not require any type of grid over the space of orthogonal matrices and does not require the simulation of any critical value.
b) The seminal paper of giacomini_kitagawa:2014 [GK] develops a novel and generally applicable robust Bayesian approach to conduct inference about a specific coefficient of the impulse-response function in a set-identified SVAR. In terms of our notation, their procedure can be described as follows. One takes posterior draws from $\mu$ and evaluates, at each posterior draw, the functions $\underline{v}(\mu), \overline{v}(\mu)$ by solving a nonlinear program. Their credible set is a numerical (grid-search) approximation to the smallest interval that covers $100(1-\alpha) \%$ of the posterior realizations of the identified set.
Our baseline procedure will be typically faster to implement than the GK robust procedure (since our baseline projection only needs to solve two nonlinear programs). The price to pay for the reduced computational cost is the excess of robust Bayesian credibility. Our calibrated projection requires a similar amount of work as the GK robust method (both procedures evaluate the bounds of the identified set for each posterior draw). Our main contribution relative to giacomini_kitagawa:2014 is that our calibrated projection allows for non-conservative simultaneous credibility statements over different horizons, different variables, and different shocks.
c) GMM:2015 [GMM1] establish the differentiability of the bounds of $\underline{v}(\mu), \overline{v}(\mu)$ for a class of SVAR models that impose restrictions only on the responses to one structural shock. Based on the differentiability results, they propose a `delta-method' confidence interval given by the plug-in estimators of the bounds plus or minus standard errors multiplied by the normal critical value $r$. It can be shown that, in large samples, the `delta-method' procedure in GMM1 is equivalent to a projection region based on a Wald ellipsoid for $\mu$ with radius $r^2$.\footnote{See details in Appendix C of gafarov2025projectioninferencesetidentifiedsvars.}
This paper studies the $n$-dimensional Structural Vector Autoregression with $p$ lags; i.i.d. structural innovations---denoted $\varepsilon_{t}$---distributed according to $F$; and unknown $n \times n$ structural matrix $B$:
see Lutkepohl:2007, p. 362.
The reduced-form parameters of the SVAR model are defined as the vectorized autoregressive coefficients and the half vectorized covariance matrix of reduced-form residuals: $$\mu \equiv ( \text{vec}(A)', \text{vech}(\Sigma)' )^{\prime} \in \mathbb{R}^{d}, \quad \text{where} \quad A \equiv (A_1, A_2, \ldots , A_p), \quad \Sigma \equiv BB'.$$ These parameters can be estimated directly from the data using least squares: $$\widehat{\mu}_{T} \equiv ( \text{vec}(\widehat{A}_{T})', \text{vech}(\widehat{\Sigma}_{T})' )^{\prime},$$ where
with $X_t \equiv (Y_{t-1}', \ldots , Y_{t-p}')'$ and $\widehat{\eta}_t \equiv Y_t - \widehat{A}_{T} X_t$.
A common formula for the asymptotic variance of $\widehat{\mu}_{T}$ in stationary models is: $$\widehat{\Omega}_{T} \equiv V_{T} \Big( \frac{1}{T} \sum_{t=1}^{T} \text{vec} \Big( [\widehat{\eta}_t X_t', \widehat{\eta}_t\widehat{\eta}_t' - \widehat{\Sigma}_{T} ] \Big) \text{vec} \Big( [\widehat{\eta}_t X_t', \widehat{\eta}_t\widehat{\eta}_t' - \widehat{\Sigma}_{T} ] \Big)' V_{T}'$$ where $$ V_T \equiv
,$$ and $L_n$ is the matrix of dimension $n(n+1)/2 \times n^2$ such that $vech(\Sigma)=L_n vec(\Sigma)$, see Lutkepohl:2007, p. 662 equation A.12.1.
The SVAR parameters $(A_1, \ldots , A_p, B, F)$ define a probability measure, denoted $P$, over the data observed by the econometrician. The measure $P$ is assumed to belong to some class $\mathcal{P}$ which we describe in this section.\\
We state a simple high-level assumption concerning the asymptotic behavior of the $1-\alpha$ Wald confidence ellipsoid for $\mu$, which is defined as:
\footnotetext{The radius $\chi^2_{d,1-\alpha}$ denotes the $1-\alpha$ quantile of a central $\chi^2$ distribution with $d$ degrees of freedom.} The first assumption requires uniform consistency in level (over the class $\mathcal{P}$) of the Wald confidence set for the reduced-form parameters. That is:
Assumption (ref) holds if the class $\mathcal{P}$ under consideration contains only uniformly stable VARs where the error distributions, $F$, have uniformly bounded fourth moments.\footnote{A class $\mathcal{P}$ that satisfies Assumption (ref) could be written by using a uniform version of the conditions in Lutkepohl:2007, p. 73. This is, there are positive constants $c_1, c_2, c_3, c_4$ such that:
Other possible definitions of $\mathcal{P}$ can be given by generalizing Theorem 3.5 in Chen2015 to either multivariate linear processes with i.i.d. innovations or to martingale difference sequences. For joint inference on VAR parameters $\mu$ in non-stationary case see gafarov2024wildinferencewildsvars.} Assumption (ref) turns out to be sufficient to conduct frequentist inference on the structural parameters of a set-identified SVAR, defined as follows. \\
{ Coefficients of the Structural Impulse-Response Function:} Given the autoregressive coefficients $A \equiv (A_1, A_2, \ldots , A_p)$ define, recursively, the nonlinear transformation $$C_k(A) \equiv \sum_{m=1}^{k} C_{k-m}(A) \: A_{m}, \quad k \in \mathbb{N},$$ where $C_0= \mathbb{I}_{n}$ and $A_{m}=0$ if $m>p$; see Lut90IRF, p. 116.
In this section we show that, under Assumption (ref), it is possible to `project' the $1-\alpha$ Wald confidence set for $\mu$ to conduct frequentist inference about the coefficients of the structural impulse-response function and the function itself in set-identified models.
{ Set-Identified SVARs:} As mentioned in the introduction, the SVAR allows researchers to transform the reduced-form parameters, $\mu \equiv (\text{vec}(A)',\text{vech}(\Sigma)^{\prime})^{\prime}$, into the structural parameters of interest, $\lambda_{k,i,j}(A,B)$. The parameter $\mu$ determines a unique value of $A$; however, several values of $B$ are compatible with $\Sigma$ (any $B$ such that $BB'=\Sigma$). This indeterminacy of $B$ implies there are multiple values of $\lambda_{k,i,j}(A,B)$ that are compatible with one value of $\mu$.
{{The Identified Set and its Bounds:}} It is common in applied macroeconomic work to impose restrictions on the matrix $B \in \mathbb{R}^{n \times n}$ in order to limit the range of a structural coefficient of interest, $\lambda_{k,i,j}$ (taking $\mu$ as given). Mathematically, a set of restrictions on $B$---that we denote as $\mathcal{R}(\mu)$---can be interpreted as a subset of $\mathbb{R}^{n \times n}$. This leads to the following definition:
Table (ref) presents a list of the most common restrictions, $\mathcal{R}(\mu)$, used in SVAR analysis (all of which can be handled by our frequentist approach described below).\footnote{Our approach can further handle restrictions on the FEVD as in Volpicella2022 and ranking restrictions of amir2021identification. One can also extend our approach to alternative normalization assumptions provided the identified set is bounded, see read2024set.} Throughout the paper, we focus on the case with non-mutually contradictory identification restrictions under probability measure $P$. In other words, the structural restrictions are correctly specified, and the identified set is non-empty. It is not critical for the validity of the inference procedure, but it guarantees non-empty confidence sets in sufficiently large samples. If the identified set is truly empty, the projection confidence set could also be empty, but it remains valid since an empty set is always a subset of any set. An empty confidence set would imply model misspecification. Practitioners could omit some of the constraints to restore non-emptiness.\footnote{In the case of misspecified moment inequalities, 10.1093/restud/rdad033 provides a valid procedure for projection inference on pseudo-parameters. Their procedure can also be adapted to our setup for inference on pseudo-parameters. } Note that under Assumption (ref), we allow for sequences of drifting DGP $P_T$ and the corresponding $\mu_T$ that correspond to non-singleton sets $\mathcal{I}_{H}^{\mathcal{R}} (\mu_T)$, but with a singleton limit as considered in gafarov2014identification. One can also accommodate both strong and weak proxy-IV variables in our setup by explicitly including them in the VAR system and imposing appropriate short run restrictions.
{ Projection Approach:} A key feature of set-identified SVARs is that the bounds of the identified set depend on a finite-dimensional parameter. `Projecting\textquoteright\:down the $1-\alpha$ Wald ellipsoid for $\mu$ seems a natural approach to conduct inference on the structural impulse response function. The first result in this paper establishes the frequentist uniform validity of projection inference.
The idea of `projecting' a confidence set for a parameter $\mu$ to conduct inference about a lower dimensional parameter $\lambda$ has been used extensively in econometrics; see scheffe1953method, dufour1990exact, and Dufour:2005, Dufuour:2007 for some examples. In addition to its conceptual simplicity, one advantage of the projection approach is that its validity does not require special conditions on the identifying restrictions that can be imposed by practitioners. For instance, we do not need to assume that $\underline{v}_{k,i,j}(\cdot)$ and $\overline{v}_{k,i,j}(\cdot)$ are continuous or differentiable functions of the reduced-form parameters.
The problem of conducting inference on the whole impulse-response function (and not only on one specific coefficient) has been a topic of recent interest, both from the Bayesian and frequentist perspective, as exemplified below. For Bayesian set-identified SVARs with only sign restrictions, Kilian_Inoue:2013 report the vector of structural impulse-response coefficients with highest posterior density (based on a prior on reduced-form parameters and a uniform prior on rotation matrices). They propose a Bayesian credible set (represented by shotgun plots) that characterizes the joint uncertainty about a given collection of structural impulse-response coefficients.
For frequentist point-identified SVARs, Inoue:2014 propose a bootstrap procedure that allows the construction of asymptotically valid confidence regions for any subset of structural impulse responses. To the best of our knowledge, our projection approach is the first frequentist procedure for set-identified SVARs that provides confidence regions for any collection of structural coefficients (response of different variables, to different shocks, over different horizons).
It is important to note that uhlig:2005's approach to conduct inference on set-identified SVARs does not provide credible sets for vectors of the structural parameters. The same is true for the Bayesian approaches described in the recent work of Arias:2015 and Baumeister_Hamilton:2014, as well as the approaches of Moon-Granziera-Schorfheide:2013 and giacomini_kitagawa:2014.
A common concern in set-identified models is whether the suggested inference approach is valid only for the identified parameter, $\lambda^H$, or also for its identified set $\mathcal{I}_{H}^{\mathcal{R}} (\mu)$. Note that the second-to-last inequality in the proof of Theorem (ref) imply that our projection region covers the identified set of any vector of coefficients $\lambda^H$.
This section analyzes the robust credibility of projection as the sample size grows large. \\
{ Bayesian Set-up:} In a Bayesian SVAR the distribution of the structural innovations is fixed and treated as a known object. A common choice---which we follow in this section---is to assume that $F \sim \mathcal{N}_n(0, \mathbb{I}_{n})$. We discuss how to relax this restriction after stating Assumption (ref).
Let $P^*$ denote some prior for the structural parameters $(A_1, \ldots ,A_p, B)$ and let $\lambda^{H}(A,B)$ $\in \mathbb{R}^{H}$ denote the vector of structural coefficients of interest. For a given square root of $\Sigma \equiv B B^{\prime}$ define the `rotation\textquoteright\:matrix $Q \equiv \Sigma^{-1/2} B$. It is well known that a prior $P^*$ can be written as $(P^*_{\mu}, P^*_{Q|\mu})$, where $P^*_{\mu}$ is a prior on the reduced-form parameters, and $P^*_{Q|\mu}$ is a prior on the rotation matrix, conditional on $\mu$. Following this notation, let $\mathcal{P}(P^*_{\mu})$ denote the class of prior distributions such that $\mu\sim P^*_{\mu}$.
We are interested in characterizing the smallest posterior probability that the set $\text{CS}_{T}(1-\alpha; \lambda^H)$ could receive, allowing the researcher to vary the prior for $Q$:
The event of interest is whether the structural coefficients $\lambda^H (A,B)$ (treated as random variables in the Bayesian Set-up) belong to the projection region, after conditioning on the data. This event would typically be referred to as the credibility of $\text{CS}_{T}(1-\alpha; \lambda^H)$ (see berger1985statistical, p. 140). We would like to find the smallest credibility of projection when different priors over $Q$ are considered as in the pioneering work of Kitagawa:2012. We follow the recent work of giacomini_kitagawa:2014 and refer to ((ref)) as the robust Bayesian credibility of the set $\text{CS}_T (1-\alpha, \lambda^H)$.
Let $f(Y_1, . . . , Y_T | \mu)$ denote the Gaussian statistical model for the data (which depends solely on the reduced-form parameters) and let $o_{p}(1; Y_1, \ldots Y_T | \mu )$ denote a random variable such that $\lim_{T \rightarrow \infty} P_{Y_1, \ldots, Y_{T}|\mu}(|o_{p}(1; Y_1, \ldots Y_T | \mu )|$ $> \epsilon)=0$ for all $\epsilon>0$ when the distribution of the data is conditioned on $\mu$.
{ Main Assumption for Bayesians:} Robust credibility can be viewed as a random variable (as it depends on $Y_1, \ldots, Y_{T}$). We use the following high-level assumption to characterize its asymptotic behavior:
Assumption (ref) requires the prior over the reduced-form parameters (and the statistical model) to be regular enough to guarantee that the asymptotic Bayesian credibility of the $1-\alpha$ Wald ellipsoid converges in probability to $1-\alpha$. Thus, our high-level assumption is implied by the Bernstein-von Mises Theorem (Dasgupta08, p. 291) for the reduced-form parameter $\mu$.
Since the Gaussian statistical model $f(Y_1, \ldots Y_{T} | \mu_0)$ can be shown to be Locally Asymptotically Normal (LAN) whenever $A_0$ is stable and $\Sigma_0$ has full rank, Theorem 1 and 2 in ghosal1995convergence (GGS) imply that Assumption (ref) will be satisfied whenever $P^*_{\mu}$ has a continuous density at $\mu_0$ with polynomial majorants.\footnote{In Online Appendix S1 we verify an `almost sure\textquoteright\:version of Assumption (ref) for a Gaussian SVAR for the Normal-Wishart priors suggested in Uhlig:1994 and a confidence set for $\mu$ based on the formula for the asymptotic variance $\widehat{\Omega}_{T}$ that obtains in the Gaussian model Lutkepohl:2007. } In fact, the same theorems could be used to establish Assumption (ref) for non-Gaussian SVARs that are LAN and satisfy the regularity conditions of Ibragimov:2013 (IH), as long as $\text{CS}_{T}(1-\alpha; \mu)$ is centered at the Maximum Likelihood estimator of $\mu$ and $\widehat{\Omega}_{T}$ is replaced by the model\textquoteright s inverse information matrix. An alternative approach to establish Assumption (ref) using a different set of primitive conditions can be found in Ben:2016.
We now establish the robust Bayesian credibility of projection as $T \rightarrow \infty$.
This means that---given any prior that satisfies Assumption (ref)---our projection region can be interpreted, in large samples, as a robust $1-\alpha$ credible region for the impulse-response function and its coefficients.
The projection approach generates conservative regions for both a frequentist and a robust Bayesian. For a frequentist, the large-sample coverage may be strictly above the desired confidence level. For a robust Bayesian, the asymptotic robust credibility of the nominal $1-\alpha$ projection region may be strictly above $1-\alpha$.
This section applies the approach in Kaido_Molinari_Stoye:2014 to eliminate the excess of robust Bayesian credibility in a computationally tractable way. We focus on calibrating the robust credibility of our projection region to be exactly equal to $1-\alpha$ (either in a finite sample for a given prior on $\mu$, or in large samples for a large class of priors on $\mu$).
Given a vector $\Lambda^{H}= \{\lambda_{k_h,i_h,j_h}\}_{h=1}^{H}$ of structural coefficients of interest and its corresponding nominal $1-\alpha$ projection region, the calibration exercise is based on the following result:
This means that in order to calibrate the robust credibility of projection, it is sufficient to choose $1-\alpha^*(Y_1, \ldots, Y_T)$ to guarantee that exactly $\alpha\%$ of the bounds of the identified set for the different structural coefficients in $\lambda^H$ fall outside the projection region.
{ Calibration Algorithm:} The calibration algorithm we propose consists in finding a nominal level $1-\alpha^*(Y_1, \ldots, Y_{T})$ such that the posterior probability of the event: $$ [\underline{v}_{k_1,i_1,j_1}(\mu), \overline{v}_{k_1,i_1,j_1}(\mu)] \times \ldots \times [\underline{v}_{k_h,i_h,j_h}(\mu), \overline{v}_{k_h,i_h,j_h}(\mu)] \subseteq \text{CS}_{T}(1-\alpha^*, \lambda^{H}) $$ equals $1-\alpha$ under the posterior distribution associated with the prior $P^*_{\mu}$ or under a suitable large-sample approximation for the posterior, such as $\mu | Y_1, \ldots Y_{T} \sim \mathcal{N}_{d}(\widehat{\mu}_{T}, \widehat{\Omega}_{T}/T)$.\footnote{The Gaussian approximation for the posterior will eliminate projection bias asymptotically, provided a Bernstein-von Mises Theorem for $\mu$ holds. We establish this result in Online Appendix S2. } \\
The calibration algorithm is the following:
It is also possible to show that whenever the bounds of the identified set for each $\lambda_h $ are differentiable, there is a sense in which our calibration algorithm also removes the excess of frequentist coverage (see details in gafarov2025projectioninferencesetidentifiedsvars). \\
This subsection discusses the implementation of the baseline projection region: $$\text{CS}_{T}(1-\alpha; \lambda_{k,i,j}) \equiv \Big[ \inf_{\mu \in \text{CS}_{T}(1-\alpha, \mu)} \underline{v}_{k,i,j}(\mu) \: , \: \sup_{\mu \in \text{CS}_{T}(1-\alpha, \mu)} \overline{v}_{k,i,j}(\mu) \Big].$$ We note that both the upper and lower bounds of this confidence interval can be thought of as solutions to a pair of `nested' optimization problems.
The first optimization problem---that we refer to as the inner optimization---solves for $\overline{v}_{k,i,j}(\mu)$ and $\underline{v}_{k,i,j}(\mu)$. These functions correspond to the largest and smallest value of the structural impulse response $\lambda_{k,i,j}$ given a set of restrictions and a vector of reduced-form parameters $\mu$. The second optimization problem---that we refer to as the outer optimization problem---solves for the maximum value of $\overline{v}_{k,i,j}(\cdot)$ and the minimum value of $\underline{v}_{k,i,j}(\cdot)$ over the $(1-\alpha)$ Wald confidence ellipsoid, CS$_T(1-\alpha,\mu)$.
{ Implementation:} Our proposal is to combine the inner and outer problems into a single mathematical program that gives the bounds of the projection confidence interval directly. The upper bound can be found by solving:
The lower bound of the projection confidence interval can be found analogously. Importantly, the simple reformulation in ((ref)) allows us to base the implementation of our projection region upon state-of-the-art solution algorithms for optimization problems. For most applications, including the one considered in the paper, restrictions $B \in \mathcal{R}(\mu)$ are smooth functions. This allows one to use a simple SQP/IP algorithm (see details in Online Appendix S3). For non-smooth constraints $\mathcal{R}(\mu)$, one can use a more computationally demanding global search algorithm (for example, a genetic algorithm) instead.
As an example, we consider the demand-supply SVAR model studied in Section 5 of Baumeister_Hamilton:2014 [henceforth, BH]. We fit a 6-lag VAR to U.S. data on growth rates of real labor compensation, $\Delta w_t$, and total employment, $\Delta n_t$, from 1970:Q1 to 2014:Q2.\footnote{Our selection is based on the fact that 6 is the smallest number of lags such that CS$(68\%;\mu)$ does not contain unstable VAR coefficients and non-invertible reduced-form covariance matrices. $68\%$ confidence sets are frequently used in applied macroeconomic research. The Bayes Information Criteria and the Akaike Information Criteria both select less than six lags.}
Using our notation, the demand-supply SVAR can be written as: $$
= A_1
+ \ldots + A_6
+ B
,$$ BH set-identify an expansionary demand and supply shock by means of the following sign restrictions: $$ B \equiv
\quad satisfies \quad
.$$ The sign restrictions state that a demand shock increases both real labor compensation and total employment, while a supply shock lowers wages but raises employment.
In this model, the short-run wage elasticity of labor supply (identified from a demand shock) is defined as: $$ \alpha \equiv b_2/b_1$$ Likewise, the short-run wage elasticity of labor demand (identified from a supply shock) is defined as: $$\beta \equiv b_4/b_3 $$ Finally, the long-run impact of a demand shock on employment is given by: $$ \gamma \equiv e_2'(\mathbb{I}_{n}-\sum_{p=1}^{6} A_p )^{-1}Be_1.$$
BH impose three additional restrictions. The first two are elasticity bounds motivated by the findings of different empirical studies. Hamermesh:1996, Akerlof:07, lichter14 provide bounds on the wage elasticity of labor demand. chetty:11, Whalen:2012 provide bounds on the wage elasticity of labor supply. The third and final restriction arises from imposing lower and upper bounds on the long-run impact of a demand shock on employment.
BH incorporate the restrictions in the form of priors on the structural parameters, but we treat the constraints as additional sign restrictions. Let $t_v$ denote the standard $t$ distribution with $v$ degrees of freedom. Table (ref) summarizes the way in which BH incorporate prior information:
Thus, summarizing, our version of the BH model has 10 sign restrictions:
where the parameter $V$ is allowed to take the values $\{.01, .1, 1\}$ as in p. 1992 of BH.
Using our SQP/IP local solution algorithm, we compute the 68% projection confidence intervals for the cumulative response of wages and employment to the structural shocks in the model (20 consecutive quarters and setting $V=1$). In addition to the projection region, we compute the 68% Bayesian credible set following the implementation in both uhlig:2005 and BH.
Figure (ref) shows the projection region as a solid blue line and the standard Bayesian credible set (based on BH priors) as a gray-shaded area. Online Appendix S4 Figure 3 shows the boundaries of the projection region as a solid blue line and the Bayesian credible set based on uhlig:2005's priors as a gray-shaded area. The 68% credible sets differ substantially depending on the specification of prior beliefs. Such sensitivity is the main motivation for our projection approach. In this example, the length of the credible sets for the cumulative response of employment seems to differ by a factor of at least two. The projection region seems quite large compared to the credible sets. This could be a consequence of either the robustness of projection or its conservativeness. To disentangle these effects, we calibrate projection to guarantee that it has exact robust Bayesian credibility in the next subsection.
We investigate the computational feasibility of our projection by comparing its computing time with standard Bayesian methods.\footnote{To get a fair sense of the computational cost, none of the global algorithms were parallelized.} Since the global methods are initialized at the local solution, these procedures take at least as much time as SQP/IP. Among the three global methods considered, the Genetic Algorithm takes the longest. Brute-force grid search (which refers to grid search on CS$_T(1-\alpha, \mu)$ to optimize $\underline{v}_{k,i,j}(\mu)$ and $\overline{v}_{k,i,j}(\mu)$) with only 1,000 draws from $\mu \in \mathcal{\mathbb{R}}^{27}$ takes about 6 times longer than the baseline SQP/IP and generates substantially smaller bounds. We further compare the accuracy across a range of local and global solution methods. For this application, it seems that none of the global algorithms improve on the local solution obtained from SQP/IP. For details, see Online Appendix S3 Table 1 and Figure 2.
The key restriction used to set-identify an expansionary demand shock in the illustrative example is that it must increase wages and employment upon impact. According to the credible sets in Figures (ref) and Online Appendix S4 Figure 3, the demand shock has---in fact---noncontemporaneous effects on these two variables (every quarter over a 5 year horizon). Our calibrated projection confirms that there are medium-run effects of demand shocks on employment but suggests that the non-zero effects on wages beyond the first two quarters could be an artifact of prior beliefs.
A similar observation is true for supply shocks. Our calibrated projection suggests that the decrease in wages five years after an expansionary supply shock is robust to the choice of prior on the set-identified parameters. The medium-run effects of supply shocks on employment lack this robustness.\
{ Implementation of our Calibrated Projection:} We close this subsection providing further details about the computational demands of our calibration exercise.
Instead of working with a specific posterior for $\mu$, we calibrated the projection relying on the large-sample approximation $\mu | Y_1, \ldots ,Y_{T} \sim \mathcal{N}_{d}(\widehat{\mu}_{T}, \widehat{\Omega}_{T}/T)$. Taking draws from this model is straightforward and does not require any special sampling technique (as a Markov Chain Monte-Carlo). Figure (ref) uses M=100,000 draws.
As described in our calibration algorithm, for each of the draws of $\mu$ (denoted $\mu^*_m$), and for each horizon $k\in\{0,1,2, \ldots 20\}$, variable $i \in \{\text{wage},\text{employment}\}$ and shock $j \in \{\text{demand shock},\text{supply shock}\}$ we solved two mathematical programs to generate: $$ [\underline{v}_{k, i, j}(\mu^*_m), \overline{v}_{k, i, j}(\mu^*_m)].$$ Computing the bounds of the identified set for all the combinations $(k,i,j)$ given $\mu^*_m$ took approximately 9 seconds. Generating the boxes and the black dashed lines in Figure (ref) took approximately 5 hours using 50 parallel Matlab `workers' on a computer cluster.\footnote{Calibrating projection to guarantee frequentist coverage at one point in the parameter space took us 76 hours using the 50 parallel Matlab workers in the same computer cluster.} Notice that we chose M=100,000 for illustrative purposes, and the calibration results are barely different for M=1,000, which takes 3 minutes using the same computer cluster (or 2.5 hours without parallelization at all).
After generating the bounds of the identified set, the calibration exercise adjusts the nominal level of projection to simultaneously contain $68\%$ of the draws from the bounds of the identified set for each combination $(k,i,j)$.\footnote{To do this, we ran the baseline projection SQP/IP algorithm for different nominal confidence levels. An efficient calibration algorithm that requires only a few iterations over the nominal level is the combination of bisection with secant and interpolation, as provided by Matlab's fzero function. For a reasonably low tolerance of $\eta=0.001$, we need 15 iteration steps. With each step taking about 734 seconds (see Online Appendix Table 1), steps 3 through 5 take about 1 hour.} The calibrated confidence level for the Wald ellipsoid is $1.85 \cdot 10^{-4}$% instead of the original 68%.
A practical concern regarding standard Bayesian inference for set-identified Structural Vector Autoregressions is the fact that prior beliefs continue to influence posterior inference even when the sample size is infinite. Motivated by this observation, this paper studied the properties of projection inference for set-identified SVARs.
A nominal $1-\alpha$ projection region collects all the structural parameters of interest that are compatible with the VAR reduced-form parameters in a nominal $1-\alpha$ Wald ellipsoid. By construction, projection inference does not rely on the specification of prior beliefs for set-identified parameters.
We showed that---under mild assumptions concerning the asymptotic behavior of estimators and posterior distributions for the reduced-form parameters--- projection produces regions with frequentist coverage and asymptotic robust Bayesian credibility of at least $1-\alpha$.
The main drawback of our projection region is that it is conservative. For a frequentist, the large-sample coverage is strictly above the desired confidence level. For a robust Bayesian, the asymptotic robust credibility of the nominal $1-\alpha$ projection region is strictly above $1-\alpha$.
We used the calibration idea described in Kaido_Molinari_Stoye:2014 to eliminate the excess of robust Bayesian credibility. The calibration procedure consists of drawing the reduced-form parameters, $\mu$, from its posterior distribution (or a suitable large-sample Gaussian approximation); evaluating the functions $\underline{v}(\mu), \overline{v}(\mu)$ for each draw of $\mu$; and, finally, decreasing the nominal level of the projection region until it contains exactly $(1-\alpha)\%$ of the values of $\underline{v}(\mu), \overline{v}(\mu)$. The calibration exercise required more work than the baseline projection, but it is computationally feasible (and easily parallelizable).
We implemented our projection confidence set in a demand/supply SVAR of the U.S. labor market. The main set-identifying assumptions were sign restrictions on contemporaneous responses. Standard Bayesian credible sets suggested that the medium-run response of wages and employment to structural shocks behave similar to the contemporaneous responses. Our projection region (baseline and calibrated) showed that only the qualitative effects of demand shocks on employment and the qualitative effects of supply shocks on wages are robust to the choice of prior. Our projection approach is a natural complement for the Bayesian credible sets that are commonly reported in applied macroeconomic work.