EconBase
← Back to paper

Identification and Inference Under Narrative Restrictions

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

109,287 characters · 21 sections · 38 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Identification and Inference Under Narrative Restrictions

abstractWe consider structural vector autoregressions subject to `narrative restrictions', which are inequality restrictions on functions of the structural shocks in specific periods. These restrictions raise novel problems related to identification and inference, and there is currently no frequentist procedure for conducting inference in these models. We propose a solution that is valid from both Bayesian and frequentist perspectives by: 1) formalizing the identification problem under narrative restrictions; 2) correcting a feature of the existing (single-prior) Bayesian approach that can distort inference; 3) proposing a robust (multiple-prior) Bayesian approach that is useful for assessing and eliminating the posterior sensitivity that arises in these models due to the likelihood having flat regions; and 4) showing that the robust Bayesian approach has asymptotic frequentist validity. We illustrate our methods by estimating the effects of US monetary policy under a variety of narrative restrictions.

JEL classification: C32, E52

Keywords: Frequentist coverage, global identification, identified set, multiple priors

Introduction

Estimating the dynamic causal effects of structural shocks is a key challenge in macroeconomics. A common approach to this problem is to use a structural vector autoregression (SVAR) with sign or zero restrictions on the model's structural parameters. Recently, a number of papers have augmented these restrictions with restrictions that involve the values of the structural shocks in specific periods. For example, Antolin-Diaz_Rubio-Ramirez_2018 (AR18) propose restricting the signs of structural shocks and their contributions to the change in particular variables in certain historical episodes. Ludvigson, Ma and Ng (2018)\nocite{Ludvigson_Ma_Ng_2018} independently propose restricting the sign or magnitude of the structural shocks in specific periods. A burgeoning empirical literature has adopted similar restrictions, including BenZeev_2018, Furlanetto_Robstad_2019, Cheng_Yang_2020, Inoue_Kilian_2020, Kilian and Zhou (2020a, 2020b)\nocite{Kilian_Zhou_2020a}\nocite{Kilian_Zhou_2020b}, Laumer_2020, Redl_2020, Zhou_2020 and Ludvigson, Ma and Ng (2020)\nocite{Ludvigson_Ma_Ng_2020}. The fact that these restrictions are placed on the shocks rather than the parameters raises novel problems related to identification, estimation and inference. This paper clarifies the nature of these problems and proposes a solution that is valid from both Bayesian and frequentist perspectives.

Henceforth, we refer to any restrictions that can be written as inequalities involving structural shocks in particular periods as `narrative restrictions' (NR). An example of NR are `shock-sign restrictions', such as the restriction in AR18 that the US economy was hit by a positive monetary policy shock in October 1979. This is when the Federal Reserve markedly increased the federal funds rate following Paul Volcker becoming chairman, and is widely considered an example of a positive monetary policy shock (e.g., Romer_Romer_1989). AR18 also consider `historical-decomposition restrictions', such as the restriction that the change in the federal funds rate in October 1979 was overwhelmingly due to a monetary policy shock. This is an inequality restriction that simultaneously constrains the historical decomposition of the federal funds rate with respect to all structural shocks in the SVAR. Other restrictions on the structural shocks also fit into this framework. For example, we additionally consider `shock-rank restrictions', such as the restriction that the monetary policy shock in October 1979 was the largest positive realization of this shock in the sample period.

From a frequentist perspective, NR are fundamentally different from traditional identifying restrictions, such as sign restrictions on impulse responses (e.g., Uhlig_2005). Under normally distributed structural shocks, traditional sign restrictions induce set-identification, because they generate a set-valued mapping from the SVAR's reduced-form parameters to its structural parameters that represents observational equivalence (i.e., an identified set). This set-valued mapping corresponds to the flat region of the structural-parameter likelihood and, by the definition of observational equivalence (e.g., Rothenberg_1971), does not depend on the realization of the data. NR also result in the structural-parameter likelihood possessing flat regions and hence generate a set-valued mapping from the reduced-form parameters to the structural parameters. Crucially, this mapping depends not only on the reduced-form parameters, but also on the realization of the data. The data-dependence of this mapping implies that the standard concept of an identified set does not apply. In turn, this means that: 1) it is unclear whether NR are point- or set-identifying restrictions; and 2) there is no known valid frequentist procedure to conduct inference in these models.\footnote{Ludvigson et al. (2018, in press)\nocite{Ludvigson_Ma_Ng_2018}\nocite{Ludvigson_Ma_Ng_2020} conduct inference using a bootstrap procedure, but its frequentist validity is unknown.}

From a Bayesian perspective, AR18 and the empirical papers that adopt their approach conduct standard (single-prior) Bayesian inference under NR in much the same way as under traditional sign restrictions. However, we highlight two features of this approach that can spuriously affect inference. First, the conditional likelihood used by AR18 to construct the posterior (distribution) implies that, for some types of NR, a component of the prior (distribution) is updated only in the direction that makes the NR unlikely to hold ex ante. This occurs because the numerator of the conditional likelihood -- the likelihood of the reduced-form VAR -- is flat with respect to the orthonormal matrix that maps reduced-form VAR innovations into structural shocks, whereas the denominator -- the ex ante probability that the NR hold -- depends on this matrix. Second, standard Bayesian inference under NR may be sensitive to the choice of prior when the NR yield a likelihood with flat regions. A flat likelihood implies that the conditional posterior of the orthonormal matrix is proportional to its conditional prior whenever the likelihood is nonzero. Posterior inference may therefore be sensitive to the choice of conditional prior for the orthonormal matrix. This is a problem that also occurs in set-identified models under traditional restrictions (e.g., Poirier_1998).

To address the above issues, we study identification under NR and propose a framework for conducting estimation and inference that is potentially appealing to both Bayesians and frequentists. We proceed in four main steps. First, we formalize the identification problem under NR. Second, we propose a simple modification of the existing Bayesian approach that eliminates the source of posterior distortion arising under NR. The modification is to use the unconditional likelihood, rather than the conditional likelihood, to construct the posterior. Third, as a tool for assessing and/or eliminating posterior sensitivity occurring due to the likelihood having flat regions, we propose a robust (multiple-prior) Bayesian approach to estimation and inference. Finally, we show that the robust Bayesian approach has frequentist validity in large samples.

To the best of our knowledge, this is the first paper to formally study identification under general NR. Plagborg-Moller_Wolf_2020b suggest that shock-sign restrictions, in particular, could in principle be recast as an external instrument (or `proxy') and used to point-identify impulse responses in a proxy SVAR or local projection framework. We explore this idea in Appendix (ref) and highlight the potential sensitivity of this approach to the realization of the unrestricted shocks in the time periods that enter the NR. Petterson, Seim and Shapiro (2020)\nocite{Petterson_Seim_Shapiro_2020} derive bounds for a slope parameter in a single equation given restrictions on the plausible magnitude of the residuals, but the restrictions are over the entire sample and the setting is non-probabilistic.

We make two main contributions to the study of identification under NR. First, we provide a necessary and sufficient condition for global identification of an SVAR under NR and show that this condition is satisfied in a simple bivariate example with a single shock-sign restriction. That is, in contrast with traditional sign restrictions, NR may be formally point-identifying despite generating a set-valued mapping from reduced-form to structural parameters in any particular sample. However, this point-identification result does not deliver a point estimator, because the observed likelihood is almost always flat at the maximum. Second, to develop a frequentist-valid procedure for inference, we introduce the notion of a `conditional identified set'. The conditional identified set extends the standard notion of an identified set to a setting where identification is defined in a repeated sampling experiment conditional on the set of observations entering the NR. This provides an interpretation for the set-valued mapping induced by the NR as the set of observationally equivalent structural parameters in such a conditional frequentist experiment.

In terms of inference under NR, this paper makes contributions from both a Bayesian and a frequentist point of view.

The paper's contribution to Bayesian inference is to address the issues associated with the current approach to standard Bayesian inference under NR. First, we advocate using the unconditional likelihood -- the joint probability of observing the data and the NR being satisfied -- when constructing the posterior, rather than the conditional likelihood. Regardless of the type of NR imposed, the unconditional likelihood is flat with respect to the orthonormal matrix that maps reduced-form VAR innovations into structural shocks. This removes the source of posterior distortion that arises due to conditioning on the NR holding. Standard Bayesian inference under the unconditional likelihood requires a simple change to existing computational algorithms. Second, to address posterior sensitivity to the choice of prior, we adapt the robust Bayesian approach of Giacomini_Kitagawa_2020a (GK) to a setting with NR.

In the context of an SVAR under traditional identifying restrictions, the robust Bayesian approach of GK involves decomposing the prior for the structural parameters into a prior for the reduced-form parameters, which is revised by the data, and a conditional prior for the orthonormal matrix given the reduced-form parameters, which is unrevisable. Considering the class of all conditional priors for the orthonormal matrix that are consistent with the identifying restrictions generates a class of posteriors, which can be summarized by a set of posterior means (an estimator of the identified set) and a robust credible region. This removes the source of posterior sensitivity.\footnote{Giacomini, Kitagawa and Read (2019)\nocite{Giacomini_Kitagawa_Read_2019} extend this approach to proxy SVARs where the parameters of interest are set-identified using external instruments.}

We show that this approach can also be used to summarize posterior sensitivity under NR, since the unconditional likelihood at the realized data possesses flat regions and the posterior can therefore be sensitive to the choice of prior, as in standard set-identified models. There are, however, some modifications needed to account for the novel features of the NR. In particular, one cannot use a conditional prior for the orthonormal matrix to impose the NR due to the data-dependent mapping between reduced-form and structural parameters. However, by considering the class of all conditional priors consistent with any traditional identifying restrictions (if present), one can trace out all possible posteriors that are consistent with the traditional restrictions and the NR. This is because traditional restrictions truncate the support of the conditional prior, while NR truncate the support of the likelihood. Consequently, the posterior given any particular conditional prior is only supported on the common support of the conditional prior and the likelihood.

If the researcher has a credible conditional prior, we recommend reporting the standard Bayesian posterior under the unconditional likelihood together with the robust Bayesian output. This allows other researchers to assess the extent to which posterior inference may be driven by prior choice. In the absence of a credible conditional prior, the robust Bayesian output should be reported as an alternative to the standard Bayesian posterior.

The paper's contribution to frequentist inference is to provide an asymptotically valid approach to inference under NR, which, to the best of our knowledge, was not previously available. To explore the asymptotic frequentist properties of our robust Bayesian procedure, we assume a fixed number of NR. This assumption is empirically relevant given that applications typically impose no more than a handful of NR. We provide conditions under which the robust credible region provides asymptotically valid frequentist coverage of the conditional identified set for the impulse response. Since the conditional identified set is guaranteed to include the true impulse response, the robust credible region also provides valid coverage of the true impulse response. Our robust Bayesian approach should therefore appeal to Bayesians as well as frequentists.

We illustrate our methods by estimating the effects of monetary policy shocks in the United States. We find that posterior inferences about the response of output obtained under restrictions based on the October 1979 episode may be sensitive to the choice of conditional prior for the orthonormal matrix. In contrast, under an extended set of restrictions constructed by AR18 based on multiple historical episodes, output falls with high posterior probability following a positive monetary policy shock regardless of the choice of conditional prior. We also estimate the set of output responses that are consistent with the restriction that the monetary policy shock in October 1979 was the largest positive realization of the shock in the sample period. Compared with the extended set of restrictions, this shock-rank restriction results in broadly similar robust posterior inferences about the output response.

\noindentOutline. The remainder of the paper is structured as follows. Section (ref) highlights the econometric issues that arise when imposing NR using a simple bivariate example. Section (ref) describes the general SVAR($p$) framework. Section (ref) formally analyzes identification under NR and introduces the concept of a conditional identified set. Section (ref) discusses how to conduct standard and robust Bayesian inference under NR. Section (ref) explores the frequentist properties of the robust Bayesian approach. Section (ref) contains the empirical application and Section (ref) concludes. The appendices contain proofs and other supplemental material.

\noindentGeneric notation: For the matrix $\mathbf{X}$, $\mathrm{vec}(\mathbf{X})$ is the vectorization of $\mathbf{X}$ and $\mathrm{vech}(\mathbf{X})$ is the half-vectorization of $\mathbf{X}$ (when $\mathbf{X}$ is symmetric). $\mathbf{e}_{i,n}$ is the $i$th column of the $n\times n$ identity matrix, $\mathbf{I}_{n}$. $\mathbf{0}_{n\times m}$ is a $n\times m$ matrix of zeros. $1(.)$ is the indicator function. $\lVert . \rVert$ is the Euclidean norm.

Bivariate example

This section sets out the econometric issues that arise when imposing NR using the simplest possible SVAR as an example. Consider the SVAR($0$) $\mathbf{A}_{0}\mathbf{y}_{t} = \bm{\varepsilon}_{t}$, for $t=1,\ldots,T$, where $\mathbf{y}_{t} = (y_{1t},y_{2t})'$ and $\bm{\varepsilon}_{t} = (\varepsilon_{1t},\varepsilon_{2t})'$ with $\bm{\varepsilon}_{t} \overset{iid}{\sim} N(\mathbf{0}_{2\times 1},\mathbf{I}_{2})$. We abstract from dynamics for ease of exposition, but this is without loss of generality. The orthogonal reduced form of the model reparameterizes $\mathbf{A}_{0}$ as $\mathbf{Q}'\bm{\Sigma}_{tr}^{-1}$, where $\bm{\Sigma}_{tr}$ is the lower-triangular Cholesky factor (with positive diagonal elements) of $\bm{\Sigma} = \mathbb{E}(\mathbf{y}_{t}\mathbf{y}_{t}') = \mathbf{A}_{0}^{-1}\left(\mathbf{A}_{0}^{-1}\right)'$. We parameterize $\bm{\Sigma}_{tr}$ directly as

equation[equation omitted — 162 chars of source]

and denote the vector of reduced-form parameters as $\bm{\phi} = \mathrm{vech}(\bm{\Sigma}_{tr})$. $\mathbf{Q}$ is an orthonormal matrix in the space of $2\times 2$ orthonormal matrices, $\mathcal{O}(2)$:

equation[equation omitted — 335 chars of source]

where the first set is the set of `rotation' matrices and the second set is the set of `reflection' matrices.

Given the `sign normalization' $\mathrm{diag}(\mathbf{A}_{0}) \geq \mathbf{0}_{2\times 1}$, the set of values for $\mathbf{A}_{0}$ that are consistent with the reduced-form parameters in the absence of additional restrictions is

multline[multline omitted — 737 chars of source]

Shock-sign restrictions

Consider the `shock-sign restriction' that $\varepsilon_{1k}$ is nonnegative for some $k \in \left\{1,\ldots,T\right\}$:

equation[equation omitted — 231 chars of source]

Given the realization of the data in period $k$, Equation ((ref)) implies that the restricted structural shock can be written as a function $\varepsilon_{1k}(\theta,\bm{\phi},\mathbf{y}_{k})$. Under the sign normalization and the shock-sign restriction, $\theta$ is restricted to the set

multline[multline omitted — 441 chars of source]

Since $y_{1k}$ and $y_{2k}$ enter the inequalities characterising this set, the shock-sign restriction induces a set-valued mapping from $\bm{\phi}$ to $\theta$ that depends on the realization of $\mathbf{y}_{k}$. For example, if $\sigma_{21} < 0$, $\sigma_{21}y_{1k}-\sigma_{11}y_{2k} > 0$ and $y_{1k} > 0$,

equation[equation omitted — 248 chars of source]

The direct dependence of this mapping on the realization of the data implies that the standard notion of an identified set -- the set of observationally equivalent structural parameter values given the reduced-form parameters -- does not apply. Consequently, it is not obvious whether existing frequentist procedures for conducting inference in set-identified models are valid under NR. Moreover, it is unclear whether the restrictions are, in fact, set-identifying in a formal frequentist sense. We formally analyze identification under NR in Section (ref).

When conducting Bayesian inference, AR18 construct the posterior using the conditional likelihood, which is the likelihood of observing the data conditional on the NR holding. Letting $\mathbf{y}^{T} = (\mathbf{y}_{1}',\ldots,\mathbf{y}_{T}')'$ represent a realization of the random variable $\mathbf{Y}^{T}$, the conditional likelihood is

equation[equation omitted — 392 chars of source]

The numerator in the first term is a function of $\bm{\phi}$ and $\mathbf{y}^{T}$, while the denominator is equal to 1/2, because the marginal distribution of $\varepsilon_{1k}$ is standard normal. The conditional likelihood therefore depends on $\theta$ only through the indicator function $1\left(\varepsilon_{1k}(\theta,\bm{\phi},\mathbf{y}_{k}) \geq 0\right)$. This indicator function truncates the likelihood, with the truncation points depending on $\mathbf{y}_{k}$. To illustrate, the left panel of Figure (ref) plots the likelihood given different realizations of the data drawn from a data-generating process with $\sigma_{21} < 0$ and assuming for simplicity that the econometrician knows $\bm{\phi}$.\footnote{The data-generating process assumes $\mathbf{A}_{0} =

bmatrix[bmatrix omitted — 35 chars of source]

$, which implies that $\theta = \arcsin(0.5\sigma_{22})$ with $\mathbf{Q}$ equal to the rotation matrix. We assume the time series is of length $T=3$ and draw sequences of structural shocks such that $\varepsilon_{1,1} \geq 0$. $T$ is a small number to control Monte Carlo sampling error in the exercises below. The analysis with known $\bm{\phi}$ replicates the situation with a large sample, where the likelihood for $\bm{\phi}$ concentrates at the truth. The assumption that $\bm{\phi}$ is known also facilitates visualizing the likelihood, which otherwise is a function of four parameters.} The conditional likelihood is flat over the region for $\theta$ satisfying the shock-sign restriction and is zero outside this region. The support of the nonzero region depends on the realization of $\mathbf{y}_{k}$.

figure[figure omitted — 649 chars of source]

The flat likelihood function implies that the posterior will be proportional to the prior in the region where the likelihood function is nonzero, and it will be zero outside this region. The standard approach to Bayesian inference in SVARs identified via sign restrictions assumes a uniform (or Haar) prior over $\mathbf{Q}$, as does the approach in AR18.\footnote{See, for example, Uhlig_2005, Rubio-Ram\'{i}rez, Waggoner and Zha (2010)\nocite{Rubio-Ramirez_Waggoner_Zha_2010}, Baumeister_Hamilton_2015 and Arias, Rubio-Ram\'{i}rez and Waggoner (2018)\nocite{Arias_Rubio-Ramirez_Waggoner_2018}.} In the bivariate example, this is equivalent to a prior for $\theta$ that is uniform over the interval $[-\pi,\pi]$. This prior implies that the posterior for $\theta$ is also uniform over the interval for $\theta$ where the likelihood function is nonzero.

The impact impulse response of $y_{1t}$ to a positive standard-deviation shock $\varepsilon_{1t}$ is $\eta \equiv \sigma_{11}\cos\theta$. The right panel of Figure (ref) plots the posterior for $\eta$ induced by a uniform prior over $\theta$ given the same realizations of the data for which the likelihood was plotted in the left panel. The uniform posterior for $\theta$ induces a posterior for $\eta$ that assigns more probability mass to more-extreme values of $\eta$. This highlights that even a `uniform' prior may be informative for parameters of interest, which is also the case under traditional sign restrictions (Baumeister_Hamilton_2015). One difference is that the conditional prior under sign restrictions is never updated by the data, whereas the support and shape of the posterior for $\eta$ under NR may depend on the realization of $\mathbf{y}_{k}$ through its effect on the truncation points of the likelihood, so there may be some updating of the conditional prior by the data. For example, when $\sigma_{21} < 0$, $\sigma_{21}y_{1k}-\sigma_{11}y_{2k} > 0$ and $y_{1k} > 0$,

equation[equation omitted — 243 chars of source]

However, the conditional prior is not updated at values of $\theta$ corresponding to the flat region of the likelihood. Posterior inference about $\eta$ may therefore still be sensitive to the choice of prior, as in standard set-identified SVARs.

Historical-decomposition restrictions

The historical decomposition is the contribution of a particular structural shock to the observed unexpected change in a particular variable over some horizon. The contribution of the first shock to the change in the first variable in the $k$th period is

equation[equation omitted — 191 chars of source]

while the contribution of the second shock is

equation[equation omitted — 184 chars of source]

Consider the restriction that the first structural shock in period $k$ was positive and (in the language of AR18) the `most important contributor' to the change in the first variable, which requires that $|H_{1,1,k}(\theta,\bm{\phi},\mathbf{y}_{k})| \geq |H_{1,2,k}(\theta,\bm{\phi},\mathbf{y}_{k})|$. Under these restrictions and the sign normalization, $\theta$ must satisfy a set of inequalities that depends on $\bm{\phi}$ and $\mathbf{y}_{k}$. As in the case of the shock-sign restriction, this set of restrictions generates a set-valued mapping from $\bm{\phi}$ to $\theta$ that depends on $\mathbf{y}_{k}$.\footnote{See Appendix A for this set of inequalities. It is more difficult to analytically characterize the induced mapping than in the shock-sign example, so we do not pursue this.}

Let $\mathcal{D}(\theta,\bm{\phi},\mathbf{y}_{k}) = 1\{\varepsilon_{1k}(\theta,\bm{\phi},\mathbf{y}_{k}) \geq 0, |H_{1,1,k}(\theta,\bm{\phi},\mathbf{y}_{k})| \geq |H_{1,2,k}(\theta,\bm{\phi},\mathbf{y}_{k})|\}$ represent the indicator function equal to one when the NR are satisfied and equal to zero otherwise, and let $\tilde{\mathcal{D}}(\theta,\bm{\phi},\bm{\varepsilon}_{k}) = 1\{\varepsilon_{1k} \geq 0, |\tilde{H}_{1,1,k}(\theta,\bm{\phi},\varepsilon_{1k})| \geq |\tilde{H}_{1,2,k}(\theta,\bm{\phi},\varepsilon_{2k})|\}$ represent the indicator function for the same event in terms of the structural shocks rather than the data. The conditional likelihood function given the restrictions is then

equation[equation omitted — 405 chars of source]

As in the case of the shock-sign restriction, the numerator of the first term does not depend on $\theta$. In contrast, the probability in the denominator now depends on $\theta$ through the historical decomposition. Intuitively, changing $\theta$ changes the impulse responses of $y_{1t}$ to the two shocks and thus changes the ex ante probability that $|\tilde{H}_{1,1,k}(\theta,\bm{\phi},\varepsilon_{1k})| \geq |\tilde{H}_{1,2,k}(\theta,\bm{\phi},\varepsilon_{2k})|$. The conditional likelihood therefore depends on $\theta$ both through this probability and through the indicator function determining the truncation points of the likelihood. Consequently, the likelihood function is not necessarily flat when it is nonzero.

To illustrate, the left panel of Figure (ref) plots the conditional likelihood evaluated at a random realization of the data satisfying the restrictions using the same data-generating process as above and assuming that $\bm{\phi}$ is known. The probability in the denominator of the conditional likelihood is approximated by drawing 1,000,000 realizations of $\bm{\varepsilon}_{k}$ and computing the proportion of draws satisfying the restrictions at each value of $\theta$. This probability is plotted in the right panel of Figure (ref). The likelihood is again truncated according to a set-valued mapping from $\bm{\phi}$ and $\mathbf{y}_{k}$ to $\theta$, but an important difference from the case with the shock-sign restriction is that the likelihood is no longer flat within the region where it is nonzero. In particular, the conditional likelihood has a maximum at the value of $\theta$ that minimizes the ex ante probability that the NR are satisfied (within the set of values of $\theta$ that are consistent with the restrictions). The posterior for $\theta$ induced by a uniform prior will therefore assign greater posterior probability to values of $\theta$ that yield a lower ex ante probability of satisfying the NR.

figure[figure omitted — 747 chars of source]

If we view the narrative event as a part of the observables and its probability of occurring depends on the parameter of interest, conditioning on the narrative event implies that we are conditioning on a non-ancillary statistic. When conducting likelihood-based inference, conditioning on a non-ancillary statistic is undesirable, because it represents a loss of information about the parameter of interest. The probability that the shock-sign restriction is satisfied is independent of the parameters, so the event that the restriction is satisfied is ancillary. In the case where there is also a restriction on the historical decomposition, the probability that the NR are satisfied depends on $\theta$, so the event that the NR are satisfied is not ancillary. Conditioning on this non-ancillary event results in the likelihood no longer being flat, but the shape of the likelihood is fully driven by the inverse probability of the conditioning event. That is, the loss of information for $\theta$ can be viewed as distorting the shape of the posterior in the sense that the prior is updated toward values of $\theta$ that make the event that the NR are satisfied less likely ex ante. We therefore advocate forming the likelihood without conditioning on the restrictions holding.

The joint (or unconditional) likelihood of observing the data and the NR holding is obtained by multiplying the conditional likelihood by the probability that the NR are satisfied:

equation[equation omitted — 352 chars of source]

Conditional on being nonzero, the unconditional likelihood is flat with respect to $\theta$. The unconditional likelihood depends on $\theta$ only through the points of truncation. To illustrate, Figure (ref) plots the unconditional likelihood given the same realization of the data used to plot the conditional likelihood. As in the case of the shock-sign restriction, the flat unconditional likelihood implies that posterior inference may be sensitive to the choice of prior. We describe our approach to addressing this posterior sensitivity in Section (ref).

General framework

This section describes the general SVAR($p$) and outlines the restrictions that we consider.

SVAR($p$)

Let $\mathbf{y}_{t}$ be an $n\times 1$ vector of endogenous variables following the SVAR($p$) process:

equation[equation omitted — 152 chars of source]

where $\mathbf{A}_{0}$ is invertible and $\bm{\varepsilon}_{t}\overset{iid}{\sim} N(\mathbf{0}_{n\times 1},\mathbf{I}_{n})$ are structural shocks. The initial conditions $(\mathbf{y}_{1-p},...,\mathbf{y}_{0})$ are given. We omit exogenous regressors (such as a constant) for simplicity of exposition, but these are straightforward to include. Letting $\mathbf{x}_{t} = (\mathbf{y}_{t-1}',\ldots,\mathbf{y}_{t-p}')'$ and $\mathbf{A}_{+} = (\mathbf{A}_{1},\ldots,\mathbf{A}_{p})$, rewrite the SVAR($p$) as

equation[equation omitted — 143 chars of source]

$(\mathbf{A}_{0},\mathbf{A}_{+})$ are the structural parameters. The reduced-form VAR($p$) representation is

equation[equation omitted — 118 chars of source]

where $\mathbf{B} = (\mathbf{B}_{1},\ldots,\mathbf{B}_{p})$, $\mathbf{B}_{l}=\mathbf{A}_{0}^{-1}\mathbf{A}_{l}$ for $l=1,\ldots,p$, and $\mathbf{u}_{t} = \mathbf{A}_{0}^{-1}\bm{\varepsilon}_{t} \overset{iid}{\sim} N(\mathbf{0}_{n\times 1},\bm{\Sigma})$ with $\bm{\Sigma} = \mathbf{A}_{0}^{-1}(\mathbf{A}_{0}^{-1})'$. $\bm{\phi} = (\mathrm{vec}(\mathbf{B})',\mathrm{vech}(\bm{\Sigma})')' \in \bm{\Phi}$ are the reduced-form parameters. We assume that $\mathbf{B}$ is such that the VAR($p$) can be inverted into an infinite-order vector moving average (VMA($\infty$)) representation.\footnote{The VAR($p$) is invertible into a VMA($\infty$) process when the eigenvalues of the companion matrix lie inside the unit circle. See Hamilton_1994 or Kilian_Lutkepohl_2017.}

As is standard in the literature that considers set-identified SVARs, we reparameterize the model into its orthogonal reduced form (e.g., Arias et al. (2018)\nocite{Arias_Rubio-Ramirez_Waggoner_2018}):

equation[equation omitted — 158 chars of source]

where $\bm{\Sigma}_{tr}$ is the lower-triangular Cholesky factor of $\bm{\Sigma}$ (i.e. $\bm{\Sigma}_{tr}\bm{\Sigma}_{tr}'=\bm{\Sigma}$) with diagonal elements normalized to be non-negative, $\mathbf{Q}$ is an $n\times n$ orthonormal matrix and $\mathcal{O}(n)$ is the set of all such matrices. The structural and orthogonal reduced-form parameterizations are related through the mapping $\mathbf{B} = \mathbf{A}_{0}^{-1}\mathbf{A}_{+}$, $\bm{\Sigma} = \mathbf{A}_{0}^{-1}(\mathbf{A}_{0}^{-1})'$ and $\mathbf{Q} = \bm{\Sigma}_{tr}^{-1}\mathbf{A}_{0}^{-1}$ with inverse mapping $\mathbf{A}_{0} = \mathbf{Q}'\bm{\Sigma}_{tr}^{-1}$ and $\mathbf{A}_{+} = \mathbf{Q}'\bm{\Sigma}_{tr}^{-1}\mathbf{B}$.

The VMA($\infty$) representation of the model is

equation[equation omitted — 198 chars of source]

where $\mathbf{C}_{h}$ is the $h$th term in $(\mathbf{I}_{n} - \sum_{l=1}^{p}\mathbf{B}_{l}L^{l})^{-1}$ and $L$ is the lag operator. $\mathbf{C}_{h}$ is defined recursively by $\mathbf{C}_{h} = \sum_{l=1}^{\min\{k,p\}}\mathbf{B}_{l}\mathbf{C}_{h-l}$ for $h \geq 1$ with $\mathbf{C}_{0} = \mathbf{I}_{n}$. The $(i,j)$th element of the matrix $\mathbf{C}_{h}\bm{\Sigma}_{tr}\mathbf{Q}$, which we denote by $\eta_{i,j,h}(\bm{\phi},\mathbf{Q})$, is the horizon-$h$ impulse response of the $i$th variable to the $j$th structural shock:

equation[equation omitted — 187 chars of source]

where $\mathbf{c}_{i,h}'(\bm{\phi}) = \mathbf{e}_{i,n}'\mathbf{C}_{h}\bm{\Sigma}_{tr}$ is the $i$th row of $\mathbf{C}_{h}\bm{\Sigma}_{tr}$ and $\mathbf{q}_{j} = \mathbf{Q}\mathbf{e}_{j,n}$ is the $j$th column of $\mathbf{Q}$.

Narrative restrictions

In the absence of any identifying restrictions, it is well-known that $\mathbf{Q}$ is set-identified. Consequently, functions of $\mathbf{Q}$, such as the impulse responses, are also set-identified. Imposing traditional identifying restrictions on the SVAR is equivalent to restricting $\mathbf{Q}$ to lie in a subspace of $\mathcal{O}(n)$. It is conventional to impose a `sign normalization' on the structural shocks. We normalize the diagonal elements of $\mathbf{A}_{0}$ to be non-negative, so a positive value of $\varepsilon_{it}$ is a positive shock to the $i$th equation in the SVAR at time $t$. The sign normalization implies that $\mathrm{diag}(\mathbf{Q}'\bm{\Sigma}_{tr}^{-1}) \geq \mathbf{0}_{n\times 1}$.

It is common to impose sign restrictions on the impulse responses (e.g., Uhlig_2005) or on the structural parameters themselves. For example, the restriction that the horizon-$h$ impulse response of the $i$th variable to the $j$th shock is nonnegative is $c_{i,h}'(\bm{\phi})\mathbf{q}_{j} \geq 0$, which is a linear inequality restriction on a single column of $\mathbf{Q}$ that depends only on the reduced-form parameter $\bm{\phi}$. Restrictions on elements of $\mathbf{A}_{0}$ take a similar form.

In contrast, NR constrain the values of the structural shocks in particular periods. The structural shocks are

equation[equation omitted — 121 chars of source]

The shock-sign restriction that the $i$th structural shock at time $k$ is positive is

equation[equation omitted — 227 chars of source]

We can treat $\mathbf{u}_{t}$ as observable given $\bm{\phi}$ and the data, so we suppress the dependence of $\mathbf{u}_{t}$ on $\bm{\phi}$ and $(\mathbf{y}_{t}',\mathbf{x}_{t}')'$ for notational convenience. The restriction in ((ref)) is a linear inequality restriction on a single column of $\mathbf{Q}$. In contrast with traditional sign restrictions, the shock-sign restriction depends directly on the data through the reduced-form VAR innovations.

In addition to shock-sign restrictions, AR18 consider restrictions on the historical decomposition, which is the cumulative contribution of the $j$th shock to the observed unexpected change in the $i$th variable between periods $k$ and $k+h$:

equation[equation omitted — 387 chars of source]

One example of a restriction on the historical decomposition is that the $j$th structural shock was the `most important contributor' to the change in the $i$th variable between periods $k$ and $k+h$, which requires that $|H_{i,j,k,k+h}| \geq \max_{l \neq j} |H_{i,l,k,k+h}|$. Another example is that the $j$th structural shock was the `overwhelming contributor' to the change in the $i$th variable between periods $k$ and $k+h$, which requires that $|H_{i,j,k,k+h}| \geq \sum_{l \neq j} |H_{i,l,k,k+h}|$. From Equation ((ref)), it is clear that these restrictions are nonlinear inequality constraints that simultaneously constrain every column of $\mathbf{Q}$ and that depend on the realizations of the data in particular periods in addition to the reduced-form parameters.

Other restrictions also naturally fit into this framework. For instance, we can consider restrictions on the relative magnitudes of a particular structural shock in different periods. We refer to these restrictions as `shock-rank restrictions', since they imply a (possibly partial) ordering of the shocks. As an example, one could impose that the $i$th shock in period $k$ was the largest positive realization of this shock in the observed sample. This requires that $\varepsilon_{ik}(\bm{\phi},\mathbf{Q},\mathbf{u}_{k}) \geq \max_{t\neq k}\{\varepsilon_{it}(\bm{\phi},\mathbf{Q},\mathbf{u}_{t})\}$, which can be expressed as a system of $T-1$ linear inequality restrictions on a single column of $\mathbf{Q}$: $(\bm{\Sigma}_{tr}^{-1}(\mathbf{u}_{k}-\mathbf{u}_{t}))'\mathbf{q}_{i} \geq 0$ for $t \neq k$. Alternatively, one could impose that the $i$th shock in period $k$ was the largest-magnitude realization of that shock, or $|\varepsilon_{ik}(\bm{\phi},\mathbf{Q},\mathbf{u}_{k})| \geq \max_{t\neq k}\left\{|\varepsilon_{it}(\bm{\phi},\mathbf{Q},\mathbf{u}_{t})|\right\}$. If $\varepsilon_{ik}(\bm{\phi},\mathbf{Q},\mathbf{u}_{k}) \geq 0$, this would require that $(\bm{\Sigma}_{tr}^{-1}(\mathbf{u}_{k}-\mathbf{u}_{t}))'\mathbf{q}_{i} \geq 0$ and $(\bm{\Sigma}_{tr}^{-1}(\mathbf{u}_{k}+\mathbf{u}_{t}))'\mathbf{q}_{i} \geq 0$ for $t \neq k$, which is a system of $2(T-1)$ linear inequalities constraining $\mathbf{q}_{i}$. These restrictions could also be applied to a subset of the observations rather than the full sample (e.g., $\varepsilon_{ik}(\bm{\phi},\mathbf{Q},\mathbf{u}_{k}) > \varepsilon_{it}(\bm{\phi},\mathbf{Q},\mathbf{u}_{t})$ for some $t \in \{1,\ldots,T\}$).\footnote{Similar to the shock-rank restrictions we describe, BenZeev_2018 imposes a restriction on the timing of the maximum three-year average of a particular shock, as well as restrictions on the sign and relative magnitudes of this three-year average in specific periods. Restrictions on averages of shocks can also be implemented in the framework we consider.}

The collection of NR can be represented in the general form $N(\bm{\phi},\mathbf{Q},\mathbf{Y}^{T}) \geq \mathbf{0}_{s\times 1}$, where $s$ is the number of restrictions. As an illustration, consider the case where there is a single shock-sign restriction in period $k$, $\varepsilon_{1k}(\bm{\phi},\mathbf{Q},\mathbf{u}_{k}) \geq 0$, as well as the restriction that the first structural shock was the most important contributor to the change in the first variable in period $k$. Then,

equation[equation omitted — 416 chars of source]

Traditional sign and zero restrictions can also be applied alongside NR. We follow AR18 by explicitly allowing for sign restrictions on impulse responses and on elements of $\mathbf{A}_{0}$. We denote such sign restrictions by $S(\bm{\phi},\mathbf{Q}) \geq \mathbf{0}_{\tilde{s}\times 1}$, where $\tilde{s}$ is the number of traditional sign restrictions. It is straightforward to additionally allow for zero restrictions, including `short-run' zero restrictions (as in Sims_1980), `long-run' zero restrictions (as in Blanchard_Quah_1989), or restrictions arising from external instruments (as in Mertens_Ravn_2013 and Stock_Watson_2018); for example, see GK and Giacomini et al. (2019)\nocite{Giacomini_Kitagawa_Read_2019}.

Conditional and unconditional likelihoods

When constructing the posterior of the SVAR's parameters, AR18 use the likelihood conditional on the NR holding. Define

align*[align* omitted — 492 chars of source]

The likelihood conditional on $D_N = 1$ can be written as

equation[equation omitted — 203 chars of source]

$f(\mathbf{y}^T|\bm{\phi})$ is the joint density of the data given $\bm{\phi}$ (i.e., the likelihood function of the reduced-form VAR), which depends only on $\bm{\phi}$ and the data. The indicator function $D_N(\bm{\phi},\mathbf{Q},\mathbf{y}^T)$ is equal to one when the NR are satisfied and is equal to zero otherwise. This determines the truncation points of the likelihood. $r(\bm{\phi},\mathbf{Q})$ is the ex ante probability that the NR are satisfied. This will be a constant when there are only shock-sign or shock-rank restrictions; for example, if there are $s$ shock-sign restrictions, $r(\bm{\phi},\mathbf{Q}) = (1/2)^{s}$. In contrast, when there are restrictions on the historical decomposition, this probability will depend on $\bm{\phi}$ and $\mathbf{Q}$.

Consider the case where $\bm{\phi}$ is known, which will be the case asymptotically because $\bm{\phi}$ is point-identified. When $r(\bm{\phi},\mathbf{Q})$ depends on $\mathbf{Q}$, the conditional likelihood will be maximized at the value of $\mathbf{Q}$ that minimizes $r(\bm{\phi},\mathbf{Q})$ (within the set of values of $\mathbf{Q}$ that satisfy the restrictions). The posterior based on this likelihood will therefore place higher posterior probability on values of $\mathbf{Q}$ that result in a lower ex ante probability that the restrictions are satisfied. As discussed in Section (ref), this is an artefact of conditioning on a non-ancillary event, which represents a loss of information about the parameters.

We therefore advocate constructing the likelihood without conditioning on the NR holding. The unconditional likelihood (the joint distribution of the data and $D_N$) can be expressed as

align[align omitted — 455 chars of source]

For any value of $\bm{\phi}$ such that $\mathbf{y}^{T}$ is compatible with the NR, there will be a set of values of $\mathbf{Q}$ that satisfy the restrictions, which depend on the data, but the value of the unconditional likelihood will be the same for all values of $\mathbf{Q}$ within this set. The conditional posterior of $\mathbf{Q}|\bm{\phi},\mathbf{y}^{T}$ will therefore be proportional to the conditional prior for $\mathbf{Q}|\bm{\phi}$ in these regions. Given a fixed number of NR, the likelihood will possess flat regions even with a time-series of infinite length, so posterior inference may be sensitive to the choice of conditional prior for $\mathbf{Q}$, even asymptotically (which is also the case for the conditional likelihood when the restrictions are ancillary). This motivates considering Bayesian inferential procedures that are robust to the choice of unrevisable conditional prior for $\mathbf{Q}$, which we explore in Section (ref).

Discussion

In this section, we briefly discuss the distributional assumptions for the structural shocks and the mechanism that generates the NR.

Distributional assumptions

Practitioners may be concerned about the robustness of inference with respect to deviations from the assumption of standard normal shocks. For instance, one could worry that the periods in which the NR are imposed are `unusual' in the sense that the structural shocks in these periods were drawn from a distribution with, say, inflated variance or fat tails. The unconditional likelihood depends on the normality assumption only through $f(\mathbf{y}^T|\bm{\phi})$. By omitting terms in $f(\mathbf{y}^T|\bm{\phi})$ corresponding to the periods in which the NR are imposed, one can conduct inference that is robust to the distributional assumption about the shocks in these particular periods. To illustrate, consider the case where NR are imposed in period $k$ only and assume the likelihood function for $\mathbf{y}^{T}$ takes the form

equation[equation omitted — 159 chars of source]

where

equation[equation omitted — 291 chars of source]

and $w(\mathbf{y}_{k} - \mathbf{Bx}_k)$ is an unknown, potentially non-normal, density. Replacing $f(\mathbf{y}^T|\bm{\phi})$ in Equation ((ref)) with $v(\left\{\mathbf{y}_{t} - \mathbf{Bx}_t\right\}_{t\neq k} | \bm{\phi})$ yields an `unconditional partial likelihood' that does not depend on the distribution of $\bm{\varepsilon}_{k}$, but that is still truncated by the NR. This would potentially result in a loss of information relative to a likelihood that correctly specifies the distribution of the shocks in period $k$. However, when NR are imposed in only a few periods, this loss of information is likely to be small. In contrast, the conditional likelihood approach cannot leave fully unspecified the distribution of the restricted structural shocks, because computing $r(\bm{\phi},\mathbf{Q})$ requires specifying this distribution.

Concerns about heteroscedasticity or non-normality may also be alleviated by recognizing that the distributional assumption will become irrelevant asymptotically. The set of values of $\mathbf{Q}$ with non-zero unconditional likelihood depends only on $\bm{\phi}$, which summarizes the second moments of the data, and the realization of the data in the periods in which the NR are imposed. Under regularity assumptions, the likelihood (and thus the posterior) of $\bm{\phi}$ will converge to a point at the true value of $\bm{\phi}$ asymptotically regardless of whether the true data-generating process is a VAR with homoscedastic normal shocks.\footnote{See Plagborg-Moller_2019 for a discussion of this point in the context of a structural VMA model.} The set of values of $\mathbf{Q}$ with non-zero likelihood will therefore converge asymptotically to the same set regardless of whether the distributional assumption is correct.

Mechanism generating NR

Note that we do not explicitly model the mechanism responsible for revealing the information underlying the NR (i.e., whether $D_{N} = 1$ or $D_{N} = 0$) or the mechanism determining the periods in which this information is revealed (e.g., the identity of $k$ in examples above), which is consistent with the papers that impose these restrictions. If the revelation of this information depends on the data, the likelihood will be misspecified. The exact implications of this misspecification for estimation or inference will depend on assumptions about the mechanism revealing the narrative information. Exploring the consequences of such misspecification may be an interesting area for further work. In the bivariate example of Section (ref), if the identity of $k$ is randomly determined independently of $\bm{\varepsilon}_{1},\ldots,\bm{\varepsilon}_{T}$, we can interpret the current analysis conditional on $k$.

Identification under NR

This section formally analyzes identification in the SVAR under NR. Section (ref) considers whether NR are point- or set-identifying in a frequentist sense. Section (ref) introduces the notion of a `conditional identified set', which extends the standard notion of an identified set to the setting where the mapping from reduced-form to structural parameters depends on the realization of the data. This provides an interpretation of the mapping induced by the NR. Additionally, we make use of this object when showing the frequentist validity of our robust Bayesian procedure in Section (ref).

Point-identification under NR

Denoting the true parameter value by $(\bm{\phi}_0, \textbf{Q}_0)$, point-identification for the parametric model ((ref)) requires that there is no other parameter value $(\bm{\phi},\mathbf{Q}) \neq (\bm{\phi}_0,\mathbf{Q}_0)$ that is observationally equivalent to $(\bm{\phi}_0,\mathbf{Q}_0)$.\footnote{$(\bm{\phi},\mathbf{Q}) \neq (\bm{\phi}_0,\mathbf{Q}_0)$ is observationally equivalent to $(\bm{\phi}_0,\mathbf{Q}_0)$ if $p(\mathbf{Y}^{T},D_N=d | \bm{\phi},\mathbf{Q}) = p(\mathbf{Y}^{T},D_N=d | \bm{\phi}_0,\mathbf{Q}_0)$ holds for all $\mathbf{Y}^T$ and $d \in \{0,1 \}$.}

To assess the existence or non-existence of observationally equivalent parameter points, we analyze a statistical distance between $p(\mathbf{y}^{T},D_N=d | \bm{\phi},\mathbf{Q})$ and $p(\mathbf{y}^{T},D_N=d | \bm{\phi}_0,\mathbf{Q}_0)$ that metrizes observation equivalence. Specifically, in the current setting where the support of the distribution of observables can depend on the parameters, it is convenient to work with the Hellinger distance:

align[align omitted — 555 chars of source]

and $\mathbf{Y}$ is the sample space for $\mathbf{Y}^{T}$. As is known in the literature on minimum distance estimation (see, for example, Basu, Shioya and Park (2011)\nocite{Basu_Shioya_Park_2011}), $(\bm{\phi},\mathbf{Q})$ and $(\bm{\phi}_0,\mathbf{Q}_0)$ are observationally equivalent if and only if $HD(\bm{\phi},\mathbf{Q}) = 0$ or, equivalently, $\mathcal{H}(\bm{\phi},\mathbf{Q}) = 1$.

We similarly define the Hellinger distance for the conditional likelihood as

align[align omitted — 375 chars of source]

The next proposition analyzes the conditions for $\mathcal{H}(\bm{\phi}, \mathbf{Q})=1$ and $\mathcal{H}_c(\bm{\phi}, \mathbf{Q})=1$, and shows that observational equivalence of $(\bm{\phi},\mathbf{Q})$ and $(\bm{\phi}_0,\mathbf{Q}_0)$ boils down to geometric equivalence of the set of reduced-form VAR innovations satisfying the NR.

propositionLet $(\bm{\phi}_0, \mathbf{Q}_0)$ be the true parameter value and let $\mathbf{U} \equiv \mathbf{U}(\mathbf{y}^{T};\bm{\phi}) = (\mathbf{u}_{1}',\ldots,\mathbf{u}_{T}')'$ collect the reduced-form VAR innovations. Define \begin{equation} \mathcal{Q}^{\ast} \equiv \left\{ \begin{matrix} \mathbf{Q} \in \mathcal{O}(n) : \{ \mathbf{U}:N(\bm{\phi}, \mathbf{Q}, \mathbf{Y}^{T}) \geq \mathbf{0}_{s \times 1} \} = \{ \mathbf{U}:N(\bm{\phi}_0, \mathbf{Q}_0, \mathbf{Y}^{T}) \geq \mathbf{0}_{s \times 1} \} \\ up to $f(\mathbf{Y}^T| \bm{\phi}_0)$-null set, \ \mathrm{diag}( \mathbf{Q}'\bm{\Sigma}_{tr}^{-1}) \geq \mathbf{0}_{n \times 1} \end{matrix} \right\}. \notag \end{equation} The unconditional likelihood model ((ref)) and the conditional likelihood model ((ref)) are globally identified (i.e., there are no observationally equivalent parameter points to $(\bm{\phi}_0, \mathbf{Q}_0)$) if and only if $\mathcal{Q}^{\ast}$ is a singleton. If the parameter of interest is an impulse response to the $j$th structural shock, $\eta_{i,j,h}(\bm{\phi}, \mathbf{Q})$, as defined in ((ref)), then $\eta_{i,j,h}(\bm{\phi}, \mathbf{Q})$ is point-identified if the projection of $\mathcal{Q}^{\ast}$ onto its $j$th column vector is a singleton.
proofSee Appendix B.

This proposition provides a necessary and sufficient condition for global identification of SVARs by NR. As shown in the proof in Appendix B, $\mathcal{Q}^{\ast}$ defined in this proposition corresponds to the observationally equivalent $\mathbf{Q}$ matrices given $\bm{\phi} = \bm{\phi}_0$, but, importantly, it does not correspond to any flat region of the observed likelihood (the conditional identified set in Definition (ref) below).

To illustrate this point, consider the simple bivariate example of Section 2 with the NR ((ref)), where $\mathbf{y}_t$ itself is the reduced-form error, so $\mathbf{U}$ in Proposition (ref) can be set to $\mathbf{y}_k$. Given $\bm{\phi}$, the set of $\mathbf{y}_k \in \mathbb{R}^2$ satisfying the NR is the half-space given by

equation[equation omitted — 248 chars of source]

The condition for point-identification shown in Proposition (ref) is satisfied if no $\theta' \neq \theta$ can generate the half-space of $\mathbf{y}_k$ identical to ((ref)). Such $\theta'$ cannot exist, since a half-space passing through the origin $(a_1,a_2) \mathbf{y}_k \geq 0$ can be indexed uniquely by the slope $a_1/a_2$ and ((ref)) implies the slope $\sigma_{11}^{-1}(\sigma_{22} (\tan \theta)^{-1} - \sigma_{21})$ is a bijective map of $\theta$ on a constrained domain due to the sign normalization. Figure (ref) plots the Hellinger distances in this bivariate example under the shock-sign restriction ((ref)) and the historical decomposition restriction. For both the conditional and unconditional likelihood, the Hellinger distances are minimized uniquely at the true $\theta$, which is consistent with our point-identification claim for $\theta$.\footnote{Under the restriction on the historical decomposition, a notable difference between the conditional and unconditional likelihood cases is the slope of the Hellinger distance around the minimum. The Hellinger distance of the unconditional likelihood yields a steeper slope than the conditional likelihood. This indicates the loss of information for $\theta$ in the conditional likelihood due to conditioning on the non-ancillary event.}

figure[figure omitted — 423 chars of source]

Proposition (ref) also provides conditions under which $(\bm{\phi},\mathbf{Q})$ is not globally identified, but a particular impulse response is. To give an example of this, consider an SVAR with $n > 2$ and with a shock-sign restriction on the first shock in period $k$. Given $\bm{\phi}$, the set of $\mathbf{u}_{k} \in \mathbb{R}^{n}$ satisfying the NR is a half-space defined by $\mathbf{q}_{1}'\bm{\Sigma}_{tr}^{-1}\mathbf{u}_{k} \geq 0$. The set of values of $\mathbf{u}_{k}$ satisfying this inequality is indexed uniquely by $\mathbf{q}_{1}$ given $\bm{\Sigma}_{tr}$ at its true value, so there are no values of $\mathbf{Q}$ that are observationally equivalent to $\mathbf{Q}_{0}$ with $\mathbf{q}_{1} \neq \mathbf{Q}_{0}\mathbf{e}_{1,n}$. Any value for the remaining $n-1$ columns of $\mathbf{Q}$ such that they are orthogonal to $\mathbf{Q}_{0}\mathbf{e}_{1,n}$ will generate the same half-space for $\mathbf{u}_{k}$, so $\mathcal{Q}^{\ast}$ is not a singleton and the SVAR is not globally identified. However, the projection of $\mathcal{Q}^{\ast}$ onto its first column is a singleton, so $\eta_{i,1,h}(\bm{\phi}, \mathbf{Q})$ is globally identified.

Although a single NR can deliver global identification in the frequentist sense, the practical implication of this theoretical claim is not obvious. The observed unconditional likelihood is almost always flat at the maximum, so we cannot obtain a unique maximum likelihood estimator for the structural parameter. As a result, the standard asymptotic approximation of the sampling distribution of the maximum likelihood estimator is not applicable. The SVAR model with NR possesses features of set-identified models from the Bayesian standpoint (i.e., flat regions of the likelihood). However, strictly speaking, it can be classified as a globally identified model in the frequentist sense when the condition of Proposition (ref) holds.

Conditional identified set

It is well-known that traditional sign restrictions deliver set-identification of $\mathbf{Q}$ (or, equivalently, the structural parameters). Given the reduced-form parameter $\bm{\phi}$ -- which is point-identified -- there are multiple observationally equivalent values of $\mathbf{Q}$, in the sense that there exists $\mathbf{Q}$ and $\tilde{\mathbf{Q}} \neq \mathbf{Q}$ such that $p(\mathbf{y}^{T}|\bm{\phi},\mathbf{Q}) = p(\mathbf{y}^{T}|\bm{\phi},\tilde{\mathbf{Q}})$ for every $\mathbf{y}^{T}$ in the sample space. The identified set for $\mathbf{Q}$ given $\bm{\phi}$ contains all such observationally equivalent parameter points, and is defined as

equation[equation omitted — 260 chars of source]

The identified set is a set-valued map only of $\bm{\phi}$, which carries all the information about $\mathbf{Q}$ contained in the data.

The complication in applying this definition of the identified set in SVARs when there are NR is that the reduced-form VAR parameters no longer represent all information about $\mathbf{Q}$ contained in the data; by truncating the likelihood, the realizations of the data entering the NR contain additional information about $\mathbf{Q}$. To address this, we introduce a refinement of the definition of an identified set.

definitionLet $N \equiv N(\bm{\phi},\mathbf{Q},\mathbf{y}^{T}) \geq \mathbf{0}_{s\times 1}$ represent a set of NR in terms of the parameters and the data. (i) The conditional identified set for $\mathbf{Q}$ under NR is \begin{equation} \mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N) = \{\mathbf{Q} \in \mathcal{O}(n): N(\bm{\phi},\mathbf{Q},\mathbf{y}^{T}) \geq \mathbf{0}_{s\times 1}\}. \end{equation} The conditional identified set for the impulse response $\eta = \eta_{i,j,h}(\bm{\phi},\mathbf{Q})$ under NR is defined by projecting $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ via $\eta_{i,j,h}(\bm{\phi},\mathbf{Q})$: \begin{equation} CIS_{\eta}(\bm{\phi}|\mathbf{y}^{T},N) = \{\eta_{i,j,h}(\bm{\phi},\mathbf{Q}) : \mathbf{Q} \in \mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N) \}. \end{equation} (ii) Let $\mathbf{s}: \mathbf{Y} \to \mathbb{R}^S$ be a statistic. We call $\mathbf{s}(\mathbf{Y}^T)$ a sufficient statistic for the conditional identified set $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ if the conditional identified set for $\mathbf{Q}$ depends on the sample $\mathbf{y}^T$ through $\mathbf{s}(\mathbf{y}^T)$; i.e., there exists $\tilde{\mathcal{Q}}(\bm{\phi}|\cdot,N)$ such that \begin{equation} \mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N) = \tilde{\mathcal{Q}}(\bm{\phi}|\mathbf{s}(\mathbf{y}^{T}),N) \end{equation} holds for all $\bm{\phi} \in \bm{\Phi}$ and $\mathbf{y}^T \in \mathbf{Y}$.

Unlike the standard identified set $\mathcal{Q}(\bm{\phi}|S)$, the conditional identified set $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ depends on the sample $\mathbf{y}^T$ because of the aforementioned data-dependent support of the likelihood. In terms of the observed likelihood, however, they share the property that the likelihood is flat on the (conditional) identified set. Hence, given the sample $\mathbf{y}^{T}$ and the reduced-form parameters $\bm{\phi}$, all values of $\mathbf{Q}$ in $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ fit the data equally well and, in this particular sense, they are observationally equivalent.

When the NR concern shocks in only a subset of the time periods in the data, the conditional identified set under these NR depends on the sample only through a few observations entering the NR. The sufficient statistics $\textbf{s}(\mathbf{y}^T)$ defined in Definition (ref)(ii) represent such observations. For instance, in the toy example of Section (ref), the conditional identified set depends only on the observations in period $k$, so $\mathbf{s}(\mathbf{y}^T) =\mathbf{y}_k$. If we extend the example of Section (ref) to the SVAR($p$), the shock-sign restriction in Equation ((ref)) can be expressed as

equation[equation omitted — 202 chars of source]

Hence, the conditional identified set $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ depends on the data only through $(\mathbf{y}_k^{\prime}, \mathbf{x}_k^{\prime})'= (\mathbf{y}_k^{\prime},\mathbf{y}_{k-1}^{\prime}, \cdots, \mathbf{y}_{k-p}^{\prime} )'$, so we can set $\mathbf{s}(\mathbf{y}^T)=(\mathbf{y}_k^{\prime},\mathbf{y}_{k-1}^{\prime}, \cdots, \mathbf{y}_{k-p}^{\prime} )'$.

If the conditional distribution of $\mathbf{Y}^T$ given $\mathbf{s}(\mathbf{Y}^T) = \mathbf{s}(\mathbf{y}^T)$ is nondegenerate, we can consider a frequentist experiment (repeated sampling of $\mathbf{Y}^T$) conditional on the sufficient statistics set to the observed value. In this conditional experiment, we can view the conditional identified set $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ as the standard identified set in set-identified models, since it no longer depends on the data in the conditional experiment where $\mathbf{s}(\mathbf{y}^{T})$ is fixed. This is the reason that we refer to $\mathcal{Q}(\bm{\phi}|\mathbf{y}^{T},N)$ as the conditional identified set. In Section (ref) below, we show the frequentist validity of the robust-Bayes credible region by establishing conditional coverage of the conditional identified set for an impulse response.

Posterior inference under NR

This section presents approaches to conducting posterior inference in SVARs under NR. Section (ref) discusses how to modify the standard Bayesian approach in AR18 to use the unconditional likelihood rather than the conditional likelihood. Section (ref) explains how to conduct robust Bayesian inference under NR, which further addresses the issue of posterior sensitivity due to the flat unconditional likelihood. Section (ref) describes how to numerically implement the robust Bayesian procedure.

Standard Bayesian inference

AR18 propose an algorithm for drawing from the uniform-normal-inverse-Wishart posterior of $(\bm{\phi},\mathbf{Q})$ given a set of traditional sign restrictions and NR. This is the posterior induced by a normal-inverse-Wishart prior over $\bm{\phi}$ and an unconditionally uniform prior over $\mathbf{Q}$. The algorithm proceeds by drawing $\bm{\phi}$ from a normal-inverse-Wishart distribution and $\mathbf{Q}$ from a uniform distribution over $\mathcal{O}(n)$, and checking whether the restrictions are satisfied. If the restrictions are not satisfied, the joint draw is discarded and another draw is made. If the restrictions are satisfied, the ex ante probability that the NR are satisfied at the drawn parameter values is approximated via Monte Carlo simulation. Once the desired number of draws are obtained satisfying the restrictions, the draws are resampled with replacement using as importance weights the inverse of the probability that the NR are satisfied.\footnote{Based on the results in Arias_Rubio-Ramirez_Waggoner_2018, AR18 argue that their algorithm draws from a normal-generalized-normal posterior over the SVAR's structural parameters $(\mathbf{A}_{0},\mathbf{A}_{+})$ induced by a conjugate normal-generalized-normal prior, conditional on the restrictions.}

This algorithm essentially draws from the posterior under the unconditional likelihood and then uses importance sampling to transform these draws into draws from the posterior given the conditional likelihood. To draw from the uniform-normal-inverse-Wishart posterior using the unconditional likelihood to construct the posterior, one therefore simply needs to omit the importance-sampling step from this algorithm. Approximating the probability used to construct the importance weights requires Monte Carlo integration, which can be computationally expensive, particularly when the NR constrain the structural shocks in multiple periods. Omitting the importance-sampling step can therefore ease the computational burden of drawing from the posterior. However, as discussed above, standard Bayesian inference under the unconditional likelihood may be sensitive to the choice of conditional prior for $\mathbf{Q}|\bm{\phi}$, because the likelihood possesses flat regions.

By rejecting draws that do not satisfy the restrictions, the algorithm described above places more weight on draws of $\bm{\phi}$ that are less likely to satisfy the restrictions under the uniform distribution over $\mathcal{O}(n)$. As discussed in Uhlig_2017, one may instead prefer to use a prior that is conditionally uniform over $\mathbf{Q}|\bm{\phi}$. To draw from the posterior of $(\bm{\phi},\mathbf{Q})$ under the unconditional likelihood given an arbitrary prior over $\bm{\phi}$ and a conditionally uniform prior over $\mathbf{Q}|\bm{\phi}$, one can repeat Step 2 of Algorithm 1 in Section (ref).

Robust Bayesian inference

This section explains how to conduct robust Bayesian inference about a scalar-valued function of the structural parameters under NR and traditional sign restrictions. The approach can be viewed as performing global sensitivity analysis to assess whether posterior conclusions are robust to the choice of prior on the flat regions of the likelihood. We assume that the object of interest is a particular impulse response $\eta$, although the discussion in this section also applies to any other scalar-valued function of the structural parameters, such as the forecast error variance decomposition or the historical decomposition.

Let $\pi_{\bm{\phi}}$ be a prior over the reduced-form parameter $\bm{\phi} \in \bm{\Phi}$, where $\bm{\Phi}$ is the space of reduced-form parameters such that $\mathcal{Q}(\bm{\phi}|S)$ is non-empty. A joint prior for $(\bm{\phi},\mathbf{Q}) \in \bm{\Phi}\times \mathcal{O}(n)$ can be written as $\pi_{\bm{\phi},\mathbf{Q}} = \pi_{\mathbf{Q}|\bm{\phi}}\pi_{\bm{\phi}}$, where $\pi_{\mathbf{Q}|\bm{\phi}}$ is supported only on $\mathcal{Q}(\bm{\phi}|S)$. When there are only traditional identifying restrictions, $\pi_{\mathbf{Q}|\bm{\phi}}$ is not updated by the data, because the likelihood function is not a function of $\mathbf{Q}$. Posterior inference may therefore be sensitive to the choice of conditional prior, even asymptotically. As discussed above, a similar issue arises under NR. The difference under NR is that $\pi_{\mathbf{Q}|\bm{\phi}}$ is updated by the data through the truncation points of the unconditional likelihood. However, at each value of $\bm{\phi}$, the unconditional likelihood is flat over the set of values of $\mathbf{Q}$ satisfying the NR. Consequently, the conditional posterior for $\mathbf{Q}|\bm{\phi},\mathbf{Y}^{T}$ is proportional to the conditional prior for $\mathbf{Q}|\bm{\phi}$ at each $\bm{\phi}$ whenever the conditional identified set for $\mathbf{Q}$ given $(\bm{\phi}, \mathbf{Y}^T)$ is nonempty.

Rather than specifying a single prior for $\mathbf{Q}|\bm{\phi}$, the robust Bayesian approach of GK considers the class of all priors for $\mathbf{Q}|\bm{\phi}$ that are consistent with the traditional identifying restrictions:

equation[equation omitted — 169 chars of source]

Notice that we cannot impose the NR using a particular conditional prior on $\mathbf{Q}|\bm{\phi}$ due to the data-dependent mapping from $\bm{\phi}$ to $\mathbf{Q}$. However, by considering all possible conditional priors for $\mathbf{Q}|\bm{\phi}$ that are consistent with the traditional identifying restrictions, we trace out all possible conditional posteriors for $\mathbf{Q}|\bm{\phi},\mathbf{Y}^{T}$ that are consistent with the traditional identifying restrictions and the NR. This is because the NR truncate the unconditional likelihood function and the traditional identifying restrictions truncate the prior for $\mathbf{Q}|\bm{\phi}$, so the posterior for $\mathbf{Q}|\bm{\phi},\mathbf{Y}^{T}$ is supported only on the values of $\mathbf{Q}$ that satisfy both sets of restrictions.

Given a particular prior for $(\bm{\phi},\mathbf{Q})$ and using the unconditional likelihood, the posterior is

align[align omitted — 416 chars of source]

The final expression for the posterior makes it clear that any prior for $\mathbf{Q}|\bm{\phi}$ that is consistent with the traditional identifying restrictions is in effect further truncated by the NR (through the likelihood) once the data are realized. Generating this posterior using every prior within the class of priors for $\mathbf{Q}|\bm{\phi}$ generates a class of posteriors for $(\bm{\phi},\mathbf{Q})$:

equation[equation omitted — 324 chars of source]

Marginalizing each posterior in this class of posteriors induces a class of posteriors for $\eta$, $\Pi_{\eta|\mathbf{Y}^{T},D_{N} = 1}$. Each prior within the class of priors $\Pi_{\mathbf{Q}|\bm{\phi}}$ therefore induces a posterior for $\eta$. Associated with each of these posteriors are quantities such as the posterior mean, median and other quantiles. For example, as we consider each possible prior within $\Pi_{\mathbf{Q}|\bm{\phi}}$, we can trace out the set of all possible posterior means for $\eta$. This will always be an interval, so we can summarize this `set of posterior means' by its endpoints:

equation[equation omitted — 211 chars of source]

where $l(\bm{\phi},\mathbf{Y}^{T}) = \inf \{\eta(\bm{\phi},\mathbf{Q}):\mathbf{Q} \in \mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)\}$, $u(\bm{\phi},\mathbf{Y}^{T}) = \sup \{\eta(\bm{\phi},\mathbf{Q}):\mathbf{Q} \in \mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)\}$ and

equation[equation omitted — 147 chars of source]

is the set of values of $\mathbf{Q}$ that are consistent with the traditional identifying restrictions and the NR. In contrast, in GK the set of posterior means is obtained by finding the infimum and supremum of $\eta(\bm{\phi},\mathbf{Q})$ over $\mathcal{Q}(\bm{\phi}|S)$ and averaging these over $\pi_{\bm{\phi}|\mathbf{Y}^{T}}$. The important difference from GK is that the current set of posterior means depends on the data not only through the posterior for $\bm{\phi}$ but also through the set of admissible values of $\mathbf{Q}$ under the NR. As a result, in contrast with GK, we cannot interpret the set of posterior means ((ref)) as a consistent estimator for the identified set for $\eta$ (which is not well-defined, as we discussed above). Nevertheless, the set of posterior means still carries a robust Bayesian interpretation similar to GK in that it clarifies posterior results that are robust to the choice of prior on the non-updated part of the parameter space (i.e., on the flat regions of the likelihood).

As in GK, we can also report a robust credible region with credibility level $\alpha$, which is the shortest interval estimate for $\eta$ such that the posterior probability put on the interval is greater than or equal to $\alpha$ uniformly over the posteriors in $\Pi_{\eta|\mathbf{Y}^{T}, D_{N} = 1}$ (see Proposition 1 of GK). One may also be interested in posterior lower and upper probabilities, which are the infimum and supremum, respectively, of the probability for a hypothesis over all posteriors in the class.

GK provide conditions under which their robust Bayesian approach has a valid frequentist interpretation, in the sense that the robust credible region is an asymptotically valid confidence set for the true identified set. For the same reason as mentioned above, however, frequentist validity of the robust credible region does not immediately extend to the NR case. We provide conditions under which the robust credible region has a valid frequentist interpretation in Section (ref).

Numerical implementation of robust Bayesian approach

This section describes a general algorithm to implement our robust Bayesian procedure under NR. GK propose numerical algorithms for conducting robust Bayesian inference in SVARs identified using traditional sign and zero restrictions. Their Algorithm 1 uses a numerical optimization routine to obtain the lower and upper bounds of the identified set at each draw of $\bm{\phi}$. Obtaining the bounds via numerical optimization is not generally applicable under the class of NR considered in AR18, since the constraints on the historical decomposition are not differentiable everywhere in $\mathbf{Q}$. We therefore adapt Algorithm 2 of GK, which approximates the bounds of the identified set at each draw of $\bm{\phi}$ using Monte Carlo simulation.

Algorithm 1. Let $N(\bm{\phi},\mathbf{Q},\mathbf{Y}^{T}) \geq \mathbf{0}_{s\times 1}$ be the set of NR and let $S(\bm{\phi},\mathbf{Q}) \geq \mathbf{0}_{\tilde{s}\times 1}$ be the set of traditional sign restrictions (excluding the sign normalization). Assume the object of interest is $\eta_{i,j^{*},h} = c_{i,h}'(\bm{\phi})\mathbf{q}_{j^{*}}$.

itemize• Step 1: Specify a prior for $\bm{\phi}$, $\pi_{\bm{\phi}}$, and obtain the posterior $\pi_{\bm{\phi}|\mathbf{Y}^{T}}$. • Step 2: Draw $\bm{\phi}$ from $\pi_{\bm{\phi}|\mathbf{Y}^{T}}$ and check whether $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$ is empty using the subroutine below. \begin{itemize} • Step 2.1: Draw an $n\times n$ matrix of independent standard normal random variables, $\mathbf{Z}$, and let $\mathbf{Z} = \tilde{\mathbf{Q}}\mathbf{R}$ be the QR decomposition of $\mathbf{Z}$.\footnote{This is the algorithm used by Rubio-Ramirez_Waggoner_Zha_2010 to draw from the uniform distribution over $\mathcal{O}(n)$, except that we do not normalize the diagonal elements of $\mathbf{R}$ to be positive. This is because we impose a sign normalization based on the diagonal elements of $\mathbf{A}_{0} = \mathbf{Q}'\bm{\Sigma}_{tr}^{-1}$ in Step 2.2.} • Step 2.2: Define \begin{equation*} \mathbf{Q} = \left[\mathrm{sgn}((\bm{\Sigma}_{tr}^{-1}\mathbf{e}_{1,n})'\tilde{\mathbf{q}}_{1})\frac{\tilde{\mathbf{q}}_{1}}{\lVert \tilde{\mathbf{q}}_{1}\rVert},\ldots, \mathrm{sgn}((\bm{\Sigma}_{tr}^{-1}\mathbf{e}_{n,n})'\tilde{\mathbf{q}}_{n})\frac{\tilde{\mathbf{q}}_{n}}{\lVert \tilde{\mathbf{q}}_{n}\rVert}\right], \end{equation*} where $\tilde{\mathbf{q}}_{j}$ is the $j$th column of $\tilde{\mathbf{Q}}$. • Step 2.3: Check whether $\mathbf{Q}$ satisfies $\mathbf{N}(\bm{\phi},\mathbf{Q},\mathbf{Y}^{T}) \geq \mathbf{0}_{s\times 1}$ and $S(\bm{\phi},\mathbf{Q}) \geq \mathbf{0}_{\tilde{s}\times 1}$. If so, retain $\mathbf{Q}$ and proceed to Step 3. Otherwise, repeat Steps 2.1 and 2.2 (up to a maximum of $L$ times) until $\mathbf{Q}$ is obtained satisfying the restrictions. If no draws of $\mathbf{Q}$ satisfy the restrictions, approximate $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$ as being empty and return to Step 2. \end{itemize} • Step 3: Repeat Steps 2.1--2.3 until $K$ draws of $\mathbf{Q}$ are obtained. Let $\{\mathbf{Q}_{k}, k=1,...,K\}$ be the $K$ draws of $\mathbf{Q}$ that satisfy the restrictions and let $\mathbf{q}_{j^{*},k}$ be the $j^{*}$th column of $\mathbf{Q}_{k}$. Approximate $[l(\bm{\phi},\mathbf{Y}^{T}),u(\bm{\phi},\mathbf{Y}^{T})]$ by $[\min_{k}\mathbf{c}_{i,h}'(\bm{\phi})\mathbf{q}_{j^{*},k}$, $\max_{k}\mathbf{c}_{i,h}'(\bm{\phi})\mathbf{q}_{j^{*},k}]$. • \textbf{Step 4}: Repeat Steps 2--3 $M$ times to obtain $[l(\bm{\phi_{m}},\mathbf{Y}^{T}),u(\bm{\phi_{m}},\mathbf{Y}^{T})]$ for $m=1,...,M$. Approximate the set of posterior means using the sample averages of $l(\bm{\phi_{m}},\mathbf{Y}^{T})$ and $u(\bm{\phi_{m}},\mathbf{Y}^{T})$. • \textbf{Step 5}: To obtain an approximation of the smallest robust credible region with credibility $\alpha \in (0,1)$, define $d(\eta,\bm{\phi},\mathbf{Y}^{T}) = \max\{|\eta-l(\bm{\phi},\mathbf{Y}^{T})|,|\eta-u(\bm{\phi},\mathbf{Y}^{T})|\}$ and let $\hat{z}_{\alpha}(\eta)$ be the sample $\alpha$-th quantile of $\{d(\eta,\bm{\phi_{m}},\mathbf{Y}^{T}), m=1,...,M\}$. An approximated smallest robust credible interval for $\eta_{i,j^{*},h}$ is an interval centered at $\arg \min_{\eta} \hat{z}_{\alpha}(\eta)$ with radius $\min_{\eta}\hat{z}_{\alpha}(\eta)$.

}

Algorithm 1 approximates $[l(\bm{\phi},\mathbf{Y}^{T}),u(\bm{\phi},\mathbf{Y}^{T})]$ at each draw of $\bm{\phi}$ via Monte Carlo simulation. The approximated set will be too narrow given a finite number of draws of $\mathbf{Q}$, but the approximation error will vanish as the number of draws goes to infinity. The algorithm may be computationally demanding when the restrictions substantially truncate $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$, because many draws of $\mathbf{Q}$ from $\mathcal{O}(n)$ may be rejected at each draw of $\bm{\phi}$. However, the same draws of $\mathbf{Q}$ can be used to compute $l(\bm{\phi},\mathbf{Y}^{T})$ and $u(\bm{\phi},\mathbf{Y}^{T})$ for different objects of interest, which cuts down on computation time. For example, the same draws of $\mathbf{Q}$ can be used to compute the impulse responses of all variables to all shocks at all horizons of interest. They can also be used to compute other parameters by replacing $\eta_{i,j^{*},h}$ with some other function, such as the forecast error variance decomposition, an element of $\mathbf{A}_{0}$, the historical decomposition or the structural shocks themselves in particular periods.\footnote{Impulse responses to a unit shock -- rather than a standard-deviation shock -- can be computed as in Algorithm 3 of Giacomini et al. (2019)\nocite{Giacomini_Kitagawa_Read_2019}.} Step 3 is parallelizable, so reductions in computing time are possible by distributing computation across multiple processors. Other algorithms may be computationally more efficient than Algorithm 1 in particular cases. We discuss these in Appendix (ref).

Frequentist coverage under a few NR

In this section, we show that the robust Bayes credible region attains asymptotically valid frequentist coverage in a setting where the number of NR is small relative to the length of the sampled periods in a sense that we make precise in the next assumption. This assumption is empirically relevant given that applications typically impose these restrictions in at most a handful of periods.

assumption(fixed-dimensional $\mathbf{s}(\mathbf{Y}^T)$): The conditional identified set under NR has sufficient statistics $\mathbf{s}(\mathbf{Y}^T)$, as defined in Definition (ref)(ii), and the dimension of $\mathbf{s}(\mathbf{Y}^T)$ does not depend on $T$.

Let $(\bm{\phi}_{0},\mathbf{Q}_{0})$ be the true parameter values. We view the sample $\mathbf{Y}^T$ as being drawn from $p(\mathbf{Y}^T|\bm{\phi}_0)$. Let $p(\mathbf{Y}^T|\bm{\phi}_0, \mathbf{s})$ be the conditional distribution of the sample $\mathbf{Y}^T$ given the sufficient statistics for the conditional identified set $\mathbf{s}= \mathbf{s}(\mathbf{Y}^T)$ at $\bm{\phi}= \bm{\phi}_0$. We denote by $p(\mathbf{s}|\bm{\phi}_0)$ the distribution of the sufficient statistics $\mathbf{s}(\mathbf{Y}^T)$ at $\bm{\phi}=\bm{\phi}_0$. The next assumption assumes that in the conditional experiment given $\mathbf{s}(\mathbf{Y}^T)$, the sampling distribution for the maximum likelihood estimator $\hat{\bm{\phi}} \equiv \arg \max_{\bm{\phi}} p(\mathbf{Y}^T|\bm{\phi})$ centered at $\bm{\phi}_0$ and the posterior for $\bm{\phi}$ centered at $\hat{\bm{\phi}}$ asymptotically coincide.

assumption(Conditional Bernstein-von Mises property for $\bm{\phi}$): For $p(\mathbf{s}|\bm{\phi}_0)$-almost every $\mathbf{s}$ and $p(\mathbf{Y}^T|\bm{\phi}_0, \mathbf{s})$-almost every sampling sequence $\mathbf{Y}^T$, the posterior for $\sqrt{T}(\bm{\phi} - \hat{\bm{\phi}})$ asymptotically coincides with the sampling distribution of $\sqrt{T}(\hat{\bm{\phi}} - \bm{\phi}_0)$ with respect to $p(\mathbf{Y}^T|\bm{\phi}_0, \mathbf{s})$, as $T \to \infty$, in the sense stated in Assumption 5(i) in GK.

This is a key assumption for establishing the asymptotic frequentist validity of the robust credible region under NR. It holds, for instance, when $\mathbf{s}(\mathbf{y}^T)$ corresponds to one or a few observations in the whole sample, as we had in the toy example of Section (ref). In this case, the influence of $\mathbf{s}(\mathbf{y}^T)$ vanishes in the conditional sampling distribution of $\sqrt{T}(\hat{\bm{\phi}} - \bm{\phi}_0)$ as $T \to \infty$, as the latter asymptotically agrees with the asymptotically normal sampling distribution for the maximum likelihood estimator with variance-covariance matrix given by the inverse of the Fisher information matrix. By the well-known Bernstein-von Mises theorem for regular parametric models, the posterior for $\sqrt{T}(\bm{\phi} - \hat{\bm{\phi}})$ asymptotically agrees with this sampling distribution.

The last assumption requires convexity and smoothness of the conditional identified set, and is analogous to Assumption 5(ii) of GK for standard set-identified models.

assumption(Almost-sure convexity and smoothness of the impulse response identified set): Let $\widetilde{CIS}_{\eta}(\bm{\phi}|\mathbf{s}(\mathbf{Y}^T),N)$ be the conditional identified set for $\eta$ with the sufficient statistics $\mathbf{s}(\mathbf{Y}^T)$. For $p(\mathbf{Y}^{T}|\bm{\phi}_0)$-almost every $\mathbf{Y}^T$, $\widetilde{CIS}_{\eta}(\bm{\phi}|\mathbf{s}(\mathbf{y}^T),N)$ is closed and convex, $\widetilde{CIS}_{\eta}(\bm{\phi}|\mathbf{s}(\mathbf{y}^T),N) = [\tilde{\bm{\ell}}(\bm{\phi}, \mathbf{s}(\mathbf{Y}^T)), \tilde{\mathbf{u}}(\bm{\phi}, \mathbf{s}(\mathbf{Y}^T))]$, and its lower and upper bounds are differentiable in $\bm{\phi}$ at $\bm{\phi} = \bm{\phi}_0$ with nonzero derivatives.

Propositions (ref)--(ref) in Appendix (ref) provide primitive conditions for Assumption (ref) to hold in the case where there are shock-sign restrictions. Imposing Assumptions (ref), (ref) and (ref), we obtain the following theorem.

theoremFor $\gamma \in (0,1)$, let $\widehat{C}_{\alpha}^{\ast}$ be the volume-minimizing robust credible region for $\eta$ with credibility $\alpha$,\footnote{The volume-minimizing robust credible region $\widehat{C}_{\alpha}^{\ast}$ is defined as a shortest interval among the connected intervals $C_{\alpha}$ satisfying \begin{equation*} P_{\mathbf{Y}^T|\mathbf{s},\bm{\phi}}(\widetilde{CIS}_{\eta}(\bm{\phi}_0 |\mathbf{s}(\mathbf{Y}^T),N) \subset C_{\alpha} |\mathbf{s}(\mathbf{Y}^T), \bm{\phi}_0) \geq \alpha. \end{equation*} See Proposition 1 in GK for a procedure to compute the volume-minimizing credible region.} which satisfies \begin{equation} \inf_{\pi \in \Pi_{\bm{\phi},\mathbf{Q}|\mathbf{Y}^{T},D_{N} = 1}} \pi( \widehat{C}_{\alpha}^{\ast} ) = \pi_{\bm{\phi}|\mathbf{Y}^{T},D_{N} = 1}(CIS_{\eta}(\bm{\phi}|\mathbf{Y}^{T},N) \subset \hat{C}_{\alpha}^{\ast}|\mathbf{Y}^T, D_{N} = 1) = \alpha. \end{equation} Under Assumptions (ref), (ref), and (ref), $\widehat{C}_{\alpha}^{\ast}$ attains asymptotically valid coverage for the true impulse response, $\eta_0$, conditional on $\mathbf{s}(\mathbf{Y}^T)$. \begin{multline} \liminf_{T \to \infty} P_{\mathbf{Y}^T|\mathbf{s},\bm{\phi}}( \eta_0\in \widehat{C}_{\alpha}^{\ast} |\mathbf{s}(\mathbf{Y}^T), \bm{\phi}_0) \geq \\ \lim_{T \to \infty} P_{\mathbf{Y}^T|\mathbf{s},\bm{\phi}}(\widetilde{CIS}_{\eta}(\bm{\phi}_0|\mathbf{s}(\mathbf{Y}^{T}),N ) \subset \widehat{C}_{\alpha}^{\ast} |\mathbf{s}(\mathbf{Y}^T), \bm{\phi}_0) = \alpha. \end{multline} Accordingly, $\widehat{C}_{\alpha}^{\ast}$ attains asymptotically valid coverage for $\eta_0$ unconditionally, \begin{equation} \liminf_{T \to \infty} P_{\mathbf{Y}^T|\bm{\phi}}( \eta_0\in \widehat{C}_{\alpha}^{\ast} | \bm{\phi}_0) \geq \lim_{T \to \infty} P_{\mathbf{Y}^T|\bm{\phi}}(\widetilde{CIS}_{\eta}(\bm{\phi}_0|\mathbf{s}(\mathbf{Y}^{T}),N ) \subset \widehat{C}_{\alpha}^{\ast} | \bm{\phi}_0) = \alpha. \end{equation}
proofSee Appendix B.

This theorem shows that the robust credible region of GK applied to the SVAR model with NR attains asymptotically valid frequentist coverage for the true impulse response as well as the conditional impulse-response identified set. Even if the point-identification condition of Proposition (ref) holds for the impulse response, it is not obvious if the standard Bayesian credible region can attain frequentist coverage. This is because the Bernstein-von Mises theorem does not seem to hold for the impulse response due to the non-standard features of models with NR.

One could also consider asymptotics under an increasing number of restrictions. We conjecture that, under certain assumptions about how the NR are generated, the class of posteriors for $\eta$ will converge to a point mass at its true value. An implication is that the posterior mean under any conditional prior for $\mathbf{Q}$ that places probability one on the identified set would be consistent for the true value. This result would be an interesting contrast to the case under traditional set-identifying restrictions, where the location of the posterior mean within the identified set is determined purely by the conditional prior and the posterior quantiles lie strictly within the identified set (e.g., Moon_Schorfheide_2012). We do not analyze this case here, since the assumption that there is a fixed number of restrictions seems to be of primary interest. However, Appendix (ref) provides numerical evidence in support of our conjecture. We leave formal investigation of this conjecture for future work.

Empirical application: the dynamic effects of a monetary policy shock

AR18 estimate the effects of monetary policy shocks on the US economy using a combination of sign restrictions on impulse responses and NR. The reduced-form VAR is the same as that used in Uhlig_2005. The model's endogenous variables are real GDP, the GDP deflator, a commodity price index, total reserves, non-borrowed reserves (all in natural logarithms) and the federal funds rate; see Arias, Caldara and Rubio-Ram\'{i}rez (2019)\nocite{Arias_Caldara_Rubio-Ramirez_2019} for details on the variables. The data are monthly and run from January 1965 to November 2007. The VAR includes 12 lags and we include a constant.

As NR, AR18 impose that the monetary policy shock in October 1979 was positive and that it was the overwhelming contributor to the unexpected change in the federal funds rate in that month. This was the month in which the Federal Reserve markedly and unexpectedly increased the federal funds rate following the appointment of Paul Volcker as chairman of the Federal Reserve, and is widely considered to be an example of a positive monetary policy shock (e.g., Romer_Romer_1989). The traditional sign restrictions considered in Uhlig_2005 are also imposed. Specifically, the response of the federal funds rate is restricted to be non-negative for $h=0,1,\ldots,5$ and the responses of the GDP deflator, the commodity price index and nonborrowed reserves are restricted to be nonpositive for $h=0,1,\ldots,5$.

We assume a Jeffreys' (improper) prior over the reduced-form parameters, $\pi_{\bm{\phi}} = \pi_{\mathbf{B},\bm{\Sigma}} \propto |\bm{\Sigma}|^{-\frac{n+1}{2}}$, which is truncated so that the VAR is stable. The posterior for the reduced-form parameters, $\pi_{\bm{\phi}|\mathbf{Y}^{T}}$, is then a normal-inverse-Wishart distribution, from which it is straightforward to obtain independent draws (for example, see DelNegro_Schorfheide_2011). We obtain 1,000 draws from the posterior of $\bm{\phi}$ such that the VAR is stable and $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$ is non-empty. We use Algorithm 1 with $K = 10,000$ draws of $\mathbf{Q}$ at each draw of $\bm{\phi}$ to approximate $l(\bm{\phi},\mathbf{Y}^{T})$ and $u(\bm{\phi},\mathbf{Y}^{T})$. If we cannot obtain a draw of $\mathbf{Q}$ satisfying the restrictions after 100,000 draws of $\mathbf{Q}$, we approximate $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$ as being empty at that draw of $\bm{\phi}$.

We explore the sensitivity of posterior inference to the choice of prior for $\mathbf{Q}|\bm{\phi}$ when the unconditional likelihood is used to construct the posterior. For brevity, we report only the impulse responses of the federal funds rate and real GDP to a positive standard-deviation monetary policy shock (Figure (ref)). As a point of comparison, we report results obtained using a conditionally uniform prior for $\mathbf{Q}|\bm{\phi}$. Under this prior, the 68 per cent highest posterior density credible intervals for the response of real GDP exclude zero at horizons greater than a year or so.\footnote{The results are not directly comparable to those presented in Figure 6 of AR18. First, we present responses to a standard-deviation shock, whereas AR18 describe their responses as being to a 25 basis point shock (although, from close inspection of their Figure 6, it is evident that this normalization is not imposed correctly, because the impact response of the federal funds rate fans out around zero). Second, we use a prior for $\mathbf{Q}$ that is conditionally uniform given $\bm{\phi}$, whereas AR18 use a prior that is unconditionally uniform.} In contrast, the 68 per cent robust credible intervals include zero at all horizons. Under the single prior, the posterior probability that the output response is negative two years after the shock is 95 per cent. In contrast, the posterior lower probability of this event -- the smallest probability over the class of posteriors generated by the class of priors -- is only 54 per cent. The results suggest that posterior inference about the effect of monetary policy on output can be sensitive to the choice of (unrevisable) prior for $\mathbf{Q}|\bm{\phi}$.

figure[figure omitted — 906 chars of source]

AR18 also consider an alternative set of restrictions. Specifically, they impose that the monetary policy shock was: positive in April 1974, October 1979, December 1988 and February 1994; negative in December 1990, October 1998, April 2001 and November 2002; and the most important contributor to the observed unexpected change in the federal funds rate in these months. The choice of these dates is based on a synthesis of information from different sources, including the chronology of monetary policy actions from Romer_Romer_1989, an updated series of the monetary policy shocks constructed using Greenbook forecasts in Romer_Romer_2004, the high-frequency monetary policy surprises from G\"{u}rkaynak, Sack and Swanson (2005)\nocite{Gurkaynak_Sack_Swanson_2005}, and minutes from Federal Open Markets Committee meetings. Under this extended set of restrictions, the set of posterior means and the robust credible interval are tightened noticeably, particularly at shorter horizons (Figure (ref)). The posterior lower probability of a negative output response two years after the shock is now 80 per cent, compared with 54 per cent under the October 1979 restrictions.

Finally, we investigate how posterior inference about the output response is affected by replacing AR18's extended set of restrictions with a shock-rank restriction. Specifically, we estimate the set of output responses that are consistent with the restriction that the monetary policy shock in October 1979 was the largest positive realization of the monetary policy shock in the sample period.\footnote{The large number of inequality constraints and tight conditional identified set induced by the shock-rank restriction poses computational challenges when using Algorithm 1. Accordingly, we use an alternative algorithm to obtain the results. The algorithm adapts an algorithm in Amir-Ahmadi_Drautzburg_2021 and is described in Appendix (ref).} This restriction appears plausible given that the change in the federal funds rate in October 1979 was more positive than the change in the federal funds rate in the other periods identified by AR18 as containing notable monetary policy shocks (Table (ref)). The shock-rank restriction somewhat shrinks the set of posterior means and robust credible regions relative to those obtained under the restrictions on the historical decomposition. Nevertheless, the two sets of restrictions lead to similar (robust) posterior inferences about the output response. The 68 per cent robust credible intervals include zero at all horizons under both sets of restrictions. The posterior lower probability that output falls two years after the shock is 73 per cent under the shock-rank restriction, compared with 80 per cent under the restriction on the historical decomposition.

table[table omitted — 757 chars of source]
figure[figure omitted — 955 chars of source]

In general, $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$ may be empty at particular values of $\bm{\phi}$. The proportion of draws of $\bm{\phi}$ where $\mathcal{Q}(\bm{\phi}|\mathbf{Y}^{T},N,S)$ is empty can therefore be used to assess the plausibility of the restrictions (see GK). Under the October 1979 restrictions, the posterior plausibility of the restrictions is one (i.e., every draw of $\bm{\phi}$ has a nonempty conditional identified set). In contrast, the posterior plausibility under AR18's extended set of restrictions is 53 per cent, while it is only 17 per cent under the shock-rank restriction.

Conclusion

Directly restricting the values of structural shocks to be consistent with historical narratives offers a potentially useful approach to disciplining SVARs, but raises novel issues related to identification and inference. These restrictions generate a set-valued mapping from the model's reduced-form parameters to its structural parameters that depends on the realization of the data entering the restrictions. This means that these restrictions do not fit neatly into the existing framework for analyzing identification in SVARs. In particular, we show that these restrictions may be point-identifying in a frequentist sense. We also highlight issues associated with existing standard Bayesian approaches to estimation and inference. Conditioning on the restrictions holding may result in the posterior placing more weight on parameters that yield a lower ex ante probability that the restrictions are satisfied. We therefore advocate using the unconditional likelihood when constructing the posterior. However, the observed unconditional likelihood will almost always possess flat regions, which implies that a component of the prior will not be updated by the data. Posterior inference may therefore be sensitive to the choice of prior. To address this, we provide robust Bayesian tools to assess or eliminate the sensitivity of posterior inference to the choice of prior. We also provide conditions under which these tools have a valid frequentist interpretation, so our approach should appeal to both Bayesians and frequentists.

While we focus on SVARs in the paper, our analysis could be extended to other settings. For example, Plagborg-Moller_Wolf_2020a explain how to impose traditional SVAR identifying restrictions in the local projection framework under the assumption that the structural shocks are invertible. In Appendix (ref) we briefly discuss how NR could also be imposed within the local projection framework, but we leave a formal analysis of this problem to future research.