Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,922 characters · 23 sections · 80 citation commands
Instrumental Variable Identification of Dynamic Variance Decompositions
}
Keywords: external instrument, impulse response function, invertibility, proxy variable, variance decomposition. JEL codes: C32, C36.
In recent years, and in parallel to popular microeconometric identification strategies, empirical practice in applied macroeconometrics has turned towards “external” sources of plausibly exogenous variation. Such external instrumental variables (IVs, or proxy variables) are now routinely used to estimate causal effects through a simple Two-Stage Least Squares version of Local Projections Jorda2005,Ramey2016. Appealingly, this approach is valid even without the assumption of invertibility -- the ability to recover structural shocks from current and past (but not future) values of the observed macro variables Nakamura2017WP,Stock2018.
However, applied researchers are often not just interested in dynamic causal effects, but also want to learn about a particular shock's contribution to macroeconomic fluctuations Christiano1999,Beaudry2006,Smets2007. If the IV is a perfect measure of the underlying structural macro shock, then the desired variance decompositions are readily computed from standard Local Projection regression output Gorodnichenko2017. In many applications, though, it is likely that external shock measures are contaminated by substantial measurement error, causing attenuation bias. For example, Gertler2015 use high-frequency changes in asset prices around monetary policy announcements as credible instruments for monetary shocks; since these instruments at best capture a subset of all monetary shocks, simple direct regressions on the IV are likely to substantially understate the importance of monetary disturbances. Up to this point, the only possible alternative approach was to combine the IV with conventional Structural Vector Autoregressive (SVAR) methods Stock2008,Mertens2013, thus automatically imposing the otherwise unnecessary and empirically dubious invertibility assumption.
In this paper, we show precisely to what extent external instruments are informative about shock importance. Throughout, we consider an unrestricted linear moving average model, disciplined only by IVs. This model nests conventional, invertible SVARs, as well as essentially all linearized macro models. We prove three main results. First, without further restrictions, the variance decomposition of the instrumented shock's contribution to macroeconomic fluctuations is interval-identified, with informative lower and upper bounds. Second, if the researcher is willing to impose the assumption of recoverability -- i.e., that the shock is spanned by current, past and future values of the observed macro variables -- then both variance decompositions and historical decompositions (the shock's contribution to realized fluctuations) are point-identified. Third, we derive a simple Granger causality pre-test for invertibility that we show exploits the strongest possible testable implication. We complement this set of theoretical results with an extensive code suite that implements all our inference procedures.
We adopt the exact same structural vector moving average (SVMA) model with external IVs as in Stock2018, but focus on variance decompositions, rather than impulse responses. The key identifying assumption of this model is the availability of external instruments that correlate with the shock of interest, but are otherwise dynamically uncorrelated with all other macro shocks. Importantly, the IVs may be contaminated by classical measurement error. Stock2018 show that, in this SVMA-IV model, relative impulse responses (which normalize the impact effect) are point-identified, cf. also Mertens2015. While such relative impulse responses do not require identification of the scale of the underlying shock, scale inevitably matters for variance and historical decompositions, and so lies at the heart of the identification challenge we face in this paper.
We bound the importance of the instrumented structural shock from above and from below by viewing the model as a dynamic measurement error model. Our question is: Given the second moments (autocovariances) of the macro variables and the IVs, what can be said about (forecast or unconditional) variance decompositions? The identification challenge is that we do not know the signal-to-noise ratio of the IV a priori; however, we prove that it is possible to bound this ratio using the moments of the data. At one extreme, our lower bound corresponds to the previously discussed approach of treating the IV as the shock (zero measurement error). If -- as seems likely in practice -- the IV is actually not perfect, then this lower bound may substantially understate the true importance of the shock. At the other extreme, given that we observe a certain degree of co-movement between the IV and the macro observables at various leads and lags, we know that measurement error also cannot be too pervasive. We translate this intuition into formal bounds and prove that these bounds are sharp, i.e., they exhaust all the information about variance decompositions contained in the second moments of the data.
We also characterize the set of additional assumptions that researchers could impose to point-identify both variance and historical decompositions. Here our main result is that point identification obtains if the instrumented shock is assumed to be recoverable, i.e., spanned by all lags and leads of the endogenous macro variables. Appealingly, recoverability obtains in any macro model with as many observables as shocks; in particular, it holds even in many models with news and noise shocks, unlike the strictly stronger (and, as we show, testable) invertibility assumption made in SVAR analysis Leeper2013.
We provide the applied researcher with an easy-to-use code suite that constructs confidence intervals for all parameters of interest. In a first step, we use a reduced-form VAR in macro variables and IVs as a convenient tool for approximating the second moments of the data. The second step then constructs sample analogues of our identification bounds and inserts these into the confidence procedure of Imbens2004; alternatively, we also provide confidence intervals valid under the additional point-identifying restriction of recoverability. We prove that our confidence intervals have asymptotically valid frequentist coverage under weak nonparametric conditions on the data generating process.
To demonstrate the feasibility and applicability of our procedures, we bound the importance of monetary shocks for inflation dynamics in the U.S. We employ the high-frequency IV proposed by Gertler2015, mentioned above. As discussed in Ramey2016, the rising importance of forward guidance since the early 1990s is likely to invalidate the invertibility assumption and so threatens consistency of the standard SVAR-IV estimator used by Gertler2015. Indeed, we find that the data are consistent with substantial non-invertibility. Applying our robust methodology, we find that monetary shocks are almost irrelevant for aggregate inflation in our post-1990 sample: The 90% confidence intervals for the forecast variance contribution of monetary shocks rules out values above 8% at all horizons. Thus, to the extent that inflation is a monetary phenomenon, it is so because of the systematic part of U.S. monetary policy, not because of its erratic conduct.
Finally, we use a series of analytical and quantitative examples to give intuition for why, in spite of its weak identifying assumptions, our method will often manage to give very tight upper bounds on shock importance, consistent with our findings in the monetary application.
\paragraph{Literature.} Plagborg2020 prove that the invertibility-robust Local Projection IV impulse response estimator has the same estimand as a recursive SVAR that includes the IV and orders it first. This paper complements our other work by analyzing the identification of variance and historical decompositions, which requires completely different mathematical arguments.
Non-invertibility and its effects on SVAR identification have received substantial attention in recent years (see the references in PlagborgMoller2019). Previous work has emphasized that, in the empirically relevant case of foresight about economic fundamentals or policy (“news”), conventional SVAR analysis invariably fails: Rational expectations equilibria create non-invertible SVMA representations, and so SVARs cannot correctly recover the structural shocks Leeper2013,Wolf2018. In contrast, non-invertibility poses no challenge to the methods developed in this paper. We also show that, in the SVMA-IV model, the degree of invertibility is set-identified. Our proposed test of invertibility is related to the Granger causality tests developed in SVAR settings by Giannone2006 and Forni2014. Finally, the weaker notion of “recoverability” studied here has independently been proposed by Chahrour2018 outside the context of external IV identification.\footnote{Recoverability is formally equivalent to the assumption that the structural shock is spanned by current and future reduced-form VAR forecast errors. Such dynamic rotations of $u_t$ have been exploited in non-IV settings by Lippi1994, Mertens2010, and Forni2017EJ,Forni2017AEJMacro.}
\paragraph{Outline.} (ref) defines the SVMA-IV model and the parameters of interest, and states the identification problem. (ref) derives our identification results. (ref) gives a practical overview of our procedures and their implementation. (ref) applies the procedures to bound the importance of monetary shocks. (ref) illustrates the usefulness and interpretation of the upper bound on shock importance through analytical examples. (ref) compares the finite-sample performance of our procedures to the SVAR-IV approach through simulations. (ref) concludes. Proofs of our main results are relegated to (ref). The Matlab code suite and a supplemental appendix are available online.\footnote{\url{https://github.com/mikkelpm/svma_iv}}
We begin by defining the econometric model and the parameters of interest. Then we state the identification problem.
Following Stock2018, we assume a SVMA-IV model. This model allows for an unrestricted linear shock transmission mechanism and, unlike standard SVAR analysis, does not require shocks to be invertible. We also assume the availability of valid external IVs (proxy variables) -- variables that correlate with the shock of interest, but not with the other shocks. For notational clarity, we assume throughout that all time series below have zero mean and are strictly non-deterministic.
First, we define the SVMA model, which places no restrictions on the linear transmission of the vector of shocks $\varepsilon_t$ to the vector of observed endogenous variables $y_t$.
The $(i,j)$ element $\Theta_{i,j,\ell}$ of the moving average coefficient matrix $\Theta_\ell$ is the impulse response of variable $i$ to shock $j$ at horizon $\ell$. The $j$-th column of $\Theta_\ell$ is denoted by $\Theta_{\bullet, j,\ell}$ and the $i$-th row by $\Theta_{i, \bullet,\ell}$. The full-rank assumption guarantees a nonsingular stochastic process. This condition requires $n_\varepsilon \geq n_y$, but -- crucially -- we do not assume that the number of shocks $n_\varepsilon$ is known. The mutual orthogonality of the shocks is the standard assumption in empirical macroeconomics. The model is semiparametric in that we place no a priori restrictions on the coefficients of the infinite moving average, except to ensure a valid stochastic process. In particular, the infinite-order SVMA model (ref) is consistent with all discrete-time Dynamic Stochastic General Equilibrium (DSGE) models and all stable SVAR models for $y_t$.
Second, we assume the availability of one or more external IVs for the shock of interest, with the shock of interest specified to be the first one, $\varepsilon_{1,t}$. Each of the $n_z$ IVs $z_t = (z_{1,t},\dots,z_{n_z,t})'$ are assumed to correlate with the first shock but not the other shocks, after controlling for lagged variables: For all $i=1,\dots,n_z$,
where $\tilde{z}_{i,t}$ is the population residual from projecting $z_{i,t}$ on all lags of $\lbrace z_t,y_t \rbrace$. The key exclusion restriction is that the shock of interest $\varepsilon_{1,t}$ is the only contemporaneous shock to correlate with the IVs $z_t$. Thus, $\tilde{z}_t$ is a proxy for $\varepsilon_{1,t}$ (up to scale) that is contaminated by classical measurement error. This is a strong assumption that must be carefully defended in applications. Ramey2016 and Stock2018 survey the extensive applied literature that has constructed plausibly valid external IVs for various shocks.
Using linear projection notation, we can equivalently express the IV exclusion restrictions (ref) as the assumption that the IVs $z_t$ are proportional to the shock of interest $\varepsilon_{1,t}$ plus classical measurement error $v_t$ (and possibly lagged observed variables).
The interpretation of external IVs as noisy measures of true shocks is discussed in Mertens2013 and Stock2018. In our notation (ref), the scale parameter $\alpha$ (along with the residual variance-covariance matrix $\Sigma_v$) measures the overall strength of the IVs, while the unit-length vector $\lambda$ determines which IVs are stronger than others. The assumptions on the coefficients $\Psi_\ell$ and $\Lambda_\ell$ ensure stationarity. We emphasize that the linearity of equation (ref) is not a structural assumption; it arises from a linear projection (as in the “first stage” of cross-sectional IV). In particular, (ref) is consistent with the IV being a binary or censored series, since such a variable can still satisfy the moment conditions (ref) that are equivalent with equation (ref).
Since we restrict attention to identification from second moments, we may without loss of generality simplify notation by assuming that all disturbances are Gaussian.
The Gaussianity assumption is strictly for notational convenience. We could instead have maintained the above white noise assumptions (which allow for conditional heteroskedasticity) and phrased all our results using linear projection notation. The sole meaningful restriction is that we only exploit second moments of the data for identification, as is standard in the applied macro literature, and without loss of generality for Gaussian data.\footnote{If we were to take the assumption of i.i.d. shocks seriously, and the shocks were not Gaussian, higher-order moments of the data would be informative about the parameters. However, we agree with most of the literature that the assumption of i.i.d. shocks is too strong due to the likely presence of stochastic volatility.} We drop the Gaussianity assumption when developing inference procedures in (ref).
Finally note also that (ref) together imply that the $(n_y+n_z)$-dimensional data vector $(y_t',z_t')'$ is strictly stationary.
We are interested in the propagation of the first structural shock $\varepsilon_{1,t}$ to the macroeconomic aggregates $y_t$. This section lists the parameters of interest to the applied macroeconomist.
\paragraph{Impulse responses.} As discussed above, the $(i,1)$ element $\Theta_{i,1,\ell}$ of the moving average coefficient matrix $\Theta_\ell$ is the impulse response of variable $i$ to shock $1$ at horizon $\ell$. We distinguish such absolute impulse responses from relative impulse responses $\Theta_{i,1,\ell}/\Theta_{1,1,0}$, which give the response of $y_{i,t+\ell}$ to a shock to $\varepsilon_{1,t}$ that increases $y_{1,t}$ by one unit on impact.
\paragraph{Invertibility and recoverability.} The shock $\varepsilon_{1,t}$ is said to be invertible if it is spanned by past and current (but not future) values of the endogenous variables $y_t$: $\varepsilon_{1,t} = E(\varepsilon_{1,t} \mid \lbrace y_\tau\rbrace_{-\infty<\tau\leq t})$. This condition may or may not hold in a given moving average model (ref), depending on the impulse response parameters $\Theta_\ell$. Conventional SVAR analysis invariably imposes invertibility, since the SVAR model obtains from the additional assumptions that $n_\varepsilon=n_y$ and that $\Theta(L)$ has a one-sided inverse, so the shocks $\varepsilon_t = \Theta(L)^{-1}y_t$ are spanned by current and past data. However, in many structural macro models, at least some of the shocks cannot be recovered from only lagged macro observables, i.e., the moving average representation is noninvertible. For example, this is often the case in models with news (anticipated) shocks or noise (signal extraction) shocks Blanchard2013,Leeper2013. Furthermore, if $n_\varepsilon>n_y$, it is impossible for all shocks to be invertible.
A continuous measure of the degree of invertibility is the $R^2$ value in a population regression of the shock on past and current observed variables (Sims2006; Forni2018). More generally, we define
the population R-squared value in a projection of the shock of interest on data up to time $t+\ell$ (recall that $\operatorname*{Var}(\varepsilon_{1,t})=1$). If the shock is invertible in the sense of the previous paragraph, then $R_0^2=1$. Hence, if $R_0^2<1$, then no SVAR model can generate the impulse responses $\Theta(L)$, although the model is nearly consistent with SVAR structure if $R_0^2 \approx 1$ Wolf2018.
A weaker condition than invertibility is that the shock of interest is recoverable from all leads and lags of the endogenous variables -- that is, if $E(\varepsilon_{1,t} \mid \lbrace y_\tau \rbrace_{-\infty<\tau<\infty})=\varepsilon_{1,t}$, or equivalently if $R_\infty^2 = 1$. A sufficient condition is that $n_\varepsilon=n_y$, since then $\Theta(L)$ automatically has a two-sided inverse Brockwell1991, and thus the shocks $\varepsilon_t = \Theta(L)^{-1}y_t$ are spanned by current, past, and future data. This is the case in many DSGE models with news (i.e., anticipated) shocks Leeper2013.
\paragraph{Variance decompositions.} Variance decompositions are the key parameters of interest in this paper. We focus in the main text on the forecast variance ratio (FVR), where the FVR for the shock of interest for variable $i$ at horizon $\ell$ is defined as \[\mathit{FVR}_{i,\ell} \equiv 1 - \frac{\operatorname*{Var}(y_{i,t+\ell} \mid \lbrace y_\tau \rbrace_{-\infty<\tau\leq t}, \lbrace \varepsilon_{1,\tau} \rbrace_{t<\tau<\infty})}{\operatorname*{Var}(y_{i,t+\ell} \mid \lbrace y_\tau \rbrace_{-\infty<\tau\leq t})} = \frac{\sum_{m=0}^{\ell-1} \Theta_{i,1,m}^2}{\operatorname*{Var}(y_{i,t+\ell} \mid \lbrace y_\tau \rbrace_{-\infty<\tau\leq t})}.\] The FVR measures the reduction in the econometrician's forecast variance that would arise from being told the entire path of future realizations of the first shock. The larger this measure is, the more important is the first shock for forecasting variable $i$ at horizon $\ell$. The FVR is always between 0 and 1.
(ref) defines and provides identification analysis for two additional variance decomposition concepts. First, the forecast variance decomposition (FVD) is like the FVR but instead conditions on the history of all past shocks $\lbrace \varepsilon_\tau \rbrace_{-\infty<\tau\leq t}$, rather than the history of observables $\lbrace y_\tau \rbrace_{-\infty<\tau\leq t}$. Under invertibility, the FVR and FVD are identical (since then the information set $\lbrace y_\tau \rbrace_{-\infty<\tau\leq t}$ equals the information set $\lbrace \varepsilon_\tau \rbrace_{-\infty<\tau\leq t}$), explaining why the previous SVAR literature has not distinguished between the two. Second, we consider the unconditional frequency-specific variance decomposition (VD) of Forni2018.
\paragraph{Historical decomposition.} The historical decomposition of variable $y_{i,t}$ at time $t$ attributable to the shock of interest is defined as $E(y_{i,t} \mid \lbrace \varepsilon_{1,\tau} \rbrace_{-\infty<\tau\leq t}) = \sum_{\ell=0}^\infty \Theta_{i,1,\ell}\varepsilon_{1,t-\ell}$.
Our goal for the remainder of the paper is to answer the question: Given (ref), what do the second moments (autocovariances) of the data $(y_t',z_t')'$ say about the parameters of interest defined above? In particular, can we test whether the shock $\varepsilon_{1,t}$ is invertible?
Stock2018 showed that relative impulse responses are point-identified in the SVMA-IV model. To see this transparently, consider the case with a single IV, so $\lambda=1$. Since
the absolute impulse responses $\Theta_{i, 1,\ell}$ for all variables $i$ and all horizons $\ell$ are identified up to the single scale parameter $\alpha$. Thus, the relative impulse responses $\Theta_{i, 1,\ell}/\Theta_{1, 1,0}$ are point-identified, as $\alpha$ drops out from the fraction.
The main challenge addressed in this paper is that (partial) identification of variance and historical contributions requires (partial) identification of the absolute impulse responses, and thus of the scale parameter $\alpha$.
This section contains our main theoretical identification results. Readers who are primarily interested in practical implementation are encouraged to skip ahead to (ref). For exposition, we start in (ref) by deriving results for a simple static version of our SVMA-IV model. We then turn to the general dynamic model in (ref), applying the static results to the frequency domain representation of the data. We initially focus on the case with a single IV, but we discuss the straight-forward extension to multiple IVs in (ref).
To build intuition, consider a static version of the SVMA-IV model with a single instrument:
Here $\alpha,\sigma_v \geq 0$ are scalars, $\xi_t \equiv \sum_{j=2}^{n_\varepsilon} \Theta_{\bullet,j,0}\varepsilon_{j,t}$ is an $n_y$-dimensional random vector that captures all the structural shocks other than the one of interest, and $\Sigma_\xi \equiv \operatorname*{Var}(\xi_t)$.\footnote{While the static model is primarily intended to provide intuition about the analysis of the SVMA-IV model, the results in this subsection are directly relevant for identification in the more restrictive SVAR model with an external IV. In that framework, $y_t$ would denote the $n_y$ reduced-form VAR residuals, which are linear functions of the vector $\varepsilon_t$ of $n_\varepsilon$ contemporaneous structural shocks.}
Our main parameter of interest is the Forecast Variance Ratio \[\textit{FVR}_{i,1} = 1 - \frac{\operatorname*{Var}(y_{i,t} \mid \varepsilon_{1,t})}{\operatorname*{Var}(y_{i,t})} = \frac{\Theta_{i,1,0}^2}{\operatorname*{Var}(y_{i,t})}.\] This is just the population R-squared value in the (infeasible) regression of $y_{i,t}$ on $\varepsilon_{1,t}$. Since $\operatorname*{Cov}(y_{i,t},z_t) = \alpha\Theta_{i,1,0}$, it is easy to see that the FVR is identified up to a factor $1/\alpha^2$:
Thus, we ask: What does the variance-covariance matrix of the data $(y_t',z_t)'$ say about the scale parameter $\alpha^2$?
Our key insight is that the static model is nothing but a multivariate classical measurement error model: Whereas we would like to measure the R-squared value from a regression of $y_t$ on $\varepsilon_{1,t}$, we only observe the noisy proxy $z_t$ for the “regressor”. Intuitively, the contribution of $\varepsilon_{1,t}$ to $y_t$ is not point-identified because the signal-to-noise ratio $\alpha^2/\sigma_v^2$ of the proxy $z_t$ is not known a priori. For example, upon observing a small correlation between the IV and macro observables, we do not know whether this correlation is small because of measurement error or because the shock is unimportant. Nevertheless, the moments of the data are informative about the signal-to-noise ratio. At one extreme, the IV can never be more than perfect -- at best, there is no measurement error (infinite signal-to-noise ratio). At the other extreme, the signal-to-noise ratio cannot be zero, since then the IV would not correlate at all with macro observables. We now formalize this intuition.\footnote{Our bounds do not follow from existing results in the literature on measurement error in linear regression Klepper1984, since our parameters of interest are not regression coefficients.}
\paragraph{Lower bound on shock importance.} We begin with a lower bound on the importance of the shock (and so on the amount of measurement error), or equivalently an upper bound on $\alpha^2$. To derive this bound, simply observe that \[\alpha^2 \leq \alpha^2 + \sigma_v^2 = \operatorname*{Var}(z_t).\] This inequality binds when there is no measurement error in the IV, i.e., when $\sigma_v=0$.
Mapping this upper bound on $\alpha^2$ into a lower bound on the FVR via (ref), we get
The lower bound corresponds to the population R-squared value in a regression of $y_{i,t}$ on $z_t$, that is, a regression which treats the IV as if it were a perfect measure of the shock $\varepsilon_{1,t}$ (up to scale). The attenuation bias imparted by the measurement error $v_t$ implies that this regression yields a lower bound on the true FVR.
\paragraph{Upper bound on shock importance.} To derive the upper bound on the importance of the shock (and on the amount of measurement error), or equivalently the lower bound on $\alpha$, define first $z_t^\dagger \equiv E(z_t \mid y_t)$ and $\varepsilon_{1,t}^\dagger \equiv E(\varepsilon_{1,t} \mid y_t)$. Then, by standard linear projection algebra, we must have \[\operatorname*{Var}(z_t^\dagger) = \alpha^2\operatorname*{Var}(\varepsilon_{1,t}^\dagger) \leq \alpha^2 \operatorname*{Var}(\varepsilon_{1,t}) = \alpha^2.\] Intuitively, $\alpha^2=\operatorname*{Var}(E(z_t \mid \varepsilon_{1,t}))$ is the explained sum of squares from a projection of $z_t$ on the shock $\varepsilon_{1,t}$. This must weakly exceed the explained sum of squares $\operatorname*{Var}(z_t^\dagger)=\operatorname*{Var}(E(z_t \mid y_t))$ from a projection of $z_t$ on $y_t$, simply because the variables in $y_t$ are effectively noisy measures of the shock $\varepsilon_{1,t}$ contaminated by other structural shocks $\xi_t$, and uncorrelated with $v_t$. In other words, the explanatory power of the variables $y_t$ for the IV $z_t$ puts a lower bound on the possible signal-to-noise ratio $\alpha^2/\sigma_v^2=\alpha^2/(\operatorname*{Var}(z_t)-\alpha^2)$. The inequality above binds when the shock is invertible ($\varepsilon_{1,t}^\dagger \equiv E(\varepsilon_{1,t} \mid y_t)=\varepsilon_{1,t}$), i.e., when the macro observables $y_t$ explain as much of the variation in the IV as the shock $\varepsilon_{1,t}$ itself does.
Mapping the lower bound on $\alpha^2$ into an upper bound on the FVR via (ref), we get
The upper bound corresponds to treating the projection $z_t^\dagger = \alpha \varepsilon_{1,t}^\dagger$ of the IV on the macro observables as a perfect measure of the shock (up to scale). This is correct if indeed the shock were invertible ($\varepsilon_{1,t}^\dagger=\varepsilon_{1,t}$), but otherwise overstates the importance of the shock. Intuitively, unless the shock is in fact invertible, the upper bound mistakenly attributes too much of the lack of co-movement between $y_t$ and $z_t$ to measurement error (rather than the actual limited importance of $\varepsilon_{1,t}$).
Whereas the lower bound (ref) on the FVR for variable $i$ does not depend on the entire set of observed macro aggregates $y_t$, the upper bound (ref) decreases monotonically as we add more variables to the vector $y_t$. In particular, the upper bound equals the trivial bound of 1 if there is only one observable ($n_y=1$), since in this case we cannot rule out that the scalar time series $y_t$ is driven entirely by the first shock, with the imperfect correlation between $y_t$ and $z_t$ purely caused by measurement error.\footnote{Mathematically, when $y_t$ is a scalar, then $z_t^\dagger = E(z_t \mid y_t) \propto y_t$, so $\operatorname*{Corr}(y_t,z_t^\dagger)= \pm 1$.} However, when $n_y \geq 2$, the upper bound is generally below 1. We present an analytical example in (ref) that shows how the addition of extra observables helps sharpen identification, and clarifies the conditions under which we can expect the upper bound to be close to the true FVR.
\paragraph{Identified set.} The bounds $\alpha^2 \in [\operatorname*{Var}(z_t^\dagger),\operatorname*{Var}(z_t)]$ are sharp, i.e., exploit all information contained in the second moments of the data, in the following sense. Suppose we are given any non-singular variance-covariance matrix for the data $(y_t',z_t)'$, as well as any value of $\alpha^2$ in our interval. We can then choose appropriate values of the remaining parameters such that the model matches the given variance-covariance matrix of the data.\footnote{ This is achieved by the choices $\Theta_{\bullet,1,0} = \frac{1}{\alpha}\operatorname*{Cov}(y_t,z_t)$, $\sigma_v^2 = \operatorname*{Var}(z_t) - \alpha^2$, and $\Sigma_\xi = \operatorname*{Var}(y_t) - \frac{1}{\alpha^2}\operatorname*{Cov}(y_t,z_t)\operatorname*{Cov}(y_t,z_t)'$. This choice of $\sigma_v^2$ is nonnegative since $\operatorname*{Var}(z_t) \geq \alpha^2$, and (ref) in (ref) implies that the choice of $\Sigma_\xi$ is a positive semidefinite matrix since $\alpha^2 \geq \operatorname*{Var}(z_t^\dagger) = \operatorname*{Cov}(z_t,y_t)\operatorname*{Var}(y_t)^{-1}\operatorname*{Cov}(y_t,z_t)$.}
Under what conditions are the bounds on $\alpha^2$ -- and thus on the FVR -- likely to be tight (i.e., close to the true FVR)? We can express the identified set for $1/\alpha^2$ in terms of the underlying model parameters as follows: \[\frac{1}{\alpha^2} \in \left[\frac{\alpha^2}{\alpha^2 + \sigma_v^2} \times \frac{1}{\alpha^2}\;,\; \frac{1}{\operatorname*{Var}(E(\varepsilon_{1,t} \mid y_t))} \times \frac{1}{\alpha^2}\right].\] The lower bound is closer to the true value $1/\alpha^2$ when the actual signal-to-noise ratio $\alpha^2/\sigma_v^2$ is larger, i.e., when the IV is stronger. The upper bound is closer to the true FVR when the degree of invertibility $R_0^2 = \operatorname*{Var}(\varepsilon_{1,t}^\dagger) = \operatorname*{Var}(E(\varepsilon_{1,t} \mid y_t))$ is larger, i.e., when the macro variables $y_t$ are more informative about the hidden shock $\varepsilon_{1,t}$. Finally, the identified set is never empty, and it collapses to a point only in case of a perfect IV and invertibility.
\paragraph{Point identification.} Point identification obtains if the researcher assumes either that the IV is perfect ($\sigma_v=0$), in which case the lower bound for the FVR binds, or that the shock of interest is invertible ($\varepsilon_{1,t}^\dagger=\varepsilon_{1,t}$), in which case the upper bound binds.
We now analyze identification in the general dynamic model of (ref). The key idea in our proofs is to apply the logic of the static model frequency-by-frequency to the frequency domain representation of the data.
As in the static case, we begin in this section by characterizing the identified set for the scale parameter $\alpha$. While not economically interesting in itself, this scale parameter is ultimately key to identification of our actual parameters of interest. We maintain (ref) throughout, but for the moment consider the case of a single IV ($n_z=1$), leaving the generalization to (ref). That is, $z_t$ is a scalar and $\lambda=1$ in equation (ref). We write $\Sigma_v^{1/2}=\sigma_v \geq 0$, a scalar.
\paragraph{Preliminaries.} It will prove convenient to define the IV projection residual that removes any dependence on lagged observed variables:
Note that $\tilde{z}_t$ is serially uncorrelated by construction.
Next, we need to define our notation for spectral density matrices. For any two jointly stationary vector time series $a_t$ and $b_t$ of dimensions $n_a$ and $n_b$, respectively, define the $n_a \times n_b$ cross-spectral density matrix function Brockwell1991 \[s_{ab}(\omega) \equiv \frac{1}{2\pi}\sum_{\ell=-\infty}^\infty e^{-i\omega\ell}\operatorname*{Cov}(a_t,b_{t-\ell}),\quad \omega \in [0,2\pi].\] For any vector time series $a_t$, we denote its spectrum by $s_a(\omega) \equiv s_{aa}(\omega)$.
\paragraph{Lower bound on shock importance.} We again begin with a lower bound on shock importance (or the amount of measurement error), which corresponds to an upper bound on the scale parameter $\alpha$. As in the static model, we find
Thus, once we look at the residualized IV in (ref), the bound construction works as in the static case, with the boundary $\alpha = \alpha_{UB}$ corresponding to a perfect IV.
\paragraph{Upper bound on shock importance.} For the upper bound on shock importance (or the lower bound on $\alpha^2$), we apply a version of the argument from the static case to the joint spectrum of the data at every frequency. First, as in the static case, we define the projections of $\tilde{z}_t$ and $\varepsilon_{1,t}$, respectively, just now onto all leads and lags of the endogenous variables $y_t$:
Note that $\tilde{z}_t^\dagger = \alpha \varepsilon_{1,t}^\dagger$, since the measurement error $v_t$ is dynamically uncorrelated with $y_t$. Applying the same logic as in the static case at an arbitrary frequency $\omega \in [0,2\pi]$, we have
The last equality uses that the shock $\varepsilon_{1,t}$ is white noise with variance 1. Similar to the static case, the inequality above arises because the “explained sum of squares” $s_{\varepsilon_1^\dagger}(\omega)$ from a frequency-specific projection of the shock $\varepsilon_{1,t}$ on all leads and lags of the macro observables $y_t$ must be less than the “total sum of squares” $s_{\varepsilon_1}(\omega)$.\footnote{Brockwell1991 show that $s_{\tilde{z}^\dagger}(\omega) = s_{y\tilde{z}}(\omega)^*s_y(\omega)^{-1}s_{y\tilde{z}}(\omega)$ and $s_{\varepsilon_1^\dagger}(\omega) = s_{y\varepsilon_1}(\omega)^*s_y(\omega)^{-1}s_{y\varepsilon_1}(\omega)$. Since the joint spectrum is positive semidefinite, $s_{\varepsilon_1}(\omega) \geq s_{\varepsilon_1^\dagger}(\omega)$ for all $\omega$.} Exploiting the inequality (ref) at all frequencies, we obtain the lower bound
The bound binds if at some frequency $\omega \in [0,\pi]$ the observed macro aggregates are perfectly informative about the hidden shock $\varepsilon_{1,t}$. This is the natural dynamic, frequency-domain analogue of the condition in the static case, where we required the static $y_t$ to be perfectly informative about $\varepsilon_{1,t}$. If the macro aggregates are in fact not perfectly informative about the shock at any frequency, then the lower bound attributes too much of the (frequency-by-frequency) lack of co-movement between $y_t$ and $z_t$ to measurement error.
\paragraph{The identified set.} The main theoretical result of this paper is that the above bounds $\alpha_{LB}^2,\alpha_{UB}^2$ are sharp.
Recall that the previous discussion has already shown that any value of $\alpha^2 \notin [\alpha_{LB}^2,\alpha_{UB}^2]$ is impossible. The proposition strengthens this result to say that, given the second moments of the data, we cannot rule out any values of $\alpha^2$ in the interval $[\alpha_{LB}^2,\alpha_{UB}^2]$.\footnote{The proposition does not cover the knife-edge case $\alpha=\alpha_{LB}$ due to economically inessential technicalities.}
To interpret the identified set, we proceed as in the static model and express the interval in terms of the underlying model parameters. We focus on the identified set for $\frac{1}{\alpha^2}$, as this transformation is again the most relevant one for identifying the FVR and degree of invertibility/recoverability, as shown below. We can write the identified set for $1/\alpha^2$ as
As in the static case, the lower bound is larger (and closer to the true $\frac{1}{\alpha^2}$) when the instrument is stronger in the sense of a higher signal-to-noise ratio $\alpha^2/\sigma_v^2$. The upper bound is again smaller (and closer to the true $\frac{1}{\alpha^2}$) when the data are more informative about the shock of interest. The relevant notion of informativeness, however, is now more complicated than in the static case, for two reasons: First, we now exploit the explanatory power of all leads and lags of the macro aggregates when forming the projection $\varepsilon_{1,t}^\dagger \equiv E(\varepsilon_{1,t} \mid \lbrace y_\tau \rbrace_{-\infty<\tau<\infty})$; and second, we consider all frequencies of the data separately. The upper bound is close to the truth as long as the leads and lags of $y_t$ are highly informative about the frequency-$\overline{\omega}$ fluctuations of the shock at some frequency $\overline{\omega}$ (e.g., in the long run $\overline{\omega} \approx 0$), in the sense that the spectral density of the projection residual $\varepsilon_{1,t}- \varepsilon_{1,t}^\dagger$ vanishes at this frequency. This does not require the macro variables to be informative about the shock at all frequencies (e.g., in the short run $\overline{\omega} \approx \pi$). We illustrate this point in (ref).
Similar to the static case, the identified set for $\frac{1}{\alpha^2}$ does not collapse to a point unless the instrument is perfect and there exists a frequency $\overline{\omega}$ for which the data are perfectly informative about the frequency-$\overline{\omega}$ cyclical component of the shock.
\paragraph{Practical upper bound on shock importance.} In practice, we do not recommend exploiting the sharp lower bound on $\alpha$ for estimation and inference. The reason is that $\alpha_{LB}$ in equation (ref) equals the supremum of a function, which depends on the spectral density matrix of the data. Nonparametric estimation of the supremum of an unknown function is highly challenging given the moderate sample sizes available to applied macroeconomists Gafarov2018. For this reason, our implementation in (ref) instead uses the weaker bound
Since $\underline{\alpha}^2$ is given by an integral of the spectrum as opposed to a supremum, its point estimator defined in (ref) is consistent and asymptotically normal, as shown in (ref).\footnote{Methods from the moment inequality literature could be applied to develop confidence intervals that exploit our sharp lower bound $\alpha_{LB}^2$ Andrews2013, Andrews2017, Chernozhukov2013. We leave this more complicated option to future work. Alternatively, if researchers have a strong a priori reason to believe that the shock is likely to be particularly important at certain frequencies, then they may fix frequency bounds $[\omega_1, \omega_2]$ and compute the integral in (ref) by integrating over this interval only.}
Since $\operatorname*{Var}(\tilde{z}_t^\dagger) = \alpha^2 \operatorname*{Var}(\varepsilon_{1,t}^\dagger) = \alpha^2 \times R_\infty^2$, the weaker lower bound on $\alpha^2$ will nevertheless be close to the truth if the shock of interest is close to being recoverable ($R_\infty^2 \approx 1$), and thus in particular if the shock is close to being invertible ($R_0^2 \approx 1$). In (ref) we show by example that the bound $\underline{\alpha}^2$ binds in a model with news shocks, which cannot be analyzed using conventional SVAR-IV methods that assume invertibility.
Given the identified set for $\frac{1}{\alpha^2}$, it is now straight-forward to derive identified sets for variance decompositions as well as the degree of invertibility and recoverability.
\paragraph{Variance decompositions.} The FVR satisfies
Hence, as in the static case, the identified set for $\mathit{FVR}_{i,\ell}$ equals the identified set for $\frac{1}{\alpha^2}$, scaled by the (point-identified) second fraction on the far right-hand side above. As discussed previously, and as in the static case, the lower bound for the FVR depends on the strength of the IV, and the upper bound on the FVR depends on the informativeness of the macro variables for the shock of interest. Adding more variables to the vector $y_t$ of endogenous observables always leads to a weakly narrower identified set (in percentage terms, since the parameter $\mathit{FVR}_{i,\ell}$ itself also changes when we change the vector $y_t$). Unlike in the static case, the upper bound in the dynamic case is generally below 1 even if we only observe a single macro time series ($n_y=1$), as shown by example in (ref).
(ref) derives bounds on the other variance decomposition concepts (VD and FVD) introduced in (ref). Bounding the FVD in particular requires more work.
\paragraph{Degree of invertibility & recoverability.} The definition (ref) of $R_\ell^2$ implies
Since the variance on the right-hand side above is point-identified, the identified sets for the degree of invertibility ($\ell = 0$) and the degree of recoverability ($\ell=\infty$) follow immediately from the identified set for $\frac{1}{\alpha^2}$.
From the sharp bounds on $R_0^2$ and $R_\infty^2$, we can also derive testable conditions under which the distribution of the observable data is consistent with invertibility or recoverability.
According to (ref), $\varepsilon_{1,t}$ is certain to be noninvertible if and only if $\tilde{z}_t$ Granger causes $y_t$ (which is equivalent with the condition that $z_t$ Granger causes $y_t$). This result will be the basis for the pre-test of invertibility in (ref). Note, however, that a finding of Granger non-causality need not imply that $R_0^2=1$; the identified set for $R_0^2$ always includes values below 1. (ref) additionally implies that $\varepsilon_{1,t}$ is certain to be non-recoverable if and only if $\tilde{z}_t^\dagger$, defined in (ref), is serially correlated at some lag.\footnote{We leave the development of a practical statistical test of recoverability to future research.}
\paragraph{Absolute impulse responses.} For completeness, we note that the identified set for the absolute impulse response $\Theta_{i,1,\ell}$ is obtained by scaling the identified set for $\frac{1}{\alpha}$, cf. equation (ref). This extends existing results on the point-identification of relative impulse responses Stock2018, as discussed at the end of (ref).
As we have seen, without further restrictions, our various parameters of interest are only interval-identified, albeit with informative bounds. In this section we complement those results by stating a menu of sufficient conditions, each of which guarantees point identification of the FVR and historical decompositions.
\paragraph{Informative instruments.} Point identification obtains if the researcher is willing to assume that the instrument is perfect, i.e., $\sigma_v=0$. In this case the lower bounds on the FVR and degree of invertibility/recoverability bind. Indeed, since the instrument equals the shock up to scale, $\tilde{z}_t=\alpha\varepsilon_{1,t}$, the FVR and historical decompositions are easily computed through regressions Jorda2005,Gorodnichenko2017. Note that the assumption that the IV is perfect is not testable.
\paragraph{Informative macro aggregates.} The second set of sufficient conditions relates to the informativeness of the macro aggregates $y_t$ for the hidden shock $\varepsilon_{1,t}$. In this category, our weakest condition for point identification is that the data $y_t$ is perfectly informative about $\varepsilon_{1,t}$ at some frequency, i.e., the spectral density of the projection residual $\varepsilon_{1,t}-\varepsilon_{1,t}^\dagger$ vanishes at some frequency $\overline{\omega}$. Then $\alpha=\alpha_{LB}$, so the FVR and degree of invertibility/recoverability are identified. This assumption is not testable.
A stronger but more easily interpretable assumption is recoverability, i.e., $\varepsilon_{1,t}^\dagger \equiv E(\varepsilon_{1,t} \mid \lbrace y_\tau\rbrace_{-\infty<\tau<\infty})=\varepsilon_{1,t}$. This assumption is testable, cf. (ref). As explained in (ref), recoverability is restrictive, but it is a meaningfully weaker requirement than invertibility in many economic applications, such as in the news shock model in (ref) below. In particular, it is satisfied whenever there are as many shocks as variables, $n_\varepsilon=n_y$. Under recoverability, the shock itself can be identified as $\varepsilon_{1,t} = \frac{1}{\alpha} \tilde{z}_t^\dagger$, so the historical decomposition $E(y_{i,t} \mid \lbrace \varepsilon_{1,\tau} \rbrace_{-\infty<\tau \leq t})=E(y_{i,t} \mid \lbrace \tilde{z}_\tau^\dagger \rbrace_{-\infty<\tau \leq t})$ is also identified.
To conclude, we briefly extend the analysis to a model with multiple IVs for the shock of interest ($n_z\geq 2$). This extended multiple-IV model is testable, unlike the single-IV model. As in the single-IV case, define the projection residual
(ref) shows that the testable implication of the multiple-IV model is that the cross-spectrum $s_{y\tilde{z}}(\omega)$ has a rank-1 factor structure. The validity of the multiple-IV model can be rejected if and only if this factor structure fails.
When the multiple-IV model is consistent with the distribution of the data, then the identification analysis can be reduced to the single-IV case in (ref). Specifically, (ref) shows that (i) $\lambda$ is point-identified, and (ii) the identified sets for $\alpha$, variance decompositions, and the degree of invertibility are the same as the identified sets that exploit only the scalar instrument
Intuitively, $\breve{z}_t \propto E(\varepsilon_{1,t} \mid \tilde{z}_t)$. Because $\breve{z}_t$ is a linear combination of all $n_z$ instruments, the identified sets are narrower than if we had used any one instrument $z_{k,t}$ in isolation.
In (ref) we also derive sharp bounds in the more general case of multiple instruments being correlated with multiple structural shocks, as in Mertens2013.
We now describe the practical implementation of our inference procedures for variance decompositions and our test of invertibility of the shock of interest. To keep the exposition self-contained, we review some of the conclusions from (ref). For ease of notation we focus on the case with a single IV $z_t$ in this section. The generalization to multiple instruments is straight-forward, cf. (ref).
The key (ref) in (ref) showed that, without further assumptions, variance decompositions and the degree of invertibility are only partially identified. That is, even if the sample size were infinite so we knew the autocovariance function of the observed data $W_t \equiv (y_t',z_t)'$ perfectly, we would not be able to exactly pinpoint the true values of these parameters. However, we were able to derive informative bounds on the parameters of interest. The remainder of this section gives an overview of how to compute those bounds in practice, how to do inference on the identified set, and how the additional a priori assumption of recoverability allows for consistent point estimation of all economic parameters of interest.
As mentioned in the introduction, a Matlab code suite that implements all steps below is available online.
Our bounds are simple functions of the autocovariances of the data $W_t \equiv (y_t',z_t)'$. The first step of our procedure is thus to estimate this autocovariance function. Though various estimators could in principle be used in conjunction with our identification results, we choose here to approximate the distribution of the observed data with a finite-order VAR (we discuss nonparametric consistency below). Note that this is an approximation of the reduced-form dynamics of the data; we do not need to assume a structural VAR model and the restrictive invertibility assumption that goes with it. Since our analysis in (ref) assumes stationarity, the data should be appropriately transformed and detrended prior to the analysis.\footnote{Because our analysis relies heavily on the spectral density matrix of the data, it is not straight-forward to extend our procedures to work directly with non-stationary data (without prior transformation/detrending). We leave this important topic to future research.}
As a first step, we select the VAR lag length $p$ by a standard information criterion, such as the Akaike Information Criterion (AIC). We then estimate a VAR($p$) model for the data $W_t=(y_t',z_t)'$ by OLS. Finally, we compute the VAR-implied estimates of the autocovariances and cross-covariances of $y_t$ and the projection residual $\tilde{z}_t \equiv z_t - E(z_t \mid \lbrace y_\tau,z_\tau \rbrace_{-\infty < \tau \leq t-1})$. Denote these estimates by $\widehat{\operatorname*{Var}}(\tilde{z}_t)$, $\widehat{\operatorname*{Cov}}(\tilde{z}_t, y_{t+h})$, and $\widehat{\operatorname*{Cov}}(y_t, y_{t-h})$ for $h=0,1,\dots$; see (ref) for explicit formulas.
Though our identification bounds below are valid irrespective of the invertibility of the shocks, some researchers may wish to have available a convenient pre-test of the null hypothesis of invertibility. (ref) showed that the distribution of the data is consistent with the shock of interest $\varepsilon_{1,t}$ being invertible if and only if the IV $z_t$ does not Granger cause the vector $y_t$ of macro observables. Intuitively, if the shock is invertible, then lags of the macro observables $y_t$ capture all the forecasting power of lags of the shock $\varepsilon_{1,t}$; hence, lags of the IV $z_t$ (a noisy measure of $\varepsilon_{1,t}$) do not contribute anything to forecasting. We can test the null hypothesis of no Granger causality in the following standard way:
Non-rejection should not be interpreted as strong evidence in favor of invertibility: Any valid test of invertibility necessarily has trivial power against some non-invertible alternatives, since it is possible for $z_t$ not to Granger cause $y_t$ even if the shock $\varepsilon_{1,t}$ is non-invertible.\footnote{Stock2018 develop an invertibility test which directs power against alternatives with impulse response functions that differ substantially from the invertible null. It is not immediately clear whether their test has power against all falsifiable non-invertible alternatives, as our proposed test does.}
We now describe how to estimate our identification bounds for the Forecast Variance Ratio (FVR) and the degrees of invertibility and recoverability. Bounds on Variance Decompositions (VDs) and Forecast Variance Decompositions (FVDs) are provided in (ref).
The bounds all depend on the two scalar quantities \[\hat{\underline{\alpha}}^2 \equiv \widehat{\operatorname*{Var}}(E(\tilde{z}_t \mid \lbrace y_\tau \rbrace_{-\infty<\tau<\infty})),\quad \hat{\bar{\alpha}}^2 \equiv \widehat{\operatorname*{Var}}(\tilde{z}_t),\] which are lower and upper bounds for $\alpha^2$, cf. (ref) and the discussion surrounding (ref). An explicit formula for the somewhat non-standard projection variance $\hat{\underline{\alpha}}^2$ -- as well as other, similar projection variances mentioned below -- is given in (ref).
\paragraph{Point identification/estimation under recoverability.} Finally, our analysis in (ref) showed that it is possible to point-identify many of the parameters of interest if the researcher is willing to impose additional a priori assumptions. In particular, if we are willing to assume that the shock is recoverable -- i.e., $R_\infty^2=1$ -- then the upper bound for $\mathit{FVR}_{i,\ell}$ in (ref) is a consistent estimator of the true FVR, as argued in (ref). As discussed in (ref), recoverability is a mathematically and economically weaker assumption than the invertibility assumption required by conventional SVAR-IV analysis.
In (ref) we prove that the above-mentioned bounds are jointly asymptotically normal under weak nonparametric regularity conditions on the data generating process (DGP). We assume neither that the true DGP is a finite-order VAR, nor that the shocks are Gaussian. This argument requires the VAR lag length $p=p_T$ used for estimation to diverge with the sample size $T$ at an appropriate rate.
Since the bounds are asymptotically normal, we can use standard arguments to construct confidence sets Imbens2004. Consider any one of the partially identified parameters discussed above and denote the estimates of its bounds by the generic notation $[\hat{\underline{\theta}},\hat{\bar{\theta}}]$. We then use a conventional bootstrap for VAR models Kilian2017 to generate bootstrap samples of the bound estimates $\hat{\underline{\theta}}$ and $\hat{\bar{\theta}}$, and let $\hat{\underline{q}}_{\beta}$ and $\hat{\bar{q}}_{\beta}$ denote the bootstrap $\beta$-quantiles of the lower and upper bounds, respectively. Then the interval $[\hat{\underline{q}}_{\beta/2},\hat{\bar{q}}_{1-\beta/2}]$ is a valid $1-\beta$ confidence interval for the identified set of the parameter in question.\footnote{The validity requires that the VAR bootstrap procedure is consistent. For example, the bootstrap must take into account the conditional heteroskedasticity of the data. See Kilian2017 for a menu of procedures, formal results, and regularity conditions.} That is, the probability that the confidence interval contains the entire identified set is greater than or equal to $1-\beta$ asymptotically; in particular, this confidence interval is therefore also a valid confidence set for the parameter itself.\footnote{In principle one could construct narrower confidence intervals that only guarantee coverage of the parameter itself (not the identified set), as in Imbens2004 and Stoye2009, and we do this in our Matlab code suite. However, the decrease in length appears to be minimal in realistic applications.} Under the additional point-identifying assumption that the shock is recoverable, the FVR is consistently estimated by the upper bound $\hat{\bar{\theta}}$, so we can construct a $1-\beta$ confidence interval as $[\hat{\bar{q}}_{\beta/2},\hat{\bar{q}}_{1-\beta/2}]$.
Because VAR inference is subject to well-known small-sample biases Kilian2017, we recommend that the following alternative formulas be used. Let $\hat{\underline{\theta}}^*$ and $\hat{\bar{\theta}}^*$ denote the average bootstrap draws of $\hat{\underline{\theta}}$ and $\hat{\bar{\theta}}$. Then we report the bias-corrected point estimate $[2\hat{\underline{\theta}}-\hat{\underline{\theta}}^*,2\hat{\bar{\theta}}-\hat{\bar{\theta}}^*]$ of the bounds, as well as Hall's percentile confidence interval $[2\hat{\underline{\theta}}-\hat{\underline{q}}_{1-\beta/2}, 2\hat{\bar{\theta}} -\hat{\bar{q}}_{\beta/2}]$. Similar corrections can be applied in the case of point identification via recoverability.
To illustrate our method, we revisit an old question: the importance of monetary shocks for U.S. macro fluctuations. Our main result is that monetary shocks are of limited importance for post-1990 aggregate dynamics, especially for inflation. The application illustrates that our upper bound on variance decompositions can yield surprisingly sharp inference, despite the weakness of our identifying assumptions.
\paragraph{Background.} Gertler2015 construct an external instrument for monetary shocks from high-frequency changes in asset prices in very short time windows around FOMC announcements, following earlier work by Kuttner2001, Cochrane2002, and Gurkaynak2005. While Gertler2015 focus on estimation of relative IRFs, we will seek to quantify shock importance, taking the validity of their instrument as given.\footnote{Caldara2019 compute FVDs for a similar specification, assuming an SVAR model. Their estimates of the importance of monetary shocks for inflation are somewhat larger than our upper bounds.} This setting is ideal for illustrating the appeal of our method, for two reasons.
First, measurement error is likely to be substantial. Intuitively, while short time windows around FOMC meetings may be a clean way of isolating some monetary shocks, all shocks occurring outside of that window are necessarily missed. Moreover, financial data are subject to noise due to market microstructure effects and uninformed traders. Treating the IV as the shock -- as in the method of Gorodnichenko2017, which is equivalent to our lower bound -- will then understate the importance of monetary shocks due to attenuation bias.\footnote{Formally, let the total monetary shock consist of two independent components, $\varepsilon_{1,t} \equiv \bar{\varepsilon}_{1,t} + \tilde{\varepsilon}_{1,t}$, where $\bar{\varepsilon}_{1,t}$ captures those shocks that occur inside FOMC announcement windows. Assume $\lbrace\bar{\varepsilon}_{1,t}, \tilde{\varepsilon}_{1,t} \rbrace$ are independent of $\lbrace \varepsilon_{2,t},\dots,\varepsilon_{n_\varepsilon,t}\rbrace$. If $z_t = \bar{\varepsilon}_{1,t} + \bar{v}_t$, where the noise $\bar{v}_t$ is independent of $\lbrace \bar{\varepsilon}_{1,t},\tilde{\varepsilon}_{1,t}, \varepsilon_{2,t},\dots,\varepsilon_{n_\varepsilon,t}\rbrace$, then the IV moment conditions (ref) are satisfied. For the case $\bar{v}_t=0$, our results in (ref) imply that the Gorodnichenko2017 FVR estimator will be biased downward by a factor of $\operatorname*{Var}(\bar{\varepsilon}_{1,t}) \in [0,1]$.} Second, non-invertibility is a threat to SVAR-IV analysis. For example, Ramey2016, citing the increasing prevalence of forward guidance in the conduct of U.S. monetary policy, cautions against the conventional SVAR-IV approach. In contrast, our partial identification approach does not require the shock to be invertible (or even recoverable).
\paragraph{Model.} Our specification largely follows Gertler2015, except that we do not impose a SVAR structure. We consider four endogenous macro variables $y_t$: output growth (log growth rate of industrial production), inflation (log growth rate of CPI inflation), the Federal Funds Rate (FFR), and the Excess Bond Premium of Gilchrist2012 as a measure of the non-default-related corporate bond spread. For robustness, we also try replacing the FFR with the 1-year Treasury rate, as in Gertler2015. The external IV $z_t$ is constructed from changes in 3-month-ahead futures prices written on the FFR, where the changes are measured over short time windows around Federal Open Market Committee monetary policy announcement times.\footnote{See Gertler2015 for details on the construction of the IV and a discussion of the exclusion restriction. Nakamura2018 argue that the monetary shock identified using this IV partially captures revelation of the Federal Reserve's superior information about economic fundamentals. This is related to the idea in Campbell2012 that monetary policy communication can be both “Delphic” and “Odyssean”. (ref) shows that our FVR bounds can generally be interpreted as bounding the importance of the particular linear combination of shocks that tend to hit during FOMC announcements, e.g., a weighted sum of “Delphic” and “Odyssean” shocks.} Data are monthly from January 1990 to June 2012. The AIC selects $p=6$ lags in the reduced-form VAR. We use 1,000 bootstrap draws from a homoskedastic recursive residual VAR bootstrap.
\paragraph{Results.} The data are consistent with substantial non-invertibility. (ref) shows point estimates and 90% confidence intervals for the identified sets of the degree of invertibility and the degree of recoverability, either using the FFR or the 1-year rate as the interest rate variable. When we use the FFR, we can reject invertibility at the 10% level, since the confidence set for the degree of invertibility excludes 1. When we use the 1-year rate, we cannot outright reject invertibility, but the confidence set is still consistent with very low degrees of invertibility.\footnote{The p-values for the Granger causality pre-test of invertibility in (ref) are 0.0001 (FFR) and 0.390 (1-year rate). Note that Stock2018 fail to reject invertibility in a somewhat different specification.} Since the data cannot rule out a low degree of invertibility in either case, we proceed with our invertibility-robust SVMA-IV analysis. The data are similarly consistent with a wide range of values for the degree of recoverability.
(ref) shows partial identification robust confidence intervals for the forecast variance ratio of the four endogenous macro variables with respect to the monetary shock. We report point estimates and confidence intervals for the identified sets at each horizon separately. We focus here on the specification with the FFR instead of the 1-year rate, since our quantitative conclusions are if anything even starker with the latter observable. At all forecast horizons, the 90% confidence intervals rule out FVRs above 31% for output growth and 8% for inflation. At forecast horizons up to 6 months, we can rule out that the monetary shock accounts for more than 19% of the forecast variance of the Excess Bond Premium. However, we cannot rule out that the monetary shock is an important contributor to medium- or long-run forecasts of the bond premium. On the other hand, we cannot rule out that the monetary shock is completely unimportant either.
Our analysis reveals that the weak assumptions of the SVMA-IV model suffice to obtain tight upper bounds on the forecast variance contribution of monetary shocks for several variables, especially inflation. This is despite the finding by Stock2018 that standard errors for impulse response functions are large in this application. Many commentators have documented a recent divorce between inflation and output dynamics Hall2011; our results document a similar divorce in dynamics conditional on monetary policy shocks in post-1990 data. Although this finding echoes previous SVAR work Christiano1999,Ramey2016, our identifying assumptions are weaker -- we merely impose validity of the IV.\footnote{(ref) reports variance decompositions obtained from a conventional SVAR-IV procedure. These results confirm the limited importance of the monetary shock, though under stronger identifying assumptions.} We conclude that, if inflation is a monetary phenomenon, it is so because of the systematic component of monetary policy, not because of erratic policy shocks.
\paragraph{Other application: Oil news shocks.} In (ref) we show that our method also yields highly informative upper bounds on the importance of international oil supply news shocks for the U.S. and global business cycles. We use an IV constructed by Kaenzig2021 from OPEC announcements.\footnote{Our empirical specification otherwise differs somewhat from his because we work with stationarity-transformed variables and restrict the sample to the period where the IV is available.} We find the oil news shock to be highly non-invertible, causing conventional SVAR-IV analysis to reach several spurious conclusions.
In this section we consider three simple analytical examples that illustrate how our identification bounds depend on the characteristics of the data. (ref) argued that the tightness of our lower bound on variance decompositions depends solely on the strength of the IV. We here show that, under stylized but empirically motivated assumptions, the upper bound can be expected to be highly informative, in the sense that it at worst mildly overstates the instrumented shock's contribution to macroeconomic fluctuations.
Throughout this section we assume the availability of a single IV $z_t = \alpha\varepsilon_{1,t} + \sigma_v v_t$, and then consider different illustrative toy models for the $y_t$ variables, as specified below. In (ref) we extend the analytical intuition below to the much richer quantitative DSGE model developed by Smets2007.
As our first example, consider the static model
with $n_y=2$ observables: inflation ($y_{1,t}$) and the monetary policy interest rate ($y_{2,t}$). We think of $\varepsilon_{1,t}$ as a conventional monetary policy shock. If we only observed inflation $y_{1,t}$ in addition to an IV, we would not be able to rule out that the monetary shock drives all the variation in this single variable, as explained in (ref). How can adding the interest rate to the data set help tighten the identified set for the FVR of inflation?
To gain economic intuition, let the reduced-form moments of the data be given by
where $\zeta \geq 0$. We see that, even with the amount of measurement error unknown, the IV $z_t$ reveals the signs and the relative magnitudes of the co-movement in observables induced by monetary shocks: The shock $\varepsilon_{1,t}$ moves inflation and interest rates in opposite directions Uhlig2005, while the unconditional correlation of interest rates and inflation is $\rho$.
We consider three instructive special cases. For the first two, we set $\zeta = 1$; applying our identification analysis, we then get the bounds
Now suppose first that $\rho = -1$; that is, interest rates and inflation are not just perfectly negatively correlated conditional on monetary shocks, but also unconditionally. In that case the upper bound for the FVR of both variables equals 1: The data cannot rule out that the correlation of the IV with macro observables is imperfect purely because of measurement error. Second, suppose that $\rho = 1$; that is, interest rates and inflation are perfectly positively correlated in the data. Then our upper bound for the FVR suddenly equals zero: The monetary shock induces co-movement patterns that we never see in the data, so it cannot possibly explain any observed macro fluctuations. Third, if instead $\rho = 0$, then our upper bounds are, for any $\zeta \geq 0$,
Suppose that nominal rates respond much more to the monetary shock than inflation does, i.e., $\zeta \gg 1$. Then the upper bound on the inflation variance decomposition $\mathit{FVR}_{1,0}$ is very small; intuitively, since the IV reveals that the monetary shock moves interest rates by much more than inflation, but both have the same unconditional variance, the monetary shock cannot possibly be an important driver of inflation. This third example rationalizes the findings in our application to monetary shocks in (ref): The IV $z_t$ correlates much more with interest rates than with prices, yet prices are not commensurately less volatile than interest rates, so monetary shocks cannot account for much of the volatility in prices.
This example shows that our upper bound on the FVR is close to the true value if either the shock is very prominent (so that the bound of one is not far from the truth) or if the shock induces somehow atypical co-movements of the various observed macro aggregates. This second condition is equivalent to the shock being prominent for some linear combination of the macro observables $y_t$, which is equivalent to the shock being nearly invertible in this static model. Thus, the preceding arguments agree with the analysis in (ref).
Whereas the previous example illustrated how the availability of several macro time series sharpens identification in a static context, we now show how the dynamics of individual time series can do the same. Consider the univariate but dynamic model
with $n_y = 1$. That is, we observe a single variable $y_t$ driven by $n_\varepsilon$ independent AR(1) processes. To fix ideas, we think of $\varepsilon_{1,t}$ as a technology shock and $y_t$ as aggregate output.
Now assume that long-run fluctuations in output $y_t$ are exclusively driven by the technology shock $\varepsilon_{1,t}$; that is, consider the limit $\rho_1 \rightarrow 1$, while fixing $|\rho_j|<1$ for all $j \geq 2$. In this case, the sharp lower bound on $\alpha^2$ converges to the truth:\footnote{Note that $s_{y\tilde{z}}(0) = \frac{\alpha}{2\pi}\sum_{\ell=0}^\infty \rho_1^\ell = \frac{\alpha}{2\pi} \times \frac{1}{1-\rho_1}$ and $s_y(0) = \frac{1}{2\pi}(\sum_{j=1}^{n_\varepsilon}\sum_{\ell=0}^\infty \rho_j^\ell)^2 = \frac{1}{2\pi}(\sum_{j=1}^{n_\varepsilon} \frac{1}{1-\rho_j})^2$. (ref) then implies $2\pi s_{\tilde{z}^\dagger}(0) = s_{y\tilde{z}}(0)^2/s_y(0) = ((1-\rho_1)s_{y\tilde{z}}(0))^2/((1-\rho_1)s_y(0)) \to \alpha^2$ as $\rho_1\to 1$.} \[\lim_{\rho_1 \to 1} \alpha_{LB}^2 = \lim_{\rho_1 \to 1} 2 \pi \sup_{\omega \in [0,\pi]} s_{\tilde{z}^\dagger}(\omega) = \lim_{\rho_1 \to 1} 2 \pi s_{\tilde{z}^\dagger}(0) = \alpha^2.\] Intuitively, at spectral frequency zero, all fluctuations in $y_t$ are driven by the technology shock. Loosely speaking, applying a low-pass filter to $y_t$ therefore isolates the fluctuations caused by the technology shock. Leads and lags of this low-pass filtered series are thus highly correlated with the IV, putting a lower bound on the signal-to-noise ratio in the IV. This is why the sharp upper bound for the FVR converges to the true value.
The example reveals that cross-restrictions over time can be highly informative even if the shock of interest is neither invertible nor recoverable. Intuitively, for the sharp upper bound on the FVR to bind, our method only needs the shock to dominate at some frequency; the across-frequency restrictions then do the rest, exactly like the cross-variable restrictions in the static example above.\footnote{As mentioned in (ref), for finite-sample statistical reasons, we recommend the use of a weaker lower bound $\underline{\alpha}^2$ on $\alpha^2$ in place of the sharp bound $\alpha_{LB}^2$. This weaker lower bound does not converge to $\alpha^2$ as $\rho_1 \to 1$, unless $\varepsilon_{1,t}$ is recoverable. However, as discussed in (ref), researchers may leverage a strong prior belief about the low-frequency importance of shocks by computing the integral in (ref) for a pre-specified range of (low) frequencies.}
In the third example, we show how our method deals with non-invertible news shocks. First discussed in Pigou1927, news shocks have recently received much attention as drivers of macroeconomic fluctuations Beaudry2006,Beaudry2014,Jaimovich2009,Schmitt2012. Unfortunately, foresight of economic agents complicates conventional SVAR-based analysis since it induces equilibria with non-invertible MA representations Leeper2013. In contrast, our methods are valid irrespective of invertibility.
To illustrate, consider a moving average model of order 1 with $n_y = n_\varepsilon = 2$:
where $\zeta > 1$. As is well known, this assumption implies that the moving average representation is non-invertible. We think of $\varepsilon_{1,t}$ as a monetary forward guidance shock: The shock moves inflation and nominal interest rates by more tomorrow (when the shock directly hits the monetary policy rule) than today (when the news is revealed).
The conventional SVAR-IV approach mis-measures the FVR because of non-invertibility. By standard arguments Leeper2013 the reduced-form VAR residuals equal
where the degree of invertibility equals
Since SVAR procedures assume that the structural shocks $\varepsilon_t$ can be obtained as linear functions of the reduced-form residuals $u_t$, equation (ref) shows that any SVAR analysis will conflate the explanatory power of the shock $\varepsilon_{1,t}$ with that of its lags. As a consequence, (ref) shows that the SVAR-IV estimand of the FVR overstates the contribution of the shock to one-step-ahead forecasts:
Clearly, the population bias of the SVAR-IV estimand worsens as the degree of invertibility $R_0^2$ decreases to 0 Forni2018. In the oil news shock application in (ref) we demonstrate that the SVAR-IV bias can be large in practice.
In contrast, our identification bounds are valid irrespective of invertibility, since we do not assume that $\varepsilon_{1,t}$ can be recovered as a function of only the contemporaneous VAR residuals $u_t$. In fact, in this model with as many observables as shocks, both shocks $\varepsilon_t=(\varepsilon_{1,t},\varepsilon_{2,t})'$ are recoverable.\footnote{In particular, $\varepsilon_{t} = -R_0^2 \Theta_0^{-1} u_t - \left(1 - R_0^2 \right)\sum_{\ell=1}^\infty \left( -\zeta\right)^{-\ell} \Theta_0^{-1} u_{t+\ell}$.} Hence, if we exploit this knowledge, we can even point-identify the shock as $\varepsilon_{1,t} \propto z_t^\dagger = E(z_t \mid \lbrace y_\tau \rbrace_{-\infty<\tau<\infty})$. The key is that our method can use the future values of nominal rates and inflation, $y_\tau$, $\tau \geq t$, to recover the forward guidance shock $\varepsilon_{1,t}$ at time $t$. In so doing, it effectively realigns the information sets of the economic agents and of the econometrician, sidestepping the invertibility problem.
We finish by showing that our inference procedures have good finite-sample performance in simulations. Our methods continue to work well in non-invertible models, unlike the conventional SVAR-IV procedure.
\paragraph{DGP.} We adopt a variant of the DGP in Kilian2011 and assume that the macro aggregates $y_t$ follow a structural VARMA($p$,1) model:
We consider $n_y = 2$ macro variables, $p=1$ autoregressive lag (with one exception discussed below), and set $\Xi_1 = \left(
\right)$. For the MA part, we consider $n_\varepsilon = 2$ shocks (which are thus both recoverable) and set $\Theta_0 = chol \left(
\right),$ where ``chol'' denotes the lower triangular Cholesky decomposition. As in \cref{sec:illustration_news}, $\zeta$ is a scalar parameter that governs the degree of invertibility, with $\zeta>1$ implying non-invertibility. We add an external instrument $z_t$ for the shock of interest $\varepsilon_{1,t}$:
Notice that we have normalized $\alpha = 1$. Finally, the measurement error and structural shocks are i.i.d. Gaussian and orthogonal as in (ref).
We run Monte Carlo experiments for nine different parameterizations of the above DGP. Specifically, we consider various deviations from a baseline parametrization. In our benchmark, we set $\rho_y = 0.5$, $\rho_z = \rho_{zy} = 0$, $\zeta = 0$, $\sigma_v = 1$, and sample size $T = 250$. We then consider variations with more autoregressive persistence (either $\rho_y = 0.9$, or $\rho_z = 0.8$ and $\rho_{zy} = 0.3$), an invertible MA component ($\zeta=0.5$), a non-invertible MA component ($\zeta = 2$), a weaker instrument ($\sigma_v = 2$), and different sample sizes ($T = 100$, $T = 500$). Finally, we allow for richer dynamics, with $p=4$ and $\Xi_j = \frac{1}{j^2} \Xi_1$ for $j = 2,3,4$.
\paragraph{Results.} Our parameters of interest are the degree of invertibility $R_0^2$ and the FVR for variable $y_{2,t}$ at horizons $1$ and $4$. We conduct $5,000$ Monte Carlo repetitions per DGP, and construct confidence intervals at the 90% level using $1,000$ bootstrap draws per simulation. We use a homoskedastic recursive residual bootstrap. The reduced-form VAR lag length is selected using AIC, and we use Hall's percentile bootstrap confidence interval, cf. (ref).
(ref) shows that the partial identification robust SVMA-IV confidence sets defined in (ref) achieve coverage rates close to or exceeding the desired level of 90% throughout. We report coverage rates for both the population identified sets (columns “Set”) and for the underlying parameters (columns “Param”). The coverage rate for the parameter is never below 86.8% in any case. The coverage rate for the identified set is mostly close to 90% and at worst 82.9% in our experiments. We also report the coverage rates of conventional SVAR-IV bootstrap confidence intervals for the FVR (columns “SVAR”). The coverage distortions of our SVMA-IV procedures are almost always smaller than those of the SVAR-IV procedure. Most notably, our procedures have acceptable coverage even in the non-invertible case ($\zeta=2$), whereas the SVAR-IV procedure under-covers severely in this case.\footnote{We acknowledge, however, that in DGPs with only mild non-invertibility, SVAR-IV procedures may be preferable to our more robust SVMA-IV procedure, since the former procedure has fewer parameters to estimate and will be only mildly biased (cf. (ref)).}
We make the following additional remarks. First, coverage deteriorates slightly with noisier/weaker instruments ($\sigma_v=2$), as expected. Our inference methods are not robust to arbitrarily weak instruments ($\sigma_v \to \infty$); we leave this issue to future work. Second, we face some well-known parameter-at-the-boundary issues. For most experiments, $R_0^2 = 1$. This explains the over-coverage of confidence intervals for this parameter and, less so, for the overall identified set. Similar problems would arise if the true FVR were close to $0$. Third, for more persistent DGPs, the AIC tends to select an insufficient number of lags, resulting in moderate under-coverage, in particular for the FVRs at horizon 4. For example, in the experiment with $p = 4$ autoregressive lags, the AIC selects an average lag length of $2.2$.
Applied macroeconomists have recently turned to external sources of exogenous variation to identify dynamic causal effects. Though such external instruments or proxies are frequently used to estimate impulse responses, existing methods did not allow researchers to quantify the contribution of individual shocks to business-cycle fluctuations -- a question of first-order interest in traditional business-cycle analysis. We fill this gap by providing identification results and inference techniques for variance decompositions, historical decompositions, and the degree of invertibility. Our methods require neither the absence of measurement error in the external instrument, nor the often dubious assumption that the instrumented shock is invertible (as assumed in conventional SVAR analysis). We prove that the importance of the instrumented shock is generally interval-identified. Point identification can be achieved if the shock is known to be recoverable -- a substantively weaker assumption than invertibility. We provide a software package that implements all steps of our inference procedures. Applying our method to U.S. data, we are able to establish a tight upper bound on the importance of monetary shocks for recent inflation dynamics, despite our weak identifying assumptions.