Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,370 characters · 16 sections · 91 citation commands
Beyond Validity: SVAR Identification Through the Proxy Zoo
\doublespacing
Identification through external instruments---or proxies---is widely used in structural vector autoregressions (SVARs). By exploiting variables that are correlated with a specific structural shock while remaining orthogonal to all others, the proxy-SVAR approach delivers point-identified impulse responses while maintaining a flexible reduced-form specification mertens2013dynamic,stock2018identification. This approach has been adopted across a wide range of applications, including monetary policy gertler2015monetary, fiscal policy mertens2013dynamic, oil shocks kanzig2021macroeconomic, uncertainty shocks angelini2019exogenous, and financial shocks ottonello2025financial, among many others.
At the same time, this empirical success has brought into focus a fundamental tension. The orthogonality requirement is often difficult to justify and impossible to test directly. This has led to a proliferation of proxies intended to identify the same structural shock---a proxy zoo---constructed using different data sources and methodologies, and to a central debate over proxy exogeneity. Recent work documents non-negligible correlations between widely used proxies and non-target shocks, challenging exact exogeneity as a maintained assumption. For example, schlaak2023monetary raise concerns about the exogeneity of high-frequency monetary policy surprises, while the narrative shocks of romer2004new have long been argued to reflect oil-related disturbances barnichon2025innovations.
Instead of resolving disputes over proxy exogeneity, we develop a framework for robust identification in SVARs with a proxy zoo.\footnote{Although we develop the framework in the context of SVARs, the approach extends naturally to local projections with external instruments following plagborg-moller2021local.} Central to identification is a set of generalized ranking restrictions (GRR) on the relative correlation of each proxy with the target and non-target shocks, governed by a continuous proxy-quality parameter. We show three main results. First, we characterize the identified set for structural impulse responses under GRR combined with standard sign and narrative restrictions. Second, we show how to partially identify the proxy-quality parameter using the joint information contained in the proxy zoo and these additional identifying restrictions. Third, we develop a suite of sensitivity and diagnostic tools in the spirit of manski2003partial that allow researchers to assess transparently how substantive conclusions depend on assumptions about proxy exogeneity.
The identifying power of the GRR is shaped by two distinct forces: (i) the proxy-quality parameter and (ii) proxy complementarity. The proxy-quality formulation delivers transparency and robustness. It requires researchers to state explicitly the degree of exogeneity violations they are willing to tolerate, with the conventional valid-IV assumption emerging only as a limiting case with zero contamination. The resulting identified sets contain the true impulse responses whenever the true proxy quality exceeds the imposed level. Complementarity arises when proxies carry distinct, falsifiable information masten2021salvaging, restricting different directions in the space of admissible structural representations. Strikingly, when proxies are sufficiently complementary, point identification can be recovered even when all of them are contaminated.\footnote{Because informative identification is possible without valid instruments, the framework encourages the construction of new proxies under weaker conditions than classical exogeneity, broadening the scope of external information that can be brought to bear on structural identification.}
Crucially, the quality parameter governing the GRR need not be calibrated ex ante. Instead, we exploit two sources of falsifiable information in the proxy zoo to derive an application-specific upper bound on proxy quality. First, once auxiliary economic restrictions, such as sign or narrative constraints, are imposed, the data themselves restrict the range of contamination levels compatible with those restrictions. Second, when proxies contain conflicting information, each proxy rules out structural representations that others would admit, so that joint feasibility disciplines admissible contamination levels without parametric modeling of endogeneity nguyen2025bayesian.
Having established an upper bound on proxy quality, we recommend reporting two diagnostic measures. First, the breakdown frontier, which records the minimum level of proxy quality required to sustain empirical claims of interest. Second, proxy informativeness, which quantifies the contribution of each proxy to tightening the identified set and detects redundancy among proxies. These diagnostics complement standard sensitivity analysis by clarifying how empirical conclusions depend on the composition of the proxy zoo, thereby connecting proxy-SVAR identification to recent advances in sensitivity analysis and research transparency vanderweele2017sensitivity,andrews2020transparency,masten2020inference. In our monetary policy application, these diagnostics reveal substantial heterogeneity in the identifying content of commonly used proxies, with a small subset driving most of the tightening of the identified set and several others contributing little additional information once the proxy zoo is considered jointly.
A simulation study based on a medium-scale dynamic stochastic general equilibrium (DSGE) model smets2007shocks further demonstrates that treating contaminated proxies as valid instruments---such as proxies constructed from sign-restriction procedures baumeister2019structural,jarocinski2020deconstructing---can produce substantial bias while preserving the expected sign of responses, making contamination difficult to detect from “puzzles” alone.\footnote{The issue does not stem from sign restrictions as an identification device, but from common procedures that map sign-restricted SVARs into a single proxy series for external use. These constructions need not preserve orthogonality with non-target shocks.} In contrast, our method recovers the true impulse responses within robust identified sets. Empirically, we illustrate the method using a rich set of U.S.\ monetary policy proxies drawn from narrative, high-frequency based, and model-based approaches. We show that conventional proxy-SVAR estimates are sensitive to the choice of proxy, while our set-identified approach delivers informative and transparent bounds even without assuming any proxy to be exogenous. In this application, the data imply a finite upper bound on the proxy-quality parameter, indicating that the joint restrictions imposed by the proxy zoo are not compatible with exact exogeneity for all proxies simultaneously.
\noindentLiterature. Our framework contributes to several strands of the literature. First, it nests popular identification schemes in SVARs as special cases. By varying the proxy-quality parameter, our GRR reduces to the least restrictive external-variable constraints ludvigson2021uncertainty, to the more informative ranking-based identification braun2023identification, and to the classical proxy-SVAR framework based on exact exogeneity mertens2013dynamic. Our auxiliary restrictions encompass both conventional sign restrictions on impulse responses uhlig2005what and narrative-based restrictions antolin-diaz2018narrative, making explicit how empirical conclusions depend on identifying assumptions.
Second, the paper relates to recent work addressing proxy contamination while preserving identification. Prominent approaches exploit volatility changes schlaak2023monetary,angelini2025invalid, higher-order moment restrictions keweloh2025estimating, combinations thereof carriero2024blended, or innovations-powered inference barnichon2025innovations. Our framework differs by treating all proxies as potentially contaminated and exploiting their joint restrictions for identification of a single shock, rather than relying on auxiliary statistical structure.
Third, our partial identification of proxy quality connects to the literature on overidentification tests, dating back to sargan1958estimation, and their recent applications to assess proxy exogeneity schlaak2023monetary,bruns2024testing,angelini2025test,angelini2025invalid. In contrast, we exploit the direction of endogeneity bias---revealed by overidentifying restrictions---to learn about the parameter of interest masten2021salvaging.
Finally, the framework extends the proxy-SVAR approach to sensitivity analysis manski2003partial,andrews2017measuring,masten2020inference, and relates to recent advances on “plausibly exogenous” instruments conley2012plausibly and regression sensitivity analysis kiviet2020testing,cinelli2020making.
Outline. The remainder of the paper is organized as follows. Section (ref) provides the motivating descriptive evidence that illustrates potential contamination of popular monetary policy proxies. Section (ref) presents the general framework for set-identified proxy-SVAR with the generalized ranking restrictions. The proxy quality parameter is partially identified in Section (ref), based on which a set of diagnostic tools for sensitivity analysis is discussed in Section (ref). Section (ref) presents the results of a simulation study and Section (ref) revisits the monetary policy application. Section (ref) concludes.
We illustrate the empirical challenges of conventional proxy-SVAR identification using the case of U.S.\ monetary policy shocks. Rather than assessing the validity of individual proxies, this section documents that exact exogeneity imposed jointly across a proxy zoo is a strong maintained assumption. We show that monetary policy proxies both deliver heterogeneous impulse responses and exhibit nontrivial correlations with proxies for other structural shocks.
We estimate a standard seven-variable monthly VAR with industrial production (INDPRO), CPI inflation (CPIAUCSL), the unemployment rate (UNRATE), the CRB commodity price index (CRBPI), the one-year Treasury rate (GS1), financial market prices (S&P500), and the excess bond premium (EBP), using data from 1973m1 to 2019m12.\footnote{The specification includes 12 lags and a constant.} The specification follows miranda-agrippino2021transmission and is widely used in empirical analyses of U.S.\ monetary policy.
We consider eight external instruments that are among the most widely used in the literature: the narrative series of romer2004new (RR); the target-factor proxy of gurkaynak2005actions (GSST);\footnote{Constructed from the authors' replication of gurkaynak2005actions.} the high-frequency innovation of gertler2015monetary (GK); the series of nakamura2018highfrequency (NS);\footnote{Obtained from the supplementary materials of brennan2025monetary.} the unified shock measure of bu2021unified (BRW);\footnote{We use the cumulative-sum proxy following bu2021unified.} the proxy of miranda-agrippino2021transmission (MR); the sign-restriction-based measure of jarocinski2020deconstructing (JK); and the high-frequency surprise of bauer2023reassessment (BS). These proxies are constructed under different principles---narrative analysis, high-frequency surprises, factor decomposition---but are all intended to identify the same underlying monetary policy shock, even though they may capture different aspects of policy actions and information releases.
(ref) reports the impulse responses obtained from point-identified proxy-SVARs when each proxy is used individually. The resulting responses differ substantially across instruments, both in magnitude and in dynamic shape. In some cases, the implied policy rate path displays patterns that are not typically associated with contractionary monetary policy, such as a negative response of short-term interest rates at medium horizons. Such heterogeneity does not permit conclusions about the validity or invalidity of any particular proxy; rather, it highlights the sensitivity of point-identified inference to the choice of instrument.
A key implication of the conventional approach is that, if all proxies satisfied the identifying assumptions and were sufficiently informative, they would deliver identical impulse responses. Substantial discrepancies across instruments therefore indicate that inference based on a single proxy may be fragile in practice. As an illustration, two influential proxies, MR and BS, disagree on the sign of the identified structural shock in 43.75% of the months between 1979 and 2019; see (ref) in the appendix. This disagreement is not diagnostic of invalidity, but it underscores the fragility of inference based on any single proxy and motivates identification strategies that remain informative when proxies provide conflicting signals.
In principle, exogeneity restrictions can be assessed in overidentified systems using tests of overidentifying restrictions sargan1958estimation,hansen1982large. However, the interpretation of such tests relies on maintaining that at least one proxy is exactly exogenous kiviet2020testing, which becomes a strong maintained assumption when all available proxies may be subject to some degree of contamination.
As complementary descriptive evidence, we examine pairwise correlations between the monetary policy proxy zoo and a broad set of proxies for other structural shocks, including oil supply, financial, fiscal, and technology shocks.\footnote{We consider 19 non-monetary proxies from kanzig2021macroeconomic, baumeister2019structural, ottonello2025financial, ramey2018government, fisher2010using, benzeev2017chronicle, leeper2012quantitative, beaudry2006stock, fernald2014quarterly, benzeev2015investmentspecific, benzeev2018what, and francis2014flexible. Some series are available only at quarterly frequency; for these, we sum the monetary policy instruments within quarters.} Under an ideal benchmark where structural shocks are orthogonal and measurement errors are independent, exogeneity implies zero correlation across proxies for distinct shocks. While non-zero correlations can arise from sampling variation or common information, we treat these patterns as descriptive indicators of potential contamination risks.
(ref) reports the correlation-significance map. Two broad patterns emerge. First, fiscal and financial shocks appear to be common sources of contamination. Most monetary proxies exhibit strong positive correlations with the government spending news shock of benzeev2017chronicle (ranging from 0.15 to 0.43). Correlations are even more prominent with the ramey2018government military news shock (ranging from 0.15 to 0.66), except for the RR proxy. Similarly, most monetary proxies show negative correlations with the sign-restricted financial proxy of ottonello2025financial, while some---specifically RR and BRW---also correlate significantly with the high-frequency-based financial proxy. Second, idiosyncratic contamination patterns are prevalent. For instance, the NS proxy correlates significantly with oil news shocks and various TFP measures, whereas the BRW proxy displays uniquely high correlations with tax shocks (0.56) and defense spending news (0.66). Taken together, these patterns illustrate that exact exogeneity imposed on individual proxies---or jointly across a proxy zoo---may be delicate in practice, motivating identification strategies that remain informative under limited forms of contamination.
Consider a standard $n$-variate SVAR($p$) model, with $p<\infty$:
where $Y_t$ is the vector of observables, $ u_{t} $ is the vector of reduced-form innovations, $\epsilon_t$ is the vector of structural shocks, and $B$ is the structural matrix capturing the contemporaneous effects of the structural shocks on the variables $Y_t$. We impose the following standard regularity conditions.
Assumption (ref) collects standard conditions ensuring that a structural VAR is well-defined. Part (ref) imposes stability of the reduced form, which guarantees the existence of a causal VAR representation and well-defined impulse responses. Part (ref) specifies that structural shocks are orthonormal innovations, which normalizes shock variances to unity and ensures shocks correspond to economically distinct primitive forces. Finally, Part (ref) requires $B$ to be invertible. This ensures that the reduced-form innovations have a non-degenerate structural representation and that each structural shock corresponds to a distinct direction in the space of reduced-form errors stock2018identification. Identification procedures, including sign restrictions and external instruments, restrict the set of admissible structural matrices.
Without loss of generality, we focus on identifying the dynamic causal effects of the first structural shock $ \epsilon_{1,t} $. The impulse response of variable $i$ to shock $\epsilon_1$ at horizon $h$ is given by the $(i,1)$-th element of $ C_{h}B $, where
and $A_\ell=0_{n\times n}$ for $\ell>p$; see kilian2017structural for a detailed derivation. Thus, the response of variable $i$ to shock $j$ at horizon $h$ can be written as
where $e_i$ denotes the $i$-th elementary basis vector of $\mathbb{R}^n$. Since each $C_h$ depends only on the autoregressive coefficients, which are consistently estimable by OLS, identification reduces to recovering (a column of) the structural matrix $B$.
\paragraph{Point identification via valid instruments.} In the presence of a valid instrument $ m_{t} $ for the target shock $\epsilon_{1,t}$, point identification of the first column of $B$ can be achieved via the proxy-SVAR approach mertens2013dynamic. Specifically, if $ \mathbb{E}[m_{t}\epsilon_{1,t}] \neq 0 $ and $ \mathbb{E}[m_{t}\epsilon_{j,t}] = 0 $ for all $ j \geq 2 $, then
The structural responses to shock $\epsilon_{1,t}$ are thus identified up to a scaling factor by combining the second-moment restrictions with the valid-IV assumption. However, instrument validity is often questionable in practice. With contamination---$\mathbb{E}[ m_t \epsilon_{j,t}] \neq 0$ for some $j \geq 2$---the standard approach identifies only a linear combination of structural responses:
Identification breaks down unless responses to different non-target shocks are proportionally similar, which is unlikely in practice.
\paragraph{Set identification approach.} Instead of pursuing point identification, we adopt a partial identification framework that characterizes the set of structural matrices $B$ compatible with both the statistical properties of the reduced-form VAR and economic restrictions. The reduced-form VAR model already provides $n(n+1)/2$ second-moment restrictions from the reduced-form covariance matrix:
While insufficient to uniquely pin down all $ n^2 $ elements of $ B $, these restrictions provide a useful starting point. Let $L$ denote the unique lower-triangular Cholesky factor of $\Sigma$ such that $\Sigma = LL'$, and let $\mathcal{O}(n) = \{Q \in \mathbb{R}^{n \times n}: QQ' = I_n\}$ denote the set of orthogonal matrices of dimension $n$. Then any matrix $\tilde{B} = LO$ with $O \in \mathcal{O}(n)$ automatically satisfies the second-moment restrictions:
Thus, $B$ is identified only up to an orthogonal rotation $O$. Set identification reduces to characterizing the orthogonal matrices $O$ consistent with additional economic restrictions introduced below.
For future usage, we define the proxy moment vector $M_\ell \coloneqq \mathbb{E}[L^{-1}u_t\, m_{\ell,t}]$, and we collect the VAR coefficients, the covariance elements, and the proxy moments into a single vector:
where $\text{vec}(\cdot)$ stacks matrix columns and $\text{vech}(\cdot)$ stacks the lower-triangular elements.
We introduce an identification strategy that exploits external proxy variables while relaxing the conventional valid-IV assumption. Let $m_t=(m_{1,t},\ldots,m_{k,t})'$ denote a vector of $k$ external variables intended to contain information about the target shock $\epsilon_{1,t}$. We refer to the collection $(m_{1,t},\ldots,m_{k,t})$ as the proxy zoo.
The classical proxy-SVAR framework requires each proxy to be uncorrelated with all non-target shocks, that is, Equation (ref) holds. In many applications, however, this assumption is empirically difficult to maintain: variables designed to capture a specific structural shock often respond, at least weakly, to other economic forces. Once validity is relaxed, the relevant question is therefore not whether contamination is present, but whether it is sufficiently limited for the proxy to remain informative about the target shock.
The following example illustrates this issue and motivates our approach.
Example (ref) makes clear that, once validity is abandoned, identification hinges on the relative strength of a proxy's correlation with the target shock relative to non-target shocks. We formalize the comparison between these correlations through a proxy-specific quality parameter that bounds how strongly a proxy may co-move with non-target shocks relative to the target shock. Importantly, this parameter is not calibrated ex ante: in later sections, we show that it is itself partially identified from the joint restrictions implied by the data and the maintained identifying assumptions.
Assumption (ref) imposes standard regularity conditions on the proxies, as commonly assumed in the proxy-SVAR literature mertens2013dynamic. Assumption (ref) formalizes the intuition in Example (ref) by parameterizing proxy quality in terms of the relative magnitude of correlations with the target and non-target shocks. Specifically, for a given proxy, the parameter $\tau_{\ell,0}$---defined as the minimum $\tau_{\ell,j,0}$ across non-target shocks---summarizes how large the proxy's correlation with the target shock is relative to its correlations with non-target shocks. In Example (ref), $\tau_{\ell,0}=3$, reflecting that the proxy's correlation with the monetary-policy shock exceeds its correlation with each non-target shock by a factor of three, at least. Any finite value of $\tau_{\ell,0}$ permits contamination while indexing how informative the proxy is about the target shock. Assumption (ref) is operationalized as a set of ranking restrictions on the rotation matrix $O.$
\paragraph{Operationalizing the proxy assumption.} The contamination assumption (ref) is stated in terms of population correlations between proxies and structural shocks, which are unobservable. The next result expresses this as explicit constraints on $O$ and a consistently estimable moment vector.
Since $M_\ell$ is consistently estimable, this proposition yields a set of implementable linear inequality constraints on $O$. We refer to the linear inequality constraints in (ref) as the generalized ranking restrictions (GRR).
\paragraph{Relationship to existing assumptions.} By allowing researchers to explicitly specify the proxy quality parameter $ \tau_{\ell,0} \in [0,\infty] $, our contamination assumption encompasses several popular identifying assumptions. First, Example (ref) shows that the limiting case $\tau_{\ell,0} = \infty$ corresponds to the conventional valid-IV assumption mertens2013dynamic,stock2018identification. Assuming a finite $ \tau_{\ell,0} $ then strictly weakens the exogeneity requirement. Second, setting $\tau_{\ell,0} = 1$ recovers the ranking assumption of braun2023identification as a special case, where the proxy must be at least as correlated with the target shock as with any other shock. Finally, specifying $\tau_{\ell,0} = 0$ leads to the “external variable constraints” of ludvigson2021uncertainty, which discard the interpretation of $m_{\ell,t}$ as a proxy for the target shock and only require $m_{\ell,t}$ to be correlated with the shock of interest.\footnote{Allowing for a positive slack term on the right-hand side of (ref) strengthens this requirement by imposing a minimum correlation threshold without invoking any validity assumption.} These assumptions may be unnecessarily strong or weak---e.g., misspecifying $\tau_{\ell,0}=1$ for a valid IV will lead to a valid but unnecessarily wide set of responses. Our framework allows researchers to discipline the degree of contamination and conduct sensitivity analysis, as specified in Section (ref).
We combine the ranking restrictions with economically motivated sign restrictions. These constraints reflect the researcher's prior knowledge about either (i) the direction of specific impulse responses, or (ii) the sign of structural shocks at specific historical dates. We show that both restrictions can be cast as linear inequalities on columns of the rotation matrix $O$.
\paragraph{IRF-based sign restrictions.} Given the definition of impulse response in Equation (ref), restrictions of the form $\theta_{i,j,h} \ge 0$---requiring the response of variable $i$ to shock $j$ at horizon $h$ to be non-negative---translate into
We collect indices of such restrictions in $\mathcal R_{\mathrm{IRF}} \subset \{1,\ldots,n\}^2 \times \{0,\ldots,H\}$. For notational brevity, we suppress the dependence of sign restriction vectors $r_{i,j,h}$ on reduced-form parameters $\phi$. This general formulation encompasses several standard identification schemes, as illustrated below.
\paragraph{Narrative restrictions.} Following antolin-diaz2018narrative, we may constrain shock signs at specific dates. Suppose narrative evidence implies $\epsilon_{j,t_0}>0$ for some shock $j$ at date $t_0$. Using the structural mapping $\epsilon_{t} = O' L^{-1} u_{t}$, we have
This narrative information translates into the restriction
We collect such restrictions in $\mathcal R_{\mathrm{nar}} \subset \{1,\ldots,n\} \times \mathbb T$, where $\mathbb T$ denotes the set of dates with narrative information. Again, the dependence of $r_{j,t_0}$ on $\phi$ is suppressed when the context is clear.
\paragraph{Admissible set.} Combining IRF and narrative restrictions, the sign-restriction-admissible set is
Each $O \in \mathcal{F}_{\textup{sign}}(\phi)$ generates a structural matrix $B = LO$ consistent with the imposed sign restrictions. The associated set for the target structural vector is
where $\mathbb{S}^{n-1} \coloneqq \{x \in \mathbb{R}^n: \|x\| = 1\}$ denotes the unit $(n-1)$-sphere. Throughout, we assume that the restrictions are correctly specified:
Assumption (ref) is standard in the sign restrictions literature uhlig2005what,baumeister2015sign,antolin-diaz2018narrative, which ensures the feasible set $\mathcal{F}_{\textup{sign}}(\phi_0)$ is non-empty. Note that non-emptiness is necessary but not sufficient for valid identification---misspecified sign restrictions may yield non-empty sets of responses that exclude the true structural impulse response. Assumption (ref) rules out such misspecification.
We achieve identification by combining generalized ranking restrictions with sign restrictions.\footnote{Practitioners may not always have prior knowledge on sign restrictions. In this case, we recommend imposing self-sign restrictions only as in (ref), which serve as a normalization without loss of generality.} In practice, researchers rarely possess reliable prior information to calibrate proxy-specific quality parameters $\tau_{\ell,0}$. We therefore impose a common quality parameter $\tau$ across all proxies and work with the homogeneous GRR:
The feasible set depends on a collection of reduced-form parameters through (ref) and (ref). Since $\{ C_h \}_{h=0}^{H}$, $L$, and $ \{ M_{\ell} \}_{\ell=1}^{k} $ are continuous functions of $\phi$, the feasible set defines a correspondence in $(\tau, \phi)$.
Throughout, the dependence on $\phi$ is often suppressed and reintroduced only in the asymptotic analysis.
Under Assumptions (ref)-(ref), if the researcher specifies $\tau = \tau_0 \coloneqq \min_{\ell=1,\dots,k}\tau_{\ell,0}$, then the true impulse responses satisfy $\theta_{i,1,h}^0 \in \Theta_{i,h}(\tau_0)$ for all $i$ and $h$, achieving valid identification. Moreover, the researcher's choice of $\tau$ determines the width and validity of the identified set through the following monotonicity property.
This proposition states that as $ \tau $ increases, the identified set shrinks. This reflects a fundamental trade-off between identifying power and robustness: Specifying $\tau > \tau_0$ yields tighter bounds but risks excluding the truth, while specifying $\tau < \tau_0$ guarantees coverage of the truth at the cost of reduced precision.
Crucially, the true proxy quality $\tau_0$ is not identified without further assumptions---analogous to the impossibility of testing instrument validity using outcome data alone. We address this through two complementary approaches. Section (ref) exploits complementarities between the proxy zoo and sign restrictions to partially identify an outer set $[0, \overline{\tau}]$ containing $\tau_0$. Section (ref) conducts sensitivity analysis by reporting $\Theta_\tau$ as $\tau$ varies over this plausible range, following established practice in partial identification andrews2017measuring,vanderweele2017sensitivity,masten2020inference.
We now describe how to compute the identified set in practice. Let $\hat{\phi}$ denote a consistent estimator of the reduced-form parameters $\phi$ defined in (ref), obtained via OLS for the VAR coefficients, $(\hat A_1,...,\hat A_p)$, and sample moments for $\Sigma$ and $\{M_\ell\}_{\ell=1}^k$, i.e.,
where $\hat L=\mathrm{Chol}(\hat\Sigma)$ the Cholesky factor of the residual covariance matrix. Further denote $\hat C_h$ the reduced-form impulse responses in (ref) and $\hat r_{i,j,h}:= r_{i,j,h}(\hat{\phi})$ the sample analogues of the sign-restriction vectors $r_{i,j,h}$.
Given a proxy quality parameter $\tau$, the identified set $\Theta_{i,h}(\tau) = \Big[\underline{\theta}_{i,1,h}(\tau),\, \overline{\theta}_{i,1,h}(\tau)\Big]$ is estimated by solving constrained optimization problems over the orthogonal group. The lower bound is
subject to
The upper bound $\hat{\overline{\theta}}_{i,1,h}(\tau)$ is obtained by replacing is obtained by replacing $\inf$ with $\sup$ in (ref), subject to the same constraints (ref). The optimization can be efficiently solved using off-the-shelf nonlinear programming solvers such as fmincon in MATLAB; see Appendix (ref) for computational details.
The generalized ranking restrictions (ref) depend on a proxy quality parameter $\tau$ that governs a sharp trade-off: larger values tighten identified sets but impose stronger assumptions. Valid identification requires the researcher to specify $\tau \leq \tau_0$, where $\tau_0$ is the true (unknown) proxy quality.
Rather than calibrating $\tau$ ex ante, we exploit two sources of falsifiable information to derive an empirically grounded upper bound $\bar{\tau}$ such that $\tau_0 \in [0,\bar{\tau}]$. First, when proxies contradict IRF-based sign or narrative restrictions, this conflict reveals that proxies cannot be arbitrarily strong instruments relative to their contamination. Second, when multiple proxies provide conflicting directional information, their joint feasibility restricts admissible values of $\tau_0$. Together, these sources of disagreement discipline the quality parameter without requiring parametric assumptions about the endogeneity mechanism.
The next result formalizes this partial identification approach. Let $\alpha_{\ell,v}$ denote the angle between the proxy moment vector $M_\ell$ and a generic conformable vector $v$.
Theorem (ref) quantifies the degree of conflicting information within the complete set of identification assumptions: the generalized ranking restrictions encoded in the proxy moment vectors $\{ M_{\ell} \}_{\ell=1}^{k}$ and the sign restrictions represented by $ \mathcal{G}_{\textup{sign}} $. The bound $\bar{\tau}$ tightens when this conflict intensifies, revealing that high proxy quality is incompatible with the data and maintained restrictions.
To build intuition, recall from Example (ref) that perfect exogeneity assumption ($\tau_0 = \infty$) requires the first column of the rotation matrix to align exactly with the proxy moment: there must exist $ q \in \mathcal{G}_{\textup{sign}} $ such that $ q \propto M_\ell$. This alignment condition clarifies why disagreement---either between proxies and sign restrictions, or among multiple proxies---forces finite bounds on proxy quality, in a manner analogous to overidentification in classical IV settings sargan1958estimation,masten2021salvaging.
\paragraph{Single proxy, consistent with sign restrictions.} Consider first a single proxy ($k=1$) whose moment vector $M_1$ lies within the sign-feasible set, so that $\min_{q \in \mathcal{G}_{\textup{sign}}} \alpha_{1,q} = 0$.\footnote{Without loss of generality, we consider proxy covariance vectors normalized to unit length.} Since some $q \in \mathcal{G}_{\textup{sign}}$ aligns with $M_1$ (e.g., $ q=M_1 $), the alignment poses no contradiction with the sign restrictions, and thus the data permit $ \tau_0 $ to be arbitrarily large. This is reflected by $ \bar{\tau}=\sqrt{n-1}\cot(0)=\infty $.
\paragraph{Single proxy conflicting with sign restrictions.} When the proxy moment vector lies outside the sign-feasible cone---so that $\min_{q \in \mathcal{G}{\textup{sign}}} \alpha_{1,q} > 0$---perfect alignment must violate the sign restrictions. This follows from the fact that every $ q \propto M_1 $ (under the assumption of $\tau_0 = \infty$) does not lie within the sign-restricted set $\mathcal{G}_{\textup{sign}}$. This inconsistency forces $\tau_0 < \infty$: the proxy must contain some contamination for the restrictions to remain jointly feasible. The larger the angular distance, the tighter the bound.
\paragraph{Multiple conflicting proxies.} Consider a proxy zoo with $k\geq 2$. Even when all proxies are individually consistent with sign restrictions ($M_{\ell}\in \mathcal{G}_{\textup{sign}}$ for all $ \ell=1,\ldots,k $), disagreement among proxies disciplines $\tau_0$. If $ M_{\ell}\neq M_{s} $ for some pair of proxies $\ell$ and $s$, then any $ q \propto M_{\ell} $ must satisfy $ q \not\propto M_{s} $, implying contamination in at least the $s$-th proxy. Therefore, when proxies point in different directions, their joint feasibility imposes a finite upper bound on proxy quality.
A visual illustration of these cases is given by (ref) in the appendix.
The complementarity mechanism described above has a striking theoretical implication: with sufficiently complementary proxies, the collective identifying power can shrink the feasible set to a single point. In this limiting case, the zoo achieves point identification without requiring any individual proxy to be a valid instrument.
Theorem (ref) illustrates that point identification is possible purely through an informational conflict between contaminated sources: in the ideal case, two complementary proxies suffice for point identification. While perfect complementarity is unlikely in practice, this result highlights complementarity as an alternative channel for sharp identification to exogeneity assumptions.
Having established the upper bound for the proxy quality $ \tau_0 \in [0,\overline{\tau}] $, this section develops formal tools for sensitivity analysis and diagnostic assessment of the proxy zoo. We first characterize breakdown values that demarcate when specific substantive conclusions become supported by the data vanderweele2017sensitivity,masten2020inference. We then introduce proxy diagnostics to evaluate the information content of individual proxies relative to the zoo. Finally, we provide practical guidelines for implementing sensitivity analysis in empirical work.
Researchers often seek to establish specific substantive conclusions---such as whether a monetary policy shock raises output, or the magnitude of the effect exceeds a certain threshold. In the current framework, such conclusions may or may not hold as the proxy quality assumption $ \tau_0 $ varies. We formalize this sensitivity through the concept of breakdown value.
The breakdown value $\tau^{\ast}(c)$ represents the weakest possible proxy quality assumption required to support the claim $c$. If the researcher believes $\tau_0 \geq \tau^{\ast}$, then the claim $c$ is consistent with the identified set; otherwise, the claim lacks support.
We illustrate popular claims through the following examples. Throughout, we denote by $\mathcal{H}$ some horizon set and without loss of generality consider generic variable indices $ i $ and $ j $. Moreover, we suppress the dependence of the breakdown value $ \tau^{\ast} $ on claim-specific parameters such as the horizon set $ \mathcal{H} $.
\paragraph{Breakdown frontiers.} For multiple claims or claims involving additional parameters---such as the horizon set $ \mathcal{H} $---it is useful to visualize the breakdown frontier as the locus of parameter pairs where conclusions change. The frontier partitions the assumption/conclusion space into regions where different conclusions are supported, providing a comprehensive sensitivity map.
When multiple proxies are available, researchers naturally ask: Which proxies are most informative? Is the proxy zoo delivering genuine complementarity, or are some proxies redundant? How sensitive are the conclusions to specific proxies? We develop a set of diagnostic tools to answer these questions.
We start by defining the informativeness of a proxy zoo. Formally, let $\mathcal{M}$ be a set of proxies and denote the identified set obtained using this proxy set as $\Theta_{i,h}(\mathcal{M},\tau)$. We further denote the identified set using sign and narrative restrictions only by
We then measure the information content of the proxy zoo $\mathcal{M}$ as follows.
Intuitively, $\kappa_{i,h}(\mathcal{M},\tau)$ captures the reduction in the width of the identified set achieved by using the proxy set $\mathcal{M}$, expressed as a fraction of the baseline width implied by sign restrictions alone. At the population level, the width of the identified set is non-negative and weakly decreases when a proxy zoo is introduced; hence, $\kappa(\mathcal{M},\tau) \in [0,1] $ by construction. A value of $\kappa(\mathcal{M},\tau)$ close to 1 indicates that the proxy zoo $\mathcal{M}$ substantially tightens the identified set, while $\kappa(\mathcal{M},\tau) \approx 0$ suggests that the zoo provides little information beyond the sign restrictions. In practice, the measure may slightly exceed 1 because standard numerical solvers (e.g., fmincon in MATLAB) are not guaranteed to attain global optima when computing the bounds of $\Theta_{i,h}(\mathcal{M},\tau)$.
\paragraph{Information Analysis.} Given Definition (ref), we can directly compare the information content across different proxy zoos. For two generic proxy zoos $\mathcal{M}$ and $\widetilde{\mathcal{M}}$, define the relative informativeness of zoo $\mathcal{M}$ compared with $\widetilde{\mathcal{M}}$ as
For any nested proxy zoos, i.e., $\widetilde{\mathcal{M}}\subset \mathcal{M}$, theoretical monotonicity implies $\kappa(\mathcal{M},\tau) \ge \kappa(\widetilde{\mathcal{M}},\tau)$.\footnote{In finite-sample implementations, numerical solvers used to compute the bounds of $\Theta_{i,h}(\mathcal{M},\tau)$ may converge to local optima, leading to mild violations of monotonicity. We interpret such cases as indicating a negligible marginal contribution from the additional proxies in $\mathcal{M}\setminus\widetilde{\mathcal{M}}$.} For non-nested proxy zoos, $\Delta(\mathcal{M},\widetilde{\mathcal{M}},\tau)$ provides a purely comparative measure without a monotonic interpretation.
\paragraph{Leave-one-proxy-out (LOPO) Information Analysis.} A particularly useful special case is the leave-one-proxy-out (LOPO) analysis. Let $\mathcal{M}^{(-\ell)}$ denote the set of all proxies in the zoo $\mathcal{M}$ excluding proxy $\ell$. The informativeness of the $\ell$-th proxy is quantified by
This measure captures the marginal contribution of proxy $\ell$ to tightening the identified set. A large value of $\Delta_\ell(\tau)$ indicates that proxy $\ell$ is highly informative conditional on the remaining proxies, while values close to zero indicate redundancy.
We conclude with a practical workflow for implementing sensitivity analysis in applied work. \paragraph{Step 1: Establish the feasible range.} Compute the upper bound $\bar{\tau}$ from Theorem (ref). This defines the plausible range $\tau \in [0, \bar{\tau}]$ over which to conduct sensitivity analysis.\footnote{If $\bar{\tau} = \infty$, use a large $ \overline{\tau} $ for the following sensitivity analysis.}
\paragraph{Step 2: Compute identified sets on a grid.} For a fine grid $\{\tau_1, \ldots, \tau_G\} \subset [0, \bar{\tau}]$, solve the optimization problems in (ref)--(ref) to obtain $\{\Theta_{i,h}(\tau_g,\hat{\phi})\}_{g=1}^G$. Plot the bounds $[\underline{\theta}_{i,h}(\tau,\hat{\phi}), \overline{\theta}_{i,h}(\tau,\hat{\phi})]$ as functions of $\tau$ for target variables $i$ and horizons $h$.
\paragraph{Step 3: Report breakdown frontiers.} For substantive conclusions of interest (e.g., “output rises on impact”, “price response is negative”), compute and report $\tau^{\ast}(c)$ using Definition (ref). Interpret these as “minimal contamination tolerance” required to support each conclusion.
\paragraph{Step 4: Conduct proxy diagnostics.} Compute proxy zoo informativeness $\kappa$ as in (ref) and LOPO information measures $\{\Delta_\ell(\tau)\}_{\ell=1}^k$ as in (ref). Report the sensitivity of conclusions to proxies in the zoo.
\paragraph{Step 5: Compare to point-identified benchmarks.} When a particular proxy is of interest, e.g., treated as a valid IV in the benchmark model, compute the point-identified impulse responses under $\tau = \infty$. Compare it to the robust bounds under finite $\tau \in [\tau^{\ast}(c),\overline{\tau}]$. This quantifies the cost of relaxing exogeneity and highlights where conclusions are robust versus fragile.
In this section, we illustrate the proposed framework using a simulation study based on a medium-scale dynamic stochastic general equilibrium (DSGE) model. We generate data from the model of smets2007shocks for $T=5000$ periods and estimate a VAR$(12)$ on seven observable variables: output growth, consumption growth, investment growth, wage growth, inflation, the nominal interest rate, and employment.\footnote{We consider a large sample size to isolate identification challenges from finite-sample uncertainty.}
A common practice in the empirical literature is to recover structural shock series using set-identified methods---most notably sign restrictions---and subsequently use the recovered shock series as external proxies in further analyses baumeister2019structural,jarocinski2020deconstructing. To mirror this practice, we first identify the monetary policy shock using standard sign restrictions in the spirit of uhlig2005what. Specifically, we impose correct sign restrictions on the responses to a contractionary monetary policy shock for up to two horizons, leaving the response of output growth unrestricted.
(ref) reports the identified set under the above sign restrictions. Perhaps surprisingly, the identified sets of output responses remain large and relatively uninformative, although the sign restrictions do reduce the set of admissible impulse responses. Moreover, the identified sets contain the true impulse responses, which is consistent with the implication of Proposition (ref) when GRR are absent.
Given the set of admissible structural representations, we construct proxy series for the monetary policy shock following two standard approaches. First, we define a Median-$B$ proxy, obtained from the structural shock series associated with the element-wise median of admissible contemporaneous impact matrices. Second, we define a Closest-$B$ proxy, corresponding to the admissible impact matrix that is closest (in Euclidean distance) to the element-wise median.
While both proxies are consistent with the imposed sign restrictions, they are not guaranteed to be orthogonal to other structural shocks in the data generating process. (ref) illustrates this point by reporting correlations between the constructed proxies and the true structural shocks of the DSGE model. Even with the very large sample size, both proxies exhibit non-negligible correlations with non-monetary shocks, including technology and wage markup shocks.\footnote{With realistic sample sizes, such contamination becomes even more prominent.} Moreover, there is no systematic pattern indicating which non-target shocks are more likely to load on the proxy, nor which construction method yields a uniformly cleaner proxy.
To illustrate the failure of point identification given contaminated proxies, we follow mertens2013dynamic and estimate the structural impulse responses using the Median-$ B $ proxy. As (ref) shows, the resulting impulse response estimates are severely biased. Moreover, the signs happen to be consistent with the truth, making it difficult to detect identification failure from conventional “puzzling” responses alone.
We next apply the GRR, following Section (ref), to characterize identification in the presence of such contamination. Rather than imposing exact orthogonality, we consider relaxation sets indexed by the proxy quality parameter $\tau$, which bounds the proxy's covariance with non-target shocks relative to the target shock. Identified sets are computed over a grid $\tau \in [0,20]$. The true value $\tau_0=3.5438$, computed from the simulation metadata, is used only for validation and is not observable in applications.
(ref) reports the identified sets of impulse responses obtained using the Median-$B$ proxy. The red line shows the true impulse response, while the orange line corresponds to the point-identified proxy-SVAR that treats the proxy as exactly exogenous.\footnote{In population, the point-identified IRF that treats the contaminated proxy as exactly exogenous is given by Equation (ref)}
Two patterns emerge. First, for a wide range of values of $\tau$, the identified sets contain the true impulse response, which is consistent with Proposition (ref). As $\tau$ increases and stronger assumptions are imposed, the identified sets shrink toward the point-identified proxy-SVAR estimate. Second, the sensitivity analysis shows that informative and robust inference is possible without imposing exact exogeneity of the proxy.
This experiment highlights that proxies constructed from set-identified procedures need not satisfy exact orthogonality with non-target shocks, even in large samples. The proposed framework provides a formal and transparent way to conduct identification in such environments.
Finally, we note a sharp contrast in computational burden between the simulation and empirical settings. The simulation relies on repeated sampling over admissible rotations, whereas in the empirical application with eight proxies the identified sets are computed via deterministic constrained optimization, as in (ref). For a fixed value of $\tau$ (for example, at a breakdown value), computing impulse-response bounds takes approximately seven seconds on a standard machine, indicating that the proposed approach is computationally tractable even with a rich proxy zoo.
In this section, we apply the proposed framework to revisit the effects of monetary policy shocks. We identify the set of impulse responses using the eight proxies described in Section (ref), imposing self-sign restriction in (ref) and the generalized ranking restrictions in Equation (ref). Since the self-sign restriction serves only to resolve sign indeterminacy, identification relies entirely on the information contained within the set of proxies (the “proxy zoo”).
Following our empirical guidelines, we first report the identified impulse responses over a grid of proxy quality parameters $ \tau \in [0,\overline{\tau}) $ in (ref). Three key insights emerge from this analysis. First, leveraging the proxy zoo allows us to partially identify the upper bound of the proxy quality parameter, $ \overline{\tau}=2.15 $. This result implies that at least some proxies in the set violate the strict exogeneity assumption, corroborating recent skepticism regarding proxy validity bruns2024testing. Second, under relatively conservative assumptions regarding proxy quality---for example, $ \tau=0 $ (corresponding to external variable restrictions, as in ludvigson2021uncertainty) or $ \tau=1 $ (corresponding to ranking restrictions, as in braun2023identification)---the identified sets remain wide and uninformative, despite the inclusion of multiple proxies. Third, as we strengthen the identification assumption by increasing $ \tau $, the identified sets shrink as expected. Notably, the identified sets become informative at finite values of $ \tau $.
We formalize this intuition through a breakdown value analysis. Suppose we postulate that industrial production and inflation must fall contemporaneously following a monetary tightening. Our analysis reveals that this claim holds only if we impose $ \tau\ge \tau ^{\ast}= 1.88 $. Intuitively, this requires the proxies to be, on average, approximately $ 1.9 $ times more correlated with the monetary policy shock than with any other structural shock. Collectively, these results suggest that for applied researchers, the standard assumption of valid IVs (infinite $ \tau $) may be excessively strong and is unnecessary for establishing certain qualitative conclusions.
Next, we examine the sensitivity of these empirical claims to the composition of the proxy set. To this end, we conduct a leave-one-proxy-out (LOPO) analysis (see Section (ref)), fixing the proxy quality parameter at the breakdown value $ \tau^{\ast} $. Recall that any assumption weaker than $ \tau^{\ast} $ is insufficient to support the claim that both industrial production and inflation contract following the shock.\footnote{Importantly, a wide identified set does not necessarily imply that the claims of interest are rejected; rather, it indicates that there is insufficient information to draw a definitive conclusion.} By excluding one proxy at a time, we assess the stability of this claim and the informativeness of individual proxies.
(ref) illustrates the resulting identified sets. The LOPO analysis demonstrates that the identified sets remain stable as long as the nakamura2018highfrequency (NS) shock is included in the zoo. Specifically, the identified sets for all other leave-one-out combinations (e.g., excluding GK or RR) cluster tightly around the benchmark full-zoo results. In sharp contrast, excluding the NS proxy---and, to a lesser extent, the BRW proxy---causes the identified sets to expand dramatically, particularly for price responses. This visual inspection provides a straightforward qualitative assessment of the sensitivity of the empirical claims to the composition of the proxy zoo.
This qualitative evidence is corroborated by the quantitative information measures reported in (ref). Two patterns emerge. First, the last column shows that the full zoo with eight proxies substantially tightens identification, reducing the widths of the identified sets by 86% relative to the sign-only baseline. This marked improvement is expected, given that the baseline imposes only uninformative self-sign restrictions. Second, excluding individual proxies reveals significant heterogeneity in information content. To quantify this, the second row reports the relative information $ \Delta_{\ell}(\tau^{\ast}) $. As is clear, the NS shock is by far the most informative proxy in the zoo: excluding it substantially reduces the information of the zoo by 21 percentage points. In contrast, removing the miranda-agrippino2021transmission (MR) shock has no material effect on the identified set.\footnote{The information content of the zoo excluding MR ($0.87$) is slightly higher than the information of the full zoo ($0.86$). This discrepancy arises because the optimization problem guarantees only local optima.}
Overall, the LOPO exercise clarifies the critical assumptions driving the empirical findings. If researchers believe that the NS proxy is significantly more contaminated than implied by $ \tau^{\ast} $---for example, due to new evidence of contamination---the remaining proxy zoo would lack the information necessary to identify contractionary effects of monetary policy.
This paper develops a framework for robust identification with external proxies that relaxes exact proxy exogeneity and treats proxy contamination as a matter of degree, rather than a binary property. By combining generalized ranking restrictions (GRR) with auxiliary economic constraints, we characterize identified sets of impulse responses, derive upper bounds on proxy quality using falsifiable information, and propose diagnostic tools that make transparent how empirical conclusions depend on assumptions about proxy contamination and the composition of the proxy zoo. We illustrate the robustness and transparency of the proposed framework using contaminated proxies in a simulated New Keynesian DSGE model, where the conventional proxy-SVAR approach induces substantial bias. In the empirical analysis, we show how the zoo of monetary policy proxies can be reconciled with economic theory under substantially weaker assumptions than those required by conventional IV approaches.
Several directions for future research follow naturally. First, while the present analysis takes the proxy zoo as given, the geometric structure underlying our results highlights the role of proxy complementarity in sharpening identification with external information. Proxies contribute to identification by imposing distinct, non-redundant restrictions on admissible structural representations, so that the resulting identified sets reflect the joint consistency of external information across the proxy zoo, rather than hinging on any single proxy in isolation. Formalizing proxy construction as a search for complementarity among proxies, instead of a pursuit of perfect exogeneity, is an important avenue for future work.
Second, the partial identification of the proxy quality parameter raises the possibility of formal diagnostic procedures. In the empirical application, the fact that the upper bound on proxy quality is finite indicates that the joint restrictions implied by the proxy zoo are incompatible with exact exogeneity for all proxies. Adapting this framework to develop flexible exogeneity tests is a natural extension.
Finally, although our analysis focuses on proxy-SVARs, the core insight that proxy complementarity within a zoo can deliver narrow and economically meaningful identified sets even under endogeneity is more general. Extending the framework to regression settings with mildly endogenous instruments, where the geometric structure is less explicit, may help clarify how information contained in the direction of endogeneity bias across multiple instruments can be exploited for identification.
Taken together, these results emphasize that the identifying content of external information is inherently joint, and that identification depends less on the validity of individual proxies than on the structure and interaction of the proxy zoo as a whole.
\afterpage