Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,389 characters · 22 sections · 29 citation commands
=0pt =0pt plus .5=0pt plus .5=.3Relative Bias Under Imperfect Identification in Observational Causal Inference
\allsectionsfont
Keywords\quad sensitivity analysis \textbullet proximal inference \textbullet instrumental variables
Causal inference in observational studies requires identification assumptions which account for potential confounding that occurs from non-random assignment. The most familiar set identification assumptions is that of no unobserved confounding, also known as ignorable treatment assignment, or selection on observables (SOO): given observed covariates, potential outcomes are independent of treatment. In many observational settings, however, SOO is considered implausible. The presence of unobserved confounders can bias treatment effects which are estimated under an SOO assumption.
Alternative identification strategies have been proposed as an alternative to the selection-on-observables assumptions. Researchers may try to find an instrumental variable (IV) which is associated with treatment but satisfies an exclusion restriction with the outcome. Under certain additional assumptions, an IV allows for estimating the effect of interest. More recently, proximal inference has been developed as an alternative identification approach miao2015identification, miao2018identifying, cui2024semiparametric. Instead of assuming researchers measure all relevant confounders, proximal inference assumes researchers have access to two informative proxies: a treatment proxy and an outcome proxy. If these proxies satisfy certain conditional independence assumptions, not unlike the exclusion restriction in IV, then a treatment effect can be estimated, even though the proxies themselves are confounded.
The existence of three distinct identification approaches leads to natural questions about how the corresponding estimators compare in practice. Much work has been conducted to compare the relative efficiency of the different approaches andrews2019weak. In particular, weak instruments, and by extension, weak proxies, can inflate the variance of IV and proximal estimators, such that SOO estimators will have lower mean squared error even in settings when the IV and proximal assumptions hold, and the SOO assumptions do not.
Fewer studies have examined how the bias of different identification strategies compares. The nature of identification assumptions is that they are untestable; in real-world observational data, it is unlikely that any set of identifying conditions, SOO, IV, or proximal, will hold exactly. It is important, then, to understand the bias that may result from plausible violations of each set of identifying assumptions, and how this bias compares between identification strategies. Such a comparison is useful not only in picking an identification strategy, but also in understanding the assumptions made as part of a primary causal analysis under a single strategy.
Recent literature has introduced different sensitivity frameworks to evaluate the sensitivity of estimators to potential violations in the underlying identifying assumptions for SOO rosenbaum1987sensitivity, frank2000impact, imbens2003sensitivity, tan2006distributional, carnegie2016assessing, ding2019decomposing, zhao2019sensitivity, cinelli2020making, hong2021did, dorn2023sharp, huang2025variance, for IV kang2021ivmodel, freidling2022optimization, cinelli2025iv, and for proximal cobzaru2024bias, to name a few examples. However, the different frameworks and approaches make it challenging to compare bias across identification strategies.
This paper introduces a framework for researchers to reason about the robustness of different identification approaches under the realistic setting that none of the identification assumptions hold exactly. Our framework enables researchers to carefully consider the trade-offs made between using a traditional SOO, IV, and proximal assumptions, and understand the conditions under which each identification strategy will provide the best estimates of the underlying effect.
First, in Section (ref), we derive expressions for the bias of both IV and proximal estimators in terms of the violations of their underlying assumptions. We show that the bias that arises from violations in one set of identification assumptions can impact the bias in alternative identification approaches. Specifically, the bias that arises in IV will additionally depend on violations in the SOO assumptions, while the bias for a proximal estimator will depend on both violations in the IV assumptions and the SOO assumptions.
Second, in Sections (ref) and (ref), we propose a set of sensitivity tools in the form of numerical and visual summary measures. To compare the sensitivities of the various identification approaches, we introduce a generalization of standard robustness values called the total robustness value, which measures the minimum amount of total violation of a set of identification assumptions that result in a certain amount of bias. Then, to connect the identifying assumptions themselves, we show that the bias of both IV and proximal inference can be written as a function of just the selection-on-observables parameters and the data. This unification allows us to relate the three identification strategies on a single bias contour plot.
Our sensitivity framework can aid practitioners in several ways. Often, applied researchers identify a causal effect of interest, consider plausible identification strategies, and then collect data to enable inference under their chosen strategy. In contrast to existing sensitivity analyses, which work with a single identification strategy, our proposed measures are constructed to directly account for the nested dependency in violations across the different identification assumptions. This allows researchers to directly compare different identification strategies and choose one most appropriate for their data.
The proposed framework is still useful, however, for researchers who have already settled on an identification strategy for their primary analysis. As both the bias expressions in Section (ref) and the unification in Section (ref) show, the SOO, IV, and proximal assumptions are deeply interrelated. Even when the SOO assumptions are not plausible, the value of the SOO estimate places certain restrictions on plausible values for the different sensitivity parameters. For example, the observed data may imply that either the IV assumptions hold or a certain minimum amount of treatment confounding must exist. If an analyst is confident about the direction of confounding bias, that may also immediately rule out the plausibility of certain other identifying assumptions. Thus our framework provides for better interrogation and integration of identifying assumptions, allowing researchers to consider their assumptions from a different angle and move beyond fully a priori reasoning about their plausibility.
To illustrate our discussion and demonstrate the proposed sensitivity analyses in a concrete setting, we re-analyze data from hk2022poland, who studied the effects of state surveillance on resistance to the Polish regime, as measured by group and individual protests. The authors digitized records from the Służba Bezpieczeństwa (SB; Department of Security) from 1949--1989, covering the entire existence of the Polish People's Republic. These records include the number of secret agents assigned to each municipality each year. The authors focus on the region of Upper Silesia, where data was particularly comprehensive. They also collected data on protests organized by the Solidarność (Solidarity) movement, which was the center of resistance to the governing regime in the 1980s, as well as indirect data on individual refusals to work on Saturdays. The research question is whether additional surveillance of a municipality caused either group or individual protest to increase or decrease.
Hager and Krakowski's primary analysis is based on an instrumental variable: the number of local Catholic priests who were secretly cooperating with the SB. The police intentionally recruited priests, who in turn recruited other local informants. However, priests were randomly assigned to parishes by the Catholic Church, making the number of corrupt priests plausibly exogenous. The authors include a detailed study and discussion of the validity of the instrument, which we also discuss below. As a robustness check, they also conducted a cross-sectional analysis under a SOO assumption. Both analyses found that state surveillance increased group protests but decreased individual protests. Here, we study the sensitivity of both the IV and SOO findings, and also conduct a sensitivity analysis for proximal inference using the same data.
This section provides an overview of three common identification strategies for the ATE in cross-sectional observational settings: selection-on-observables, instrumental variables, and proximal inference. We express these strategies in terms of a common statistical model, which we introduce next, and we discuss the plausibility of the different identifying assumptions in the context of our running example.
We study the following partially linear structural model for an outcome \(Y\), treatment \(Z\), vector of covariates \(X\), unobserved confounder \(U\), and two additional variables \(W_Y\) and \(W_Z\).\footnote{Because we are studying the behavior of three different identification strategies, we require a model for all four variables: \(Y\), \(Z\), \(W_Y\), and \(W_Z\). In settings where researchers are only interested in selection-on-observables, they only need to assume a model for \(Y\). Similarly, for IV, a model for \(W_Z\) and \(Y\) are sufficient.}
Above, \(\varepsilon_y, \varepsilon_z, \varepsilon_{w_z}\), and \(\varepsilon_{w_y}\) are independent noise terms, and \(f_y\), \(f_z\), \(f_{w_y}\), and \(f_{w_z}\) are generic functions of the covariates \(X\). The estimand of interest is the treatment effect, represented by the coefficient in front of \(Z\) in the equation for \(Y\), which we denote \(\tau\).\footnote{The stable unit treatment value assumption (SUTVA) is also implied by this setup. In settings where the treatment \(Z\) is binary, we can map the structural models to the potential outcomes framework, where the potential outcomes for \(Y\) under control and treatment, which we write as \(Y(0)\) and \(Y(1)\), can be obtained by substituting \(Z=0\) and \(Z=1\) into the equation for \(Y\). As a result, the average treatment effect (ATE) is \(\operatorname{\mathbb{E}}[Y(1) - Y(0)] = \tau\).} This can be thought of as a special case of the fully non-parametric setting studied in miao2018identifying. For expositional clarity, the rest of the paper works with the reduced model that (nonparametrically) partials out \(X\); this also has the effect of centering the remaining variables at zero. In a slight abuse of notation, we will use the same notation for the variables and coefficients in the reduced model. Figure (ref) illustrates the setup and the assumptions discussed below. We take a population view of the linear models. Thus, the `hat' above different terms (i.e., \(\hat Z\) and \(\hat W_Y\)) is used to denote that these are predicted values (at the population-level), not that they correspond to sample estimates. \newline
Remark. Throughout the paper, we assume that \(W_Z\) and \(W_Y\) are pre-treatment variables. In other words, we do not account for settings when researchers are using post-treatment negative controls for potential candidates of the treatment proxy \(W_Y\), or settings where \(W_Z\) (i.e., the instrument or outcome proxy) is potentially post-treatment. The framework can be adapted to accommodate these settings, though this would require either accounting for potential post-treatment bias from erroneously adjusting for these variables in a selection-on-observables setting, or accounting for different sets of covariates across the different identification strategies.
The ATE \(\tau\) can be estimated from Eq. (ref) only if \(U\) is observed, which it is not. The simplest way to estimate \(\tau\) in the presence of an unobserved \(U\) is to assume that the observed covariates \(X\) contain all relevant confounders, an assumption known as selection-on-observables (SOO). In our model, SOO means that at least one of the following two assumptions holds.
Assumption (ref) corresponds to \(\gamma_u=0\) in Eq. (ref), and Assumption (ref) corresponds to \(\beta_u=0\). Either of these assumptions is sufficient for \(\tau\) to be identified, and estimated via a regression on the observed variables, with \(\tau_\mathrm{soo}\) the coefficient on \(Z\) in this regression. The following standard result, whose proof is omitted, records this fact.
In practice, the selection-on-observables assumptions cannot be tested and are often untenable. Researchers are often constrained by the pre-treatment covariates that are available in observational contexts. In the presence of latent, mismeasured, or unknown confounders, relying on a SOO identification strategy will generally result in biased estimates.
Instrumental variables have long been used in contexts where researchers believe that confounding remains after controlling for observed covariates. Informally, this identification approach relies on on finding a variable \(W_Z\), the instrument, that can explain variation in the treatment \(Z\), but can only affect the outcome \(Y\) through its relationship with \(Z\). The variation in \(W_Z\) can then be used to identify the treatment effect \(\tau\), even in the presence of an unobserved confounder \(U\). While IV no longer requires Assumptions (ref) or (ref), identification under our model requires that \(W_Z\) satisfy the following two assumptions.
An instrument \(W_Z\) is considered a valid instrument if both exclusion restriction, as well as exogeneity, holds. In the context of Eq. (ref), Assumption (ref) corresponds to \(\varphi_u=0\) and Assumption (ref) corresponds to \(\beta_{w_z}=0\). The instrument must also be relevant, meaning that \(\gamma_{w_z}\neq 0\), so it has some effect on \(Z\). Given a valid instrument, the most common way to estimate the treatment effect is via two-stage least squares, which will recover \(\tau\). In general settings beyond the model in Eq. (ref), additional assumptions about treatment effect homogeneity or monotonicity are required to identify a causal effect.
As Section (ref) describes, in our running example from hk2022poland, the authors used the number of local Catholic priests who were secretly cooperating with the SB as an instrumental variable. They show the number of corrupted priests is strongly correlated with the number of SB officers, and justify both IV identification assumptions using both substantive knowledge and auxiliary data. However, like all identification approaches, there still remain existing mechanisms that could violate either assumption. For example, if the Catholic Church reserved assignments to important cities, such as those with a history of protest or surveillance, for priests that they judged to be more resistant to recruitment, then the exogeneity assumption would no longer hold. Similarly, the authors note that the possibility of indoctrination of parishioners by corrupted priests would result in a violation of exclusion restriction. We provide more discussion in Section (ref). All in all, like many other real-world settings, the complex social dynamics in hk2022poland make it unlikely that both IV assumptions hold exactly, and warrant investigation of the sensitivity of the estimates to violations of either assumption.
While IV allows researchers to relax the selection-on-observables assumptions in exchange for two instrumental variables assumptions, finding a valid instrument can be challenging. In particular, an instrument \(W_Z\) must explain variation in \(Z\), but be exogenous. Recently, miao2015identification proposed an alternative identification strategy, proximal inference, which relaxes the IV assumption of exogeneity.
Proximal inference relies on finding two `proxy variables': (1) an outcome proxy \(W_Y\) that is unrelated to the treatment but is related to the outcome, and (2) a treatment proxy \(W_Z\) that is unrelated to the outcome, but is related to the treatment. The treatment proxy is similar to an instrumental variable, but does not need to be exogenous. As such, proximal inference can be viewed as an alternative to instrumental variables, in settings when researchers are not confident they have an exogenous instrument.
More formally, the proximal inference approach of miao2015identification replaces SOO or IV assumptions with two assumptions about the proxy variables.
When both Assumptions (ref) and (ref) hold, we say that \(W_Y\) and \(W_Z\) are valid proxy variables. In the context of Eq. (ref), Assumption (ref) corresponds to \(\gamma_{w_y}=0\) and \(\alpha_{w_z}=0\), and Assumption (ref) means \(\beta_{w_z}=0\). Notice that Assumption (ref) is exactly the same as Assumption (ref).
Like instrumental variables, a two-stage least squares procedure can be used to estimate the treatment effect. However, instead of regressing the treatment indicator \(Z\) with the instrument \(W_Z\), the first stage relies on regressing the outcome proxy \(W_Y\) with \(Z\) and \(W_Z\). Because \(W_Y\) has no direct causal connection with \(Z\) and \(W_Z\) under Assumption (ref), any correlation between the two sets of variables can only be due to the unobserved confounder \(U\), and thus the first stage amounts to estimating a proxy for \(U\). That estimated \(\hat W_Y\) is then included in the second-stage, along with \(Z\). When both \(W_Y\) and \(W_Z\) are valid proxies, the estimated treatment effect under proximal inference (denoted \(\tau_{\mathrm{prox}}\)) will recover the average treatment effect \(\tau\).
hk2022poland did not originally run a causal analysis under proximal assumptions. Yet, as described above, both SOO and IV identification strategies face plausible challenges to their validity, which proximal identification might address. The treatment proxy \(W_Z\) is the count of corrupted priests, corresponding to the instrument under IV. For the outcome proxy \(W_Y\), we argue that the indicator for former Russian occupation is a reasonable candidate. Hager and Krakowski explain that Russia implemented “russification” policies in these areas “which were intended to destroy Polish communities and identities,” in turn making protest against the Polish regime less likely. A mechanism for affecting the amount of surveillance is less obvious; the authors suggest in passing that Russia's experience in the occupied territories may have made administering spy networks easier, but we find this mechanism less plausible. Indeed, the occupation variable has a partial correlation with the group protest outcome about three times larger than the correlation with the treatment, after controlling for other covariates, and a ten-times larger correlation with the individual protest outcome than treatment.
Proximal identification is no panacea, however. Valid inference still hinges on the exclusion restriction holding, and the outcome proxy must satisfy two additional exclusion restrictions. Specifically, Russian occupation cannot directly affect the number of corrupted priests or the amount of surveillance. While plausible for some of the reasons Hager and Krakowski outline in their discussion of the exogeneity of the instrument, it is not unreasonable to assume that the same weakening of Polish communities in occupied areas might affect which priests are assigned and their success in recruiting informants. Thus, as with the other two identification strategies, it remains important to examine the sensitivity of estimates to violations of the identifying assumptions.
Table (ref) summarizes the three identification approaches, their assumptions, and the regressions used in estimation. It also displays the estimated effect of surveillance on individual protests in our running example, under each estimation approach. The scale of the effects is difficult to interpret due to several layers of rescaling and pre-processing in the underlying data. Here, we have rescaled both the outcome and treatment variables to have unit variance, to ease comparison and discussion of the estimates. Thus a value of \(0.3\) below means a one-s.d. increase in surveillance is associated with an \(0.3\)-s.d. increase in protest.\footnote{The same estimate corresponds to a value of \(0.068\) in the original paper.}
Like hk2022poland, we find opposite signs for the effect on group and individual protest. Notice that the estimates are all statistically significant at traditional levels. One immediately apparent aspect of the estimates in Table (ref) is that the three approaches disagree on the effect of individual protest.\footnote{The authors additionally analyzed an outcome of `group protests', which is a count of protests organized by the Solidarność (Solidarity) movement. We provide the additional analysis in the Appendix.}
\setstretch{1.0}
\setstretch{1.0}
This section presents decompositions of the bias that arises from violations in any of the underlying identification assumptions. We find that identification strategies are more sensitive to violations of their assumptions when there is more unobserved confounding. In other words, violations in Assumption (ref) will amplify any violations in the alternative identification assumptions.
The results here are derived for the population estimates. Because our model (Eq. (ref)) is linear, however, they hold equally when all coefficients are interpreted as their finite-sample versions. Throughout the paper, we follow the notation in cinelli2020making and use superscripted \(\perp\) to indicate partialling out, such that for any variables \(A\), \(B\), and \(C\), \(A^{\perp B, C} := A - \operatorname{\mathbb{E}}[A\mid B, C].\) We define \(\operatorname{sd}(A) := \sqrt{\operatorname{\mathbb{V}}(A)}\) and \(\operatorname{corr}(A, B) := \mathrm{Cov}(A, B)/(\operatorname{sd}(A)\operatorname{sd}(B))\). We make use of partial correlations, \(R_{A\sim B\mid C} := \operatorname{corr}(A^{\perp C}, B^{\perp C})\), and the resulting partial \(R^2\) values (i.e., \(R^2_{A\sim B\mid C} = \operatorname{corr}^2(A^{\perp C}, B^{\perp C})\)). Finally, we will use \textcolor{sooblue}{blue} to highlight parameters related to selection-on-observables assumptions, \textcolor{ivpink}{pink} for parameters related to instrumental variables assumptions, and \textcolor{proxred}{brown} for parameters related to proximal inference assumptions.
For selection-on-observables to be a valid identification approach, there cannot be an omitted confounder \(U\) that affects both the outcome \(Y\) and the treatment assignment process \(Z\). In the presence of an omitted \(U\), the bias for \(\tau_\mathrm{soo}\) will be a function of \(U\)'s relationship with \(Y\) and \(Z\), respectively. This can be expressed in terms of the coefficients in Eq. (ref) as
Lemma (ref) contains the full derivation.
Consider a potential confounder of state capacity in the running example. Then, \(\textcolor{sooblue}{\beta_u}\) represents the relationship between state capacity and the outcome (either group or individual protest) that is not already captured by the observed proxies of state capacity (i.e., number of schools). Similarly, \(\textcolor{sooblue}{\gamma_u}\) represents the residual imbalance of state capacity across the treatment and control groups, after controlling for the observed covariates. If researchers believe that proxies of state capacity sufficiently capture the variation in state capacity, then we expect \(\textcolor{sooblue}{\beta_u}\) and \(\textcolor{sooblue}{\gamma_u}\) to be relatively small in magnitude, and little bias to result for not controlling for the latent variable of state capacity. In contrast, researchers could be worried that the measurement error in using the number of schools as a proxy of state capacity could be related to protests, in which case, we might expect there to be more bias from omitting a variable like state capacity.
Importantly, Eq. (ref) formalizes the well-studied result that if omitted confounder \(U\) is related to both the outcome process and the treatment assignment process, there will be bias imbens2003sensitivity. Even if \(U\) is imbalanced between the treatment and control groups, if it does not explain variation in the outcome, then omitting it will not lead to bias, since \(\beta_u = 0\). Similarly, if \(U\) explains variation in the outcome but does not explain variation in the treatment assignment process, then \(\gamma_u = 0\), while it may be advantageous to control for \(U\) for precision, doing so is not strictly necessary to eliminate bias.
Theorem (ref) presents a decomposition of the bias of the instrumental variables estimate. The bias decomposition highlights that there are dependencies between the bias of the \(\tau_\mathrm{iv}\) and the selection-on-observables assumptions. In other words, even though \(\tau_\mathrm{iv}\) relies on different identification assumptions as \(\tau_\mathrm{soo}\), the bias in \(\tau_\mathrm{iv}\) can be amplified by violations in the selection-on-observables assumptions. Appendix (ref) further discusses the special case when researchers are willing to assume exogeneity of the instrument holds exactly, and wish to examine sensitivity to the exclusion restriction.
The bias in Eq. (ref) depends on three key terms: (a) a scaling factor depends on the strength of the instrument, (b) bias from violations in the exclusion restriction, as captured by \(\textcolor{ivpink}{\beta_{w_z}}\), and (c) the bias from violations in instrument exogeneity, captured by the term \(\textcolor{sooblue}{\beta_{u}} \cdot \textcolor{ivpink}{\varphi_u}\).
The first term is a scaling factor, which is comprised of fully observable quantities. While the scaling factor does not depend on violations in the underlying IV assumptions, it will amplify any potential violations. Notably, the scaling factor depends on the strength of the underlying instrumental variable as measured by \(R_{Z\sim W_Z\mid W_Y}\). In settings when there is a weak instrument, \(R_{Z\sim W_Z\mid W_Y}\) will be low, and even in the presence of very weak violations in the underlying assumptions, the overall bias incurred by an IV estimator will be relatively large. This is on top of any inferential challenges, such as high variance and non-normality, that arise from weak instruments even when the IV assumptions hold exactly andrews2019weak.
The second term is \(\textcolor{ivpink}{\beta_{w_z}}\), which measures how related the instrument \(W_Z\) is to the outcome \(Y\), after conditioning on \(\{W_Y, Z, U\}\). This is a direct representation of violations in the exclusion restriction.
Finally, the last term represents the bias incurred from violations of exogeneity and is represented by the the product between \(\textcolor{sooblue}{\beta_{u}}\) and \(\textcolor{ivpink}{\varphi_u}\). \(\textcolor{ivpink}{\varphi_u}\) represents the residual variation in \(W_Z\) that can be explained by the omitted confounder \(U\), after controlling for observed covariates, and directly maps to the violation of Assumption (ref). Interestingly, \(\textcolor{ivpink}{\varphi_u}\) is scaled by \(\textcolor{sooblue}{\beta_u}\), which corresponds to violations in the underlying SOO assumptions. We can informally think of the SOO parameters as controlling the overall `difficulty' of the identification problem. When \(U\) is a relatively weak confounder such that \(\textcolor{sooblue}{\beta_u}\) is close to zero, then there will be less exposure to bias for both \(\tau_\mathrm{soo}\) and \(\tau_\mathrm{iv}\). In this setting, \(\tau_\mathrm{iv}\) will be more robust to violations of exogeneity. Conversely, more outcome confounding directly increases the effect of violations of exogeneity.
Theorem (ref) presents a decomposition of the bias of the proximal estimate. Like the results in Section (ref), the bias of \(\tau_\mathrm{prox}\) will implicitly depend on not only the degree to which the proximal assumptions are violated, but also the selection-on-observables and IV assumptions.
The theorem formalizes that the bias of a proximal estimator is going to be made up of two parts: the first corresponds to the bias from an invalid treatment proxy, and the second corresponds to bias from an invalid outcome proxy. Both sources of bias are scaled by values that corespond to the strength of the respective proxies.
To start, the first driver of bias results from an invalid treatment proxy, as measured by \(\textcolor{ivpink}{\beta_{w_z}}\). This parameter is identical to the parameter from the IV estimator, which corresponded to the degree to which the exclusion restriction has been violated. This clarifies that while proximal estimators allow researchers to relax the exogeneity assumption associated with an instrumental variable, violations of the exclusion restriction impact both approaches similarly. This term is scaled by a term that depends on how much variation in the constructed \(\hat W_Y\) can be attributed to both \(W_Z\) and \(Z\).
The second driver of bias arises from an invalid outcome proxy. A valid outcome proxy requires that \(W_Y\) is not related to both \(Z\) and \(W_Z\), after conditioning on \(U\) (and other pre-treatment covariates). This is represented by \(\gamma_{w_y}\) and \(\varphi_{w_y}\), which measure how related \(W_Y\) is to \(Z\) and \(W_Z\), respectively, controlling for both \(U\) and other pre-treatment covariates. Violations from an invalid outcome proxy are amplified by a scaling factor. The scaling factor depends primarily on the strength of \(W_Y\) as an outcome proxy as measured by \(\alpha_u\) and the SOO parameter \(\textcolor{sooblue}{\beta_u}\). Analogously to IV, when \(U\) strongly affects the outcome, there is a greater sensitivity to bias from an invalid outcome proxy.
We can equivalently re-write the bias expressions derived in Theorem (ref) and Theorem (ref) in terms of partial correlations instead of regression coefficients and residual variances. Doing so results in a representation of the bias that is invariant to the scale of the underlying variables. In Appenidx A.4, we present the bias expressions in terms of partial correlations. Table (ref) provides a summary of the different coefficients and the corresponding partial \(R^2\) values.
While re-writing the bias in terms of partial correlation terms allows for a representation of the bias that is represented by scale-invariant parameters, the bias expressions are challenging to reason about in practice. The bias results in the previous section, Theorem (ref) and Theorem (ref), highlight how violations of the different identifying assumptions interact and lead to bias in estimates of \(\tau\). This is further exacerbated by the \(R^2\) representations. In particular, these coefficients and variances are not variation-independent---the range of possible values for a coefficient depends on the values of other coefficients. Additionally, it is difficult to make comparisons across identification strategies, because each bias expression is parametrized in terms of a different set of coefficients and variances.
In the following section, we address these limitations by reparameterizing the problem in terms of four unitless partial correlations, three of which directly measure the violation of certain identifying assumptions.
In this section, we introduce a variation independent parameterization of the bias for all three of the identification strategies. We show that the variation independent parameterization can be directly mapped back to the partial \(R^2\) values that measure the degree of violation of the different identifying assumptions. Doing so allows us to define a numerical summary measure called the robustness value that quantifies the minimum violation of the identifying assumptions needed to result in a substantively meaningful change in the result. The proposed measure generalizes the robustness value introduced in cinelli2020making to the IV and proximal case, and allows researchers to compare the sensitivities across the different identification approaches. To aid in the interpretation of the RV, we propose a benchmarking approach, which allows researchers to reason about the plausibility of potential violations in their underlying identification assumptions.
Any approach to comparing identification strategies must require parametrizing the estimates, the bias, and the identifying assumptions on a common space. To start, because each estimator is simply the coefficient from a linear regression, we can write each estimator in terms of the population covariance matrix \(\Sigma \in \ensuremath{\mathbb{R}}^{7 \times 7}\) of \((Y, Z, W_Z, W_Y, \hat Z, \hat W_Y, U)\): \[ \tau_\mathrm{soo} = \Sigma_{zw_zw_y,zw_zw_y}^{-1}\Sigma_{zw_zw_y,y}, \quad \tau_\mathrm{iv} = \Sigma_{\hat zw_y,\hat zw_y}^{-1}\Sigma_{\hat zw_y,y}, \qand \tau_\mathrm{prox} = \Sigma_{z\hat w_y,z\hat w_y}^{-1}\Sigma_{z\hat w_y,y}, \] where subscripts here indicate the submatrix of \(\Sigma\) that corresponds to the variables in the subscript; i.e., \(\Sigma_{zw_zw_y,y}\) is the 3-by-1 submatrix containing the covariance of \((Z, W_Z, W_Y)\) with \(Y\). Similarly, the true effect \(\tau\) can be expressed as \(\tau = \Sigma_{zw_zw_yu,zw_zw_yu}^{-1}\Sigma_{zw_zw_yu,y}\).
While the different estimators do not depend on \(U\), \(\tau\) depends on \(U\). Because the observed data fixes most of the matrix \(\Sigma\), only the submatrix \(\Sigma_{u,yzw_zw_y}\) is unidentified.\footnote{ Because \(\hat Z\) and \(\hat W_y\) are deterministic functions of observed variables only, the submatrix \(\Sigma_{u,\hat z\hat w_y}\) will depend deterministically on the other observed entries as well as \(\Sigma_{u,yzw_zw_y}\), leaving only the submatrix \(\Sigma_{u,yzw_zw_y}\) as being truly unidentified.} Thus, another view of sensitivity analysis is to vary the entries in the submatrix \(\Sigma_{u,yzw_zw_y}\) to evaluate how \(\tau\) changes.
The four entries in \(\Sigma_{u,yzw_zw_y}\) are not unconstrained---they must be such that \(\Sigma\) is positive semidefinite. Without loss of generality, we can set \(\Sigma_{uu}=1\), since the scale of the confounder does not affect any observables. We then parametrize \(\Sigma_{u,yzw_zw_y}\) by a set of partial correlations of \(U\) and each observable variable: \[ \rho \coloneq (R_{U\sim W_Y}, \textcolor{ivpink}{R_{U\sim W_Z\mid W_Y}}, \textcolor{sooblue}{R_{U\sim Z\mid W_Y,W_Z}}, \textcolor{sooblue}{R_{U\sim Y\mid Z,W_Z,W_Y}}). \] Notice that the latter two partial correlations are exactly the SOO sensitivity parameters, and the second, \(\textcolor{ivpink}{R_{U\sim W_Z\mid W_Y}}\), measures the violation of Assumption (ref). The first parameter, \(R_{U\sim W_Y}\), measures the strength of the outcome proxy, but does not directly measure the violation of any identifying assumption.
Critically, this choice of partial correlations admits a bijective mapping to the Cholesky decomposition of \(\Sigma\) and thus to \(\Sigma\) itself lewandowski2009generating. As a result, unlike the entries in \(\Sigma_{u,yzw_zw_y}\), these parameters are variation independent: any values of \(\rho\in[-1, 1]^4\) corresponds to a valid covariance matrix \(\Sigma(\rho)\).
The mapping \(\rho\mapsto\Sigma\) fixes the true effect \(\tau\). It also determines the value of the various partial correlations that measure how much each identifying assumption is violated. In other words, we can re-express each of the partial correlation values that correspond to the different violations in identifying assumptions as functions of \(\rho\). For example, the violation of the exclusion restriction corresponds to the value of \(\textcolor{ivpink}{R_{Y\sim W_Z\mid Z,W_Y,U}}\), and can be written as \[ \textcolor{ivpink}{R_{Y\sim W_Z\mid Z,W_Y,U}}(\rho) = \frac{\Sigma(\rho)^{-1}_{y,w_z}}{ \sqrt{\Sigma(\rho)^{-1}_{y,y}\Sigma(\rho)^{-1}_{w_z,w_z}}}. \] Thus, even though the exclusion restriction violation \(\textcolor{ivpink}{R_{Y\sim W_Z\mid Z,W_Y,U}}\) does not appear anywhere in \(\rho\), unlike the SOO sensitivity parameters, it is still related nonlinearly to the SOO parameters and the other entries in \(\rho\), due to the common parametrization.
The reparametrization of the bias in terms of \(\rho\) enables the definition of the robustness value, one measure of the sensitivity of an estimate to violations of identifying assumptions. The robustness value was first introduced by cinelli2020making as the minimum violation of the identifying assumptions “to change the research conclusions.”
The robustness value depends on an analyst-chosen bias threshold \(b\), representing the smallest change in a causal estimate so that is considered meaningful.\footnote{ So that, for example, a true value of \(\tau=\tau_\mathrm{iv}-b\) is considered substantively different to the causal estimate \(\tau_\mathrm{iv}\) actually obtained. A common choice is \(b\) equal to the estimate itself, or the estimate plus or minus two standard errors, and thus large enough to reverse the sign of the causal effect. Alternatively, \(b\) could be chosen to be equal to some multiple of the estimate's standard error. For a Normal sampling distribution, a bias of one standard error reduces the coverage of nominally 95% confidence intervals to around 80%; a two-s.e. bias reduces coverage to around 50%. This is similar to the notion of a `killer confounder', first presented in huang2024sensitivity and hartman2024sensitivity.} To begin, for a fixed \(\rho\), we define the vector of parameters associated with each identification strategy as follows: \[
\] Because the SOO sensitivity parameters are part of \(\rho\) itself, \(A_\mathrm{soo}\) is simply comprised of the last two entries in \(\rho\); however, \(A_\mathrm{iv}\) and \(A_\mathrm{prox}\) are nonlinear functions of \(\rho\). While the bias of both the IV and proximal estimators additionally depend on other partial correlations, \(A_\mathrm{iv}(\rho)\) and \(A_\mathrm{prox}(\rho)\) represent the partial correlations that directly map to the underlying identification assumptions.
We can then formally define the robustness value (RV) as the minimum bound on the violations of the identifying assumption that is sufficient to guarantee absolute bias no more than \(b\):
Here, \(P\subseteq [-1,1]^4\) is a possibly restricted set of allowable values of \(\rho\), which we discuss below.
Eq. (ref) captures the notion that violations smaller than the RV guarantee bias smaller than the critical threshold. The robustness value measures a kind of safety guarantee: if all of the assumptions are violated less than the RV, as measured by the partial \(R^2\) values, then the bias will be smaller than \(b\). A smaller robustness value means that a smaller violation of the identifying assumptions could be sufficient to result in a threshold bias of \(b\). In contrast, a larger robustness value means that a larger violation of the identifying assumptions is needed to obtain the same bias.
Because of the common parametrization, it is possible to compare the robustness values of different identification approaches. In an extreme case, if the RV for one identification strategy is very small, while it relatively large for the others, then this means that a smaller violation in the assumptions of that strategy could result in a substantial amount of bias; this can warrant caution in evaluating the credibility of the estimate from that strategy. However, each RV measures violations of different assumptions, and so when RV values are closer, direct comparisons may be more difficult. We develop a benchmarking procedure for the RV based on observed covariates in Section (ref) to aid in these comparisons.
The proposed robustness value generalizes the robustness value defined in cinelli2020making, which was introduced for the selection-on-observables setting.\footnote{cinelli2020making define the RV through a specific formula, but motivate it in the way defined here, and for the SOO setting that they studied, their explicit definition and the implict one in Eq. (ref) coincide.} However, some properties of the RV for SOO do not carry over to other strategies. In the SOO setting, the solution to the minimization problem in Eq. (ref) requires that the two SOO sensitivity parameters (the last two components of \(\rho\)) be equal. That may not be the case with the RV when applied to IV or proximal strategies. Another difference is that the SOO bias is monotonically increasing in each parameter. For IV and proximal, however, the addition and subtraction in their bias expressions means it is possible for assumption violations to cancel out. Thus, if an assumption is violated more than \(\mathrm{RV}(b)\), there is no guarantee that the bias will exceed \(b\) for IV and proximal, depending on the signs of the partial correlations.
One difficulty with the RV for IV and proximal is the behavior of the bias when treatment confounding \(\rho_3 (i.e., R^2_{Z\sim U\mid W_Z,W_Y}\)) is very high (near 1). In particular, when \(\rho_3\) is near 1, the bias of both IV and proximal can be extremely sensitive to small changes in the other parameters, even if those parameters are relatively small in magnitude. In other words, any miniscule change in other parameters in \(\rho\) result in extreme changes in the value of \(\tau\). This can be seen visually in a traditional sensitivity contour plot where the contours of bias are extremely dense around the point \((1, 0)\). The high derivative of the bias in this region may lead to smaller robustness values. We show this pattern in Figure (ref) for our running example by carrying out the optimization in Eq. (ref) while fixing \(\rho_3\) to different values along its possible range. The RV for SOO is minimized at an intermediate value of \(\rho_3\)---in fact, exactly where \(\rho_3=\rho_4\). However, the RV for IV and proximal tends to 0 as \(\rho_3\to 1\). This also highlights that in settings when there is a large degree of confounding in the treatment assignment process, the bias of IV and proximal estimates will be extremely sensitive to any violations in their underlying identification assumptions.
This motivates restricted choices of \(P\), the allowable values of \(\rho\) in Eq. (ref). Restricting \(\rho_3\) to values away from this region avoids this pathology. As a default, we recommend restricting \(\rho_3^2<0.95\), as in most applied settings it is unlikely that a single confounder could explain more than 95% of treatment variation. This restriction still allows for a large degree of confounding in the treatment assignment process, while avoiding the region where the bias is extremely sensitive to small changes in \(\rho\). Fortunately, we have found in examining plots like Figure (ref) for other outcomes, and for other DGPs generated by samping \(\Sigma\) uniformly from the set of all correlation matrices, that patterns like the one in Figure (ref) are common: while there is variation in the RV for IV and proximal across values of \(\rho_3\), the relative positions of the RV values for the different identification strategies are relatively stable. Researchers are unlikely to incorrectly conclude, for instance, that their IV estimate is more or less robust than their SOO estimate due to a restriction on \(\rho_3\).
To estimate the RV in practice, researchers can use numerical optimization with sample estimates of the covariance matrix. Our parametrization \(\Sigma(\rho)\) is differentiable, and so researchers can use reverse-mode automatic differentiation to calculate exact gradients for the minimization problem in Eq. (ref). The bias equality constraint can be handled by penalization. Researchers can quantify estimation uncertainty by bootstrapping, as we demonstrate below.
The RV provides a helpful, single number summary of how much the underlying assumptions can be violated before risking a substantively meaningful change in our research conclusion, one that accounts for the nested dependency structure across the different identification assumptions. However, it cannot be used to determine which identification strategy has less (or more) bias necessarily, because the actual value of \(\rho\) is unknown. However, a small RV does signify that even a small violation in any underlying assumption could result in a substantial amount of bias, and warrants potential caution in evaluating the credibility of an estimated causal effect.
While the RV allows researchers to quantify and compare sensitivities, it measures different assumptions for each identification strategy, across which comparison may not always be straightforward. One approach to aid in these comparisons is benchmarking, which allows researchers to use the observed covariates as hypothetical unobserved confounders, to calibrate plausible values for the partial correlations parametrizing assumption violations. This strategy cannot be directly applied to the partial correlations in Section (ref), because the unobserved confounder \(U\) is often part of the conditioning set.\footnote{ When \(X\) and \(U\) are both in the conditioning set, treating a single covariate \(X_j\) as the confounder \(U\) does not change the value of the partial correlation at all.} The shared population covariance matrix parameterization here, however, allows us to benchmark the sensitivity parameters and the TRV for all the identification strategies.
For the \(j\)-th covariate (where \(j \in \{1, ..., |X|\}\)), we define \[ \hat \rho^{(j)} = \qty( \hat R_{X^{(j)} \sim W_Y \mid X^{-(j)}}, \hat R_{X^{(j)} \sim W_Z \mid W_Y, X^{-(j)}}, \hat R_{X^{(j)} \sim Z \mid W_Y, W_Z, X^{-(j)}}, \hat R_{X^{(j)} \sim Y \mid Z, W_Z, W_Y, X^{-(j)}} ), \] where \(X^{-(j)}\) is the set of covariates except the \(j\)-th. In this way, \(\hat \rho^{(j)}\) represents a set of partial correlations that would correspond to an omitted confounder \(U\) with equivalent confounding strength as an observed covariate \(X^{(j)}\). Then, because each partial correlation term in the \(A_\mathrm{iv}\) and \(A_\mathrm{prox}\) is a function of the elements of \(\rho\), we can directly estimate the benchmarked parameters, as well as the corresponding maximum assumption violation that would occur from omitting a confounder \(U\) with equivalent confounding strength to \(X^{(j)}\) (denoted by \(\norm{A(\hat \rho^{(j)})}^2_\infty\)). The benchmarked total assumption violation can be compared to the RV to evaluate how much stronger (or weaker) a hypothetical \(U\) would have to be to result in an assumption violation that would saturate the minimum violation strength represented by the RV.
Benchmarking is most useful in settings when researchers have strong substantive priors about the types of covariates that are prognostic of different aspects of the data generating process. In general, we caution that while benchmarking allows researchers to quantify the relative strength of a confounder in terms of observed covariates, these types of statements cannot be used to rule out the existence of confounding.
To see the proposed sensitivity summary measures in action, we return to our running example from hk2022poland, and calculate robustness values for the individual protest outcome under each identification strategy. For the bias threshold \(b\), we use the respective point estimates, so that a bias of \(b\) would move the estimated effect to zero. This choice of \(b\) favors proximal and especially IV, since their estimates errors are larger in magnitude.\footnote{ For readers concerned about the use of a different \(b\) across strategies, the appendix contains an analysis of the group protest outcome, where fortuitously all estimates, and thus all \(b\), are nearly equal.}
Table (ref) shows these bias thresholds and their corresponding total robustness values for each identification strategy. The table also shows the minimizing value of \(\rho\) from the optimization routine.\footnote{ We used a very small penalty on the total magnitude of \(\rho\) in order to encourage a unique solution to the optimization problem, but due to the limiations of numerical optimization, these values of \(\rho\) should be interpreted with some caution.} To account for uncertainty, we use a fractional-weighted bootstrap and report the median absolute deviation of the estimates across bootstrap iterations. The fractional-weighted bootstrap, a kind of Bayesian bootstrap, avoids collinearity issues when projecting out indicator variables that are only present in a few observations xu2020applications, and the use of MAD is appropriate given the heavy skew of some of the bootstrap distributions, which results from the quantities being bounded below by 0.
\setstretch{1.0}
\setstretch{1.0}
For both outcomes, selection-on-observables has the largest robustness value across the identification strategies; for the individual protest outcome, \(\mathrm{RV}_\mathrm{soo}=0.331\). In other words, if the unobserved confounder explains less than \(33\%\) of the variation in both treatment and outcome, then the bias would not exceed \(b=-0.296\). In comparison, the total robustness values for IV and proximal inference are much smaller. For IV, \(\mathrm{RV}_\mathrm{iv}=0.0173\) means that a violation of the exclusion restriction and exogeneity as small as \(0.0173\) on the \(R^2\) scale is sufficient to create bias as large as \(b=-0.805\). For proximal, \(\mathrm{RV}_\mathrm{prox}=0.011\) also leaves little room for even small violations of the proximal assumptions. This does not necessarily mean that there is a greater amount of bias in \(\tau_\mathrm{iv}\) and \(\tau_\mathrm{prox}\). However, it does mean that any realistic deviation from the proximal assumptions could result in a substantial amount of bias, and researchers should be cautious when drawing causal conclusions from \(\tau_\mathrm{iv}\) and \(\tau_\mathrm{prox}\).
To help reason about whether it is plausible that the magnitude of violations captured by the RVs could occur, we perform the benchmarking discussed above, estimate the maximum assumption violation that would occur for an omitted \(U\) that is as strong as an observed covariate \(X^{(j)}\). Table (ref) shows that across the different covariates, the maximum assumption violations for IV and proximal are consistently larger than the benchmarked maximum assumption violations for selection-on-observables. More specifically, SOO's RV is around 7 times larger than the largest benchmarked maximum assumption violation, whereas the RVs for IV and proximal are are around 5--9 times smaller than the benchmarked maximum assumption violations. Appendix (ref) contains benchmarked values for each identifying assumption, and a parallel analysis of the group protest outcome, where we observe similar patterns. Notably, even though all three estimators have nearly identical point estimates, both IV and proximal have very small RVs and so may be susceptible to bias in the presence of even small violations of their underlying assumptions.
\setstretch{1.0}
{
}
\setstretch{1.0}
The common parametrization from Section (ref) has deeper implications for comparing estimators. In particular, when researchers choose an identification strategy, they are implicitly expressing a belief about the degree of violations that must be present in alternative identification strategies. In this section, we show that all three estimator's biases can be re-expressed in terms of the standard SOO sensitivity parameters \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Z, W_Y}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Z, W_Y}}\}\). This means, for example, that if a researcher believes that the IV identification assumptions hold, then a certain magnitude of unobserved confounding must also be present. In some cases, this degree of confounding may be implausible. Researchers might then conclude that the IV assumptions cannot hold exactly. The bias re-expression also allows us to visualize the relative bias of each approach on a single bias contour plot.
Section (ref) parametrized the data-generating process in terms of four partial correlations, and used this parametrization to define and evaluate a set of numerical sensitivity summary measures. However, for the purposes of evaluating bias in \(\tau_\mathrm{soo}\), \(\tau_\mathrm{iv}\), and \(\tau_\mathrm{prox}\), we can reduce the number of parameters to just two.
Specifically, we can re-write the bias of \(\tau_\mathrm{iv}\) and \(\tau_\mathrm{prox}\) with the selection-on-observables parameters \(\textcolor{sooblue}{R_{Y\sim U\mid Z,W_Z,W_Y}}\) and \(\textcolor{sooblue}{R_{Z\sim U\mid W_Z,W_Y}}\). The key observation is that for a particular \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z}}\}\), the bias of the selection-on-observables estimate will be fixed, which implicitly also fixes the value of \(\tau\), the true treatment effect. As such, we can re-express the bias of \(\tau_\mathrm{iv}\) and \(\tau_\mathrm{prox}\) as
Both of these expressions depend only on the observed data and the value \(\mathrm{Bias}(\tau_\mathrm{soo}; \textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z})}\), which is fixed by the parameters \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z}}\}\). Note that the partial correlations corresponding to other assumptions, like the exclusion restriction \(R_{Y\sim W_Z\mid Z,W_Y,U}\) do depend on other parts of \(\rho\); it is only the bias of each estimate that can be reduced to exactly the SOO parameters.
This unification is advantageous for several reasons. First, Eq. (ref) gives us a route to translate assumptions made for one identification strategy into the language of treatment and outcome confounding. Suppose a researcher believes that their IV strategy is highly plausible, perhaps in part due to the IV-specific sensitivity analyses developed in this paper, and furthermore they are able to estimate the likely direction of IV bias. As a result, they might believe that most likely \(0<\mathrm{Bias}(\tau_\mathrm{iv})<0.1\). Eq. (ref) then connects this belief to an equivalent belief about likely values of \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z}}\}\). The researcher's judgement of the plausibility of these values can further inform their understanding of the IV bias, or even their choice of identification strategy.
Second, Eq. (ref) can be directly solved for the set of SOO sensitivity parameters where each identification strategy has lower bias than the others. Because Eq. (ref) is linear in all bias terms, each estimator will dominate the others on a particular interval of values of \(\tau\), corresponding in turn to a particular interval of values of \(\mathrm{Bias}(\tau_\mathrm{soo}; \textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z})}\). These intervals are exactly obtained by cutting the real line at the midpoints between the values of \(\tau_\mathrm{soo}\), \(\tau_\mathrm{iv}\), and \(\tau_\mathrm{prox}\). For example, if \(\tau_\mathrm{soo}=0.1\), \(\tau_\mathrm{iv}=0.4\), and \(\tau_\mathrm{prox}=0.2\), then SOO dominates for the \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z}}\}\) corresponding to any true \(\tau<0.15\), proximal inference dominates sensitivity parameters yielding \(0.15<\tau<0.3\), and IV dominates for parameters giving \(\tau>0.3\). We can map these intervals to the traditional sensitivity plot, discussed in the next section, as contour lines that bound regions where each estimator dominates. These regions can be visualized together, providing insight into the assumptions and sensitivity of each identification approach.
Readers will note that the manipulations in Eq. (ref) are in no way reliant on using the SOO sensitivity parameters specifically: the same re-expression could be applied in terms of the partial correlations for IV or proximal bias. However, the bias expressions for IV and proximal are substantially more complicated, and rely on unobservable partial correlations with no direct connection to the IV and proximal assumptions. Moreover, the subtraction that is present in those bias expressions means that bias is not monotonic in either sensitivity parameter the way that it is for SOO, which further complicates interpretation and visualization. Finally, treatment and outcome confounding are in many ways the core challenges of causal inference, and researchers are familiar with considering confounding mechanisms in these terms. Thus we focus specifically on the SOO parametrization, as we anticipate it will be easiest to reason about for applied researchers.
The unified bias framework also allows us to extend the standard SOO sensitivity plot and allow researchers to visually compare the relative biases across the identification strategies. Traditionally, bias contour plots have been employed to evaluate SOO bias by varying the two sensitivity parameters \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z}}\}\). A contour line corresponding to \(\tau=0\) is often plotted to help identify which values of sensitivity parameters would change the sign of the estimate. The benchmarked sensitivity parameters are often also plotted to help researchers reason about what could be plausible magnitudes of the underyling sensitivity parameters.
The augmented sensitivity plot we propose builds on this existing visualization in three ways. First, because the sign of the bias, and thus the dominating estimator, depends on the sign of the sensitivity parameters, we generate the sensitivity plot using the signed parameters rather than their square. Second, as described in the preceding section, we separate the contour lines according to the regions where each identification strategy dominates, indicating this partition by lines and colors. Third, we plot a contour line that corresponds to zero bias in either the proximal or IV estimate. If the assumptions of either of those strategies hold exactly, then the SOO sensitivity parameters must lie along the corresponding contour.
While our primary focus throughout the paper has been on the bias of different identification strategies, researchers may instead be interested in the relative mean-squared error of the different approaches. Notably, the same re-expression in Eq. (ref) can be applied for mean-squared error, and a similar augmented sensitivity plot can be produced. In some cases this may show that the zero-bias contour for proximal or IV is actually contained in the lowest-MSE region for another estimator. In other words, that even if the IV or proximal assumptions hold exactly, another estimator may be preferred due to lower mean-squared error. This phenomenon has been observed for IV estimates when the instrument is weak bound1995problems, hahn2005estimation. See Appendix (ref) for more discussion.
We now return to our running example, where we generate the proposed augmented sensitivity plot, including the benchmarked values of \(\{\textcolor{sooblue}{R_{Y \sim U \mid Z, W_Y, W_Z}}, \textcolor{sooblue}{R_{Z \sim U \mid W_Y, W_Z}}\}\) based on the observed covariates (see Figure (ref)) If an analyst had originally designed their observational study around IV, for instance, then their beliefs about the values of the SOO sensitivity parameters must lie close to the curve labeled “No IV Bias” in Figure (ref). In other words, belief about an identification strategy implies a relatively narrow set of beliefs on the values of the sensitivity parameters. Practitioners can compare that set of parameter values to the benchmark values of the covariates, which are plotted as gray crosses in Figure (ref). In this case, the benchmark values lie squarely in the region labeled “SOO least bias.” The largest benchmarked \(R_{Z\sim U\mid W_Z,W_Y}\) is \(0.0378\) and corresponds to (log) municipality population. In contrast, the smallest possible value of \(R_{Z\sim U\mid W_Z,W_Y}\) consistent with a valid proximal strategy is \(0.413\)--around \(11\) times larger. For IV, the factor grows to \(15\). Both of those factors assume worst-case outcome confounding, with \(R_{Y\sim U\mid Z,W_Z,W_Y}=1\), which is also several times larger than the confounding associated with any of the observed covariates.
Thus if a researcher is to proceed with a proximal or IV strategy, they must believe that the unobserved confounder is significantly stronger than any of the observed covariates' benchmark values, and that the treatment and outcome confounding are related in a way as to fall close to one of the no-bias contour lines in Figure (ref). One immediate qualitative implication is that treatment and outcome confounding must have the same sign for a proximal or IV approach to produce less bias. If they do not, then the sensitivity parameters lie in quadrants II or IV of Figure (ref), where the SOO estimate always dominates.
Figure (ref) additionally provides insight on the total robustness values in Section (ref). Near the origin, where the SOO assumptions approximately hold, the contour lines are sparsely spaced (note that the plotted contours in Figure (ref) are on a logarithmic scale). Thus, even moderate violations of the SOO assumptions lead to small changes in the value of \(\tau\), and thus small amounts of bias. In contrast, the “No Proximal Bias” and “No IV Bias” lines in Figure (ref) are much farther from the origin, where the contour lines are denser. Smaller changes in the sensitivity parameters in these regions lead to correspondingly larger changes in bias, which in turn lowers the total robustness values.
From Eq. (ref), the distance of the “No Proximal Bias” and “No IV Bias” lines from the origin is directly related to the difference between the point estimates for these identification strategies and the SOO estimate. The more the estimates differ, the more confounding is necessarily implied under each set of identification assumptions. As both Theorem (ref) and Theorem (ref) show, the bias of the IV and proximal estimates grows with the amount of outcome confounding. Thus a larger observed difference between the SOO estimate and IV or proximal estimates implies that the outcome confounding is larger, which in turn will generally translate to lower robustness.
This paper proposes a unified sensitivity framework for researchers to consider the relative bias of different identification strategies. While the unified model imposes some constraints compared to a fully nonparametric setup, the relative simplicity also clarifies the key drivers of bias under IV and proximal identification, and highlights connections between each estimator and their biases. As Section (ref) and Section (ref) show, the dimensionality of unidentified parts of the causal model is limited, and this creates (usually nonlinear) relationships between the various causal assumptions. Thus, by comparing benchmarking values and point estimates under each strategy, more can be said about each of the identifying assumptions. The bias expressions in Section (ref) and the proposed total robustness values also help researchers understand the sensitivity of each approach to violations of these assumptions.
There are several interesting avenues of future work. One natural direction would be to explore the extent to which these findings hold in a fully nonparametric setting. As chernozhukov2024long show for SOO, key intuition and structure from the linear setting can carry over with little modification to a nonparametric model. The two-stage estimation of IV and proximal, however, complicates such an analysis and would warrant careful investigation.
Second, recent work in observational causal inference has introduced an idea known as design sensitivity for observational studies, which allows researchers to consider sensitivity to potential violations in the underlying identification assumptions a priori rosenbaum2004design, rosenbaum2010design, huang2025design. Design sensitivity quantifies the tradeoffs in robustness that arise from different `design' choices (i.e., choosing an estimand, treatment definitions, study populations). However, much of this work focuses explicitly on the SOO setting. Future work could build on the unified sensitivity framework presented in this paper to consider design sensitivity across different identification strategies.