EconBase
← Back to paper

Sensitivity Analysis for Linear Estimators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

71,088 characters · 13 sections · 53 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sensitivity Analysis for Linear Estimators

\onehalfspacing

abstractWe propose a novel sensitivity analysis framework for linear estimators with identification failures that can be viewed as seeing the wrong outcome distribution. Our approach measures the degree of identification failure through the change in measure between the observed distribution and a hypothetical target distribution that would identify the causal parameter of interest. The framework yields a sensitivity analysis that generalizes existing bounds for Average Potential Outcome (APO), Regression Discontinuity (RD), and instrumental variables (IV) exclusion failure designs. Our partial identification results extend results from the APO context to allow even unbounded likelihood ratios. Our proposed sensitivity analysis consistently estimates sharp bounds under plausible conditions and estimates valid bounds under mild conditions. We find that our method performs well in simulations even when targeting a discontinuous and nearly infinite bound.

Introduction

Many important estimators in economics are weighted averages of observed outcomes. These estimators leverage powerful identifying restrictions to target meaningful estimands. In observational settings, these identifying restrictions may not be satisfied: many treatment choices can be selected on unobservables, manipulators can strategically sort across treatment cutoffs, and instruments can directly affect outcomes.

In this paper, we propose a framework for the sensitivity analysis of identification failures for linear estimators. The framework proceeds as follows. First, we define a target distribution: a synthetic distribution over observed variables that would enable a standard estimator to point-identify the causal estimand of interest. Second, we consider a structural model that implies restrictions on the divergence between the target population and the population the practitioner observes. Third, we leverage work from the literature in statistics and distributionally robust optimization to estimate bounds implied by the restrictions from the second step. The framework is especially powerful when the implied restrictions correspond to a family of restrictions on the Radon-Nikodym derivative (the generalization of likelihood ratio to allow point masses) between the distribution of observables under the target distribution and under the distribution the observed data is drawn from.

Our proposed framework has several empirical advantages. The approach typically yields a sensitivity analysis: under the strongest restriction, the bounds reduce to the standard point estimate; under the weakest restriction, the bounds reduce to a worst-case exercise. Between the two extremes, partial identification bounds can often be estimated using modern tools from the literature on distributionally robust optimization (DRO). The partial identification bounds can often be written as an explicit closed form moment of the observed data.

We illustrate our framework for $L_\infty$ bounds on the Radon–Nikodym derivative between the target and observed distributions. We extend existing results from the average potential outcome (APO) literature where distributional restrictions correspond to limits on treatment selection tan2006distributional, masten2018identification, zhao2019sensitivity, dorn2022sharp. We extend this literature to obtain a simple closed-form characterization of the identified set, even with unbounded changes in measure. The plug-in estimator is consistent for the sharp bounds when given consistent estimates of the primary analysis nuisance functions and a certain conditional outcome quantile function. When the conditional outcome quantile function is inconsistent, the plug-in estimator yields valid bounds that are too wide rather than too narrow.

We apply our framework to yield novel results for three applications. In our first application, we study bounds on the APO with a treatment that may be selected on an unobserved confounder. Our framework nests unconfoundedness, Manski-type bounds that restrict only the support of the unobserved potential outcomes, and tan2006distributional's Marginal Sensitivity Model as special cases. As a corollary, we obtain a simpler characterization of bounds under masten2018identification's conditional c-dependence model.

In our second application, we study bounds on causal effects from sharp RD with one-sided manipulation. As mccrary2008manipulation notes, the RD assumption of no manipulation is testable. When mccrary2008manipulation's test fails, gerard2020bounds propose a worst-case bound on the conditional average treatment effect for non-manipulators: a Conditional Local Average Treatment Effect (CLATE). We show that a stronger restriction on manipulation choice allows us to obtain meaningful bounds on the more standard conditional average treatment effect (CATE), which to the best of our knowledge has been an open issue in the literature. Our framework provides a sensitivity analysis by nesting an unconfoundedness-type assumption and the gerard2020bounds assumption as extreme cases.

In our third application, we study treatment effects with an instrument that fails the exclusion restriction. When the instrument is allowed to affect the outcome directly, we consider estimation of a generic weighted average of local average treatment effects (LATEs) across instrument values. In the continuous outcome case, we contribute a sensitivity analysis for a measure of exclusion failure that is unit-free. The proposed bounds are simple and tractable, but can be wider than the bounds implied by the original model.

We illustrate the value of our approach by simulation. We study estimation of APO bounds under masten2018identification's conditional c-dependence model, which restricts the difference between observed and true propensities to be at most $c$. $c = 0$ corresponds to unconfoundedness; in our example, the bounds discontinuously become unbounded as $c$ crosses $0.1$. We implement a simple plug-in estimator and percentile bootstrap that leverage a fixed grid of quantile estimates across bootstraps and values of $c$. We find that a plug-in boostrap approach achieves excellent bias and coverage properties for small $c$; increasingly over-covers as $c$ grows towards $0.1$; and nearly achieves nominal coverage of the real line at the most difficult $c = 0.1$ case.

Related Work

We now mention some related work.

Our framework unifies ideas from the operations research and economics literature. Two closely related frameworks are bertsimas2022distributionally and ChristensenAndConnault, both of which are limited to discrete covariates. In the discrete covariate case, relative to bertsimas2022distributionally, we propose justifying distributional distances in terms of underlying structural models of economic failure and propose constructing target distributions that apply in settings beyond average treatment effects, and offer a specific analysis of $L_\infty$ bounds rather than other distributional distances. Relative to ChristensenAndConnault in the discrete covariate case, we analyze a nominally different class of identification failures and provide closed-form bounds for specifically linear estimators.

Our work is related to the recent literature on sensitivity analysis for IPW estimators, which relates to our first application. A sensitivity analysis is an approach to partial identification that begins from assumptions that point-identify the causal estimand of interest and then considers increasing relaxations of those assumptions MolinariChapter. Our analysis is an extension of dorn2022sharp's sharp characterization of bounds under tan2006distributional's marginal sensitivity model. tan2022modelassisted and frauen2023sharp previously extended this characterization to families that bound the Radon-Nikodym derivative of interest. We generalize these results to also include unbounded Radon-Nikodym derivative, so that we can include a compact characterization of bounds under masten2018identification's conditional c-dependence model as a special case. There is rich work in this literature under other sensitivity assumptions like $f$-divergences and Total Variation distance. These other assumptions also fit within our framework, because our target distribution constructions are independent of the $L_\infty$ sensitivity assumptions that we analyze.

Our other two applications relate to existing work on sensitivity analysis. Our proposed sensitivity analysis for sharp regression discontinuity (RD) applies when data on the running variable fails tests for manipulation mccrary2008manipulation, otsu2013estimation, bugni2021testing. Our proposal nests both an exogeneity-type assumption and gerard2020bounds's bounds as special cases. There is other work in the RD context on partial identification bounds under manipulation rosenman2019optimized, ishihara2020manipulation but to our knowledge, our proposal is the first sensitivity analysis for manipulation. There are sensitivity analysis for exclusion failure with instrumental variables ramsahai2012causal, van2018beyond, masten2021salvaging, freidling2022sensitivity, but to our knowledge our proposal is the first sensitivity analysis whose underlying assumptions are invariant to invertible transformations of variables.

Notation. We use $\bar{\mathds{R}}$ to refer to the extended real number line $\mathds{R} \cup \{ -\infty, \infty \}$. For a real-valued random variable $Z$ and a distribution $\mathbb{Q}$, we use the notation $E_{\mathbb{Q}}[Z] = \int z d \mathbb{Q}(z)$ and we write that the expected value of $Z$ exists under $\mathbb{Q}$ if $E_{\mathbb{Q}}[|Z|]$ is finite or $E_{\mathbb{Q}}[Z]$ is well-defined as exactly one of positive or negative infinity. We write that the expected value of $Z$ exists (without specifying the distribution) if the expected value exists under the observed distribution $\mathbb{P}^{\textup{Obs}}$, where $\mathbb{P}^{\textup{Obs}}$ is defined below. Similarly, we sometimes suppress the dependence of expectations when referring to the expectation under the observed distribution $\mathbb{P}^{\textup{Obs}}$. We write $\{ a, b \}_+ = \max\{a, b\}$ and $\{a, b\}_- = \min\{a, b\}$, and abuse notation by writing $\{a\}_+ = \max\{a, 0\}$. For random variables $Y \in \mathds{R}^1$ and $R \in \mathds{R}^d$ and a function $t: \mathds{R}^d \to [0, 1]$, we refer to “the" conditional quantile function $Q_{t(R)}(Y \mid R)$, which is any minimizer of $E_{\mathbb{P}^{\textup{Obs}}}[ t(R) \{ Y - Q(R), 0 \}_+ - (1-t(R)) \{ Y - Q(R), 0\}_- ]$ in functions $Q$ from the domain of $R$ to the extended real line $\mathds{R} \cup \{-\infty, \infty\}$. To consolidate notation, when $b$ is infinite, we evaluate the interval $[a, b]$ as $[a, \infty)$. If our estimation procedure calls for an estimate of a nuisance function $f$ that depends on another nuisance function $g$, we use the notation $\hat{f}$ to denote the full estimated nuisance function, which may include a composition of nuisance estimates. We use $1\{ . \}$ to refer to the indicator function that takes value 1 if true and 0 otherwise.

Framework for Sensitivity Analysis

In the following section, we illustrate our generic framework for sensitivity analysis: we propose linking a structural model of identification failures to an implied statistical distribution bound and then conducting partial identification under the distributional bound. We illustrate our approach with a family of $L_\infty$ distributional distances that is especially tractable for estimation. We illustrate how our framework applies to specific structural models on (ref).

General Setting

We begin with our general partial identification setting. We use the average potential outcome application as a running example.

There is a target distribution $\mathbb{P}^{\textup{Target}}$ over $(R, Y)$ that point identifies a causal estimand of interest, where $Y \in \mathds{R}^1$ is the outcome of interest and $R \in \mathds{R}^d$ are observable quantities that do not include $Y$. We call $R$ the regressors. However, the researcher only has access to data from an observed distribution $\mathbb{P}^{\textup{Obs}}$. $\mathbb{P}^{\textup{Obs}}$ is also a distribution over $(R, Y)$, but $\mathbb{P}^{\textup{Obs}}$ may not be able to point-identify the causal estimand under the researcher's preferred estimator, for example if $R$ is a selected treatment or if $R$ includes an instrument that fails the exclusion restriction.

manualassumption{Support} The marginal distribution of $R$ is the same under $\mathbb{P}^{\textup{Target}}$ and $\mathbb{P}^{\textup{Obs}}$. The distribution $\mathbb{P}^{\textup{Target}}$ is absolutely continuous with respect to $\mathbb{P}^{\textup{Obs}}$.

(ref) has two components. The first component is that $\mathbb{P}^{\textup{Obs}}$ identifies the target distribution of regressors $R$, but may have a different conditional distribution of outcomes $Y$. The second component is an absolute continuity assumption that ensures the existence of a Radon-Nikodym derivative $\frac{d \mathbb{P}^{\textup{Target}}}{d \mathbb{P}^{\textup{Obs}}}$. The assumption accommodates unbounded Radon-Nikodym derivatives, for example $Y \mid R \sim Unif(0, 1)$ under $\mathbb{P}^{\textup{Obs}}$ and $f_{Y}(y \mid R) = 1\{y \in (0, 1]\} y^{-1/2} / 2$ under $\mathbb{P}^{\textup{Target}}$. The absolute continuity assumption is stronger than necessary for our results, and could be reduced to assuming that the support of $(Y, R)$ under $\mathbb{P}^{\textup{Target}}$ is a subset of the support under $\mathbb{P}^{\textup{Obs}}$. Absolute continuity rules out $\mathbb{P}^{\textup{Target}}$ possessing mass points where $\mathbb{P}^{\textup{Obs}}$ lacks mass points, so that such distributions can only correspond to limits of distributions within (ref).

Our framework is inspired by the literature on sensitivity analysis for average treatment effects. In this case, there is a distribution $\mathbb{P}^{\textup{True}}$ over $(X, T, Y(1), Y(0))$, but we only observe the distribution $\mathbb{P}^{\textup{Obs}}$ over $(X, T, Y = Y(T))$. The regressors are $R = (X, T)$. For simplicity, we illustrate our approach for the Average Potential Outcome (APO) $E_{\mathbb{P}^{\textup{True}}}[ E_{\mathbb{P}^{\textup{True}}}[ Y(1) \mid R] ]$. The researcher can accurately estimate $\mathbb{P}^{\textup{True}}( T \mid X )$, but cannot observe $\mathbb{P}^{\textup{True}}( Y(1) \mid X, T = 0)$. As a result, our proposed approach targets a distribution $\mathbb{P}^{\textup{Target}}$ that first samples $(X, T) \sim \mathbb{P}^{\textup{True}}$ and then samples $Y \mid X, T$ from the distribution of $Y(T) \mid X$ under $\mathbb{P}^{\textup{True}}$. If a causal estimator like an inverse propensity weighted (IPW) estimator were applied to $\mathbb{P}^{\textup{Target}}$ instead of $\mathbb{P}^{\textup{Obs}}$, the researcher would identify the APO under $\mathbb{P}^{\textup{True}}$. Our later examples illustrate the applicability of this framework for other causal estimands.

We refer to our target as the causal estimand. We restrict ourselves to estimands that correspond to linear estimators.

definitionThe causal estimand is $\psi_0 = E_{\mathbb{P}^{\textup{Target}}}[ \lambda(R) Y ]$ for some real-valued function $\lambda$. The primary estimand is $E_{\mathbb{P}^{\textup{Obs}}}[ \lambda(R) Y]$.

We generally assume that $\lambda$ is identified and the researcher has conducted a primary analysis that consistently estimates $\lambda$. For example, in the APO case, the $\lambda$ function is the inverse propensity $\frac{T}{\mathbb{P}^{\textup{Obs}}(T = 1 \mid X)}$. When the researcher is able to make strong enough assumptions that $\mathbb{P}^{\textup{Target}}$ is equivalent to $\mathbb{P}^{\textup{Obs}}$, then the primary estimand will be equal to the causal estimand. When the researcher is only able to bound the difference between $\mathbb{P}^{\textup{Target}}$ and $\mathbb{P}^{\textup{Obs}}$, then the researcher can partially identify the causal estimand through a restriction on the Radon-Nikodym derivative between the two distributions.

lemmaSuppose (ref) holds and $\lambda(R) Y$ is integrable under $\mathbb{P}^{\textup{Target}}$. Then $$\psi_0 = E_{\mathbb{P}^{\textup{Obs}}} \left[ \lambda(R) Y \frac{\textup{d} \mathbb{P}^{\textup{Target}}( Y \mid R) }{d \mathbb{P}^{\textup{Obs}}(Y \mid R)} \right].$$
proof(ref) is a standard result for Importance Sampling and an immediate property of the Radon-Nikodym derivative. All other proofs can be found in (ref).

This reweighing characterization is useful for partial identification because any non-negative putative outcome reweighing $\bar{W}$ that satisfies $E[ \bar{W} \mid R] = 1$ almost surely will correspond to a well-defined distribution $\mathbb{Q}$ over $(R, Y)$.

The Radon-Nikodym derivative characterization in (ref) often maps to interpretable quantities. For example, in the APO case, the Radon-Nikodym derivative maps to odds ratios as follows:

align*[align* omitted — 422 chars of source]

As a result, there is a bijection between structural models of treatment selection and causal models of treatment effects. More generally, principled restrictions on an underlying treatment selection model may imply, or be equivalent to, restrictions on the Radon-Nikodym derivative.

Partial Identification Assumption and Result

In the this subsection, we characterize the sharp bounds on the identified set under an $L_\infty$ restriction on the largest and smallest Radon-Nikodym derivative.

definition[Sensitivity assumption] For any pair of functions $\underline{w}, \bar{w}: \mathds{R}^d \to \bar{\mathds{R}}$ satisfying $0 \leq \underline{w}(R) \leq 1 \leq \bar{w}(R)$ almost surely, we define $\mathcal{M}( \underline{w}, \bar{w} )$ as the set of distributions $\mathbb{Q}$ over $(R, Y)$ satisfying $\lambda(R) \frac{d \mathbb{Q}( R, Y )}{d \mathbb{P}^{\textup{Obs}}( R, Y)} \in [\lambda(R) \underline{w}(R), \lambda(R) \bar{w}(R)]$ almost surely.

This family has several advantages. The restrictions on $\frac{d \mathbb{Q}}{d \mathbb{P}^{\textup{Obs}}}$ decouple across values of $R$, enabling tractable characterizations of the identified set. The family nests both strong observational assumptions and Manski-type bounds as limits. Point identification corresponds to the case $\underline{w}(R) = \bar{w}(R) = 1$ almost surely. Manski-type bounds that only restrict the support of $Y$ correspond to $\bar{w}(R) = \infty$ with domain-appropriate $\underline{w}(R)$. In between, the causal estimand is only point identified. When the outcome $Y$ is binary, the restriction can equivalently be viewed as a restriction on the conditional mean of $Y \mid R$. We show below that, as in the tan2006distributional model that inspired this generalization, the resulting bounds are highly tractable for estimating sharp and valid bounds.

We adapt standard notation from HoAndRosen.

definitionThe identified set is $\mathcal{I}(\underline{w}, \bar{w} ) = \{ E_{\mathbb{Q}}[ \lambda(R) Y ] \mid \mathbb{Q} \in \mathcal{M}( \underline{w}, \bar{w} ) \}$. The sharp bounds on the identified set are $\psi^-(\underline{w}, \bar{w} ) = \inf_{\psi \in \mathcal{I}(\underline{w}, \bar{w} )} \psi$ and $\psi^+(\underline{w}, \bar{w} ) = \sup_{\psi \in \mathcal{I}(\underline{w}, \bar{w} )} \psi$. A valid identified set is a weak superset of $\mathcal{I}( \underline{w}, \bar{w} )$. We call bounds that are strictly weaker than the sharp bounds, conservative.

We abuse notation and simply write $\mathcal{I}$, $\psi^-$, and $\psi^+$ as shorthand for the identified set and bounds under a generic model family. (ref) corresponds to the statistical partial identification bounds implied by an underlying structural model. As we illustrate in our applications, a structural model underlying (ref) can sometimes, but not always, imply a narrower bound.

We will require some nuisance functions to characterize the identified set. In particular:

definitionThe threshold probability is $\tau(R) = \frac{\bar{w}(R) - 1}{\bar{w}(R) - \underline{w}(R)}$. The threshold quantiles are $Q^+(R) \equiv Q_{\tau(R)}(\lambda(R) Y \mid R)$ and $Q^-(R) \equiv Q_{1-\tau(R)}(\lambda(R) Y \mid R)$. The likelihood shifting term is $a(\underline{w}, \bar{w}, s) = (\bar{w} - \underline{w}) 1\{ s > 0 \} - (1-\underline{w})$. For a function $\bar{Q}: \mathds{R}^d \to \bar{\mathds{R}}$, define the pseudo-outcome $$\phi^+(\underline{w}, \bar{w}, r, y \mid \bar{Q}) \equiv \lambda(r) y + ( \lambda(r) y - \bar{Q}(r)) a( \bar{w}(r), \underline{w}(r), \lambda(r) y - \bar{Q}(r)),$$ and define $\phi^-(\underline{w}, \bar{w}, r, y \mid \bar{Q})$ analogously. Define the indicator variables $F = 1\{ \bar{w}(R) \text{ finite}\}$, $G^+ = 1\{ \lambda(R) Y - Q^+(R) > 0 \}$, $G^- = 1\{ \lambda(R) Y - Q^-(R) < 0 \}$, and $H = 1\{ \tau(R) < 1 \}$.

When $\bar{w}(R)$ is finite, the formula for $\tau(R)$ can be found in tan2022modelassisted and frauen2023sharp.

We will make additional regularity assumptions.

manualassumption{Moments} The expected values of $F |Q^+(R)|$, $F |Q^-(R)|$, $|\phi^+(\underline{w}, \bar{w}, R, Y \mid Q^+)|$ and $|\phi^-(\underline{w}, \bar{w}, R, Y \mid Q^-)|$ exist.

We allow infinite $E_{\mathbb{P}^{\textup{Obs}}}[ Q^+(R) ]$ for completeness with unbounded outcomes. While $\bar{w}$ can be infinite, we rule out heavy tails that would yield infinite $E_{\mathbb{P}^{\textup{Obs}}}[ 1\{ \bar{w}(R) \text{ finite}\} Q^+(R) ]$.

Our work focuses on the upper bound of the identified set.

theoremSuppose Assumptions (ref) and (ref) hold. Then the sharp upper bound on the identified set satisfies: \begin{align} \psi^+( w, \bar{w} ) & = E_{\mathbb{P}^{Obs}}\left[ \lambda(R) Y + ( \lambda(R) Y - Q^+(R) ) a( \bar{w}(R), w(R), \lambda(R) Y - Q^+(R) ) \right]. \end{align} The lower bound $\psi^-( \underline{w}, \bar{w} ) $ follows symmetrically, i.e., \begin{align} \psi^-( w, \bar{w} ) & = E_{\mathbb{P}^{Obs}}\left[ \lambda(R) Y + ( \lambda(R) Y - Q^-(R) ) a( \bar{w}(R), w(R), Q^-(R) - \lambda(R) Y ) \right]. \end{align} Further, the identified set is convex.

The theorem states that the sharp upper bound can be written as a closed-form moment: after the constituent components such as $\lambda(.), Q^+(.),$ and $a(.)$ have been estimated, no further optimization problem needs to be solved.

A proof sketch is in order. We abuse notation to write $\psi^+ = E_{\mathbb{P}^{\textup{Obs}}}[\psi^+(R)]$, where $\psi^+(r)$ is a pointwise upper bound function satisfying:

align*[align* omitted — 228 chars of source]

and $\phi^+(R) \equiv E_{\mathbb{P}^{\textup{Obs}}}[ \phi^+(\underline{w}, \bar{w}, R, Y \mid Q^+) \mid R ]$. The claim is $E_{\mathbb{P}^{\textup{Obs}}}[\psi^+(R)] = E_{\mathbb{P}^{\textup{Obs}}}[ \phi^+(R) ]$. We prove the stronger claim that $\psi^+(R) = \phi^+(R)$ almost surely. For simplicity, this sketch will ignore the possibility that $\lambda(R) Y = Q^+(R)$ with positive probability and drop “almost sure" caveats. As dvds note, $\psi^+(R)$ can be mapped to a simple DRO problem over $(W - \underline{w}(R)) / (1 - \underline{w}(R)) \in [0, (\bar{w}(R) - \underline{w}(R)) / (1-\underline{w}(R))]$. The solution is $\psi^+(R) = E_{\mathbb{P}^{\textup{Obs}}}[ \underline{w}(R) \lambda(R) Y + (1-\underline{w}(R)) CVaR_{\tau(R)}^+(R) \mid R]$, where $CVaR^+_{\tau(R)} = E\left[ Q^+ + \frac{\{Y - Q^+\}_+}{1-\tau(R)} \mid R \right]$ is the level-$\tau(R)$ conditional value at risk of $\lambda(R) Y$. We split on the event $\tau(R) < 1$. By previous work tan2022modelassisted, frauen2023sharp, $\tau(R) < 1$ implies $\psi^+(R) = \phi^+(R)$. We extend these results for the case $\tau(R) = 1$. On that event, $\phi^+(R)$ evaluates to $E[ \underline{w}(R) \lambda(R) Y + (1 - \underline{w}(R)) Q^+(R) \mid R ]$ and $Q^+(R) = CVaR_{1}^+(R)$. As a result, $\psi^+(R) = \phi^+(R)$ even if $\bar{w}(R)$ is infinite with positive probability.

One could characterize the partial identification bounds using the conditional value at risk directly, but the reweighing characterization (ref) is valuable for sensitivity analysis. We present the worst-case Radon-Nikodym derivative here for convenience.

lemmaSuppose $\bar{w}(R)$ is bounded. Then one can construct a distribution $\mathbb{Q}^+ \in \mathcal{M}( \underline{w}, \bar{w} )$ and a random variable satisfying $\gamma(R) \in [\underline{w}(R), \bar{w}(R)]$ almost surely such that $\psi^+ = E_{\mathbb{Q}^+}[ \lambda(R) Y]$ and: \begin{equation} \frac{d \mathbb{Q}^+(R, Y)}{d \mathbb{P}^{Obs}(R, Y)} = W^{*}_{sup} =\begin{cases} \bar{w}(R) & if \quad \lambda(R)Y > Q^+(R) \\ w(R) & if \quad \lambda(R)Y < Q^+(R) \\ \gamma(R) & if \quad \lambda(R)Y = Q^+(R). \end{cases} \end{equation}

A researcher with a primary estimate of $\lambda(R)$ and who is capable of quantile regression can easily construct a plug-in estimate of $\psi^+$. Even if the researcher fails to consistently estimate the additional nuisance parameter $Q^+$, they can still easily estimate valid bounds:

lemmaSuppose Assumptions (ref) and (ref) hold and there is a putative quantile function $\bar{Q}$ such that the expectation of $|\phi^+(\underline{w}, \bar{w}, r, y \mid \bar{Q})|$ exists. Then replacing $Q^+$ with the potentially incorrect quantile function $\bar{Q}$ yields valid bounds: \begin{align*} \psi^+( w, \bar{w} ) & \leq E_{\mathbb{P}^{Obs}}\left[ \lambda(R) Y + ( \lambda(R) Y - \bar{Q}(R) ) a( \bar{w}(R), w(R), \lambda(R) Y - \bar{Q}(R) ) \right]. \end{align*} An analogous result holds for the lower bound $\psi^-$.

The argument follows by tan2022modelassisted's argument writing $(\lambda(R) Y - \bar{Q}(R)) a( \hdots )$ in terms of the Quantile Regression check function. $Q^+(R)$ is the minimizer by classic arguments.

Note that a similar recipe as in this section extends beyond the family we consider. For other assumptions, the model family in (ref) would change, the sharp upper bound in (ref) would change and may lack a closed form, and the validity (ref) may or may not hold, but the fundamental logic would carry through.

Applications

In this section, we illustrate how to apply our framework to several potential failures of identifying assumptions: APOs with selection on unobservables, RD with manipulation, and IV without exclusion. We also discuss some other applications of our framework. We show that the sharp statistical bounds under our approach contain all restrictions implied by the underlying structural model for APO selection and RD manipulation, but not IV exclusion failure.

Average Potential Outcomes

In this section, we illustrate our framework in a case that generalizes existing results: bounds on average potential outcomes under restrictions on selection on unobservables. Many of our results are restatements of recent work, but we include an extension of partial identification bounds to unbounded Radon-Nikodym derivatives.

We assume there is a distribution $\mathbb{P}^{\textup{True}}$ over $(X, T, \{Y(t)\}, U)$, where $X$ are controls, $T$ is a discrete treatment, $Y(t)$ is the potential outcome corresponding to treatment level $t$, and $U$ are potential unobserved confounders. We only observe the coarsened distribution $\mathbb{P}^{\textup{Obs}}$ over $(X, T, Y = Y(T))$, where $Y$ is the observed outcome. We write $I_t = 1\{T = t\}$.

We tailor our application to inverse propensity weighting (IPW) estimation of the APO. (We consider average treatment effects at the end of the section.) We write the observable propensity $e(X) = \mathbb{P}^{\textup{Obs}}(T = 1 \mid X)$, where we assume $e(X) \in (0, 1)$ almost surely. The primary estimate is the IPW estimator that first estimates the $e(X)$ and then estimates the APO as $E_{\mathbb{P}^{\textup{Obs}}}[ \lambda(R) Y ]$, where $\lambda(R) = I_1 / e(X)$. When unconfoundedness holds given the observed covariates so that $I_t \rotatebox[origin=c]{90}{$\models$} Y(t) \mid X$, the IPW strategy consistently estimates the causal estimand $E_{\mathbb{P}^{\textup{True}}}[Y(1)] $. However, we will only assume unconfoundedness holds if the researcher had access to both the observed and unobserved covariates: $I_t \rotatebox[origin=c]{90}{$\models$} Y(t) \mid X, U$.\footnote{This is without loss of generality by setting $U = \{ Y(t) \}$.}

The target distribution $\mathbb{P}^{\textup{Target}}$ is defined as $\mathbb{P}^{\textup{Target}}(X, T, Y) \equiv \mathbb{P}^{\textup{Obs}}(X, T) \mathbb{P}^{\textup{True}}(Y \mid X, T)$. The distribution of $R$ is the same under $\mathbb{P}^{\textup{Target}}$ and $\mathbb{P}^{\textup{Obs}}$ and the distribution $\mathbb{P}^{\textup{Target}}$ satisfies $E_{\mathbb{P}^{\textup{True}}}[Y(1)] = E_{\mathbb{P}^{\textup{Target}}}[\lambda(R) Y]$, so that $E_{\mathbb{P}^{\textup{Target}}}[\lambda(R) Y]$ is the causal estimand. Notice that $\mathbb{P}^{\textup{Target}}$ and $\mathbb{P}^{\textup{Obs}}$ may have different distributions of $Y \mid X$: the equivalent construction of $\mathbb{P}^{\textup{Obs}}$ could be factored as $\mathbb{P}^{\textup{Obs}}(X, T, Y) = \mathbb{P}^{\textup{Obs}}(X, T) \mathbb{P}^{\textup{Obs}}(Y \mid X, T)$.

Many researchers study models that imply there are functions $\ell(X), u(X)$ such that $$\ell(X) \leq \frac{e(X, U) / (1-e(X, U))}{e(X) / (1-e(X))} \leq u(X)$$ almost surely Manski1990, tan2006distributional, AronowLeeInterpretable, masten2018identification, zhao2019sensitivity, dvds, tan2023sensitivity, frauen2023sharp. This model implies pointwise restrictions on $\mathbb{P}^{\textup{True}}(T = 1 \mid X, Y(1))$ and the Radon-Nikodym derivative $\frac{d \mathbb{P}^{\textup{Target}}(Y \mid X, T=1)}{d \mathbb{P}^{\textup{Obs}}(Y \mid X, T=1)}$. In particular, the selection assumption would imply the following almost sure bound:

align*[align* omitted — 226 chars of source]

Unconfoundedness corresponds to the special case $\ell(X) = u(X) = 1$. Manski bounds correspond to the special case $\ell(X) = 0$, $u(X) = \infty$. tan2006distributional's Marginal Sensitivity Model corresponds to $\underline{w}(X, t) = \mathbb{P}^{\textup{Obs}}(T = t \mid X) + \mathbb{P}^{\textup{Obs}}(T \neq t \mid X) \Lambda^{-1}$ and $\bar{w}(X, t) = \mathbb{P}^{\textup{Obs}}(T = t \mid X) + \mathbb{P}^{\textup{Obs}}(T \neq t \mid X) \Lambda$. basit2023riskratiobased's risk ratio marginal sensitivity model, that restricts $e(X, U) \in \left[ \Gamma_0^{-1} e(X), \Gamma_1^{-1} e(X) \right]$, corresponds to $\underline{w}(R) = \Gamma_1$ and $\bar{w}(R) = \Gamma_0$ when $e(X) \leq \Gamma_1$. masten2018identification's conditional c-dependence model, that restricts $e(X,U) \in \left[ e(X) - c, e(X) + c\right] \cap [0, 1]$, corresponds to $\underline{w}(X, t) = \mathbb{P}^{\textup{Obs}}(T = t \mid X) + \mathbb{P}^{\textup{Obs}}(T \neq t \mid X) \frac{\max\{0, 1-(e(X)+c)\} / \min\{ 1, e(X) + c \}}{(1-e(X)) / e(X)}$ and $\bar{w}(X, t) = \mathbb{P}^{\textup{Obs}}(T = t \mid X) + \mathbb{P}^{\textup{Obs}}(T \neq t \mid X) \frac{\min\{1, 1-(e(X)-c)\} / \max\{ 0, e(X) - c \}}{(1-e(X)) / e(X)}$.

(ref) yields valid bounds on the identified set with the following quantities:

propositionIn the APO case, our method can be implemented on $E_{\mathbb{P}^{\textup{True}}}[Y(1)]$ with: \begin{align*} \lambda(R) = \frac{I_1}{e(X)}, \quad w(R) = e(X) + (1-e(X)) u(X)^{-1}, and \bar{w}(R) = e(X) + (1-e(X)) \ell(X)^{-1}. \end{align*}

The result in (ref) extends the analysis of tan2022modelassisted, frauen2023sharp to allow $\ell(X)$ to be arbitrarily small or equal to zero. As a result, it includes masten2018identification's conditional c-dependence assumption so long as $\lambda(R) Y$ is integrable under the target distribution. The characterization of bounds under conditional c-dependence is formally simpler than masten2018identification's characterization: their characterization involves an integral over a transformation of the full quantile regression function. Our approach also yields estimates of valid bounds under only the IPW estimation assumptions. An interesting question for future work is whether MastenPoirierZhang's proposed estimator for conditional c-dependence possesses similar validity guarantees.

The result is intuitive. The $w$ bounds would be $1$ under the unconfoundedness assumption $\ell(X) = u(X) = 1$. As $e(X)$ gets closer to one, we see a greater share of the treated potential outcomes and the $w$ bounds grow closer to one. An analogous approach can be used to obtain $E[Y(0)] = E[Y(1-Z)/(1-e(X,Y(0)))]$.

The resulting bounds are sharp for the underlying structural model for the APO and average treatment effect (ATE) estimands. The only restrictions on the distribution of $(Y, e(X, U)) \mid X, T=1$ implied by $\mathbb{P}^{\textup{Obs}}$ are the requirements that $E_\mathbb{P}^{\textup{Obs}}[ 1 / e(X, U) \mid X, T=1] = 1/e(X)$ almost surely. If we instead solved for the set of feasible $E_{\mathbb{P}^{\textup{Obs}}}[ I_1 Y / \bar{E} ]$ subject to $\ell(X) \leq \frac{\bar{E} / (1-\bar{E})}{e(X) / (1-e(X))} \leq u(X)$ and $E_\mathbb{P}^{\textup{Obs}}[ 1 / \bar{E} \mid X, T=1] = 1/e(X)$ almost surely, we would obtain a convex set with the same bounds. For the ATE with two potential treatment statuses, an extension of dorn2022sharp's sharpness argument shows sharpness continues to hold under pointwise restrictions on the true propensity.

Regression Discontinuity

In this section, we introduce a novel sensitivity analysis for sharp RD designs. Our framework allows us to bound standard causal estimands that previous work could not meaningfully quantify.

We work under a slight modification of gerard2020bounds's model. There is a full distribution $\mathbb{P}^{\textup{True}}$ over $(M, X(1), X(0), Y(1), Y(0), T)$, where $M$ is a variable corresponding to manipulation status, $X(m)$ is a potential running variable corresponding to manipulation status $M = m$, $Y(t)$ is a potential outcome corresponding to treatment status $T = t$, and $T \in \{0, 1\}$ is the treatment status. We face the fundamental problem of causal inference and only observe the coarsening $\mathbb{P}^{\textup{Obs}}$ over $(X = X(M), Y = Y(T), T)$.

We study sharp RD designs. We assume there is a cutoff $c$ such that $\mathbb{P}^{\textup{True}}(T = 1 \mid X > c) = 1$, $\mathbb{P}^{\textup{True}}(T = 1 \mid X < c) = 0$. This mimics assumption (RD) of hahn2001identification.\footnote{gerard2020bounds extend their approach to fuzzy RD, where there are complex shape restrictions.} It is observationally testable and often obvious in applications: if $X$ is the net reported vote share of an election candidate, then the candidate wins if and only if $X > 0$. We use the notation “$X = c$" to mean that we take the limit as $\varepsilon \to 0$ of $X \in [c - \varepsilon, c + \varepsilon]$. Similarly, we let $X = c^+$ denote the limit corresponding to $X \in (c,c+\varepsilon]$ and let $X=c^-$ correspond to $X \in [c-\varepsilon,c)$.

We assume that observations can be partitioned into manipulators and non-manipulators. We assume manipulation is one-sided, and without loss of generality assume manipulators choose treatment ($F_{X \mid M=1}(c) = 0$). An interesting direction for future work are bounds in which manipulation can occur in both directions. As in gerard2020bounds, the probability of treated observations being manipulated, $$\eta = \mathbb{P}^{\textup{True}}( M = 1 \mid X = c^+ ),$$ is identified.\footnote{gerard2020bounds call this quantity $\tau$.} We assume appropriate continuity of potential outcomes and running variables given the manipulation status. We relegate the details to (ref) in the appendix. The assumptions imply that non-manipulator average treatment effects would be identified by the change in non-manipulator outcomes across the cutoff $c$. However, when $\eta > 0$ so that there is manipulation, the distribution of $Y \mid X = c^+$ includes manipulators' treated potential outcomes, so that standard treatment effects are not point-identified.

Estimands that we may be interested in include the conditional average treatment effect (CATE), the conditional local average treatment effect (CLATE), and the conditional average treatment effect on the treated (CATT).

equation[equation omitted — 273 chars of source]

We call $E\left[ Y(1) - Y(0) \mid X = c, M = 0 \right]$ the CLATE because it is an average treatment effect at the cutoff among the population for whom the treatment is randomly assigned at the cutoff, conditional on being at the cutoff.

We write the causal estimands of interest in terms of conditional expectations of \textcolor{black}{observed variables} and \textcolor{black}{potential outcomes} as follows:

lemmaOur main estimands of interest have the following expressions: \begin{align*} \psi_{CATE}& = \frac{1}{2 - \textcolor{black}{\eta}} \textcolor{black}{E\left[ Y \mid X = c^+ \right]} + \frac{1-\textcolor{black}{\eta}}{1-2\textcolor{black}{\eta}} \textcolor{black}{E_{\mathbb{P}^{True}} \left[ Y(1) \mid X = c, M = 0 \right]} \\ & \qquad - \frac{\textcolor{black}{\eta}}{2-\textcolor{black}{\eta}} \textcolor{black}{E_{\mathbb{P}^{True}} \left[ Y(0) \mid X = c, M = 1 \right]} - \left(1-\frac{\textcolor{black}{\eta}}{2-\textcolor{black}{\eta}}\right) \textcolor{black}{E\left[ Y \mid X = c^- \right]} \\ \psi_{CATT} & = \frac{\textcolor{black}{E\left[ (2 T - 1) Y \mid X = c \right]} - \frac{\textcolor{black}{\eta}}{2-\textcolor{black}{\eta}} \textcolor{black}{E_{\mathbb{P}^{True}}[Y(0) \mid X = c, M = 1]}}{\textcolor{black}{\mathbb{P}^{Obs}( X = c^+ \mid X = c)}} \\ \psi_{CLATE} & = \textcolor{black}{E_{\mathbb{P}^{True}}[Y(1) \mid X = c, M = 0]} - \textcolor{black}{E[Y \mid X = c^-]}. \end{align*}

Most quantities in the expressions above are point-identified. The only unidentified objects are the conditional expectations \textcolor{black}{$E_{\mathbb{P}^{\textup{True}}} \left[ Y(1) \mid X = c, M = 0 \right]$} and \textcolor{black}{$E_{\mathbb{P}^{\textup{True}}} \left[ Y(0) \mid X = c, M = 1 \right]$}.

The specific target distribution $\mathbb{P}^{\textup{Target}}$ will depend on the estimand of interest as follows. For sets $\mathcal{Y} \subset \mathds{R}$, define:\footnote{A formal construction would define the appropriate distributions at $X=x$ based on the distribution conditional on $|x-c|$ and then take the limit as $x \to c$ from either side, which would yield well-defined distributions by the maintained continuity assumptions.}

align*[align* omitted — 553 chars of source]

Then the target distribution for estimated $Est \in \{CATE, CLATE, CATT\}$ is defined as $\mathbb{P}^{\textup{Target}}(X, T, Y) \equiv \mathbb{P}^{\textup{Obs}}(X, T) \mathbb{P}^{\textup{Target}}_{Est}(Y \mid X, T)$. The distribution of $R$ is the same under $\mathbb{P}^{\textup{Target}}$ and $\mathbb{P}^{\textup{Obs}}$ and $E_{\mathbb{P}^{\textup{Target}}}[\lambda(R) Y]$ is equal to the target estimand for $\lambda(R)$ we define below.

Suppose we have a structural model that implies there are values $\Lambda_0, \Lambda_1 \geq 1$ such that for $t = 1, 0$:

align[align omitted — 306 chars of source]

Define $\mathcal{M}(\Lambda_0, \Lambda_1)$ as the set of distributions $\mathbb{Q}$ over $(M, X(1), X(0), Y(1), Y(0), T)$ that marginalize to the distribution of $(X(M), Y(T), T = 1\{ X(M) > c\})$ under $\mathbb{P}^{\textup{True}}$, satisfy the restrictions of this section, and satisfy $\frac{\mathbb{Q}(M = 1 \mid Y(t), X=c)}{\mathbb{Q}(M = 1 \mid Y(t), X=c)} / \frac{\mathbb{P}^{\textup{Obs}}(M = 1 \mid X = c)}{\mathbb{P}^{\textup{Obs}}(M = 0 \mid X = c)} \in [\Lambda_t^{-1}, \Lambda_t]$ almost surely.

More complicated functional forms could be accomodated at the cost of additional notation. This functional form is inspired by tan2006distributional's Marginal Sensitivity Model. $\Lambda_t = 1$ corresponds to an unconfoundedness-type case in which the difference in regression values yields the CATE. $\Lambda_t = \infty$ corresponds to gerard2020bounds's assumption, in which the distribution of $Y(0) \mid M = 1, X=c$ is only constrained by the support of the potential outcome. In between, larger values of $\Lambda$ accommodate larger degrees of manipulation on potential outcomes.

propositionIn the RD case, our method can be implemented for the CLATE $E_{\mathbb{P}^{\textup{True}}}[ Y(1) \mid X = c, M = 0] - E_{\mathbb{P}^{\textup{Obs}}}[ Y \mid X = c^-]$ with the following values: \begin{align*} \lambda(R) & = \frac{1\{ X = c^+ \}}{\mathbb{P}^{Obs}(X = c^+ \mid X = c)} - \frac{1\{ X = c^- \}}{\mathbb{P}^{Obs}(X = c^- \mid X = c)} \\ w(R) & = \frac{1\{ X = c^+ \}}{1 - \eta + \eta \Lambda_1} + 1\{ X = c^- \}, \qquad \bar{w}(R) = \frac{1\{ X = c^+ \}}{1 - \eta + \eta \Lambda_1^{-1}} + 1\{ X = c^- \}; \end{align*} our method can be implemented for the CATT $E_{\mathbb{P}^{\textup{True}}}[Y(1) - Y(0) \mid X = c^+]$ with the same $\lambda(R)$ and: \begin{align*} w(R) & = 1\{ X = c^+ \} + 1\{ X = c^- \} \left( (1-\eta) + \eta \Lambda_0^{-1} \right), \qquad \bar{w}(R) = 1\{ X = c^+ \} + 1\{ X = c^- \} \left( (1-\eta) + \eta \Lambda_0 \right); \end{align*} and our method can be implemented for the CATE $E_{\mathbb{P}^{\textup{True}}}[Y(1) - Y(0) \mid X = c]$ with the same $\lambda(R)$ and: \begin{align*} w (R) &= 1\{ X=c^+ \} \left[ \frac{1}{2-\eta} + \frac{1-\eta}{1-2\eta} \frac{1}{1-\eta+\eta\Lambda_1} \right] + 1\{ X=c^- \} \left[ \left( 1 - \frac{\eta}{2-\eta}\right) + \frac{\eta}{2-\eta} \Lambda_0^{-1} \right] \\ \bar{w} (R) &= 1\{ X=c^+ \} \left[ \frac{1}{2-\eta} + \frac{1-\eta}{1-2\eta} \frac{1}{1-\eta+\eta\Lambda_1^{-1}} \right] + 1\{ X=c^- \} \left[ \left( 1 - \frac{\eta}{2-\eta}\right) + \frac{\eta}{2-\eta} \Lambda_0 \right]. \end{align*}

The connection between (ref) and the characterization from (ref) comes from correspondences that we derive in Appendix (ref). Note that even when $\Lambda_1$ is infinite, the CLATE bounds are finite, which allows gerard2020bounds to obtain non-trivial CLATE bounds. Our CLATE bounds for $\Lambda_1=\infty$ are identical to the bounds in gerard2020bounds for sharp RD. When $\Lambda_1$ is finite, we are also able to obtain meaningful CATE bounds.

Our analysis so far bounds manipulation on each potential outcome separately. There turns out to be no additional information available from a structural model that bounds manipulation on both potential outcomes simultaneously.

propositionSuppose there is a finite $\Lambda \geq 1$ such that $\Lambda_1 = \Lambda_0 = \Lambda$. Let $\mathcal{M}'(\Lambda)$ be the set of distributions $\mathbb{Q} \in \mathcal{M}(\infty, \infty)$ satisfying: \begin{align} \left. \frac{\mathbb{Q}( M = 1 \mid Y(1), Y(0), X = c)}{\mathbb{Q}( M = 0 \mid Y(1), Y(0), X = c)} \right/ \frac{\mathbb{P}^{Obs}( M = 1 \mid X = c)}{\mathbb{P}^{Obs}( M = 0 \mid X = c)} \in [\Lambda^{-1}, \Lambda]. \end{align} Then $\mathcal{M}'(\Lambda) \subseteq \mathcal{M}(\Lambda, \Lambda)$. Further, take any distributions $\mathbb{Q}_1, \mathbb{Q}_0 \in \mathcal{M}(\Lambda, \Lambda)$ and write $\psi_t = E_{\mathbb{Q}_t}[Y(t) \mid X = c, M = 1-t]$. Then there is a distribution $\mathbb{Q}' \in \mathcal{M}'(\Lambda)$ satisfying $E_{\mathbb{Q}'}[Y(t) \mid X = c, M = 1-t] = \psi_t$ for $d = 1, 0$.

(ref) shows that the bounds are sharp, i.e., separate bounds on both potential outcomes suffice to bound our causal estimands of interest. The quantities $E_{\mathbb{P}^{\textup{True}}}[ Y(t) \mid X = c, M = 1-t]$ still identify our causal estimands of interest.

The high-level logic of (ref) is fundamentally similar to Section (ref). Any given pair of Radon-Nikodym derivatives may not be simultaneously achievable for both potential outcomes. However, the pair corresponding to the worst-case bounds are simultaneously achievable. As a result, the identified set is the same whether we bound manipulation on one potential outcome or both potential outcomes simultaneously.

Instrumental Variables

We consider bounds on local average treatment effects (LATEs) in instrumental variable models under restrictions on violations of the exclusion restriction. Our previous examples correspond to forms of selection on unobservables. We now show our framework also applies to bounds on local average treatment effects (LATEs) in instrumental variable models under restrictions on violations of the exclusion restriction.

Consider a canonical IV setup. We assume there is a distribution $\mathbb{P}^{\textup{True}}$ over $\left(Z,\{T\left(z\right)\},\left\{ Y\left(t,z\right)\right\} ,X\right)$, where $Z$ is a binary instrument, $Y(t, z)$ is the potential outcome with treatment status $t$ and instrument status $z$, and $T(z)$ is the potential treatment given the instrument $z$ is $T(z)$. However, we only observe the coarsening $\mathbb{P}^{\textup{Obs}}$ over $\left(Z,T=T\left(z\right),Y=Y\left(T,Z\right),X\right)$. We will maintain that the instrument is randomly assigned ($Y(t, z) \rotatebox[origin=c]{90}{$\models$} Z \mid X$) and monotonicity holds ($T(1) \geq T(0)$). Under monotonicty, there are three treatment response groups: always-takers (At, $T(1) = T(0) = 1$), never-takers (Nt, $T(1) = T(0) = 0$), and compliers (Co, $1 = T(1) > T(0) = 0$). The regressors are $R = (T, Z, X)$.

We will not impose the exclusion restriction. If exclusion holds, then $Y\left(t,1\right)=Y\left(t,0\right)$. If exclusion fails, then the standard conditional IV estimand, $(E[Y \mid Z=1, X] - E[Y \mid Z=0, X]) / (E[T \mid Z=1, X] - E[T \mid Z=0, X])$, will incorrectly assign any direct effect of the instrument on outcomes to treatment effects.

We target an instrument-weighted average LATE given by:

align*[align* omitted — 137 chars of source]

where $\eta(X)$ reflects covariate weighting and $\omega(1 \mid X) = 1 - \omega(0 \mid X)$ is a researcher-estimated instrument weighting function in $[0, 1]$. When exclusion holds, this class includes the average complier treatment effect for $\eta(X) = 1 / \mathbb{P}^{\textup{Obs}}(Co)$. When exclusion fails, this specification also allows the researcher to target a particular weighted average of treatment effects across instrument statuses. For example, $\omega(Z \mid X) = 1$ targets the treatment effect at $Z = 1$ and $\omega(Z \mid X) = E[Z \mid X]$ targets the treatment effect at the average observed instrument status.

We rewrite the causal estimand $\psi$ in terms of conditional expected outcomes as follows:

align*[align* omitted — 178 chars of source]

where $\lambda(Z, X, T) = \eta(X) \left( \frac{Z - \mathbb{P}^{\textup{Obs}}(Z = 1 \mid X)}{\mathbb{P}^{\textup{Obs}}(Z = 1 \mid X) \mathbb{P}^{\textup{Obs}}(Z = 0 \mid X)} \right)$. Always-takers and never-takers are included in this statement of $\psi$ for convenience, but their potential outcomes cancel out for $\lambda(R)$ because $E[\lambda(R) \mid T(1), T(0), X] = 0$.

The target distribution $\mathbb{P}^{\textup{Target}}$ is constructed as a marginal distribution. For $v \in \{0, 1\}$ and $\mathcal{Y} \subset \mathds{R}$, define the distribution $\mathbb{P}^{\textup{Target}}$, which re-draws a value $v$ from a Bernoulli($\omega(1 \mid X)$) distribution in order to achieve appropriate outcome weighting, as follows: $$\mathbb{P}^{\textup{Target}}(X, T, Z, Y) \equiv \sum_{v, t_1, t_0 \in \{0, 1\}} \mathbb{P}^{\textup{True}}(X, t, Z, T(1) = t_1, T(0) = t_0) \omega(v \mid X) \mathbb{P}^{\textup{True}}( Y(T, v) \in \mathcal{Y} \mid X, T(1), T(0)).$$ The distribution of $R$ is the same under $\mathbb{P}^{\textup{Target}}$ and $\mathbb{P}^{\textup{Obs}}$. Further, by abadie2003semiparametric's argument, the distribution satisfies $E_{\mathbb{P}^{\textup{Target}}}[\lambda(R) Y] = E_{\mathbb{P}^{\textup{True}}}[ \lambda(R) E_{\mathbb{P}^{\textup{True}}}[ Y(T, Z) \mid X, Co] ] = \psi$, so that $E_{\mathbb{P}^{\textup{Target}}}[\lambda(R) Y]$ is the causal estimand.

Suppose we have a structural model that implies there are functions $\ell_{Nt}(X), u_{Nt}(X)$, $\ell_{At}(X), u_{At}(X)$, $\ell_{Co}^{1}(X), u_{Co}^{1}(X)$, $\ell_{Co}^0(X), u_{Co}^0(X)$ such that:

align*[align* omitted — 618 chars of source]

and suppose further that all four Radon-Nikodym derivatives are finite and strictly positive. Exclusion corresponds to the case $\ell = u = 1$. Worst-case bounds that only restrict the support of the potential outcomes correspond to $\ell = 0$, $u = \infty$. When $Y$ is binary, ramsahai2012causal proposes a sensitivity analysis that places a bound on $\mathbb{P}^{\textup{True}}(Y=1\mid X, T, Z=1, U) - \mathbb{P}^{\textup{True}}(Y=1\mid X, T, Z=0, U)$ for some unobserved $U$, which immediately translates to bounds on always- and never-taker likelihood ratios by taking $U = \{ T(1), T(0) \}$ and implies bounds on complier likelihood ratios. This object can be decomposed as a convolution of several of the probability objects above, so its interpretation is less transparent in a causal framework with potential treatments. Alternatively, $\Lambda$ bounds on the effect of $Z$ on the odds of a binary potential outcome given $X$ and potential treatments imply $\Lambda$ bounds on the odds ratios here, though narrower partially identified regions could be obtained by leveraging estimated regression functions.

As we show in Appendix (ref), these objects further imply bounds on $\frac{d \mathbb{P}^{\textup{True}}( Y(T, Z) \mid Co, X)}{d \mathbb{P}^{\textup{Obs}}( Y \mid T, Z, X)}$ for each value of $T$ and $Z$.

propositionSuppose we write the implied bounds from the worst-case $\ell$ and $u$ applied to the formulas from Appendix (ref) as \begin{align*} w (t,z|X) &\leq \frac{d \mathbb{P}^{Target}( Y \mid X, T=t, Z=z)}{d \mathbb{P}^{Obs}( Y \mid X, T=t, Z=z)} \leq \bar{w}( t,z|X), \end{align*} where we write $\underline{w}(t,z|X) = \bar{w}(t,z|X) = 1$ for values such that $\mathbb{P}^{\textup{Obs}}(T=t, Z=z \mid X) = 0$. Then our method can be implemented on $E_{\mathbb{P}^{\textup{True}}}[ \eta(X) 1\{Co\} \sum_z \omega(z \mid X) ( Y(1, z) - Y(0, z) )]$ with the following values: \begin{align*} \lambda(R) & = \eta(X) \frac{Z - \mathbb{P}^{Obs}(Z = 1 \mid X)}{\mathbb{P}^{Obs}(Z = 1 \mid X) \mathbb{P}^{Obs}(Z = 0 \mid X)}, \quad \underline{w}(R) = \underline{w}(T, Z \mid X), \quad \bar{w}(R) = \bar{w}(T, Z \mid X). \end{align*}

Unlike the previous examples, the characterization in (ref) is conservative.

propositionSuppose the observed distribution follows $X = 1$; $Z \mid X \sim Bern(0.5)$; $T \mid Z, X \sim Bern( Z / 2 )$; and $Y \mid X, Z, T \sim Unif(-1, 1)$. Suppose we are interested in the average complier treatment effect at $z=1$, i.e. $\eta(x) = 2$ and $\omega(z \mid x) = z$. Suppose a structural model implies lower bounds of $\ell(x) = 1$ and $u(x) = \infty$ for all groups. Then the structural model implies the sharp bounds are the singleton $\{0\}$, but the bounds from (ref) are $[-1, 1]$.

Intuitively, the structural model implicitly includes cross-restrictions between the $Co$, $At$, and $Nt$ Radon-Nikodym derivatives. In this case, the $\frac{d \mathbb{P}^{\textup{Target}}(Y \mid T=0, Z=0)}{d \mathbb{P}^{\textup{Obs}}( Y \mid T=0, Z=0)}$ includes a product of Radon-Nikodym derivatives for compliers and never-takers, which must satisfy further constraints that our approach omits for the purpose of ease of use.

We discuss some other applications in (ref).

Implementation

We now provide theoretical guarantees for consistency of a plug-in estimator. We then illustrate the effective performance of our procedure for simulations in the conditional c-dependence environment.

Estimation and Inference

This subsection shows that natural plug-in estimators achieve standard asymptotics under reasonable conditions. We focus on the upper bound for exposition. The target object is: \[ T \equiv E_{\mathbb{P}^{\textup{Obs}}}\left[\lambda(R)Y+\left(\lambda(R)Y-Q_{\tau(R)}\left(\lambda(R)Y|R\right)\right)a(\underline{w}(R),\bar{w}(R),\lambda(R)Y,Q_{\tau(R)}(\lambda(R)Y\mid R))\right], \] where \[ a(\underline{w},\bar{w},\lambda y,q)\equiv\left(\bar{w}-\underline{w}\right)1\left\{ \lambda y>q\right\} -(1-\underline{w}). \] This is a statistical quantity that depends on the observed distribution $\mathbb{P}^{\textup{Obs}}$ alone, so we now suppress the dependence on $\mathbb{P}^{\textup{Obs}}$ for concision.

We use hatted objects to denote the estimated objects and $\hat{E}$ to denote the sample mean. \[ \hat{T}\equiv \hat{E}\left[\hat{\lambda}(R)Y+\left(\hat{\lambda}(R)Y-\hat{Q}_{\hat{\tau}(R)}\left(\hat{\lambda}(R)Y|R\right)\right)a(\hat{\underline{w}}(R),\hat{\bar{w}}(R),\hat{\lambda}(R)Y,\hat{Q}_{\hat{\tau}(R)}(\hat{\lambda}(R)Y\mid R))\right] \]

Consistency follows under natural conditions. Further, we have a form of one-sided robustness.

propositionIf we have iid sampling, finite second moments, and $\hat{\lambda}\xrightarrow{p}\lambda, \hat{Q}\xrightarrow{p}Q, \hat{\bar{w}} \xrightarrow{p} \bar{w}, \hat{\underline{w}} \xrightarrow{p} \underline{w}$, then $\hat{T} \xrightarrow{p} T$. Further, even if $\hat{Q} \xrightarrow{p} \bar{Q} \ne Q$, there exists some $\hat{T}^*$ such that $\hat{T} \geq \hat{T}^*$ and $\hat{T}^* \xrightarrow{p} T$.

The first part of the proposition is immediate by applying the continuous mapping theorem and the law of large numbers for iid observations. The $a(\cdot)$ dependence on an indicator function only introduces a kink point rather than a discontinuity, so the function is still continuous. The assumptions of the proposition are made at a high level, so that we can accommodate various forms of consistent estimators for functions such as $\hat{\bar{w}}$. In particular, we can accommodate machine learning nuisance estimators, as we do in our simulation. Note that our consistency assumption may not be achievable when the quantiles are finite but unbounded, as at one point in our simulation.

The second part of the proposition states a useful robustness guarantee. Even if the quantile is not estimated correctly, the resulting estimated bounds will be valid: too wide for the estimated sensitivity model rather than too narrow. This validity property corresponds to (ref)'s one-sided validity guarantee.

For inference, we use a standard percentile bootstrap in our implementation. Namely,

enumerate• For every $b = 1, \cdots,B $, \begin{enumerate} • Sample $n$ observations from the data iid with replacement to get the bootstrap data. • Calculate $t^{(b)} \equiv \sqrt{n} (\hat{T}^{(b)} - \hat{T})$, where $\hat{T}^{(b)}$ is the estimator that uses the boostrapped data. \end{enumerate} • Let $G_N$ denote the CDF of $t^{(b)}$ and $q_\alpha$ denote that $\alpha$ quantile of $G_N$. For a size $1-\alpha$ confidence interval, use $[\hat{T} - \frac{1}{\sqrt{n}} q_{1-\alpha/2} , \hat{T} - \frac{1}{\sqrt{n}} q_{\alpha/2}] =CI(\alpha)$

The bootstrap will have coverage at least as large as nominal under standard conditions, like smoothness of the $\hat{w}$ estimates, even if the quantile estimator tends to an inconsistent limit. The core argument is that an infeasible boostrap estimator that replaces the estimated $1 \{ \hat{\lambda}^{(b)} Y > \hat{Q}^+ \}$ with the true $1 \{ \lambda Y > Q^+ \}$ in the construction of $a$ would be valid and have weakly more aggressive confidence intervals. As a result, the quantiles do not even need to be reestimated in the bootstraps dorn2022sharp. Further, estimation error in the quantiles exhibit a second-order influence on the estimated bounds, so that the confidence intervals can asymptotically achieve the nominal rate under moderate conditions dvds. A generic proof of bootstrap consistency is outside the scope of this paper. The presence of extreme quantiles or kink points, as in masten2018identification's conditional c-dependence model, may call for more exotic bootstraps and a subtle proof of bootstrap validity.

Simulation

We illustrate the procedure using a simulation in the conditional c-dependence context of Section (ref).

The observed distribution $\mathbb{P}^{\textup{Obs}}$ over $(X, Z, Y$) is $X \sim U[-\eta, \eta]$; $Z \mid X \sim Bern( 1 / (1 + exp(-X)) )$; and $Y \mid X, Z \sim \mathcal{N}((2 + X)(Z - 1), 1)$, where $\eta$ is chosen so that the support of $e(X)$ is $[0.1, 0.9]$.

We consider estimation at $c = 0, 0.01, ..., 0.1$. When $c > 0.1$, the identified set is unbounded. Conversely, we show in Appendix (ref) that for all $c \in [0, 0.1)$, the identified set remains uniformly bounded. Estimation for small $c$ reduces to a fairly standard problem. As $c$ approaches $0.1$, estimation becomes more difficult and the relevant extremal quantiles introduce a “delicate problem" for estimation MastenPoirierZhang. Once $c$ crosses $0.1$, the identified set becomes the real line and (ref) no longer applies. Still, we evaluate simulation performance at this extreme case by testing how often the $c = 0.1$ confidence intervals yield the true infinite bounds. We call this just-barely-infinite case $c = 0.1 + \epsilon$.

The bounds in the general program are estimated by plug-in. We estimate $\lambda$, $\tau$, $\underline{w}$, and $\bar{w}$ by plugging in propensity estimates $\hat{e}(X)$. The propensities are estimated using well-specified logistic regression of $Z$ on $X$. We estimate the quantile function using quantile regression on 99 grid points ($\tau = 0.01, ..., 0.99$) and estimate the extreme quantiles ($\tau = 0, 1$) from the known infinite bounds on the conditional outcome distribution. The quantile regression grid is estimated by regressing $\hat{\lambda} Y$ on $Z$ interacted with both $X$ and $\hat{\lambda}(X)$ and is re-used across values of $c$ and bootstraps. The underlying quantile regression is estimated in two ways: linear quantile regression and random forests with three-fold cross validation. For each observation, we take the quantile regression estimate by linearly interpolating the grid at the estimated $\hat{\tau}$.

Inference proceeds by a standard percentile bootstrap. For a given dataset, we can redraw observations with replacement. For every bootstrap draw $b$, we reestimate propensities via one-step updating, and reestimate the weight bounds accordingly. We then estimate bootstrap upper and lower bounds $\hat{\psi}^+_{b}$ and $\hat{\psi}^-_{b}$. We do not reestimate the quantile regression grid between bootstraps, which introduces a squared outward bias. The 95% confidence interval for the identified set is the set bounded by the 2.5th quantile of the $\hat{\psi}^-_{b}$ draws and the 97.5th quantile of the $\hat{\psi}^+_{b}$ draws. The 95% confidence intervals for the one-sided lower and upper bounds is the set bounded by the 5th quantile of the $\hat{\psi}^-_{b}$ draws or the 95th quantile of the $\hat{\psi}^+_{b}$ draws. The quantiles of the estimated bounds $\{ \psi^-_b\}, \{ \psi^+_b\}$ then form the confidence interval for the identified set.

We compare our estimates and confidence intervals to the true bounds. We obtain a close estimate of the true bounds numerically. In particular, we average the closed-form identified set for $E[Y(1) - Y(0) \mid X]$ over one million draws of $X$. With the true $\psi^+$ and $\psi^-$ essentially known, we can assess the validity of our estimation and inference procedures.

For a given sensitivity parameter $c$, we run 1,000 simulations of the data with 2,000 observations. Within each simulation, we take 1,000 bootstrap draws and calculate the bounds for each bootstrap draw. We also evaluate the coverage rate of the $c = 0.1$ estimates for slightly larger $c$, in which case the true identified set is the reals. We denote the resulting coverage rates in tables as $c = 0.1 + \epsilon$.

figure[figure omitted — 460 chars of source]
figure[figure omitted — 456 chars of source]

We present our mean and median bound estimates using linear and random forest quantiles in Figures (ref) and (ref), respectively. In Figure (ref), our median bound estimates generally track the true bounds. Our mean bound estimates roughly track the identified set until $c$ gets close enough to $0.1$ that some simulations produce infinite estimated bounds (\protect0.2 % of simulations at ). As $c$ gets close to $0.1$ and the most extreme $\tau(R)$ values get close to one, our median estimates become slightly too wide. This phenomenon likely reflects the robustness of our characterization with respect to quantile errors, which are especially likely when applying our discrete grid to extreme $\bar{w}$ values. In the random forest estimates represented in Figure (ref), the estimates are somewhat more conservative, but still track the true bounds reasonably closely. Since we have a linear parametric model, we would expect the parametric quantile regression to be more accurate than a non-parametric method and to produce narrower bounds.

table[table omitted — 1,402 chars of source]

Our coverage results are reported in (ref). Under nominal coverage, the coverage rate would be \protect93.6% to 96.3% with 95% probability. With linear quantile regression, the coverage rate is close to nominal for small and moderately sized values of $c$. There is some over-coverage as $c$ approaches $0.1$. When $c = 0.1$, the identified set rests on a knife edge: as $c$ approaches $0.1$ from below, the lower and upper bounds tend towards $1.5$ and $2.5$, but for any $c$ above $0.1$, the identified set is the full real line. We find that in this case the confidence intervals cover the true (finite) identified set in \protect98.7% of simulations and cover the near (infinite) identified set in \protect77.3% of cases of simulations. (At $c=0.10$, \protect66.4% of bound estimates are unbounded.) The bounds with quantile forest estimates remain valid, but tend to be more conservative. As before, this phenomenon can be attributed to well-specified parametric quantile regression being more accurate and generating narrower bounds that produce coverage closer to the nominal rates.

Conclusion

This paper proposes a novel sensitivity analysis framework for identification failures for linear estimators. By placing bounds on the distributional distance between the observed distribution and a target distribution that identifies the causal parameter of interest, we obtain sharp and tractable analytic bounds. This framework generalizes existing sensitivity models in RD and IPW and motivates a new sensitivity model for IV exclusion failures. We provide new results on sharp and valid sensitivity analysis that allow even unbounded likelihood ratios. We illustrate how our framework and partial identification results contribute to three important applications, including new procedures for sensitivity analysis for the CATE under RD with manipulation and for instrumental variables with exclusion.