Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
71,937 characters · 18 sections · 54 citation commands
A General Approach to Relaxing Unconfoundedness
JEL classification: C14; C18; C21; C25; C51
Keywords: Identification, Treatment Effects, Partial Identification, Sensitivity Analysis, Unconfoundedness
\onehalfspacing
A large literature studies the identification and estimation of treatment effects when a binary treatment is randomly assigned conditional on covariates. This assumption is called unconfoundedness, conditionally independent treatment assignment, or ignorability, among other terms. With observational data it is often considered very strong, however, so a corresponding literature has developed to relax this assumption. These papers use a variety of different classes of relaxations of unconfoundedness. That is, there are different ways of formalizing the idea that treatment is “almost” randomly assigned, given the covariates. This variation raises a question: How do these different relaxations compare to each other? This question is important because empirical researchers are often concerned that the number of robustness checks they must consider is constantly growing; if some of these checks are related, however, then that relationship can potentially be used to simplify the overall analysis. Moreover, mathematically related analyses do not necessarily provide “independent” evidence of robustness, a second motivation for better understanding the relationships between different relaxations of an assumption.
With that aim, this paper makes two main contributions. First, we define a general class of relaxations, which includes several previous approaches as special cases. Second, we derive closed form, analytical identification results for treatment effects under this general class of relaxations. This paper therefore unifies several disparate identification results in the literature. In doing so, we also provide a variety of new identification results, because we study an extensive list of parameters, including quantile treatment effects (QTEs) and the distribution of treatment effects (DTEs), whereas most existing papers focus solely on average-type treatment effects. These new results were previously unknown even for the specific types of relaxations that have been considered before. We give a precise discussion of how our results compare to the previous literature in the next subsection.
In section (ref) we set up the baseline treatment effects model and define the target parameters we study. We define our general class of relaxations at the start of section (ref). We show how this class relates to previous relaxations in sections (ref) and (ref). In section (ref) we derive general analytical identification results for marginal cdfs of potential outcomes and monotonic functionals of those cdfs. We apply those results in section (ref) to obtain analytical bounds on various treatment effect parameters. We conclude in section (ref).
\nocite{Rosenbaum1995, Rosenbaum2002, Rosenbaum2017}
A vast literature studies unconfoundedness; we do not attempt a comprehensive review here. Instead we discuss the most closely related prior work. Nonparametric relaxations of unconfoundedness were pioneered by Paul Rosenbaum's work (see his 2002 or 2017 books for a survey, for example). His work focuses on sensitivity analysis within the context of finite sample randomization inference (c.f., chapter 5 of ImbensRubin2015). Much of the subsequent literature has instead focused on large population level identification analysis. In particular, inspired by Rosenbaum's approach, Tan2006 proposed the marginal sensitivity model (MSM), a specific nonparametric relaxation of unconfoundedness (which we review in section (ref)). Given this relaxation, Tan showed that bounds on parameters of interest can be characterized as the solutions to optimization problems with infinitely many constraints, but did not provide any formal results, proofs, or closed form expressions for these bounds. ZhaoSmallBhattacharya2019 derived non-sharp bounds on the average potential outcome $\ensuremath{\mathbb{E}}(Y_x)$ and the average treatment effect (ATE) under the MSM, but also did not derive closed form expressions for these bounds. DornGuo2023 strengthened that result by deriving sharp bounds on $\ensuremath{\mathbb{E}}(Y_x)$, ATE, and the average effect of treatment on the treated (ATT) under the MSM, but again without closed form expressions. DornGuoKallus2024 subsequently refined that result by obtaining closed form expressions for sharp bounds on $\ensuremath{\mathbb{E}}(Y_x)$ and ATE under the MSM, in addition to developing the concept of double-validity and double sharpness. Tan2024 gives alternative sharp bound expressions for the ATE under the MSM. KallusZhou2018 studied policy learning under the MSM, which is related to identification of the average weighted welfare (what they call the “policy value”), but they do not derive population bounds on this parameter.
This existing literature on the MSM largely focuses on average potential outcomes $\ensuremath{\mathbb{E}}(Y_x)$ or the ATE. Our paper provides the first sharp bounds on a wide variety of target parameters under the MSM, including the quantile treatment effect (QTE), the quantile treatment effect on the treated (QTT), the distribution of treatment effects (DTE), and the average weighted welfare (AWW). Moreover, for many of the parameters we study, our bounds are closed form. The existence of closed form expressions simplifies the construction of estimation and inference procedures, and also allows us to analytically examine how the bounds depend on the distribution of the observed data, and thus which features of the data lead results to be robust.
MastenPoirier2016,MastenPoirier2018 proposed an alternative relaxation of unconfoundedness called conditional $c$-dependence, and derived closed form sharp bounds on a variety of treatment effect parameters under this relaxation, including $\ensuremath{\mathbb{E}}(Y_x)$, ATE, ATT, the QTE, and the DTE (in MastenPoirier2019BF). In the current paper we extend these identification results to a class of parameters that also includes the average weighted welfare (AWW), weighted average treatment effects, and to quantiles of the distribution of conditional average treatment effects (QCATE). Those earlier papers also restricted attention to continuous or binary outcomes whereas our new results apply for any distribution of the outcome, including mixed continous-discrete distributions. We also show how the conditional $c$-dependence relaxation is related to the marginal sensitivity model.
While our general class of relaxations includes several previously proposed relaxations of unconfoundedness, there are alternative relaxations where it is not yet clear if they can be accommodated by our class. This includes BonviniKennedy2022, who derive closed-form sharp bounds on ATE under a mixture-style relaxation, and HuangPimentel2025, who derive closed-form non-sharp bounds on ATT under an assumption about how much unobserved variables can affect the variance of odds ratios similar to those that arise in the MSM; in appendix A.6 they derive sharp but non-closed form bounds on the ATT under the same relaxation. It also includes DingVanderWeele2016 and VanderWeeleDing2017, who derive closed-form non-sharp bounds on the causal relative risk under assumptions about relative risks involving latent confounders; Sjolander2024 derives closed-form sharp bounds under the same relaxation.
There are several related papers that provide general methods for deriving bounds. DornYap2024 show how to derive analytical expressions for sharp bounds on parameters that can be written as certain weighted averages of outcomes under a restriction on a generalized likelihood ratio. Like us, they show that their class of relaxations includes several previous relaxations (such as the MSM of Tan2006 and conditional $c$-dependence of MastenPoirier2018). Whereas we only study relaxations of unconfoundedness, they also show how to use their results to do sensitivity analysis for instrumental variables and regression discontinuity designs. Their analysis of unconfoundedness, however, focuses on average potential outcomes and ATE, whereas we also study parameters like the QTE and DTE. A large literature in econometrics has studied how to derive identified sets for a variety of parameters under a variety of assumptions when all observed variables are discretely distributed; see, for example, Torgovitsky2018 and GuRussellStringham2024, and the references therein. Duarte2024 uses similar ideas to numerically compute identified sets for a variety of sensitivity analyses when all variables are discretely distributed. RambachanCostonKennedy2023 provide a variety of general sensitivity analyses for binary outcomes. In contrast to this literature, we obtain analytical sharp bounds without any restriction on the distribution of the outcome variable, which allows the outcome to be continuously distributed or even mixed continuous-discretely distributed.
Several prior papers also discuss the relationship between various relaxations of unconfoundedness. MastenPoirier2023EJ discuss mean independence, quantile independence assumptions, and a weaker version of quantile independence that they call $\mathcal{U}$-independence. Their focus is on interpreting relaxations in terms of treatment selection models, rather than providing identification results for a broad class of relaxations. ZhaoSmallBhattacharya2019 discuss the relationship between the MSM and Rosenbaum's sensitivity model. For binary outcomes, RambachanCostonKennedy2023 relate their relaxation to the MSM, Rosenbaum's sensitivity model, conditional $c$-dependence, and an approach called Tukey's factorization.
For random a variable $A$ and a random vector $B$, we let $F_{A \mid B}(a \mid b) \coloneqq \ensuremath{\mathbb{P}}(A \leq a \mid B=b)$ denote the conditional cdf. For $\tau \in (0,1)$, we let $Q_{A \mid B}(\tau \mid b) \coloneqq \inf\{a \in \ensuremath{\mathbb{R}} : F_{A \mid B}(a \mid b) \geq \tau\}$ denote the left-inverse of this cdf, that is, its conditional quantile function.
We are interested in the causal impact of a binary treatment variable $X \in \{0,1\}$ on an outcome variable $Y$. Let $(Y_1,Y_0)$ be potential outcomes under treatment and no treatment respectively. Denote the realized outcome by \[ Y = X Y_1 + (1-X) Y_0. \] Let $W$ be a vector of covariates with support $\operatorname*{supp}(W) \subseteq \ensuremath{\mathbb{R}}^{d_W}$. We use $p_{x|w}$ to denote $\ensuremath{\mathbb{P}}(X=x \mid W=w)$. $p_{1|w}$ thus denotes the propensity score. We assume realizations of $(Y,X,W)$ are observed by the researcher. Our identification analysis abstracts from sampling uncertainty and assumes the joint distribution of $(Y,X,W)$ is known.
Throughout the paper we maintain the following assumption, which is a strict overlap assumption. It is also sometimes called strict positivity.
With observational data, a commonly imposed assumption is unconfoundedness. It is also called selection on observables, ignorability, or conditional independence. This assumption states that potential outcomes are independent of treatment given covariates $W$. This conditional independence is either imposed jointly on both potential outcomes, or on each potential outcome separately:
Under Assumption (ref) and either version of unconfoundedness given in equation (ref), it is well known that the pair of distribution functions $(F_{Y_1 \mid W},F_{Y_0 \mid W})$ is point-identified via \[ F_{Y_x \mid W}(y \mid w) = \ensuremath{\mathbb{P}}(Y \leq y \mid X=x, W=w) \] for $x \in \{0,1\}$. This point-identification implies that many parameters that summarize aspects of the distribution of $(Y_1,Y_0,X,W)$ are point-identified. Specifically, parameters that can be expressed as functionals of $F_{Y_1 \mid W}$, $F_{Y_0 \mid W}$, and the known distribution of observables $(Y,X,W)$, are point-identified. For example, we can point-identify the Conditional Average Treatment Effect (CATE) as defined by $\text{CATE}(w) \coloneqq \ensuremath{\mathbb{E}}[Y_1 - Y_0 \mid W=w]$ because it can be written as
However, parameters that depend on other aspects of the distribution of potential outcomes may only be partially identified. For example, consider the distribution function of $Y_1 - Y_0$, the unit-level treatment effect:
This parameter depends on the structure of the dependence between the two potential outcomes, which is unknown from either version of the unconfoundedness assumption. As discussed in FanPark2010, $F_{Y_1 - Y_0}(z)$ is partially identified and sharp bounds can be recovered in terms of the distribution of $(Y,X,W)$.
To help classify treatment effect parameters, consider the decomposition
where $C_{1,0 \mid X,W}(\cdot,\cdot \mid x,w)$ is a copula that characterizes the dependence between $Y_1$ and $Y_0$ conditional on $(X,W) = (x,w)$. By Sklar's Theorem (Sklar1959), such a copula exists. Given that $F_{Y,X,W}$ is known, we consider treatment effect parameters that can be written as a function of $(F_{Y_1 \mid X,W},F_{Y_0 \mid X,W},C_{1,0 \mid X,W},F_{Y,X,W})$. We denote these parameters through the functional
and we give several examples below. The dependence of $\theta$ on some of its arguments is suppressed if the functional is constant with respect to them. We will sometimes denote these parameters as functionals of $(F_{Y_1 \mid W},F_{Y_0 \mid W},F_{Y,X,W})$ rather than $(F_{Y_1 \mid X,W},F_{Y_0 \mid X,W},F_{Y,X,W})$. These two formulations are equivalent due to the relationship
which holds for all $(y,x,w) \in \ensuremath{\mathbb{R}} \times \{0,1\}\times \operatorname*{supp}(W)$.
Next we consider eleven example target parameters. Our results give sharp bounds for all eleven parameters under our general class of assumptions, including the marginal sensitivity model as a special case. For many of these parameters we also obtain analytical, closed form expressions for the bound functions.
The parameters in examples (ref)--(ref) only depend on the distribution of potential outcomes through their marginal distributions given $(X,W)$, while the parameters in examples (ref)--(ref) also depend on their copulas. Under overlap and unconfoundedness, the parameters in (ref)--(ref) are all point-identified. The parameters (ref)--(ref) are partially identified under overlap and unconfoundedness since the conditional copulas $C_{1,0 \mid X,W}$ are not identified from the joint distribution of $(Y,X,W)$. In other words, these parameters depend on the type of dependence between $Y_1$ and $Y_0$, and unconfoundedness does not reveal any information about this dependence. For example, FanPark2010 show the identified set for $F_{Y_1 - Y_0}(z)$, the DTE, is an interval and they provide a closed-form expression for its lower and upper bounds.
However, if unconfoundedness fails, all these parameters will be partially identified. The identified set for parameters that are partially identified under unconfoundedness becomes larger under failures of unconfoundedness.
We now consider relaxations of the unconfoundedness assumption. We will consider two related, general relaxations of unconfoundedness that encompass several disparate relaxations that were studied in the literature. We begin by considering a class of assumptions on the probabilities of treatment when conditioning on covariates $W$ and one of the potential outcomes.
This assumption restricts the manner in which potential outcomes affect the treatment probability $p_x(y,w)$, which we call a latent propensity score. We will use $c$-dependence assumptions to conduct sensitivity analyses for unconfoundedness. Here the sensitivity parameters are $\underline{c}(w,\eta)$ and $\overline{c}(w,\eta)$, which we refer to as bound functions. Like the notation in RambachanCostonKennedy2023, we let $\eta$ be a possibly infinite-dimensional nuisance parameter that is point-identified from the distribution $F_{Y,X,W}$. The bound functions are also allowed to depend on the covariate value $w$. In principle, we can also allow the bounds to differ across $x \in \{ 0,1 \}$, but we do not include an $x$ subscript for simplicity. The specification of $\underline{c}(w,\eta)$ and $\overline{c}(w,\eta)$ is left implicit, which allows them to be functions of low-dimensional or scalar sensitivity parameters. For example, the marginal sensitivity model of Tan2006, which depends on a single sensitivity parameter, can be viewed as a special case of marginal $c$-dependence. We show this in section (ref).
We can also see that setting $\underline{c}(w,\eta) = \overline{c}(w,\eta) = p_{1|w}$ yields unconfoundedness as a special case of marginal $c$-dependence, while letting $(\underline{c}(w,\eta),\overline{c}(w,\eta))$ approach $(0,1)$ implies that no restrictions on the dependence between $X$ and $Y_x$ (given covariates) are imposed. Note that we restrict the propensity score $p_{1|w}$ to lie within our specified bounds for $p_x(Y_x,w)$. If the propensity score were outside these bounds, then the assumption would be misspecified because, by the law of iterated expectations, $p_{1|w} = \ensuremath{\mathbb{E}}[\ensuremath{\mathbb{P}}(X=1 \mid Y_x,W = w) \mid W = w] \in [\underline{c}(w,\eta),\overline{c}(w,\eta)]$.
We also consider a closely related class of assumptions that restricts the probability of treatment given both potential outcomes.
Joint $c$-dependence with bound functions $\underline{c}(w,\eta)$ and $\overline{c}(w,\eta)$ implies marginal $c$-dependence with the same bound functions. This is due to the law of iterated expectations.
We next show that several unconfoundedness relaxations from recent related literature can be viewed as special cases of either marginal or joint $c$-dependence.
Tan2006 proposed the Marginal Sensitivity Model (MSM), which restricts the odds ratio between propensity scores and treatment probabilities that also condition on the potential outcome $Y_x$, for $x = 0,1$. For $x \in \{0,1\}$, let
denote this odds ratio. When $Y_x$ is continuously distributed with respect to the Lebesgue measure, this ratio can also be expressed as ratios of conditional densities of $Y_x \mid X=1,W$ and $Y_x \mid X=0,W$.
Tan2006's MSM posits known bounds for these odds ratios.
$\Lambda$ is a scalar sensitivity parameter. Setting $\Lambda = 1$ is equivalent to assuming unconfoundedness, and increasing $\Lambda$ allows for more dependence of latent propensity scores on potential outcomes. Variants of Tan2006's MSM whose odds ratios condition on both potential outcomes have also been considered. Similarly, these odds ratios may instead condition on an abstract confounder $U$ rather than potential outcomes. See, for example, DornGuo2023 and DornGuoKallus2024 for recent examples of these two variants. To distinguish it from the case where one conditions on the potential outcomes one at a time, we call the version that conditions on both potential outcomes the Joint Sensitivity Model (JSM).
We now state generalizations of the MSM and JSM that allow their odds ratios to have arbitrary bounds, as opposed to bounds that have product equal to 1. We also allow their bounds to depend on covariates or nuisance parameters. We will continue distinguishing between marginal sensitivity models, which condition on one potential outcome at a time, and joint sensitivity models, which condition on both potential outcomes.
We can see that the MSM is a special case of the GMSM by setting $[\underline{\Lambda}(w,\eta),\overline{\Lambda}(w,\eta)] = [\Lambda^{-1},\Lambda]$. Similarly, the JSM is a special case of the GJSM.
The GMSM is equivalent to marginal $c$-dependence because, for each bound function pair $[\underline{c}(w,\eta),\overline{c}(w,\eta)]$ under marginal $c$-dependence, there exists exactly one corresponding bound function pair $[\underline{\Lambda}(w,\eta), \overline{\Lambda}(w,\eta)]$ under the GMSM. The same link exists between joint $c$-dependence and the GJSM. We show this in the following proposition.
This proposition shows that the marginal $c$-dependence is equivalent to the generalized marginal sensitivity model. Similarly, joint $c$-dependence is equivalent to the generalized marginal sensitivity model.
MastenPoirier2018 studied a relaxation of unconfoundedness they called conditional $c$-dependence, which assumed symmetric bounds on the latent propensity score.
This is a special case of marginal $c$-dependence where the bounds equal
Here the nuisance parameter is $p_{1|(\cdot)}$, the propensity score function. Unconfoundedness is obtained by setting $c = 0$, while the no-assumption bounds are obtained for $c$'s equal to or larger than $\sup_{w \in \operatorname*{supp}(W)} \max\{p_{1|w},p_{0|w}\}$. MastenPoirier2018 provided closed-form expressions for sharp bounds on the CQTE, CATE, ATE, QTE, and ATT when potential outcomes are continuously distributed or binary. MastenPoirierZhang2020 describe flexible parametric estimators of these bounds and provide nonstandard inference methods.
Next we derive sharp bounds on a large class of target parameters under the relaxations described in section (ref). We will study a class of parameters that includes all eleven examples in section (ref) as special cases. Specifically, we will compute these parameters' sharp bounds, or their identified set, under marginal and joint $c$-dependence, which are equivalent to the GMSM and GJSM, respectively.
Before studying our general class of treatment effect parameters, we first consider bounds on the distribution functions of each potential outcome, given covariates. These cdfs are building blocks for these parameters and, as we will see, analytical bounds on these cdfs will directly map into analytical bounds on these parameters.
Specifically, we begin by analyzing the conditional cdf of the potential outcome $Y_x$ given the covariate value $w$, $F_{Y_x \mid W}(y \mid w) \coloneqq \ensuremath{\mathbb{P}}(Y_x \leq y \mid W=w)$. Under either marginal or joint $c$-dependence, we can show that this cdf is bounded above and below by two cdfs which form an envelope for $F_{Y_x \mid W}(y \mid w)$ for all values of $(y,w) \in \ensuremath{\mathbb{R}} \times \operatorname*{supp}(W)$.
Define the following functions:
Viewed as functions of $y$, these four functions are cdfs since they are nondecreasing, right-continuous, and their limits as $y \rightarrow -\infty,+\infty$ equal 0 and 1, respectively. We show these four cdfs form bounds for $F_{Y_x \mid W}$ under marginal or joint $c$-dependence.
We note a few properties of these bounds. The bounds for $F_{Y_x \mid W}(y \mid w)$ collapse to a point if either $\underline{c}(w,\eta) = p_{1|w}$ or $\overline{c}(w,\eta) = p_{1|w}$. Also note that $F_{Y \mid X,W}(y \mid x,w)$ always lies within the bounds for $F_{Y_x \mid W}(y \mid w)$. This is because $c$-dependence never rules out unconfoundedness, and unconfoundedness implies that the distribution of $Y$ given $(X,W) = (x,w)$ equals that of $Y_x$ given $W = w$.
These cdf bounds also yield cdf bounds under the GMSM or GJSM since they are equivalent to $c$-dependence as shown in Proposition (ref). These bounds are also valid for the standard MSM or JSM, as they are special cases of marginal or joint $c$-dependence.
We now show these cdf bounds are sharp, or that they cannot be improved upon. This is the case under marginal or joint $c$-dependence. The cdf of $Y_x \mid W=w$ can also lie in the interior of these bounds, as we show that any convex linear combination of the upper and lower cdf bounds can be attained.
Before establishing this, let $\mathcal{C}$ denote the set of all bivariate copulas and let
denote the collection of all bivariate copulas across all treatment and covariate values $(x,w) \in \{0,1\} \times \operatorname*{supp}(W)$. We also say that a distribution function for $(Y_1,Y_0) \mid X,W$ is compatible with the observed distribution $F_{Y,X,W}$ if
for all $w\in\operatorname*{supp}(W)$.
We now define the identified set for the distribution of $(Y_1,Y_0) \mid X,W$ from the observable distribution $F_{Y,X,W}$ under $c$-dependence. This set consists of all conditional cdfs and copulas that imply a distribution for $(Y_1,Y_0) \mid X,W$ that is both compatible with the data distribution $F_{Y,X,W}$ and with a $c$-dependence condition.
In our later derivations, we sometimes refer to the identified set for $(F_{Y_1 \mid W},F_{Y_0 \mid W},C_{1,0 \mid X,W})$ instead, which we denote by $\mathcal{I}_0^{i}(F_{Y,X,W};c)$ for $i \in \{\text{marg}, \text{joint} \}$. Via equation (ref), this set can be viewed as an affine transformation of the identified set for $(F_{Y_1 \mid X,W}, F_{Y_0 \mid X,W},C_{1,0 \mid X,W})$.
We now show some key properties of the cdfs and copulas in these identified sets. We begin with marginal $c$-dependence.
This theorem shows that the four pairs of cdfs $(\underline{F}_{Y_1 \mid W},\underline{F}_{Y_0 \mid W})$, $(\overline{F}_{Y_1 \mid W},\underline{F}_{Y_0 \mid W})$, $(\underline{F}_{Y_1 \mid W},\overline{F}_{Y_0 \mid W})$, and $(\overline{F}_{Y_1 \mid W},\overline{F}_{Y_0 \mid W})$ are part of the identified set. This is obtained by varying $(\varepsilon,\gamma)$ over $\{(1,1),(0,1),(1,0),(0,0)\}$. We show this by explicitly constructing latent propensity scores $p_1(Y_1,w)$ and $p_0(Y_0,w)$ that lie in $[\underline{c}(w,\eta),\overline{c}(w,\eta)]$ almost surely under the implied distribution of $(Y_1,Y_0) \mid W = w$. These propensity scores have a switching structure where they equal the lower/upper bound for low values of $Y_x$ and the upper/lower bound for large values of $Y_x$. For example, the propensity score $p_1(Y_1,w)$ associated with cdf upper bound $\overline{F}_{Y_1 \mid W}$ equals
where
and
Note that $\overline{A}_1 \in [\underline{c}(w,\eta),\overline{c}(w,\eta)]$. We denote $\lim_{q \nearrow \overline{Q}_1} \overline{F}_{Y_1 \mid W}(q \mid w)$ by $\overline{F}_{Y_1 \mid W}(\overline{Q}_1- \mid w)$. The propensity scores associated with the cdf bounds $\underline{F}_{Y_1 \mid W}$, $\overline{F}_{Y_0 \mid W}$, and $\underline{F}_{Y_0 \mid W}$ can all be found in Appendix (ref).
This switching structure of the latent propensity score was observed for conditional $c$-dependence by MastenPoirier2018, and for the MSM in Proposition 2 of DornGuo2023. Our sharpness proof implies that latent propensity score $p_1(Y_1,w)$ satisfies an integral constraint, namely that $\ensuremath{\mathbb{E}}[p_1(Y_1,W) \mid W = w]= \ensuremath{\mathbb{P}}(X=1 \mid W = w)$, in order to ensure it is compatible with the observed propensity score.
Theorem (ref) also shows that any convex linear combinations of these four cdf pairs lies in the identified set. As a result, the identified set for $F_{Y_x \mid W}(y \mid w)$ is the entire closed interval $[\underline{F}_{Y_x \mid W}(y \mid w),\overline{F}_{Y_x \mid W}(y \mid w)]$. Moreover, the identified set for the pair $(F_{Y_1 \mid W}(y \mid w),F_{Y_0 \mid W}(y \mid w))$ is the Cartesian product of their individual identified sets, meaning that fixing or knowing the conditional distribution of one potential outcome does not affect the identified set of the distribution of the other potential outcome.
Finally, Theorem (ref) proves that no conditional copulas are ruled out by marginal $c$-dependence. For example, marginal $c$-dependence allows $Y_1$ and $Y_0$ to be independent, comonotonic\footnote{This is also referred to as rank invariance. For example, see the discussion in HeckmanSmithClements1997.}, or countermonotonic given $X$ and $W$.
A similar result is obtained under joint $c$-dependence.
The only difference between the two theorems concerns the dependence structures between $Y_1$ and $Y_0$. Theorem (ref) shows that all copulas are compatible with marginal $c$-dependence, while our proof of Theorem (ref) only exhibits one copula for each pair of conditional distributions $(\varepsilon \underline{F}_{Y_1 \mid W} + (1-\varepsilon) \overline{F}_{Y_1 \mid W},\gamma \underline{F}_{Y_0 \mid W} + (1-\gamma) \overline{F}_{Y_0 \mid W})$.
The sharp bounds provided in theorems (ref) and (ref) can be used to deliver analytical expressions for sharp bounds on a large class of treatment effect parameters. We first define the identified set for a parameter $\theta$ defined in (ref).
These sets are the parameter values consistent with the known distribution of observables $F_{Y,X,W}$ and with a $c$-dependence condition. Without restrictions on how $\theta$ depends on the distribution of potential outcomes, these sets may take various shapes.
We focus on a class of scalar estimands that can be ordered with respect to first order stochastic dominance.
Next we define our target class of parameters.
Following Manski1997a, monotonic parameters are also called $D$-parameters, or $D_1$-parameters. Also see Manski2003 or Stoye2010 who consider parameters that are increasing with respect to second-order stochastic dominance.
As an example, consider a parameter $\theta(F_{Y_1 \mid W})$ that is increasing in $F_{Y_1 \mid W}$ and suppose $c$-dependence holds. Then
since $\overline{F}_{Y_1 \mid W} \preceq F_{Y_1 \mid W} \preceq \underline{F}_{Y_1 \mid W}$, which holds by Lemma (ref). This interval cannot be made narrower since the cdf bounds $[\underline{F}_{Y_1 \mid W},\overline{F}_{Y_1 \mid W}]$ are sharp by theorems (ref) and (ref). Therefore, the identified set for $\theta(F_{Y_1 \mid W})$ is a subset of this closed interval that always contains its two endpoints. The interior of this interval is also part of the identified set if the functional $\theta$ is continuous in the sense that $\varepsilon \mapsto \theta(\varepsilon \underline{F}_{Y_1 \mid W} + (1-\varepsilon) \overline{F}_{Y_1 \mid W})$ is continuous. This type of continuity is implied by the continuity of the mapping $F \mapsto \theta(F)$ under the sup-distance metric.
Assuming monotonicity of the parameter will help derive properties of its identified set. Monotonicity is a substantive restriction, but all eleven parameters from Section (ref) satisfy it. This is formally established in Lemma (ref) below. We begin by considering monotonic parameters that do not depend on copulas.
This theorem shows that substituting the upper/lower cdf bounds delivers sharp bounds for any parameter that is monotonic in the first-order stochastic dominance sense. The result is derived under an assumption that the parameter is increasing in $F_{Y_1 \mid W}$ and decreasing in $F_{Y_0 \mid W}$, but it immediately generalizes to parameters that are increasing or decreasing in either or both conditional cdfs. For example, the cdf pair $(\overline{F}_{Y_1 \mid W},\underline{F}_{Y_0 \mid W})$ will maximize a parameter that is decreasing in $F_{Y_1 \mid W}$ and increasing in $F_{Y_0 \mid W}$, and the cdf pair $(\overline{F}_{Y_1 \mid W},\overline{F}_{Y_0 \mid W})$ will maximize (minimize) a parameter that is decreasing (increasing) in both $F_{Y_1 \mid W}$ and $F_{Y_0 \mid W}$. The identified set for these parameters always contains endpoints $\theta(\overline{F}_{Y_1 \mid W},\underline{F}_{Y_0 \mid W},F_{Y,X,W})$ and $\theta(\underline{F}_{Y_1 \mid W},\overline{F}_{Y_0 \mid W},F_{Y,X,W})$. It also contains all the values between these endpoints whenever the mapping $\theta$ is continuous in the appropriate sense.
We document the monotonicity of various building blocks for parameters of interest in the following technical lemma. We omit covariates $W$ for simplicity here, except in part 4 on QCATE because that parameter requires covariates to be nontrivial.
Using this lemma, all eight parameters that are invariant to copulas are bounded by substituting the upper or lower cdf bounds from Theorem (ref). This allows us to compute analytical bounds for these parameters.
We explore these analytical bounds by focusing on five of our examples to illustrate these expressions. The first three parameters are independent from the copula, while the last two are copula-dependent.
From Lemma (ref).1, we have that the ATE satisfies
This interval equals the identified set by the monotonicity and continuity of the expectation functional which was established in Lemma (ref).1. The lower and upper bounds can be obtained by calculating $\int y \, d\overline{F}_{Y_x \mid W}(y \mid w)$ and $\int y \, d\underline{F}_{Y_x \mid W}(y \mid w)$ for $x \in \{0,1\}$. Via the quantile transformation, these bounds can also be written as integrals of $\underline{Q}_{Y_x \mid W}(u \mid w)$ and $\overline{Q}_{Y_x \mid W}(u \mid w)$ over $u \in (0,1)$. Thus the ATE bounds can be written as integrals of quantiles. Via Lemma (ref) in Appendix (ref), we show that these quantile integrals can be converted into conditional expectations of outcomes given that they exceed or fall short of a fixed conditional quantile. These are equivalent to Conditional Value at Risk (CVaR) measures that appear in DornGuoKallus2024. These bounds are stated explicitly in equations (ref)--(ref) in Appendix (ref) in the general case. When $Y \mid X,W$ is continuously distributed, we obtain simpler expressions for these bounds that we give here:
and
Note that the dependence of $(\underline{c},\overline{c})$ on $(W,\eta)$ was suppressed for convenience.
We now consider bounds on the quantile treatment effect for a fixed quantile $\tau \in (0,1)$. By Lemma (ref).2, the functional $\theta_\text{QTE}$ is increasing in $F_{Y_1 \mid W}$ and decreasing in $F_{Y_0 \mid W}$. Therefore, by Theorem (ref), $\text{QTE}(\tau)$ has the following sharp bounds:
where $\overline{Q}_{Y_x}$ is the left-inverse of cdf $\underline{F}_{Y_x}(\cdot) \coloneqq \ensuremath{\mathbb{E}}[\underline{F}_{Y_x \mid W}(\cdot \mid W)]$ for $x \in \{0,1\}$. Analogously, $\underline{Q}_{Y_x}$ is the left-inverse of cdf $\overline{F}_{Y_x}(\cdot) \coloneqq \ensuremath{\mathbb{E}}[\overline{F}_{Y_x \mid W}(\cdot \mid W)]$. Analytical expressions for the unconditional cdf bounds for the treated potential outcome are given by
and similar expressions can be obtained for $Y_0$. The left-inverses of the previous expressions yield bounds on quantiles of $Y_1$ and $Y_0$, and which can be used to compute the QTE bounds.
Consider a policy $\omega: \operatorname*{supp}(W) \to [0,1]$ that treats units with covariate value $w$ with probability $\omega(w)$. The average welfare in a population under such policy is given by
By adapting Lemma (ref).1, this functional is increasing in $F_{Y_1 \mid W}$, increasing in $F_{Y_0 \mid W}$, and continuous in the sense defined in the lemma. Therefore, by Theorem (ref), its identified set is the closed interval given by
An analytical expression for these bounds can be obtained by substituting in the expressions for the cdf bounds in the previous functionals. When $Y \mid X,W$ is continuously distributed, the bounds are given by
and
We now consider identification of the parameters in examples (ref)--(ref) which all depend on the copulas $C_{1,0 \mid X,W}$. Even under unconfoundedness these parameters are not point-identified. Relaxing unconfoundedness will yield larger identified sets for these parameters when compared to the unconfoundedness baseline. We will focus on marginal $c$-dependence since it does not restrict the dependence structure between the potential outcomes.
Consider identification of the joint cdf $F_{Y_1,Y_0}(y_1,y_0)$ under marginal $c$-dependence. Consider the functional
Fix the conditional copula function $C_{1,0 \mid X,W}(\cdot, \cdot \mid \cdot, \cdot)$. Then by Lemma (ref).4, this functional is decreasing in $F_{Y_0 \mid X,W}(y_0 \mid 1,w)$ and $F_{Y_1 \mid X,W}(y_1 \mid 0,w)$. Thus it is bounded below by \[ \theta_\text{CDF}(\underline{F}_{Y_1 \mid X,W},\underline{F}_{Y_0 \mid X,W},C_{1,0 \mid X,W},F_{Y,X,W};y_1,y_0) \] and above by \[ \theta_\text{CDF}(\overline{F}_{Y_1 \mid X,W},\overline{F}_{Y_0 \mid X,W},C_{1,0 \mid X,W},F_{Y,X,W};y_1,y_0). \] Moreover, by Theorem (ref), these bounds are sharp.
Since $C_{1,0 \mid X,W}$ is unknown, we then compute the maximum and minimum of these bounds over the set of copulas that are consistent with marginal $c$-dependence; this is simply the set of all copulas. The Fr\'echet-Hoeffding bounds show that all copulas $C$ satisfy
for all $(u,v) \in [0,1]^2$. The copula bounds $\underline{C}$ and $\overline{C}$ are themselves copulas. Combining these facts, we obtain the following analytical bounds on the joint cdf of potential outcomes.
The bounds in (ref) are themselves cdfs, so these bounds can be attained simultaneously for all $(y_1,y_0) \in \ensuremath{\mathbb{R}}^2$. The bounds for $F_{Y_1,Y_0}$ under unconfoundedness are obtained as a special case when $\underline{c} = p_{1 \mid W} = \overline{c}$, which implies that $\overline{F}_{Y_x \mid X,W} = \underline{F}_{Y_x \mid X,W} = F_{Y \mid X,W}(\cdot \mid x,\cdot)$. Making this substitution in equation (ref) yields these bounds under unconfoundedness.
Identification of this parameter under unconfoundedness was studied in FanPark2010, by applying results first shown in Makarov1982 and later studied in WilliamsonDowns1990. MastenPoirier2019BF also studied this parameter under conditional $c$-dependence and under a range of assumptions on copulas for $(Y_1,Y_0)$. By Lemma 2.1 in FanPark2010, the cdf of $Y_1 - Y_0$ given $(X,W) = (x,w)$ satisfies
and these bounds are sharp for any pair of cdfs $(F_{Y_1 \mid X,W},F_{Y_0 \mid X,W})$. These bounds are decreasing in $F_{Y_1 \mid X,W}$ and increasing in $F_{Y_0 \mid X,W}$, therefore substituting the upper/lower cdf bounds for $F_{Y_x \mid X,W}$ results in sharp bounds for $F_{Y_1 - Y_0 \mid X,W}$ under $c$-dependence. This was established in Lemma (ref).6. Integrating these bounds over the marginal distribution of $(X,W)$ yields sharp bounds for the unconditional cdf of $Y_1 - Y_0$. This result is summarized in the following proposition.
This expression involves two one-dimensional optimization problems, but the objective functions are known, closed-form functionals of the distribution of the observables. Bounds on the QDTE can be obtained as a corollary by taking the left-inverse of the cdf bounds.
In this paper we proposed a general class of relaxations of unconfoundedness, and showed how it includes several previous approaches as special cases. We then derived closed form identification results for many different target parameters under this general class of relaxations. There are at least three natural next steps. First, in this paper we focused on population level identification results. Corresponding estimation and inference results can likely be derived by using standard sample analog estimators and arguments, but we leave the details to future work. Second, it would be interesting to explore whether our bounds have either the double-sharpness or double-validity properties defined in DornGuoKallus2024, and if not, whether alternative bounds that had these properties could be derived. Third, it would be interesting to extend our results to independence assumptions beyond unconfoundedness, such as IV exogeneity (e.g., section 4 of MastenPoirier2018a).