Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
49,776 characters · 16 sections · 38 citation commands
Identification of Treatment Effects under Conditional Partial Independence
JEL classification: C14; C18; C21; C51
Keywords: Treatment Effects, Conditional Independence, Unconfoundedness, Selection on Observables, Sensitivity Analysis, Nonparametric Identification, Partial Identification
\onehalfspacing
The treatment effect model under conditional independence is widely used in empirical research. The conditional independence assumption states that, after conditioning on a set of observed covariates, treatment assignment is independent of potential outcomes. This assumption has many other names, including unconfoundedness, ignorability, exogenous selection, and selection on observables. It delivers point identification of many parameters of interest, including the average treatment effect, the average effect of treatment on the treated, and quantile treatment effects. ImbensRubin2015 provide a recent overview of this literature.
Without additional data, like the availability of an instrument, the conditional independence assumption is not refutable: The data alone cannot tell us whether the assumption is true. Moreover, conditional independence is often considered a strong and controversial assumption. Consequently, empirical researchers may wonder: How credible are treatment effect estimates obtained under conditional independence?
In this paper, we address this concern by studying what can be learned about treatment effects under a nonparametric class of assumptions that are weaker than conditional independence. While there are many ways to weaken independence, we focus on just one, which we call conditional $c$-dependence.\footnote{See MastenPoirier2016 for an analysis and discussion of several other approaches. We refer to any of these approaches as partial independence assumptions.} This assumption states that the probability of being treated given observed covariates and an unobserved potential outcome is not too far from the probability of being treated given just the observed covariates. We use the sup-norm distance, where the scalar $c$ denotes how much these two conditional probabilities may differ. This class of assumptions nests both the conditional independence assumption and the opposite end of no constraints on treatment selection.
In our first main contribution, we derive sharp bounds on conditional cdfs that are consistent with conditional $c$-dependence. This result can be used in many models, including the treatment effects model we study here.\footnote{See MastenPoirier2016 for several other applications of this result.} In that model, as our second main contribution, we derive identified sets for many parameters of interest. These include the average treatment effect, the average effect of treatment on the treated, and quantile treatment effects. These identified sets have simple, analytical characterizations. Empirical researchers can use these identified sets to examine how sensitive their parameter estimates are to deviations from the baseline assumption of conditional independence. We illustrate this sensitivity analysis in a brief numerical example.
In the rest of this section, we review the related literature. As discussed in section 22.4 of ImbensRubin2015, a large literature starting with the seminal work of RosenbaumRubin1983sensitivity relaxes conditional independence by modeling the conditional probabilities of treatment assignment given both observable and unobservable variables parametrically. This literature also typically imposes a parametric model on outcomes. This includes LinPsatyKronmal1998, Imbens2003, and AltonjiElderTaber2005,AltonjiElderTaber2008. An important exception is RobinsRotnitzkyScharfstein2000, who relax parametric assumptions on outcomes. They continue to use parametric models for treatment assignment probabilities, however, when applying their results. Our work builds on this literature by developing fully nonparametric methods for sensitivity analysis. Our new methods can ensure that empirical findings of robustness do not rely on auxiliary parametric assumptions.
We are aware of only two previous analyses in this sensitivity analysis literature which develop fully nonparametric methods. The first is IchinoMealliNannicini2008, who avoid specifying a parametric model by assuming that all observed and unobserved variables are discretely distributed, so that their joint distribution is determined by a finite dimensional vector. Unlike our approach, theirs rules out continuous outcomes. It also involves many different sensitivity parameters, while our approach uses only one sensitivity parameter.
The second is Rosenbaum1995,Rosenbaum2002, who proposes a sensitivity analysis within the context of randomization inference for testing the sharp null hypotheses of no unit level treatment effects for all units in one's dataset. Since this approach is based on finite sample randomization inference (c.f., chapter 5 of ImbensRubin2015), rather than population level identification analysis, this is quite different from the approaches discussed above and from what we do in the present paper. Like our results, however, his approach does not impose a parametric model on treatment assignment probabilities.
A large literature initiated by Manski has studied identification problems under various assumptions which typically do not point identify the parameters (e.g., Manski2007). In the context of missing data analysis, Manski2016 suggested imposing a class of assumptions which includes conditional $c$-dependence. He did not, however, derive any identified sets under this assumption. Several papers study partial identification of treatment response under deviations from mean independence assumptions, rather than the statistical independence assumption we start from. ManskiPepper2000,ManskiPepper2009 relax mean independence to a monotonicity constraint in the conditioning variable, while HotzMullinSanders1997 suppose mean independence only holds for some portion of the population. These relaxations and conditional $c$-dependence are non-nested. Moreover, these papers focus on mean potential outcomes, while we also study quantiles and distribution functions. Finally, Manski's original no assumptions bounds for average treatment effects (Manski1989,Manski1990) are obtained as a special case of our conditional $c$-dependence ATE bounds when $c$ is sufficiently large.
We study the standard potential outcomes model with a binary treatment. In this section we setup the notation and some maintained assumptions. We define our parameters of interest and state the key assumption which point identifies them: random assignment of treatment, conditional on covariates. We discuss how we relax this assumption. We derive identified sets under these relaxations in section (ref). We conclude this section by suggesting a few ways to interpret our deviations from conditional independence.
Let $Y$ be an observed scalar outcome variable and $X \in \{ 0,1\}$ an observed binary treatment. Let $Y_1$ and $Y_0$ denote the unobserved potential outcomes. As usual the observed outcome is related to the potential outcomes via the equation
Let $W \in \supp(W)$ denote a vector of observed covariates, which may be discrete, continuous, or mixed. Let $p_{x \mid w} = \Prob(X=x \mid W=w)$ denote the observed generalized propensity score (Imbens2000). We consider both continuous and binary potential outcomes. We begin with the continuous outcome case, where we maintain the following assumption on the joint distribution of $(Y_1,Y_0,X,W)$.
\setcounter{partialIndepSection}{1}
By equation (ref), \[ F_{Y \mid X,W}(y \mid x,w) = \Prob(Y_x \leq y \mid X=x,W=w) \] and hence A(ref).(ref) implies that the distribution function of $Y \mid X=x,W=w$ is also strictly increasing and continuous. By the law of iterated expectations, the marginal distributions of $Y$ and $Y_x$ have the same properties as well. We consider the binary outcome case on page (ref).
A(ref).(ref) states that the conditional support of $Y_x$ given $X=x',W=w$ does not depend on $x'$, and that this support is a possibly infinite closed interval. The first equality is a `support independence' assumption, which is implied by the standard conditional independence assumption. Since $Y \mid X=x,W=w$ has the same distribution as $Y_x \mid X=x,W=w$, this implies that the support of $Y \mid X=x,W=w$ equals that of $Y_x \mid W=w$. Consequently, the endpoints $\underline{y}_x(w)$ and $\overline{y}_x(w)$ are point identified. A(ref).(ref) is a standard overlap assumption.
Define the conditional rank random variables $R_1 = F_{Y_1 \mid W}(Y_1 \mid W)$ and $R_0 = F_{Y_0 \mid W}(Y_0 \mid W)$. For any $w \in \supp(W)$, $R_1 \mid W=w$ and $R_0 \mid W=w$ are uniformly distributed on $[0,1]$, since $F_{Y_1 \mid W}(\cdot \mid w)$ and $F_{Y_0 \mid W}(\cdot \mid w)$ are strictly increasing. Moreover, by construction, both $R_1$ and $R_0$ are independent of $W$. The value of unit $i$'s conditional rank random variable $R_x$ tells us where unit $i$ lies in the conditional distribution of $Y_x \mid W=w$. We occasionally use these variables throughout the paper.
It is well known that the conditional distributions of potential outcomes $Y_1 \mid W$ and $Y_0 \mid W$ and therefore the marginal distributions of $Y_1$ and $Y_0$ are point identified under the following assumption:
These marginal distributions are immediately point identified from \[ F_{Y_x \mid W}(y \mid w) = F_{Y \mid X,W}(y \mid x,w) \] and \[ F_{Y_x}(y) = \int_{\supp(W)} F_{Y \mid X,W}(y \mid x,w) \; dF_W(w). \] Consequently, any functional of $F_{Y_1 \mid W}$ and $F_{Y_0 \mid W}$ is also point identified under the conditional independence assumption. Leading examples include the average treatment effect, \[ \text{ATE} = \Exp(Y_1 - Y_0), \] and the $\tau$-th quantile treatment effect, \[ \text{QTE}(\tau) = Q_{Y_1}(\tau) - Q_{Y_0}(\tau), \] where $\tau \in (0,1)$. The goal of our identification analysis is to study what can be said about such functionals when conditional independence partially fails. To do this we define the following class of assumptions, which we call conditional $c$-dependence.
Under the conditional independence assumption $X \independent Y_x \mid W$, \[ \P(X=1 \mid Y_x = y_x,W=w) = \Prob(X=1 \mid W=w) \] for all $y_x \in \supp(Y_x \mid W=w)$ and all $w \in \supp(W)$. Conditional $c$-dependence allows for deviations from this assumption by allowing the conditional probability $\Prob(X=1 \mid Y_x = y_x, W=w)$ to be different from the propensity score $p_{1 \mid w}$, but not too different. This class of assumptions nests conditional independence as the special case where $c=0$. Moreover, when $c \geq \max \{ p_{1 \mid w}, p_{0 \mid w} \}$, from (ref) we see that conditional $c$-dependence imposes no constraints on $\Prob(X=1 \mid Y_x=y_x, W=w)$. Values of $c$ strictly between zero and $\max\{p_{1 \mid w},p_{0 \mid w}\}$ lead to intermediate cases. These intermediate cases can be thought of as a kind of limited selection on unobservables, since the value of one's unobserved potential outcome $Y_x$ is allowed to affect the probability of receiving treatment.
Beginning with RosenbaumRubin1983sensitivity, many papers use parametric models for unobserved conditional probabilities similar to $\P(X=1 \mid Y_x = y_x,W=w)$ to model deviations from conditional independence. For example, see RobinsRotnitzkyScharfstein2000 and Imbens2003. In contrast, conditional $c$-dependence is a nonparametric class of assumptions. Our results therefore ensure that empirical findings of robustness do not depend on any auxiliary parametric assumptions.
By invertibility of $F_{Y_x \mid W}(\cdot \mid w)$ for each $x \in \{0,1\}$ and $w \in \supp(W)$ (assumption A(ref).(ref)), equation (ref) is equivalent to \[ \sup_{r_x \in [0,1]} | \Prob(X = 1 \mid R_x = r_x,W=w) - \Prob(X=1 \mid W=w) | \leq c. \tag{\ref{eq:c-indep1}$^\prime$} \] Using this result, we obtain the following characterization of conditional $c$-dependence.
Proposition (ref) shows that conditional $c$-dependence is equivalent to a constraint on the sup-norm distance between the joint pdf of $(X,R_x) \mid W$ and the product of the marginal distributions of $X \mid W$ and $R_x \mid W$. Although we do not pursue this here, this alternative characterization also suggests how to extend this concept to continuous treatments. Finally, we note that another equivalent characterization of conditional $c$-dependence obtains by replacing $X=1$ with $X=0$ in equation (ref).
Throughout the rest of the paper we impose conditional $c$-dependence between $X$ and the potential outcomes given covariates:
Interpreting the deviations from one's baseline assumption is an important part of any sensitivity analysis. In this subsection we give several suggestions for how to interpret our sensitivity parameter $c$ in practice.
Our first suggestion, going back to the earliest sensitivity analysis of CornfieldEtAl1959, and used more recently in Imbens2003, AltonjiElderTaber2005,AltonjiElderTaber2008, and Oster2016, is to use the amount of selection on observables to calibrate our beliefs about the amount of selection on unobservables. To formalize this idea in our context, recall that conditional $c$-dependence is defined using a distance between two conditional treatment probabilities: the usual propensity score $\Prob(X=1 \mid W=w)$ and that same conditional probability, except also conditional on an unobserved potential outcome $Y_x$. Hence the question is: How much does adding this extra conditioning variable affect the conditional treatment probability?
In the data, $Y_x$ is unobserved, so we cannot answer this question directly. But we can examine the impact of adding additional observed covariates on conditional treatment probabilities (assuming $K \equiv \dim(W) \geq 1$). Specifically, suppose we partition our vector of covariates $W$ into $(W_{-k},W_k)$ where $W_k$ is the $k$th component of $W$ and $W_{-k}$ is a vector of the remaining $K-$1 components. Define \[ \overline{c}_k = \sup_{w_{-k}} \, \sup_{w_k} | \Prob(X=1 \mid W = (w_{-k},w_k) ) - \Prob(X=1 \mid W_{-k}=w_{-k}) | \] where we take suprema over $w_k \in \supp(W_k \mid W_{-k} = w_{-k})$ and $w_{-k} \in \supp(W_{-k})$. $\overline{c}_k$ tells us the largest amount that the observed conditional treatment probabilities with and without the variable $W_k$ can differ.\footnote{If $K=1$ then one can compare $\Prob(X=1 \mid W=w)$ with $\Prob(X=1)$.} Less formally, it is a measure of the marginal impact of including the $k$th variable on treatment assignment, given that we have already included the vector $W_{-k}$. Similarly to CornfieldEtAl1959 and the subsequent literature, the idea is that if adding an extra observed variable creates variation $\overline{c}_k$, then it might be reasonable to expect that adding the unobserved variable $Y_x$ to our conditioning set may also create variation $\overline{c}_k$. In practice, one can compute and examine $\overline{c}_k$ for each $k$.
Our next suggestion is a variation which incorporates information on the distribution of $W$. Define \[ p_{1 \mid W}(w_{-k},w_k) = \Prob(X=1 \mid W=(w_{-k},w_k) ) \] and \[ p_{1 \mid W_{-k}}(w_{-k}) = \Prob(X=1 \mid W_{-k}=w_{-k}). \] Rather than examining the largest point in the support of the random variable \[ | p_{1 \mid W}(W_{-k},W_k) - p_{1 \mid W_{-k}}(W_{-k}) | \] we could also consider quantiles of this distribution, such as the 50th, 75th, or 90th percentiles. One could also plot the distribution of this random variable for each $k$.
These suggestions are a kind of nonparametric version of the implicit partial $R^2$'s used by Imbens2003 in his parametric model, or of the logit coefficients used by RosenbaumRubin1983sensitivity. The overall idea is the same: We are trying to measure the partial effect of adding an extra conditioning covariate on the conditional probability of treatment.
Our final suggestion reiterates a point made by Rosenbaum (Rosenbaum2002rejoinder, section 7): Precise quantitative interpretations of sensitivity parameters like $c$ are not always necessary. We can perform qualitative comparisons of robustness across different studies and datasets by comparing the corresponding bound functions, as we do in our numerical illustration on page (ref). Imbens2003 made a similar point, stating that “not \ldots\ all evaluations are equally sensitive to departures from the exogeneity assumption” (page 126). Such rankings of studies in terms of their robustness may help one aggregate findings across different studies. We leave a formal study of this kind of robustness-adjusted meta-analysis to future work.
\endinput
In this section, we study identification of treatment effects under conditional $c$-dependence. To do so, we start by deriving bounds on cdfs under generic $c$-dependence. We then apply these results to obtain sharp bounds on various treatment effect functionals.
In this subsection we consider the relationship between a generic scalar random variable $U$ and a binary variable $X \in \{ 0,1 \}$. We derive sharp bounds on the conditional cdf of $U$ given $X$ when (1) the marginal distributions of $U$ and $X$ are known and (2) $X$ is $c$-dependent with $U$, meaning that
In the next subsection we will condition on $W$ and apply this general result with $U = R_x$, the conditional rank variable to obtain sharp bounds on various treatment effect parameters.
Let $F_{U \mid X}(u \mid x) = \Prob(U\leq u \mid X=x)$ denote the unknown conditional cdf of $U$ given $X=x$. Let $F_U(u) = \Prob(U \leq u)$ denote the known marginal cdf of $U$. Let $p_x = \Prob(X=x)$ denote the known marginal probability mass function of $X$.
Define \[ \overline{F}_{U \mid X}^c(u \mid x) = \min\left\{F_U(u) + \frac{c}{p_x}\min\{F_U(u),1-F_U(u)\}, \, \frac{F_U(u)}{p_x}, \, 1\right\} \] and \[ \underline{F}_{U \mid X}^c(u \mid x) = \max\left\{F_U(u) - \frac{c}{p_x} \min\{F_U(u),1-F_U(u)\}, \, \frac{F_U(u)-1}{p_x} +1, \, 0 \right\}. \]
The proof of theorem (ref), along with all other proofs, is given in the appendix. In this proof, we note that there are two constraints on the conditional distribution of $U \mid X$. The first is the $c$-dependence constraint. The second is the fact that the marginal distributions of $U$ and $X$ are known, and hence the conditional cdfs must satisfy a law of total probability constraint. This result is therefore a variation on the decomposition of mixtures problem. See CrossManski2002, Manski2007 chapter 5, and MolinariPeski2006 for further discussion.
Theorem 1 has three conclusions. First, we show that the functions $\overline{F}_{U \mid X}^c(\cdot \mid x)$ and $\underline{F}_{U \mid X}^c(\cdot \mid x)$ bound the unknown conditional cdf $F_{U \mid X}(\cdot \mid x)$ uniformly in their arguments. Second, we show that these bounds are functionally sharp in the sense that the joint identified set for the two conditional cdfs $(F_{U \mid X}(\cdot \mid 1),F_{U \mid X}(\cdot \mid 0))$ contains linear combinations of the bound functions $\overline{F}_{U \mid X}^c(\cdot \mid x)$ and $\underline{F}_{U \mid X}^c(\cdot \mid x)$. Finally, we remark that this functional sharpness implies pointwise sharpness.
Importantly, the bound functions $\overline{F}_{U \mid X}^c(\cdot \mid x)$ and $\underline{F}_{U \mid X}^c(\cdot \mid x)$ are piecewise linear functions with simple analytical expressions. These bounds are proper cdfs and can be attained, as stated above. As $c$ approaches zero, these bounds for $F_{U \mid X}(u \mid x)$ collapse to the conditional cdf $F_U(u)$. When $c$ exceeds $\max\{p_{0},p_{1}\}$, the $c$-dependence constraint is not binding. Consequently, the cdf bounds simplify to \[ \overline{F}_{U \mid X}^c(u \mid x) = \min\left\{\frac{F_U(u)}{p_{x}}, \, 1 \right\} \qquad \text{and} \qquad \underline{F}_{U \mid X}^c(u \mid x) = \max\left\{\frac{F_U(u)-1}{p_{x}} + 1, \, 0\right\}. \] These bounds can be interpreted as the no assumptions bounds since the only constraint imposed on the cdfs is that they satisfy the law of total probability.
Figure (ref) shows several examples of the bound functions $\overline{F}_{U \mid X}^c(\cdot \mid x)$ and $\underline{F}_{U \mid X}^c(\cdot \mid x)$. In this example we let $U \sim \text{Unif}[0,1]$ and $p_1 = 0.75$. We let $c= 0.1$, $0.4$, and $0.9$, which represent three qualitative regions for $c$: (a) $c < \min \{ p_1, p_0 \}$, where both bounds are strictly between 0 and 1 on the interior of the support, (b) $\min \{ p_1, p_0 \} \leq c < \max \{ p_1, p_0 \}$, where one bound is strictly between 0 and 1 on the interior of the support, but the other is not, and (c) $c \geq \max \{ p_1, p_0 \}$, where we simply obtain the no assumptions bounds.
Next we study identification of various treatment effects under conditional $c$-dependence. Throughout most of this section we focus on continuous outcomes. We study binary outcomes on page (ref). We end with a numerical illustration of the identified set for ATE as a function of $c$, and discuss how this set depends on features of the observed distribution of the data.
Under conditional independence, $c=0$, the marginal conditional distribution functions $F_{Y_0 \mid W}$ and $F_{Y_1 \mid W}$ are point identified. For $c > 0$, these functions are partially identified. In this case, we derive sharp bounds on these cdfs in proposition (ref) below.
Define
and
for all $y\in [\underline{y}_x(w),\overline{y}_x(w))$. For $y < \underline{y}_x(w)$, define these cdf bounds to be zero. For $y \geq \overline{y}_x(w)$, define these cdf bounds to be one.
The following proposition shows that the functions (ref) and (ref) are sharp bounds on the cdf of $Y_x \mid W=w$ under conditional $c$-dependence.
The proof of this result relies importantly on our general result, theorem (ref). Similarly to that result, proposition (ref) has three conclusions. First we show that the functions (ref) and (ref) bound $F_{Y_x \mid W}(\cdot \mid w)$ uniformly in their arguments. Second, we show that these bounds are functionally sharp. As in theorem (ref), sharpness is subtle because $\mathcal{F}_{Y_x \mid w}^c$ is not the identified set for $F_{Y_x \mid W}(\cdot \mid w)$---it contains some cdfs which cannot be attained. For example, it contains cdfs with jump discontinuities on their support, which violates A(ref).(ref). We could impose extra constraints to obtain the sharp set of cdfs but this is not required for our analysis.
That said, the bound functions (ref) and (ref) used to define $\mathcal{F}_{Y_x \mid W}^c$ are sharp for the function $F_{Y_x \mid W}(\cdot \mid w)$ in the sense that there are cdfs $F_{Y_x \mid W}(\cdot \mid w; \epsilon, \eta)$ which are (a) attainable, (b) can be made arbitrarily close to the bound functions, and (c) continuously vary between the lower and upper bounds. The bound functions themselves are not always continuous and so violate A(ref).(ref). This explains the presence of the $\eta$ variable. It also explains why the endpoints of the bounds in our results below may not be attainable. If $c$ is small enough, the bound functions can be attainable, but we do not enumerate these cases for simplicity.
We use attainability of the functions $F_{Y_x \mid W}(\cdot \mid w; \epsilon, \eta)$ to prove sharpness of identified sets for various functionals of $F_{Y_1 \mid W}$ and $F_{Y_0 \mid W}$. For example, as the third conclusion to proposition (ref), we already stated the pointwise-in-$y$ sharpness result for the evaluation functional. In general, we obtain bounds for functionals by evaluating the functional at the bounds (ref) and (ref). Sharpness of these bounds then follows by applying proposition (ref).
Next we derive identified sets for functionals of the marginal distribution of potential outcomes given covariates. We begin with the conditional quantile treatment effect: \[ \text{CQTE}(\tau \mid w) = Q_{Y_1 \mid W}(\tau \mid w) - Q_{Y_0 \mid W}(\tau \mid w). \] By integrating these bounds over $\tau$ from zero to one, we will obtain sharp bounds for the conditional average treatment effect: \[ \text{CATE}(w) = \Exp(Y_1 \mid W=w) - \Exp(Y_0 \mid W=w). \] Finally, averaging these bounds over the marginal distribution of $W$ yields sharp bounds on ATE.
We first give closed form expressions for bounds on the quantile function of the potential outcome $Y_x$ given $W=w$. Define
and
The following proposition and corollary formalize these results.
The bounds (ref) are also sharp for the function $\text{CQTE}(\cdot \mid \cdot)$ in a sense similar to that used in theorem (ref) and proposition (ref); we omit the formal statement for brevity. This functional sharpness delivers the following result.
All of these bounds are defined directly from equations (ref) and (ref), or averages of those equations. Those equations have simple analytical expressions, which makes all of these bounds quite tractable. These bounds are all monotonic in $c$, as illustrated in figure (ref) of our numerical example. In particular, as $c$ goes to zero, the CQTE bounds collapse to the point $Q_{Y \mid X,W}(\tau \mid 1,w) - Q_{Y \mid X,W}(\tau \mid 0,w)$ while the $\text{CATE}(w)$ bounds collapse to the point $\Exp(Y \mid X=1, W=w) - \Exp(Y \mid X=0,W=w)$ and the ATE bounds collapse to $\Exp[ \Exp(Y \mid X=1, W) - \Exp(Y \mid X=0,W) ]$.
We can also derive bounds on the unconditional $\text{QTE}(\tau)$ by first deriving bounds on the unconditional cdfs of $Y_1$ and $Y_0$ and inverting them. These unconditional cdf bounds obtain by integrating the conditional bounds of proposition (ref) over $w$. Inverses here denote the left inverse.
As with our earlier results, all three results in this corollary are also functionally sharp. And again all these bounds collapse to a single point as $c$ approaches zero.
\endinput
The average effect of treatment on the treated is \[ \text{ATT} = \Exp(Y_1 - Y_0 \mid X=1). \] Under conditional independence, $\text{ATT} = \Exp[ \text{CATE}(W) \mid X=1 ]$. That is, we average CATE over the distribution of covariates $W$ within the treated group, whereas ATE is the unconditional average of CATE over the covariates, $\text{ATE} = \Exp[ \text{CATE}(W) ]$. In corollary (ref) we showed that the bounds on ATE under conditional $c$-dependence are simply the average of our CATE bounds over the marginal distribution of $W$. Hence a natural first approach for obtaining bounds on ATT is to average our CATE bounds over the distribution of $W \mid X=1$, just as we do in the baseline case of $c=0$. For $c > 0$, however, this approach is not correct. This follows since, when conditional independence fails, potential outcomes are not independent of treatment assignment, even conditional on covariates. That is, for $x \in \{0,1\}$, the distribution of $Y_x \mid X=1,W=w$ is not necessarily the same as the distribution of $Y_x \mid X=0,W=w$. Hence the parameters $\Exp(Y_1-Y_0 \mid W=w)$ and $\Exp(Y_1 - Y_0 \mid X=1,W=w)$ are not necessarily equal. Below, we derive the correct identified set for ATT under conditional $c$-dependence.
As in the baseline case of conditional independence, equation (ref) immediately implies that $\Exp(Y_1 \mid X=1)$ is point identified by $\Exp(Y \mid X=1)$ without any assumptions on the dependence between $Y_1$ and $X$. Hence only $\Exp(Y_0 \mid X=1)$ will be partially identified under deviations from conditional independence. Thus we relax A(ref).(ref) and A(ref) as follows.
By the law of iterated expectations and some algebra, \[ \Exp(Y_0 \mid X=1) = \frac{\Exp(Y_0) - p_0 \Exp(Y \mid X=0)}{p_1}. \] Hence bounds on the conditional mean can be obtained from bounds on the unconditional mean $\Exp(Y_0)$. Let \[ \underline{E}_0^c(w) = \int_0^1 \underline{Q}_{Y_0}^c(\tau \mid w) \; d\tau \qquad \text{and} \qquad \overline{E}_0^c(w) = \int_0^1 \overline{Q}_{Y_0}^c(\tau \mid w) \; d\tau \] denote bounds on $\Exp(Y_0 \mid W=w)$. Averaging these over the marginal distribution of $W$ yields bounds on $\Exp(Y_0)$, denoted by \[ \underline{E}_0^c = \Exp \big( \underline{E}_0^c(W) \big) \qquad \text{and} \qquad \overline{E}_0^c = \Exp \big( \underline{E}_0^c(W) \big). \]
Bounds for the average effect of treatment on the untreated, $\text{ATU} = \Exp(Y_1 - Y_0 \mid X=0)$, can be obtained similarly.
Here we drop the continuity assumption A(ref).(ref) and instead consider binary potential outcomes $Y_x$. We replace the support assumption A(ref).(ref) by the following.
This assumption is equivalent to $\Prob(Y_x = 1 \mid X=x', W=w) \in (0,1)$ for all $x,x' \in \{ 0,1 \}$ and $w \in \supp(W)$. Let $p_{1 \mid x,w} = \Prob(Y=1 \mid X=x,W=w)$. Define \[ \overline{P}^c_x(1 \mid w) = \min\left\{\frac{p_{1 \mid x,w}p_{x \mid w}}{p_{x \mid w} - c}\indicator(p_{x \mid w}>c)+\indicator(p_{x \mid w}\leq c), \, p_{1 \mid x,w}p_{x \mid w} + (1-p_{x \mid w})\right\} \] and \[ \underline{P}^c_x(1 \mid w) = \frac{ p_{1 \mid x,w}p_{x \mid w} }{\min\{p_{x \mid w} + c,1\}}. \]
Bounds for $\Prob(Y_x=0 \mid W=w)$ obtain immediately by taking complements. Averaging over the marginal distribution of $W$ yields \[ \Prob(Y_x=1) \in \left[ \Exp \left( \underline{P}_x^c(1 \mid W) \right), \, \Exp \left( \overline{P}_x^c(1 \mid W) \right) \right] \] with a sharp interior.
Bounds for average treatment effects $\Prob(Y_1=1) - \Prob(Y_0=1)$ can be obtained by combining the bounds for each separate probability $\Prob(Y_x=1)$, $x \in \{ 0,1 \}$, similarly to equation (ref).
\endinput
We conclude this section with a brief numerical illustration. For $x=0,1$ and $w=0,1$, suppose the density of $Y \mid X=x,W=w$ is \[ f_{Y \mid X,W}(y \mid x,w) = \frac{1}{\gamma_X x + \gamma_W w + \sigma} \phi_{[-4,4]} \left( \frac{y - (\pi_X x + \pi_W w)}{\gamma_X x + \gamma_W w + \sigma} \right) \] where $\phi_{[-4,4]}$ is the pdf for the truncated standard normal on $[-4,4]$. $X$ and $W$ are binary with \[ \Prob(X=1) = p_1, \quad \Prob(W=1) = q, \quad \text{and} \quad \Prob(X=1 \mid W=w) = p_{1 \mid w} \] for $w=0,1$. We let $(\pi_X,\pi_W) = (1,1)$, $(\gamma_X,\gamma_W) = (0.1,0.1)$, $p_1 = 0.5$, and $q = 0.5$ in all dgps. We specify the choice of $p_{1 \mid w}$ and $\sigma$ below.
Under the conditional independence assumption, this dgp implies that treatment effects are heterogeneous, with an average treatment effect of $\text{ATE} = \pi_X = 1$. To examine the sensitivity of this finding to partial failure of conditional independence, figure (ref) shows identified sets for ATE under conditional $c$-dependence for $c$ from zero to one. First consider the solid lines, which are the same in both plots. These correspond to the dgp with $(p_{1 \mid 1},p_{1 \mid 0}) = (0.6, 0.4)$ and $\sigma = 0.965$. We see that ATE under conditional independence is positive, and that this conclusion is robust to deviations of up to about $c=0.26$ from independence, but not to larger deviations.
Next we vary the dgp parameters to examine how our identified sets depend on features of the distribution of $(Y,X,W)$. Imbens2003 performed similar dgp comparisons for his method using empirical datasets. In the left plot we change the observed propensity score $p_{1 \mid w}$ while holding all other parameters fixed. Relative to the solid lines, if we increase the variation in the observed propensity score by setting $(p_{1 \mid 1}, p_{1 \mid 0}) = (0.9,0.1)$ then the bounds widen for most values of $c$, as shown by the dashed lines. In particular, the conclusion that ATE is positive now only holds for $c$'s less than about $0.085$. Conversely, if we eliminate the variation in the observed propensity score by setting $(p_{1 \mid 1}, p_{1 \mid 0}) = (0.5, 0.5)$ then the bounds shrink for most values of $c$, as shown by the dotted lines. The conclusion that ATE is positive now holds for slightly more values of $c$ than under the baseline dgp used for the solid lines.
Next consider the right plot. Here we change $R^2$ in the regression of $Y$ on $(1,X,W)$ while holding the observed propensity score fixed at $(p_{1 \mid 1},p_{1 \mid 0}) = (0.6,0.4)$. We vary the value of $R^2$ by varying $\sigma$. The solid lines have $R^2$ equal to 30% (since $\sigma = 0.965$). Relative to these lines, if we decrease $R^2$ to 15%, then the bounds widen for all values of $c$, as shown by the dashed lines. The conclusion that ATE is positive becomes less robust. Conversely, if we increase $R^2$ to 60%, then the bounds shrink for all values of $c$, as shown by the dotted lines. The conclusion that ATE is positive becomes more robust.
The shape of the bounds depends on other features of the distribution of $(Y,X,W)$ as well. For example, if $\pi_X$ increases then all the identified sets shift upward. Hence, holding all else fixed, a larger ATE implies that the sign of ATE will be point identified under weaker independence assumptions. Similar analyses can also be done with other parameters of interest, like $\text{QTE}(\tau)$ for various values of $\tau$. Here we merely illustrate the kinds of objects empirical researchers can compute using the results we develop in this paper.
In this paper we studied conditional $c$-dependence, a nonparametric approach for weakening conditional independence assumptions. We used this concept to study identification of treatment effects when the conditional independence assumption partially fails, but no further data---like observations of an instrument---are available. We derived identified sets under conditional $c$-dependence for many parameters of interest, including average treatment effects and quantile treatment effects. These identified sets have simple, analytical characterizations. These analytical identified sets lend themselves to sample analog estimation and inference via the existing literature on inference under partial identification (see CanayShaikh2017 for a survey). Our identification results can be used to analyze the sensitivity of one's results to the conditional independence assumption, without relying on auxiliary parametric assumptions.
Several questions remain. First, we focused on identification of $D$-parameters (Manski2003, page 11). Many other parameters, like the variance of potential outcomes or Gini coefficients, are not $D$-parameters. Nonetheless, the cdf and mean bounds we derived can be used as a direct input into theorem 2 of Stoye2010 to derive explicit, analytical bounds on these spread parameters. In future work it would be helpful to obtain precise expressions for these spread parameter bounds. Finally, while we have given several suggestions for how to interpret conditional $c$-dependence, there are likely other possibilities. For example, one could adapt Rosenbaum and Silber's RosenbaumSilber2009 `amplification' approach to our setting. Incorporating this or other ideas from the extensive literature on conditional treatment probabilities would be a helpful addition to our nonparametric sensitivity analysis.
\endinput
\singlespacing