EconBase
← Back to paper

Choosing Exogeneity Assumptions in Potential Outcome Models

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,399 characters · 19 sections · 31 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Choosing Exogeneity Assumptions in Potential Outcome Models

abstractThere are many kinds of exogeneity assumptions. How should researchers choose among them? When exogeneity is imposed on an unobservable like a potential outcome, we argue that the form of exogeneity should be chosen based on the kind of selection on unobservables it allows. Consequently, researchers can assess the plausibility of any exogeneity assumption by studying the distributions of treatment given the unobservables that are consistent with that assumption. We use this approach to study two common exogeneity assumptions: quantile and mean independence. We show that both assumptions require a kind of non-monotonic relationship between treatment and the potential outcomes. We discuss how to assess the plausibility of this kind of treatment selection. We also show how to define a new and weaker version of quantile independence that allows for monotonic treatment selection. We then show the implications of the choice of exogeneity assumption for identification. We apply these results in an empirical illustration of the effect of child soldiering on wages.

JEL classification: C14; C18; C21; C25; C51

Keywords: Selection on Unobservables, Nonparametric Identification, Treatment Effects, Partial Identification, Sensitivity Analysis

\onehalfspacing

Introduction

Exogeneity is a critical assumption in much structural or causal empirical work. In general, exogeneity refers to assumptions on the statistical dependence between an observable term and an unobservable term.\footnote{Throughout this paper we use `exogeneity' in the same sense as the treatment effects literature; for example, see Imbens2004. This is related to but distinct from the Cowles Commission definition of an exogenous variable as a variable that is determined outside of the model under consideration. See the discussion in HendryMorgan1995, Imbens1997, and Heckman2000.} There are many such assumptions, however, including zero correlation, median independence, and full statistical independence. The choice of a formal definition of exogeneity is not innocuous. Different definitions have different substantive interpretations and different implications for identification, rates of convergence, asymptotic distributions, efficiency bounds, and overidentification.

In the context of potential outcome models, there has been debate over the appropriate choice of exogeneity assumption. HeckmanIchimuraTodd1998 assume potential outcomes are mean independent of treatment, conditional on covariates. They justify this focus on mean independence by arguing that “conditional independence assumptions...are far stronger than the mean-independence conditions typically invoked by economists” (page 262). Imbens2004 agrees that “this [mean independence] assumption is unquestionably weaker [than full independence]”, but argues that “in practice it is rare that a convincing case is made for the weaker [mean independence] assumption 2.3 without the case being equally strong for the stronger [full independence assumption].” The justification he provides is that “the weaker assumption is intrinsically tied to functional-form assumptions, and as a result one cannot identify average effects on transformations of the original outcome (such as logarithms) without the stronger assumption.”

In this paper, we contribute to this debate as follows: We recommend that researchers focus directly on the substantive economic interpretation of the assumption, and the plausibility of its restrictions, rather than assess assumptions based on what they can be used to identify (as in the functional form dependency critique of mean independence) or on mathematical orderings of what implies what (as in the observation that statistical independence implies mean independence, but not vice versa). Specifically, we focus on two of the most common forms of exogeneity assumptions used, quantile independence and mean independence. We provide several results to help researchers assess the plausibility of these exogeneity assumptions. First, in section (ref), we provide a brief informal discussion of the motivations for making exogeneity assumptions. There we note that exogeneity assumptions are typically made on structural unobservables, variables that satisfy some kind of policy or treatment invariance property, like potential outcomes or unobserved ability. In this case, the form of exogeneity depends on the form of treatment selection. Consequently, the plausibility of any exogeneity assumption can be assessed by examining the distributions of treatment given the unobservables that are consistent with that assumption.

Next, in section (ref), we characterize these distributions of treatment given the unobservables that are consistent with either quantile independence or mean independence. This characterization shows that both quantile independence and mean independence require specific kinds of non-monotonic treatment selection. In section (ref) we show how to modify the quantile independence assumption to create a new, weaker exogeneity assumption which allows for more plausible forms of treatment selection, including monotonic treatment selection. We call this assumption $\mathcal{U}$-independence. In section (ref) we derive identified sets for the average effect of treatment for the treated (ATT) and the quantile treatment effect for the treated (QTT) parameters under both quantile independence and $\mathcal{U}$-independence. These identified sets have a simple, closed form characterization, which makes them easy to use in practice. By comparing these identified sets we show that the identifying power of quantile independence comes from the fact that it only allows for a restrictive kind of non-monotonic selection on unobservables. In section (ref) we examine these differences in an empirical illustration of the effects of child soldiering on wages based on unconfoundedness. We show that the baseline results are generally robust under quantile independence relaxations of unconfoundedness, but not under $\mathcal{U}$-independence relaxations. This difference highlights both the practical importance of choosing exogeneity assumptions and how our approach can help researchers make this choice. Finally, in section (ref) we use a Roy model to further illustrate how researchers can assess the plausibility of non-monotonic selection on unobservables.

Choosing Exogeneity Assumptions

There are many different kinds of exogeneity assumptions available to researchers. Manski1988 and Powell1994 catalog some of the most common forms, including zero correlation, mean independence, quantile independence, conditional symmetry, statistical independence, and a variety of index conditions. Other kinds of exogeneity assumptions have since been defined, including mean monotonicity (e.g., ManskiPepper2000,ManskiPepper2009, chapter 2 of Manski2003), approximate mean independence (Manski2003, section 9.4), stochastic dominance assumptions (e.g., BlundellEtAl2007), and quantile uncorrelation (KomarovaSeveriniTamer2012), among many others. How should researchers choose among these many options? In this section we discuss one approach to answering this question.

The answer depends on whether the exogeneity assumption is made on a structural unobservable or a reduced form unobservable. This distinction goes back to the earliest work on simultaneous equation models in econometrics (see Hausman1983 for a survey) but has been used in recent work as well (e.g. BlundellMatzkin2014). Structural unobservables are variables that satisfy some kind of policy or treatment invariance property, like potential outcomes, unobserved ability, or preferences. Reduced form unobservables are functions of the structural unobservables, and possibly other variables in the model, like realized treatment. Since quantile independence is often imposed on the relationship between treatment variables and structural unobservables, we focus on that case. We briefly discuss exogeneity assumptions for reduced form unobservables in appendix (ref), along with some additional background and examples.

Let $X$ denote the observed, realized treatment, and let $Y_x$ be a potential outcome where $x$ is a logically possible value of treatment. Since, by definition, potential outcomes have a meaning and interpretation that does not depend on the realized treatment $X$, any stochastic dependence between them must reflect some form of treatment selection on $Y_x$. That is, it must reflect some form of selection on unobservables, which is described by the distribution of $X \mid Y_x$. Many exogeneity assumptions, such as quantile independence and mean independence, are defined as constraints on the distribution of $Y_x \mid X$. Consequently, to assess the plausibility of these assumptions, we recommend examining the set of distributions of $X \mid Y_x$ that are consistent with the given constraint on the distribution of $Y_x \mid X$. This allows researchers to use the large literature on treatment selection to assess the plausibility of various exogeneity assumptions. Note that this discussion and recommendation extend to more general structural or causal models where one is considering exogeneity assumptions between a structural unobservable $U$ and an observed variable $X$; $U = Y_x$ is the specific case we focus on here.

Descriptive Analysis and Causal Models

Before proceeding, it is important to emphasize the scope of the process we just described: We are interested in assumptions about the dependence structure between observable and unobservable variables in causal models. This does not include research whose end goal is a description of the joint distribution of observed random variables, and which does not aim to make causal statements. Such descriptive research studies the relationship between observed variables. For example, suppose $Y$ is an observed outcome and $X$ is an observed covariate. We might define $E$ to be the residual from a linear projection of $Y$ onto $(1,X)$. We can then ask about the statistical relationship between $E$ and $X$: It satisfies zero correlation by construction, but not necessarily other restrictions like mean independence or statistical independence. In this case, however, the joint distribution of $(E,X)$ is always point identified. Hence, at the population level, the precise relationship between these variables is always known. Consequently, there is no need to make or choose assumptions about the stochastic relationship between $E$ and $X$. In contrast, in causal models, exogeneity assumptions typically have substantial identifying power for causal effects or structural parameters.

Characterizing Exogeneity Assumptions

In this section, we present our main characterization results. We provide results for two of the most common exogeneity assumptions: quantile independence and mean independence. We focus on binary treatments throughout the paper; we generalize our results to multi-valued discrete and continuous treatments in appendix (ref). All of our results also hold if one conditions on an additional vector of observed covariates, as is typically the case in empirical applications, but we omit these for simplicity.

The Potential Outcomes Model

Let $X \in \{0,1\}$ be a binary treatment variable. Let $(Y_1,Y_0)$ denote unobserved potential outcomes. We observe the scalar outcome variable

equation[equation omitted — 68 chars of source]

Let $p_x = \ensuremath{\mathbb{P}}(X=x)$ for $x \in \{0,1\}$. We impose the following assumption on the joint distribution of $(Y_1,Y_0,X)$.

partialIndepAssumpFor each $x,x' \in \{0,1\}$: \begin{enumerate} • $Y_x \mid X=x'$ has a strictly increasing and continuous distribution function on its support, $\operatorname*{supp}(Y_x \mid X=x')$. • $\operatorname*{supp}(Y_x \mid X=x') = \operatorname*{supp}(Y_x) = [\underline{y}_x,\overline{y}_x]$ where $-\infty \leq \underline{y}_x < \overline{y}_x \leq \infty$. • $p_x > 0$. \end{enumerate}

Via A(ref).(ref), we restrict attention to continuously distributed potential outcomes. A(ref).(ref) states that the unconditional and conditional supports of $Y_x$ are equal, and are a possibly infinite closed interval. We maintain A(ref).(ref) for simplicity, but it can be relaxed using similar derivations as in MastenPoirier2016. A(ref).(ref) is an overlap assumption.

In potential outcome models, statistical independence between potential outcomes and treatment is sometimes assumed:

equation[equation omitted — 95 chars of source]

This assumption, equivalent to random assignment of treatment, can be made for $x =0$, $x=1$, or both. When this assumption is made conditional on covariates, it is often called unconfoundedness. As discussed in section (ref), however, there are many other kinds of exogeneity assumptions available in the literature, which are all weaker than full statistical independence. The purpose of this paper is to help researchers choose between these different kinds of exogeneity assumptions. To that end, in the next two subsections we provide characterization results that can help researchers assess the plausibility of these exogeneity assumptions. We then study the identifying power of these different assumptions in section (ref).

A Class of Quantile Independence Assumptions

Quantile independence of the potential outcome $Y_x$ from $X$ at quantile $\tau$ holds if

equation[equation omitted — 109 chars of source]

This assumption is often imposed at a single quantile. For example, imposing (ref) at $\tau = 0.5$ yields median independence. If this holds for all $\tau \in (0,1)$, then $Y_x$ and $X$ are statistically independent. Therefore quantile independence is a relaxation of independence.

It is often more natural to work with cdfs, an inverse of the quantile function.\footnote{For example, see assumption QI on page 731 of Manski1988 or equation (1.7) on page 2452 of Powell1994. Definitions using cdfs and those using quantiles directly (e.g., via equation (ref)) are often equivalent. Throughout this paper we use “quantile independence” to mean the cdf-based definition, as is common in the literature.} Say $Y_x$ is $\tau$-cdf independent of $X$ if

equation[equation omitted — 98 chars of source]

Note that $Y_x$ is quantile independent of $X$ at quantile $\tau$ if and only if $Y_x$ is $Q_{Y_x}(\tau)$-cdf independent of $X$ for continuously distributed $Y_x$. This motivates the following definition.\footnote{See BelloniChenChernozhukov2017 and ZhuZhangXu2017 for similar generalizations of quantile independence.}

definitionLet $\mathcal{T}$ be a subset of $\ensuremath{\mathbb{R}}$. Say $Y_x$ is $\mathcal{T}$-independent of $X$ if the cdf independence condition (ref) holds for all $\tau \in \mathcal{T}$.

With binary treatments, the dependence structure between $X$ and the potential outcome $Y_x$ is fully characterized by the function \[ p(y_x) = \ensuremath{\mathbb{P}}(X=1 \mid Y_x=y_x), \] which we call the latent propensity score. Full statistical independence of $Y_x$ and $X$, or $\mathcal{T}$-independence with $\mathcal{T} = [\underline{y}_x,\overline{y}_x]$, is equivalent to this latent propensity score being constant: \[ p(y_x) = \ensuremath{\mathbb{P}}(X=1) \] for almost all $y_x \in \operatorname*{supp}(Y_x)$. Analogously, $\mathcal{T}$-independence for $\mathcal{T} \subsetneq [\underline{y}_x,\overline{y}_x]$ is weaker than full independence, and thus partially restricts the shape of $p(y_x)$. The following theorem characterizes the set of latent propensity scores consistent with $\mathcal{T}$-independence.

theorem[Average value characterization] Suppose $X$ is binary and A(ref).1 holds. Then $Y_x$ is $\mathcal{T}$-independent of $X$ if and only if \begin{equation} \ensuremath{\mathbb{E}} \big( p(Y_x) \mid Y_x \in (t_1,t_2) \big) = \ensuremath{\mathbb{P}}(X=1) \end{equation} for all $t_1, t_2 \in \mathcal{T} \cup \{ \underline{y}_x,\overline{y}_x \}$ with $t_1 < t_2$.

The proof, along with all others, is in appendix (ref). Theorem (ref) says that $\mathcal{T}$-independence holds if and only if for every interval with endpoints in $\mathcal{T} \cup \{ \underline{y}_x,\overline{y}_x \}$ the average latent propensity score over $Y_x \in (t_1,t_2)$ equals the overall average of the latent propensity score, which is $\ensuremath{\mathbb{P}}(X=1)$. Also note that $\ensuremath{\mathbb{E}}(p(Y_x) \mid Y_x \in (t_1,t_2)) = \ensuremath{\mathbb{P}}(X=1 \mid Y_x \in (t_1,t_2))$.

To illustrate theorem (ref), suppose $\mathcal{T} = \{ 0.5 \}$ and $\ensuremath{\mathbb{P}}(X=1) = 0.5$. Further suppose that $Y_x$ is uniformly distributed on $[0,1]$ to simplify the figures. Here we have just a single nontrivial cdf independence condition: median independence. Figure (ref) plots three different latent propensity scores which are consistent with $\mathcal{T}$-independence under this choice of $\mathcal{T}$; that is, which are consistent with median independence. This figure illustrates several features of such latent propensity scores: The value of $p(y_x)$ may vary over the entire range $[0,1]$. $p$ does not need to be symmetric about $y_x = 0.5$, nor does it need to be continuous. It does need to satisfy equation (ref) over the intervals $(t_1,t_2) = (0,0.5)$ and $(t_1,t_2) = (0.5,1)$. Finally, as suggested by the pictures, $p$ must actually be nonmonotonic; we show this in corollary (ref) next.

figure[figure omitted — 384 chars of source]
corollarySuppose $X$ is binary and A(ref).1 holds. Suppose the latent propensity score $p$ is weakly monotonic and not constant on $(\underline{y}_x,\overline{y}_x)$. Then, for all $\tau \in (\underline{y}_x,\overline{y}_x)$, $Y_x$ is not $\tau$-cdf independent of $X$.

Corollary (ref) shows that any quantile independence assumption rules out all types of monotonic selection, except for the trivially monotonic constant $p(y_x) = \ensuremath{\mathbb{P}}(X=1)$ implied by full statistical independence.

Next we show that imposing multiple quantile independence conditions imposes further non-monotonicity. Say that a function $f$ changes direction at least $K$ times if there exists a partition of its domain into $K$ intervals such that $f$ is not monotonic on each interval.

corollarySuppose $X$ is binary and A(ref).1 holds. Suppose $Y_x$ is $\mathcal{T}$-independent of $X$. Suppose there exists a version of $p$ without removable discontinuities. Partition $(\underline{y}_x,\overline{y}_x)$ by the sets $\mathcal{Y}_1 = (t_0,t_1)$, $\mathcal{Y}_k = [t_{k-1},t_k)$ for $k=2,\ldots,K$ with $t_0 = \underline{y}_x$, $t_K = \overline{y}_x$, and such that for each $k$ there is a $\tau_k \in \mathcal{T} \cap \mathcal{Y}_k$. Suppose $p$ is not constant over each set $\mathcal{Y}_k$, $k=1,\ldots,K$. Then $p$ changes direction at least $K$ times.

This result says that such latent propensity scores must oscillate up and down at least $K$ times (we assume $p$ does not have removable discontinuities to rule out trivial direction changes). For example, as in figure (ref), suppose we continue to have $\ensuremath{\mathbb{P}}(X=1) = 0.5$ and $Y_x \sim \text{Unif}[0,1]$ but we add a few more isolated $\tau$'s to $\mathcal{T}$. Figure (ref) shows several latent propensity scores consistent with $\mathcal{T}$-independence when $\mathcal{T}$ has several isolated elements. Consider the figure on the left, with $\mathcal{T} = \{ 0.25, 0.5, 0.75 \}$. Partition $(0,1) = (0,0.4) \cup [0.4,0.6) \cup [0.6,1)$. Then $p$ is not monotonic over each partition set, and each partition set contains one element of $\mathcal{T}$: $0.25 \in (0,0.4)$, $0.5 \in [0.4,0.6)$, and $0.75 \in [0.6,1)$. There are $K=3$ partition sets, and hence the corollary says $p$ must change direction at least 3 times. We see this in the figure since there are 3 interior local extrema. A similar analysis holds for the figure on the right. Overall, these triangular and sawtooth latent propensity scores illustrate the oscillation required by corollary (ref).

figure[figure omitted — 420 chars of source]

We document one more feature: As long as there is some interval that is not in $\mathcal{T}$ then there is a latent propensity score that takes the most extreme values possible, 0 and 1.

corollarySuppose $X$ is binary and A(ref).1 and A(ref).3 hold. Suppose $[\underline{y}_x,\overline{y}_x] \setminus \mathcal{T}$ contains a non-degenerate interval. Then there exists a latent propensity score which is consistent with $\mathcal{T}$-independence of $Y_x$ from $X$ and for which the sets \[ \{y_x \in [\underline{y}_x,\overline{y}_x]: p(y_x) = 0\} \qquad \text{and} \qquad \{y_x \in [\underline{y}_x,\overline{y}_x]: p(y_x) = 1\} \] have positive Lebesgue measure.

Consequently, $\mathcal{T}$-independence allows for a kind of extreme imbalance, where there is a positive mass of potential outcome values that only appear in the treatment group and another positive mass of potential outcome values that only appear in the control group.

Characterizing Mean Independence

Mean independence is another commonly used exogeneity assumption. For example, HeckmanIchimuraTodd1998 assume potential outcomes are mean independent of treatments, conditional on covariates. As in the previous section, we characterize the constraints this assumption places on the conditional distribution of $X$ given $Y_x$.

definitionSay $Y_x$ is mean independent of $X$ if $\ensuremath{\mathbb{E}}(Y_x \mid X=0) = \ensuremath{\mathbb{E}}(Y_x \mid X=1)$.

From definition (ref) and Bayes' rule, it immediately follows that

equation[equation omitted — 160 chars of source]

assuming $\ensuremath{\mathbb{E}}(Y_x) \neq 0$. Theorem (ref) showed that quantile independence constrains the unweighted average value of the latent propensity score over certain subintervals of its domain. In contrast, equation (ref) shows that mean independence constrains a weighted average value of the latent propensity score over its entire domain. Equation (ref) can be extended to multi-valued and continuous $X$ as in our analysis of quantile independence in section (ref); we omit this extension for brevity.

Although mean independence imposes a different constraint on the latent propensity score than quantile independence, it also requires non-constant latent propensity scores to be non-monotonic.

propositionSuppose $X$ is binary and A(ref).1 holds. Suppose $\ensuremath{\mathbb{E}}(|Y_x|) < \infty$. Suppose the latent propensity score $p$ is weakly monotonic and not constant on the interior of its domain. Then $Y_x$ is not mean independent of $X$.

Therefore, assuming mean independence rules out all types of monotonic selection, except for the trivially monotonic constant $p(y_x) = \ensuremath{\mathbb{P}}(X=1)$ implied by full statistical independence.

Discussion

To place these results in context, consider the case where $X$ is an indicator for completing college and $Y_0$ denotes a person's earnings if they do not complete college. $X$ is often thought to be endogenous due to its relationship with ability, which is captured by $Y_0$. In this example, $p(y_0)$ is the proportion of people who complete college, among those with a fixed level of non-graduate earnings. Corollary (ref) states that any quantile independence condition rules out nonconstant, weakly monotonic treatment selection. Similarly, proposition (ref) implies that mean-independence rules out this monotonic treatment selection. They would thus rule out that the proportion who attend college is weakly increasing in the level of non-graduate earnings, unless we assume that college attendance and non-graduate earnings are statistically independent, in which case $p(y_0)$ is constant.

Corollary (ref) would require the probability of attending college to oscillate (or be constant) to accommodate multiple quantile independence conditions. For example, if $K$ quantile independence conditions hold, there must exist $K$ non-graduate earnings thresholds where the effect of non-graduate earnings on college attendance changes sign. Alternatively, college attendance can again be statistically independent of non-graduate earnings. We discuss the plausibility of these oscillations in the context of a Roy Model in section (ref).

Corollary (ref) characterizes another feature of the latent propensity scores allowed by quantile independence. In our returns to schooling example, a finite number of quantile independence restrictions allow for a strictly positive proportion of people with non-graduate earnings levels for which nobody attends college ($p(y_0) = 0$), and for which everyone attends college ($p(y_0) = 1$). Depending on the context, the existence of these two groups may or may not appear plausible. If it appears implausible, using an assumption that allows for their existence is inefficient compared to an assumption that rules out their existence. Since ruling out their existence is a stronger assumption, it would result in weakly narrower identified sets compared to the overly conservative set obtained under quantile independence alone. To rule out their existence, one could add a constraint such as $p(y_0) \in [c_1, c_2]$ for prespecified $0 < c_1 \leq c_2 < 1$ and derive the identified set for the relevant parameter under quantile independence and this added constraint. Throughout this paper, we emphasize this approach of tailoring exogeneity conditions to the likely and unlikely features of the treatment selection function $p(y_0)$ in the given application. Precisely finding the assumption that best captures the nonexistence of those groups is beyond the scope of this paper, however, since it is application dependent.

The Identifying Power of Different Exogeneity Assumptions

In this section we study the implications of the choice of exogeneity assumption for identification. Our first result in section (ref) shows that quantile independence imposes a constraint on the average value of a latent propensity score. We use this characterization to motivate an assumption weaker than quantile independence, which we call $\mathcal{U}$-independence. The difference between these two assumptions is that quantile independence imposes some additional average value constraints on $p(y_x)$ that $\mathcal{U}$-independence does not. In particular, $\mathcal{U}$-independence allows for monotonic treatment selection. Hence the difference between identified sets obtained under these two assumptions is a measure of the identifying power of these additional average value constraints, which are the features of quantile independence that require the latent propensity score to be non-monotonic. Unlike quantile independence, it is not clear to us how to naturally weaken mean independence to allow for monotonic latent propensity scores while still retaining an interpretable assumption with identifying power. For that reason, in this section we only study identification under quantile independence and its weaker version, $\mathcal{U}$-independence.

We focus on two parameters: The average treatment effect for the treated,

align*[align* omitted — 157 chars of source]

and the quantile treatment effect for the treated,

align*[align* omitted — 142 chars of source]

for $q \in (0,1)$. To analyze treatment on the treated parameters, we only need to make assumptions on the relationship between $Y_0$ and $X$. Our analysis can easily be extended to parameters like ATE by imposing $\mathcal{T}$- or $\mathcal{U}$-independence between $Y_1$ and $X$ as well as between $Y_0$ and $X$.

Under statistical independence $Y_0 \mathbin{ \mathpalette{\@indep}{} } X$, both the ATT and QTT are point identified. Under $\mathcal{T}$- or $\mathcal{U}$-independence, however, they are generally partially identified. We derive identified sets for ATT and QTT under both classes of exogeneity assumptions in this section. These sets have simple explicit expressions which make them easy to use in practice. In section (ref) we compare these identified sets in an empirical application. In this application, the estimated identified sets are significantly larger under $\mathcal{U}$-independence, implying that the additional average value constraints inherent in $\mathcal{T}$-independence have substantial identifying power.

Weakening Quantile Independence

Throughout this section, we focus on the case where $\mathcal{T}$ is an interval. In this case, we show that latent propensity scores consistent with $\mathcal{T}$-independence have two features: (a) they are flat on $\mathcal{T}$ and (b) they are non-monotonic outside the flat regions, such that the average value constraint (ref) is satisfied. We use this finding to motivate a weaker assumption which retains feature (a) but drops feature (b). We call this assumption $\mathcal{U}$-independence. This new weaker assumption has two uses: First, we can use it as a tool for understanding quantile independence itself. Specifically, by comparing identified sets under quantile independence and under the weaker $\mathcal{U}$-independence we will learn the identifying power of the average value constraints on the latent propensity score. Second, $\mathcal{U}$-independence can be used by itself as a method for relaxing statistical independence and performing sensitivity analysis. We illustrate both of these uses below and in our empirical analysis of section (ref).

We begin with the following corollary to theorem (ref).

corollarySuppose $X$ is binary and that A(ref) holds for $Y_0$. Let $\mathcal{T} = [a,b]\subseteq [\underline{y}_0,\overline{y}_0]$. Then $\mathcal{T}$-independence of $Y_0$ from $X$ implies \begin{equation} \ensuremath{\mathbb{P}}(X=1 \mid Y_0=y_0) = \ensuremath{\mathbb{P}}(X=1) \end{equation} for almost all $y_0 \in \mathcal{T}$.

Corollary (ref) shows that $\mathcal{T}$-independence requires the latent propensity score to be constant on $\mathcal{T}$ and equal to the overall unconditional probability of being treated. The first property---that the latent propensity score is flat on $\mathcal{T}$---means that random assignment holds within the subpopulation of units whose untreated outcomes are in the set $\mathcal{T}$; that is, $X \mathbin{ \mathpalette{\@indep}{} } Y_0 \mid \{ Y_0 \in \mathcal{T} \}$. Corollary (ref) can be generalized to allow $\mathcal{T}$ to be a finite union of intervals, but we omit this for brevity.

This corollary motivates the following definition.

definitionLet $\mathcal{U} \subseteq [\underline{y}_x,\overline{y}_x]$ be an interval. Say that $Y_x$ is $\mathcal{U}$-independent of $X$ if $\ensuremath{\mathbb{P}}(X=1 \mid Y_x=y_x) = \ensuremath{\mathbb{P}}(X=1)$ for almost all $y_x \in \mathcal{U}$.

Importantly, unlike $\mathcal{T}$-independence, $\mathcal{U}$-independence allows for monotonic treatment selection. Corollary (ref) shows that $\mathcal{T}$-independence implies $\mathcal{U}$-independence with $\mathcal{U} = \mathcal{T}$. The converse does not hold since $\mathcal{T}$-independence requires additional average value constraints to hold, by theorem (ref). In particular, $\mathcal{U}$-independence implies the average value constraint \[ \ensuremath{\mathbb{E}}(p(Y_x) \mid Y_x \in (t_1,t_2)) = \ensuremath{\mathbb{P}}(X=1) \tag{\ref{eq:averageValueCondition_main}} \] for all $t_1,t_2 \in \mathcal{U}$, because it requires that $p(u)$ is constant on $\mathcal{U}$. But $\mathcal{T}$-independence also requires that (ref) holds for choices of $t_1$ and $t_2$ in $\mathcal{U} \cup \{\underline{y}_x,\overline{y}_x\}$. That is, $t_1$ and $t_2$ can equal the end points $\underline{y}_x$ or $\overline{y}_x$. Hence it imposes an average value constraint outside of the set $\mathcal{U}$. For example, figure (ref) shows two latent propensity scores. One satisfies $\mathcal{T}$-independence, but the other only satisfies $\mathcal{U}$-independence. Finally, note that $\mathcal{U}$-independence is a nontrivial assumption only when $\ensuremath{\mathbb{P}}(Y_x \in \mathcal{U}) > 0$. Conversely, $\mathcal{T}$-independence is nontrivial even when $\mathcal{T}$ is a singleton.

figure[figure omitted — 566 chars of source]

The Identified Sets For ATT and $\text{QTT}(q)$

In this subsection we derive sharp bounds on the ATT and $\text{QTT}(q)$ under both $\mathcal{T}$- and $\mathcal{U}$-independence. To do so, it suffices to derive bounds on $Q_{Y_0 \mid X}(q \mid 1)$.

We show the validity of the following bounds in proposition (ref) below. Let $\mathcal{T} = \mathcal{U} = [Q_{Y_0}(a),Q_{Y_0}(b)]$ for $0 < a \leq b < 1$. The $\mathcal{T}$-independence bounds are then defined by \[ \overline{Q}_{Y_0 \mid X}^\mathcal{T}(\tau \mid 1) =

casesQ_{Y \mid X}(a \mid 0) & for $\tau \in (0,a]$ \\ Q_{Y \mid X}(\tau \mid 0) & for $\tau\in(a,b]$ \\ Q_{Y \mid X}(1 \mid 0) & for $\tau\in(b,1)$,

\qquad Q_{Y_0 \mid X}^\mathcal{T}(\tau \mid 1) =

casesQ_{Y \mid X}(0 \mid 0) & for $\tau\in(0,a]$ \\ Q_{Y \mid X}(\tau \mid 0) & for $\tau\in(a,b]$ \\ Q_{Y \mid X}(b \mid 0) & for $\tau \in(b,1)$.

\] We let $Q_{Y \mid X}(0 \mid x) = \underline{y}_x$ and $Q_{Y \mid X}(1 \mid x) = \overline{y}_x$. For $\mathcal{U}$-independence, there are two cases. First consider the lower bound. If $(1-(b-a))p_1 \leq a$, \[ Q_{Y_0 \mid X}^\mathcal{U}(\tau \mid 1) =

casesQ_{Y \mid X}(0 \mid 0) & for $\tau \in (0,1-(b-a)]$ \\ Q_{Y \mid X}\left( \tau + \dfrac{b-1}{p_0} \mid 0 \right) & for $\tau \in (1-(b-a), 1)$.

\] If $(1-(b-a))p_1 \geq a$, \[ Q_{Y_0 \mid X}^\mathcal{U}(\tau \mid 1) =

casesQ_{Y \mid X}(0 \mid 0) & for $\tau \in \left(0,\dfrac{a}{p_1} \right]$ \\ Q_{Y \mid X} \left(\tau - \dfrac{a}{p_1} \mid 0 \right) & for $\tau \in \left( \dfrac{a}{p_1},\dfrac{a}{p_1} + b-a \right]$ \\ Q_{Y \mid X}(b-a \mid 0) & for $\tau \in \left( \dfrac{a}{p_1} + b-a,1 \right)$.

\] Next consider the upper bound. If $(1-(b-a))p_0 \leq a$, \[ \overline{Q}_{Y_0 \mid X}^\mathcal{U}(\tau \mid 1) =

casesQ_{Y \mid X}(1-(b-a) \mid 0) &for $\tau \in \left( 0,1-(b-a) - \dfrac{1-b}{p_1} \right]$ \\ Q_{Y \mid X} \left( \tau +\dfrac{1-b}{p_1} \mid 0 \right) &for $\tau \in \left( 1-(b-a) - \dfrac{1-b}{p_1}, 1 - \dfrac{1-b}{p_1} \right]$ \\ Q_{Y \mid X}(1 \mid 0) &for $\tau \in \left( 1 - \dfrac{1-b}{p_1}, 1 \right)$.

\] If $(1-(b-a))p_0 \geq a$, \[ \overline{Q}_{Y_0 \mid X}^\mathcal{U}(\tau \mid 1) =

casesQ_{Y \mid X} \left(\tau + \dfrac{a}{p_0} \mid 0 \right) & for $\tau \in (0, b-a]$ \\ Q_{Y \mid X}(1 \mid 0) & for $\tau \in (b-a,1)$.

\]

propositionLet A(ref) hold. Suppose $Y_0$ is $\mathcal{T}$-independent of $X$ with $\mathcal{T} = [Q_{Y_0}(a),Q_{Y_0}(b)]$, $0 < a \leq b < 1$. Suppose the joint distribution of $(Y,X)$ is known. Let $q \in (0,1)$. Then \begin{equation} Q_{Y_0 \mid X}(q \mid 1) \in \left[ Q_{Y_0 \mid X}^\mathcal{T}(q \mid 1), \, \overline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1) \right]. \end{equation} Moreover, the interior of the set in equation (ref) equals the interior of the identified set. Finally, the proposition also holds if we replace $\mathcal{T}$ with $\mathcal{U}$ .

$\mathcal{T}$-independence of $Y_0$ from $X$ with $\mathcal{T} = [Q_{Y_0}(a),Q_{Y_0}(b)]$ is equivalent to the quantile independence assumptions $Q_{Y_0 \mid X}(\tau \mid x) = Q_{Y_0}(\tau)$ for all $\tau \in [a,b]$, by A(ref). The bounds (ref) are also sharp for the function $Q_{Y_0 \mid X}(\cdot \mid 1)$ in a sense similar to that used in proposition (ref) in the appendix; we omit the formal statement for brevity. This functional sharpness delivers the following result.

corollarySuppose the assumptions of proposition (ref) hold. Let $\ensuremath{\mathbb{E}}( | Y_0 |) < \infty$. Then $\ensuremath{\mathbb{E}}(Y_0 \mid X=1)$ lies in the set \[ \left[ \underline{\ensuremath{\mathbb{E}}}^\mathcal{T}(Y_0 \mid X=1), \,\overline{\ensuremath{\mathbb{E}}}^\mathcal{T}(Y_0 \mid X=1) \right] \equiv \left[ \int_0^1 \underline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1) \; dq, \, \int_0^1 \overline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1) \; dq \right]. \] Moreover, the interior of this set equals the interior of the identified set for $\ensuremath{\mathbb{E}}(Y_0 \mid X=1)$. Finally, the corollary also holds if we replace $\mathcal{T}$ with $\mathcal{U}$.

By proposition (ref) we have that $\mathcal{T}$-independence implies that $\text{QTT}(q)$ lies in the set \[ \left[ Q_{Y \mid X}(q \mid 1) - \overline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1), \; Q_{Y \mid X}(q \mid 1) - \underline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1) \right] \] and that the interior of this set equals the interior of the identified set for $\text{QTT}(q)$. Likewise for $\mathcal{U}$-independence. If $q \in \mathcal{T}$, then $\text{QTT}(q)$ is point identified under $\mathcal{T}$-independence; this follows immediately from our bound expressions above. This result---that a single quantile independence condition can be sufficient for point identifying a treatment effect---was shown by Chesher2003. A similar result holds in the instrumental variables model of ChernozhukovHansen2005 and the LATE model of ImbensAngrist1994. See the discussion around assumption 4 in section 1.4.3 of MellyWuthrich2017.

By corollary (ref) we have that $\mathcal{T}$-independence implies that the ATT lies in the set \[ \left[ \ensuremath{\mathbb{E}}(Y \mid X=1) - \overline{\ensuremath{\mathbb{E}}}^\mathcal{T}(Y_0 \mid X=1), \; \ensuremath{\mathbb{E}}(Y \mid X=1) - \underline{\ensuremath{\mathbb{E}}}^\mathcal{T}(Y_0 \mid X=1) \right] \] and that the interior of this set equals the interior of the identified set for the ATT. Likewise for $\mathcal{U}$-independence. Furthermore, in appendix (ref) we show that these ATT bounds have simple analytical expressions, obtained from integrating our closed form expressions for the bounds on $Q_{Y_0 \mid X}(q \mid 1)$.

Empirical Illustration: The Effect of Child Soldiering on Wages

In this section we use our results to study the impact of relaxing the unconfoundedness assumption in an empirical study of the effects of child soldiering on wages. We do this using both the $\mathcal{T}$- and $\mathcal{U}$-independence relaxations of statistical independence. We find that the identified sets are significantly larger under $\mathcal{U}$-independence. This implies that the average value constraints imposed by $\mathcal{T}$-independence have substantial identifying power; recall that these constraints are the features of quantile independence that require the latent propensity score to be non-monotonic. In particular, the baseline empirical results are generally quite robust under the $\mathcal{T}$-independence relaxation of unconfoundedness, but not the under $\mathcal{U}$-independence relaxation. This difference highlights the importance of the choice of exogeneity assumptions in practice, and how researchers can use their beliefs about the form of latent selection to assist in this choice.

Background

By collecting extensive survey data, BlattmanAnnan2010 study the impact of child abductions during a twenty year war in Uganda, where “an unpopular rebel group has forcibly recruited tens of thousands of youth” (page 882). Although they consider a variety of outcome variables, we focus on the impact of abduction on later life wages.

The main identification problem is that selection into military service is typically non-random. They argue, however, that forced recruitment in Uganda led to conditional random assignment of military service. They condition on two variables: (1) Prewar household size, because larger households were less likely to be raided by small bands of rebels, and (2) Year of birth, because abduction levels varied over time, so that some youth ages were more likely to be abducted than others. Hence their identification strategy is based on unconfoundedness, conditioning on these two variables. Although their qualitative evidence supporting unconfoundedness is compelling, this assumption is still nonrefutable. We therefore use our results to assess the sensitivity of relaxing unconfoundedness on their empirical conclusions.

Sample Definition

The data comes from phase 1 of SWAY, the Survey of War Affected Youth in northern Uganda (see AnnanBlattmanHorton2006). This phase has 1216 males born between 1975 and 1991. We look at the subsample of units who (1) have wage data available and (2) earned positive wages. This leaves us with 448 observations. Let $Y$ denote log wage. We define treatment $X$ to be an indicator that the person was not abducted. We include the two covariates discussed above, age when surveyed and household size in 1996. We omit other covariates for simplicity.

Age has 17 support points, household size has 21 support points, and treatment has 2 support points. Hence there are 714 total conditioning variable cells, relative to our sample size of 448 observations. To ensure that our conditional quantile estimators are reasonably smooth in the quantile index, we collapse these conditioning variables into 8 cells. Specifically, we replace age with a binary indicator of whether one is above or below the median age. Likewise, we replace household size with a binary indicator of whether one lived in a household with above or below median household size. This gives 8 total conditioning variable cells, with approximately 55 observations each.

Baseline Analysis

First we present estimates of conditional average treatment effects for the treated (CATT) and conditional quantile treatment effects for the treated (CQTTs), under the unconfoundedness assumption. For brevity we focus on the covariate cell $w$ = (age, household size) = (above median, above median). This group has the largest baseline effects of treatment, meaning that being abducted lowered their later life wages by the largest. Specifically, our estimate of the CATT for this group is 0.57. Our CQTT estimates are 0.67 for $\tau = 0.25$, 0.54 for $\tau = 0.5$, and 0.56 for $\tau = 0.75$. Note that our sample size is small, with 121 observations in this cell. We omit standard errors here because the purpose of this section is to illustrate the methods developed in our paper.

Sensitivity Analysis

To check the robustness of these baseline point estimates to failure of unconfoundedness, we estimate identified sets for the CATT and CQTT using our results from section (ref). To highlight the importance of the choice of relaxation, we consider sets $\mathcal{T} = \mathcal{U}$. In this case, corollary (ref) shows that $\mathcal{T}$-independence implies $\mathcal{U}$-independence. Hence identified sets using $\mathcal{T}$-independence must necessarily be weakly contained within identified sets using only $\mathcal{U}$-independence. We explore the magnitude of this difference in the data. Since $\mathcal{T}$-independence is simply $\mathcal{U}$-independence combined with some additional average value constraints, the difference between these identified sets tells us the identifying power of these additional average value constraints.

Specifically, we use the choice $\mathcal{T} = \mathcal{U} = [\delta,1-\delta]$ for $\delta \in [0,0.5]$. For $\delta = 0$, this choice corresponds to full conditional independence $Y_0 \mathbin{ \mathpalette{\@indep}{} } X \mid W=w$ under both classes of assumptions. For $\delta = 0.5$, this choice corresponds to median independence for $\mathcal{T}$-independence, and no assumptions for $\mathcal{U}$-independence. Values of $\delta$ between 0 and $0.5$ yield conditional partial independence between $Y_0$ and $X$ for both classes of assumptions.

figure[figure omitted — 379 chars of source]

Figure (ref) shows estimated identified sets for both CATT and $\text{CQTT}(\tau)$ as $\delta$ varies from $0$ to $0.5$, and for $\tau \in \{ 0.25, 0.5, 0.75 \}$. These are sample analog estimates, where $\widehat{Q}_{Y \mid X,W}(\cdot \mid x,w)$ is estimated by inverting a kernel based estimate of $F_{Y \mid X,W}(\cdot \mid x,w)$. First consider the plot on the top left, which shows the estimated $\text{CQTT}(0.5)$ bounds. The dashed lines are the identified sets under $\mathcal{T}$-independence. Since median independence of $Y_0$ from $X$ conditional on $W=w$ is sufficient to point identify the conditional median $Q_{Y_0 \mid X,W}(0.5 \mid 1, w)$, median independence is also sufficient to point identify the CQTT at $0.5$. Hence the identified set is a singleton for all $\delta \in [0,0.5]$. This singleton equals 0.54, the baseline estimate. Next consider the solid lines. These are the estimated identified sets under $\mathcal{U}$-independence. When $\delta = 0.5$, $\mathcal{U}$-independence does not impose any constraints on the model, and hence we obtain the no assumption bounds, which are quite wide: $[-3.42, 3.42]$. If we decrease $\delta$ a small amount, thus making the $\mathcal{U}$-independence constraint nontrivial, the estimated identified set does not change. In fact, we can impose random assignment for about the middle 50% of units (i.e., $\mathcal{U} = [0.25,0.75]$, or $\delta = 0.25$) and still we only obtain the no assumption bounds. Consequently, for intervals $\mathcal{T} \subseteq [0.25,0.75]$, the point identifying power of $\mathcal{T}$-independence is due solely to the constraint it imposes on the average value of the latent propensity score outside the interval $\mathcal{T}$, rather than the constraint that random assignment holds for units in the middle of the distribution of $Y_0$.

Next define \[ \delta^\mathcal{U}_\text{bp}(\tau) = \sup \{ \delta \in [0,0.5] : \text{LB}^\mathcal{U}(\tau, \delta) \geq 0 \} \] where $\text{LB}^\mathcal{U}(\tau,\delta)$ is the lower bound of the identified set for $\text{CQTT}(\tau)$ under $\mathcal{U}$-independence with $\mathcal{U} = [\delta,1-\delta]$. Define $\delta^\mathcal{T}_\text{bp}(\tau)$ analogously. This value $\delta^\mathcal{U}_\text{bp}(\tau)$ is a breakdown point: It is the largest amount we can relax full independence while still being able to conclude that the treatment effect is nonnegative. For $\tau = 0.5$, the estimated breakdown point for $\text{CQTT}(0.5)$ is 0.103. Thus we can allow randomization to fail for about 20.6% of units while still being able to conclude that $\text{CQTT}(0.5)$ is nonnegative. In contrast, as mentioned above, the breakdown point for $\mathcal{T}$-independence is always 0.5.

Next consider the lower two plots of figure (ref). These plots show estimated identified sets for $\text{CQTT}(0.25)$ on the left and $\text{CQTT}(0.75)$ on the right. There are two main differences between these plots and that of $\text{CQTT}(0.5)$: First, the $\mathcal{U}$-independence upper and lower bounds are not symmetric. Nonetheless, the qualitative robustness conclusions are similar. For example, $\widehat{\delta}^\mathcal{U}_\text{bp}(0.25)$ is 0.122 and $\widehat{\delta}^\mathcal{U}_\text{bp}(0.75)$ is 0.071. So conclusions about smaller quantiles are slightly more robust than conclusions about larger quantiles. Second, the $\mathcal{T}$-independence identified sets are no longer always singletons. In particular, we obtain non-singleton bounds when $\delta > 0.25$. However, conclusions under the $\mathcal{T}$-independence relaxation are substantially more robust than conclusions under the $\mathcal{U}$-independence relaxation. Specifically, $\widehat{\delta}^\mathcal{T}_\text{bp}(0.25)$ is 0.437. This is about 3.5 times as large as $\widehat{\delta}^\mathcal{U}_\text{bp}(0.25)$. Similarly, $\widehat{\delta}^\mathcal{T}_\text{bp}(0.75)$ is 0.25. This is also about 3.5 times as large as $\widehat{\delta}^\mathcal{U}_\text{bp}(0.75)$.

Finally consider the plot on the top right of figure (ref), which shows estimated identified sets for CATT. First consider the $\mathcal{T}$-independence relaxation, the dashed lines. The CATT is no longer point identified under median independence, or any set $\mathcal{T} \subsetneq (0,1)$ of quantile independence conditions; that is, the CATT is partially identified for all $\delta > 0$. Nonetheless, even median independence alone has substantial identifying power: For $\delta = 0.5$, the estimated identified set under median independence is $[-1.36, 2.06]$, whereas the no assumption bounds are $[-3.42, 3.42]$. Thus the width of the bounds has been cut in half. For $\delta > 0$, $\mathcal{U}$-independence has non-trivial identifying power, as shown in the solid lines. However, comparing the length of these bounds to the length to the $\mathcal{T}$-independence bounds, we see that imposing the average value constraint outside the interval $[\delta,1-\delta]$ again has substantial identifying power: the $\mathcal{T}$-independence bounds are anywhere from 50% ($\delta = 0.5)$ to almost 100% (arbitrarily small $\delta$) smaller than the $\mathcal{U}$-independence bounds. That is, the difference in lengths increases as we get closer to independence (as $\delta$ gets smaller). Thus conclusions about CATT are substantially more sensitive to small deviations from independence which do not impose the average value constraint outside the interval $[\delta,1-\delta]$, compared with small deviations which do impose that constraint. A second way to see this is to compare the breakdown points under the two relaxations. Define \[ \delta^\mathcal{U}_\text{bp} = \sup \{ \delta \in [0,0.5] : \text{LB}^\mathcal{U}(\delta) \geq 0 \} \] where $\text{LB}^\mathcal{U}(\delta)$ is the lower bound of the identified set for CATT under $\mathcal{U}$-independence with $\mathcal{U} = [\delta,1-\delta]$. Define $\delta^\mathcal{T}_\text{bp}$ analogously. As shown in the plots above, $\widehat{\delta}^\mathcal{U}_\text{bp} = 0.063$ while $\widehat{\delta}^\mathcal{T}_\text{bp} = 0.222$. Thus the breakdown point under $\mathcal{T}$-independence is again about 3.5 times as large as the breakdown point under $\mathcal{U}$-independence.

Empirical Conclusions

In this section we used our identification results to study the robustness of conclusions about CATT and CQTTs to failures of unconfoundedness. Our baseline point estimates suggest that child abduction and forced military service has a negative effect on later life wages, for those children who were older when they were abducted and who came from larger households. This holds both on average (from the CATT) and across the distribution of treatment effects (as seen in the CQTTs). We then asked: How sensitive are these conclusions to failures of unconfoundedness? We saw that using the $\mathcal{T}$-independence relaxation, these conclusions are generally robust to large relaxations of unconfoundedness. However, using the $\mathcal{U}$-independence relaxation, these conclusions appear much more sensitive. As we earlier discussed, the difference arises from the additional average value constraints that $\mathcal{T}$-independence imposes. Those constraints are the features of quantile independence that require the latent propensity score to be non-monotonic. Thus it is critical to assess the plausibility of those additional constraints when deciding between these two forms of exogeneity assumptions to use for assessing sensitivity.

In this empirical context, a monotonic latent propensity score arises when youths who have larger potential earnings when they're abducted (larger $Y_0$) are more likely to be abducted. If youths are targeted for abduction because of their innate or pre-existing skills, which would generally lead to large $Y_0$, then this would be a form of monotonic selection that would not be allowed for by the $\mathcal{T}$-independence relaxation, but would be allowed for by the $\mathcal{U}$-independence relaxation. So if we are concerned that unconfoundedness fails due to this kind of non-random selection, then $\mathcal{U}$-independence is a more appropriate choice for assessing sensitivity than $\mathcal{T}$-independence. Given this choice, the baseline results still hold under mild relaxations of unconfoundedness, since we saw that $\mathcal{U}$-independence breakdown points were generally around $\delta = 0.1$. But the baseline results no longer hold for larger relaxations; in this case, the data are inconclusive.

The Treatment Selection Implications of a Roy Model

As we emphasized, there is a direct mapping between exogeneity assumptions and the allowed forms of treatment selection. At one extreme, full independence assumes no selection at all of $X$ on $Y_x$, and therefore $p(y_x)$ is constant. On the other hand, weaker exogeneity assumptions allow for a class of deviations that one wishes to be robust against. Since this class is often not explicitly specified, we refer to such deviations as latent selection models. Our main results in section (ref) characterize the set of latent selection models allowed by quantile and mean independence restrictions.

In this section, we consider a class of Roy Models and examine the relationship between their implied treatment selection functions and the exogeneity assumptions of section (ref). We discuss different assumptions on the economic primitives which lead these models to be either consistent or inconsistent with quantile or mean independence restrictions. We only consider single-agent models, but similar analyses can likely be done for multi-agent models.

Suppose we are again interested in identifying the average treatment effect for the treated parameter

align*[align* omitted — 157 chars of source]

As in section (ref), its identification depends on our assumptions about the stochastic relationship between $X$ and $Y_0$. Suppose agents choose treatment to maximize their outcome:

equation[equation omitted — 76 chars of source]

This is the classical Roy model (see HeckmanVytlacil2007part1). This assumption specifies how treatment $X$ relates to $Y_0$. Specifically, consider the latent propensity score

align*[align* omitted — 133 chars of source]

The second line follows by our Roy model treatment choice assumption. Thus the shape of $p$ depends on the joint distribution of $(Y_1,Y_0)$. We classify these distributions into two possible cases, based on a concept called regression dependence (which is formally defined in definition (ref) in appendix (ref)).

enumerate• First suppose $(Y_1,Y_0)$ is such that $Y_1$ is regression dependent on $Y_0$. This implies that $p$ is monotonic. Corollary (ref) and proposition (ref) therefore imply that no mean or quantile independence conditions of $Y_0$ on $X$ can hold unless $X \mathbin{ \mathpalette{\@indep}{} } Y_0$. This occurs when $X$ is degenerate, as when treatment effects $Y_1-Y_0$ are constant, or more generally when $(Y_1 - Y_0) \mathbin{ \mathpalette{\@indep}{} } Y_0$. In particular, any mean or quantile independence condition of $Y_0$ on $X$ rules out bivariate normally distributed $(Y_1,Y_0)$, again unless $X \mathbin{ \mathpalette{\@indep}{} } Y_0$. • Next suppose $(Y_1,Y_0)$ is such that $Y_1$ is not regression dependent on $Y_0$. For example, let $Y_1 = Y_0 + \mu(Y_0) - \varepsilon$ where $\mu$ is a deterministic function and $\varepsilon \sim \ensuremath{\mathcal{N}}(0,1)$, $\varepsilon \mathbin{ \mathpalette{\@indep}{} } Y_0$. Then \[ p(y_0) = \ensuremath{\mathbb{P}}(X=1 \mid Y_0 = y_0) = \Phi[ \mu( y_0 ) ], \] where $\Phi$ is the standard normal cdf. If $\mu$ is non-monotonic then $p$ will also be non-monotonic. For this joint distribution of potential outcomes, the unit level treatment effects $Y_1-Y_0$ conditional on the baseline outcome $Y_0=y_0$ are distributed $\ensuremath{\mathcal{N}}( \mu(y_0), 1)$. Hence non-monotonicity of $\mu$ implies that the mean of this distribution of treatment effects is not monotonic. For instance, suppose the outcome is earnings and treatment is completing college. Let \begin{align*} \mu(y_0) &> 0 \qquadif $y_0 \in (\alpha,\beta)$ \\ \mu(y_0) &\leq 0 \qquadif $y_0 \in (-\infty,\alpha] \cup [\beta,\infty)$ \end{align*} for $-\infty < \alpha < \beta < \infty$. Then people with sufficiently small or sufficiently large earnings when they do not complete college do not benefit from completing college, on average. People with moderate earnings when they do not complete college, on the other hand, do typically benefit from completing college. This kind of joint distribution of potential outcomes combined with the Roy model assumption (ref) on treatment selection produces non-monotonic latent propensity scores. We just gave one example joint distribution of $(Y_1,Y_0)$ where regression dependence fails. More generally, theorem 5.2.10 on page 196 of Nelsen2006 characterizes the set of copulas for which $Y_1$ is regression dependent on $Y_0$, when both are continuously distributed. This result therefore also tells us the set of copulas where $Y_1$ is not regression dependent on $Y_0$. Among these copulas, $\mathcal{T}$-independence (or, analogously, mean independence) of $Y_0$ from $X$ will specify a further subset of allowed dependence structures. The precise set is given by all copulas which lead to latent propensity scores that satisfy the average value constraint.

Whether one of these cases is plausible depends on the specific application at hand. For example, Heckman, Smith, and Clements' HeckmanSmithClements1997 study the Job Training Partnership Act (JTPA). They find that “plausible impact distributions require high measures of positive dependence [of $Y_1$ on $Y_0$]” (page 506). This suggests that case 1 is more relevant for their data, and hence it may be unlikely that any quantile or mean independence holds in their setting.

Conclusion

In this paper we gave several results to help researchers assess the plausibility of quantile and mean independence assumptions on structural unobservables like potential outcomes. Keep in mind, however, that when doing identification analysis it is not necessary to choose a single exogeneity assumption. For example, researchers may want to consider a variety of exogeneity assumptions in this step, as part of a sensitivity analysis. We illustrated this in sections (ref) and (ref). The choice of which exogeneity assumptions to consider is still determined by considering the kinds of treatment selection we want to allow for, as discussed in section (ref). Conversely, there may be situations where researchers do not find it plausible to impose any kind of exogeneity assumption. In this case we often can still learn something about the parameters of interest, as in the classical no assumption bounds of Manski1990. In this paper we focused on the case where the researcher does want to impose some kind of exogeneity assumption, however. In this case, we hope that our results can help researchers better select the most appropriate exogeneity assumptions for their settings.

\singlespacing

\allowdisplaybreaks