Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
69,627 characters · 16 sections · 80 citation commands
Interpreting Quantile Independence
JEL classification: C14; C18; C21; C25; C51
Keywords: Nonparametric Identification, Treatment Effects, Partial Identification, Sensitivity Analysis
\onehalfspacing
A large literature has studied identification and estimation of structural econometric models under quantile independence rather than full independence.\footnote{Important early work includes KoenkerBassett1978, who introduced quantile regression, Manski1975, who studied discrete choice models under median independence, and Manski1985, who extended that analysis to quantile independence.} This literature has faced a longstanding question: How can one substantively interpret and judge the credibility of a given set of quantile independence conditions? For example, Manski1988 notes that “Quantile independence restrictions sometimes make researchers uncomfortable. The assertion that a given quantile of $u$ does not vary with $x$ may lead one to ask: Why this quantile but not others?”
In this paper, we answer this question by providing a treatment assignment characterization of quantile independence conditions. Specifically, we consider the relationship between a continuous unobserved variable $U$ and an observed treatment variable $X$, which may be binary, discrete, or continuous. For example, $U$ could be an unobserved structural variable like ability, or an unobserved potential outcome. In section (ref) we consider the binary $X$ case. There the dependence structure between $X$ and $U$ is fully characterized by the conditional probability \[ p(u) = \ensuremath{\mathbb{P}}(X = 1 \mid U=u). \] If $U$ were observed, this would be an ordinary propensity score. Since $U$ is not observed, however, we call this a latent propensity score; for brevity we will often refer to it simply as a propensity score. Constant propensity scores correspond to full statistical independence while non-constant propensity scores represent deviations from full independence. In this sense, any exogeneity assumption weaker than full independence allows for certain kinds of selection on unobservables. Our main theorem characterizes the set of propensity scores consistent with a set of quantile independence conditions. That is, we fully describe the kinds of selection on unobservables that quantile independence does and does not allow. This result shows that quantile independence imposes a set of constraints on average values of the propensity score $p$. We then describe several properties of propensity scores which satisfy these constraints. Most notably, non-constant propensity scores which are consistent with a single quantile independence condition must be non-monotonic. Furthermore, if multiple isolated quantile independence conditions hold, then non-constant propensity scores must also oscillate up and down. These results do not depend on a specific econometric model, and hence apply any time one makes a quantile independence assumption. We use our main result to compare the constraints quantile independence imposes on selection with the constraints imposed by mean independence. In particular, we show that mean independence also requires non-constant propensity scores to be non-monotonic.
In section (ref) we generalize our main theorem on characterization of quantile independence to discrete and continuous treatments. As in the binary treatment case, quantile independence is equivalent to a set of average value constraints on the distribution of treatment given the unobservable. In both the discrete and continuous case, this constraint implies a non-monontonicity result. Specifically, treatment cannot be regression dependent on the unobservable; equivalently, the distribution of treatment given unobservables cannot be stochastically monotonic in the unobservable.
To understand the restrictiveness of these non-monotonicity constraints, and therefore the plausibility of quantile independence assumptions, in section (ref) we study two simple economic models of treatment selection: One for continuous treatments and one for binary treatments. In both models we show that standard assumptions on the economic primitives imply treatment selection rules that are monotonic in the unobservable. Therefore, by our characterization results, these selection models with the standard assumptions are incompatible with quantile independence. We then discuss how quantile independence can be retained by considering alternative assumptions on the economic primitives. Consequently, researchers using quantile independence constraints must argue that such alternative assumptions are plausible in the empirical setting under consideration.
While the primary contribution of this paper is about the interpretation of quantile independence, in section (ref) we return to the binary treatment case to study the implications of our results for identification. We consider the case where an interval of quantile independence conditions hold, like $[0.25,0.75]$. We show that the latent propensity score must be flat on that interval. Our characterization result shows that the latent propensity score must also satisfy an average value constraint outside of that interval, which implies that the latent propensity score is non-monotonic outside of that interval. To isolate the identifying power of this average value constraint, we define a concept weaker than quantile independence, which we call $\mathcal{U}$-independence. This concept simply specifies that, on the set $\mathcal{U}$, the latent propensity score equals the overall unconditional probability of being treated. Hence the latent propensity score is flat on $\mathcal{U}$. While quantile independence on the set $\mathcal{U}$ also imposes this flatness constraint, $\mathcal{U}$-independence does not constrain the average value of the latent propensity score outside of this set. Thus the difference between identified sets derived under the two assumptions can be ascribed solely to this average value constraint, which is what requires latent propensity scores to be non-monotonic.
To understand the identifying power of the average value constraint, one must first specify an econometric model and a parameter of interest. While our quantile independence results are relevant for many different models (such as those cited in our literature review below), we focus on a simple but important model: the standard potential outcomes model of treatment effects with a binary treatment. We adapt the analysis of MastenPoirier2017 to derive identified sets for the average treatment effect for the treated (ATT) and the quantile treatment effect for the treated (QTT) under either an interval of quantile independence assumptions, or under $\mathcal{U}$-independence. We then compare these identified sets in a numerical illustration. In this illustration, the identified sets are significantly larger under $\mathcal{U}$-independence, implying that the average value constraint has substantial identifying power.
Quantile independence of $U$ from $X$ constraints the distribution of the unobservable $U$ conditional on the observable $X$. Our analysis is essentially a study of what these constraints on $U \mid X$ imply about the distribution of $X \mid U$. Hence it relies on the fact that, like mean independence, quantile independence treats these two variables asymmetrically. While this asymmetry has long been noted in the literature (for example, page 85 of Manski1988book), we are unaware of any prior study of its implications.\footnote{A large literature in statistics studies dependence concepts; for example, see Joe1997. Many of these concepts are asymmetric. This literature studies the properties of and relations between these concepts. Quantile independence is not one of the commonly studied concepts in this literature.} Instead, most prior research states various constraints on the joint distribution of $U$ and $X$ as a menu of options, with little guidance for choosing between them (for example, see Manski1988 or section 2 of Powell1994). Mean independence is sometimes argued to be undesirable since it is not invariant to strictly monotone transformations of the variables (for example, see page 8 of Imbens2004). This non-invariance concern does not apply to quantile independence restrictions, or the stronger assumption of full independence.
A key point of our paper is that the choice of any assumption weaker than full independence depends on the form of selection on unobservables one wishes to allow. We have emphasized this by characterizing the form of selection on unobservables allowed by quantile independence conditions. In principle, a similar analysis can be done for other kinds of exogeneity assumptions, like zero correlation, mean independence, or conditional symmetry.
In this paper we emphasize the interpretation of a set of quantile independence conditions which are strictly weaker than full independence. As the most common case, a large literature has studied the identifying power of a single quantile independence condition. For example, Manski1988 performs an identification analysis and notes that “If, in fact, other quantiles are also independent of $x$, this information is ignored.” To give an idea of the breadth of models where quantile independence is used, we review a subset of papers which perform similar identification analyses; we omit papers which use quantile independence conditions only as a characterization of a full independence assumption.\footnote{In particular, our results are not relevant if one is interested solely in descriptive quantile regressions, rather than causal effects or structural functions. Sasaki2015 analyzes the relationship between quantile regressions and structural functions in detail.} We also omit papers primarily on estimation theory. See QRhandbook2017 for a comprehensive overview of quantile methods.
Settings where quantile independence is used as an identifying assumption include: binary response models with interval measured regressors (ManskiTamer2002), discrete response models with exogenous regressors (Manski1985, Torgovitsky2015), discrete response models with endogenous regressors (BlundellPowell2004, Chesher2010), discrete games (Tang2010, WanXu2014, Kline2015), IV quantile treatment effect models (ChernozhukovHansen2005, Chesher2007worldCongress, ChernozhukovHansenWuthrich2017), triangular nonseparable models (Chesher2003,Chesher2005,Chesher2007,Chesher2007worldCongress), generalized instrumental variable models (ChesherRosen2017), panel data models (Koenker2004, GalvaoKato2017), censored regression models (Powell1984,Powell1986, HonoreKhanPowell2002, HongTamer2003), social interaction models (BrockDurlauf2007), bargaining models (MerloTang2012), and transformation models (Khan2001).
One of our main results is that quantile independence assumptions impose a non-monotonicity condition on treatment selection. In contrast, monotonicity conditions of various kinds are often viewed as plausible, and are widely used throughout the econometrics literature. These assumptions include monotonicity of treatment on potential outcomes (Matzkin1994, Manski1997, AltonjiMatzkin2005), monotonicity of unobservables on potential outcomes (Matzkin2003, Chesher2003,Chesher2005), monotonicity of an instrument on potential treatment (ImbensAngrist1994), monotonicity of unobservables on potential treatment (ImbensNewey2009), mean potential outcomes conditional on an instrument are monotonic in the instrument (ManskiPepper2000,ManskiPepper2009), quantiles of potential outcomes conditional on an instrument are monotonic in the instrument (Giustinelli2011), monotonicity of potential outcomes in actions of other players in a game (KlineTamer2012, Lazzati2015), and many others. See ChetverikovSantosShaikh2017 for a recent survey of the econometrics literature on shape restrictions, including monotonicity. When applying the main result in our paper to the treatment effects model, our non-monotonicity result concerns the relationship between a realized treatment and an unobservable, like a potential outcome.
In this section, we present our main characterization result. Let $X$ be an observable random variable and $U$ an unobservable random variable. For example, in section (ref) we study a treatment effects model where $X$ is a binary treatment and $U$ is an unobserved potential outcome. In this section, however, we generally remain agnostic as to the interpretation of these variables. Finally, note that all results in this section continue to hold if one conditions on an additional vector of observed covariates $W$, as is typically the case in empirical applications.
Quantile independence of $U$ from $X$ is based on the well known result that statistical independence between $U$ and $X$ is equivalent to \[ Q_{U \mid X}(\tau \mid x) = Q_U(\tau) \] for all $\tau \in (0,1)$ and all $x \in \operatorname*{supp}(X)$. Existing research typically focuses on two extreme assumptions: a single quantile independence condition holds, such as $Q_{U \mid X}(0.5 \mid x) = Q_U(0.5)$ for all $x \in \operatorname*{supp}(X)$, or all quantile independence conditions hold (statistical independence). We study a class of assumptions which includes both of these cases. It is often more natural to work with cdfs. Say $U$ is $\tau$-cdf independent of $X$ if
holds for $u=\tau$ and for all $x \in \operatorname*{supp}(X)$. This motivates the following definition.\footnote{See BelloniChenChernozhukov2017 and ZhuZhangXu2017 for similar generalizations of quantile independence.}
We assume that $U$ is continuously distributed throughout this paper. In this case, we can without loss of generality normalize its distribution to be uniform on $[0,1]$. This follows since $U$ is $\tau$-cdf independent of $X$ if and only if $F_U(U)$ is $F_U(\tau)$-cdf independent of $X$, for continuously distributed $U$. For the normalized variable $F_U(U)$, equation $\eqref{cdfIndependence}$ is nontrivial only for $\tau \in (0,1)$, and hence it suffices to let $\mathcal{T}$ be a subset of $(0,1)$.
We now focus on the binary treatment case, $X \in \{ 0,1 \}$. We extend our results to the discrete and continuous $X$ case in section (ref). Recall that the dependence structure between $X$ and $U$ is fully characterized by the latent propensity score \[ p(u) = \ensuremath{\mathbb{P}}(X=1 \mid U=u). \] Full statistical independence, $X \protect\mathpalette{\protect\independenT}{\perp} U$, is equivalent to this propensity score being constant: \[ p(u) = \ensuremath{\mathbb{P}}(X=1) \] for almost all $u \in \operatorname*{supp}(U)$. Consequently, any non-constant propensity score represents a deviation from full independence. $\mathcal{T}$-independence restricts the form of these deviations. The following theorem characterizes the set of propensity scores consistent with $\mathcal{T}$-independence.
The proof, along with all others, is in appendix (ref). Theorem (ref) says that $\mathcal{T}$-independence holds if and only if for every interval with endpoints in $\mathcal{T} \cup \{ 0,1 \}$ the average value of the propensity score over that interval equals the overall average of the propensity score, since \[ \int_0^1 p(u) \; du = \ensuremath{\mathbb{P}}(X=1). \] This overall average is just the unconditional probability of being treated.
Given our assumption that $U$ is continuously distributed, specifying $U \sim \text{Unif}[0,1]$ is a normalization which simply rescales the latent propensity score's domain. If $V$ is our original continuously distributed variable and $U \equiv F_V(V)$ is our scaled variable, then the constraint (ref) on $p(u)$ can be translated into a constraint on the original latent propensity score via the equation $\ensuremath{\mathbb{P}}(X=1 \mid V=v) = p(F_V(v))$ for almost all $v \in \operatorname*{supp}(V)$.
To illustrate theorem (ref), suppose $\mathcal{T} = \{ 0.5 \}$ and $\ensuremath{\mathbb{P}}(X=1) = 0.5$. Here we have just a single nontrivial cdf independence condition, median independence. Figure (ref) plots three different propensity scores which are consistent with $\mathcal{T}$-independence under this choice of $\mathcal{T}$; that is, which are consistent with median independence. This figure illustrates several features of such propensity scores: The value of $p(u)$ may vary over the entire range $[0,1]$. $p$ does not need to be symmetric about $u = 0.5$, nor does it need to be continuous. Finally, as suggested by the pictures, $p$ must actually be nonmonotonic; we show this in corollary (ref) next.
Corollary (ref) shows that a non-constant propensity score must be non-monotonic if it is to satisfy a $\tau$-cdf independence condition. This result can be extended as follows. Say that a function $f$ changes direction at least $K$ times if there exists a partition of its domain into $K$ intervals such that $f$ is not monotonic on each interval.
This result essentially says that such propensity scores must oscillate up and down at least $K$ times (we assume $p$ does not have removable discontinuities to rule out trivial direction changes). For example, as in figure (ref), suppose we continue to have $\ensuremath{\mathbb{P}}(X=1) = 0.5$ but we add a few more isolated $\tau$'s to $\mathcal{T}$. Figure (ref) shows several propensity scores consistent with $\mathcal{T}$-independence for larger choices of $\mathcal{T}$. Consider the figure on the left, with $\mathcal{T} = \{ 0.25, 0.5, 0.75 \}$. Partition $[0,1] = [0,0.4) \cup [0.4,0.6) \cup [0.6,1]$. Then $p$ is not monotonic over each partition set, and each partition set contains one element of $\mathcal{T}$: $0.25 \in [0,0.4)$, $0.5 \in [0.4,0.6)$, and $0.75 \in [0.5,1]$. There are $K=3$ partition sets, and hence the corollary says $p$ must change direction at least 3 times. We see this in the figure since there are 3 interior local extrema. A similar analysis holds for the figure on the right. Overall, these triangular and sawtooth propensity scores illustrate the oscillation required by corollary (ref).
One final feature we document is that as long as there is some interval which is not in $\mathcal{T}$ then there is a propensity score which takes the most extreme values possible, 0 and 1.
Like quantile independence, mean independence is commonly used to weaken statistical independence. For example, HeckmanIchimuraTodd1998 assume potential outcomes are mean independent of treatments, conditional on covariates. Our main result, theorem (ref), allows us to compare the kinds of constraints on selection on unobservables imposed by quantile independence with the constraints imposed by mean independence. In this subsection, we briefly explore this comparison.
Say $U$ is mean independent of $X$ if $\ensuremath{\mathbb{E}}(U \mid X=x) = \ensuremath{\mathbb{E}}(U)$ for all $x \in \operatorname*{supp}(X)$. The following result follows immediately from this definition.
In particular, for comparison with theorem (ref), if $U \sim \text{Unif}[0,1]$ then equation (ref) simplifies to \[ \int_0^1 2u \, p(u) \; du = \ensuremath{\mathbb{P}}(X=1). \] Theorem (ref) showed that quantile independence constrains the unweighted average value of the latent propensity score over certain subintervals of its domain. In contrast, proposition (ref) shows that mean independence constrains a weighted average value of the latent propensity score over its entire domain. Proposition (ref) can be extended to multi-valued and continuous $X$ similar to our analysis of quantile independence in section (ref); we omit this extension for brevity.
Although mean independence imposes different a constraint on the latent propensity score than quantile independence, it also requires non-constant latent propensity scores to be non-monotonic.
Thus far we have focused on binary $X$. In this section we extend our main characterization result (theorem (ref)) to both multi-valued discrete and continuous $X$. As in the binary $X$ case, our results show that the deviations from independence allowed by quantile independence require a kind of non-monotonic selection on unobservables.
We begin with the continuous case.
The interpretation of equation (ref) is similar to the binary $X$ case: $\mathcal{T}$-independence holds if and only if, for each possible level of treatment $x$, and for each interval with endpoints in $\mathcal{T} \cup \{ 0,1 \}$ the average value of the conditional probability of receiving treatment larger than $x$ given the unobservable equals the overall unconditional probability of receiving treatment larger than $x$. Notice that, by adding $-1$ to each side of equation (ref), this constraint can equivalently be seen as a constraint on the conditional cdf $F_{X \mid U}$.
As in the binary $X$ case, the constraint (ref) imposes a non-monotonicity condition.
For example, suppose $X$ is level of education, $x$ is completing college, and $U$ is ability. Then any nontrivial $\mathcal{T}$-independence condition implies that at some point increasing ability lowers the probability of getting more than a college education.
The monotonicity condition of corollary (ref) dates back to Tukey1958 and Lehmann1966, who give the following definition.
Thus corollary (ref) states that we cannot simultaneously have quantile independence of $U$ on $X$ and regression dependence of $X$ on $U$ (except when $X \protect\mathpalette{\protect\independenT}{\perp} U$).
LehmannRomano2005 call positive regression dependence an “intuitive meaning of positive dependence”. To support this claim, Lehmann1966 gave the following simple sufficient conditions for regression dependence: If one can write $X = \pi_0 + \pi_1 U + V$ where $\pi_0$ and $\pi_1$ are constants and $V$ is a random variable independent of $U$, then $X$ is regression dependent on $U$ if $\pi_1 \neq 0$. In particular, if $X$ and $U$ are jointly normally distributed then they are regression dependent so long as they have nonzero correlation. While these are special cases, theorem 5.2.10 on page 196 of Nelsen2006 provides a general characterization of regression dependence in terms of the copula between $X$ and $U$, when both variables are continuous. In particular, if $C_{X,U}(x,u)$ is the copula for $(X,U)$, $X$ is regression dependent on $U$ if and only if $C_{X,U}(x,\cdot)$ is concave for any $x \in [0,1]$.
Regression dependence is also known as stochastic monotonicity, since it is equivalent to the set of cdfs $\{ F_{X \mid U}(\cdot \mid u) : u \in \operatorname*{supp}(U) \}$ being either increasing or decreasing in the first order stochastic dominance ordering. A large literature studies tests of stochastic monotonicity. For example, LeeLintonWhang2009 state that stochastic monotonicity is “of interest in many applications in economics”, and provide many references which use such monotonicity assumptions. See DelgadoEscanciano2012, HsuLiuShi2016, and Seo2017 for further work on testing stochastic monotonicity. In a different application of stochastic monotonicity, BlundellEtAl2007 study the classic problem of identifying the distribution of potential wages, given that wages are only observed for workers. Following ManskiPepper2000, they argue that stochastic monotonicity assumptions are often plausible. They specifically consider stochastic monotonicity of wages on labor force participation status, as well as stochastic monotonicity of wages on an instrument. They furthermore provide a detailed analysis of when stochastic monotonicity assumptions may not be plausible.
Corollary (ref) shows that any assumption of $\mathcal{T}$-independence of $U$ on $X$ rules out stochastic monotonicity of $X$ on $U$. Thus, if one wants to allow for a class of deviations from independence which includes stochastically monotonic selection, assumptions of quantile independence of $U$ on $X$ should not be used. Conversely, if one makes a quantile independence assumption of $U$ on $X$, one should argue why stochastically non-monotonic selection models are the deviations of interest. We discuss these issues further in section (ref).
The following theorem extends our characterization results to the discrete $X$ case.
This result has a similar interpretation as our previous results for binary $X$ and continuous $X$. First, we have the following corollary.
The interpretation is analogous to corollary (ref). Second, all of the interpretations given in section (ref) apply to the probabilities $\ensuremath{\mathbb{P}}(X = x_k \mid U=u)$ for $k \in \{1,\ldots,K \}$. In particular, these conditional probabilities must be non-monotone. This result is primarily relevant for the lowest treatment level ($k=1$) and the highest treatment level ($k=K$), since non-monotonicity of the middle probabilities would be implied, for example, by a simple ordered threshold crossing model, like $X = x_k$ if $\alpha_k \leq U \leq \alpha_{k+1}$ for constants $\alpha_k \leq \alpha_{k+1}$, $k \in \{ 1,\ldots,K \}$.
Many econometric models obtain point identification via quantile independence restrictions, rather than statistical independence. These results are often motivated solely by the fact that quantile independence is weaker than statistical independence, and hence results using only quantile independence are `more robust' than those using statistical independence. As we have emphasized, any deviation from statistical independence of $U$ and $X$ allows for certain forms of treatment selection, in the sense that the distribution of $X \mid U=u$ depends nontrivially on $u$. Thus the choice of an assumption weaker than statistical independence depends on the class of deviations one wishes to be robust against. Since this class is often not explicitly specified, we refer to such deviations as latent selection models. Our main results in sections (ref) and (ref) characterize the set of latent selection models allowed by quantile independence restrictions.
In this section, we study two standard econometric selection models. We discuss different assumptions on the economic primitives which lead these models to be either consistent or inconsistent with quantile independence restrictions. We only consider single-agent models, but similar analyses can likely be done for multi-agent models.
We first consider a selection model discussed by ImbensNewey2009.\footnote{Pakes1994 and BlundellMatzkin2014 give additional examples of economic selection models which yield strict monotonicity in a first stage unobservable.} Consider a population of people deciding how much education to obtain. Let $Y$ denote earnings, $X$ the chosen level of education, and $U$ ability. Let $m(x,u)$ denote the earnings production function. Hence earnings, education, and ability jointly satisfy \[ Y = m(X,U). \] Let $c(x,z)$ denote the cost of obtaining education level $x$. $Z$ denotes a variable which shifts cost and is known to agents. Suppose agents do not perfectly know their own ability, but instead observe a noisy signal $V$ of $U$. Agents know the joint distribution of $(U,V)$. Suppose agents choose $X$ to maximize expected earnings, minus costs. Thus agents solve the problem
The following proposition is a variation of a result stated by ImbensNewey2009.
Assumption 1 constrains the earnings production function. Higher education increases earnings, with diminishing marginal returns. Higher ability also increases earnings. Importantly, earnings and ability are complementary. As ImbensNewey2009 mention, a Cobb-Douglas production function satisfies these assumptions. Assumption 2 constrains the cost function. Cost is increasing in earnings with increasing marginal cost. Assumption 3 implies that $Z$ has no information about agents' true ability $U$. Assumption 4 formalizes the idea that $V$ is a signal of $U$. See Milgrom1981 for the definition and further discussion of the strict MLRP. The strict MLRP is also sometimes called strict affiliation. AtheyHaile2002,AtheyHaile2007 and PinkseTan2005 discuss strict affiliation in the context of auction models. The MLRP, and hence the strict MLRP, implies that $V$ is positive regression dependent on $U$.
Proposition (ref) gives conditions under which the treatment selection equation has the form \[ X = h(Z,V) \] where $h(z,\cdot)$ is strictly increasing for each $z \in \operatorname*{supp}(Z)$. This is a common restriction imposed in the control function literature. The following proposition shows that this monotonicity restriction combined with regression dependence of the signal $V$ on ability $U$ implies regression dependence of chosen education $X$ on ability $U$.
This corollary is perhaps not surprising, since one would not typically consider $X$ to be `exogenous' in the model above. Indeed, ImbensNewey2009 go on to assume that $Z$ is observable and then use its variation to identify treatment effects. To reiterate our previous points, however: A quantile independence assumption allows for selection on unobservables, since it is weaker than full independence. Proposition (ref) shows that the form of this allowed selection is not compatible with the selection model described above. On the other hand, if one of the assumptions of proposition (ref) fails, then $X$ might not be regression dependent on $U$, and hence a quantile independence condition might hold. In particular, the assumption that $V$ and $U$ satisfy the strict MLRP (which implies regression dependence of $V$ on $U$) could perhaps be dropped. Researchers using quantile independence assumptions should argue why the class of selection models compatible with the quantile independence conditions---as specified in our characterization theorems---are the deviations of interest.
Let $X \in \{ 0,1 \}$ be a binary treatment and $Y_1$ and $Y_0$ denote potential outcomes. We study identification of this model in section (ref). Here we study the class of latent selection models consistent with quantile independence. Suppose agents choose treatment to maximize their outcome:
This is the classical Roy model (see HeckmanVytlacil2007). Suppose we are interested in identifying treatment on the treated parameters. Then identification depends on our assumptions about the stochastic relationship between $X$ and $Y_0$. In particular, one might consider assuming that some quantile of $Y_0$ is independent of $X$. As we have discussed, such an assumption constrains the latent selection model of $X$ given $Y_0$. Specifically, consider the latent propensity score,
The second line follows by our Roy model treatment choice assumption. Thus regression dependence of $Y_1$ on $Y_0$ implies that $p$ is monotonic and hence no quantile independence conditions of $Y_0$ on $X$ can hold (except when $Y_1 \protect\mathpalette{\protect\independenT}{\perp} Y_0$ or if $X$ is degenerate, as when treatment effects $Y_1-Y_0$ are constant). In particular, any quantile independence condition of $Y_0$ on $X$ rules out bivariate normally distributed $(Y_1,Y_0)$, unless $Y_1 \protect\mathpalette{\protect\independenT}{\perp} Y_0$.
There are, however, joint distributions of $(Y_1,Y_0)$ such that $Y_1$ is not regression dependent on $Y_0$. For example, let $Y_0 \sim \ensuremath{\mathcal{N}}(0,1)$ and $Y_1 = Y_0 + \mu(Y_0) - \varepsilon$ where $\mu$ is a deterministic function and $\varepsilon \sim \ensuremath{\mathcal{N}}(0,1)$, $\varepsilon \protect\mathpalette{\protect\independenT}{\perp} Y_0$. Then \[ p(y_0) = \ensuremath{\mathbb{P}}(X=1 \mid Y_0 = y_0) = \Phi[ \mu( y_0 ) ], \] where $\Phi$ is the standard normal cdf. If $\mu$ is non-monotonic then $p$ will also be non-monotonic. For this joint distribution of potential outcomes, the unit level treatment effects $Y_1-Y_0$ conditional on the baseline outcome $Y_0=y_0$ are distributed $\ensuremath{\mathcal{N}}( \mu(y_0), 1)$. Hence non-monotonicity of $\mu$ implies that the mean of this distribution of treatment effects is not monotonic. For instance, suppose the outcome is earnings and treatment is completing college. Suppose
for $-\infty < \alpha < \beta < \infty$. Then people with sufficiently small or sufficiently large earnings when they do not complete college do not benefit from completing college, on average. People with moderate earnings when they do not complete college, on the other hand, do typically benefit from completing college. Put differently, if potential earnings is an increasing deterministic function of ability, then low and high ability people do not benefit from completing college; only middle ability people do. This kind of joint distribution of potential outcomes combined with the Roy model assumption (ref) on treatment selection produce non-monotonic latent propensity scores.
While that is just one example joint distribution of $(Y_1,Y_0)$ where regression dependence fails, theorem 5.2.10 on page 196 of Nelsen2006 characterizes the set of copulas for which $Y_1$ is regression dependent on $Y_0$, when both are continuously distributed. This result therefore also tells us the set of copulas where $Y_1$ is not regression dependent on $Y_0$. Among these copulas, $\mathcal{T}$-independence of $Y_0$ from $X$ will specify a further subset of allowed dependence structures. The precise set is given by all copulas which lead to latent propensity scores that satisfy the average value constraint. One could conversely pick a set of allowed copulas and use theorem (ref) to obtain a set of quantile independence conditions that might hold. This would allow one to obtain an identified set for parameters like the average treatment effect for the treated under the given constraints on the set of copulas, although we do not pursue this here.
As in the continuous treatment case, researchers using quantile independence assumptions should argue that the set of primitives---like the joint distributions of $(Y_1,Y_0)$ in the Roy model---allowed by quantile independence are the deviations of interest. In other cases, quantile independence may be implausible, as in Heckman, Smith, and Clements' HeckmanSmithClements1997 empirical analysis of the Job Training Partnership Act (JTPA), who find that “plausible impact distributions require high measures of positive dependence [of $Y_1$ on $Y_0$]” (page 506).
Our main result in section (ref) shows that quantile independence imposes a constraint on the average value of a latent propensity score. In this section, we study the implications of this result for identification. We first use our characterization to motivate an assumption weaker than quantile independence, which we call $\mathcal{U}$-independence. The only difference between these two assumptions is that quantile independence imposes the average value constraint while $\mathcal{U}$-independence does not. Hence the difference between identified sets obtained under these two assumptions is a measure of the identifying power of the average value constraint, which is the feature of quantile independence that requires the latent propensity score to be non-monotonic.
To compute such identified sets and perform such a comparison, one must first specify an econometric model and a parameter of interest. While this can be done in many different models, we focus on a simple but important model: the standard potential outcomes model of treatment effects with a binary treatment. We adapt the analysis of MastenPoirier2017 to derive identified sets for the average treatment effect for the treated (ATT) and the quantile treatment effect for the treated (QTT) under both $\mathcal{T}$- and $\mathcal{U}$-independence. We then compare these identified sets in a numerical illustration. In this illustration, the identified sets are significantly larger under $\mathcal{U}$-independence, implying that the average value constraint has substantial identifying power.
Throughout this section, we focus on the case where $\mathcal{T}$ is an interval. In this case, we show that latent propensity scores consistent with $\mathcal{T}$-independence have two features: (a) they are flat on $\mathcal{T}$ and (b) they are non-monotonic outside the flat regions, such that the average value constraint (ref) is satisfied. We use this finding to motivate a weaker assumption which retains feature (a) but drops feature (b). We call this assumption $\mathcal{U}$-independence. While one may consider this a reasonable assumption, our primary motivation for studying $\mathcal{U}$-independence is as a tool for understanding quantile independence.
We begin with the following corollary to theorem (ref).
Corollary (ref) shows that $\mathcal{T}$-independence requires the latent propensity score to be flat on $\mathcal{T}$ and equal to the overall unconditional probability of being treated. The first property---that the latent propensity score is flat on $\mathcal{T}$---means that random assignment holds within the subpopulation of units whose unobservables are in the set $\mathcal{T}$; that is, $X \protect\mathpalette{\protect\independenT}{\perp} U \mid \{ U \in \mathcal{T} \}$. Corollary (ref) can be generalized to allow $\mathcal{T}$ to be a finite union of intervals, but we omit this for simplicity.
This corollary motivates the following definition.
Corollary (ref) shows that $\mathcal{T}$-independence implies $\mathcal{U}$-independence with $\mathcal{T}=\mathcal{U}$. The converse does not hold since $\mathcal{T}$-independence furthermore requires the average value constraint to hold, by theorem (ref). For example, figure (ref) shows two latent propensity scores. One satisfies $\mathcal{T}$-independence, but the other only satisfies $\mathcal{U}$-independence. Finally, note that $\mathcal{U}$-independence is a nontrivial assumption only when $\ensuremath{\mathbb{P}}(U \in \mathcal{U}) > 0$. Conversely, $\mathcal{T}$-independence is nontrivial even when $\mathcal{T}$ is a singleton.
The following result extends corollary (ref) to allow $X$ to be multi-valued or continuous.
Although we focus on binary $X$ for the remainder of this section, this corollary suggests that we can generalize our definition of $\mathcal{U}$-independence to allow multi-valued or continuous treatments by specifying $F_{X \mid U}(x \mid u) = F_X(x)$ for all $x \in \ensuremath{\mathbb{R}}$ and almost all $u \in \mathcal{U}$.
Let $Y_1$ and $Y_0$ denote unobserved potential outcomes. Let $X \in \{ 0,1\}$ be an observed binary treatment. We observe the scalar outcome variable
All results hold if we condition on an additional vector of observed covariates $W$. For simplicity we omit these covariates. Let $p_x = \ensuremath{\mathbb{P}}(X=x)$ for $x \in \{0,1\}$. We impose the following assumption on the joint distribution of $(Y_1,Y_0,X)$.
\setcounter{partialIndepSection}{1}
Via A(ref).(ref), we restrict attention to continuously distributed potential outcomes. A(ref).(ref) states that the unconditional and conditional supports of $Y_x$ are equal, and are a possibly infinite closed interval. This assumption implies that the endpoints $\underline{y}_x$ and $\overline{y}_x$ are point identified. We maintain A(ref).(ref) for simplicity, but it can be relaxed using similar derivations as in MastenPoirier2016. A(ref).(ref) is an overlap assumption.
We focus on two parameters: The average treatment effect for the treated, \[ \text{ATT} = \ensuremath{\mathbb{E}}(Y_1 - Y_0 \mid X=1), \] and the quantile treatment effect for the treated, \[ \text{QTT}(q) = Q_{Y_1 \mid X}(q \mid 1) - Q_{Y_0 \mid X}(q \mid 1), \] for $q \in (0,1)$. Treatment on the treated parameters are particularly simple to analyze since the distribution of $Y_1 \mid X=1$ is point identified directly from the observed distribution of $Y \mid X=1$. Hence we only need to make assumptions on the relationship between $Y_0$ and $X$. Our analysis can be extended to parameters like ATE by imposing $\mathcal{T}$- or $\mathcal{U}$-independence between $Y_1$ and $X$ as well as between $Y_0$ and $X$.
In this subsection we derive sharp bounds on the ATT and $\text{QTT}(q)$ under both $\mathcal{T}$- and $\mathcal{U}$-independence. Since $F_{Y_1 \mid X}(\cdot \mid 1) = F_{Y \mid X}(\cdot \mid 1)$, \[ \text{QTT}(q) = Q_{Y \mid X}(q \mid 1) - Q_{Y_0 \mid X}(q \mid 1). \] Hence it suffices to derive bounds on $Q_{Y_0 \mid X}(q \mid 1)$. Let $0 < a \leq b < 1$. We define the functions $\overline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1)$, $\underline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1)$, $\overline{Q}_{Y_0 \mid X}^\mathcal{U}(q \mid 1)$, and $\underline{Q}_{Y_0 \mid X}^\mathcal{U}(q \mid 1)$ in appendix (ref). These are piecewise functions which depend on $a$, $b$, $p_1$, and $Q_{Y \mid X}$.
$\mathcal{T}$-independence of $Y_0$ from $X$ with $\mathcal{T} = [Q_{Y_0}(a),Q_{Y_0}(b)]$ is equivalent to the quantile independence assumptions $Q_{Y_0 \mid X}(\tau \mid x) = Q_{Y_0}(\tau)$ for all $\tau \in [a,b]$, by A(ref). The bounds (ref) are also sharp for the function $Q_{Y_0 \mid X}(\cdot \mid 1)$ in a sense similar to that used in proposition (ref) in appendix (ref); we omit the formal statement for brevity. This functional sharpness delivers the following result.
By proposition (ref) we have that $\mathcal{T}$-independence implies that $\text{QTT}(q)$ lies in the set \[ \left[ Q_{Y \mid X}(q \mid 1) - \overline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1), \; Q_{Y \mid X}(q \mid 1) - \underline{Q}_{Y_0 \mid X}^\mathcal{T}(q \mid 1) \right] \] and that the interior of this set is sharp. Likewise for $\mathcal{U}$-independence. If $q \in \mathcal{T}$, then $\text{QTT}(q)$ is point identified under $\mathcal{T}$-independence (as is immediate from our bound expressions in appendix (ref)). This result---that a single quantile independence condition can be sufficient for point identifying a treatment effect---was shown by Chesher2003. A similar result holds in the instrumental variables model of ChernozhukovHansen2005 and the LATE model of ImbensAngrist1994. See the discussion around assumption 4 in section 1.4.3 of MellyWuthrich2017.
By corollary (ref) we have that $\mathcal{T}$-independence implies that ATT lies in the set \[ \left[ \ensuremath{\mathbb{E}}(Y \mid X=1) - \overline{\ensuremath{\mathbb{E}}}^\mathcal{T}(Y_0 \mid X=1), \; \ensuremath{\mathbb{E}}(Y \mid X=1) - \underline{\ensuremath{\mathbb{E}}}^\mathcal{T}(Y_0 \mid X=1) \right] \] and that the interior of this set is sharp. Likewise for $\mathcal{U}$-independence. Furthermore, in appendix (ref) we show that these ATT bounds have simple analytical expressions, obtained from integrating our closed form expressions for the bounds on $Q_{Y_0 \mid X}(q \mid 1)$.
By corollary (ref), $\mathcal{T}$-independence implies $\mathcal{U}$-independence for $\mathcal{T} = \mathcal{U}$. Hence identified sets using $\mathcal{T}$-independence must necessarily be weakly contained within identified sets using only $\mathcal{U}$-independence, when $\mathcal{T} = \mathcal{U}$. In this subsection, we use a numerical illustration to explore the magnitude of this difference. Since $\mathcal{T}$-independence is simply $\mathcal{U}$-independence combined with the average value constraint, the size difference between these identified sets tells us the identifying power of the average value constraint.
For $x \in \{0,1 \}$, suppose the density of $Y \mid X = x$ is \[ f_{Y \mid X}(y \mid x) = \frac{1}{\gamma x + 1} \phi_{[-4,4]} \left( \frac{y - \pi x}{\gamma x + 1} \right) \] where $\phi_{[-4,4]}$ is the pdf for the truncated standard normal on $[-4,4]$. $X$ is binary with $\ensuremath{\mathbb{P}}(X=1) = 0.5$. Let $\gamma = 0.1$ and $\pi = 1$.
Under the full independence assumption $Y_0 \protect\mathpalette{\protect\independenT}{\perp} X$, this dgp implies that treatment effects are heterogeneous, with an average treatment effect for the treated of $\text{ATT} = \pi = 1$. The quantile treatment effect for the treated at $q = 0.5$ also equals $\pi = 1$ under full independence. We continuously relax full independence by considering $\mathcal{T}$- and $\mathcal{U}$-independence with the choice $\mathcal{T} = \mathcal{U} = [\delta,1-\delta]$ for $\delta \in [0,0.5]$. For $\delta = 0$, this choice corresponds to full independence under both classes of assumptions. For $\delta = 0.5$, this choice corresponds to median independence for $\mathcal{T}$-independence, and no assumptions for $\mathcal{U}$-independence. Values of $\delta$ between 0 and $0.5$ yield partial independence between $Y_0$ and $X$ for both classes of assumptions.
Figure (ref) shows identified sets for both ATT and $\text{QTT}(0.5)$ as $\delta$ varies from $0$ to $0.5$. First consider the plot on the left, which shows the $\text{QTT}(0.5)$ bounds. The dashed lines are the identified sets under $\mathcal{T}$-independence. Since median independence of $Y_0$ from $X$ is sufficient to point identify the conditional median $Q_{Y_0 \mid X}(0.5 \mid 1)$, median independence is also sufficient to point identify the QTT at $0.5$. Hence the identified set is a singleton for all $\delta \in [0,0.5]$. Next consider the solid lines. These are the identified sets under $\mathcal{U}$-independence. When $\delta = 0.5$, $\mathcal{U}$-independence does not impose any constraints on the model, and hence we obtain the no-assumptions bounds, which are quite wide: $[-3,5]$. If we decrease $\delta$ a small amount, thus making the $\mathcal{U}$-independence constraint nontrivial, the identified set does not change. In fact, we can impose random assignment for the middle 50% of units (i.e., $\mathcal{U} = [0.25,0.75]$, which is $\delta = 0.25$) and still we only obtain the no-assumptions bounds. Consequently, for intervals $\mathcal{T} \subseteq [0.25,0.75]$, the point identifying power of $\mathcal{T}$-independence is due solely to the constraint it imposes on the average value of the latent propensity score outside the interval $\mathcal{T}$, rather than the constraint that random assignment holds for units in the middle of the distribution of $Y_0$.
Next consider the plot on the right of figure (ref), which shows the ATT bounds. The dashed lines are the identified sets under $\mathcal{T}$-independence. The ATT is no longer point identified under median independence, or any set $\mathcal{T} \subsetneq (0,1)$ of quantile independence conditions; that is, the ATT is partially identified for all $\delta > 0$. Nonetheless, even median independence alone has substantial identifying power: For $\delta = 0.5$, the identified set under median independence is $[-1,3]$, whereas the no-assumptions bounds are $[-3,5]$. Thus the length of the bounds has been cut in half. For $\delta > 0$, $\mathcal{U}$-independence has non-trivial identifying power, as shown in the solid lines. However, comparing the length of these bounds to the length to the $\mathcal{T}$-independence bounds, we see that imposing the average value constraint outside the interval $[\delta,1-\delta]$ again has substantial identifying power: the $\mathcal{T}$-independence bounds are anywhere from 50% ($\delta = 0.5)$ to almost 100% (arbitrarily small $\delta$) smaller than the $\mathcal{U}$-independence bounds. That is, the difference in lengths increases as we get closer to independence (as $\delta$ gets smaller). Thus conclusions about ATT are substantially more sensitive to small deviations from independence which do not impose the average value constraint, compared with small deviations which do impose that constraint.
In this paper we studied the interpretation of quantile independence of an unobserved variable $U$ from an observed variable $X$. We considered binary, discrete, and continuous $X$. We characterized sets of such quantile independence assumptions in terms of their constraints on the distribution of $X$ given $U$. This characterization shows that quantile independence requires non-monotonic treatment selection. For example, if any quantile independence conditions on $U \mid X$ hold then for (a) binary $X$ the probability of receiving treatment given $U=u$ must be non-monotonic in $u$, while for (b) continuous $X$ the distribution of $X \mid U=u$ cannot be stochastically monotonic in $u$. Moreover, in a numerical illustration we show that the average value constraint (which imposes this non-monotonicity) has substantial identifying power, by comparing quantile independence with a weaker version we call $\mathcal{U}$-independence.
Any class of deviations from statistical independence of $U$ and $X$ allows for certain forms of treatment selection, in the sense that the distribution of $X \mid U=u$ depends nontrivially on $u$. Thus the choice of such deviations from independence should be driven by the class of desired selection rules one wishes to allow for. For quantile independence, we characterized this class of selection rules. This class of selection rules may be of interest in some empirical applications, but not in others. Either way, researchers should justify their choice. If the class of selection rules allowed by quantile independence is deemed undesirable, several alternatives exist. These include $\mathcal{U}$-independence as defined in this paper, and $c$-dependence as defined in MastenPoirier2017. $c$-dependence constrains the probability of receiving treatment given one's unobservables to be not too far from the overall probability of receiving treatment, and allows for monotonic selection.
In section (ref) we considered several standard selection models. We linked their structure to the presence or absence of monotonic treatment selection, and therefore to the plausibility of quantile independence assumptions. It would be helpful to perform a more extensive analysis for additional models. In particular, in section (ref) we only considered the relationship between a treatment variable and an unobservable. Many models instead use quantile independence to relax statistical independence between an instrument and an unobservable. Consequently, these models allow instruments to be selected on the unobservables. Our main results in sections (ref) and (ref) characterize the kinds of distributions of instruments given unobservables allowed by quantile independence. Detailing how these distributions relate to economic models of instrument selection would help researchers assess the plausibility of quantile independence assumptions involving instruments.
In section (ref), we studied the identifying power of the average value constraint in the standard potential outcomes model with a binary treatment. It would be helpful to perform a similar analysis for other parameters in that model, like the ATE or QTE, and also for different models altogether. In particular, quantile independence assumptions are widely used in discrete response models (following Manski1975,Manski1985). While our main results in sections (ref) and (ref) already apply to the interpretation of quantile independence in these models, an identification analysis analogous to that in section (ref) would explain the importance of the average value constraint---which requires non-monotonic treatment selection---in obtaining point identification.
\singlespacing