Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
30,808 characters · 5 sections · 34 citation commands
Self-selection into treatment is a common challenge in causal inference. One approach, pioneered by heckman1976common, is to impose a model on the selection process. Another approach is to invoke the assumptions of imbens1994identification and use instrumental variables to identify the local average treatment effect (LATE). vytlacil2002independence shows that the two approaches are equivalent, even though the LATE approach does not provide an explicit model of the selection process. Specifically, vytlacil2002independence finds that the monotonicity and independence conditions imposed in imbens1994identification together imply a nonparametric binary choice model, in which the instrument and the unobserved heterogeneity are additively separable in the latent index. When conditioning covariates are included, say, for instrument validity, the selection model representation can be established on a given value of covariates. Namely, holding fixed the value of covariates, imposing a (nonparametric) selection model is no stronger than imposing the LATE assumptions.
However, in most empirical settings, a fully nonparametric analysis, conditioning on each value of the covariates, is prohibitively data-demanding. It is typical to pool observations with different characteristics and to incorporate covariates in the selection model (for example, carneiro2011estimating and cornelissen2018benefits). These empirical works conduct global analysis that explicitly models covariates while the theoretical analysis of vytlacil2002independence is local in the sense that covariates are fixed at a constant level.\footnote{The terminology “local” also appears in the work by dahl2020nevertoolate with a different meaning. They consider the weakening of the LATE assumption based on the outcome distributions rather than the covariates. } This paper aims at filling this gap.\footnote{In a recent paper, kline2019heckits compares the selection model and the LATE approach when covariates are present. While they have established that, in the absence of covariates, the selection model and the LATE approach will yield numerically identical estimates, they also point out that their equivalence result does not hold when the covariates are introduced at least in the way covariates are usually modeled in empirical works. This paper addresses the same issue by characterizing the set of selection models that are equivalent to the LATE model when covariates are present.}
Our result extends the representation in vytlacil2002independence to settings where the covariates are not held fixed. We show that the conditional LATE (CLATE) model abadie2003semiparametric has a threshold-crossing representation in which the instruments are separated from the covariates in the latent index. Loosely speaking, the separability is a result of the monotonicity condition in the CLATE model that requires the instruments to affect the potential treatment status in the same direction for all individuals.\footnote{heckman2005structural suggests renaming the monotonicity condition as uniformity because “it is a condition across people than the shape of a function for a particular person.” } In particular, the direction of the monotonicity condition is the same across all values of the covariates, and thus a latent index representation that separates out the covariates is appropriate in this case.
The separability result implies that it is possible to uniformly rank the instruments' values by the propensity score across the covariates' values. Such a ranking has two practical usages. First, it can be used as a single index to construct testable implications for the CLATE model. Second, we can use the ranking to relabel the values of the instruments so that an increase in the ranking never causes individuals to drop the treatment, regardless of their covariates values. Our representation result also implies that the techniques developed in the conditional LATE (CLATE) framework, such as the identification analysis in abadie2003semiparametric, can be applied to selection models that impose separability, and vice versa. As a corollary of the representation theorem. We reformulate and refine the testable implications of heckman2005structural for the marginal treatment effect framework.
The remaining of the paper is organized as follows. In section (ref), we present the global latent index representation of the LATE model, which is our main result. Section (ref) discusses some implications of the representation. Section (ref) generalizes the result to the case of ordered-discrete choice selection models. The last section concludes.
We first introduce the CLATE model. Let the binary variable $D$ be the receipt of treatment so that $D=1$ denotes the treatment status, and $D=0$ denotes the untreated status. The potential outcomes under the treated and untreated status are denoted by $Y_1$ and $Y_0$, respectively. The actual outcome observed by the econometrician is $Y = DY_1 + (1-D)Y_0$. Let $X$ be a random vector containing variables that could potentially affect both the outcome and the treatment choice. The covariates are introduced in the model to make the validity of the instruments plausible. Denote $\mathcal{X}$ as the support of $X$. Let the random vector $Z$ be the collection of variables that affect the treatment choice $D$ but not the potential outcomes. These variables are referred to as instruments or excluded variables. Note that under this specification, $Z$ and $X$ are disjoint sets of variables. Denote $\mathcal{Z}$ as the support of $Z$. For each value $z$ of the instrument, let $D_z$ be the counterfactual treatment status if $Z$ were externally set to $z$. The realized treatment can be represented as $D = D_Z = \sum_{z \in \mathcal{Z}} \mathbf{1}\{Z=z\}D_z$.
To avoid measure-theoretic technicalities, we assume both $\mathcal{Z}$ and $\mathcal{X}$ are countable. We further assume that for any $x \in \mathcal{X}$, $P(D=1 \mid X=x) \in (0,1)$ and $P(D=1 \mid Z=z,X=x)$ is not a trivial function of $z$. This means that there exist both treated and untreated individuals, given each value of the covariates value. This assumption is also imposed in vytlacil2002independence. The assumptions of the CLATE model are listed as follows.
Assumption (ref) requires the instrument to be “as good as randomly assigned" conditional on the covariates. Assumption (ref) is the monotonicity condition that is typically required in the LATE literature. Together, the two assumptions form the CLATE framework. Note that the exclusion restrictions of the instrument on the outcome is already embedded in the notation of the potential outcomes.
We discuss the monotonicity condition in more detail. This condition is global as it requires the direction of monotonicity to be the same across different values of $x$. A weaker and conditional version of monotonicity would be to impose, for any $(z,z') \in \mathcal{Z}^2$, and for each $x$ locally, either
For any $x \in \mathcal{X}$, we can consider the individual with $P(D_z > D_{z'} \mid X=x) = 1$ as the complier and the individual with $P(D_z < D_{z'} \mid X=x) = 1$ as the defier. Then under the local monotonicity condition ((ref)), it is possible that for some $x$, there are compliers but no defier; while for other $x$, there are defiers but no complier. This notion of local monotonicity can be found, for example, in kolesar2013estimation and sloczynski2020should. However, it is important to have uniformity in the direction of monotonicity in order to obtain the separability result in the global representation.
The main result of this paper is Theorem (ref).
This representation result achieves separability between the instrument and covariates in the treatment choice process. The function $m$ ranks the values of the instrument. Moreover, this ranking is invariant to changes in the covariates and is identified up to an increasing transformation. We further explain this ranking in the next section.
The form of Equation ((ref)) is to emphasize the separation between the instrument $Z$ and the covariates $X$. Alternatively, we can define $\tilde{U} = q(X,U)$, and write the selection equation as
where $(Y_1,Y_0,\tilde{U}) \perp Z \mid X$. Representation ((ref)) is used in the proof of Corollary (ref). Note that the separation between $Z$ and $X$ in representation ((ref)) and ((ref)) holds inside the indicator function, and it does not necessarily imply that propensity score is additively separable in $Z$ and $X$. For example, consider the simple treatment selection equation $\mathbf{1}\{Z + X \geq U\}$, where $U \mid (Z,X) \sim N(0,1)$. In this case, the propensity score is equal to $\pi(z,x) = \Phi(z+x)$, with $\Phi$ being the distribution function of the standard normal distribution. This particular propensity is not additively separable between its two arguments.\footnote{That being said, non-separabilities of $Z$ and $X$ in the propensity score may lead to a contradiction to the monotonicity assumption. For example, this can happen if the marginal effect of increasing $Z$ is positive for some $(X,Z)=(x,z)$ but negative for another value $(X,Z)=(x',z)$. For example, if the propensity score is $\pi(z,x) = \Phi(zx)$, then the monotonicity assumption would be violated if $x$ can take both positive and negative values.}
The intuition of the Theorem is explained along with the following proof, where we make use of the idea presented in vytlacil2006note.
The separability property between $Z$ and $X$ in the choice equation implies a rank-invariance property of the ranking of the instrument in terms of the propensity score $$\pi(z,x) \equiv P(D=1 \mid Z=z,X=x).$$ The following corollary also discusses the identification of the function $m$ from the propensity score.
This first implication means that the function $m$ provides an observable ordering of the instrument values by their strength of pushing individuals to take up the treatment. That is, in the CLATE model, we can rank the instrument values by their effectiveness of inducing individuals into the treatment status. This ordering remains invariant under different values of $X$ because the monotonicity is assumed to be global (Assumption (ref)). However, we do not impose the “normalization” that a higher value of the instrument always leads to more treatment take-ups, so the function $m$ need not be increasing.
The second implication uses the identified $m$ to derive a set of testable implications of the CLATE model. This set of testable implications is a refinement of the testable implications of the marginal treatment effect framework derived in heckman2005structural as the role of Z is fully summarized by the function $m$.\footnote{Notice that notation “$Z$” in heckman2005structural is different from ours as it represents the joint set of the instruments and covariates. By contrast, $Z$ only contains the excluded instruments in our paper. Accordingly, the set of testable implications we derive is also stronger in that $m$ only depends on the excluded instrument, which is a result of the monotonicity condition imposed by the CLATE model.} The testable implications are also analogous to those presented in Equation (3.3) in kitagawa2015test except that, here, $m(Z)$, but not $Z$ itself, enters the conditioning set. That is, the testable implications we derived do not restrict the direction of the effect of $Z$ on the treatment take-up. The distinction appears as we do not explicitly assume that no defier exists as kitagawa2015test does. We only assume that defiers and compliers can not both exist. That is, as stated in Assumption (ref), we leave the direction of monotonicity unspecified. \footnote{If the function $m$ were known and increasing, the testable implication in this paper would essentially reduce to Equation (3.3) in kitagawa2015test, except that $Z$ can possibly be non-binary in our case. Combined with the testing procedure proposed in Section 3.1 in his paper to handle multivalued instruments, we may likewise design a test for our implication.}
This section extends the representation result in Section (ref) to incorporate multiple ordered levels of treatment. The argument follows from the equivalence results in vytlacil2006ordered. Let there be $K$ possible levels of treatment. Now the treatment $D$ takes values in an ordered set $\{1,\cdots,K\}$. The counterfactual treatment $D_z$'s are defined accordingly. The corresponding potential outcomes are denoted by $(Y_1,\cdots,Y_K)$.
The CLATE assumptions are modified to incorporate the ordered multiplicity in treatment levels. Although we have a different definition of $D$, the statement of the monotonicity condition does not change.
This is basically the conditional version of the representation result in vytlacil2006ordered. The main point is that even though the random thresholds $U_1,\cdots,U_K$ covariates with $X$, the latent index $m(Z)$ does not explicitly depend on $X$. Again, this is because Assumption (ref) requires that the direction of monotonicity has to be the same across all values of $X$.
This paper shows that the CLATE model has a latent index representation in which the instrument and the covariates are separable in the treatment choice equation. On the theoretical side, the result more rigorously links the CLATE model to the latent index representation when covariates are present. On the practical side, the result establishes conditions when methods from the two pieces of literature can be used interchangeably. For example, one can employ the nonparametric estimator in frolich2007nonparametric as robustness checks for the structural estimates in selection models. For future works, one can consider extending this result into the unordered monotonicity model heckman2018unordered.