Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
77,530 characters · 18 sections · 74 citation commands
Identifying Effects of Multivalued Treatments
\onehalfspacing
Since the seminal work of heckman1979sample, selection problems have been one of the main themes in both empirical economics and econometrics. One popular approach in the literature is to rely on instruments to uncover the patterns of the self-selection into different levels of treatments, and thereby to identify treatment effects. The main branches of this literature are the local average treatment effect (LATE) framework of LATE1994 and the local instrumental variables (LIV) framework of MIV2005.
The LATE and LIV frameworks emphasize different parameters of interest and suggest different estimation methods. However, they both focus on binary treatments, and restrict selection mechanisms to be “monotonic”. vytlacil2002independence establishes that the LATE and LIV approaches rely on the same monotonicity assumption. For binary treatment models, these approaches require that selection into treatment be governed by a single index crossing a threshold.
Many real-world selection problems are not adequately described by single-crossing models. The literature has developed ways of dealing with less restrictive models of assignment to treatment. angrist1995two analyze ordered choice models. HUV2006,HUV2008 show how (depending on restrictions and instruments) a variety of treatment effects can be identified in discrete choice models that are additively separable in instruments and errors. More recently, heckmanpinto-pdt define an “unordered monotonicity” condition that is weaker than monotonicity when treatment is multivalued. They show that given unordered monotonicity, several treatment effects can be identified.
Even the most generally applicable of these approaches can still only deal with models of treatment that are formally analogous to an additively separable discrete choice model, as proved in Section 6 of heckmanpinto-pdt. The key condition is that the data contain changes in instruments that create only {\em one-way flows\/} in or out of the treatment cells the analyst is interested in. In binary treatment models, this is exactly the meaning of monotonicity: there cannot be both compliers and defiers, so that LATE estimates the average treatment effect on compliers\footnote{deChaisemartin shows that under a weaker condition, LATE estimates the average treatment effect on a specific subset of the compliers.}. Things are somewhat more complex in multivalued treatment models. Unless selection only depends on one function of the instruments, there exist changes in instruments that generate two-way flows in and out of any treatment cell. Unordered monotonicity requires that we observe some\/ changes in instruments that only induce one way-flows.
This is still too restrictive for important applications. For instance, many transfer programs (or many educational tests) rely on several criteria and combine them in complex ways to assign agents to treatments; and agents add their own objectives and criteria to the list. An additively separable discrete choice model may not describe such a selection mechanism. To see this, start from a very simple and useful application: the double hurdle model, which treats agents only if each of {\em two\/} indices passes a threshold\footnote{See e.g. Poirier80 for a parametric version of this model.}. While this is a binary treatment model, the existence of two thresholds makes it non-monotonic: if a change in instruments increases a threshold but reduces the other, some agents move into the treatment group and some move out of it.
The double hurdle model is still unordered monotonic, as any change in instruments that moves the two thresholds in the same direction only creates one-way flows. Now let us change the structure of the model slightly: there are still two thresholds, but we only treat agents who are above one threshold and below the other. As we will see in Section (ref), {\em any\/} change in instruments that moves both thresholds generates two-way flows, and standard approaches to identification fail. This model of {\em selection with two-way flows\/} cannot be represented by a discrete choice model; it is formally equivalent to a discrete choice model with three alternatives in which the analyst only observes partitioned choices (e.g. the analyst only observes whether alternative 2 is chosen or not). Our identification results apply to this variant of the double hurdle model, and to all treatment models generated by a finite family of threshold-crossing rules. In fact, one way to describe our contribution is that it encompasses all additively separable discrete choice models in which the analyst only observes a partition of the set of alternatives.
To illustrate the applicability of our framework, assume that assignment to treatment can be described by a random utility model of choice. Now imagine that, as is common in practice, the analyst only observes choices between {\em sets\/} of treatments: e.g., various vocational programs have been aggregated into a “training” category in her dataset. Our methods allow identification of the effect of these different training programs on outcomes, provided that continuous instruments shift their mean utilities. Variables such as distance to the locations of the training centers or other components of the “full cost” of treatment could serve as instruments in this application. For another example, consider a dynamic sequence of treatments such as the curriculum of a college student or the career of a worker. This could be represented as a “decision tree” in which various threshold-crossing rules govern the path of the individual through time. Again, this type of model can be analyzed using the techniques in this paper. Here we could use measures of performances of the worker, or the grades of the student, as (quasi) continuous instruments in order to infer the effect of each of the possible paths on outcomes. We study related examples more formally in Section (ref).
Our analysis allows selection to be determined by a vector of threshold-crossing rules. Each of these rules compares a scalar unobservable to a threshold; these unobservables can be correlated with each other and with potential outcomes. We proceed in two steps. First assume that the thresholds are known to the analyst. We use their values as control variables to deal with multidimensional unobserved heterogeneity. One important difference with the unidimensional case is that in our setting LATE-type estimators can only recover a mixture of causal parameters on groups that cross different thresholds, and are therefore harder to interpret. We establish conditions under which one can identify a generalized version of the marginal treatment effects (MTE) of MIV2005, as well as the probability distribution of unobservables governing the selection mechanism, and more aggregated treatment effects such as the average treatment effect (ATE), quantile treatment effects, the average treatment effect on the treated (ATT), and the policy-relevant treatment effect (PRTE).
Since thresholds often are not known a priori, the second step requires identifying them from the data. This is highly model-specific and the family of models encompassed in this paper is too large and diverse to allow for a general result. We limit our discussion to a few applications; in particular, we provide what we believe are new identification theorems for the double-hurdle model.
We give a detailed comparison of our paper to the existing literature in Section (ref). Let us here mention a few points in which our paper differs from the literature. Unlike imbens2000role, hirano2004propensity, cattaneo2010efficient, and Yang-et-al:16, we allow for selection on unobservables. gautier-hoderlein study binary treatment when selection is driven by a rule that is linear in a vector of unobservable heterogeneity. Lewbel2016 consider a different non-monotonic rule for binary treatment to identify the average treatment effect. These two papers break monotonicity in different ways than ours. We focus on the point identification of marginal treatment effects, unlike the research on partial identification (see e.g. manski1990-AERpp, manski1997 and manski-pepper-2000). Chesher:03, HoderleinMammen2007, FHMV2008, imbens2009identification, DF2015, and Torgovitsky2015 study models with continuous endogenous regressors. Each of these papers develops identification results for various parameters of interest. Our paper complements this literature by considering multivalued (but not continuous) treatments with more general types of selection mechanisms.
HV2007-handbook and HUV2008 and more recently heckmanpinto-pdt and pinto-jmp are more closely related to our paper. But they focus on the selection induced by multinomial discrete choice models, whereas our paper allows for more general selection problems.
The paper is organized as follows. Section (ref) sets up our framework; it motivates our central assumptions by way of examples. We present and prove our identification results in Section (ref). Section (ref) applies our results to three important classes of applications, including the models mentioned in this introduction. We relate our contributions to the literature in Section (ref). Finally, Section (ref) gives the proof of the main theorem. Some further results and details of the omitted proofs are collected in Online Appendices.
We assume throughout that treatments take values in a finite set of treatments $\mathcal{K}$. This set may be naturally ordered, as with different tax rates. But it may not be, as when welfare recipients enroll in different training schemes for instance; this makes no difference to our results. We assume that treatments are exclusive. This involves no loss of generality as treatment values could easily be redefined otherwise. We denote $K=\left\vert\mathcal{K}\right\vert$ the number of treatments, and we map the set $\mathcal{K}$ into $\{0,\ldots,K-1\}$ for notational convenience.
We denote $\{Y_k: k \in \mathcal{K} \}$ the potential outcomes. Let $D_k$ be 1 if the $k$ treatment is realized and 0 otherwise. The observed outcome and treatment are $Y := \sum_{k \in \mathcal{K}} Y_k D_k$ and $D := \sum_{k \in \mathcal{K}} k D_k$, respectively.
In addition to the covariates $\bm{X}$, observed treatment $D$ and outcomes $Y$, the data contain a random vector $\bm{Z}$ that will serve as instruments. We always condition on the value of $\bm{X}$ in our analysis of identification, and thus suppress it from the notation. Observed data consist of a sample $\{(Y_i,D_i,\bm{Z}_i): i=1,\ldots, N\}$ of $(Y,D,\bm{Z})$, where $N$ is the sample size. We denote the generalized propensity scores by $P_k(\bm{Z}) := \Pr(D = k\vert\bm{Z});$ they are directly identified from the data. Our models of treatment assignment rely on functions of the instruments $Q_j(\bm{Z})$ that are a priori unknown to the econometrician and will need to be identified. We also introduce random vectors $\bm{V}$ to represent unobserved heterogeneity.
Let $G$ denote a function defined on the support $\mathcal{Y}$ of $Y$, which can be discrete, continuous, or multidimensional. We focus on identification of the conditional counterfactual expectations $E\left(G(Y_k)\vert \bm{V}=\bm{v}\right)$ and on measures of treatment effects that can be derived from them. For example, a possible object of interest is the marginal treatment effect (MTE), defined as $E\left(Y_k-Y_l\vert \bm{V}=\bm{v}\right)$. This is similar to the MTE in the binary treatment model, in that it conditions on the value of unobserved heterogeneity in treatment. One important difference is that the link between the unobserved heterogeneity vector $\bm{V}$ and the generalized propensity scores $\Pr(D=k\vert \bm{Z})$ is now more indirect.
Aggregating up would give the mean of the counterfactual outcome $G(Y_k)$ (conditional on the omitted covariates $\bm{X}$). Once we identify $E G(Y_k)$ for each $k$, we also identify the average treatment effect $E(G(Y_k) -G(Y_j))$ between any two treatments $k$ and $j$. Alternatively, if we let $G(Y_k) = {\rm 1\kern-.40em 1}(Y_k \leq y)$ for some $y$, where ${\rm 1\kern-.40em 1}(\cdot)$ is the usual indicator function, then the object of interest is the marginal distribution of $Y_k$. This leads to the identification of quantile treatment effects.
One of our aims is to relax the usual monotonicity assumption that underlies the LATE and LIV estimators. Consider the following, simple example where $K=3$, and treatment assignment is driven by a pair of random variables $V_1$ and $V_2$ whose marginal distributions are normalized to be $U[0,1].$
To take a slightly more complicated example, consider the following entry game.
These two examples motivate the weak assumption we impose on the underlying selection mechanism. In the following we use $\bm{J}$ to denote the set $\{1,\ldots,J\}.$
The threshold conditions in Assumption (ref) have the “rectangular” form $V_j<Q_j(\bm{Z})$. Appendix (ref) discusses a more general form of linear inequalities $\bm{\beta}_j\cdot \bm{V} <Q_j(\bm{Z}).$ Note that the fact that every observation belongs to one and only one treatment group imposes further constraints. We defer discussion of these constraints to section (ref), where we show how they can be used for overidentification tests.
In this notation, the validity of the instruments translates into:
To describe the class of selection mechanisms defined in Assumption (ref) more concretely, we focus on a treatment value $k$. We define $S_j(\bm{V},\bm{Q}(\bm{Z})):={\rm 1\kern-.40em 1}(V_j <Q_j(\bm{Z}))$ for $j=1,\ldots,J$. The $\sigma$-field generated by $\{E_j(\bm{V},\bm{Q}(\bm{Z})): j=1,\ldots,J\}$ is obtained by taking unions, intersections, and complements of these $E_j$ sets. These three operations correspond to taking sums, products, and differences of their indicator functions $S_j$. Therefore the function $d_k$ referred to in Assumption (ref).(iii) can be written as an algebraic sum of products of the $S_j$ indicator functions. Let $\mathcal{L}$ denote the set of all subsets $l=\left\{l_1,\ldots,l_{\left\vertl\right\vert}\right\}$ of $\bm{J}$. Then
where the $c^k_l$ are algebraic integers. Moreover, this decomposition is unique.
Since $d_k(\bm{V},\bm{Q}(\bm{Z}))$ depends on $\bm{V}$ and $\bm{Q}(\bm{Z})$ only through $\bm{S} := \{S_{j}(\bm{V},\bm{Q}(\bm{Z})): j \in \bm{J} \}$, it will sometimes be convenient to express $d_k$ as a function of $\bm{S}$, which we denote $\mathcal{D}_k(\bm{S}).$ For example, if $J=2$, we have $\mathcal{D}_k(\bm{S}) = c^k_{\emptyset}+c^k_{\{1\}} S_1+ c^k_{\{2\}} S_2 + c^k_{\{1,2\}} S_1 S_2$ for some algebraic integers $c^k_{\emptyset},c^k_{\{1\}}, c^k_{\{2\}},$ and $c^k_{\{1,2\}}$.
To illustrate this, let us return to Example (ref), with $J=2$ and $K=3$. For $k=0$, the selection mechanism is described by the intersection $E_1\cap E_2$, whose indicator function is $\mathcal{D}_0(\bm{S}) = S_1 S_2$. Similarly, for $k=1$ we find $\mathcal{D}_1(\bm{S})=(1-S_1) (1-S_2)$. Finally, for $k=2$ we have \[ \mathcal{D}_2(\bm{S})=S_1(1-S_2)+(1-S_1)S_2=S_1+S_2-2S_1 S_2. \]
It is useful to think of the products in (ref) as alternatives in a discrete choice model. For instance, $(1-S_1)S_2$ could be interpreted as “item 1” having negative value and “item 2” having positive value. In Example (ref), $D_2=1$ informs us that the values of item 1 and of item 2 have opposite signs. In essence, we are dealing with discrete choice models with only partially observed choices. This analogy will prove useful.
The term $l=\left\{1,\ldots,J\right\}=\bm{J}$, which corresponds to the product of all $J$ indicator functions $S_j$ in (ref), plays an important role in our analysis. We will call its $c_l$ the {\em index of the treatment}.
In Example (ref), the highest order term has a coefficient $c^k_{\{1,2\}}=-1$. With $J=2$ as in Example (ref), the only treatments with a zero index are those which depend on only one threshold: e.g. ${\rm 1\kern-.40em 1}(V_1<Q_1)$. But with three or more thresholds ($J>2$), it is not hard to generate cases in which a treatment value $k$ depends on all $J$ thresholds and still has zero index, as shown in Example (ref).
When the index is zero as in Example (ref), the indicator function of the corresponding treatment $k$ has degree strictly smaller than $J$. Since Assumption (ref) rules out the uninteresting cases when treatment $k$ occurs with probability zero or one, its indicator function cannot be constant; and its leading terms have degree $m\geq 1$. We call $m$ the {\em degree\/} of treatment $k$. In Example (ref), treatment value 0 has index 0 and degree 2.
The following lemma summarizes the discussion in Sections (ref) and (ref):
In this section we fix $\bm{x}$ in the support of $\bm{X}$ and we suppress it from the notation. All the results obtained below are local to this choice of $\bm{x}$. Global (unconditional) identification results follow immediately if our assumptions hold for almost every $\bm{x}$ in the support of $\bm{X}$.
We only treat the non-zero index in the text. We make this explicit in the following assumption.
We analyze zero-index treatments in Appendix (ref).
We require that $\bm{V}$ have full support:
Note that when $J=1$, Assumptions (ref) and (ref) define the usual threshold-crossing model that underlies the LATE and LIV approaches. However, our assumptions allow for a much richer class of selection mechanisms when $J>1$. Our Example (ref) illustrates that our “multiple thresholds model” does not impose any multidimensional extension of the monotonicity condition that is implicit with a single threshold model. Even when $K=2$, so that treatment is binary, $J$ could be larger than one. This would allow for flexible treatment assignment: just modify Example (ref) to obtain the double hurdle model \[ D={\rm 1\kern-.40em 1}\left(V_1<Q_1(\bm{Z}) \mbox{ and } V_2<Q_2(\bm{Z})\right). \]
Let $f_{\bm{V}}(\bm{v})$ denote the joint density function of $\bm{V}$ at $\bm{v} \in [0,1]^{J}$. Our identification argument relies on continuous instruments that generate enough variation in the thresholds. This motivates the following three assumptions.
For any function $\psi$ of $\bm{q}$, define “local equicontinuity at $\bm{\overline{q}}$” by the following property: for any subset $I\subset \bm{J}$, the family of functions $\bm{q}_I \mapsto \psi(\bm{q}_I, \bm{q}_{-I})$ indexed by $\bm{q}_{-I} \in [0,1]^{\left\vert\bm{J}-I\right\vert}$ is equicontinuous in a neighborhood of $\bm{\overline{q}}_I.$
Assumption (ref) allows us to differentiate the relevant expectation terms. It is fairly weak: Lipschitz-continuity for instance implies local equicontinuity\footnote{It would be easy to adapt our results to cases where, for instance, $\bm{Q}$ has discontinuities. We do not pursue it in this paper.}.
The next two assumptions apply to the functions $\bm{Q}(\bm{Z})$ and in particular to their range of variation over the support $\mathcal{Z}$ of $\bm{Z}$. The functions $\bm{Q}$ are unknown in most cases, and need to be identified; in this part of the paper we assume that they are known. We will return to identification of the $\bm{Q}$ functions in Section (ref).
Assumption (ref) ensures that we can generate any small variation in $\bm{Q}(\bm{Z})$ around $\bm{q}$ by varying the instruments around $\bm{z}.$ This makes the instruments strong enough to deal with multidimensional unobserved heterogeneity $\bm{V}$.
With $J$ thresholds, Assumption (ref) requires that $\mathcal{Q}$ contains a $J$-dimensional neighborhood of $\bm{q}$. This in turn can only happen (given Assumption (ref)) if the range of variation of the instruments $\mathcal{Z}$ contains an open subset of $\mathbb R^J.$ Having $J$-dimensional continuous variation in the instruments is crucial to our approach.
For some corollaries, we use a global version of Assumptions (ref) and (ref). To state it formally, we need one last definition.
Assumption (ref) requires both that the variation in the instruments generate all possible values of the $J$ thresholds and that Assumptions (ref) and (ref) hold everywhere. We do not need this rather stringent assumption to identify the marginal treatment effects; but it is useful to derive various parameters of interest that aggregate the marginal treatment effects.
We are now ready to prove identification of $E\left(G(Y_k)\vert \bm{V}=\bm{q}\right)$ when treatment $k$ has a non-zero index. In the following theorem, for any real-valued function $\bm{q}\mapsto h(\bm{q})$, the notation \[ T h(\bm{q})\equiv \frac{\partial^J h}{\prod_{j=1}^J\partial q_j}(\bm{q}) \] refers to the $J$-order derivative that obtains by taking derivatives of the function $h$ at $\bm{q}$ in each direction of $\bm{J}$, when this derivative exists.
For two treatment values $k$ and $\ell$, define the marginal treatment effect as
The MTE function $\bm{v} \mapsto \Delta_{\text{MTE}}^{(k, \ell)} (\bm{v})$ is the average treatment effect conditional on $\bm{V}=\bm{v}$. Since $\bm{V}$ is the vector of unobservables that determine the selection mechanism, the MTE function reveals how treatment effects vary with the unobservables governing selection. As such, it captures the effect of selection and it allows the analyst to simulate counterfactual policies. It follows from Theorem (ref) that if $k$ and $\ell$ are two treatments to which all of our assumptions apply, then we can identify the marginal treatment effect of moving between these two treatments, as well as the quantile version of this MTE. We also identify the joint density function $\bm{v} \mapsto f_{\bm{V}} (\bm{v})$, which is an object of interest since it describes the dependence among elements of $\bm{V}$. Appendix (ref) extends Theorem (ref) to zero-index treatment values; it shows that similar formul\ae\ identify marginal treatment effects averaged over the missing threshold rules.
As in MIV2005, we can identify various treatment effect parameters using Theorem (ref). The following corollary shows that one can identify the average treatment effect (ATE), the average treatment effect on the treated (ATT), and the policy relevant treatment effect (PRTE) of heckman2001policy. The PRTE measures the average effect of moving from a baseline policy to an alternative policy. To define the PRTE, consider a class of policies that change $\bm{Q}$ but that do not affect $E[G(Y_k) \vert \bm{V} =\bm{v}]$. Let $D_k^\ast$ and $Y^\ast$, respectively, denote the treatment choice indicator and the outcome under a new policy $\bm{Q}^\ast$. Define $D^\ast \equiv \sum_{k \in \mathcal{K}} k D_k^\ast$.
In many applications, the range of variation of the thresholds may be limited so that Assumption (ref) will not hold. However, it is still possible to construct bounds for the ATE, ATT and PRTE if $G(Y_k)$ is bounded. For example, consider the ATE with $G(Y_k) = {\rm 1\kern-.40em 1}(Y_k \leq y)$. As shown in the proof of Theorem (ref), we can point-identify $E[G(Y_k) \vert \bm{V} =\bm{q}]f_{\bm{V}}(\bm{q}) $ by $\left(c^k_{\bm{J}}\right)^{-1} {T E\left(G(Y)D_k\vert\bm{Q}(\bm{Z}) = \bm{q}\right)}$ for each $\bm{q} \in \mathcal{\tilde{Q}}$. In addition, we know that $G(Y_k)$ lies between 0 and 1. As in manski1990-AERpp and NBERt0252, using this fact we can bound $EG(Y_k)$ within the interval defined by \[ \frac{1}{c^k_{\bm{J}}} \int_{\bm{q} \in \mathcal{S}_{\bm{Q}(\bm{Z})}} {T E\left(G(Y)D_k\vert\bm{Q}(\bm{Z}) = \bm{q}\right)} d\bm{q} \] and \[ \frac{1}{c^k_{\bm{J}}} \int_{\bm{q} \in \mathcal{S}_{\bm{Q}(\bm{Z})}} {T E\left(G(Y)D_k\vert\bm{Q}(\bm{Z}) = \bm{q}\right)} d\bm{q} + 1-\Pr\left(\bm{Q}(\bm{Z})\in \mathcal{Q}\right); \] and without further information, these bounds are sharp.
Finally, the analyst may only have discrete-valued instruments. A recent literature on the MTE focuses on this case (with binary treatment); it relies on assumed restrictions on the shape of the MTE function (see for example BeyondLATE, NBERw22363 and MST2016). In future work it would be interesting to consider relaxing these restrictions within our framework.
So far we assumed that the functions $\{\bm{Q}_j (\bm{Z}): j=1,\ldots,J\}$ were known (see Assumption (ref)). In practice we often need to identify them from the data before applying Theorems (ref) or (ref). The most natural way to do so starts from the generalized propensity scores $\{P_k(\bm{Z}): k=0,\ldots,K-1\}$, which are identified as the conditional probabilities of treatment\footnote{It would also be possible to seek identification jointly from the generalized propensity scores and from the cross-derivatives that appear in Theorems (ref) or (ref), especially when they are over-identified. We do not pursue this here.}.
First note that by definition (and by Assumption (ref)),
Note that this is a $J$-index model. ichilee:multindex consider identification of multiple index models when the indices are specified parametrically. matzkin1993-JoE, matzkin2007-advances obtains nonparametric identification results for discrete choice models\footnote{See HV2007-handbook for an application to treatment models.}; but her results only apply to a subset of the types of selection mechanisms we consider (discrete choice models when all choices are observed). Section (ref) discusses identification of the $\bm{Q}$'s through the lens of several models.
Our framework covers a wide variety of commonly used models. For simplicity, we only illustrate its usefulness on two-threshold selection models in this section. These models generate different selection patterns. Not surprisingly, the identification conditions require somewhat stronger instruments as the number of treatment values---the information available to the analyst---decreases.
Let us return to Example (ref), in which
It is useful to start with some exclusion restrictions that help us identify $Q_1(\mathbf{Z})$ and $Q_2(\mathbf{Z})$ separately from the generalized propensity scores given in (ref). Assume that
The first condition in Assumption (ref) is just a normalization of the marginal distribution of each $V_j \in \bm{V}$. The crucial part of Assumption (ref) is in the exclusion restrictions: $Z_1$ affects $Q_1$ but not $Q_2$, and $Z_2$ affects $Q_2$ but not $Q_1$. For example, if $Q_1$ and $Q_2$ represent minimum required grades on two parts of an exam, then $Z_1$ should affect only the requirement on the first part, and $Z_2$ should only affect the second part.
Suppose that the analyst has picked a point in the partially identified $(Q_1,Q_2)$ set. Using these $Q_1$ and $Q_2$, since the indices are all nonzero ($c^0_{\bm{J}}=c^1_{\bm{J}}=1$ and $c^2_{\bm{J}}=-2$) we apply Theorem (ref) to identify the joint density by
where $k=0,1,2$.
Note that $f_{V_1,V_2}(q_1,q_2)$ is overidentified; checking equality between the right-hand sides of (ref) for $k=0,1,2$ provides a specification test\footnote{Since probabilities add up to one, only one of these equalities generates a specification test.}. Similar remarks apply to the conditional expectations $E(Y_k \vert V_1=q_1, V_2=q_2)$; and as
for each $k=0,1,2$, the identification of the marginal and average treatment effects follows immediately.
In practice, $Q_1$ and $Q_2$ are only identified up to the (restricted) additive constant $C_1^0$ in Theorem (ref)(ii). As a consequence, $f_{V_1,V_2}(q_1,q_2)$ and $E(Y_k \vert V_1=q_1, V_2=q_2)$ are only identified up to the corresponding location shift in $(q_1,q_2)$. However, it is easy to check that (ref) still yields a usable specification test.
Let us now return to the double hurdle model of the introduction, where treatment is binary and the selection mechanism is governed by
and $D= 0$ otherwise.
Both treatment values have non-zero indices: $c^1_{\bm{J}}=1$ and $c^0_{\bm{J}}=-1.$ But identification of $Q_1$ and $Q_2$, which is a premise of Theorem (ref), is far from straightforward. In fact, this case is more demanding than the selection model with two-way flows in Section (ref) since we only have two treatment values. We observe the propensity score
which we denote $H(\bm{Z})$. This is a nonparametric double index model in which both the link function $F_{V_1,V_2}$ and the indices $Q_1$ and $Q_2$ are unknown; it is clearly underidentified without stronger restrictions. matzkin1993-JoE, matzkin2007-advances considers nonparametric identification and estimation of polychotomous choice models. Our multiple hurdle model has a similar but not identical structure.
In order to identify $\bm{Q}$, we assume that there exist two instruments that are excluded from one of the thresholds. More precisely, let the vector of instruments be $\bm{Z}=(Z_1,Z_2,\bm{Z}_{-12})$, with $Z_1$ and $Z_2$ scalar; we require that
To simplify notation, we fix the value of $\bm{Z}_{-12}$ and we denote $Q_1(\bm{Z}) =G_1\left(Z_1\right)$ and $Q_2(\bm{Z})=G_2(Z_2)$, where $G_1$ and $G_2$ are two unknown functions. Note that the propensity score becomes $H(Z_1,Z_2)=F_{V_1,V_2}(G_1(Z_1),G_2(Z_2))$.
We give two identification results under these exclusion restrictions. We first build on lewbel2000 and on matzkin1993-JoE, matzkin2007-advances's results to identify $\bm{Q}$ and rely on full support restrictions (conditional on the value of $\bm{Z}_{-12}$):
While Theorem (ref) requires two continuous instruments that generate all possible values of the thresholds, various additional restrictions would relax this requirement. If for instance $G_1$ and $G_2$ were linear, we would be back to the linear multiple index model of ichilee:multindex.
Theorem (ref) provides a complementary result and is useful when the instruments have limited support. It relies on a semiparametric restriction. Remember that we normalized the marginal distributions of $V_1$ and $V_2$ to be uniform over $[0,1]$; we now assume that the codependence between $V_1$ and $V_2$ is described by a strict symmetric Archimedean copula\footnote{The class of Archimedean copulas include the Clayton, Frank, and Gumbel families among others Nelsen:2006.}:
where $\phi$ belongs to the set $\Psi$ of $C^2$, strictly decreasing, and convex functions from $[0,1]$ unto $[0,+\infty]$.
Let $H_k(z_1,z_2)$ denote the derivative of the propensity score $H(z_1,z_2)$ with respect to its $k$th argument ($k = 1, 2$) and $H_{12}(z_1,z_2)$ the second-order cross derivative of $H(z_1,z_2)$. Note that the scale of $\phi$ is not identifiable in view of (ref). Furthermore, if $\phi$ is specified nonparametrically as an element of $\Psi$, the location of $\phi$ is only identifiable when the argument of $\phi$ takes values close to 1:
Our constructive identification starts by writing
for all $(z_1,z_2)$ such that $H(z_1,z_2)=h$. Once $\phi$ is identified from (ref), then we proceed to identify $G_1$ and $G_2$. The function $G_1$, for instance, would be identified by \[ \phi(G_1(z_1))=\phi(H(z_1,z^0_2))-\phi(g^0_2) \] for a fixed $z^0_2$ and a value $g^0_2$ of $G_2(z^0_2)$.
Given more a priori restrictions on the function $\phi$, identification results can be sharper. The following example illustrates this point by taking a parametric family of $\phi$.
Once $Q_1(\bm{Z})$ and $Q_2(\bm{Z})$ are identified, then under our assumptions we identify the joint density by
and the marginal treatment effect is given by
Furthermore, it follows from Corollary (ref) that under the additional Assumption (ref), the ATE, ATT and PRTE parameters are identified as well.
To conclude our examples, let us consider a two-period model of dynamic treatment where treatment assignment $D^2$ in the second period depends on the first-period treatment $D^1$ and outcome $Y^1$: \[ D^1={\rm 1\kern-.40em 1}(V_1<Q_1(\bm{Z}^1)) \; \mbox{ and } \; D^2={\rm 1\kern-.40em 1}\left(V_2<Q_2(\bm{Z}^2, D^1, Y^1)\right). \] The analyst observes $(D^1,D^2,Y^1,Y^2,\bm{Z}^1,\bm{Z}^2)$. Theorem (ref) applies to this model, provided only that the functions $Q_1$ and $Q_2$ are identified. The identification of $Q_1$ is straightforward. To identify $Q_2$, we use the results of shaikhvytlacil2011, which considers a model similar to our second-period treatment assignment. While they stress partial identification, their Remark 2.2 (p.\ 954) gives a sufficient condition for point identification. Translated in our notation, this requires that
Assumption 1 above requires that the set of instruments in the second period has a component that does not affect treatment in the first period, and whose range of variation does not depend on the propensity score of the first period. Assumption 2 adds the requirement that the ranges of the second-period propensity scores are independent of the first-period treatment, for all values of the first-period outcome.
These assumptions require overlap between treatment branches. They would not hold, for instance, in a medical trial when patients are oriented towards completely different treatments depending on how they fare early on.
The existing literature is very large; we only discuss here the most directly relevant papers.
angrist1995two consider two-stage least-squares (TSLS) estimation of a model in which the ordered treatment takes a finite number of values, and a discrete-valued instrument is available. They show that the TSLS estimator obtained by regressing outcome $Y$ on a preestimated $E(D\vert Z)$ converges to a weighted sum of {\em average causal responses\/} under some monotonicity assumption. HUV2006, HUV2008 go beyond angrist1995two by showing how the TSLS estimate can be reinterpreted in more transparent ways in the MTE framework. They also analyze a family of discrete choice models, to which we now turn.
HUV2008 consider a multinomial discrete choice model of treatment. They posit
where the $U$'s are continuously distributed and independent of $\bm{Z}.$ Then they study the identification of marginal and local average treatment effects under assumptions that are similar to ours: continuous instruments that generate enough dimensions of variation in the thresholds.
As they note, the discrete choice model with an additive structure implicitly imposes monotonicity, in the following form: if the instruments $\bm{Z}$ change in a way that increases $R_k(\bm{Z})$ relative to all other $R_l(\bm{Z})$, then no observation with treatment value $k$ is assigned to a different treatment. We make no such assumption, as Example (ref) and Figure (ref) illustrate. Our results extend those of HUV2008 to any model with identified thresholds. We consider a discrete choice model with three alternatives as an example.
There is a growing empirical literature on multivalued unordered treatments. dahl2002 develops a semiparametric Roy model for migration across U.S. states. In his empirical work, the number of unordered treatment is 51 (50 states plus the District of Columbia) and he controls for selection bias by conditioning on migration probabilities. kirkeboen2016 use discrete instruments to obtain TSLS estimates of returns to different fields of study in postsecondary education in Norway. In their setup, the unordered treatments are different fields of study. kline2016 use data from the Head Start Impact Study to estimate a semiparametric selection model. Their model has three treatment cells: Head Start, competing preschool programs, and no preschool (that is, home care).
Broadly speaking, these papers are in the same vein as Roy models and discrete choice models. Our approach complements this literature by focusing on the role of unobserved heterogeneity and the selection mechanism.
In an important recent paper, heckmanpinto-pdt introduce a new concept of monotonicity. Their “unordered monotonicity” assumption can be rephrased in our notation in the following way. Take two values $\bm{z}$ and $\bm{z}^\prime$ of the instruments $\bm{Z}$ and any treatment value $k$.
Unordered monotonicity for treatment value $k$ requires that if some observations move out of (resp.\ into) treatment value $k$ when instruments change value from $\bm{z}$ to $\bm{z}^\prime$, then no observation can move into (resp.\ out of) treatment value $k$. For binary treatments, unordered monotonicity is equivalent to the usual monotonicity assumption: there cannot be both compliers and defiers. When $K>2$, it is weaker than ordered choice. For example, suppose that there are three options $\{0, 1, 2\}$ and that a change of instruments makes option 1 less appealing. Under ordered choice, all agents who give up option 1 must fall back on option 0, or all must fall back on option 2. Unordered monotonicity allows different agents to fall back on different options. It still rules out two-way flows, that is agents moving from option 0 or 2 into option 1.
heckmanpinto-pdt show that unordered monotonicity (for well-chosen changes in instruments) is essentially equivalent to a treatment model based on rules that are additively separable in the unobserved variables---that is, the model of section (ref). In this interpretation, changes in instruments that increase the mean utility of an alternative relative to all others are unordered monotonic for that alternative, for instance. We refer the reader to Section 6 of heckmanpinto-pdt for a more rigorous discussion, and to pinto-jmp for an application to the Moving to Opportunity program.
Unlike us, heckmanpinto-pdt do not require continuous instruments; all of their analysis is framed in terms of discrete-valued instruments and treatments. Beyond this (important) difference, unordered monotonicity clearly obeys our assumptions. On the other hand, we allow for much more general models of treatment. It would be impossible, for instance, to rewrite our Examples (ref), (ref) and (ref) so that they obey unordered monotonicity. We illustrate this point using Example (ref) below.
\begin{example*1}[continued] In Example (ref), $D=2$ iff $(V_1-Q_1(\bm{Z}))$ and $(V_2-Q_2(\bm{Z}))$ have opposite signs. Note that there are two unobserved categories within $D=2$:
Each one is unordered monotonic; but because we only observe their union, $D=2$ is not unordered monotonic---increasing $Q_1$ brings more people into $2a$ but moves some out of $2b$, so that in the end we have two-way flows, contradicting unordered monotonicity. To put it differently, the selection mechanism in Example (ref) becomes a discrete choice model when each of four alternatives $d=0,1,2a,2b$ is observed; however, we only observe whether alternative $d=0$, $d=1$ or $d = 2$ is chosen in Example (ref). This amounts to an unordered monotonic treatment that is observed through a coarser information partition; this coarsening destroys unordered monotonicity. $\qed$ \end{example*1}
It is also worth commenting on other papers that break monotonicity. gautier-hoderlein consider a triangular random coefficients model for the binary treatment case. Their model is motivated by a single agent Roy model with random coefficients. Its selection mechanism is governed by \[ D = 1\{ V_1 - Z_1 - g(Z_1,\ldots,Z_L) - \sum_{j=2}^J V_j f_j (Z_j) > 0 \}, \] where $\bm{V} = (V_1, \ldots, V_J)$ is a vector of unobserved random variables, $\bm{Z} = (Z_1,\ldots,Z_J)$ is a vector of instruments that are independent of $(Y_0,Y_1,\bm{V})$, and the functions $f_2,\ldots,f_J$ and $g$ are unknown. If we limit our attention to the case of two unobservables as in the double hurdle model, then the selection equation in gautier-hoderlein reduces to \[ D = 1\{ V_1 - Z_1 - g(Z_1,Z_2) - V_2 f_2 (Z_2) > 0 \}. \] Here changes in $Z_1$ conform to monotonicity; but changes in $Z_2$ need not.
Lewbel2016 consider a different non-monotonic selection mechanism for estimating the average treatment effect. They show that the average treatment effect is identified when a binary treatment is assigned by \[ D={\rm 1\kern-.40em 1}\left( \alpha_0\leq Z + V \leq \alpha_1 \right), \] where $V$ is an unobserved random variable; $Z$ is a continuous variable that satisfies $E(Y_j\vert V,Z) = E(Y_j\vert V)$ for $j=0,1$ and $V \perp\!\!\!\perp Z$; and $\alpha_0, \alpha_1$ are unknown parameters.
Chesher:03 develops conditions to identify derivatives of structural functions in nonseparable models by functionals of quantile regression functions. In addition, FHMV2008 consider a potential outcome model with a continuous treatment. They assume a stochastic polynomial restriction and show that the average treatment effect can be identified if a suitable control function can be constructed using instruments.
imbens2009identification also consider selection on unobservables with a continuous treatment. They assume that the treatment (more generally in their paper, an endogenous variable) is given by $D=g(Z,V)$, with $g$ increasing in a scalar unobserved $V$. They identify the average structural function as well as quantile, average, and policy effects. Other more recent identification results along this line can be found in Torgovitsky2015 and DF2015 among others. One key restriction in this group of papers is the monotonicity in the scalar $V$ in the selection equation. We do not rely on this type of restriction, but we only focus on the case of multivalued treatments. Hence, our approach and those of the papers cited in this subsection are complementary.
Finally, our approach shares some features with HoderleinMammen2007. They consider the identification of marginal effects in nonseparable models without monotonicity. They show how local average structural derivatives can be identified. Like ours, their approach relies on differentiation of observed functionals. The parameters of interest they study are quite different, however, and their selection mechanism is not as explicit as ours.
Our proof has three steps. We first write conditional moments as integrals with respect to indicator functions. Then we show that these integrals are differentiable and we compute their multidimensional derivatives. Finally, we impose Assumption (ref) and we derive the equalities in the theorem.
{\bf Step 1:}
Under the assumptions imposed in the theorem, for any $\bm{q}$ in the range of $\bm{Q}$,
where the third equality follows from Assumption (ref) and the others are obvious. As a consequence,
Let $b_k(\bm{v}) \equiv E[G(Y_k) \vert \bm{V} =\bm{v}] f_{\bm{V}}(\bm{v})$ and $B_k(\bm{q})=E[ G(Y) D_k\vert\bm{Q}(\bm{Z}) = \bm{q}].$ Then (ref) takes the form \[ B_k(\bm{q}) = \int {\rm 1\kern-.40em 1}(d_k(\bm{v},\bm{q})=1) b_k(\bm{v})d\bm{v}. \] Now recall from Lemma (ref) that the indicator function of $D=k$ is a multivariate polynomial of the indicator functions $S_j$ for $j\in\bm{J}$. Moreover, \[ S_j(\bm{V}, \bm{Q}(\bm{Z}))={\rm 1\kern-.40em 1}(V_j<Q_j(\bm{Z}))=H(Q_j(\bm{Z})-V_j), \] where $H(t)={\rm 1\kern-.40em 1}(t>0)$ is the one-dimensional Heaviside function. Therefore we can rewrite the selection of treatment $k$ as
and it follows that
{\bf Step 2:}
By Assumption (ref), the function $\bm{b}$ is locally equicontinuous; and by Assumption (ref), it is defined over an open neighborhood of $\bm{q}$. This implies that all terms in (ref) are differentiable along all dimensions of $\bm{q}.$ To see this, start with dimension $j=1$. Any term $l$ in (ref) that does not contain $1$ is constant in $q_1$ and obviously differentiable. Take any other term and rewrite it as \[ A_l(q_1)\equiv c^k_l \int_{0}^{q_1} \int \left(\prod_{j\in l} H(q_{j}-v_{j})\right) b_k(v_1,\bm{v}_{-1})d\bm{v}_{-1} dv_1, \] where $\bm{v}_{-1}$ collects all directions of $\bm{v}$ in $l-\{1\}$.
Then for any $\varepsilon\neq 0$,
Since the functions $(b_k(\cdot,\bm{v}_{-1}))$ are locally equicontinuous at $q_1$, for any $\eta>0$ we can choose $\varepsilon$ such that if $\left\vertq_1-v_1\right\vert<\varepsilon,$ \[ \left\vertb_k(q_1,\bm{v}_{-1})-b_k(v_1,\bm{v}_{-1})\right\vert<\eta; \] and since the Heaviside functions are bounded above by one, we have \[ \left\vert\frac{A_l(q_1+\varepsilon)-A_l(q_1)}{\varepsilon} -c^k_l \int \left(\prod_{j\in l-\{1\}} H(q_{j}-v_{j})\right) b_k(q_1,\bm{v}_{-1})d\bm{v}_{-1}\right\vert< \left\vertc_l\right\vert \eta. \]
This proves that $A_l$ is differentiable in $q_1$ and that its derivative with respect to $q_1$, which we denote $A^1_l$, is \[ A^1_l=c_l \int \prod_{j \in l-\{1\}} H(q_j-v_j) \; b_k(q_1,\bm{v}_{-1})d\bm{v}_{-1}. \]
But this derivative itself has the same form as $A_l$. Letting $\bm{v}_{-{1,2}}$ collect all components of $\bm{v}$ except $(q_1,q_2)$, the same argument would prove that since the functions $(b_k(\cdot,\bm{v}_{-{1,2}}))$ are locally equicontinuous at $(q_1,q_2)$, the function $A^1_l$ is differentiable with respect to $q_2$ and its derivative is \[ c^k_l \int \left(\prod_{j\in l-\{1,2\}} H(q_{j}-v_{j})\right) \; b_k(q_1,q_2,\bm{v}_{-{1,2}})d\bm{v}_{-{1,2}}. \] Continuing this argument finally gives us the cross-derivative with respect to $(\bm{q}^{l})$ as \[ c^k_l \int b_k(\bm{q}^{l},\bm{v}_{-l})d\bm{v}_{-l}, \] where $\bm{v}_{-l}$ collects all components of $\bm{v}$ whose indices are not in $l$.
{\bf Step 3:}
Lemma (ref) and Assumption (ref) also imply that the leading term in the sum $\sum_l c^l_l \prod_{j\in l} H(q_{j}-v_{j})$ is \[ c^k_{\bm{J}} \prod_{j=1}^J H(q_j-v_j). \]
Now take the $J$-order derivative of $B(\bm{q})$ with respect to all $q_j$ in turn. By Lemma (ref), the highest-degree term of $B$ in $\bm{q}$ is \[ c^k_{\bm{J}} \int \left(\prod_{j=1}^{J} H(q_j-v_j)\right) b_k(\bm{v})d\bm{v} \] as $c^k_{\bm{J}}\neq 0$ under Assumption (ref); all other terms have a smaller number of indices $j$.
This term contributes a cross-derivative \[ c^k_{\bm{J}} b_k(\bm{q}), \] and all other terms generate zero-value contributions since each of them is constant in at least one of the directions $j$.
More formally,
Given Assumptions (ref) and (ref), we can apply (ref) successively to the pair of functions \[ B_k(\bm{q})= E[G(Y)D_k\vert\bm{Q}(\bm{Z}) = \bm{q}] \ \ \text{and} \ \ b_k(\bm{v})=E[G(Y_k) \vert \bm{V} =\bm{v}] f_{\bm{V}}(\bm{v}), \] as in (ref), and to the pair of functions \[ B_k(\bm{q})= \Pr[D=k\vert\bm{Q}(\bm{Z}) = \bm{q}] \ \ \text{with} \ \ b_k(\bm{v})=f_{\bm{V}}(\bm{v}). \] The first pair gives us the second equality in the Theorem, and the second pair gives us the first equality. $\qed$
{\singlespacing }