Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,750 characters · 21 sections · 114 citation commands
Identification of multi-valued treatment effects with unobserved heterogeneity
\if0 {2pt} {2pt} \fi
Keywords: Treatment effect, unobserved heterogeneity, identification, endogeneity, instrumental variable
JEL classification: C14, C21, C26
Unobserved heterogeneity in treatment effects is an essential consideration in many empirical studies in economics. As discussed in Heckman2001, for example, economic theory and applications strongly suggest that the causal effects of treatments or policy variables differ across individuals and subpopulations with the same characteristics. Quantile treatment effects characterize the heterogeneous impacts of treatments on individuals with different levels of unobserved characteristics in terms of potential outcome quantiles. Based on instrumental variable (IV) methods, local treatment effects, as introduced by Imbens1994, are the treatment effects conditional on an unobservable subpopulation for which the instrument affects treatment states.
In this paper, we establish sufficient conditions for identifying treatment effects on continuous outcomes in endogenous and multi-valued discrete treatment settings with unobserved heterogeneity using IV methods. We use only discrete instruments for identification because instruments are discrete in many empirical applications. Discrete treatments are implicitly or explicitly multi-valued in many applications. For example, households may receive different levels of transfers in anti-poverty programs, and students who wish to attend college have multiple options for choosing a college or major. For the policy-maker, it is essential to compare such multi-valued treatment effects when determining which treatment level is appropriate.
In multi-valued treatment settings, CH2005 establish the identification of quantile treatment effects (QTEs; on the observed populations) with discrete instruments, and their identification results are testable in principle. However, as we present in the next section, it is unclear how to interpret the required numerical conditions in each empirical study economically.
The main contributions of this paper are stated as follows. We establish sufficient conditions for identifying treatment effects in multi-valued treatment settings, which are easier to interpret economically than the identification conditions of CH2005. In addition, we provide closed-form expressions of the identified treatment effects. We also establish the identification of local treatment effects in multi-valued treatment settings based on our assumptions.
To illustrate the usefulness of our results, we provide three examples based on empirical research and discuss the applicability of our identification result for these examples. The first example is the effects of choosing different fields on earnings in postsecondary education. The second example is the effects of expanding access to two-year colleges on student outcomes. The third example is the effects of relocating to low-poverty neighborhoods on the outcomes of disadvantaged families living in high-poverty neighborhoods.
We provide identification conditions that take the form of monotonicity assumptions. The monotonicity assumption, introduced by Imbens1994 for the binary treatment case, has a clear interpretation in empirical studies because they can be motivated based on behavioral assumptions. The monotonicity assumption we employ for multi-valued and unordered treatments is motivated by an unordered discrete choice model, where each individual chooses the treatment option with the highest utility. Our monotonicity assumption is related to Heckman2018's “unordered monotonicity" assumption for multi-valued and unordered treatments. However, the identification condition we provide is based on a weaker assumption that holds only on a particular subset of the pairs of values the instrument can take. As highlighted in our examples, imposing the monotonicity assumption on all the instrument pairs employed in many studies, including Heckman2018, may be too strong when the instrument is multi-valued.
When the treatment is binary, VX2017 and Wuthrich2019 show identification of the treatment effects as closed-form expressions under the instrument satisfying the monotonicity assumption. Wuthrich2019 and FVX2019 develop plug-in estimators based on the closed-form expressions of the identified treatment effects. Our identification analysis covers more general settings with multi-valued treatments and instruments, regardless of whether the treatment is ordered or unordered.
The identification approach adopted here generalizes the idea of matching two distributions, which is introduced by AtheyImbens2006 and used for identifying the treatment effects in some recent studies. For the binary treatment case, VX2017 and Wuthrich2019 establish identification with binary instruments by matching the two potential outcome distributions conditional on the same subpopulation called “compliers” under the monotonicity assumption. In continuous treatment settings, torgovitsky2015, d2015, and ishihara2017 establish identification with binary instruments employing sufficiently large support of the treatment.
In multi-valued treatment settings, however, we generally cannot match the potential outcome distributions for two treatment states in the same subpopulation. The selection mechanism becomes more complicated than in the binary treatment case, and the support of the treatment is limited compared to the continuous treatment case. We overcome this difficulty by developing systems of equations with multiple potential outcome distributions. These equations are derived from relationships between the compliers that can be motivated by the discrete choice model under our monotonicity assumption. Although the distributions are conditional on different subpopulations, we show that the simultaneous equations can be solved uniquely, and the potential outcome distributions are identified under our assumptions.
We also employ our monotonicity relationships for identifying local treatment effects in multi-valued treatment settings. We establish identification when the outcome variable is continuously distributed under our monotonicity assumption, with some additional assumptions on unobservable factors, such as the rank similarity assumption.\footnote{In Section (ref), we briefly review the identification studies of local treatment effects in multi-valued treatment settings.}
The remainder of the paper is organized as follows: in Section (ref), we introduce our basic setup and provide three real-world examples. In Section (ref), we introduce our monotonicity assumption for multi-valued and unordered treatments. In Section (ref), we establish the identification of the treatment effects using our monotonicity assumption. Section (ref) provides conclusions with brief recommendations for estimation. Proofs of the main results and some auxiliary results are presented in Appendices (ref)- (ref). Some additional discussions are given in Appendices (ref)-(ref) of the Supplemental Appendix.
In this section, we introduce our basic setup with a benchmark nonseparable model. We provide three real-world examples to show that our identification approach can be applied in well-known empirical settings.
Throughout this paper, we use the notations $F_A$, $Q_A$, and $f_A$ for the unconditional cumulative distribution function (cdf), quantile function (qf), and probability density function (pdf) of a scalar-valued random variable $A$, respectively. Similarly, for a set $\mathcal{D}$ and random vectors $B$ and $C$, $F_{A|\mathcal{D} BC}(\cdot|b,c)$, $Q_{A|\mathcal{D} BC}(\cdot|b,c)$, and $f_{A|\mathcal{D} BC}(\cdot|b,c)$ denote the conditional cdf, qf, and pdf of $A$ on $\mathcal{D}\cap\{(B,C)=(b,c)\}$, respectively; let $\mathcal{D}^\circ$ denote the interior of $\mathcal{D}$.
To introduce our basic setup, we consider the following nonseparable simultaneous equations for a continuous outcome and a multi-valued endogenous treatment:
For the outcome equation, $Y$ is the outcome, $T\in\mathcal{T}$ is the multi-valued (possibly) unordered treatment where $\mathcal{T}$ contains $k+1$ values, $X\in\mathcal{X}\subset\mathbb{R}^r$ is a vector of observed covariates, and the random vector $U$ captures unobserved heterogeneity in the effect of $T$ on $Y$. For the treatment $T$, we assume a finite collection of multiple treatment statuses (either unordered or ordered) indexed by $t\in \mathcal{T}$ where, without loss of generality, $\mathcal{T}=\{0,1,2,\ldots, k\}$. For the treatment equation, $Z\in\mathcal{Z}$ is a discrete instrument where $\mathcal{Z}$ contains at least $k+1$ values, and the random vector $V$ captures unobserved factors affecting selection into treatment. For simplicity, we suppress $X$ throughout the identification analysis. All assumptions and results can be understood as conditional on $X$. The potential outcome under each treatment level $t\in\mathcal{T}$ is $Y_t=g(t,U)$, and we assume $Y_t\in\mathcal{Y}\subset\mathbb{R}$ and $E[|Y_t|]<\infty$. The potential treatment choice if $Z$ had been externally set to $z$ is $T(z)=\rho(z,V)$. In this paper, for two different treatment levels $t$ and $t'$, we are interested in the sufficient conditions for identifying the average treatment effect (ATE): $E[Y_t] - E[Y_{t'}]$ and the quantile treatment effect (QTE): $Q_{Y_t} (\tau) - Q_{Y_{t'}} (\tau)$, where $\tau\in(0,1)$. We are also interested in the sufficient conditions for identifying the local treatment effects, and we introduce them in Section (ref). For notation simplicity, we assume that the supports of the distributions of $T$, $Y_t$, $Y$, and $Z$ are equal to $\mathcal{T}$, $\mathcal{Y}$, $\mathcal{Y}$, and $\mathcal{Z}$, respectively. The results in this paper do not rely on these restrictions.
For the outcome equation (ref), we allow $U=(U_0,\ldots,U_k)'$ to be multi-dimensional, and we assume that the potential outcome is expressed as $Y_t=g(t,U_t)$, where we define $U_t:=F_{Y_t} (Y_t)$. $U_t$ is the rank variable that characterizes heterogeneity of outcomes for individuals with the same observed characteristics by relative ranking in terms of potential outcomes. For the rank variable, we assume the “rank similarity" introduced in CH2005. The rank similarity is an assumption that weakens the “rank invariance" assumption. Rank invariance assumes that $U$ is a scalar error term and that the rank variables satisfy $U_t=U_{t'}=U$ for any $t\neq t'$. However, rank similarity allows the rank variables to deviate from a common ranking $U$.
For the treatment equation (ref), we allow $V=(V_0,\ldots,V_k)'$ to be multi-dimensional for the unordered treatment $T$, and each element $V_t$ represents unobserved individual preference heterogeneity from choosing $T=t$. The treatment decision can be explained by an unordered discrete choice model, where each individual chooses the treatment option with the highest indirect utility,
where $I_t$ is the indirect utility of choosing $T=t$.\footnote{We can justify the utility maximization model of the (potential) treatment choice as in (ref) under the rank similarity assumption with an argument similar to Example 2 of chernozhukov2013quantile, where QTE is used to examine the effects of participating in a 401(k) plan.} chesher2013instrumental also employs a similar unordered choice model where $V$ is allowed to be multi-dimensional.
When the unobserved factor $V$ is a scalar random variable, the two equations (ref) and (ref) form the triangular model of chesher2005nonparametric. chesher2005nonparametric studies interval identification of the endogenous nonseparable triangular model with a discrete ordered treatment. Our research is related to chesher2005nonparametric, but our treatment equation for unordered treatments is fundamentally different from the triangular model that captures a single source of unobserved heterogeneity.
To identify these treatment effects, it suffices to identify the conditional mean and qf of the potential outcomes. We identify these under the following set of assumptions. CH2005, VX2017, and Wuthrich2019 employ a similar set of assumptions.
Assumption (ref) (i) states an expression for the potential outcome with the rank variable and imposes continuity of the potential outcome cdf. Under Assumption (ref) (i), $F_{Y_t} (y)$ is strictly increasing in $y\in\mathcal{Y}^\circ$, and $Q_{Y_t} (\tau)$ is strictly increasing in $\tau\in(0,1)$.\footnote{Lemmas (ref) and (ref) in Appendix (ref) prove these properties.} We do not assume that $Q_{Y_t} (\cdot)$ is continuous and allow $F_{Y_t} (\cdot)$ to have flat intervals. Under Assumption (ref) (i), the rank variable $U_t$ follows the uniform distribution on $(0,1)$, and $Y_t$ and $Q_{Y_t}(U_t)$ are identically distributed. Hence, we can interpret the QTE as treatment effects on individuals with the same level of unobserved heterogeneity at some level $U_t=\tau$.\footnote{We employ a slightly different definition for the rank variable from the original definition of CH2005. CH2005 define the rank variable $U_t$ as a uniformly distributed random variable on $(0,1)$ that satisfies $Y_t=Q_{Y_t}(U_t)$, and they directly assume that $Q_{Y_t} (\tau)$ is strictly increasing in $\tau\in(0,1)$. The difference does not matter in our settings because $Y_t$ and $Q_{Y_t}(U_t)$ are identically distributed.} Assumption (ref) (ii) imposes conditional independence between the potential outcomes and the instrument. Assumption (ref) (iii) states a general selection equation where the random vector $V$ captures unobserved factors affecting selection into treatment. Assumption (ref) (iv) is the rank similarity assumption. The rank similarity is arguably strong, but this condition has essential implications for identification and is consistent with many empirical situations. Assumption (ref) (v) is assumed to simplify the proofs in the main paper, and we relax Assumption (ref) (v) in Appendix (ref). See also Remark (ref) for related discussions.
The main statistical implication of Assumption (ref) is that, for each $t\in\mathcal{T}$, the qf of $Y_t$ satisfies the following nonlinear moment equation (CH2005 Theorem 1):
where $p_t(z)$ is defined as
CH2005 show that $Q_{Y_t} (\tau )$'s are identified if the following $(k+1)\times(k+1)$ matrix $\Pi'(y_0,\ldots,y_k)$ is full rank for all the values $(y_0,\ldots,y_k)$ in a set of potential solutions to the moment equations ((ref)):
where $\{z_0,z_1\ldots,z_k\}\subset\mathcal{Z}$. The identification condition of CH2005 is, in principle, directly testable. However, in each empirical study, it is not easy to check this numerical condition, which takes the form of matrices of the outcome conditional densities. It is unclear how to interpret the requirements for the endogenous variable and the instruments implied in this condition. In this paper, we establish sufficient conditions for identifying treatment effects in multi-valued treatment settings, which are easier to interpret economically than the identification conditions of CH2005. We also provide closed-form expressions of the identified treatment effects that could be used for constructing plug-in estimators.\footnote{This idea resembles that of Das2005, who develops an estimation strategy based on the closed-form expression of the regression function with discrete endogenous treatments in the nonparametric regression model with an additive error term.}
In this section, we introduce three examples based on empirical research. Throughout this paper, we consider Example I as a running example. In Section (ref), we discuss the content of our assumptions and the applicability of our identification result for Examples II and III.
The first example is the effects of choosing different fields on earnings in postsecondary education. In postsecondary education, almost all students have to choose a field of study, and earnings differ across not only universities but also fields. kirkeboen2016field study the identification and estimation of local average treatment effects (LATEs) of choosing different fields on earnings in Norway's postsecondary educational system. kirkeboen2016field find that a centralized admission process in Norway randomizes applicants into different groups, and applicants in each group are much more likely to receive an offer for each field. Based on this process, kirkeboen2016field use the predicted offers for each field as an instrument. This process also provides information on individuals' ranking of fields, and the identification analysis of kirkeboen2016field depends on each individual's next best alternative, that is, the field one would prefer if one's preferred field would not be feasible. Our results can be applied when we can find a certain instrument that effectively randomizes students into different groups, and information on individuals' ranking of fields is not necessary for our identification approach.
The second example is the effects of expanding access to two-year colleges on student outcomes. Two-year community colleges will increase the flow of young people into higher education, and expanding access to two-year colleges is expected to positively affect educational attainment and earnings. However, higher enrollment rates in two-year colleges may adversely affect student outcomes because college applicants are discouraged from paying higher tuition and enrolling directly in four-year colleges. Mountjoy2019 and ferreyra2022labor estimate the effects of expanding access to two-year colleges in the United States and Colombia using instruments, respectively. The treatment is the college applicant's decision to start college at a two-year or four-year institution or not to enroll in college. With this multi-valued treatment, they compare the two opposing effects: the positive effect on new two-year entrants who otherwise would not have enrolled in any college, and the negative effect on two-year entrants who otherwise would have started directly at a four-year institution. For the instruments, they use the distance to the nearest college. Mountjoy2019 directly uses the distance as a continuous instrument, and ferreyra2022labor use a discrete instrument that indicates whether the nearest college is located within a certain distance radius. We employ discrete instruments and show that our identification method can be applied to discrete instruments.
The third example is Moving to Opportunity (MTO), a housing experiment implemented between 1994 and 1998. The MTO experiment was designed to evaluate the effects of relocating to low-poverty neighborhoods on the outcomes of disadvantaged families living in high-poverty urban neighborhoods in the United States. This project randomly assigned housing vouchers from the Section 8 program that could be used to subsidize housing costs. Eligible families were placed in one of the following three assignment groups: experimental group, to which Section 8 housing vouchers were assigned but restricted to use their vouchers in a low-poverty neighborhood; Section 8 group, to which regular Section 8 housing vouchers were assigned without any restriction on their place of use; or the control group to which no voucher was assigned. Impact evaluations were conducted in 2002, 2009, and 2010. See Orr2003, Sanbonmatsu2011, and Shroder2012 for the detailed information on this project. For the recent studies that find evidence of neighborhood effects on adult employment, aliprantis2019evidence estimate LATEs for moving to a higher-quality neighborhood under an ordered treatment model using neighborhood quality as an observed continuous measure of the treatment variable.
Under an unordered treatment model, Pinto2015 applies the identification results of Heckman2018 for estimating the conditional means of the potential outcomes on the compliers. We show that, under some additional assumptions on unobservable factors, such as the rank similarity assumption, the ATEs and QTEs are nonparametrically identified when the outcome variable is continuously distributed.
In this section, we introduce the monotonicity assumption we employ for the multi-valued treatment case. We illustrate that the monotonicity assumption is motivated by economic analysis using a discrete choice index model introduced in Section (ref).
The monotonicity assumption is first introduced by Imbens1994 for the binary treatment case, and Heckman2018 generalize this assumption to unordered multi-valued treatment settings. We define a binary variable $D_t:=1\{T=t\}$, where $1\{\mathcal{A}\}$ is the indicator function of a set $\mathcal{A}$ and $D_t$ is an indicator function of each treatment level. Then, the observed outcome can be represented as $Y=\sum_{t=0}^kY_tD_t$. We also define $D_t(z):=1\{T(z)=t\}$ as an indicator function of each potential treatment state if $Z$ had been externally set to $z$. We define $\mathcal{P}:=\{(z,z')\in\mathcal{Z}^2:z\neq z'\}$ as a set of pairs of different values that the instrument can take. The monotonicity assumption imposes restrictions on these pairs.
For the binary treatment case, the monotonicity assumption requires that either $T(z)\leq T({z'})$ and $P(\{T(z)=0, T(z')=1\})>0$ or $T(z)\geq T({z'})$ and $P(\{T(z)=1, T(z')=0\})>0$ hold almost surely for each $(z,z')\in\mathcal{P}$. This condition implies that either $P(\{T(z)=1, T(z')=0\})=0$ or $P(\{T(z)=0, T(z')=1\})=0$ holds for each $(z,z')\in\mathcal{P}$. Under the monotonicity assumption, individuals who change their choice respond in only one direction to a change in $Z$, and the group with positive probability is called “compliers."
Heckman2018 generalize this argument to multi-valued treatment settings. For each $z\in\mathcal{Z}$ and $t\in\mathcal{T}$, they assume the monotonicity assumption for binary treatments on each binary indicator $D_t(z)$, and either $D_t(z)\leq D_t(z')$ or $D_t(z)\geq D_t(z')$ holds almost surely for each pair $(z,z')\in\mathcal{P}$. They call this assumption “unordered monotonicity” because this condition can be assumed on the unordered treatments. However, as we see in Section (ref), imposing such conditions on all the pairs $(z,z')\in\mathcal{P}$ may be too strong when the instrument is multi-valued. Hence, we employ a weaker assumption that imposes such conditions only on a subset of $\mathcal{P}$. That subset is determined differently in each situation. Recently, a similar problem has been discussed when there are multiple instruments. MTW2019,mogstad2020policy and goff2020vector consider the binary treatment case, and Mountjoy2019 considers the multi-valued unordered treatment case. As discussed in Section (ref), Mountjoy2019 employs the same type of monotonicity assumption with multiple continuous instruments. We consider a (possibly) scalar multi-valued instrument and weaken the monotonicity assumption from another perspective.
We define the compliers for the multi-valued treatment case as $\mathcal{C}_{z,z'}^t:=\{D_t(z)=0, D_t(z')=1\}$. Our monotonicity assumption is characterized by inequalities such as $D_t(z)\leq D_t(z')$ and $D_t(z)\geq D_t(z')$. We employ the following monotonicity assumption:
We call this subset $\Lambda$ “monotonicity subset." When $T$ is binary, Assumptions (ref) (i)--(iii) are the monotonicity assumption in Imbens1994. Assumption (ref) (i) strengthens Assumption (ref) (ii) and assumes that the potential outcome and treatment are jointly independent of the instrument. Assumption (ref) (ii) assumes that monotonicity inequalities hold on a particular subset $\Lambda$ of $\mathcal{P}$. Assumption (ref) (iii) is an instrument relevance condition and assumes that the compliers always exist. Under these conditions, when monotonicity inequalities hold on $(z,z')\in\Lambda$, we exclude the cases where neither $D_t(z)<D_t(z')$ nor $D_t(z)>D_t(z')$ can happen for some $t\in\mathcal{T}$. Assumption (ref) (iv) strengthens the instrument relevance condition and assumes that the compliers are sufficiently large. VX2017 employ a similar condition for the binary treatment case. Under Assumption (ref) (iv), each conditional cdf of the potential outcome on compliers strictly increases on $\mathcal{Y}^\circ$.\footnote{We show this statement in Lemma (ref) in Appendix (ref).}
In this section, we illustrate that economic analysis implies the monotonicity assumption we employ for the multi-valued treatment case. We consider Example I and use a discrete choice index model introduced in Section (ref) to motivate the monotonicity assumption. Mountjoy2019 also employs a discrete choice index model to motivate his monotonicity assumption for continuous instruments. We take an approach similar to Mountjoy2019 for the choice index model.\footnote{Heckman2018 and Pinto2015 consider a general utility maximization problem and employ revealed preference analysis to motivate their unordered monotonicity assumption. The generalized models also imply our monotonicity assumption.}
We consider a setting where individuals choose between not taking any postsecondary education or completing some postsecondary education and choose between $k$ different fields of study labeled as $1,\ldots,k$. Let $Y$ denote observed earnings. For the treatment $T$, let $T=0$ denote not taking any postsecondary education, and for $j\in\{1,\ldots,k\}$, $T=j$ denotes completing field $j$. Suppose that individuals are randomly assigned to one of the following $(k+1)$ groups; individuals assigned to group $0$ have no cost reduction and for $j\in\{1,\ldots,k\}$, the cost of choosing field $j$ is decreased for individuals assigned to group $j$. Let the instrument $Z$ represent group assignment that takes values on $\mathcal{Z}=\{0,1,\ldots,k\}$, where $Z=j$ denotes assignment to group $j$. In Example I, Assumption (ref) (i) holds because vouchers are randomly assigned. Suppose that the treatment $T$ and the instrument $Z$ are sufficiently correlated, and we assume Assumptions (ref) (iii) and (iv) unless stated otherwise.
We first consider a case where individuals choose between three alternatives. As in Mountjoy2019, we assume additive separability of the utility functions in unobservable components for simplicity and ease of visualization. As in (ref) in Section (ref), each individual chooses the treatment option with the highest indirect utility:
where the indirect utilities for each treatment option are defined as follows: \[ I_0 = 0,\quad I_1 = V_1 - \mu_1(Z), \text{ and } I_2 = V_2 - \mu_2(Z). \] The utility of not taking any postsecondary education is normalized to zero. $V_t$ is an individual's gross utility from choosing field $t$ and represents unobserved individual preference heterogeneity. $\mu_t(Z)$ is the cost of choosing field $t$, and $I_t$ is the net utility of choosing field $t$. The potential treatments under this choice index model are expressed as follows:
Figure (ref) (a) shows how these treatment choice equations (ref)-(ref) partition the two-dimensional space of unobserved preferences $(V_1,V_2)$. Individuals who choose $T=0$ have low preferences for fields 1 and 2 relative to their costs, while those who choose $T=1$ or $T=2$ have higher preferences for their treatment choice.
For the cost functions, we can naturally assume the following relationships from the group assignment of the instrument:
Relationship (ref) holds because the cost of choosing field 1 decreases for individuals assigned to group 1. Relationship (ref) holds for a similar reason. Applying the restrictions on the cost functions (ref) and (ref) to the treatment choice equations (ref)-(ref) generates eight monotonicity inequalities summarized in Table (ref).
From Table (ref), Assumption (ref) (ii) holds for $\Lambda=\{(1,0),(2,0)\}$. For the inequalities in Table (ref), kirkeboen2016field also employ $D_1(1)\geq D_1(0)$ and $D_2(2)\geq D_2(0)$, and we obtain the other inequalities from the restrictions on the cost functions (ref) and (ref). As we discuss in Section (ref), Mountjoy2019 provides the same type of inequalities as the monotonicity inequalities for $(1,0)$ and $(2,0)$ with continuous instruments.
Figures (ref) (b)--(d) visualize the monotonicity inequalities summarized in Table (ref). Figure (ref) (b) illustrates how a shift in the instrument from $Z=0$ to $Z=1$ induces the three monotonicity inequalities for $(1,0)$. Because this shift in $Z$ decreases the cost of choosing field 1 from $\mu_1(0)$ to $\mu_1(1)$ but does not change the cost of choosing field 2, some individuals find field 1 more attractive and change their choice from either $T=0$ or $T=2$ to $T=1$, but no individuals find field 1 less attractive and leave from the $T=1$ group. Hence, the expansion of the $T=1$ region induces $D_1(1)\geq D_1(0)$ and, at the same time, the shrinkage of both the $T=0$ and $T=2$ regions induces $D_0(1)\leq D_0(0)$ and $D_2(1)\leq D_2(0)$. Analogously, Figure (ref) (c) illustrates that a shift in the instrument from $Z=0$ to $Z=2$ induces the three monotonicity inequalities for $(2,0)$.
Figure (ref) (d) illustrates that a shift in the instrument from $Z=1$ to $Z=2$ does not induce any monotonicity inequality between $D_0(1)$ and $D_0(2)$, and $(1,2)$ is not contained in the monotonicity subset. Because this shift in $Z$ increases the cost of choosing field 1 from $\mu_1(1)$ to $\mu_1(2)$ and decreases the cost of choosing field 2 from $\mu_2(1)$ to $\mu_2(2)$, some individuals who find field 1 less attractive move on to the $T=0$ group and, at the same time, some individuals who find field 2 more attractive leave from the $T=0$ group. Hence, $D_0(1)>D_0(2)$ and $D_0(1)<D_0(2)$ can both happen depending on whether more individuals are induced into or out from the $T=0$ group.
We generalize the preceding argument to general $k\in\mathcal{T}$. Similar to the case of $k=2$, a discrete choice model with $k+1$ treatment options generates the following monotonicity inequalities:
Table (ref) summarizes ((ref)). From Table (ref), Assumption (ref) (ii) holds for $\Lambda=\{(1,0),\ldots,(k,0)\}$.
In this section, we establish the identification of the potential outcome distributions and the local treatment effects using our monotonicity assumption. Before establishing our main results, we first introduce a map termed “counterfactual mapping.” Counterfactual mapping, developed by VX2017, is an essential tool for identification. We then establish the identification of counterfactual mappings.
In this section, we introduce counterfactual mapping in multi-valued treatment settings. We show that identifying the potential outcome distributions follows from identifying the counterfactual mappings. Our key identification challenge is to recover the counterfactual mappings in multi-valued treatment settings.
For $s,t\in\mathcal{T}$, define $\phi_{s,t}:\mathbb{R}\to\mathbb{R}$ as $\phi_{s,t}(y):=Q_{Y_t}(F_{Y_s}(y))$. $\phi_{s,t}$ is called “counterfactual” mapping from $Y_s$ to $Y_t$ because the potential outcomes are also called “counterfactual outcomes." VX2017 define a similar mapping for the binary treatment case. From the definition, this mapping relates the quantiles of the distribution of $Y_s$ to that of $Y_t$. Under Assumption (ref) (i), this mapping is strictly increasing on $\mathcal{Y}^\circ$, and $\phi_{s,r}=\phi_{t,r}\circ\phi_{s,t}$ holds for $s,t,r\in\mathcal{T}$. Moreover, under Assumption (ref) (v), an inverse mapping $\phi_{s,t}^{-1}$ exists on $\mathcal{Y}^\circ$, and $\phi_{t,s}(y)=\phi_{s,t}^{-1}(y)$ holds for $y\in\mathcal{Y}^\circ$. Similarly, we define the unconditional counterfactual mapping for $s,t\in\mathcal{T}$ as $\phi_{s,t}(y):=Q_{Y_t}(F_{Y_s}(y))$.
The following lemma shows that the potential outcome cdfs and means can be written as compositions of the counterfactual mappings and observable distributions. VX2017 show a similar result for the binary treatment case.
Lemma (ref) follows from the rank similarity assumption. For each $s,t\in\mathcal{T}$, the rank variables $U_s$ and $U_t$, and hence $Y_s$ and $\phi_{t,s}(Y_t)$, are identically distributed conditional on $(T,Z)=(t,z)$.
Lemma (ref) implies that for each $s\in\mathcal{T}$, $E[Y_s]$ and $Q_{Y_s} (\tau )$ for $\tau\in(0,1)$ are identified as closed-form expressions if $\phi_{s,t}$ for $t\in\mathcal{T}$ is also identified as a closed-form expression. Hence, we establish sufficient conditions to identify $\phi_{s,t}$'s and derive the closed-form expressions of $\phi_{s,t}$'s.
In this section, we introduce some basic identification results under our monotonicity assumption. We first establish the identification of the compliers under our monotonicity assumption. The following lemma shows that when $D_t(z)\leq D_t({z'})$ and $P(\mathcal{C}^t_{z,z'})>0$ hold almost surely for $(z,z')\in\mathcal{P}$ and $t\in\mathcal{T}$, each probability of $\mathcal{C}^t_{z,z'}$ and the conditional cdf of $Y_t$ given $\mathcal{C}^t_{z,z'}$ are identified as closed-form expressions. Heckman2018 show a similar result under the unordered monotonicity assumption.
With Lemma (ref) at hand, we establish the identification of the counterfactual mappings. To provide intuition, we first review the identification results for the binary treatment case by VX2017. Let the treatment be binary so that $\mathcal{T}=\{0,1\}$. Observe that
holds from the definition when $T$ is binary. Suppose Assumptions (ref) and (ref) hold with $P(D_1(z)\leq D_1({z'}))=1$ and $P(\mathcal{C}^1_{z,z'})>0$. Note that under the rank similarity assumption, the rank variables $U_0$ and $U_1$ are identically distributed conditional on $\mathcal{C}^1_{z,z'}$.\footnote{We show this statement in Lemma (ref) in Appendix (ref).} Furthermore, from the definition of the counterfactual mapping, if $y\in\mathcal{Y}^\circ$ is the $\tau\in(0,1)$ quantile of the distribution of $Y_1$, then $\phi_{1,0}(y)$ is the $\tau$ quantile of the distribution of $Y_0$. Therefore, given ((ref)), we obtain the following equation for the potential outcome conditional distributions given the compliers:
Finally, $F_{Y_0|\mathcal{C}^0_{z,z'}}$ and $F_{Y_1|\mathcal{C}^1_{z,z'}}$ are identified from Lemma (ref), and $\phi_{1,0}(y)$ for $y\in\mathcal{Y}^\circ$ is identified as $\phi_{1,0}(y)=Q_{Y_0|\mathcal{C}^0_{z,z'}}(F_{Y_1|\mathcal{C}^1_{z,z'}}(y))$ by solving ((ref)) for $\phi_{1,0}$.
In the case of multi-valued treatment settings, generally, we can compare no two treatment states on the same compliers as in ((ref)). This is because the relationships between the compliers become more complicated than ((ref)), and the compliers do not generally coincide. We overcome this difficulty by developing systems of relationships between two or more compliers for multiple treatment states that can be solved simultaneously for the counterfactual mappings.
In this section, we establish the identification of counterfactual mappings using the relationships of compliers when the treatment is multi-valued. Throughout the identification analysis, we consider Example I for notation simplicity. The following argument does not rely on the settings of Example I. We first consider a case where the treatment takes three values (then we have $\mathcal{T}=\{0,1,2\}$), and assume Assumptions (ref) and (ref) hold for the subset $\Lambda$ of $\mathcal{P}$.
We first introduce “sign treatments" that characterize the type of monotonicity relationship of each pair $(z_1,z_2)$ contained in the monotonicity subset $\Lambda$. Suppose that, for $(z_1,z_2)\in\Lambda$, there uniquely exists $t{(z_1,z_2)}\in\mathcal{T}$ such that either \[ D_{t{(z_1,z_2)}}({z_1})\geq D_{t{(z_1,z_2)}}({z_2})\text{ and } D_j({z_1})\leq D_j({z_2})\,\, \text{ for }j\in\mathcal{T}\setminus\{t{(z_1,z_2)}\} \] or \[ D_{t{(z_1,z_2)}}({z_1})\leq D_{t{(z_1,z_2)}}({z_2})\text{ and } D_j({z_1})\geq D_j({z_2})\,\, \text{ for }j\in\mathcal{T}\setminus\{t{(z_1,z_2)}\} \] holds almost surely. Then we call this $t{(z_1,z_2)}$ “sign treatment of $(z_1,z_2)$," and we state “$(z_1,z_2)$ has a sign treatment" in this paper.
When the treatment takes three values, each $(z_1,z_2)\in\Lambda$ has a sign treatment. This is because if the monotonicity inequalities of $(z_1,z_2)$ are all the same, then no compliers exist, and $\Lambda$ cannot contain $(z_1,z_2)$. Suppose that $D_j({z_1})\leq D_j({z_2})$ holds almost surely for all $j=0,1,2$. Then, $D_0({z_1})\geq D_0({z_2})$ holds almost surely because $D_1({z_1})\leq D_1({z_2})$ and $D_2({z_1})\leq D_2({z_2})$ imply $1-D_0({z_1})\leq 1-D_0({z_2})$. Hence, $D_0({z_1})= D_0({z_2})$ holds almost surely. Applying a similar argument to treatment states 1 and 2 gives $D_j({z_1})=D_j({z_2})$ for each $j\in\mathcal{T}$, which violates Assumption (ref) (iii).\footnote{When the treatment takes more than three values, each $(z_1,z_2)\in\Lambda$ may not have a sign treatment. Let $\mathcal{T}=\{0,1,2,3\}$, and suppose that the following monotonicity inequalities hold for $(z_1,z_2)$: \[ D_{0}({z_1})\geq D_{0}({z_2}),\,\, D_{1}({z_1})\geq D_{1}({z_2}),\,\, D_{2}({z_1})\leq D_{2}({z_2}),\text{ and } D_3({z_1})\leq D_3({z_2}). \] Then, from the definition, $(z_1,z_2)$ does not have a sign treatment.}
With the sign treatments, we employ the following assumption and assume that two different types of monotonicity relationships exist.
Under Assumption (ref), the monotonicity subset $\Lambda$ contains two pairs of instrument values $\lambda_1$ and $\lambda_2$ such that each pair $\lambda_i$ has a sign treatment, and the two sign treatments $t({\lambda_1})$ and $t({\lambda_2})$ are different. Then, each $t({\lambda_i})$ characterizes a type of monotonicity relationship of $\lambda_i$. Assumption (ref) holds in Example I. From Table (ref) in Section (ref), $(1,0)$ and $(2,0)$ have sign treatments $t{(1,0)}=1$ and $t{(2,0)}=2$, respectively. Then, $(1,0)$ and $(2,0)$ induce different types of monotonicity relationships.
We can interpret Assumption (ref) as an instrument relevance condition that requires monotonic correlation between each of the two endogenous variables $D_{t({\lambda_1})}$ and $D_{t({\lambda_2})}$ and pairs of instrument values $\lambda_1$ and $\lambda_2$. We illustrate this point with Example I. For $i=1,2$, whether the group assignment is $i$ or 0 produces a monotonic effect only toward the field choice $D_{t{(i,0)}}=D_i$. Compared with group 0, group $i$ additionally offers a discount only to field $i$, and only the preference for field $i$ is affected by the difference between these two group assignments.
Different types of monotonicity relationships are essential for our identification analysis. We obtain key relationships between the compliers essential for identifying the counterfactual mappings based on the two different monotonicity relationships. For Example I, the proof of Lemma (ref) in Appendix (ref) derives the following key relationships between the compliers:
Relationships (ref) and (ref) correspond to ((ref)) in the binary treatment case. Applying Lemma (ref), all the functions in (ref) and (ref) except for the counterfactual mappings are identified.
We identify the counterfactual mappings by solving (ref) and (ref) simultaneously for each $y\in\mathcal{Y}^\circ$. When we fix the value of $y$ in (ref) and (ref) at any $y=y^{f}\in\mathcal{Y}^\circ$, $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are the solutions to the following nonlinear simultaneous equations of two unknown variables $y_1$ and $y_0$:
We have two equations to solve two unknowns, and $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are identified if the solution is unique at $y^{f}\in\mathcal{Y}^\circ$. The two equations ((ref)) and ((ref)) have a unique solution for $y_1$ and $y_0$ when the conditional cdfs of the potential outcome on compliers are strictly increasing on $\mathcal{Y}^\circ$. Assumption (ref) (iv) implies the strict monotonicity of the conditional cdfs, and $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are identified as follows:\footnote{See Appendix (ref) for the derivation of (ref).}
where $\phi_{1,0}^{y_f}$ with its domain $\mathcal{Y}^f\subset\mathcal{Y}$ is defined as
and we define a function $G_{1,2}^{y^f}$ as
Other counterfactual mappings on $\mathcal{Y}^\circ$ are inversions or compositions of $\phi_{2,1}$ and $\phi_{2,0}$, and they are also identified as closed-form expressions. The following lemma shows the identification of the counterfactual mappings under Assumption (ref):
We provide intuition for the identification of the counterfactual mappings. First, the two relationships (ref) and (ref) that are key conditions for identification are based on the following two relationships between compliers:
Relationships (ref) and (ref) correspond to (ref) in the binary treatment case. We show that two monotonicity relationships for $(1,0)$ and $(2,0)$ generate (ref) and (ref).
Figure (ref) visualizes these compliers relationships (ref) and (ref) implied by the separable index model introduced in Section (ref). Figure (ref) (a) visualizes how the compliers are generated from a shift in the instrument from $Z=0$ to $Z=1$. First, because $\mathcal{C}^1_{0,1}$ compliers are driven by individuals who find field 1 more attractive and leave from either $T=0$ group or $T=2$ group, we have
Obviously, the sets $\{D_1({1})=1,D_0({0})=1\}$ and $\{D_1({1})=1,D_2({0})=1\}$ are disjoint. Second, we show that
To see this, note that $\mathcal{C}^0_{1,0}$ compliers, which consist of individuals leaving from the $T=0$ group, are entirely driven by those who find field 1 more attractive. This is because, from the restrictions on the cost functions (ref) and (ref), a shift in the instrument from $Z=0$ to $Z=1$ decreases the cost of choosing field 1 but does not change the cost of choosing field 2. Analogously, $\mathcal{C}^2_{1,0}$ compliers, which consist of individuals leaving from the $T=2$ group, are also entirely driven by those who find field 1 more attractive. Therefore, (ref) holds from ((ref)) and ((ref)). By analogous logic, as visualized in Figure (ref) (b), the compliers generated from a shift in the instrument from $Z=0$ to $Z=2$ satisfy relationship (ref).
Next, we provide an intuition for identifying $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ for each $y^{f}\in\mathcal{Y}^\circ$ as the unique solution to the two equations ((ref)) and ((ref)) of two unknowns $y_1$ and $y_0$. We compare the set of (an infinite number of) solutions to each of these equations. The two equations have a unique solution when these two sets intersect at a point. Figure (ref) visualizes the solutions to ((ref)) and ((ref)) respectively in the two-dimensional space of $(y_1,y_0)$. First, the set of solutions to ((ref)) form a strictly increasing relationship from the strict monotonicity of $F_{Y_1|\mathcal{C}^1_{0,1}}$ and $F_{Y_0|\mathcal{C}^0_{1,0}}$. To see this, suppose that both $(y^-_1,y^-_0)$ and $(y^+_1,y^+_0)$ are solutions to ((ref)). If we have $y^-_1 < y^+_1$, we also have $F_{Y_1|\mathcal{C}^1_{0,1}}(y^-_1)<F_{Y_1|\mathcal{C}^1_{0,1}}(y^+_1)$ from the strict monotonicity of $F_{Y_1|\mathcal{C}^1_{0,1}}$. Then, because $F_{Y_1|\mathcal{C}^1_{0,1}}$ and $F_{Y_0|\mathcal{C}^0_{1,0}}$ are on the left and right-hand sides of ((ref)), respectively, we need $F_{Y_0|\mathcal{C}^0_{1,0}}(y^-_0)<F_{Y_0|\mathcal{C}^0_{1,0}}(y^+_0)$ for ((ref)) to hold, and this implies that $y^-_0 < y^+_0$ from the strict monotonicity of $F_{Y_0|\mathcal{C}^0_{1,0}}$. Next, from an analogous argument with the strict monotonicity of $F_{Y_1|\mathcal{C}^1_{2,0}}$ and $F_{Y_0|\mathcal{C}^0_{2,0}}$, the set of solutions to ((ref)) form a strictly decreasing relationship, where both $F_{Y_1|\mathcal{C}^1_{2,0}}$ and $F_{Y_0|\mathcal{C}^0_{2,0}}$ are on the right-hand side of ((ref)). Hence, because strictly increasing and strictly decreasing relationships intersect only once, ((ref)) and ((ref)) have a unique solution at $(\phi_{2,1}(y^f),\phi_{2,0}(y^f))$.
Next, we generalize the above argument for identifying the counterfactual mappings to the case of arbitrary $k$. With the sign treatments, we employ the following assumption and assume that $k$ different types of monotonicity relationships exist.
Under Assumption (ref), the monotonicity subset $\Lambda$ contains $k$ pairs of instrument values $\lambda_1,\ldots,\lambda_k$ such that each pair $\lambda_i$ has a sign treatment, and the sign treatments $t({\lambda_i})$'s are all different. Then, each $t({\lambda_i})$ characterizes a type of monotonicity relationship of $\lambda_i$. As in the case of $k=2$, Assumption (ref) holds in Example I. From Table (ref) in Section (ref), for $i=1,\ldots,k$, $(i,0)$ has a sign treatment $t{(i,0)}=i$. Then, for $i\neq j$, $(i,0)$ and $(j,0)$ induce different types of monotonicity relationships.
Under Assumption (ref), we obtain key relationships between the compliers essential for identifying the counterfactual mappings. For Example I, the proof of Lemma (ref) in Appendix (ref) derives the following key relationships between the compliers:
The $k$ relationships of (ref) correspond to (ref) and (ref) in the case of $k=2$. Applying Lemma (ref), all the functions in ((ref)) except for the counterfactual mappings are identified. When we fix ((ref)) at $y^{f}\in\mathcal{Y}^\circ$, the $k$ equations of ((ref)) constitute nonlinear simultaneous equations of $\phi_{k,i}(y^{f})$ for $j=0,\ldots,k-1$. We identify these values by solving the $k$ equations of ((ref)) simultaneously at $y^{f}$.
The following lemma shows the identification of the counterfactual mappings under Assumption (ref):
As discussed in Section (ref), $E[Y_s]$ and $Q_{Y_s} (\tau )$ for $\tau\in(0,1)$ are identified if $\phi_{s,t}(\cdot)$ for $t\in\mathcal{T}$ are identified. Hence, we obtain the following theorem:
This result is interesting because the proposed sufficient condition is economically interpretable. We do not need to interpret the numerical conditions on the distribution of the outcome variable when this condition is satisfied. This fact may be helpful when designing a social experiment for which the outcome data will be collected later.
In this section, we compare Assumption (ref) (or Assumption (ref) in the case of $k=2$) with the identification condition of CH2005 and discuss the case where Assumption (ref) is violated. For the binary treatment case, the closed-form expression of Wuthrich2019 is also valid when the full rank condition in CH2005 holds for all the values in $\mathcal{Y}$. VX2017 establish an identification condition weaker than the monotonicity assumption for the binary treatment case and show that the identification condition of CH2005 implies that condition.
In the multi-valued treatment setting, we can show that Assumption (ref) implies the full rank conditions in CH2005 under Assumptions (ref) and (ref) and some differentiability assumptions.\footnote{Appendix (ref) proves this statement.} On the other hand, our identification condition does not require any differentiability assumption. Assumption (ref) may be violated even when the identification condition of CH2005 is satisfied.\footnote{Appendix (ref) provides a numerical example where the full rank condition in CH2005 holds but Assumption (ref) is violated.} The full rank condition in CH2005 only requires that the moment equations (ref) are uniquely solved. Assumption (ref) clarifies the behavioral patterns of each individual for identifying the treatment effects. Assumption (ref) also implies additional testable restrictions compared with CH2005.\footnote{Appendix (ref) discusses these additional restrictions. These restrictions are equally difficult to check statistically compared with the full rank conditions in CH2005 for the binary treatment case.}
Next, we discuss the case where Assumption (ref) is violated. Even when the unordered monotonicity assumption holds, Assumption (ref) is violated if the sign treatments are the same for all the instrument pairs. We consider the $k=2$ case and let the monotonicity inequalities summarized in Table (ref) hold almost surely. In Example I, this may happen when individuals are randomly assigned to one of the three groups, where the cost of choosing field $2$ is decreased for all groups; Group 2 has the largest decrease in cost, and Group 0 has the smallest.
Then, Assumptions 1 and 2 hold, and the unordered monotonicity assumption holds because Assumption 2 (ii) holds for all the instrument pairs. However, the sign treatments of the instrument pairs are $t{(1,0)}=t{(2,0)}=t{(2,1)}=2$, and Assumption (ref) does not hold because all the instrument pairs induce the same type of monotonicity relationship.
Under this monotonicity assumption, we obtain the following two equations from the monotonicity relationships of $(1,0)$ and $(2,0)$ as we obtain ((ref)) and ((ref)) under Assumption (ref):
Whether $\phi_{2,1}(y)$ and $\phi_{2,0}(y)$ are identified from the two equations ((ref)) and ((ref)) depends on the numerical conditions on the conditional cdfs. For each $y^f\in\mathcal{Y}^\circ$, $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are the solutions to the following nonlinear simultaneous equations of two unknown variables $y_1$ and $y_0$:
As in Section (ref), we compare the set of solutions to each of these two equations, and Figure (ref) visualizes the solutions to ((ref)) and ((ref)), respectively, in the two-dimensional space of $(y_1,y_0)$. Notice that, from the strict monotonicity of the conditional cdfs, the solutions to ((ref)) and ((ref)) have both strictly decreasing relationships between $y_1$ and $y_0$, where both the conditional cdfs of $Y_1$ and $Y_0$ are on the same side of ((ref)) and ((ref)). Hence, when these two strictly decreasing relationships intersect only once, the two equations ((ref)) and ((ref)) have a unique solution at $(\phi_{2,1}(y^f),\phi_{2,0}(y^f))$. In this case, we can show that the full rank conditions in CH2005 hold under Assumptions (ref) and (ref) and some differentiability assumptions.\footnote{Appendix (ref) proves this statement.}
In this section, we identify the local treatment effects in the multi-valued treatment setting. Suppose that the monotonicity inequalities hold for $(z,z')$, and that $D_t(z)\leq D_t(z')$ holds almost surely for treatment level $t$. For two different treatment levels $t$ and $t'$ and instrument values $z$ and $z'$, the subpopulation $\{D_{t'}(z)=1,D_t(z')=1\}$ changes the treatment choice from $t'$ to $t$ if the instrument value changes from $z$ to $z'$. The LATE that compares treatment states $t$ and $t'$ conditional on $\{D_{t'}(z)=1,D_t(z')=1\}$ is $E[Y_t|D_{t'}(z)=1,D_t(z')=1]-E[Y_{t'}|D_{t'}(z)=1,D_t(z')=1]$. The local quantile treatment effect (LQTE) conditional on $\{D_{t'}(z)=1,D_t(z')=1\}$ is $Q_{Y_t|D_{t'}(z)=1,D_t(z')=1}(\tau)-Q_{Y_{t'}|D_{t'}(z)=1,D_t(z')=1}(\tau)$ where $\tau\in(0,1)$.
First, we briefly review the identification studies of local treatment effects in multi-valued treatment settings. From the definition, local treatment effects are treatment effects conditional on the compliers when the treatment is binary. As shown in Imbens1994, LATEs are identified as the standard two-stage least squares (2SLS) estimands under the monotonicity assumption. Heckman2018 establish the identification of the conditional means of the potential outcomes on the compliers. However, identifying the local treatment effects is more complex in multi-valued treatment settings because the relationship between the compliers becomes more complicated than in the binary treatment case. Angrist1995 show that the 2SLS estimand generally represents only a weighted average of LATEs with a binary instrument, even under certain monotonicity assumption.\footnote{kline2016evaluating and hull2018isolateing also derive related results in the case of a binary instrument.} kirkeboen2016field and Mountjoy2019 demonstrate that a similar result holds for 2SLS estimands even when there are as many instruments as treatments. Heckman2006 identify the LATEs generalized to compare treatment state $t$ and the set of other states with continuous instruments. Heckman2006 also establish identification on more various subpopulations that are identified with continuous instruments. For a more general class of selection models, lee2018 and Mountjoy2019 show similar identification results for the ATEs that compare two different treatment states, $t$ and $t'$.
This section establishes the identification of the LATEs and LQTEs that compare two different treatment states, $t$ and $t'$, under our monotonicity assumption. We require that the outcome variable is continuously distributed with additional assumptions, such as the rank similarity assumption, necessary for identifying the counterfactual mappings.
Note that, if $(z,z')$ has a sign treatment $t_{(z,z')}=t'$, then $P(\mathcal{C}_{z,z'}^t)=P(D_{t'}(z)=1,D_t(z')=1)$ holds for $t\in\mathcal{T}\setminus\{t'\}$ as we obtain ((ref)) for Example I in Section (ref). Under the rank similarity assumption, identifying the counterfactual mappings leads to identifying the treatment effects conditional on the compliers. The following lemma shows the identification of the conditional distribution of $Y_{t'}$ given $\mathcal{C}_{z,z'}^t$ as well as that of $Y_{t}$ given $\mathcal{C}_{z,z'}^t$:
With Lemma (ref) at hand, we obtain the following theorem that shows the identification of local treatment effects under our assumptions:
In this section, we discuss the content of our assumptions and the applicability of our identification result for Examples II and III in Section (ref). In each example, we confirm the assumptions for Theorems (ref) and (ref) with monotonicity inequalities. Suppose that the treatment $T$ and the instrument $Z$ are sufficiently correlated, and we assume Assumption (ref) and Assumptions (ref) (i), (iii), and (iv) unless stated otherwise.
We take an approach similar to Mountjoy2019 for the model and economic analysis. Mountjoy2019 also employs the discrete choice index model to motivate his monotonicity assumption for continuous instruments. Let $Y$ denote student outcome. Let the treatment $T$ denote college decision that takes values on $\{0,2,4\}$, where $T=0$ denotes no college, while $T=2$ and $T=4$ denote starting college at a two-year institution and a four-year institution, respectively. Let the two binary instruments $Z_2$ and $Z_4$ that take values on $\{0,1\}$ represent the distance to the nearest college, and $Z_2$ equals 1 when the nearest two-year college is within a $d_2$ km radius, whereas $Z_4$ equals 1 when the nearest four-year college is within a $d_4$ km radius. As discussed in Section (ref), the discrete choice index model generates the monotonicity inequalities. As in Mountjoy2019, we define the indirect utilities for each treatment option as follows:
where we assume strict monotonicity for the cost functions as follows:
This requirement is natural because students who live near a college will have a smaller cost for choosing that college. The strict monotonicity of the cost functions generates the monotonicity inequalities summarized in Table (ref) for $z_2,z_4\in\{0,1\}$.
Mountjoy2019 provides the same type of inequalities as the monotonicity inequalities in Table (ref) with continuous instruments. The monotonicity inequalities in Table (ref) with $z_2=z_4=0$ are the same as those for $(1,0)$ and $(2,0)$ in Example I in Section (ref) if we replace the values that $T$ and $Z$ take from $\{0,2,4\}$ to $\{0,1,2\}$ for $T$, and from $\{(0,0),(1,0),(0,1)\}$ to $\{0,1,2\}$ for $Z$. Example II may have more monotonicity inequalities than Example I because $z_2$ and $z_4$ take values of either 0 or 1.
From Table (ref), Assumption (ref) (ii) holds for $\{(1,z_4),(0,z_4)\},\{(z_2,1),(z_2,0)\}\in\Lambda$. From Table (ref), $\{(0,z_4),(1,z_4)\}$ and $\{(z_2,0),(z_2,1)\}$ have different sign treatments $t{\{(0,z_4),(1,z_4)\}}=2$ and $t{\{(z_2,0),(z_2,1)\}}=4$, respectively. Therefore, Assumption (ref) holds, and from Theorems (ref) and (ref), the treatment effects are identified as closed-form expressions.
We take an approach similar to Pinto2015 for the model and economic analysis. Let $Y$ denote the outcome of interest that is continuously distributed. Let the treatment $T$ denote the relocation decision at the intervention onset, where $T=0$ denotes no relocation, which is equivalent to choosing high poverty neighborhood, $T=1$ denotes medium-poverty neighborhood relocation, and $T=2$ denotes low-poverty neighborhood relocation. Let the instrument $Z$ represent voucher assignment that takes values on $\mathcal{Z}=\{a,b,c\}$, where $Z=a$ denotes no voucher (control group), $Z=b$ denotes the Section 8 voucher, and $Z=c$ denotes the experimental voucher.
As discussed in Section (ref), the discrete choice index model generates the monotonicity inequalities. As in Example I, we assume the following relationships for the cost functions under additive separability of the utility functions:
Relationships ((ref)) and ((ref)) can be interpreted in the same way as ((ref)) and ((ref)) in Example I. Applying these restrictions to the cost functions generates the monotonicity inequalities summarized in Table (ref).
Pinto2015 provides similar relationships as ((ref)) and ((ref)) with budget sets and choice restrictions equivalent to the inequalities in Table (ref). From Table (ref), Assumption (ref) (ii) holds for $\Lambda=\{(c,a),(b,c)\}$.\footnote{Pinto2015 further assumes that a neighborhood is a normal good and generates monotonicity inequalities in addition to those in Table (ref). Assumptions of Pinto2015 lead to part (A2) of Assumption (ref) for all the pairs in $\mathcal{P}$, and the unordered monotonicity assumption holds. Our assumptions are weaker than those in Pinto2015 but are sufficient to identify the treatment effects.} From Table (ref), $(c,a)$ and $(b,c)$ have different sign treatments $t{(c,a)}=2$ and $t{(b,c)}=1$, respectively. Therefore, Assumption (ref) holds, and from Theorems (ref) and (ref), the treatment effects are identified as closed-form expressions.
In this paper, we establish sufficient conditions for the identification of the treatment effects when the treatment is discrete and endogenous. We show that an appropriately constructed monotonicity assumption is sufficient, and this condition is economically interpretable. We also derive closed-form expressions of the identified treatment effects.
For the estimation procedure, Wuthrich2019 constructs an estimator in the binary case by semiparametric estimation of observable conditional cdfs, qfs, and probabilities and plugging them into his closed-form expression. A similar approach could be applied to our closed-form expressions.
Alternatively, especially for the estimation of the QTE, we can apply the existing estimation methods under structural quantile models based on the GMM objective function after checking our identification conditions. For estimation based on the GMM objective function, reliable and practically useful methods are developed, particularly for parametric structural quantile models. See CH2006, ChenLee2018, Zhu2018, and KaidoWuthrich2018 for linear-in-parameters quantile models, and see ChernozhukovHong2003 and deCastroetal2019 for nonlinear quantile models. Nonparametric estimation approaches are studied by CIN2007, HorowitzLee2007, ChenPouzo2009,ChenPouzo2012, and GS2012.