EconBase
← Back to paper

Identification of multi-valued treatment effects with unobserved heterogeneity

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

88,750 characters

Identification of multi-valued treatment effects with unobserved heterogeneity



\if0
\setlength{\abovedisplayskip}{2pt}
\setlength{\belowdisplayskip}{2pt}
\fi

	\maketitle

\begin{abstract}
In this paper, we establish sufficient conditions for identifying treatment effects on continuous outcomes in endogenous and multi-valued discrete treatment settings with unobserved heterogeneity. We employ the monotonicity assumption for multi-valued discrete treatments and instruments, and our identification condition has a clear economic interpretation. In addition, we identify the local treatment effects in multi-valued treatment settings and derive closed-form expressions of the identified treatment effects. We provide examples to illustrate the usefulness of our result.
\end{abstract}

\textit{Keywords:} Treatment effect, unobserved heterogeneity, identification, endogeneity, instrumental variable

\textit{JEL classification:} C14, C21, C26

\newpage

\section{Introduction}
Unobserved heterogeneity in treatment effects is an essential consideration in many empirical studies in economics.
As discussed in \cite{Heckman2001}, for example, economic theory and applications strongly suggest that the causal effects of treatments or policy variables differ across individuals and subpopulations with the same characteristics.
Quantile treatment effects characterize the heterogeneous impacts of treatments on individuals with different levels of unobserved characteristics in terms of potential outcome quantiles.
Based on instrumental variable (IV) methods, local treatment effects, as introduced by \cite{Imbens1994}, are the treatment effects conditional on an unobservable subpopulation for which the instrument affects treatment states.

In this paper, we establish sufficient conditions for identifying treatment effects on continuous outcomes in endogenous and multi-valued discrete treatment settings with unobserved heterogeneity using IV methods. We use only discrete instruments for identification because instruments are discrete in many empirical applications.
Discrete treatments are implicitly or explicitly multi-valued in many applications. For example, households may receive different levels of transfers in anti-poverty programs, and students who wish to attend college have multiple options for choosing a college or major.
For the policy-maker, it is essential to compare such multi-valued treatment effects when determining which treatment level is appropriate.

In multi-valued treatment settings, \cite{CH2005} establish the identification of quantile treatment effects (QTEs; on the observed populations) with discrete instruments, and their identification results are testable in principle.
However, as we present in the next section, it is unclear how to interpret the required numerical conditions in each empirical study economically.

The main contributions of this paper are stated as follows.
We establish sufficient conditions for identifying treatment effects in multi-valued treatment settings, which are easier to interpret economically than the identification conditions of \cite{CH2005}.
In addition, we provide closed-form expressions of the identified treatment effects.
We also establish the identification of local treatment effects in multi-valued treatment settings based on our assumptions.

To illustrate the usefulness of our results, we provide three examples based on empirical research and discuss the applicability of our identification result for these examples.
The first example is the effects of choosing different fields on earnings in postsecondary education.
The second example is the effects of expanding access to two-year colleges on student outcomes.
The third example is the effects of relocating to low-poverty neighborhoods on the outcomes of disadvantaged families living in high-poverty neighborhoods.

We provide identification conditions that take the form of monotonicity assumptions.
The monotonicity assumption, introduced by \cite{Imbens1994} for the binary treatment case, has a clear interpretation in empirical studies because they can be motivated based on behavioral assumptions.
The monotonicity assumption we employ for multi-valued and unordered treatments is motivated by an unordered discrete choice model, where each individual chooses the treatment option with the highest utility.
Our monotonicity assumption is related to \cite{Heckman2018}'s ``unordered monotonicity" assumption for multi-valued and unordered treatments.
However, the identification condition we provide is based on a weaker assumption that holds only on a particular subset of the pairs of values the instrument can take.
As highlighted in our examples, imposing the monotonicity assumption on all the instrument pairs employed in many studies, including \cite{Heckman2018}, may be too strong when the instrument is multi-valued.

When the treatment is binary, \cite{VX2017} and \cite{Wuthrich2019} show identification of the treatment effects as closed-form expressions under
the instrument satisfying the monotonicity assumption.
\cite{Wuthrich2019} and \cite{FVX2019} develop plug-in estimators based on the closed-form expressions of the identified treatment effects.
Our identification analysis covers more general settings with
multi-valued treatments and instruments, regardless of whether the treatment is ordered or unordered.

The identification approach adopted here generalizes the idea of matching two distributions, which is introduced by \cite{AtheyImbens2006} and used for identifying the treatment effects in some recent studies.
For the binary treatment case, \cite{VX2017} and \cite{Wuthrich2019} establish identification with binary instruments by matching the two potential outcome distributions conditional on the same subpopulation called ``compliers'' under the monotonicity assumption.
In continuous treatment settings, \cite{torgovitsky2015}, \cite{d2015}, and \cite{ishihara2017} establish identification with binary instruments employing sufficiently large support of the treatment.

In multi-valued treatment settings, however, we generally cannot match the potential outcome distributions for two treatment states in the same subpopulation.
The selection mechanism becomes more complicated than in the binary treatment case, and the support of the treatment is limited compared to the continuous treatment case.
We overcome this difficulty by developing systems of equations with multiple potential outcome distributions.
These equations are derived from relationships between the compliers that can be motivated by the discrete choice model under our monotonicity assumption.
Although the distributions are conditional on different subpopulations, we show that the simultaneous equations can be solved uniquely, and the potential outcome distributions are identified under our assumptions.

We also employ our monotonicity relationships for identifying local treatment effects in multi-valued treatment settings.
We establish identification when the outcome variable is continuously distributed under our monotonicity assumption, with some additional assumptions on unobservable factors, such as the rank similarity assumption.\footnote{In Section \ref{subsec:4.5}, we briefly review the identification studies of local treatment effects in multi-valued treatment settings.}

The remainder of the paper is organized as follows:
in Section \ref{sec:2}, we introduce our basic setup and provide three real-world examples.
In Section \ref{sec:3}, we introduce our monotonicity assumption for multi-valued and unordered treatments.
In Section \ref{sec:4}, we establish the identification of the treatment effects using our monotonicity assumption.
Section \ref{sec:6} provides conclusions with brief recommendations for estimation.
Proofs of the main results and some auxiliary results are presented in Appendices \ref{sec:a}- \ref{sec:a.5}. Some additional discussions are given in Appendices \ref{sec:a.25}-\ref{sec:g} of the Supplemental Appendix.
\section{Basic set up and motivating examples}\label{sec:2}
In this section, we introduce our basic setup with a benchmark nonseparable model.
We provide three real-world examples to show that our identification approach can be applied in well-known empirical settings.

\subsection{Notation and basic assumptions}\label{subsec:2.1}
Throughout this paper, we use the notations $F_A$, $Q_A$, and $f_A$ for the unconditional cumulative distribution function (cdf), quantile function (qf), and probability density function (pdf) of a scalar-valued random variable $A$, respectively. Similarly, for a set $\mathcal{D}$ and random vectors $B$ and $C$, $F_{A|\mathcal{D} BC}(\cdot|b,c)$,
$Q_{A|\mathcal{D} BC}(\cdot|b,c)$, and $f_{A|\mathcal{D} BC}(\cdot|b,c)$ denote the conditional cdf, qf, and pdf of $A$ on $\mathcal{D}\cap\{(B,C)=(b,c)\}$, respectively; let $\mathcal{D}^\circ$ denote the interior of $\mathcal{D}$.

To introduce our basic setup, we consider the following nonseparable simultaneous equations for a continuous outcome and a multi-valued endogenous treatment:
\begin{eqnarray}
Y &=& g(T,X,U),  \label{model.1}\\
T &=& \rho(Z,X,V). \label{model.2}
\end{eqnarray}
For the outcome equation, $Y$ is the outcome, $T\in\mathcal{T}$ is the multi-valued (possibly) unordered treatment where $\mathcal{T}$ contains $k+1$ values, $X\in\mathcal{X}\subset\mathbb{R}^r$ is a vector of observed covariates, and the random vector $U$ captures unobserved heterogeneity in the effect of $T$ on $Y$.
For the treatment $T$, we assume a finite collection of multiple treatment statuses (either unordered or ordered) indexed by $t\in \mathcal{T}$ where, without loss of generality, $\mathcal{T}=\{0,1,2,\ldots, k\}$.
For the treatment equation, $Z\in\mathcal{Z}$ is a discrete instrument where $\mathcal{Z}$ contains at least $k+1$ values, and the random vector $V$ captures unobserved factors affecting selection into treatment.
For simplicity, we suppress $X$ throughout the identification analysis.
All assumptions and results can be understood as conditional on $X$.
The potential outcome under each treatment level $t\in\mathcal{T}$ is $Y_t=g(t,U)$, and we assume $Y_t\in\mathcal{Y}\subset\mathbb{R}$ and $E[|Y_t|]<\infty$.
The potential treatment choice if $Z$ had been externally set to $z$ is $T(z)=\rho(z,V)$.
In this paper, for two different treatment levels $t$ and $t'$, we are interested in the sufficient conditions for identifying the average treatment effect (ATE): $E[Y_t] - E[Y_{t'}]$ and the quantile treatment effect (QTE): $Q_{Y_t} (\tau) - Q_{Y_{t'}} (\tau)$, where $\tau\in(0,1)$.
We are also interested in the sufficient conditions for identifying the local treatment effects, and we introduce them in Section \ref{subsec:4.5}.
For notation simplicity, we assume that the supports of the distributions of $T$, $Y_t$, $Y$, and $Z$ are equal to $\mathcal{T}$, $\mathcal{Y}$, $\mathcal{Y}$, and $\mathcal{Z}$, respectively.
The results in this paper do not rely on these restrictions.

For the outcome equation \eqref{model.1}, we allow $U=(U_0,\ldots,U_k)'$ to be multi-dimensional, and we assume that the potential outcome is expressed as $Y_t=g(t,U_t)$, where we define $U_t:=F_{Y_t} (Y_t)$.
$U_t$ is the rank variable that characterizes heterogeneity of outcomes for individuals with the same observed characteristics by relative ranking in terms of potential outcomes.
For the rank variable, we assume the ``rank similarity" introduced in \cite{CH2005}.
The rank similarity is an assumption that weakens the ``rank invariance" assumption.
Rank invariance assumes that $U$ is a scalar error term and that the rank variables satisfy $U_t=U_{t'}=U$ for any $t\neq t'$.
However, rank similarity allows the rank variables to deviate from a common ranking $U$.

For the treatment equation \eqref{model.2}, we allow $V=(V_0,\ldots,V_k)'$ to be multi-dimensional for the unordered treatment $T$, and each element $V_t$ represents unobserved individual preference heterogeneity from choosing $T=t$.
The treatment decision can be explained by an unordered discrete choice model, where each individual chooses the treatment option with the highest indirect utility,
\begin{equation}\label{eq:chcf}
T(z) = \rho(z,V) = \operatorname*{arg\,max}_{t\in\mathcal{T}} I_t(z,V_t),
\end{equation}
where $I_t$ is the indirect utility of choosing $T=t$.\footnote{We can justify the utility maximization model of the (potential) treatment choice as in \eqref{eq:chcf} under the rank similarity assumption with an argument similar to Example 2 of \cite{chernozhukov2013quantile}, where QTE is used to examine the effects of participating in a 401(k) plan.}
\cite{chesher2013instrumental} also employs a similar unordered choice model where $V$ is allowed to be multi-dimensional.

When the unobserved factor $V$ is a scalar random variable, the two equations \eqref{model.1} and \eqref{model.2} form the triangular model of \cite{chesher2005nonparametric}.
\cite{chesher2005nonparametric} studies interval identification of the endogenous nonseparable triangular model with a discrete ordered treatment.
Our research is related to \cite{chesher2005nonparametric}, but our treatment equation for unordered treatments is fundamentally different from the triangular model that captures a single source of unobserved heterogeneity.

To identify these treatment effects, it suffices to identify the conditional mean and qf of the potential outcomes.
We identify these under the following set of assumptions. \cite{CH2005}, \cite{VX2017}, and \cite{Wuthrich2019} employ a similar set of assumptions.
\begin{assumption}[Instrument independence and rank similarity]\label{asp:ivqr}
The following conditions hold:
\begin{enumerate}
\item[(i)] Potential outcomes: For each $t\in\mathcal{T}$, $Y_t$ is expressed as $Y_t=g(t,U_t)$ for some unknown function $g$ and $U_t=F_{Y_t} (Y_t)$, and $F_{Y_t} (\cdot)$ is continuous.
\item[(ii)] Independence: $\{U_t\}_{t=0}^k$ are independent of $Z$.
\item[(iii)] Selection: $T$ is expressed as $T = \rho(Z, V)$ for some unknown function $\rho$ and random vector $V$.
\item[(iv)] Rank similarity: Conditional on $(Z,V) = (z,v)$, $\{U_t\}_{t=0}^k$ are identically distributed.
\item[(v)] Outcome support: The closure of $\mathcal{Y}^\circ$ is equal to $\mathcal{Y}$, and $F_{Y_t}(\mathcal{Y}^\circ)$ does not depend on $t\in\mathcal{T}$.
\end{enumerate}
\end{assumption}
\noindent Assumption \ref{asp:ivqr} (i) states an expression for the potential outcome with the rank variable and imposes continuity of the potential outcome cdf.
Under Assumption \ref{asp:ivqr} (i), $F_{Y_t} (y)$ is strictly increasing in $y\in\mathcal{Y}^\circ$, and $Q_{Y_t} (\tau)$ is strictly increasing in $\tau\in(0,1)$.\footnote{Lemmas \ref{lem:3b} and \ref{lem:3c} in Appendix \ref{sec:a.5} prove these properties.}
We do not assume that $Q_{Y_t} (\cdot)$ is continuous and allow $F_{Y_t} (\cdot)$ to have flat intervals.
Under Assumption \ref{asp:ivqr} (i), the rank variable $U_t$ follows the uniform distribution on $(0,1)$, and $Y_t$ and $Q_{Y_t}(U_t)$ are identically distributed.
Hence, we can interpret the QTE as treatment effects on individuals with the same level of unobserved heterogeneity at some level $U_t=\tau$.\footnote{We employ a slightly different definition for the rank variable from the original definition of \cite{CH2005}.
\cite{CH2005} define the rank variable $U_t$ as a uniformly distributed random variable on $(0,1)$ that satisfies $Y_t=Q_{Y_t}(U_t)$, and they directly assume that $Q_{Y_t} (\tau)$ is strictly increasing in $\tau\in(0,1)$.
The difference does not matter in our settings because $Y_t$ and $Q_{Y_t}(U_t)$ are identically distributed.}
Assumption \ref{asp:ivqr} (ii) imposes conditional independence between the potential outcomes and the instrument.
Assumption \ref{asp:ivqr} (iii) states a general selection equation where the random vector $V$ captures unobserved factors affecting selection into treatment.
Assumption \ref{asp:ivqr} (iv) is the rank similarity assumption.
The rank similarity is arguably strong, but this condition has essential implications for identification and is consistent with many empirical situations.
Assumption \ref{asp:ivqr} (v) is assumed to simplify the proofs in the main paper, and we relax Assumption \ref{asp:ivqr} (v) in Appendix \ref{sec:g}.
See also Remark \ref{rem:clfm} for related discussions.

The main statistical implication of Assumption \ref{asp:ivqr} is that, for each $t\in\mathcal{T}$, the qf of $Y_t$ satisfies the following nonlinear moment equation (\cite{CH2005} Theorem 1):
\begin{equation}\label{eq:ti}
\sum_{t=0}^k F_{Y|TZ}(Q_{Y_t} (\tau)|t,z)p_t(z)=\tau,
\end{equation}
where $p_t(z)$ is defined as
\begin{equation}\label{eq:pt}
p_t(z):=P(T=t|Z=z).
\end{equation}
\cite{CH2005} show that $Q_{Y_t} (\tau )$'s are identified if the following $(k+1)\times(k+1)$ matrix $\Pi'(y_0,\ldots,y_k)$ is full rank for all the values $(y_0,\ldots,y_k)$ in a set of potential solutions to the moment equations (\ref{eq:ti}):
\begin{equation}\label{eq:ch2}
\Pi'(y_0,\ldots,y_k):=
\begin{pmatrix}
f_{Y|TZ}(y_0|0,z_0)p_0(z_0) & \cdots & f_{Y|TZ}(y_k|k,z_0)p_k(z_0) \\
\vdots & \ddots & \vdots \\
f_{Y|TZ}(y_0|0,z_k)p_0(z_k) & \cdots & f_{Y|TZ}(y_k|k,z_k)p_k(z_k)
\end{pmatrix},
\end{equation}
where $\{z_0,z_1\ldots,z_k\}\subset\mathcal{Z}$.
The identification condition of \cite{CH2005} is, in principle, directly testable.
However, in each empirical study, it is not easy to check this numerical condition, which takes the form of matrices of the outcome conditional densities.
It is unclear how to interpret the requirements for the endogenous variable and the instruments implied in this condition.
In this paper, we establish sufficient conditions for identifying treatment effects in multi-valued treatment settings, which are easier to interpret economically than the identification conditions of \cite{CH2005}.
We also provide closed-form expressions of the identified treatment effects that could be used for constructing plug-in estimators.\footnote{This idea resembles that of \cite{Das2005}, who develops an estimation strategy based on the closed-form expression of the regression function with discrete endogenous treatments in the nonparametric regression model with an additive error term.}

\subsection{Examples}\label{subsec:2.2ex}
In this section, we introduce three examples based on empirical research.
Throughout this paper, we consider Example I as a running example.
In Section \ref{sec:5}, we discuss the content of our assumptions and the applicability of our identification result for Examples II and III.

\subsubsection{Example I}
The first example is the effects of choosing different fields on earnings in postsecondary education.
In postsecondary education, almost all students have to choose a field of study, and earnings differ across not only universities but also fields.
\cite{kirkeboen2016field} study the identification and estimation of local average treatment effects (LATEs) of choosing different fields on earnings in Norway's postsecondary educational system.
\cite{kirkeboen2016field} find that a centralized admission process in Norway randomizes applicants into different groups, and applicants in each group are much more likely to receive an offer for each field.
Based on this process, \cite{kirkeboen2016field} use the predicted offers for each field as an instrument.
This process also provides information on individuals' ranking of fields, and the identification analysis of \cite{kirkeboen2016field} depends on each individual's next best alternative, that is, the field one would prefer if one's preferred field would not be feasible.
Our results can be applied when we can find a certain instrument that effectively randomizes students into different groups, and information on individuals' ranking of fields is not necessary for our identification approach.

\subsubsection{Example II}
The second example is the effects of expanding access to two-year colleges on student outcomes.
Two-year community colleges will increase the flow of young people into higher education, and expanding access to two-year colleges is expected to positively affect educational attainment and earnings.
However, higher enrollment rates in two-year colleges may adversely affect student outcomes because college applicants are discouraged from paying higher tuition and enrolling directly in four-year colleges.
\cite{Mountjoy2019} and \cite{ferreyra2022labor} estimate the effects of expanding access to two-year colleges in the United States and Colombia using instruments, respectively.
The treatment is the college applicant's decision to start college at a two-year or four-year institution or not to enroll in college.
With this multi-valued treatment, they compare the two opposing effects: the positive effect on new two-year entrants who otherwise would not have enrolled in any college, and the negative effect on two-year entrants who otherwise would have started directly at a four-year institution.
For the instruments, they use the distance to the nearest college.
\cite{Mountjoy2019} directly uses the distance as a continuous instrument, and \cite{ferreyra2022labor} use a discrete instrument that indicates whether the nearest college is located within a certain distance radius.
We employ discrete instruments and show that our identification method can be applied to discrete instruments.

\subsubsection{Example III}
The third example is Moving to Opportunity (MTO), a housing experiment implemented between 1994 and 1998.
The MTO experiment was designed to evaluate the effects of relocating to low-poverty neighborhoods on the outcomes of disadvantaged families living in high-poverty urban neighborhoods in the United States.
This project randomly assigned housing vouchers from the Section 8 program that could be used to subsidize housing costs.
Eligible families were placed in one of the following three assignment groups: experimental group, to which Section 8 housing vouchers were assigned but restricted to use their vouchers in a low-poverty neighborhood; Section 8 group, to which regular Section 8 housing vouchers were assigned without any restriction on their place of use; or the control group to which no voucher was assigned.
Impact evaluations were conducted in 2002, 2009, and 2010.
See \cite{Orr2003}, \cite{Sanbonmatsu2011}, and \cite{Shroder2012}
for the detailed information on this project.
For the recent studies that find evidence of neighborhood effects on adult employment, \cite{aliprantis2019evidence} estimate LATEs for moving to a higher-quality neighborhood under an ordered treatment model using neighborhood quality as an observed continuous measure of the treatment variable.

Under an unordered treatment model, \cite{Pinto2015} applies the identification results of \cite{Heckman2018} for estimating the conditional means of the potential outcomes on the compliers.
We show that, under some additional assumptions on unobservable factors, such as the rank similarity assumption, the ATEs and QTEs are nonparametrically identified when the outcome variable is continuously distributed.

\section{The (generalized) monotonicity assumption}\label{sec:3}
In this section, we introduce the monotonicity assumption we employ for the multi-valued treatment case.
We illustrate that the monotonicity assumption is motivated by economic analysis using a discrete choice index model introduced in Section \ref{sec:2}.

\subsection{Monotonicity assumption in multi-valued treatment settings}\label{subsec:3.1}
The monotonicity assumption is first introduced by \cite{Imbens1994} for the binary treatment case, and \cite{Heckman2018} generalize this assumption to  unordered multi-valued treatment settings.
We define a binary variable $D_t:=1\{T=t\}$, where $1\{\mathcal{A}\}$ is the indicator function of a set $\mathcal{A}$ and $D_t$ is an indicator function of each treatment level.
Then, the observed outcome can be represented as $Y=\sum_{t=0}^kY_tD_t$.
We also define $D_t(z):=1\{T(z)=t\}$ as an indicator function of each potential treatment state if $Z$ had been externally set to $z$.
We define $\mathcal{P}:=\{(z,z')\in\mathcal{Z}^2:z\neq z'\}$ as a set of pairs of different values that the instrument can take.
The monotonicity assumption imposes restrictions on these pairs.

For the binary treatment case, the monotonicity assumption requires that either $T(z)\leq T({z'})$ and $P(\{T(z)=0, T(z')=1\})>0$ or $T(z)\geq T({z'})$ and $P(\{T(z)=1, T(z')=0\})>0$ hold almost surely for each $(z,z')\in\mathcal{P}$.
This condition implies that either $P(\{T(z)=1, T(z')=0\})=0$ or $P(\{T(z)=0, T(z')=1\})=0$ holds for each $(z,z')\in\mathcal{P}$.
Under the monotonicity assumption, individuals who change their choice respond in only one direction to a change in $Z$, and the group with positive probability is called ``compliers."

\cite{Heckman2018} generalize this argument to multi-valued treatment settings.
For each $z\in\mathcal{Z}$ and $t\in\mathcal{T}$, they assume the monotonicity assumption for binary treatments on each binary indicator $D_t(z)$, and either $D_t(z)\leq D_t(z')$ or $D_t(z)\geq D_t(z')$ holds almost surely for each pair $(z,z')\in\mathcal{P}$.
They call this assumption ``unordered monotonicity'' because this condition can be assumed on the unordered treatments.
However, as we see in Section \ref{subsec:3.2}, imposing such conditions on all the pairs $(z,z')\in\mathcal{P}$ may be too strong when the instrument is multi-valued.
Hence, we employ a weaker assumption that imposes such conditions only on a subset of $\mathcal{P}$.
That subset is determined differently in each situation.
Recently, a similar problem has been discussed when there are multiple instruments.
\cite{MTW2019,mogstad2020policy} and \cite{goff2020vector} consider the binary treatment case, and \cite{Mountjoy2019} considers the multi-valued unordered treatment case.
As discussed in Section \ref{sec:5}, \cite{Mountjoy2019} employs the same type of monotonicity assumption with multiple continuous instruments.
We consider a (possibly) scalar multi-valued instrument and weaken the monotonicity assumption from another perspective.

We define the compliers for the multi-valued treatment case as $\mathcal{C}_{z,z'}^t:=\{D_t(z)=0, D_t(z')=1\}$.
Our monotonicity assumption is characterized by inequalities such as $D_t(z)\leq D_t(z')$ and $D_t(z)\geq D_t(z')$.
We employ the following monotonicity assumption:
\begin{assumption}[Instrument independence and monotonicity]\label{asp:mt}
There exists a subset $\Lambda$ of $\mathcal{P}$ such that the following conditions hold for each $\lambda=(z,z')\in\Lambda$:
\begin{enumerate}
\item[(i)] Independence: $(Y_t, T(z))$ for $t\in\mathcal{T}$ and $z\in\lambda$ are jointly independent of $Z$.
\item[(ii)] Monotonicity inequalities: Either $D_t(z)\leq D_t(z')$ or $D_t(z)\geq D_t(z')$ holds almost surely for each $t\in\mathcal{T}$.
\item[(iii)] Instrument relevance: Either $P(\mathcal{C}^t_{z,z'})>0$ or $P(\mathcal{C}^t_{z',z})>0$ holds for each $t\in\mathcal{T}$.
\item[(iv)] Sufficient support: The support of the conditional distribution of $Y_t$ on either $\mathcal{C}^t_{z,z'}$ or $\mathcal{C}^t_{z',z}$ is $\mathcal{Y}$.
\end{enumerate}
\end{assumption}

\noindent We call this subset $\Lambda$ ``monotonicity subset."
When $T$ is binary, Assumptions \ref{asp:mt} (i)--(iii) are the monotonicity assumption in \cite{Imbens1994}.
Assumption \ref{asp:mt} (i) strengthens Assumption \ref{asp:ivqr} (ii) and assumes that the potential outcome and treatment are jointly independent of the instrument.
Assumption \ref{asp:mt} (ii) assumes that monotonicity inequalities hold on a particular subset $\Lambda$ of $\mathcal{P}$.
Assumption \ref{asp:mt} (iii) is an instrument relevance condition and assumes that the compliers always exist.
Under these conditions, when monotonicity inequalities hold on $(z,z')\in\Lambda$, we exclude the cases where neither $D_t(z)<D_t(z')$ nor $D_t(z)>D_t(z')$ can happen for some $t\in\mathcal{T}$.
Assumption \ref{asp:mt} (iv) strengthens the instrument relevance condition and assumes that the compliers are sufficiently large.
\cite{VX2017} employ a similar condition for the binary treatment case.
Under Assumption \ref{asp:mt} (iv), each conditional cdf of the potential outcome on compliers strictly increases on $\mathcal{Y}^\circ$.\footnote{We show this statement in Lemma \ref{lem:3b} in Appendix \ref{sec:a.5}.}

\subsection{Motivating the monotonicity assumption}\label{subsec:3.2}
In this section, we illustrate that economic analysis implies the monotonicity assumption we employ for the multi-valued treatment case.
We consider Example I and use a discrete choice index model introduced in Section \ref{sec:2} to motivate the monotonicity assumption.
\cite{Mountjoy2019} also employs a discrete choice index model to motivate his monotonicity assumption for continuous instruments.
We take an approach similar to \cite{Mountjoy2019} for the choice index model.\footnote{\cite{Heckman2018} and \cite{Pinto2015} consider a general utility maximization problem and employ revealed preference analysis to motivate their unordered monotonicity assumption.
The generalized models also imply our monotonicity assumption.}

We consider a setting where individuals choose between not taking any postsecondary education or completing some postsecondary education and choose between $k$ different fields of study labeled as $1,\ldots,k$.
Let $Y$ denote observed earnings.
For the treatment $T$, let $T=0$ denote not taking any postsecondary education, and for $j\in\{1,\ldots,k\}$, $T=j$ denotes completing field $j$.
Suppose that individuals are randomly assigned to one of the following $(k+1)$ groups; individuals assigned to group $0$ have no cost reduction and for $j\in\{1,\ldots,k\}$, the cost of choosing field $j$ is decreased for individuals assigned to group $j$.
Let the instrument $Z$ represent group assignment that takes values on $\mathcal{Z}=\{0,1,\ldots,k\}$, where $Z=j$ denotes assignment to group $j$.
In Example I, Assumption \ref{asp:mt} (i) holds because vouchers are randomly assigned.
Suppose that the treatment $T$ and the instrument $Z$ are sufficiently correlated, and we assume Assumptions \ref{asp:mt} (iii) and (iv) unless stated otherwise.

We first consider a case where individuals choose between three alternatives.
As in \cite{Mountjoy2019}, we assume additive separability of the utility functions in unobservable components for simplicity and ease of visualization.
As in \eqref{eq:chcf} in Section \ref{sec:2}, each individual chooses the treatment option with the highest indirect utility:
\begin{equation*}
T(z) = \operatorname*{arg\,max}_{t\in\mathcal{T}} I_t(z,V_t),
\end{equation*}
where the indirect utilities for each treatment option are defined as follows:
\[
I_0 = 0,\quad I_1 = V_1 - \mu_1(Z), \text{ and } I_2 = V_2 - \mu_2(Z).
\]
The utility of not taking any postsecondary education is normalized to zero.
$V_t$ is an individual's gross utility from choosing field $t$ and represents unobserved individual preference heterogeneity.
$\mu_t(Z)$ is the cost of choosing field $t$, and $I_t$ is the net utility of choosing field $t$.
The potential treatments under this choice index model are expressed as follows:
\begin{eqnarray}
D_0(z) &=& 1\{V_1 < \mu_1(z), V_2 < \mu_2(z)\}, \label{eq:te1} \\
D_1(z) &=& 1\{V_1 > \mu_1(Z), V_2 - V_1 < \mu_2(z) - \mu_1(z)\}, \label{eq:te2} \\
D_2(z) &=& 1\{V_2 > \mu_2(Z), V_2 - V_1 > \mu_2(z) - \mu_1(z)\}. \label{eq:te3}
\end{eqnarray}
Figure \ref{fig:ex1.1} (a) shows how these treatment choice equations \eqref{eq:te1}-\eqref{eq:te3} partition the two-dimensional space of unobserved preferences $(V_1,V_2)$.
Individuals who choose $T=0$ have low preferences for fields 1 and 2 relative to their costs, while those who choose $T=1$ or $T=2$ have higher preferences for their treatment choice.

For the cost functions, we can naturally assume the following relationships from the group assignment of the instrument:
\begin{eqnarray}
\mu_1(0) &=& \mu_1(2)>\mu_1(1), \label{eq:cf1} \\
\mu_2(0) &=& \mu_2(1)>\mu_2(2). \label{eq:cf2}
\end{eqnarray}
Relationship \eqref{eq:cf1} holds because the cost of choosing field 1 decreases for individuals assigned to group 1.
Relationship \eqref{eq:cf2} holds for a similar reason.
Applying the restrictions on the cost functions \eqref{eq:cf1} and \eqref{eq:cf2} to the treatment choice equations \eqref{eq:te1}-\eqref{eq:te3} generates eight monotonicity inequalities summarized in Table \ref{tab:mtc1}.
\begin{table}[htb]
\caption{Monotonicity inequalities of Example I for the case of $k=2$}\label{tab:mtc1}
\centering
\begin{tabular}{|c| c c c c|} \hline
\multicolumn{1}{|c|}{} & \multicolumn{4}{|c|}{$\mathcal{T}$} \\ \hline
\multicolumn{1}{|c|}{} & & $0$ & $1$ & $2$ \\
& $(1,0)$ & $\,\,\,\,\,\,\,\,D_0(1)\leq D_0(0)\,\,\,\,\,\,\,\,$ & $\,\,\,\,\,\,\,\,D_1(1)\textcolor{red}{\geq} D_1(0)\,\,\,\,\,\,\,\,$ & $\,\,\,\,\,\,\,\,D_2(1)\leq D_2(0)\,\,\,\,\,\,\,\,$ \\
$\mathcal{P}$ & $(2,0)$ & $\,\,D_0(2)\leq D_0(0)\,\,$ & $\,\,D_1(2)\leq D_1(0)\,\,$ & $\,\,D_2(2)\textcolor{red}{\geq} D_2(0)\,\,$ \\
& $(1,2)$ & $\,\,\,\,$ & $\,\,D_1(1)\geq D_1(2)\,\,$ & $\,\,D_2(1)\leq D_2(2)\,\,$ \\ \hline
\end{tabular}
\end{table}
From Table \ref{tab:mtc1}, Assumption \ref{asp:mt} (ii) holds for
$\Lambda=\{(1,0),(2,0)\}$.
For the inequalities in Table \ref{tab:mtc1}, \cite{kirkeboen2016field} also employ $D_1(1)\geq D_1(0)$ and $D_2(2)\geq D_2(0)$, and we obtain the other inequalities from the restrictions on the cost functions \eqref{eq:cf1} and \eqref{eq:cf2}.
As we discuss in Section \ref{sec:5}, \cite{Mountjoy2019} provides the same type of inequalities as the monotonicity inequalities for $(1,0)$ and $(2,0)$ with continuous instruments.

Figures \ref{fig:ex1.1} (b)--(d) visualize the monotonicity inequalities summarized in Table \ref{tab:mtc1}.
Figure \ref{fig:ex1.1} (b) illustrates how a shift in the instrument from $Z=0$ to $Z=1$ induces the three monotonicity inequalities for $(1,0)$.
Because this shift in $Z$ decreases the cost of choosing field 1 from $\mu_1(0)$ to $\mu_1(1)$ but does not change the cost of choosing field 2, some individuals find field 1 more attractive and change their choice from either $T=0$ or $T=2$ to $T=1$, but no individuals find field 1 less attractive and leave from the $T=1$ group.
Hence, the expansion of the $T=1$ region induces $D_1(1)\geq D_1(0)$ and, at the same time, the shrinkage of both the $T=0$ and $T=2$ regions induces $D_0(1)\leq D_0(0)$ and $D_2(1)\leq D_2(0)$.
Analogously, Figure \ref{fig:ex1.1} (c) illustrates that a shift in the instrument from $Z=0$ to $Z=2$ induces the three monotonicity inequalities for $(2,0)$.

Figure \ref{fig:ex1.1} (d) illustrates that a shift in the instrument from $Z=1$ to $Z=2$ does not induce any monotonicity inequality between $D_0(1)$ and $D_0(2)$, and $(1,2)$ is not contained in the monotonicity subset.
Because this shift in $Z$ increases the cost of choosing field 1 from $\mu_1(1)$ to $\mu_1(2)$ and decreases the cost of choosing field 2 from $\mu_2(1)$ to $\mu_2(2)$, some individuals who find field 1 less attractive move on to the $T=0$ group and, at the same time, some individuals who find field 2 more attractive leave from the $T=0$ group.
Hence, $D_0(1)>D_0(2)$ and $D_0(1)<D_0(2)$ can both happen depending on whether more individuals are induced into or out from the $T=0$ group.

\begin{figure}[htbp]
\begin{tabular}{cc}
\centering
\begin{minipage}[t]{0.45\hsize}
\centering
\scalebox{0.95}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$V_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$V_2$};

\coordinate(P)at(3.5,3.5);
\draw[thick]($(XS)!(P)!(XL)$)node[below]{$\mu_1(z)$}--(P)--($(YS)!(P)!(YL)$)node[left]{$\mu_2(z)$};
\draw[thick,domain=3.5:6]plot(\x,\x)node[left]{$V_2-V_1=\mu_2(z)-\mu_1(z)$};

\draw(2,2)node{$D_0(z)=1$};
\draw(5,2)node{$D_1(z)=1$};
\draw(2,5)node{$D_2(z)=1$};
\end{tikzpicture}
}
\subcaption{Treatment choices for $Z=z$}
\end{minipage} &
\begin{minipage}[t]{0.45\hsize}
\centering
\scalebox{0.95}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$V_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$V_2$};

\coordinate(P0)at(3.5,3.5);
\draw[dashed,thick]($(XS)!(P0)!(XL)$)node[below]{$\mu_1(0)$}--(P0)--($(YS)!(P0)!(YL)$);
\draw[dashed,thick,domain=3.5:6]plot(\x,\x);

\coordinate(P1)at(2,3.5);
\draw[thick]($(XS)!(P1)!(XL)$)node[below]{$\mu_1(1)$}--(P1)--($(YS)!(P1)!(YL)$)node[left]{$\mu_2(1)$};
\draw[thick,domain=2:4.5]plot(\x,\x+1.5)node[left]{$V_2-V_1=\mu_2(1)-\mu_1(1)$};

\draw(1,2)node{$D_0(1)=1$};
\draw(4.5,2)node{$D_1(1)=1$};
\draw(1.5,5)node{$D_2(1)=1$};
\draw[<-,>=stealth,thick](2.1,1)--(3.4,1);
\draw[<-,>=stealth,thick](3.1,4.5)--(4.4,4.5);
\end{tikzpicture}
}
\subcaption{Instrument shift from $Z=0$ to 1}
\end{minipage} \\
\begin{minipage}[t]{0.45\hsize}
\centering
\scalebox{0.95}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$V_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$V_2$};

\coordinate(P0)at(3.5,3.5);
\draw[dashed,thick]($(XS)!(P0)!(XL)$)--(P0)--($(YS)!(P0)!(YL)$)node[left]{$\mu_2(0)$};
\draw[dashed,thick,domain=3.5:6]plot(\x,\x);

\coordinate(P2)at(3.5,2);
\draw[thick]($(XS)!(P2)!(XL)$)node[below]{$\mu_1(2)$}--(P2)--($(YS)!(P2)!(YL)$)node[left]{$\mu_2(2)$};
\draw[thick,domain=3.5:6]plot(\x,\x-1.5)node[above left]{$V_2-V_1=\mu_2(2)-\mu_1(2)$};

\draw(2,1.5)node{$D_0(2)=1$};
\draw(5,1.5)node{$D_1(2)=1$};
\draw(2,4.2)node{$D_2(2)=1$};
\draw[<-,>=stealth,thick](1,2.1)--(1,3.4);
\draw[<-,>=stealth,thick](4.5,3.1)--(4.5,4.4);
\end{tikzpicture}
}
\subcaption{Instrument shift from $Z=0$ to 2}
\end{minipage} &
\begin{minipage}[t]{0.45\hsize}
\centering
\scalebox{0.95}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$V_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$V_2$};

\coordinate(P1)at(2,3.5);
\draw[dashed,thick]($(XS)!(P1)!(XL)$)node[below]{$\mu_1(1)$}--(P1)--($(YS)!(P1)!(YL)$)node[left]{$\mu_2(1)$};
\draw[dashed,thick,domain=2:4.5]plot(\x,\x+1.5);

\coordinate(P2)at(3.5,2);
\draw[thick]($(XS)!(P2)!(XL)$)node[below]{$\mu_1(2)$}--(P2)--($(YS)!(P2)!(YL)$)node[left]{$\mu_2(2)$};
\draw[thick,domain=3.5:6]plot(\x,\x-1.5);

\draw(1,1)node{$D_0(2)=1$};
\draw(5,1.5)node{$D_1(2)=1$};
\draw(1.5,4.5)node{$D_2(2)=1$};
\draw[->,>=stealth,thick](2.1,1)--(3.4,1);
\draw[<-,>=stealth,thick](1,2.1)--(1,3.4);
\end{tikzpicture}
}
\subcaption{Instrument shift from $Z=1$ to 2}
\end{minipage}
\end{tabular}
\caption{Visualization of shifts in the instrument}\label{fig:ex1.1}
\end{figure}

We generalize the preceding argument to general $k\in\mathcal{T}$.
Similar to the case of $k=2$, a discrete choice model with $k+1$ treatment options generates the following monotonicity inequalities:
\begin{equation}\label{eq:exi3}
D_{i}({i})\geq D_{i}({0})\text{ and }D_j({i})\leq D_j({0})\text{ for } i=1,\ldots,k\text{ and }j\in\mathcal{T}\setminus\{ i\}.
\end{equation}
Table \ref{tab:mtc2} summarizes (\ref{eq:exi3}).
From Table \ref{tab:mtc2}, Assumption \ref{asp:mt} (ii) holds for $\Lambda=\{(1,0),\ldots,(k,0)\}$.
\begin{table}[htb]
\caption{Monotonicity inequalities of Example I}\label{tab:mtc2}
\begin{center}
\scalebox{0.83}{
\begin{tabular}{|c|c c c c c c|} \hline
\multicolumn{1}{|c|}{} & \multicolumn{6}{|c|}{$\mathcal{T}$} \\ \hline
\multicolumn{1}{|c|}{} & & $0$ & $1$ & $_{\cdots}$ & $k-1$ & $k$ \\
 &${(1,0)}$& $D_0({1})\leq D_0({0})$ & $D_1({1})\textcolor{red}{\geq} D_1({0})$ & $_{\cdots}$ &$D_{k-1}({1})\leq D_{k-1}({0})$ & $D_k({1})\leq D_k({0})$ \\
 $\mathcal{P}$&$\vdots$ & $\vdots$ & $\vdots$ & $\vdots$ & $\vdots$ & $\vdots$ \\
 &${(k-1,0)}$& $D_0({k-1})\leq D_0({0})$ & $D_1({k-1})\leq D_1({0})$ & $_{\cdots}$ & $D_{k-1}({k-1})\textcolor{red}{\geq} D_{k-1}({0})$ & $D_k({k-1})\leq D_k({0})$ \\
 &${(k,0)}$& $D_0({k})\leq D_0({0})$ & $D_1({k})\leq D_1({0})$ & $_{\cdots}$ & $D_{k-1}({k})\leq D_{k-1}({0})$ & $D_k({k})\textcolor{red}{\geq} D_k({0})$ \\ \hline
\end{tabular}
}
\end{center}
\end{table}

\section{Identification}\label{sec:4}
In this section, we establish the identification of the potential outcome distributions and the local treatment effects using our monotonicity assumption.
Before establishing our main results, we first introduce a map termed ``counterfactual mapping.''
Counterfactual mapping, developed by \cite{VX2017}, is an essential tool for identification.
We then establish the identification of counterfactual mappings.

\subsection{Counterfactual mappings and our identification challenge}\label{subsec:2.2}
In this section, we introduce counterfactual mapping in multi-valued treatment settings.
We show that identifying the potential outcome distributions follows from identifying the counterfactual mappings.
Our key identification challenge is to recover the counterfactual mappings in multi-valued treatment settings.

For $s,t\in\mathcal{T}$, define $\phi_{s,t}:\mathbb{R}\to\mathbb{R}$ as $\phi_{s,t}(y):=Q_{Y_t}(F_{Y_s}(y))$.
$\phi_{s,t}$ is called ``counterfactual'' mapping from $Y_s$ to $Y_t$ because the potential outcomes are also called ``counterfactual outcomes."
\cite{VX2017} define a similar mapping for the binary treatment case.
From the definition, this mapping relates the quantiles of the distribution of $Y_s$ to that of $Y_t$.
Under Assumption \ref{asp:ivqr} (i), this mapping is strictly increasing on $\mathcal{Y}^\circ$, and $\phi_{s,r}=\phi_{t,r}\circ\phi_{s,t}$ holds for $s,t,r\in\mathcal{T}$.
Moreover, under Assumption \ref{asp:ivqr} (v),
an inverse mapping $\phi_{s,t}^{-1}$ exists on $\mathcal{Y}^\circ$, and $\phi_{t,s}(y)=\phi_{s,t}^{-1}(y)$ holds for $y\in\mathcal{Y}^\circ$.
Similarly, we define the unconditional counterfactual mapping for $s,t\in\mathcal{T}$ as $\phi_{s,t}(y):=Q_{Y_t}(F_{Y_s}(y))$.

The following lemma shows that the potential outcome cdfs and means can be written as compositions of the counterfactual mappings and observable distributions.
\cite{VX2017} show a similar result for the binary treatment case.

\begin{lemma}[Potential outcome cdfs and means via counterfactual mappings]\label{lem:clfm}
Suppose that Assumption \ref{asp:ivqr} holds. Define $p_t(z)$ as in (\ref{eq:pt}).
Then, the following holds:
\begin{enumerate}
\item[(a)] For each $s\in\mathcal{T}$, $F_{Y_s}(y)$ for $y\in\mathcal{Y}^\circ$ can be expressed as
\begin{equation}\label{eq:clfm}
F_{Y_s}(y)=\sum_{t=0}^k F_{Y|TZ}(\phi_{s,t}(y)|t,z)p_t(z).
\end{equation}
\item[(b)] For each $s\in\mathcal{T}$, $E[Y_s]$ can be expressed as
\begin{equation}\label{eq:clfmasf}
E[Y_s]=\sum_{t=0}^k E[\phi_{t,s}(Y)|T=t,Z=z]p_t(z).
\end{equation}
\end{enumerate}
\end{lemma}
\noindent Lemma \ref{lem:clfm} follows from the rank similarity assumption.
For each $s,t\in\mathcal{T}$, the rank variables $U_s$ and $U_t$, and hence $Y_s$ and $\phi_{t,s}(Y_t)$, are identically distributed conditional on $(T,Z)=(t,z)$.

Lemma \ref{lem:clfm} implies that for each $s\in\mathcal{T}$, $E[Y_s]$ and $Q_{Y_s} (\tau )$ for $\tau\in(0,1)$ are identified as closed-form expressions if $\phi_{s,t}$ for $t\in\mathcal{T}$ is also identified as a closed-form expression.
Hence, we establish sufficient conditions to identify $\phi_{s,t}$'s and derive the closed-form expressions of $\phi_{s,t}$'s.

\begin{remark}[]\label{rem:clfm}
We only need Lemma \ref{lem:clfm} (a) to identify $E[Y_s]$ and $Q_{Y_s} (\tau )$ for $\tau\in(0,1)$.
The identification of $\phi_{s,t}(y)$ for $y\in\mathcal{Y}^\circ$ suffices for the identification of the treatment effects because we assume that $F_{Y_s} (y )$ is continuous in $y\in\mathcal{Y}$ in Assumption \ref{asp:ivqr} (i) and that the closure of $\mathcal{Y}^\circ$ is equal to $\mathcal{Y}$ in Assumption \ref{asp:ivqr} (v).
In Appendix \ref{sec:g}, we relax Assumption \ref{asp:ivqr} (v) and derive the closed-form expressions of $\phi_{s,t}$ and $F_{Y_s} (\cdot )$ on a sufficiently large subset of $\mathcal{Y}$.
\end{remark}

\subsection{Preliminary identification results}\label{subsec:4.0}
In this section, we introduce some basic identification results under our monotonicity assumption.
We first establish the identification of the compliers under our monotonicity assumption.
The following lemma shows that when $D_t(z)\leq D_t({z'})$ and $P(\mathcal{C}^t_{z,z'})>0$ hold almost surely for $(z,z')\in\mathcal{P}$ and $t\in\mathcal{T}$, each probability of $\mathcal{C}^t_{z,z'}$ and the conditional cdf of $Y_t$ given $\mathcal{C}^t_{z,z'}$ are identified as closed-form expressions.
\cite{Heckman2018} show a similar result under the unordered monotonicity assumption.

\begin{lemma}[Identification of the compliers]\label{lem:idcpl}
Assume that Assumption \ref{asp:mt} holds, and that $P(D_t(z)\leq D_t({z'}))=1$ and $P(\mathcal{C}^t_{z,z'})>0$ hold for $(z,z')\in\mathcal{P}$ and $t\in\mathcal{T}$. Define $p_t(z)$ as in (\ref{eq:pt}).
Then $P(\mathcal{C}^t_{z,z'})$ and $F_{Y_t|\mathcal{C}^t_{z,z'} }(y)$ for $y\in\mathcal{Y}$ are identified as
\begin{equation}\label{eq:15}
P(\mathcal{C}^t_{z,z'})=p_t(z')-p_t(z),\qquad p_t(z')>p_t(z),
\end{equation}
and
\begin{equation}\label{eq:16}
F_{Y_t|\mathcal{C}^t_{z,z'} }(y)=\frac{F_{Y|TZ}(y|t,z')p_t(z')-F_{Y|TZ}(y|t,z)p_t(z)}{p_t(z')-p_t(z)}.
\end{equation}
\end{lemma}

With Lemma \ref{lem:idcpl} at hand, we establish the identification of the counterfactual mappings.
To provide intuition, we first review the identification results for the binary treatment case by \cite{VX2017}.
Let the treatment be binary so that $\mathcal{T}=\{0,1\}$.
Observe that
\begin{equation}\label{eq:eqcpl}
\mathcal{C}^1_{z,z'}=\mathcal{C}^0_{z',z}
\end{equation}
holds from the definition when $T$ is binary.
Suppose Assumptions \ref{asp:ivqr} and \ref{asp:mt} hold with $P(D_1(z)\leq D_1({z'}))=1$ and $P(\mathcal{C}^1_{z,z'})>0$.
Note that under the rank similarity assumption, the rank variables
$U_0$ and $U_1$ are identically distributed conditional on $\mathcal{C}^1_{z,z'}$.\footnote{We show this statement in Lemma \ref{lem:3a} in Appendix \ref{sec:a.5}.}
Furthermore, from the definition of the counterfactual mapping, if $y\in\mathcal{Y}^\circ$ is the $\tau\in(0,1)$ quantile of the distribution of $Y_1$, then $\phi_{1,0}(y)$ is the $\tau$ quantile of the distribution of $Y_0$.
Therefore, given (\ref{eq:eqcpl}), we obtain the following equation for the potential outcome conditional distributions given the compliers:
\begin{equation}\label{eq:cp01}
F_{Y_1|\mathcal{C}^1_{z,z'}}(y)=F_{Y_0|\mathcal{C}^0_{z,z'}}(\phi_{1,0}(y))\text{ for }y\in\mathcal{Y}^\circ.
\end{equation}
Finally, $F_{Y_0|\mathcal{C}^0_{z,z'}}$ and $F_{Y_1|\mathcal{C}^1_{z,z'}}$ are identified from Lemma \ref{lem:idcpl}, and $\phi_{1,0}(y)$ for $y\in\mathcal{Y}^\circ$ is identified as $\phi_{1,0}(y)=Q_{Y_0|\mathcal{C}^0_{z,z'}}(F_{Y_1|\mathcal{C}^1_{z,z'}}(y))$ by solving (\ref{eq:cp01}) for $\phi_{1,0}$.

In the case of multi-valued treatment settings, generally, we can compare no two treatment states on the same compliers as in (\ref{eq:cp01}).
This is because the relationships between the compliers become more complicated than (\ref{eq:eqcpl}), and the compliers do not generally coincide.
We overcome this difficulty by developing systems of relationships between two or more compliers for multiple treatment states that can be solved simultaneously for the counterfactual mappings.

\subsection{Identification in multi-valued treatment settings}\label{subsec:4.1}
In this section, we establish the identification of counterfactual mappings using the relationships of compliers when the treatment is multi-valued.
Throughout the identification analysis, we consider Example I for notation simplicity. The following argument does not rely on the settings of Example I.
We first consider a case where the treatment takes three values (then we have $\mathcal{T}=\{0,1,2\}$), and assume Assumptions \ref{asp:ivqr} and \ref{asp:mt} hold for the subset $\Lambda$ of $\mathcal{P}$.

We first introduce ``sign treatments" that characterize the type of monotonicity relationship of each pair $(z_1,z_2)$ contained in the monotonicity subset $\Lambda$.
Suppose that, for $(z_1,z_2)\in\Lambda$, there uniquely exists $t{(z_1,z_2)}\in\mathcal{T}$ such that either
\[
D_{t{(z_1,z_2)}}({z_1})\geq D_{t{(z_1,z_2)}}({z_2})\text{ and }
D_j({z_1})\leq D_j({z_2})\,\, \text{ for }j\in\mathcal{T}\setminus\{t{(z_1,z_2)}\}
\]
or
\[
D_{t{(z_1,z_2)}}({z_1})\leq D_{t{(z_1,z_2)}}({z_2})\text{ and } D_j({z_1})\geq D_j({z_2})\,\, \text{ for }j\in\mathcal{T}\setminus\{t{(z_1,z_2)}\}
\]
holds almost surely. Then we call this $t{(z_1,z_2)}$ ``sign treatment of $(z_1,z_2)$," and we state ``$(z_1,z_2)$ has a sign treatment" in this paper.

When the treatment takes three values, each $(z_1,z_2)\in\Lambda$ has a sign treatment.
This is because if the monotonicity inequalities of $(z_1,z_2)$ are all the same, then no compliers exist, and $\Lambda$ cannot contain $(z_1,z_2)$.
Suppose that $D_j({z_1})\leq D_j({z_2})$ holds almost surely for all $j=0,1,2$.
Then, $D_0({z_1})\geq D_0({z_2})$ holds almost surely because $D_1({z_1})\leq D_1({z_2})$ and $D_2({z_1})\leq D_2({z_2})$ imply $1-D_0({z_1})\leq 1-D_0({z_2})$.
Hence, $D_0({z_1})= D_0({z_2})$ holds almost surely.
Applying a similar argument to treatment states 1 and 2 gives $D_j({z_1})=D_j({z_2})$ for each $j\in\mathcal{T}$, which violates Assumption \ref{asp:mt} (iii).\footnote{When the treatment takes more than three values, each $(z_1,z_2)\in\Lambda$ may not have a sign treatment.
Let $\mathcal{T}=\{0,1,2,3\}$, and suppose that the following monotonicity inequalities hold for $(z_1,z_2)$:
\[
D_{0}({z_1})\geq D_{0}({z_2}),\,\, D_{1}({z_1})\geq D_{1}({z_2}),\,\,
D_{2}({z_1})\leq D_{2}({z_2}),\text{ and }
D_3({z_1})\leq D_3({z_2}).
\]
Then, from the definition, $(z_1,z_2)$ does not have a sign treatment.}

With the sign treatments, we employ the following assumption and assume that two different types of monotonicity relationships exist.
\begin{assumption}[Existence of different types of monotonicity relationships]\label{asp:apb.2}
In the case of $\mathcal{T} = \{0,1,2\}$, the monotonicity subset $\Lambda$ contains two pairs of instrument values $\lambda_1$ and $\lambda_2$ such that the
following condition holds:
\begin{enumerate}
\item[]
There uniquely exists $t({\lambda_1})\in\mathcal{T}$ and $t({\lambda_2})\in\mathcal{T}$ with $\lambda_i=(\lambda_{i,1},\lambda_{i,2})$ and $t({\lambda_1})\neq t({\lambda_2})$ such that
\[
D_{t({\lambda_i})}({\lambda_{i,1}})\geq D_{t({\lambda_i})}({\lambda_{i,2}})\text{ and }
D_j({\lambda_{i,1}})\leq D_j({\lambda_{i,2}})\,\, \text{ for }j\in \mathcal{T}\setminus\{t({\lambda_i})\}
\]
hold almost surely for $i=1,2$.
\end{enumerate}
\end{assumption}
\noindent Under Assumption \ref{asp:apb.2}, the monotonicity subset $\Lambda$ contains two pairs of instrument values $\lambda_1$ and $\lambda_2$ such that each pair $\lambda_i$ has a sign treatment, and the two sign treatments $t({\lambda_1})$ and $t({\lambda_2})$ are different.
Then, each $t({\lambda_i})$ characterizes a type of monotonicity relationship of $\lambda_i$.
Assumption \ref{asp:apb.2} holds in Example I.
From Table \ref{tab:mtc1} in Section \ref{subsec:3.2}, $(1,0)$ and $(2,0)$ have sign treatments $t{(1,0)}=1$ and $t{(2,0)}=2$, respectively.
Then, $(1,0)$ and $(2,0)$ induce different types of monotonicity relationships.

We can interpret Assumption \ref{asp:apb.2} as an instrument relevance condition that requires monotonic correlation between each of the two endogenous variables $D_{t({\lambda_1})}$ and $D_{t({\lambda_2})}$ and pairs of instrument values $\lambda_1$ and $\lambda_2$.
We illustrate this point with Example I.
For $i=1,2$, whether the group assignment is $i$ or 0 produces a monotonic effect only toward the field choice $D_{t{(i,0)}}=D_i$.
Compared with group 0, group $i$ additionally offers a discount only to field $i$, and only the preference for field $i$ is affected by the difference between these two group assignments.

Different types of monotonicity relationships are essential for our identification analysis.
We obtain key relationships between the compliers essential for identifying the counterfactual mappings based on the two different monotonicity relationships.
For Example I, the proof of Lemma \ref{lem:apb} in Appendix \ref{sec:a} derives the following key relationships between the compliers:
\begin{eqnarray}
F_{Y_1|\mathcal{C}^1_{0,1}}(\phi_{2,1}(y))
&=& \frac{F_{Y_0|\mathcal{C}^0_{1,0}}(\phi_{2,0}(y))P(\mathcal{C}^0_{1,0})+F_{Y_2|\mathcal{C}^2_{1,0}}(y)P(\mathcal{C}^2_{1,0})}{P(\mathcal{C}^1_{0,1})}, \label{eq:ab1} \\
F_{Y_2|\mathcal{C}^2_{0,2}}(y)
&=& \frac{F_{Y_0|\mathcal{C}^0_{2,0}}(\phi_{2,0}(y))P(\mathcal{C}^0_{2,0})+F_{Y_1|\mathcal{C}^1_{2,0}}(\phi_{2,1}(y))P(\mathcal{C}^1_{2,0})}{P(\mathcal{C}^2_{0,2})}. \label{eq:ac1}
\end{eqnarray}
Relationships \eqref{eq:ab1} and \eqref{eq:ac1} correspond to (\ref{eq:cp01}) in the binary treatment case.
Applying Lemma \ref{lem:idcpl}, all the functions in \eqref{eq:ab1} and \eqref{eq:ac1} except for the counterfactual mappings are identified.

We identify the counterfactual mappings by solving \eqref{eq:ab1} and \eqref{eq:ac1} simultaneously for each $y\in\mathcal{Y}^\circ$.
When we fix the value of $y$ in \eqref{eq:ab1} and \eqref{eq:ac1} at any $y=y^{f}\in\mathcal{Y}^\circ$, $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are the solutions to the following nonlinear simultaneous equations of two unknown variables $y_1$ and $y_0$:
\begin{eqnarray}
F_{Y_1|\mathcal{C}^1_{0,1}}(y_1)
&=& \frac{F_{Y_0|\mathcal{C}^0_{1,0}}(y_0)P(\mathcal{C}^0_{1,0})+F_{Y_2|\mathcal{C}^2_{1,0}}(y^f)P(\mathcal{C}^2_{1,0})}{P(\mathcal{C}^1_{0,1})}, \label{eq:ab1.f} \\
F_{Y_2|\mathcal{C}^2_{0,2}}(y^f)
&=& \frac{F_{Y_0|\mathcal{C}^0_{2,0}}(y_0)P(\mathcal{C}^0_{2,0})+F_{Y_1|\mathcal{C}^1_{2,0}}(y_1)P(\mathcal{C}^1_{2,0})}{P(\mathcal{C}^2_{0,2})}. \label{eq:ac1.f}
\end{eqnarray}
We have two equations to solve two unknowns, and $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are identified if the solution is unique at $y^{f}\in\mathcal{Y}^\circ$.
The two equations (\ref{eq:ab1.f}) and (\ref{eq:ac1.f}) have a unique solution for $y_1$ and $y_0$ when the conditional cdfs of the potential outcome on compliers are strictly increasing on $\mathcal{Y}^\circ$.
Assumption \ref{asp:mt} (iv) implies the strict monotonicity of the conditional cdfs, and $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are identified as follows:\footnote{See Appendix \ref{sec:a.75} for the derivation of \eqref{eq:ac3,4}.}
\begin{equation}\label{eq:ac3,4}
\phi_{2,1}(y^f)=G_{1,2}^{y^f-1}(F_{Y_2|\mathcal{C}^2_{0,2}}(y^f)) \text{ and } \phi_{2,0}(y^f)=\phi_{1,0}^{y^f}(\phi_{2,1}(y^f)),
\end{equation}
where $\phi_{1,0}^{y_f}$ with its domain $\mathcal{Y}^f\subset\mathcal{Y}$ is defined as
\begin{equation}\label{eq:A}
\phi_{1,0}^{y_f}(y)
:=Q_{Y_0|\mathcal{C}^0_{1,0}}\left(\frac{F_{Y_1|\mathcal{C}^1_{0,1}}(y)P(\mathcal{C}^1_{0,1})-F_{Y_2|\mathcal{C}^2_{1,0}}(y^f)P(\mathcal{C}^2_{1,0})}{P(\mathcal{C}^0_{1,0})}\right)\text{ for }y\in\mathcal{Y}^f,
\end{equation}
and we define a function $G_{1,2}^{y^f}$ as
\begin{equation}\label{eq:46}
G_{1,2}^{y^f}(y):=\frac{F_{Y_0|\mathcal{C}^0_{2,0}}(\phi_{1,0}^{y^f}(y))P(\mathcal{C}^0_{2,0})+F_{Y_1|\mathcal{C}^1_{2,0}}(y)P(\mathcal{C}^1_{2,0})}{P(\mathcal{C}^2_{0,2})}.
\end{equation}
Other counterfactual mappings on $\mathcal{Y}^\circ$ are inversions or compositions of $\phi_{2,1}$ and $\phi_{2,0}$, and they are also identified as closed-form expressions.
The following lemma shows the identification of the counterfactual mappings under Assumption \ref{asp:apb.2}:
\begin{lemma}[Identification of counterfactual mappings from monotonicity]\label{lem:apb.2}
Suppose that Assumptions \ref{asp:ivqr}-\ref{asp:apb.2} hold for the case of $\mathcal{T} = \{0,1,2\}$,.
Then, $\phi_{s,t}(y)$ for $y\in\mathcal{Y}^\circ$ and $s,t\in\mathcal{T}$ are identified.
\end{lemma}

We provide intuition for the identification of the counterfactual mappings.
First, the two relationships \eqref{eq:excpls} and \eqref{eq:excpls2} that are key conditions for identification are based on the following two relationships between compliers:
\begin{equation}\label{eq:excpls}
\mathcal{C}^1_{0,1}=\mathcal{C}^0_{1,0}\cup\mathcal{C}^2_{1,0}\quad\text{and}\quad\mathcal{C}^0_{1,0}\cap\mathcal{C}^2_{1,0}=\varnothing,
\end{equation}
\begin{equation}\label{eq:excpls2}
\mathcal{C}^2_{0,2}=\mathcal{C}^0_{2,0}\cup\mathcal{C}^1_{2,0}\quad\text{and}\quad\mathcal{C}^0_{2,0}\cap\mathcal{C}^1_{2,0}=\varnothing.
\end{equation}
Relationships \eqref{eq:excpls} and \eqref{eq:excpls2} correspond to \eqref{eq:eqcpl} in the binary treatment case.
We show that two monotonicity relationships for $(1,0)$ and $(2,0)$ generate \eqref{eq:excpls} and \eqref{eq:excpls2}.

Figure \ref{fig:ex1.2} visualizes these compliers relationships \eqref{eq:excpls} and \eqref{eq:excpls2} implied by the separable index model introduced in Section \ref{subsec:3.2}.
Figure \ref{fig:ex1.2} (a) visualizes how the compliers are generated from
a shift in the instrument from $Z=0$ to $Z=1$.
First, because $\mathcal{C}^1_{0,1}$ compliers are driven by individuals who find field 1 more attractive and leave from either $T=0$ group or $T=2$ group, we have
\begin{equation}\label{eq:19}
\mathcal{C}^1_{0,1}=\{D_1({1})=1,D_0({0})=1\}\cup\{D_1({1})=1,D_2({0})=1\}.
\end{equation}
Obviously, the sets $\{D_1({1})=1,D_0({0})=1\}$ and $\{D_1({1})=1,D_2({0})=1\}$ are disjoint.
Second, we show that
\begin{equation}\label{eq:case1}
\mathcal{C}^0_{1,0}=\{D_1({1})=1,D_0({0})=1\}\quad\text{and}\quad\mathcal{C}^2_{1,0}=\{D_1({1})=1,D_2({0})=1\}.
\end{equation}
To see this, note that $\mathcal{C}^0_{1,0}$ compliers, which consist of individuals leaving from the $T=0$ group, are entirely driven by those who find field 1 more attractive.
This is because, from the restrictions on the cost functions \eqref{eq:cf1} and \eqref{eq:cf2}, a shift in the instrument from $Z=0$ to $Z=1$ decreases the cost of choosing field 1 but does not change the cost of choosing field 2.
Analogously, $\mathcal{C}^2_{1,0}$ compliers, which consist of individuals leaving from the $T=2$ group, are also entirely driven by those who find field 1 more attractive.
Therefore, \eqref{eq:excpls} holds from (\ref{eq:19}) and (\ref{eq:case1}).
By analogous logic, as visualized in Figure \ref{fig:ex1.2} (b), the compliers generated from a shift in the instrument from $Z=0$ to $Z=2$ satisfy relationship \eqref{eq:excpls2}.

\begin{figure}[htbp]
\centering
\begin{minipage}[b]{0.45\linewidth}
\centering
\scalebox{0.95}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$V_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$V_2$};

\coordinate(P0)at(3.5,3.5);
\draw[thick]($(XS)!(P0)!(XL)$)node[below]{$\mu_1(0)$}--(P0)--($(YS)!(P0)!(YL)$);
\draw[thick,domain=3.5:6]plot(\x,\x);

\coordinate(P1)at(2,3.5);
\draw[thick]($(XS)!(P1)!(XL)$)node[below]{$\mu_1(1)$}--(P1)--($(YS)!(P1)!(YL)$)node[left]{$\mu_2(0)$};
\draw[thick,domain=2:4.5]plot(\x,\x+1.5);

\draw(1,2)node{$D_0(1)=1$};
\draw(4.5,2)node{$D_1(0)=1$};
\draw(1.5,5)node{$D_2(1)=1$};
\draw(2.7,1.5)node{$\mathcal{C}^0_{1,0}$};
\draw(4.2,5)node{$\mathcal{C}^2_{1,0}$};
\draw(4.9,3.1)node{$\mathcal{C}^1_{0,1}$};
\draw[->,semithick](4.9,3.4)--(3.8,4.4);
\draw[->,semithick](4.9,2.8)--(3.1,2.3);
\end{tikzpicture}
}
\subcaption{Compliers between $Z=0$ and 1}
\end{minipage}
\begin{minipage}[b]{0.45\linewidth}
\centering
\scalebox{0.95}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$V_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$V_2$};

\coordinate(P0)at(3.5,3.5);
\draw[thick]($(XS)!(P0)!(XL)$)--(P0)--($(YS)!(P0)!(YL)$)node[left]{$\mu_2(0)$};
\draw[thick,domain=3.5:6]plot(\x,\x);

\coordinate(P2)at(3.5,2);
\draw[thick]($(XS)!(P2)!(XL)$)node[below]{$\mu_1(0)$}--(P2)--($(YS)!(P2)!(YL)$)node[left]{$\mu_2(2)$};
\draw[thick,domain=3.5:6]plot(\x,\x-1.5);

\draw(2,1)node{$D_0(2)=1$};
\draw(5,1.5)node{$D_1(2)=1$};
\draw(2,5.3)node{$D_2(0)=1$};
\draw(2,2.7)node{$\mathcal{C}^0_{2,0}$};
\draw(4.5,3.8)node{$\mathcal{C}^1_{2,0}$};
\draw(2.5,4.5)node{$\mathcal{C}^2_{0,2}$};
\draw[->,semithick](2.5,4.2)--(2,3.1);
\draw[->,semithick](2.9,4.5)--(4.5,4.2);
\end{tikzpicture}
}
\subcaption{Compliers between $Z=0$ and 2}
\end{minipage}
\caption{Visualization of the compliers}
\label{fig:ex1.2}
\end{figure}

Next, we provide an intuition for identifying $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ for each $y^{f}\in\mathcal{Y}^\circ$ as the unique solution to the two equations (\ref{eq:ab1.f}) and (\ref{eq:ac1.f}) of two unknowns $y_1$ and $y_0$.
We compare the set of (an infinite number of) solutions to each of these equations.
The two equations have a unique solution when these two sets intersect at a point.
Figure \ref{fig:ex1.3} visualizes the solutions to (\ref{eq:ab1.f}) and (\ref{eq:ac1.f}) respectively in the two-dimensional space of $(y_1,y_0)$.
First, the set of solutions to (\ref{eq:ab1.f}) form a strictly increasing relationship from the strict monotonicity of $F_{Y_1|\mathcal{C}^1_{0,1}}$ and $F_{Y_0|\mathcal{C}^0_{1,0}}$.
To see this, suppose that both $(y^-_1,y^-_0)$ and $(y^+_1,y^+_0)$ are solutions to (\ref{eq:ab1.f}).
If we have $y^-_1 < y^+_1$, we also have $F_{Y_1|\mathcal{C}^1_{0,1}}(y^-_1)<F_{Y_1|\mathcal{C}^1_{0,1}}(y^+_1)$ from the strict monotonicity of $F_{Y_1|\mathcal{C}^1_{0,1}}$.
Then, because $F_{Y_1|\mathcal{C}^1_{0,1}}$ and $F_{Y_0|\mathcal{C}^0_{1,0}}$ are on the left and right-hand sides of (\ref{eq:ab1.f}), respectively, we need $F_{Y_0|\mathcal{C}^0_{1,0}}(y^-_0)<F_{Y_0|\mathcal{C}^0_{1,0}}(y^+_0)$ for (\ref{eq:ab1.f}) to hold, and this implies that $y^-_0 < y^+_0$ from the strict monotonicity of $F_{Y_0|\mathcal{C}^0_{1,0}}$.
Next, from an analogous argument with the strict monotonicity of $F_{Y_1|\mathcal{C}^1_{2,0}}$ and $F_{Y_0|\mathcal{C}^0_{2,0}}$, the set of solutions to (\ref{eq:ac1.f}) form a strictly decreasing relationship, where both $F_{Y_1|\mathcal{C}^1_{2,0}}$ and $F_{Y_0|\mathcal{C}^0_{2,0}}$ are on the right-hand side of (\ref{eq:ac1.f}).
Hence, because strictly increasing and strictly decreasing relationships intersect only once, (\ref{eq:ab1.f}) and (\ref{eq:ac1.f}) have a unique solution at $(\phi_{2,1}(y^f),\phi_{2,0}(y^f))$.

\begin{figure}[htbp]
\centering
\scalebox{1.0}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$y_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$y_0$};

\draw[red,thick,domain=0.5:5.5]plot(\x,\x)node[above]{Solutions to (\ref{eq:ab1.f})};
\draw[blue,thick,domain=0.5:5.5]plot(\x,-\x+6)node[above right]{Solutions to (\ref{eq:ac1.f})};

\coordinate(P)at(3,3);
\fill(P)circle(0.075);
\draw[dashed,thick]($(XS)!(P)!(XL)$)node[below]{$\phi_{2,1}(y^f)$}--(P)--($(YS)!(P)!(YL)$)node[left]{$\phi_{2,0}(y^f)$};
\coordinate(P-)at(2,2);
\draw[dashed,thick]($(XS)!(P-)!(XL)$)node[below]{$y^-_1$}--(P-)--($(YS)!(P-)!(YL)$)node[left]{$y^-_0$};
\coordinate(P+)at(4,4);
\draw[dashed,thick]($(XS)!(P+)!(XL)$)node[below]{$y^+_1$}--(P+)--($(YS)!(P+)!(YL)$)node[left]{$y^+_0$};
\end{tikzpicture}
}
\caption{Identification image when $k=2$}
\label{fig:ex1.3}
\end{figure}

Next, we generalize the above argument for identifying the counterfactual mappings to the case of arbitrary $k$.
With the sign treatments, we employ the following assumption and assume that $k$ different types of monotonicity relationships exist.
\begin{assumption}[Existence of different types of monotonicity relationships]\label{asp:apb}
The monotonicity subset $\Lambda$ contains $k$ pairs of instrument values $\lambda_1,\ldots,\lambda_k$ such that the
following condition holds:
\begin{enumerate}
\item[]
For $i=1,\ldots,k$, there uniquely exists $t({\lambda_i})\in\mathcal{T}$ with $\lambda_i=(\lambda_{i,1},\lambda_{i,2})$ and $t({\lambda_i})\neq t({\lambda_j})$ for $i\neq j$ such that
\[
D_{t({\lambda_i})}({\lambda_{i,1}})\geq D_{t({\lambda_i})}({\lambda_{i,2}})\text{ and }
D_j({\lambda_{i,1}})\leq D_j({\lambda_{i,2}})\,\, \text{ for }j\in \mathcal{T}\setminus\{t({\lambda_i})\}
\]
hold almost surely.
\end{enumerate}
\end{assumption}
\noindent Under Assumption \ref{asp:apb}, the monotonicity subset $\Lambda$ contains $k$ pairs of instrument values $\lambda_1,\ldots,\lambda_k$ such that each pair $\lambda_i$ has a sign treatment, and the sign treatments $t({\lambda_i})$'s are all different.
Then, each $t({\lambda_i})$ characterizes a type of monotonicity relationship of $\lambda_i$.
As in the case of $k=2$, Assumption \ref{asp:apb} holds in Example I.
From Table \ref{tab:mtc2} in Section \ref{subsec:3.2}, for $i=1,\ldots,k$, $(i,0)$ has a sign treatment $t{(i,0)}=i$. Then, for $i\neq j$, $(i,0)$ and $(j,0)$ induce different types of monotonicity relationships.

Under Assumption \ref{asp:apb}, we obtain key relationships between the compliers essential for identifying the counterfactual mappings.
For Example I, the proof of Lemma \ref{lem:apb} in Appendix \ref{sec:a} derives the following key relationships between the compliers:
\begin{equation}\label{eq:4.3.4}
F_{Y_i|\mathcal{C}^i_{0,i}}(\phi_{k,i}(y))
=\frac{\sum_{j\neq i}F_{Y_j|\mathcal{C}^j_{i,0}}(\phi_{k,j}(y))P(\mathcal{C}^j_{i,0})}{P(\mathcal{C}^i_{0,i})}\text{ for }y\in\mathcal{Y}^\circ\text{ and }i=1,\ldots,k.
\end{equation}
The $k$ relationships of \eqref{eq:4.3.4} correspond to \eqref{eq:ab1} and \eqref{eq:ac1} in the case of $k=2$.
Applying Lemma \ref{lem:idcpl}, all the functions in (\ref{eq:4.3.4}) except for the counterfactual mappings are identified.
When we fix (\ref{eq:4.3.4}) at $y^{f}\in\mathcal{Y}^\circ$, the $k$ equations of (\ref{eq:4.3.4}) constitute nonlinear simultaneous equations of $\phi_{k,i}(y^{f})$ for $j=0,\ldots,k-1$.
We identify these values by solving the $k$ equations of (\ref{eq:4.3.4}) simultaneously at $y^{f}$.

The following lemma shows the identification of the counterfactual mappings under Assumption \ref{asp:apb}:
\begin{lemma}[Identification of counterfactual mappings from monotonicity]\label{lem:apb}
Suppose that Assumptions \ref{asp:ivqr}-\ref{asp:apb} hold. Then, $\phi_{s,t}(y)$ for $y\in\mathcal{Y}^\circ$ and $s,t\in\mathcal{T}$ are identified.
\end{lemma}
\begin{remark}[]\label{rem:apb}
In the proof of Lemma \ref{lem:apb} in Appendix \ref{sec:a}, we do not derive the closed-form expressions of $\phi_{s,t}$'s for identifying them.
Appendix \ref{sec:a.25} provides the closed-form expressions of $\phi_{s,t}$'s for the general $k\in\mathcal{T}$.
\end{remark}
\noindent As discussed in Section \ref{subsec:2.2}, $E[Y_s]$ and $Q_{Y_s} (\tau )$ for $\tau\in(0,1)$ are identified if $\phi_{s,t}(\cdot)$ for $t\in\mathcal{T}$ are identified.
Hence, we obtain the following theorem:
\begin{theorem}[Identification of potential outcome cdfs and means from monotonicity]\label{thm:apb}
Suppose that Assumptions \ref{asp:ivqr}-\ref{asp:apb} hold. Then, $E[Y_s]$ and $Q_{Y_s} (\tau )$ for $\tau\in(0,1)$ and $s\in\mathcal{T}$ are identified.
\end{theorem}
This result is interesting because the proposed sufficient condition is economically interpretable.
We do not need to interpret the numerical conditions on the distribution of the outcome variable when this condition is satisfied.
This fact may be helpful when designing a social experiment for which the outcome data will be collected later.

\subsection{Comparison with Chernozhukov and Hansen (2005) and violation of our assumptions}\label{subsec:4.4}
In this section, we compare Assumption \ref{asp:apb} (or Assumption \ref{asp:apb.2} in the case of $k=2$) with the identification condition of \cite{CH2005} and discuss the case where Assumption \ref{asp:apb} is violated.
For the binary treatment case, the closed-form expression of \cite{Wuthrich2019} is also valid when the full rank condition in \cite{CH2005} holds for all the values in $\mathcal{Y}$.
\cite{VX2017} establish an identification condition weaker than the monotonicity assumption for the binary treatment case and show that the identification condition of \cite{CH2005} implies that condition.

In the multi-valued treatment setting, we can show that Assumption \ref{asp:apb} implies the full rank conditions in \cite{CH2005} under Assumptions \ref{asp:ivqr} and \ref{asp:mt} and some differentiability assumptions.\footnote{Appendix \ref{sec:e} proves this statement.}
On the other hand, our identification condition does not require any differentiability assumption.
Assumption \ref{asp:apb} may be violated even when the identification condition of \cite{CH2005} is satisfied.\footnote{Appendix \ref{sec:e.2} provides a numerical example where the full rank condition in \cite{CH2005} holds but Assumption \ref{asp:apb} is violated.}
The full rank condition in \cite{CH2005} only requires that the moment equations \eqref{eq:ti} are uniquely solved.
Assumption \ref{asp:apb} clarifies the behavioral patterns of each individual for identifying the treatment effects.
Assumption \ref{asp:apb} also implies additional testable restrictions compared with \cite{CH2005}.\footnote{Appendix \ref{sec:e.3} discusses these additional restrictions. These restrictions are equally difficult to check statistically compared with the full rank conditions in \cite{CH2005} for the binary treatment case.}

\begin{remark}\label{rem:4.4}
The full rank conditions in \cite{CH2005} requires the instrument to take the same number of values as the treatment, which is also required in our settings.
\cite{Feng2019} and \cite{caetano2020} establish the identification of treatment effects using observed covariates when the instrument has smaller support than the treatment.
Our results do not rely on the existence of observed covariates.
\end{remark}

Next, we discuss the case where Assumption \ref{asp:apb} is violated.
Even when the unordered monotonicity assumption holds, Assumption \ref{asp:apb} is violated if the sign treatments are the same for all the instrument pairs.
We consider the $k=2$ case and let the monotonicity inequalities summarized in Table \ref{tab:mtn} hold almost surely.
In Example I, this may happen when individuals are randomly assigned to one of the three groups, where the cost of choosing field $2$ is decreased for all groups; Group 2 has the largest decrease in cost, and Group 0 has the smallest.
\begin{table}[htb]
\caption{Monotonicity inequalities when sign treatments are all the same}
\label{tab:mtn}
\begin{center}
\begin{tabular}{|c|c c c c|} \hline
\multicolumn{1}{|c|}{} & \multicolumn{4}{|c|}{$\mathcal{T}$} \\ \hline
\multicolumn{1}{|c|}{} & & $0$ & $1$ & $2$ \\
& $(1,0)$ & $\,\,\,\,\,\,\,\,D_0({1})\leq D_0({0})\,\,\,\,\,\,\,\,$ & $\,\,\,\,\,\,\,\,D_1({1})\leq D_1({0})\,\,\,\,\,\,\,\,$ & $\,\,\,\,\,\,\,\,D_2({1})\textcolor{red}{\geq} D_2({0})\,\,\,\,\,\,\,\,$ \\
$\mathcal{P}$ & $(2,0)$ & $\,\,D_0({2})\leq D_0({0})\,\,$ & $\,\,D_1({2})\leq D_1({0})\,\,$ & $\,\,D_2({2})\textcolor{red}{\geq} D_2({0})\,\,$ \\
& $(2,1)$ & $\,\,D_0({2})\leq D_0({1})\,\,$ & $\,\,D_1({2})\leq D_1({1})\,\,$ & $\,\,D_2({2})\textcolor{red}{\geq} D_2({1})\,\,$ \\ \hline
\end{tabular}
\end{center}
\end{table}
Then, Assumptions 1 and 2 hold, and the unordered monotonicity assumption holds because Assumption 2 (ii) holds for all the instrument pairs.
However, the sign treatments of the instrument pairs are $t{(1,0)}=t{(2,0)}=t{(2,1)}=2$, and Assumption \ref{asp:apb.2} does not hold because all the instrument pairs induce the same type of monotonicity relationship.

Under this monotonicity assumption, we obtain the following two equations from the monotonicity relationships of $(1,0)$ and $(2,0)$ as we obtain (\ref{eq:ab1}) and (\ref{eq:ac1}) under Assumption \ref{asp:apb.2}:
\begin{eqnarray}
F_{Y_2|\mathcal{C}^2_{0,1}}(y)
&=& \frac{F_{Y_0|\mathcal{C}^0_{1,0}}(\phi_{2,0}(y))P(\mathcal{C}^0_{1,0})+F_{Y_1|\mathcal{C}^1_{1,0}}(\phi_{2,1}(y))P(\mathcal{C}^1_{1,0})}{P(\mathcal{C}^2_{0,1})}, \label{eq:all1} \\
F_{Y_2|\mathcal{C}^2_{0,2}}(y)
&=& \frac{F_{Y_0|\mathcal{C}^0_{2,0}}(\phi_{2,0}(y))P(\mathcal{C}^0_{2,0})+F_{Y_1|\mathcal{C}^1_{2,0}}(\phi_{2,1}(y))P(\mathcal{C}^1_{2,0})}{P(\mathcal{C}^2_{0,2})}. \label{eq:all2}
\end{eqnarray}
Whether $\phi_{2,1}(y)$ and $\phi_{2,0}(y)$ are identified from the two equations (\ref{eq:all1}) and (\ref{eq:all2}) depends on the numerical conditions on the conditional cdfs.
For each $y^f\in\mathcal{Y}^\circ$, $\phi_{2,1}(y^f)$ and $\phi_{2,0}(y^f)$ are the solutions to the following nonlinear simultaneous equations of two unknown variables $y_1$ and $y_0$:
\begin{eqnarray}
F_{Y_2|\mathcal{C}^2_{0,1}}(y^f)
&=& \frac{F_{Y_0|\mathcal{C}^0_{1,0}}(y_0)P(\mathcal{C}^0_{1,0})+F_{Y_1|\mathcal{C}^1_{1,0}}(y_1)P(\mathcal{C}^1_{1,0})}{P(\mathcal{C}^2_{0,1})}, \label{eq:all1.f} \\
F_{Y_2|\mathcal{C}^2_{0,2}}(y^f)
&=& \frac{F_{Y_0|\mathcal{C}^0_{2,0}}(y_0)P(\mathcal{C}^0_{2,0})+F_{Y_1|\mathcal{C}^1_{2,0}}(y_1)P(\mathcal{C}^1_{2,0})}{P(\mathcal{C}^2_{0,2})}. \label{eq:all2.f}
\end{eqnarray}
As in Section \ref{subsec:4.1}, we compare the set of solutions to each of these two equations, and Figure \ref{fig:ex1.4} visualizes the solutions to (\ref{eq:all1.f}) and (\ref{eq:all2.f}), respectively, in the two-dimensional space of $(y_1,y_0)$.
Notice that, from the strict monotonicity of the conditional cdfs, the solutions to (\ref{eq:all1.f}) and (\ref{eq:all2.f}) have both strictly decreasing relationships between $y_1$ and $y_0$, where both the conditional cdfs of $Y_1$ and $Y_0$ are on the same side of (\ref{eq:all1.f}) and (\ref{eq:all2.f}).
Hence, when these two strictly decreasing relationships intersect only once,
the two equations (\ref{eq:all1.f}) and (\ref{eq:all2.f}) have a unique solution at $(\phi_{2,1}(y^f),\phi_{2,0}(y^f))$.
In this case, we can show that the full rank conditions in \cite{CH2005} hold under Assumptions \ref{asp:ivqr} and \ref{asp:mt} and some differentiability assumptions.\footnote{Appendix \ref{sec:e.2} proves this statement.}

\begin{figure}[htbp]
\centering
\scalebox{1.0}{
\begin{tikzpicture}
\coordinate(O)at(0,0);
\coordinate(XS)at(0,0);
\coordinate(XL)at(6,0);
\coordinate(YS)at(0,0);
\coordinate(YL)at(0,6);
\draw[->,>=stealth,semithick](XS)--(XL)node[below left]{$y_1$};
\draw[->,>=stealth,semithick](YS)--(YL)node[below left]{$y_0$};

\draw[red,thick,domain=1.25:4.75]plot(\x,-1.5*\x+7.5)node[above right]{Solutions to (\ref{eq:all1.f})};
\draw[blue,thick,domain=0.5:5.5]plot(\x,-0.5*\x+4.5)node[above right]{ Solutions to (\ref{eq:all2.f})};

\coordinate(P)at(3,3);
\fill(P)circle(0.075);
\draw[dashed,thick]($(XS)!(P)!(XL)$)node[below]{$\phi_{2,1}(y^f)$}--(P)--($(YS)!(P)!(YL)$)node[left]{$\phi_{2,0}(y^f)$};

\end{tikzpicture}
}
\caption{Identification image when sign treatments are all the same}
\label{fig:ex1.4}
\end{figure}

\subsection{Identification of the local treatment effects}\label{subsec:4.5}
In this section, we identify the local treatment effects in the multi-valued treatment setting.
Suppose that the monotonicity inequalities hold for $(z,z')$, and that $D_t(z)\leq D_t(z')$ holds almost surely for treatment level $t$.
For two different treatment levels $t$ and $t'$ and instrument values $z$ and $z'$, the subpopulation $\{D_{t'}(z)=1,D_t(z')=1\}$ changes the treatment choice from $t'$ to $t$ if the instrument value changes from $z$ to $z'$.
The LATE that compares treatment states $t$ and $t'$ conditional on $\{D_{t'}(z)=1,D_t(z')=1\}$ is $E[Y_t|D_{t'}(z)=1,D_t(z')=1]-E[Y_{t'}|D_{t'}(z)=1,D_t(z')=1]$.
The local quantile treatment effect (LQTE) conditional on $\{D_{t'}(z)=1,D_t(z')=1\}$ is $Q_{Y_t|D_{t'}(z)=1,D_t(z')=1}(\tau)-Q_{Y_{t'}|D_{t'}(z)=1,D_t(z')=1}(\tau)$ where $\tau\in(0,1)$.

First, we briefly review the identification studies of local treatment effects in multi-valued treatment settings.
From the definition, local treatment effects are treatment effects conditional on the compliers when the treatment is binary.
As shown in \cite{Imbens1994}, LATEs are identified as the standard two-stage least squares (2SLS) estimands under the monotonicity assumption.
\cite{Heckman2018} establish the identification of the conditional means of the potential outcomes on the compliers.
However, identifying the local treatment effects is more complex in multi-valued treatment settings because the relationship between the compliers becomes more complicated than in the binary treatment case.
\cite{Angrist1995} show that the 2SLS estimand generally represents only a weighted average of LATEs with a binary instrument, even under certain monotonicity assumption.\footnote{\cite{kline2016evaluating} and \cite{hull2018isolateing} also derive related results in the case of a binary instrument.}
\cite{kirkeboen2016field} and \cite{Mountjoy2019} demonstrate that a similar result holds for 2SLS estimands even when there are as many instruments as treatments.
\cite{Heckman2006} identify the LATEs generalized to compare treatment state $t$ and the set of other states with continuous instruments.
\cite{Heckman2006} also establish identification on more various subpopulations that are identified with continuous instruments.
For a more general class of selection models,
\cite{lee2018} and \cite{Mountjoy2019} show similar identification results for the ATEs that compare two different treatment states, $t$ and $t'$.

This section establishes the identification of the LATEs and LQTEs that compare two different treatment states, $t$ and $t'$, under our monotonicity assumption.
We require that the outcome variable is continuously distributed with additional assumptions, such as the rank similarity assumption, necessary for identifying the counterfactual mappings.

Note that, if $(z,z')$ has a sign treatment $t_{(z,z')}=t'$, then $P(\mathcal{C}_{z,z'}^t)=P(D_{t'}(z)=1,D_t(z')=1)$ holds for $t\in\mathcal{T}\setminus\{t'\}$ as we obtain (\ref{eq:case1}) for Example I in Section \ref{subsec:4.1}.
Under the rank similarity assumption, identifying the counterfactual mappings leads to identifying the treatment effects conditional on the compliers.
The following lemma shows the identification of the conditional distribution of $Y_{t'}$ given $\mathcal{C}_{z,z'}^t$ as well as that of $Y_{t}$ given $\mathcal{C}_{z,z'}^t$:

\begin{lemma}[Identification of potential outcome conditional cdfs and means given the compliers]\label{lem:apb2}
Suppose that Assumptions \ref{asp:ivqr}-\ref{asp:apb} hold, and that $P(D_t(z)\leq D_t({z'}))=1$ and $P(\mathcal{C}^t_{z,z'})>0$ hold for $t\in\mathcal{T}$ and $(z,z')\in\mathcal{P}$.
Then, for all $t'\in\mathcal{T}$, $F_{Y_{t'}|\mathcal{C}_{z,z'}^t }(y)$ for $y\in\mathcal{Y}^\circ$ and $E[Y_{t'}|\mathcal{C}_{z,z'}^t]$ can be expressed as
\begin{equation}\label{eq:l4.1}
F_{Y_{t'}|\mathcal{C}_{z,z'}^t }(y)=F_{Y_{t}|\mathcal{C}_{z,z'}^t }(\phi_{t',t}(y))
\end{equation}
and
\begin{equation}\label{eq:l4.2}
E[Y_{t'}|\mathcal{C}_{z,z'}^t]=E[\phi_{t,t'}(Y_t)|\mathcal{C}_{z,z'}^t],
\end{equation}
and $E[Y_{t'}|\mathcal{C}_{z,z'}^t]$ and $Q_{Y_{t'}|\mathcal{C}_{z,z'}^t } (\tau )$ for $\tau\in(0,1)$ are identified.
\end{lemma}

\noindent With Lemma \ref{lem:apb2} at hand, we obtain the following theorem that shows the identification of local treatment effects under our assumptions:

\begin{theorem}[Identification of local potential outcome cdfs and means]\label{thm:apb2}
Assume that Assumptions \ref{asp:ivqr}-\ref{asp:apb} hold. Then, for each $t,t'\in\mathcal{T}$, there exists $(z,z')\in\mathcal{P}$ such that  $E[Y_s|D_{t'}(z)=1,D_t(z')=1]$ and $Q_{Y_s|D_{t'}(z)=1,D_t(z')=1} (\tau )$ for $\tau\in(0,1)$ and $s\in\{t,t'\}$ are identified.
\end{theorem}

\subsection{Identification in real-world examples}\label{sec:5}
In this section, we discuss the content of our assumptions and the applicability of our identification result for Examples II and III in Section \ref{subsec:2.2ex}.
In each example, we confirm the assumptions for Theorems \ref{thm:apb} and \ref{thm:apb2} with monotonicity inequalities.
Suppose that the treatment $T$ and the instrument $Z$ are sufficiently correlated, and we assume Assumption \ref{asp:ivqr} and Assumptions \ref{asp:mt} (i), (iii), and (iv) unless stated otherwise.

\subsubsection{Example II}\label{subsec:3.3.2}
We take an approach similar to \cite{Mountjoy2019} for the model and economic analysis.
\cite{Mountjoy2019} also employs the discrete choice index model to motivate his monotonicity assumption for continuous instruments.
Let $Y$ denote student outcome.
Let the treatment $T$ denote college decision that takes values on $\{0,2,4\}$, where $T=0$ denotes no college, while $T=2$ and $T=4$ denote starting college at a two-year institution and a four-year institution, respectively.
Let the two binary instruments $Z_2$ and $Z_4$ that take values on $\{0,1\}$ represent the distance to the nearest college, and $Z_2$ equals 1 when the nearest two-year college is within a $d_2$ km radius, whereas $Z_4$ equals 1 when the nearest four-year college is within a $d_4$ km radius.
As discussed in Section \ref{subsec:3.2}, the discrete choice index model generates the monotonicity inequalities.
As in \cite{Mountjoy2019}, we define the indirect utilities for each treatment option as follows:
\begin{eqnarray*}
I_0 &=& 0, \\
I_2 &=& V_2 - \mu_2(Z_2), \\
I_4 &=& V_4 - \mu_4(Z_4),
\end{eqnarray*}
where we assume strict monotonicity for the cost functions as follows:
\begin{eqnarray}
\mu_2(0) >\mu_2(1) \quad\text{and}\quad \mu_4(0) >\mu_4(1). \label{eq:cft1}
\end{eqnarray}
This requirement is natural because students who live near a college will have a smaller cost for choosing that college.
The strict monotonicity of the cost functions generates the monotonicity inequalities summarized in Table \ref{tab:tyc} for $z_2,z_4\in\{0,1\}$.
\begin{table}[htb]
\caption{Monotonicity inequalities of Example II}\label{tab:tyc}
\begin{center}
\scalebox{0.95}{
\begin{tabular}{|c|c c c c|} \hline
\multicolumn{1}{|c|}{} & \multicolumn{4}{|c|}{$\mathcal{T}$} \\ \hline
\multicolumn{1}{|c|}{} & & $0$ & $2$ & $4$ \\
& $\{(1,z_4),(0,z_4)\}$ & $\,D_0(0,z_4)\leq D_0(1,z_4)\,$ & $\,D_2(0,z_4)\textcolor{red}{\geq} D_2(1,z_4)\,$ & $\,D_4(0,z_4)\leq D_4(1,z_4)\,$ \\
$\mathcal{P}$ & $\{(z_2,1),(z_2,0)\}$ & $\,D_0(z_2,0)\leq D_0(z_2,1)\,$ & $\,D_2(z_2,0)\leq D_2(z_2,1)\,$ & $\,D_4(z_2,0)\textcolor{red}{\geq} D_4(z_2,1)\,$ \\ \hline
\end{tabular}
}
\end{center}
\end{table}
\cite{Mountjoy2019} provides the same type of inequalities as the monotonicity inequalities in Table \ref{tab:tyc} with continuous instruments.
The monotonicity inequalities in Table \ref{tab:tyc} with $z_2=z_4=0$ are the same as those for $(1,0)$ and $(2,0)$ in Example I in Section \ref{subsec:3.2} if we replace the values that $T$ and $Z$ take from $\{0,2,4\}$ to $\{0,1,2\}$ for $T$, and from $\{(0,0),(1,0),(0,1)\}$ to $\{0,1,2\}$ for $Z$.
Example II may have more monotonicity inequalities than Example I because $z_2$ and $z_4$ take values of either 0 or 1.

From Table \ref{tab:tyc}, Assumption \ref{asp:mt} (ii) holds for $\{(1,z_4),(0,z_4)\},\{(z_2,1),(z_2,0)\}\in\Lambda$.
From Table \ref{tab:tyc}, $\{(0,z_4),(1,z_4)\}$ and $\{(z_2,0),(z_2,1)\}$ have different sign treatments $t{\{(0,z_4),(1,z_4)\}}=2$ and $t{\{(z_2,0),(z_2,1)\}}=4$, respectively.
Therefore, Assumption \ref{asp:apb} holds, and from Theorems \ref{thm:apb} and \ref{thm:apb2}, the treatment effects are identified as closed-form expressions.

\subsubsection{Example III}\label{subsec:3.3.3}
We take an approach similar to \cite{Pinto2015} for the model and economic analysis.
Let $Y$ denote the outcome of interest that is continuously distributed.
Let the treatment $T$ denote the relocation decision at the intervention onset, where $T=0$ denotes no relocation, which is equivalent to choosing high poverty neighborhood, $T=1$ denotes medium-poverty neighborhood relocation, and $T=2$ denotes low-poverty neighborhood relocation.
Let the instrument $Z$ represent voucher assignment that takes values on $\mathcal{Z}=\{a,b,c\}$, where $Z=a$ denotes no voucher (control group),
$Z=b$ denotes the Section 8 voucher, and $Z=c$ denotes the experimental voucher.

As discussed in Section \ref{subsec:3.2}, the discrete choice index model generates the monotonicity inequalities.
As in Example I, we assume the following relationships for the cost functions under additive separability of the utility functions:
\begin{eqnarray}
\mu_1(a) &=& \mu_1(c)>\mu_1(b), \label{eq:cfm1} \\
\mu_2(a) &>& \mu_2(b)=\mu_1(c). \label{eq:cfm2}
\end{eqnarray}
Relationships (\ref{eq:cfm1}) and (\ref{eq:cfm2}) can be interpreted in the same way as (\ref{eq:cf1}) and (\ref{eq:cf2}) in Example I.
Applying these restrictions to the cost functions generates the monotonicity inequalities summarized in Table \ref{tab:mto}.
\begin{table}[htb]
\caption{Monotonicity inequalities of Example III}\label{tab:mto}
\begin{center}
\begin{tabular}{|c|c c c c|} \hline
\multicolumn{1}{|c|}{} & \multicolumn{4}{|c|}{$\mathcal{T}$} \\ \hline
\multicolumn{1}{|c|}{} & & $0$ & $1$ & $2$ \\
& $(c,a)$ & $\,\,\,\,\,\,\,\,D_0({c})\leq D_0({a})\,\,\,\,\,\,\,\,$ & $\,\,\,\,\,\,\,\,D_1({c})\leq D_1({a})\,\,\,\,\,\,\,\,$ & $\,\,\,\,\,\,\,\,D_2({c})\textcolor{red}{\geq} D_2({a})\,\,\,\,\,\,\,\,$ \\
$\mathcal{P}$ & $(b,c)$ & $\,\,D_0({b})\leq D_0({c})\,\,$ & $\,\,D_1({b})\textcolor{red}{\geq} D_1({c})\,\,$ & $\,\,D_2({b})\leq D_2({c})\,\,$ \\ \hline
\end{tabular}
\end{center}
\end{table}
\cite{Pinto2015} provides similar relationships as (\ref{eq:cfm1}) and (\ref{eq:cfm2}) with budget sets and choice restrictions equivalent to the inequalities in Table \ref{tab:mto}.
From Table \ref{tab:mto}, Assumption \ref{asp:mt} (ii) holds for $\Lambda=\{(c,a),(b,c)\}$.\footnote{\cite{Pinto2015} further assumes that a neighborhood is a normal good and generates monotonicity inequalities in addition to those in Table \ref{tab:mto}.
Assumptions of \cite{Pinto2015} lead to part (A2) of Assumption \ref{asp:mt} for all the pairs in $\mathcal{P}$, and the unordered monotonicity assumption holds.
Our assumptions are weaker than those in \cite{Pinto2015} but are sufficient to identify the treatment effects.}
From Table \ref{tab:mto}, $(c,a)$ and $(b,c)$ have different sign treatments $t{(c,a)}=2$ and $t{(b,c)}=1$, respectively.
Therefore, Assumption \ref{asp:apb} holds, and from Theorems \ref{thm:apb} and \ref{thm:apb2}, the treatment effects are identified as closed-form expressions.

\section{Conclusion}\label{sec:6}
In this paper, we establish sufficient conditions for the identification of the treatment effects when the treatment is discrete and endogenous. We show that an appropriately constructed monotonicity assumption is sufficient, and this condition is economically interpretable. We also derive closed-form expressions of the identified treatment effects.

For the estimation procedure, \cite{Wuthrich2019} constructs an estimator in the binary case by semiparametric estimation of observable conditional cdfs, qfs, and probabilities and plugging them into his closed-form expression. A similar approach could be applied to our closed-form expressions.

Alternatively, especially for the estimation of the QTE, we can apply the existing estimation methods under structural quantile models based on the GMM objective function after checking our identification conditions.
For estimation based on the GMM objective function, reliable and practically useful methods are developed, particularly for parametric structural quantile models. See \cite{CH2006}, \cite{ChenLee2018}, \cite{Zhu2018}, and \cite{KaidoWuthrich2018} for linear-in-parameters quantile models, and see \cite{ChernozhukovHong2003} and \cite{deCastroetal2019} for nonlinear quantile models. Nonparametric estimation approaches are studied by \cite{CIN2007}, \cite{HorowitzLee2007}, \cite{ChenPouzo2009,ChenPouzo2012}, and \cite{GS2012}.



\section{Appendix}