EconBase
← Back to paper

Treatment Effects with Multidimensional Unobserved Heterogeneity: Identification of the Marginal Treatment Effect

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,173 characters · 18 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Treatment Effects with Multidimensional Unobserved Heterogeneity: Identification of the Marginal Treatment Effect Toshiki Tsuda Yale University

{ \singlespacing

}

abstractThis paper establishes sufficient conditions for the identification of the marginal treatment effects with multivalued treatments. Our model is based on a multinomial choice model with utility maximization. Our MTE generalizes the MTE defined in HV05ecta in binary treatment models. As in the binary case, we can interpret the MTE as the treatment effect for persons who are indifferent between two treatments at a particular level. Our MTE enables one to obtain the treatment effects of those with specific preference orders over the choice set. Further, our results can identify other parameters such as the marginal distribution of potential outcomes.

Keywords: Identification, treatment effect, multidimensional unobserved heterogeneity, multivalued treatments, endogeneity, instrumental variable.

JEL classification: C14, C31

Introduction

Assessing heterogeneity in treatment effects is important for precise treatment evaluation. The marginal treatment effect (MTE) provides rich information on heterogeneity across economic agents regarding their observed and unobserved characteristics. Further, once the MTE is estimated, researchers can obtain other treatment effects, such as the average treatment effect (ATE), the average treatment effect on the treated (ATT), and local ATE (LATE).

In this paper, we consider the multivalued treatments. While the multivalued treatments complicate the identification of treatment effects, they are often used in many applications. For example, vocational programs provide various types of training to participants, and college choice involves numerous dimensions to respond to varied incentives. The literature has developed treatment effects with multivalued treatments, such as LATE AI95JASA, MTE HRV06RES,HRV08AES, HV07HE71, HP18ecta, LS18ecta and instrumental variable quantile regression F20ar.

For the binary treatment model, HV99NAS establish the local instrumental variable (LIV) framework to identify MTE. They assume individuals decide on their choices based on the generalized Roy model that is separable in terms of observed and unobserved variables. V02ecta shows that the separable threshold-crossing model in the LIV approach plays the same role as the monotonicity assumption for identifying LATE IA94ecta.

For the identification of MTE with multivalued treatments, we examine the multiple discrete choice model based on utility maximization. In this model, the value of each treatment is the sum of an observed term and an unobserved term that represents unobserved heterogeneity. This model is a generalization of the multiple logit model and has been extensively studied in economics since the seminal work of M74B. In theoretical research, M93JE establishes sufficient conditions for the nonparametric identification of the discrete choice model. In applied research, D02ECTA employs this model to study the effect of self-selected migration on returns to college. KW16QJE use the discrete multiple-choice model as a self-selection model to analyze the Head Start program’s cost-effectiveness.

We identify the MTE with multidimensional unobserved heterogeneity, which enables us to evaluate treatment effects from multiple perspectives. For instance, we consider three valued treatments and set treatment 0 as the baseline. In this case, the model contains two-dimensional heterogeneity that consists of willingness to take treatment 1 and willingness to take treatment 2 against treatment 0. When we condition the MTE on a high value of the former heterogeneity and a low value of the latter heterogeneity, our identification result reveals the causal effects of those with the preference order treatment 1, treatment 0, and treatment 2 from top to bottom.

A comparison of our MTE with the MTE with a binary treatment reveals several intriguing similarities and discrepancies. As a similarity, our identified MTE with multivalued treatments has multidimensional heterogeneity whose each element follows a uniform distribution on $(0,1)$ while HV05ecta define the MTE with binary treatment conditional on unobserved heterogeneity that also uniformly ranges from 0 to 1. In this sense, our MTE generalizes the MTE in the binary treatment case defined by HV05ecta to the multivalued treatment case. Additionally, as in the binary case, we can interpret the MTE as the treatment effect for persons indifferent between treatment 1 and 0 and treatment 2 and 0 at a specific level. On the other hand, two MTEs have different relationships between treatments and preference order. In the binary treatment case, the MTE corresponds to marginal changes in the treatment choice because an individual's preference order exactly maps to his choice. However, in the multivalued treatment case, marginal changes in preference order do not necessarily correspond to the changes in treatments. Therefore, our MTE with multivalued treatments represents marginal changes in preferences over the choice set.

The main challenge of the identification is that the model properties prevent us from obtaining several marginal changes in treatments. In the binary treatment case, we set a threshold for selecting the treatment and take a derivative with respect to the threshold. This procedure identifies the MTE because the derivative with respect to the threshold exactly expresses a marginal change in the treatment. In the case of multivalued treatments, we set one specific treatment as the baseline and construct the multiple-choice model by comparing the other treatments with the baseline. For the identification, we set thresholds of the other treatments compared with the baseline and take derivatives with respect to these thresholds. In the case of the baseline treatment, this procedure identifies the conditional expectation because the derivative with respect to each threshold exactly expresses a marginal change in each treatment against the baseline. In the case of the other treatments, changing thresholds has indirect effects on all the other treatments and we cannot identify conditional expectations of those treatments by simply taking derivatives with respect to thresholds.

We solve indirect effects by focusing on an area of each treatment that thresholds have only a direct effect. In this model, each treatment has one threshold that has both indirect and direct effects on that treatment. We remove the indirect effects of the threshold by an ingenuous transformation that enables us to substitute those indirect effects with the marginal changes in other thresholds. By removing indirect effects with those substitutes, we can extract the direct effect from the marginal change in the threshold and identify conditional expectations of all the treatments from the multiple discrete choice model.

By identifying conditional expectations of treatments, we can obtain several treatment effects, including the MTE. Because our result identifies conditional expectations of each treatment given unobserved heterogeneity, we can obtain the MTE with multivalued treatments by taking their differences. Further, we can also obtain the marginal distribution of potential outcomes, which leads to identifying quantile treatment effects given multidimensional unobserved heterogeneity.

We also establish a sufficient condition for identifying thresholds. In the case of multivalued treatments, the connection between thresholds and propensity scores is unclear, even though the propensity score is equal to the threshold in the binary case. We assume the existence of at least one instrument that significantly and negatively affects only one treatment. This assumption enables us to identify the thresholds and is also used by LS18ecta for the identification of thresholds.

In the existing literature on the identification of the MTE with multivalued treatments, LS18ecta investigate the identification of conditional expectations given unobserved variables based on multinomial choice models characterized by a combination of separable threshold-crossing rules. They assume the existence of continuous instruments and identify several causal effects with identified thresholds. Our result complements the applicability of their main theorem by introducing a novel identification strategy.

HV07HE71 and HRV08AES expand the LIV approach to a model with multivalued treatments generated by a general unordered choice model. They study identification conditions of several types of treatment effects including the marginal treatment effect of one specified choice versus another choice. They achieve the identification of the MTE by using an identification-at-infinity type argument. Our identification strategy does not depend on the large support assumption.

With the introduction of new treatment effects for multivalued treatments, mountjoy2022community studies the effect of enrollment in 2-year community colleges on upward mobility, such as years of education and future income. Because the main focus of his paper is the effect of policy changes for 2-year entry on the outcome of 2-year colleges, he defines new treatment effects with respect to the marginal change of the instrument pertained to 2-year entry. Our MTEs are based on unobserved heterogeneities that correspond to marginal changes in preferences between treatments and can express his treatment effects.

The remainder of this paper is organized as follows: Section 2 proposes the basic settings and notation used in this study. We construct the model through comparisons between treatments. In Section 3, we explain our MTE with multivalued treatments through figures. We highlight similarities and differences between our MTEs and the MTE in the binary case. After we show the identification of the MTE, we add detailed explanations of our identification strategies. Section 4 establishes sufficient conditions for nonparametric identification of the thresholds. We relate our contributions to the literature in Section 5. In Sections 2--4, we consider the case when the number of treatments is three. Section 6 discusses the general case and identifies the MTE. Section 7 concludes. Proofs of the main results and some auxiliary results are collected in Appendix A. Appendix B provides an economic intuition of the assumption newly imposed in Section 6.

Notation. Let $: = $ denote “equals by definition,” and let a.s. denote “almost surely.” Let $\mathbbm{1}\{\cdot\}$ denote the indicator function. For random variables $X$ and $Z$, $f_{X}(\cdot)$ denotes the probability density function of $X$. $F_{X|Z}(\cdot)$ and $Q_{X|Z}(\cdot)$ denotes the distribution and quantile functions of $X$ given $Z$, respectively.

Model

Let $\mathcal{K}$ denote the set of treatments and assume the set comprising of $K(=|\mathcal{K}|)$ elements. Let $\{Y_{k}:k\in \mathcal{K}\}$ be a potential outcome. $D_{k}$ takes the value one if the agent takes treatment $k$. The observed outcome and treatment are expressed as $D=\sum_{k=0}^{K-1} kD_{k}$ and $Y=\sum_{k=0}^{K-1}D_{k}Y_{k}$, respectively. The data contains covariates $\mathbf{X}$ and instruments $\mathbf{Z}$. Throughout this article, we condition on the value of $\mathbf{X}$ and suppress it from the notation. Let the support of $Y$ and $\mathbf{Z}$ be $\mathcal{Y}\subset\mathbb{R}$ and $\mathcal{Z}\subset\mathbb{R}^{\dim(\mathbf{Z})}$, respectively.

Let $\mathbf{Q(Z)}$ denote the vector of functions of the instruments $Q_{i}(\mathbf{Z})$. Let $\mathbf{V}$ be a vector of unobserved continuous random variables. For some $k\in\mathcal{K}$, define $S_{k}(\mathbf{V},\mathbf{Q(Z)}):=\mathbbm{1}(V_{k}<Q_{k}(\mathbf{Z}))$. $\mathbf{V}$ is a vector of unobserved heterogeneity and $Q_{k}(\mathbf{Z})$ serves as a threshold for each $S_{k}$ when $\mathbf{Z}$ is given. Hence, $S_{k}$ consists of a separable threshold-crossing model as in the generalized Roy model.

We define MTE as \[ E[Y_{k}-Y_{j}|\mathbf{V}]\quad \text{ for $k,j\in\mathcal{K}$} \] and analyze sufficient conditions for the identification. As in HRV06RES,HRV08AES and HV07HE71, we consider the discrete choice model based on utility maximization. This model setting enables us to interpret the above MTE as the treatment effect with unobserved heterogeneity of preferences over the choice set. For details, see Section (ref).

Multiple Discrete Choice Model and Basic Assumptions

For each $k$, we define $R_{k}(\mathbf{Z})$ as an unknown function that maps from $\mathbb{R}^{\dim(\mathbf{Z})}$ to $\mathbb{R}$ and define $U_{k}$ as an unobserved continuous random variable whose support is $\mathbb{R}$. By extending the definition of the treatment variable in the binary treatment model, we formulate the treatment decision as follows:

equation[equation omitted — 117 chars of source]

where $\Pr((U_{k}-R_{k}(\mathbf{Z}))=(U_{j}-R_{j}(\mathbf{Z})))=0$ for $j\neq k$.

Intuitively, by interpreting $U_{k}$ and $R_{k}(\mathbf{Z})$ as unobserved and observed terms in an agent's utility, this discrete multiple-choice model states that he makes a choice based on utility maximization. From this intuition, we regard model ((ref)) as a straightforward generalization of the generalized Roy model.

Model ((ref)) has been studied extensively in economics since the seminal work of M74B. M93JE establishes sufficient conditions for the nonparametric identification of utility functions and the joint distribution function of unobserved random terms. The multinomial choice model has also been used in applied research. D02ECTA uses this model to study the effect of self-selected migration on the return to college. KW16QJE adopt the discrete multiple-choice model as a self-selection model and analyze the Head Start program’s cost-effectiveness in the presence of substitute preschools. KLM16QJE examine the effect of types of education on several gains in earnings. They find that the estimated payoffs are consistent with agents choosing fields based on the discrete multiple-choice model.

For the identification of the MTE, we construct a model through a combination of threshold-crossing models based on model ((ref)). For simplicity, through Sections 2-4, we examine the three valued treatment case, namely treatment 0, 1, and 2, and we generalize results in Section 6. Without loss of generality, we regard treatment 0 as the baseline and set $U_{0}-R_{0}(\mathbf{Z})=0$ almost surely. We construct a model with three alternatives using two indicator functions. Assume $U_{1}$ and $U_{2}$ are continuously distributed. Let

align*[align* omitted — 178 chars of source]

Set

gather*[gather* omitted — 109 chars of source]

Note that

align*[align* omitted — 149 chars of source]

and a similar argument gives

align*[align* omitted — 84 chars of source]

Two indicator functions, $S_{1}$ and $ S_{2}$, correspond to comparisons of utilities between treatment 0 and 1, and treatment 0 and 2, respectively. From model (ref), individuals take treatment 0 when the utility of treatment 0 is the highest among all the alternatives. Therefore, we obtain $D_{0}=S_{1}S_{2}$.

We introduce an indicator function that compares utilities between treatments 1 and 2. Define

equation*[equation* omitted — 155 chars of source]

By trivial calculation, we obtain

equation[equation omitted — 462 chars of source]

Hence, we have $D_{2}=(1-S_{2})S_{3}$ by definition. A similar argument reveals $D_{1}=(1-S_{1})(1-S_{3})$.

figure[figure omitted — 276 chars of source]

This model is depicted in Figure (ref). In this setting, treatment 0 has the form of the double hurdle model, namely $D_{0}=1$ if and only if $V_{1}<Q_{1}(\mathbf{Z})$ and $V_{2}<Q_{2}(\mathbf{Z})$ as in LS18ecta. Even though the double hurdle model essentially expresses the binary treatment case, we successfully specify $D_{1}$ and $D_{2}$ by introducing $S_{3}$ and construct the multiple-choice model based on utility maximization. Hence, our model is not covered by LS18ecta. For details, see Section (ref).

We introduce basic assumptions frequently required in the literature on program evaluation.

assumption$\{V_{1}<Q_{1}(\mathbf{Z})\},\{V_{2}<Q_{2}(\mathbf{Z})\}$ and $\{V_{1}<F_{U_{1}}(F_{U_{2}}^{-1}(V_{2})-F_{U_{2}}^{-1}(Q_{2}(\mathbf{Z}))+F_{U_{1}}^{-1}(Q_{1}(\mathbf{Z}))) \}$ are measurable sets.
assumption[Conditional Independence of Instruments] $Y_{0}$, $Y_{1}$, $Y_{2}$ and $\mathbf{V}=(V_{1},V_{2})^{'}$ are jointly independent of $\mathbf{Z}$.
assumption[Continuously Distributed Unobserved Heterogeneity in the Selection Mechanism] The joint distribution of $(U_{1},U_{2})$ is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{2}$.
assumption[The Existence of the Moments] For $k=0,1,2$, $E[|G(Y_{k})|]<\infty$, where $G$ is a measurable function defined on the support $\mathcal{Y}$ of $Y$, which can be discrete, continuous, or multidimensional.

Assumption (ref) ensures the existence of probability of each treatment, that is, $\Pr(D_{k}=1)$. This assumption guarantees that each treatment variable $D_{k}$ is a random variable. Assumption (ref) corresponds to the exogeneity of instruments, which plays a vital role in the identification in the literature using instrumental variables. We guarantee the existence of the probability density function of $(U_{1},U_{2})$ by Assumption (ref). Through the argument of the change of variables, we can also ensure the joint density of $\mathbf{V}$. Assumption (ref) ensures the existence of moments for each alternative. Otherwise, we cannot define conditional expectations of potential outcomes or identify MTE. Assumptions above often appear in the literature on treatment effects with endogeneity. For instance, Assumptions (ref), (ref), and (ref) correspond to Assumptions 2.1, 2.2, and 3.2 of LS18ecta, respectively. Assumption (ref) generalizes Assumption (A-3) of HRV08AES.

Identification

MTE with Multivalued Treatments

In this paper, we study the identification of the following conditional expectations

equation*[equation* omitted — 137 chars of source]

where we define $(q_{1}^{*},q_{2}^{*})$ as points where we evaluate treatment effects. When we set $G(Y)=Y$ and take the difference between two conditional expectations, we identify the following MTE:

equation[equation omitted — 125 chars of source]

By definition, each element of $(V_{1},V_{2})$ has a uniform distribution on $(0, 1)$ and $(q^{*}_{1},q_{2}^{*})$ refer to quantiles of distributions of $U_{1}$ and $U_{2}$, respectively. Hence, $V_{1}$ and $V_{2}$ mean the willingness to choose treatments 1 and 2 compared to treatment 0. For instance, a low value of $V_{1}$ implies an individual is less likely to take treatment 1 than treatment 0.

MTE (ref) provides rich information about treatment effects conditioned on individuals' preferences over the choice set. The MTE characterizes preferences among all the alternatives through the values of $(V_{1},V_{2})$. For example, when we identify MTE with a low value of $V_{1}$ and a high value of $V_{2}$, we can interpret this MTE as the average treatment effect in those who are more likely to take treatment 2 and less likely to take treatment 1 compared to treatment 0, i.e., their preferences would be treatment 2, treatment 0 and treatment 1 from top to bottom.

As another interpretation, MTE (ref) is the average treatment effect for individuals who would be indifferent between treatment 1 and 0, and treatment 2 and 0 at $(Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z}))=(q_{1}^{*}, q_{2}^{*})$. Under Assumption (ref), we can illustrate this interpretation in the following equation,

align*[align* omitted — 171 chars of source]

Note that each unobserved heterogeneity only focuses on comparing two treatments, treatment 1 and 0, and treatment 2 and 0. Therefore, MTE (ref) corresponds to the marginal changes in treatment 1 and 0, and treatment 2 and 0.

Comparing MTE ((ref)) with the MTE in the binary treatment case provides useful insights. HV05ecta define $D^{*}=1$ as the receipt of the treatment and characterize the decision rule as the generalized Roy model, that is,

equation*[equation* omitted — 74 chars of source]

where $\mu_{D}(\mathbf{Z})$ is an unknown function, which maps from $\mathbb{R}^{\dim(\mathbf{Z})}$ to $\mathbb{R}$, and $U_{D}$ is an unobserved continuous random variable. As a normalization, they innocuously assume that $U_{D}\sim U[0,1]$ and $U_{D}$ is the quantile of the willingness to participate in the treatment. HV05ecta then define MTE with binary treatment as

equation[equation omitted — 91 chars of source]

where $u_{D}\in (0,1)$. In our definition of MTE with multivalued treatments, $(V_{1},V_{2})$ precisely corresponds to $U_{D}$ in the binary treatment case. Therefore, MTE ((ref)) is a natural generalization of MTE with binary treatment to the multivalued treatment case.

HV05ecta show that treatment effects such as ATE and ATT can be expressed as a function of their MTE. Similarly, in our model, treatment effects such as ATE and ATT can be expressed as a function of our MTE.

MTE (ref) has a different interpretation from MTE (ref) due to the existence of multivalued treatments. In the binary case, whether an individual takes treatment or not precisely corresponds to his preference for its treatment. However, in the multivalued treatment case, preference orders over the choice set have additional information over revealed treatments. For example, if an individual's best treatment is treatment 2, her preference order of the choice set is treatment 2, 0, 1 or treatment 2, 1, 0. When $Q_{2}(\mathbf{Z})$ marginally changes through $R_{2}(\mathbf{Z})$, this change corresponds to the binary choice between treatment 2 and 0 or treatment 2 and 1, but the change in $R_{2}(\mathbf{Z})$ does not affect the preference between treatment 1 and 0. On the other hand, when $Q_{1}(\mathbf{Z})$ marginally changes and $Q_{2}(\mathbf{Z})$ remains fixed, her choice may remain in treatment 2 because $Q_{1}(\mathbf{Z})$ only affects the change in preference between treatment 1 and 0. Therefore, marginal changes in $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ correspond to not marginal changes in treatments but marginal changes in preferences between treatment 1 and 0, and treatment 2 and 0, respectively. MTE (ref) is the treatment effect depicting marginal changes in preferences over the choice set.

Illustration

We illustrate figures of three MTEs, $E[Y_{1}-Y_{0}|\mathbf{V}], E[Y_{2}-Y_{0}|\mathbf{V}]$ and $E[Y_{2}-Y_{1}|\mathbf{V}]$ given $V_2$ is fixed at $0.5$. We depict the MTEs in the following two cases:

Case 1: $\mathbf{(V_1,V_2)}$ are not independent of the outcome variable $\mathbf{Y_{k}}$

figure[figure omitted — 623 chars of source]

Case 1 examines three MTEs when $(V_1,V_2)$ is correlated with the outcome variable $Y_{k}$, i.e. \[V_{k}\not\!\perp\!\!\!\perp Y_{\ell}\quad \text{for $k\in\{1,2\}$ and $\ell\in\{0,1,2\}$.}\] In this case, $E[Y_{1}-Y_{0}|\mathbf{V}]$ increases and $E[Y_{2}-Y_{1}|\mathbf{V}]$ decreases with $V_1$ because people are more likely to choose treatment 1 than treatment 0 at the high value of $V_1$. Even though $V_1$ does not directly affect the difference between treatment 2 and 0, the MTE $E[Y_{2}-Y_{0}|\mathbf{V}]$ decreases slightly with $V_1$, reflecting the combined effect of the decrease in $E[Y_{2}-Y_{1}|\mathbf{V}]$ and the increase in $E[Y_{1}-Y_{0}|\mathbf{V}]$.

Case 2: $\mathbf{V_1}$ is independent of $\mathbf{Y_{0}}$ and $\mathbf{Y_{2}}$ given $\mathbf{V_2}$

figure[figure omitted — 617 chars of source]

Case 2 analyzes three MTEs when $V_1$ is independent of $Y_{0}$ and $Y_{2}$ given $V_{2}$, i.e. \[V_{1}\perp \!\!\! \perp Y_{k}\quad\text{given $V_2$}\quad \text{for $k\in\{0,2\}$.}\] In this case,the MTE $E[Y_{2}-Y_{0}|\mathbf{V}]$ does not depend on $V_1$ and is equal to $E[Y_2-Y_0|V_2=0.5]$ for any $V_{1}\in(0,1)$. Conditional independence of $V_{1}$ implies that the comparison in preference between treatment 1 and 0 does not affect treatment effect $Y_{2}-Y_{0}$. Furthermore, the difference in two other MTEs becomes constant because $E[Y_{2}-Y_{0}|\mathbf{V}]=E[Y_{2}-Y_{1}|\mathbf{V}]+[Y_{1}-Y_{0}|\mathbf{V}]$ holds for any $\mathbf{V}$.

Identification Result

We introduce assumptions to identify the MTE with multivalued treatments. Assumption (ref) is a technical assumption for the proof of the identification, such as continuity and differentiability.

assumption\ \begin{enumerate}[(1).] • For $k\in\{0,1,2\}$, $E[G(Y)D_{k}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ and $E[D_{k}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ are twice differentiable at $(q_{1}^{*}, q_{2}^{*})$. • For $k\in\{0,1,2\}$, $E[G(Y_{k})|V_{1}, V_{2}]$ and $E[D_{k}|V_{1}, V_{2}]$ are continuous on $(0,1)^{2}$. • \begin{enumerate} • For $k\in\{0,1,2\}$, $\sup_{(v_{1},v_{2})\in(0,1)^{2}}E[|G(Y_{k})||V_{1}=v_{1},V_{2}=v_{2}]$ is finite. • Conditional density functions satisfy the following: \begin{equation*} \sup_{(u_{1},u_{2})\in\mathbb{R}^{2}}\frac{f_{U_{1},U_{2}}(u_{1},u_{2})}{f_{U_{1}}(u_{1})}<\infty, \quad \sup_{(u_{1},u_{2})\in\mathbb{R}^{2}}\frac{f_{(U_{1},U_{2})}(u_{1},u_{2})}{f_{U_{2}}(u_{2})}<\infty, \end{equation*} \end{enumerate} \end{enumerate}

Assumption (ref) (ref) guarantees the existence of derivatives for each conditional expectation. This condition implicitly assumes that the value of $Q_{1}(\mathbf{Z})$ is movable while $Q_{2}(\mathbf{Z})$ is fixed and vice versa. We require Assumption (ref) (ref) to exchange differentiation and integration. Assumption (ref) (ref) holds when $G(Y_{k})$ is bounded for each $k\in\{0,1,2\}$.

Conditional on the assumption that $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ are identified, we can identify conditional expectations $E[G(Y_{0})|V_{1}=q_{1}^{*},V_{2}=q_{2}^{*}]$, $E[G(Y_{1})|V_{1}=q_{1}^{*},V_{2}=q_{2}^{*}]$ and $E[G(Y_{2})|V_{1}=q_{1}^{*},V_{2}=q_{2}^{*}]$ by partially differentiating conditional expectations $E[G(Y)D_{k}|Q_{1}(\mathbf{Z})=q_{1},Q_{2}(\mathbf{Z})=q_{2}]$ and $E[D_{k}|Q_{1}(\mathbf{Z})=q_{1},Q_{2}(\mathbf{Z})=q_{2}]$ for $k=0,1,2$.

theoremAssume Assumptions (ref) to (ref) hold. Then, conditional expectations of $G(Y_{0}),G(Y_{1})$ and $G(Y_{2})$ are given by \begin{align*} (a) \quad &E[G(Y_{0})|V_{1}=q_{1}^{*},V_{2}=q_{2}^{*}]=\frac{\Delta GD^{*}_{(1,1)}(0)}{\Delta D^{*}_{(1,1)}(0)}, \\ (b) \quad &E[G(Y_{1})|V_{1}=q_{1}^{*},V_{2}=q_{2}^{*}] \notag \\ =&-\frac{\Delta GD^{*}_{(1,1)}(1)}{\Delta D^{*}_{(1,1)}(0)}-\frac{\Delta GD^{*}_{(0,1)}(1)[\Delta D^{*}_{(1,0)}(0)+\Delta D^{*}_{(1,0)}(1)]\Delta D^{*}_{(0,2)}(1)}{\Delta D^{*}_{(1,1)}(0)(\Delta D^{*}_{(0,1)}(1))^{2}} \\ &+\frac{\Delta GD^{*}_{(0,2)}(1)[\Delta D^{*}_{(1,0)}(0)+\Delta D^{*}_{(1,0)}(1)]+\Delta GD^{*}_{(0,1)}(1)[\Delta D^{*}_{(1,1)}(0)+\Delta D^{*}_{(1,1)}(1)]}{\Delta D^{*}_{(1,1)}(0)\Delta D_{(0,1)}^{*}(1)}, \notag \\ (c) \quad &E[G(Y_{2})|V_{1}=q_{1}^{*},V_{2}=q_{2}^{*}] \\ =&-\frac{\Delta GD^{*}_{(1,1)}(2)}{\Delta D^{*}_{(1,1)}(0)}-\frac{\Delta GD^{*}_{(1,0)}(2)[\Delta D^{*}_{(0,1)}(0)+\Delta D^{*}_{(0,1)}(2)]\Delta D^{*}_{(2,0)}(2)}{\Delta D^{*}_{(1,1)}(0)(\Delta D^{*}_{(1,0)}(2))^{2}} \\ &+\frac{\Delta GD^{*}_{(2,0)}(2)[\Delta D^{*}_{(0,1)}(0)+\Delta D^{*}_{(0,1)}(2)]+\Delta GD^{*}_{(1,0)}(2)[\Delta D^{*}_{(1,1)}(0)+\Delta D^{*}_{(1,1)}(2)]}{\Delta D^{*}_{(1,1)}(0)\Delta D^{*}_{(1,0)}(2)}, \end{align*} where we define \begin{align*} \Delta GD^{*}_{(\ell,m)}(k)&:=\left.\frac{\partial^{(\ell+m)} E[G(Y)D_{k}|Q_{1}(\mathbf{Z}),Q_{2}(\mathbf{Z})]}{\partial^{\ell} Q_{1}(\mathbf{Z})\partial^{m} Q_{2}(\mathbf{Z})}\right|_{(Q_{1}(\mathbf{Z}),Q_{2}(\mathbf{Z}))=(q_{1}^{*},q_{2}^{*})}, \\ \Delta D^{*}_{(\ell,m)}(k)&:=\left.\frac{\partial^{(\ell+m)} E[D_{k}|Q_{1}(\mathbf{Z}),Q_{2}(\mathbf{Z})]}{\partial^{\ell} Q_{1}(\mathbf{Z})\partial^{m} Q_{2}(\mathbf{Z})}\right|_{(Q_{1}(\mathbf{Z}),Q_{2}(\mathbf{Z}))=(q_{1}^{*},q_{2}^{*})}. \end{align*}

For any $k\in\{0,1,2\}$ and $\ell,m\in\{0,1,2\}$ such that $\ell+m\leq 2$, we can obtain $\Delta GD^{*}_{(\ell,m)}(k)$ and $\Delta D^{*}_{(\ell,m)}(k)$ by using estimation methods such as local polynomial regression. Note that the density of $\mathbf{V}$ at $(q_{1}^{*},q_{2}^{*})$ is identified as $\Delta D^{*}_{(1,1)}(0)$. For (b) and (c), the second and third terms on the right-hand side correct indirect effects that we discuss in the following.

The identification result of conditional expectations enables us to identify measures of treatment effects.\footnote{In the binary treatment case, by using identification results of conditional expectations, CL09JE propose a semiparametric estimator of the MTE. BMW17JPE also employ results to identify MTE with discrete instruments.} For example, if we set $G(Y)=Y$ as we did previously, we obtain

equation*[equation* omitted — 114 chars of source]

If we let $G(Y)=\mathbbm{1}(Y\leq y)$ for $y\in\mathbb{R}$, we can identify

equation*[equation* omitted — 164 chars of source]

If $F_{Y_{1}|V_{1},V_{2}}$ and $F_{Y_{2}|V_{1},V_{2}}$ are invertible, we identify the quantile treatment effect by taking the difference between the two, that is,

equation*[equation* omitted — 166 chars of source]

where $\tau\in(0,1)$.

figure[figure omitted — 203 chars of source]

Our identification strategy consists of marginal changes in conditional expectations and proper corrections for indirect effects caused by a nonlinear threshold. We can identify $E[G(Y_{0})|V_{1},V_{2}]$ without being troubled by a nonlinear threshold. For instance, when $Q_{2}(\mathbf{Z})$ is fixed, the change in $Q_{1}(\mathbf{Z})$ corresponds exactly to the flow between treatment 0 and 1 over values of $V_{2}$ in $(0,q_{2})$ (flow A in Figure (ref)). Further, taking its derivative at $q_{2}$ provides the flow between treatment 0 and 2 at $(q_{1},q_{2})$ (flow C in Figure (ref)). Therefore, we can identify $E[G(Y_{0})|V_{1},V_{2}]$ by taking derivatives of $E[G(Y)D_{0}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ and $E[D_{0}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ with respect to $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ as in the binary case.

However, a nonlinear threshold in this model complicates identifications of conditional expectations about treatments 1 and 2. In order to identify expectations about treatment 1 and treatment 2 conditional on $V_1$ and $ V_2$, we need to take derivatives of $E[G(Y)D_{k}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ and $E[D_{k}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ for $k=1$ and $2$ with respect to $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ as in the case of treatment 0. Because marginal changes in $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ only indirectly affect the preference between treatment 1 and 2, we need to deal with indirect changes between treatment 1 and 2.

For instance, when identifying the conditional expectations of treatment 1, we first study area S in Figure (ref). Because $V_{1}$ is larger than $Q_{1}(\mathbf{Z})$ and $V_{2}$ is smaller than $Q_{2}(\mathbf{Z})$, this area represents those who have the preference order of treatment 1, 0, 2 from top to bottom. When $Q_{1}(\mathbf{Z})$ decreases through $R_{1}(\mathbf{Z})$, people with treatment 0 will move into area S (flow A in Figure (ref)), while this change in $R_1(\mathbf{Z})$ does not affect preferences over treatment 0 and 2. Moreover, when $Q_{2}(\mathbf{Z})$ slightly decreases through $R_{2}(\mathbf{Z})$, people in area S will change their preference orders from treatment 1, 0, 2 to treatment 1, 2, 0. Therefore, confining to area S, we can identify $E[G(Y_{1})|V_{1},V_{2}]$ by taking derivatives of $E[G(Y)D_{1}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ and $E[D_{1}|Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z})]$ with respect to $Q_1(\mathbf{Z})$ and $Q_2(\mathbf{Z})$ as in treatment 0.

A decrease in $Q_{1}(\mathbf{Z})$ generates another flow, flow B in Figure (ref). We must remove the marginal changes in treatment 1 caused by flow B. In this case, when $Q_{2}(\mathbf{Z})$ decreases through $R_{2}(\mathbf{Z})$, some individuals taking treatment 1 will move to treatment 2 (flow D in Figure (ref)), but there is no flow from treatment 1 to treatment 0 because the change in $R_{2}(\mathbf{Z})$ does not affect preferences over treatment 1 and 0. We can remove the marginal change between treatment 1 and treatment 0 by substituting flow D for flow B. Then, we can obtain marginal changes generated by $Q_{1}(\mathbf{Z})$ in area S by subtracting flow B from marginal changes generated by $Q_{1}(\mathbf{Z})$ in treatment 1. Therefore, we achieve the identification of $E[G(Y_{1})|V_{1},V_{2}]$.

Identification of thresholds $\mathbf{Q(Z)}$

In this section, we consider the identification result of $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ that enable us to estimate MTE with results in Theorems (ref). While the propensity score $P(\mathbf{Z})$ plays a role as a threshold in deciding on the treatment in the binary case, $\mathbf{Q(Z)}$ has more complicated relationships with propensity scores $E[D_{0}|\mathbf{Z}]$, $E[D_{1}|\mathbf{Z}]$ and $E[D_{2}|\mathbf{Z}]$. Because each marginal distribution of $V_{1}$ and $V_{2}$ is uniformly distributed on $(0,1)$, $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ correspond to the probability that treatment 0 is preferable to treatment 1 and the probability that treatment 0 is preferable to treatment 2, respectively. Hence, we cannot directly identify $\mathbf{Q}(\mathbf{Z})$ from $E[D_{0}|\mathbf{Z}]$, $E[D_{1}|\mathbf{Z}]$ and $E[D_{2}|\mathbf{Z}]$ because propensity scores only reveal probabilities of each treatment given $\mathbf{Z}$. \footnote{If we have the data of preferences over the choice set as in the case of KLM16QJE, we can directly identify $\mathbf {Q}(\mathbf{Z})$.}

We provide a sufficient condition for the nonparametric identification of $Q_{1}(\mathbf{Z})$ and $Q_{2}(\mathbf{Z})$ that enables us to estimate MTE with results in Theorems (ref). A sufficient condition requires the existence of at least one instrument that significantly and negatively affects only one utility. Let $Z^{[\ell]}$ denote the $\ell$-th component of $\mathbf{Z}$. Let $\mathbf{Z}^{[-\ell]}$ be all the instruments except for the $\ell$-th component.

assumptionFor $k=1,2$, there exists at least one element of $\mathbf{Z}$, say $Z^{[\ell_{k}]}$, and at least one value $a^{[\ell_{k}]}$, such that, given any $\mathbf{z}^{[-\ell_{k}]}\in\mathbf{Z}^{[-\ell_{k}]}$, \begin{equation*} \lim_{z^{[\ell_{k}]}\rightarrow a^{[\ell_{k}]}}R_{k}(z^{[\ell_{k}]},\mathbf{z}^{[-\ell_{k}]})=\infty\ \ \ \end{equation*} and $R_{j}(\mathbf{Z})$ is constant for $j\neq k$.

Assumption (ref) imposes a type of exclusion restriction. Conditional on all the regressors except $Z^{[\ell_{k}]}$, one can vary $R_{k}(\mathbf{Z})$ independently. Moreover, we assume the existence of one value $a^{[\ell_{k}]}$ such that, as $z^{[\ell_{k}]}$ approaches $a^{[\ell_{k}]}$, the value of the function $-R_{k}(\mathbf{Z})$ becomes sufficiently small given any $\mathbf{z}^{-[\ell_{k}]}$.

theoremAssume Assumptions (ref) to (ref) and (ref) hold. Then, $Q_{1}(\mathbf{Z}),Q_{2}(\mathbf{Z})$ are identified as \begin{align*} \lim_{z^{[\ell_{2}]}\rightarrow a^{[\ell_{2}]}}H(\mathbf{Z})&=Q_{1}(\mathbf{Z}), \\ \lim_{z^{[\ell_{1}]}\rightarrow a^{[\ell_{1}]}}H(\mathbf{Z})&=Q_{2}(\mathbf{Z}), \end{align*} where \begin{equation*} H(\mathbf{Z}):=\Pr(D_{1}=1|\mathbf{Z})=F_{\mathbf{V}}(Q_{1}(\mathbf{Z}),Q_{2}(\mathbf{Z})). \end{equation*}

The strategy for the identification of $\mathbf{Q}(\mathbf{Z})$ in Theorem (ref) deeply depends on the reduction to the binary treatment setting. As $z^{[\ell_{2}]}$ converges to $a^{[\ell_{2}]}$, for instance, $Q_{2}(\mathbf{Z})$ approximately approaches to 1, which implies that individuals take treatment 0 or treatment 1. Hence, we identify $Q_{1}(\mathbf{Z})$ as in the binary case.

Assumption (ref) is similar to assumptions for identifying thresholds in the existing literature. LS18ecta establish the general identification result of conditional expectations given that $\mathbf{Q}(\mathbf{Z})$ is known. They study the identification of $\mathbf{Q}(\mathbf{Z})$ for some choice models, using the information of these models. Especially, they impose a similar large support assumption for the identification of $\mathbf{Q(Z)}$ in a double hurdle model. Because treatment 0 has the form of a double hurdle model, Theorem (ref) can be considered as the identification result of thresholds in a double hurdle model, as in Theorem 4.2 in LS18ecta.

Comparison with the Existing Literature

In this section, we briefly review the existing literature regarding the identification of MTE with multivalued treatments and compare them with our results.

Lee and Salani\'e (2018)

LS18ecta employ the following model:

equation*[equation* omitted — 58 chars of source]

where

align[align omitted — 224 chars of source]

and $c_{l}^{k}$ is an integer. Let $\mathcal{J}$ be the set of choices, $\{1,\cdots,J\}$, and let $\mathcal{L}$ be the set of all the subsets of $\mathcal{J}$. Model ((ref)) can express any decision model that comprises sums, products, and differences of their indicator functions $S_{j}$.

In Section 5.2, LS18ecta apply their main theorem to the multiple discrete choice model. Example 5.2 of LS18ecta analyzes three treatments, $\mathcal{K}=\{0,1,2\}$. They define

align[align omitted — 375 chars of source]

Subsequently, they define

align[align omitted — 418 chars of source]

Based on the comparison among utilities as we did in Section (ref), they define treatments as follows:

itemize$D=0$ iff $V_{0,1}<Q_{0,1}(\mathbf{Z})$ and $V_{0,2}<Q_{0,2}(\mathbf{Z})$, • $D=1$ iff $V_{0,1}>Q_{0,1}(\mathbf{Z})$ and $V_{1,2}<Q_{1,2}(\mathbf{Z})$, • $D=2$ iff $V_{0,2}>Q_{0,2}(\mathbf{Z})$ and $V_{1,2}>Q_{1,2}(\mathbf{Z})$.

Evidently, this corresponds to the decision rule based on model ((ref)).

LS18ecta argue that their main theorems (Theorem 3.1 and Theorem A.1) can identify the MTE if $Q_{0,1}(\mathbf{Z}),Q_{0,2}(\mathbf{Z}),Q_{1,2}(\mathbf{Z})$ are identified and their Assumptions 2.1--2.2 and 3.2--3.4 hold. From this result, they state that we can identify MTE without monotonicity. Moreover, because they identify the MTE via multidimensional cross-derivatives, they do not rely on the identification-at-infinity strategy.

figure[figure omitted — 418 chars of source]

The following discussion shows that the model in Section 5.2 of LS18ecta may not be sufficient to identify the MTE. When we set $\mathbf{V}:=(V_{0,1},V_{0,2}, V_{1,2})^{'}$, by construction we have

align*[align* omitted — 181 chars of source]

This equality suggests that even if $\mathbf{V}$ is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{3}$, its support is not equal to $[0,1]^{3}$ as in Figure (ref). Consequently, $\mathbf{V}$ cannot satisfy Assumption 3.2 in LS18ecta, which requires that the joint distribution of $\mathbf{V}$ be absolutely continuous on $\mathbb{R}^{3}$ and that its support be equal to $[0,1]^{3}$.

Our model can be regarded as a double hurdle model for treatment 0. However, in a double hurdle model with three choices, we cannot define some treatments through two thresholds. For example, when treatment 0 has the form of a double hurdle model, the information in $\mathbbm{1}\{V_1 < Q_1(Z)\}$ and $\mathbbm{1}\{V_2 < Q_2(Z)\}$ is not sufficient to determine whether the agent receives either treatment 1 or treatment 2. As a result, we cannot identify the MTE with multivalued treatments. With the introduction of $S_{3}$, we successfully specify $D_{1}$ and $D_{2}$ in our model. Our model cannot be expressed in the form of model ((ref)), however, because $S_{3}$ includes the nonlinear transformation of $(V_{1},V_{2})$ and $(Q_{1}(\mathbf{Z}), Q_{2}(\mathbf{Z}))$. Consequently, Theorem 3.1 in LS18ecta is not sufficient to identify the corresponding conditional expectations of $Y_{1}$ or $Y_{2}$ in our model.

Heckman and Vytlacil (2007) and Heckman et al. (2008)

HV07HE71 and HRV08AES expand the LIV approach to model ((ref)). They study identification conditions of three treatment effects: the treatment effect of one specific choice versus the next best alternative, the treatment effect of one specific group of choices versus the other group, and the treatment effect of one specified choice versus another choice.

HV07HE71 and HRV08AES establish sufficient conditions for the identification of the following MTE:

equation[equation omitted — 136 chars of source]

for any $\ell\in\mathbb{R}$ and $j,k\in \{0,1,2\}$, such that $j\neq k$. To identify MTE (ref), they impose a large support assumption in Theorem 8 HV07HE71 and Theorem 3 HRV08AES. This large support assumption implies that one utility function, $R_{k}(\mathbf{Z})$, can take a sufficiently negative value as in Assumption (ref). Under the large support assumption, they succeed in reducing the model with three alternatives to a binary case and they achieve the identification of MTE (ref).

While we impose Assumption (ref) for the identification of thresholds, our identification strategy does not depend on the large support assumption. As a result, we achieve the identification of the MTE with two-dimensional unobserved heterogeneity $(V_{1},V_{2})$. Moreover, while their MTEs are conditioned on unobserved heterogeneity $U_{k}-U_{j}$, we do not require the information of distributions for heterogeneity.

Mountjoy (2022)

mountjoy2022community studies the effect of enrollment in 2-year community colleges on upward mobility, such as years of education and future income. With three valued treatments, he disentangles the treatment effect of 2-year college entry versus the other two treatments into two parts: the treatment effect between 2-year college entry and no college and the treatment effect between 2-year and 4-year entry. Using a novel identification approach, he identifies and estimates those treatment effects with the multivalued treatments.

Let $D_{0},D_{2}$ and $D_{4}$ denote no-college treatment, 2-year college entry treatment, and 4-year college entry treatment. He define $Z_{2}$ and $Z_{4}$ as continuous instrumental variables specific to $D_{2}$ and $D_{4}$, respectively and assume $D_{0},D_{2}$ and $D_{4}$ depend on two instruments, i.e., $D_{0}=D_{0}(z_{2},z_{4})$, $D_{2}=D_{2}(z_{2},z_{4})$ and $D_{4}=D_{4}(z_{2},z_{4})$. The observed outcome $D(z_2,z_4)$ is expressed as $D(z_2,z_4)=\sum_{k=0}^{4} kD_{k}(z_{2},z_{4})$.

Because mountjoy2022community is interested in a special case of the marginal treatment effect, that is, the effect of marginal policy changes for 2-year entry on the outcome of 2-year colleges, he defines new MTEs with respect to the marginal change of $z_{2}$. Under regularity conditions, he defines and identifies the following two treatment effects with interesting decomposition.

align[align omitted — 530 chars of source]

where he defines

equation*[equation* omitted — 153 chars of source]

MTEs (ref) and (ref) reflect the marginal change between treatments induced by the marginal change in $z_{2}$.

Our MTEs are based on unobserved heterogeneities $V_{1}$ and $V_{2}$, which refer to the preferences of treatment 1 and treatment 2 over treatment 0, respectively. Therefore, MTE (ref) corresponds to the marginal change in preferences over the choice set and is entirely different from MTEs (ref) and (ref). While mountjoy2022community does not require a decision model for the identification of his MTE, we set utilities of 2-year and 4-year entry as like model (ref). Let $R_{2}(\mathbf{Z})$ and $R_{4}(\mathbf{Z})$ denote observed terms of utilities for a 2-year college and a 4-year college, and $U_2$ and $U_4$ denote unobserved terms of utilities for a 2-year college and a 4-year college, respectively. Then, MTEs in mountjoy2022community can be expressed in terms of our MTEs as follows,

align*[align* omitted — 614 chars of source]

where we define

align*[align* omitted — 87 chars of source]

Generalization

This section generalizes the framework in Section (ref) and (ref) to identify the MTE for the discrete choice model with more than two treatments.

Model and Assumptions

In this subsection, we extend the model constructed in Section (ref). We set treatment $k\in\mathcal{K}$ as the baseline and consider the identification of the MTE of treatment $k$ versus treatment $j$ for any $j\neq k$ where $j,k\in\mathcal{K}$ and $K\geq3$.

Define, for each $i\neq k$ in $\mathcal{K}$,

align*[align* omitted — 122 chars of source]

Assume $\tilde{U}_{i,k}$ are continuously distributed. Let

align*[align* omitted — 181 chars of source]

By construction, $Q_{i}(\mathbf{Z})$, $V_{i}$ and $S_{i}$ are defined for each $i$ in $\mathcal{K}$ except $k$. Note that

align*[align* omitted — 227 chars of source]

Therefore, we obtain $D_{k}=\prod_{i\in\mathcal{K}\backslash \{k\}}S_{i}$.

Define, for each $i\neq j,k$ in $\mathcal{K}$,

equation*[equation* omitted — 195 chars of source]

By construction, we obtain

align*[align* omitted — 488 chars of source]

Hence, we have $D_{j}=\prod_{i\in\mathcal{K}\backslash \{j,k\}}S_{i}^{*}(1-S_{j})$ by definition.

We use the same strategy to identify MTE of treatment $k$ versus treatment $j$ as in Section (ref). Assumptions (ref) to (ref) correspond to Assumptions (ref) to (ref), respectively.

assumptionFor each $i\neq k$ and $\ell\neq k, j$ in $\mathcal{K}$, $\{V_{i}<Q_{i}(\mathbf{Z})\}$ and $\{V_{\ell}<F_{\tilde{U}_{k,\ell}}(F_{\tilde{U}_{k,j}}^{-1}(V_{j})-F_{\tilde{U}_{k,j}}^{-1}(Q_{j}(\mathbf{Z}))+F_{\tilde{U}_{k,\ell}}^{-1}(Q_{\ell}(\mathbf{Z}))) \}$ are measurable sets.
assumption[Conditional Independence of Instruments] $Y_{j}, Y_{k}$ and $\mathbf{V}=(V_{0},\cdots, \\ V_{k-1}, V_{k+1},\cdots,V_{K-1})^{'}$ are jointly independent of $\mathbf{Z}$.
assumption[Continuously Distributed Unobserved Heterogeneity in the Selection Mechanism] The joint distribution of $(\tilde{U}_{0,k},\cdots, \tilde{U}_{K-1,k})$ is absolutely continuous with respect to the Lebesgue measure on $\mathbb{R}^{K-1}$.
assumption[The Existence of the Moments] $E[|G(Y_{j})|]<\infty$ and $E[|G(Y_{k})|]<\infty$, where $G$ is a measurable function defined on the support $\mathcal{Y}$ of $Y$, which can be discrete, continuous, or multidimensional.
assumption\ \begin{enumerate}[(1).] • For $i\in\{k,j\}$, $E[G(Y)D_{i}|\mathbf{Q}(\mathbf{Z})]$ and $E[D_{i}|\mathbf{Q}(\mathbf{Z})]$ are twice differentiable at $\mathbf{q}^{*}$ where we define $\mathbf{Q}(\mathbf{Z}):=(Q_{0}(\mathbf{Z}),\cdots, Q_{k-1}(\mathbf{Z}), Q_{k+1}(\mathbf{Z}),\cdots, Q_{K-1}(\mathbf{Z}))^{'}$ and $\mathbf{q}^{*}:=(q_{0}^{*},\cdots, q_{k-1}^{*}, q_{k+1}^{*},\cdots,q_{K-1}^{*})^{'}$. • For $i\in\{k,j\}$, $E[G(Y_{i})|\mathbf{V}]$ and $E[D_{i}|\mathbf{V}]$ are continuous on $(0,1)^{K-1}$. • \begin{enumerate} • For $i\in\{k,j\}$, $\sup_{\mathbf{v}\in(0,1)^{(K-1)}}E[|G(Y_{i})||\mathbf{V}=\mathbf{v}]$ is finite. • For any $i\in\mathcal{K}^{\backslash\{k\}}$, density functions satisfy the following: \begin{align*} &\sup_{(\tilde{u}_{0,k},\cdots, \tilde{u}_{K-1,k})\in\mathbb{R}^{K-1}}\frac{f_{\tilde{U}_{0,k},\cdots, \tilde{U}_{K-1,k}}(\tilde{u}_{0,k},\cdots, \tilde{u}_{K-1,k})}{\prod_{\ell\in\mathcal{K}\backslash\{k, i\}}f_{\tilde{U}_{\ell,k}}(\tilde{u}_{\ell,k})}<\infty. \end{align*} \end{enumerate} \end{enumerate}

In studying the general case, we need to introduce the following assumption that innocuously holds in the case $K=3$.

assumptionThere exists at least one treatment $b^{*}$ in $\mathcal{K}^{\backslash \{j,k\}}$ such that the following conditions hold for any $i\in\mathcal{K}^{\backslash \{j,k\}}$: \begin{enumerate}[(1).] • The density ratio $\left.f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))\right/ f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}))$ is differentiable at $q_{i}^{*}$ and $q^{*}_{b^{*}}$. • The following terms are known, \begin{align*} &\frac{f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}^{*}))}{f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}^{*}))},& &\left.\frac{\partial}{\partial q_{i}}\left(\frac{f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))}{f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}^{*}))}\right)\right|_{q_{i}=q_{i}^{*}},& \\ &\left.\frac{\partial}{\partial q_{b}^{*}}\left(\frac{f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))}{f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}^{*}))}\right)\right|_{q_{b}^{*}=q_{b^{*}}^{*}},& &\left.\frac{\partial^{2}}{\partial q_{b}^{*}\partial q_{i}}\left(\frac{f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))}{f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}^{*}))}\right)\right|_{(q_{i},q_{b^{*}})=(q_{i}^{*}, q_{b^{*}}^{*})}.& \end{align*} \end{enumerate}

Assumption (ref) is less restrictive than the assumption requiring identification of the density ratio between \(\tilde{U}_{i,k}\) and \(\tilde{U}_{j,k}\), as well as its derivative, for each \(i \in \mathcal{K}^{\setminus\{j,k\}}\). This weaker condition is automatically satisfied when \(K = 3\) because the density ratio is equal to 1.

Focusing on the area $D=j$, as in the case of \(K = 3\), a marginal change in \(Q_j\) induces flows not only into the area where \(D = k\) but also into areas where \(D = \ell\) for any \(\ell \in \mathcal{K}^{\setminus \{j,k\}}\). As outlined in Section 3.3, these indirect effects must be individually offset by substituting marginal changes of \(Q_j\) in areas other than \(D = k\) and \(D = j\). When considering substitution effects from each area, the adjustment cost must be scaled by the density ratio specific to that area. However, since the data cannot distinguish indirect effects from individual areas, these influences are aggregated into a single sum.

For cases where \(K \geq 4\), it becomes necessary to identify the density ratios specific to each area. By contrast, when \(K = 3\), there is only one remaining area other than \(D = k\) and \(D = j\), and its density ratio can be directly offset. Therefore, this assumption is unnecessary when \(K = 3\).\footnote{Economic sufficient conditions for identifying these density ratios are provided in Appendix (ref).}

Identification Result of MTE

Conditional on the assumption that $\mathbf{Q}(\mathbf{Z})$ is identified, we can identify conditional expectations $E[G(Y_{k})|\mathbf{V}=\mathbf{q}^{*}]$, $E[G(Y_{j})|\mathbf{V}=\mathbf{q}^{*}]$ and $f_{\mathbf{V}}(\mathbf{q}^{*})$ by partially differentiating conditional expectations $E[G(Y)D_{i}|\mathbf{Q}(\mathbf{Z})=\mathbf{q}]$, $E[D_{i}|\mathbf{Q}(\mathbf{Z})=\mathbf{q}]$ and density ratios, $\left.f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))\right/f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}))$ for $i\in\mathcal{K}^{\setminus\{k,j\}}$.

theoremLet Assumptions (ref) to (ref) hold. Then, the conditional expectations of $G(Y_{k}),G(Y_{j})$ are given by \begin{align*} &E[G(Y_{k})|\mathbf{V}=\mathbf{q}^{*}] \\ =&\left.\frac{\partial E[G(Y)D_{k}|\mathbf{Q}(\mathbf{Z})]}{\partial \mathbf{Q}(\mathbf{Z})}\right|_{\mathbf{Q(Z)}=\mathbf{q}^{*}}\left/\frac{\partial E[D_{k}|\mathbf{Q}(\mathbf{Z})]}{\partial \mathbf{Q}(\mathbf{Z})}\right|_{\mathbf{Q(Z)}=\mathbf{q}^{*}}, \\ &E[G(Y_{j})|\mathbf{V}=\mathbf{q}^{*}] \notag \\ =& \frac{\partial^{K-2}}{\partial \mathbf{Q}_{-j}(\mathbf{Z})}\left(\frac{\sum_{i\neq k,j}^{K-1}\left(\frac{f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))}{f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}))}\times\left.\frac{\partial}{\partial q_{i}}E[G(Y)D_{j}|\mathbf{Q}(\mathbf{Z})=\mathbf{q}]\right|_{(\mathbf{q}_{-j}, q_{j})=(\mathbf{q}_{-j}, q_{j}^{*})} \right)}{\sum_{i\neq k,j}^{K-1}\left(\frac{f_{\tilde{U}_{i,k}}(F_{\tilde{U}_{i,k}}^{-1}(q_{i}))}{f_{\tilde{U}_{b^{*},k}}(F_{\tilde{U}_{b^{*},k}}^{-1}(q_{b^{*}}))}\times\left.\frac{\partial}{\partial q_{i}}E[D_{j}|\mathbf{Q}(\mathbf{Z})=\mathbf{q}]\right|_{(\mathbf{q}_{-j}, q_{j})=(\mathbf{q}_{-j}, q_{j}^{*})} \right) } \right. \\ \times& \left.\left.\frac{\partial}{\partial q_{j}}E[(D_{j}+D_{k})|\mathbf{Q}(\mathbf{Z})=\mathbf{q}]\right|_{(\mathbf{q}_{-j}, q_{j})=(\mathbf{q}_{-j}, q_{j}^{*})} \right. \\ -&\left.\left.\left.\frac{\partial}{\partial q_{j}}E[G(Y)D_{j}|\mathbf{Q}(\mathbf{Z})=\mathbf{q}]\right|_{(\mathbf{q}_{-j}, q_{j})=(\mathbf{q}_{-j}, q_{j}^{*})}\right)\right|_{\mathbf{q}=\mathbf{q}^{*}} \left/\frac{\partial E[D_{k}|\mathbf{Q}(\mathbf{Z})]}{\partial \mathbf{Q}(\mathbf{Z})}\right|_{\mathbf{Q(Z)}=\mathbf{q}^{*}} \end{align*} where we define $\mathbf{a}_{-i}$ as the vector that removes $a_{i}$ from the original vector, namely $\mathbf{a}_{-i}=(a_{1},\cdots, a_{i-1}, a_{i+1}, \cdots, a_{n})$.

Conclusion

We study the identification of MTE with multivalued treatments. Our model is based on a multinomial choice model with utility maximization. We establish sufficient conditions for the identification of marginal treatment effects with multidimensional unobserved heterogeneity, which reveals treatment effects conditioned on the willingness to participate in treatments against a specific treatment. Our MTE generalizes the MTE defined in HV05ecta in binary treatment models and our identification strategy does not depend on the large support assumption required by HV07HE71 and HRV08AES. We also establish a sufficient condition for identifying thresholds.

\singlespace