Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
99,760 characters · 21 sections · 35 citation commands
Treatment Effects of Multi-Valued Treatments in Hyper-Rectangle Model
In economic applications, researchers frequently encounter situations involving treatments with multiple levels. For instance, participants in labor training programs may receive varying types or intensities of training based on their individual characteristics. Similarly, in medical practice, patients may be categorized into different care groups depending on specific indicators of their health conditions. In such scenarios, traditional binary treatment models become inadequate, and researchers must instead rely on multi-valued treatment models. These models allow for a more accurate assessment of heterogeneous treatment effects, which is essential for effective policy evaluation and the optimal design of interventions.
However, applying these models in practice involves significant challenges. A key difficulty arises from the presence of unobserved heterogeneity, which complicates the relationship between treatment selection and outcomes. Since individuals may be grouped into treatments based on unobserved factors correlated with potential outcomes, standard methods often fail to provide reliable identification of treatment effects. Additionally, multi-valued treatments typically involve more complex selection mechanisms compared to binary treatments. As the number of treatment levels increases, characterizing selection behavior becomes increasingly demanding, and assumptions about treatment assignment need to be carefully justified. Traditional methods for identifying treatment effects often depend on strict assumptions such as strong functional forms, restrictive monotonicity conditions, or independence assumptions that may not hold in realistic empirical contexts. These limitations highlight the need for approaches that impose fewer and more credible restrictions while maintaining rigorous identification.
A key approach in recent literature on multi-valued treatment models involves the construction of a hyper-rectangle framework combined with instrumental variables to achieve identification. Specifically, lee2018identifying shows how the interaction between treatment assignment thresholds and the distribution of unobserved heterogeneity determines treatment selection. They further develop methods to identify the marginal treatment response, which serves as a fundamental building block for identifying various treatment effects under multi-valued treatments. Their approach demonstrates a flexible and practical path toward identification, providing a foundation for subsequent econometric analyses.
This paper builds upon previous studies by proposing a generalized framework for identifying marginal treatment responses within multi-valued treatment models. Specifically, I extend the hyper-rectangle model introduced by lee2018identifying with a customized set of assumptions. My approach considers scenarios where either the treatment selection thresholds or the distribution of unobserved heterogeneity is unknown, where one of them can be identified given the knowledge of the other one.
To facilitate identification, I introduce a decomposition of the treatment assignment mechanism by defining a concept named the "leading term". This decomposition simplifies the analysis by clearly structuring how instrumental variables interact with unobserved heterogeneity. Further, I distinguish between cases based on the rank of these leading terms. When a leading term has full rank, the model achieves point identification of the marginal treatment responses. In contrast, if no leading term achieves full rank, I propose an additional ranked treatment assumption, enabling set identification under more realistic and flexible conditions.
Additionally, I develop a hypothesis testing method tailored explicitly to evaluate the effectiveness of policy interventions based on the identified treatment effects. This testing procedure provides policymakers with a practical tool to rigorously assess how changes in treatment assignment rules influence aggregate outcomes, thus expanding the practical applicability of multi-valued treatment models.
The literature on treatment effect identification has historically focused on binary treatment settings, where individuals are assigned either to a treatment or control group. A seminal contribution by imbens1994identification provided conditions for identifying local average treatment effects using instrumental variables, clarifying the role of compliance behavior in treatment assignment. angrist1996identification formalized the instrumental variable approach within the Rubin Causal Model, explicitly characterizing the instrumental variable estimand as the average causal effect for the subgroup of compliers. Subsequently, heckman1997instrumental systematically examined the identifying assumptions necessary for estimating average treatment effects and treatment effects on the treated.
Building on these foundations, abadie2003semiparametric developed semiparametric instrumental variable estimators to identify average treatment effects under weaker assumptions. Furthermore, imbens2004nonparametric extended the analysis by reviewing various nonparametric methods to estimate the treatment effects and analyzing the plausibility of key assumptions. hahn1998role studies the average treatment effect and average treatment effect on the treated, clarifying the role of the propensity score in efficient estimation under unconfoundedness. heckman2005structural introduces the concept of marginal treatment effects to unify the nonparametric literature on treatment effects with structural econometric estimation, allowing for heterogeneity in treatment responses. mogstad2018using show how instrumental variables methods can identify policy-relevant treatment parameters beyond the subpopulation directly affected by the instruments through a unified framework based on marginal treatment effects.
While much of the literature on treatment effects focuses on point identification, recent methodological advancements have focused on relaxing stringent assumptions associated with traditional instrumental variable frameworks, where treatment effects only admit partial identification. chen2023differential propose an approach using differential treatment effects to partially identify average or heterogeneous treatment effects under unmeasured confounding, along with a two-stage inference procedure to conduct statistical inference when point identification is infeasible.
Extending beyond the binary treatment context, recent econometric research on multi-valued treatment models captures more realistic scenarios involving multiple intervention levels. imbens2000role extends the propensity score methodology from the binary treatment setting to multi-valued treatments, facilitating the estimation of average causal effects. cattaneo2010efficient develops efficient semiparametric estimators for multi-valued treatment effects defined by a collection of possibly over-identified non-smooth moment conditions when the treatment assignment is under ignorability. Further developments by heckman2018unordered introduce an unordered monotonicity assumption to identify treatment effects of multi-valued treatments without imposing a strict hierarchy among treatments. Collectively, these studies offer rigorous methodological tools and clarify essential identification issues in treatment effect models.
Beyond estimation, an important strand of the treatment effect literature focuses on hypothesis testing, particularly assessing whether a treatment has any impact and whether treatment effects vary across subpopulations. crump2008nonparametric develop nonparametric tests on whether average treatment effect is zero, as well as detecting conditional average treatment effect heterogeneity across subpopulations. wu2021randomization develop a version of the Fisher randomization test adapted for weak null hypotheses that do not imply sharp potential outcome restrictions. Their studentized test statistic achieves finite-sample exactness under the sharp null and retains asymptotic validity under the weak null, offering a model-free approach robust to test treatment effect heterogeneity. Under instrumental variable frameworks, abadie2002bootstrap proposes a bootstrap procedure to test distributional hypotheses of treatments effects, including tests of equality of distributions and stochastic dominance. More recently, chernozhukov2018generic proposes general inference methods based on machine learning proxies, facilitating estimation and testing of heterogeneous treatment effects in high-dimensional randomized experiments.
As mentioned above, it is possible that treatment effects can only be partially identified under scenarios when strong identification assumptions are relaxed. There is a growing body of work on hypothesis testing in set-identified frameworks. A foundational contribution by imbens2004confidence proposes confidence intervals that asymptotically cover the true value of the parameter with fixed probability rather than cover the entire identified region, and its exact coverage probabilities converge uniformly to their nominal value. beresteanu2008asymptotic develop a limit theory for estimators of identification regions, based on set-valued random variables and convergence in the Hausdorff metric, allowing for construction of valid confidence regions for set-identified parameters. romano2010inference propose an approach to construct uniformly valid confidence regions for identified sets defined through general objective functions. galichon2009test design a testing framework for non-identifying model restrictions that can be inverted to form confidence sets. Their approach complements the moment inequality-based procedures of chernozhukov2007estimation, offering additional flexibility for hypothesis testing when model is incomplete. Together, these contributions provide a comprehensive foundation for conducting rigorous inference and testing in treatment effect models as well as set-identified features under a wide range of identifying assumptions.
This paper contributes to the literature on multi-valued treatment effect identification in econometrics in several ways. First, this paper generalizes the hyper-rectangle model introduced by lee2018identifying, broadening its applicability. Specifically, I only require the knowledge of either the distribution of unobserved heterogeneity or treatment assignment thresholds, and consider those two scenarios. Additionally, my framework allows for a more flexible treatment selection mechanism, relaxing the requirement that treatment assignment must depend on all dimensions of unobserved heterogeneity. This generalized structure accommodates richer empirical settings and makes the model more applicable.
Second, I introduce the concept of leading terms to systematically analyze treatment assignment mechanisms. This approach distinguishes between scenarios where the leading term is of full rank or not, establishing clear conditions for identification. In particular, when a full-rank leading term is unavailable, I propose a novel ranked treatment assumption that achieves set identification of marginal treatment responses. This assumption aligns closely with realistic empirical settings, where higher treatment intensities typically yield systematically larger or smaller treatment effects. Thus, my framework extends the practical applicability and relevance of multi-valued treatment effect models.
Third, the paper contributes methodologically by developing a hypothesis test framework tailored to assess policy interventions' effectiveness. The approach I propose allows researchers and policymakers to evaluate rigorously whether changes in treatment assignment mechanisms yield statistically significant improvements in aggregate outcomes. This advancement bridges the gap between econometric theory and practical policy evaluation, providing a direct tool for policy analysis and decision-making.
The rest of the paper is organized as follows. Section (ref) introduces the hyper-rectangle model and formalizes the treatment assignment process. Section (ref) presents the identification of treatment assignment threshold or distribution of unobserved heterogeneity. Section (ref) describes the strategy of identifying marginal treatment responses in different scenarios. Section (ref) discusses a group of identifiable treatment effects besides marginal treatment responses. Section (ref) develops the hypothesis test for evaluating the effectiveness of policies. Finally, Section (ref) concludes with a discussion of implications and future directions.
This section presents the micro-econometric model which constitutes the core analytical framework for assessing treatment assignments in this study. While the foundational structure of this model is inspired by and primarily conforms to lee2018identifying, the model incorporates unique adjustments and a customized set of assumptions, which were conceived to cater to the specific context and requirements of my study.
To provide a brief overview, a standard treatment model encompasses an observed outcome, denoted by $Y \in \mathds{R}$, and the observed treatment, symbolized as $D=1,2,...,T$. Accompanying these, we have observed covariates, denoted by $X \in \mathds{R}^{\bar B}$, where $\bar B \in \mathds{N}^+$ signifies the dimension of $X$. Complementing these entity is a random vector $V=(V_1,...,V_J) \in \mathds{R}^J$ that accounts for unobserved heterogeneity. In this context, $J \in \mathds{N}^+$ represents the dimension of $V$. I restrict the outcome $Y$ to be strictly positive and bounded above, that is, $0<Y<\bar Y$ for some $\bar Y\in \mathds{R}^+$. I also confine the unobserved heterogeneity $V$ to the interval $[0,1]^J$. These constraints simplify the ensuing mathematics without significantly undermining the generality of the model.
Additionally, for each outcome, we observe an instrumental variable, denoted by $Z=(Z_1,...,Z_{\bar{W}}) \in \mathds{R}^{\bar{W}}$, where ${\bar{W}} \in \mathds{N}^+$ is the dimension of $Z$. These instruments play a crucial role in identification and estimation strategies in the later analysis.
The observed data is composed of a sample $\{ (Y^o,D^o,Z^o,X^o): o=1,...,N_o \}$, with $N_o \in \mathds{N}^+$ denoting the sample size. For the sake of notational simplicity, I suppress the conditioning on $X$ in subsequent discussions, and all results should be interpreted as conditional on $X$.
To formalize the relationship between the observed outcome and treatment, let $Y_k$, $k=1,...,T$ denote the potential outcome under treatment $k$, and let $D_k \equiv \mathds{1}\{D=k\}$, $k=1,...,T$. The observed outcome $Y$ can thus be expressed as a sum of potential outcomes weighted by their respective treatment indicators: $Y=\sum\limits_{k=1}^T Y_kD_k$.
The legitimacy of the chosen instruments $Z$ is substantiated through the following assumption:
The objective of this analysis is to estimate the Marginal Treatment Response (MTR), defined as $E[Y_k|V=v]$, which captures the expected potential outcome $Y_k$ given a specific realization of the unobserved heterogeneity $V=v$. I impose the continuity of the MTR in the Data Generating Process (DGP):
This continuity assumption ensures the mathematical tractability of the model and allows us to employ a host of econometric techniques to analyze the data.
Next, I delve into the mechanism determining the treatment variable $D$. This is controlled by the confluence of a series of conditions. Specifically, the conditions entail a set of inequalities involving unobserved heterogeneity $V$: $V_1<Q_1(Z)$ or $V_1 \geq Q_1(Z)$, ..., $V_J<Q_J(Z)$ or $V_J \geq Q_J(Z)$. Here, $Q(Z)=(Q_1(Z),...,Q_J(Z))$ represents a vector of functions of the instrumental variable $Z$, and it acts as the threshold for $V$. I impose a key assumption about the support of the threshold, which I restrict to be the open interval $(0,1)$. This assumption is formalized as follows:
Assumption (ref) precludes uninteresting scenarios in which the threshold $Q_j(Z)$ reaches a boundary point, rendering an explicit threshold ineffective in influencing the treatment assignment. In other words, it ensures that all realizations of $Q(Z)$ lie within the interior of its range of variation. Besides, the open interval of $Q(\mathcal{Z})$ also indicates any realization of $Q(Z)$ belongs to the interior of its range of variation. Furthermore, the assumption that the support of $Q(Z)$ is dense in $(0,1)$ guarantees that changes in the instrument $Z$ can bring about the entire spectrum of variation in the threshold $Q(Z)$. This is critical for the identification analysis to follow, as it ensures that there is sufficient exogenous variation in the instrument to trace out the treatment effect of interest. In the following discussions, I will separately consider the cases where the threshold function $Q(Z)$ is known and unknown.
In this model, treatment $D=k$ is selected if and only if the specified inequalities are satisfied. To formalize this, consider the $\sigma$-algebra $\sigma_{\{ V,Q(Z) \}}$ generated by the set
I put forward the following assumption
Any set in the $\sigma$-algebra corresponds to taking unions, intersections, and complementation of the sets in ((ref)). Therefore, we can envision the treatment model as being constructed on a hyperplane, where every $V_j$ forms one dimension. Given that $V_j\in [0,1]$ and is partitioned by $Q_j(Z)$, each hyper-rectangle is formed from the intersection of chosen $V_j<Q_j(Z)$ or $V_j\geq Q_j(Z)$, $j=1,...,J$. Each treatment is linked to a combination of one or several hyper-rectangles in this hyperplane. The following example demonstrates this style of treatment determination, providing a more concrete understanding of the framework presented above.
To precisely illustrate the treatment selection process, I introduce a function $d_k(V,Q(Z))$ such that the treatment indicator can be represented as $\mathds{1}\{ D=k \}=d_k(V,Q(Z))$. For the sake of simplification, denote
Given that $\mathds{1}\{ V_j \geq Q_j(Z) \} = 1- \mathds{1}\{ V_j < Q_j(Z) \}$, it is observed that for each treatment $k$, the indicator function $d_k(V,Q(Z))$ equates to the summation, product, and difference of selected $\mathds{1}\{ V_j<Q_j(Z) \}$, $j=1,..., J$, represented in a polynomial form.
Now, let $\mathcal{L}$ denote the set of all non-empty subsets $l$ of $\mathbf{J} \equiv \{ 1,..,J \}$. In this context, $d_k(V,Q(Z))$ can be articulated according to how the hyper-rectangle for treatment $k$ is constructed. Mathematically, it is expressed as
where $e_l^k \in \{0,1\}$ signifies the existence of term $l$ in the set for treatment $k$, and $r_{lj}^k \in \{0,1\}$ demonstrates whether $V_j<Q_j(Z)$ or $V_j\geq Q_j(Z)$ is involved. Upon polynomial expansion, $d_k(V,Q(Z))$ can be distinctly expressed in a decomposed form as
where $c_l^k \in \mathds{Z}$ denotes the integer coefficient of the term $l$ for treatment $k$.
Reverting to Example (ref), we can lucidly express each treatment in accordance with Equation (ref) and (ref) respectively.
In addition, it is assumed that each realization of $(V,Z)$ must be associated with one and only one treatment. Formally, this can be expressed as follows:
This assumption reinforces the exclusivity of the treatment assignment, precluding any void or overlap in treatments for any given combination of $V$ and $Z$. Moreover, it ensures that the expression given in Equation (ref) is uniquely determined. Therefore, the treatment determination mechanism can be expressed as a function:
Furthermore, I introduce the following assumption:
Assumption (ref) essentially states that we preclude the possibility of any dimension $j$ that does not contribute to the diversity of treatment assignments. This assumption is made without loss of any heterogeneity in our model. If a certain dimension of unobserved heterogeneity is not influential in determining a treatment, it can be eliminated from both $V$ and $Q(Z)$ without fundamentally altering the model. In doing so, we reduce the dimensionality from $J$ to $J-1$. This reduction simplifies the mathematical manipulations and clarifies the interpretations of our model.
An important aspect to note is that any $V_j$, $V_j'$ ($j\neq j'$) can be correlated or can even be identical. This implication is significant as the hyper-rectangle model can seamlessly incorporate the truncated model for treatment determination. Specifically, for a one-dimensional variable $V\in [0,1]$, it is possible to have two or more truncation points that divide the uniform interval into three or more segments, resulting in multi-valued treatments.
Figure (ref) provides an example with $J=4$, where $V_1=V_2$ and $V_3=V_4$. The scenario can then be represented in a two-dimensional plane instead of a four-dimensional space.
However, in order to incorporate this scenario within the framework of my model, it is necessary to introduce an "outside treatment option" $D=5$ to account for all cases not captured by this plane, such as $V_1<Q_1(Z)$, $V_2\geq Q_2(Z)$. Thus, in the scenario depicted by Figure (ref), the total number of treatments should be $T=5$. \footnote{ The classical unconfoundedness assumption, $(Y_k)_{k=1}^J \perp\!\!\!\perp D \mid X$, is not required in this framework. The unobserved heterogeneity $V$ is correlated with $Y_k$, and also enters the treatment assignment rule since $D$ is a deterministic function of $(V,Z)$. Conditioning on $X$ and $Z$ does not break the dependence between $D$ and $Y_k$. }
To facilitate the estimation process, I will employ the term "term $l$" to denote each $c_l^k \prod_{j\in l} S_j(V,Q(Z))$ with $c_l^k \neq 0$ in Equation (ref). Consequently, $d_k(V,Q(Z))$ can be viewed as a finite combination of the term $l\in \mathcal{L}$.
Let $l_j=\mathds{1}\{ j\in l \}$, a term can be succinctly represented by a vector with coefficient as $l=c_l^k (l_1,...,l_J)$. For instance, in Example (ref), the $D=1$ case implies only one term $l=(0,1,1)$, and the $D=2$ case illustrates six terms, namely, $l^1=(0,1,0)$, $l^2=(0,0,1)$, $l^3=-(1,1,0)$, $l^4=-(1,0,1)$, $l^5=-2(0,1,1)$, and $l^6=2(1,1,1)$, as implied by Equation (ref). To further clarify, I propose the following definitions related to the term:
I revisit Example (ref) to illustrate the point of leading term. In the case of $D=1$, the only term present is the leading term, which holds a rank of $2$. In the case of $D=2$, there is a sole leading term $l^6$, with a rank of $3$. From the definitions, some noteworthy propositions can be yielded, which shed light on the fundamental characteristics of treatment decompositions.
Notably, the rank of the leading term could be less than the number of dimensions involved in treatment determination. An illustration of this can be found in the treatment scenario depicted in Figure (ref).
The corresponding analytical representation of the treatment is:
This expression implies three leading terms, $(1,1,0)$, $(1,0,1)$, and $(0,1,1)$. Interestingly, the rank of each of these leading terms is only two, instead of three. Intuitively, this reduction in complexity is attributed to the perfect predictability of one rectangle from another in this particular treatment determination mechanism, effectively reducing the degrees of freedom in treatment assignments. This feature underlines the flexibility of our model and its ability to accommodate a variety of treatment determination mechanisms.
In Section (ref), I introduced the threshold function $Q(Z)$, the knowledge status of which, known or unknown, significantly influences the strategies for empirical analysis. This section addresses each case in turn, and provide a methodology for identifying the threshold function, setting the stage for subsequent analyses.
I first turn to the scenario where the threshold function $Q(Z)$ is explicitly known. For illustrative purposes, consider Example (ref), in which the minimum grades required for passing the tests are clearly specified in the program descriptions. While these thresholds may differ depending on the candidate's attributes, they are fully observable to the researcher.
The explicit knowledge of the threshold function bestows significant analytical flexibility. Importantly, it eliminates the necessity for assuming prior knowledge of the distribution of unobserved heterogeneity. Instead, we can identify this distribution directly from the data, negating the need to pre-suppose it as a known entity. This approach simplifies the empirical model and bolsters its tractability. The nuances of this advantage and its subsequent implications will be comprehensively discussed in the following discourse.
Our analysis begins with a fundamental assumption regarding the distribution of the unobserved heterogeneity $V$. We denote the probability distribution function of this unobserved heterogeneity as $f_V$, and its cumulative distribution function as $F_V$. In order to discern $f_V$ within our model, we necessitate the following condition:
For any treatment $k$, the conditional probability of treatment assignment, denoted as $\Pr(D=k|Q(Z)=q)$, can be directly observed from the data as a function of $q$, where $q=(q_1,...,q_J)$ constitutes the realization of $Q(Z)$. Additionally, the following relationship holds:
where the second equality is justified by the conditional joint independence of $V$ and $Z$ stipulated in Assumption (ref). Given that $d_k$ is an indicator function, $\mathds{1}\{ d_k(v,q)=1 \} = d_k(v,q)$. Hence, with respect to Equation (ref), the conditional probability can be reformulated as:
Equation (ref) provides a decomposition of the conditional probability of being assigned to treatment $k$ into its constituent terms. However, the existence of a treatment with a full rank leading term, versus every treatment's leading term being non-full rank, will significantly impact subsequent analysis. The upcoming content of this section will explore them in greater details.
A common and particularly manageable case arises when there exists at least one treatment with a full rank leading term. In such scenarios, researchers can directly identify the distribution of the unobserved heterogeneity without additional stringent assumptions. This specific case is extensively discussed by lee2018identifying, and their results can be directly applied to our context.
In Equation (ref), the behavior of $S_j$ as an indicator function for the threshold, as implied by Equation (ref), implies that $q$ linearly modulates the area of integration. Furthermore, Assumption (ref) establishes the continuity of $f_V$. Assumption (ref), in turn, ensures that Equation (ref) is well defined within an open neighborhood of $q$ for all $q\in (0,1)^J$. Consequently, we can infer that Equation (ref) is differentiable across all dimensions of $q$.\footnote{For further proof details, please refer to Section 6 of lee2018identifying.}
Without loss of generality, I suppose it is treatment $k$ that has a full rank leading term. Denote this full rank leading term of as $\tilde{l}$. A differentiation of $\Pr(D=k|Q(Z)=q)$ with respect to all the dimensions of $q=(q_1,...,q_J)$ will render all terms outside $\tilde{l}$ null as they lack at least one dimension of $q$. Thus, we obtain:
From this derivative, we obtain:
which identifies the probability density function $f_V(v)$ almost everywhere throughout its entire support, $v\in(0,1)^J$.
In scenarios where a full rank leading term is not present on any treatment in the model, the direct identification of $f_V$ from any specific treatment exposure becomes infeasible. Instead, by emulating the analysis structure in Section (ref), we can learn certain marginal distribution functions via differentiation.
Consider a leading term $l$ for treatment $k$. Let $i_1,...,i_{|l|}$ be the numbering of $l$'s elements which are equal to one, that is, $l_{j}=1$ for $j\in \{ i_1,...,i_{|l|} \}$ and $l_{j}=0$ for the other cases. We denote $\{ -i_1,...,-i_{J-|l|} \}$ as the complement set of $\{ i_1,...,i_{|l|} \}$ in $\{1,...,J\}$.
For simplicity in notation, we introduce the following sets:
As per Equation (ref), because $l$ is a leading term, differentiating with respect to dimensions in $I_l^+$ will cancel all terms other than $l$ in the summation over $l\in \mathcal{L}$, as implied by Proporsition (ref). Thus, we find:
We refer to Equation (ref) as the distribution identification equation, as it provides us with a marginal distribution of $f_V(v)$ on the dimension $I_l^+$. In the pursuit of further simplicity in notation, we introduce $v_{I_l^-} \equiv (v_{-i_1}, \cdots, v_{-i_{J-|l|}})$ and $ q_{I_l^+} \equiv (q_{i_1}, \cdots, q_{i_{|l|}})$. Therefore, the left hand side of Equation (ref) can be re-expressed in a concise manner as
Here, $f_{I_l^+}$ is a notation I introduce to encapsulate this marginal probability density function.
By integrating the distribution identification equations from all the leading terms of all treatments, we can discern new insights. Suppose we have two leading terms, $l$ and $l'$, where $l \subseteq l'$. It follows from the definition of leading terms that they must originate from distinct treatments. The distribution identification equation of term $l$ is consequently implied by that of term $l'$, as long as we undertake an additional integration over the dimensions in $I_l'^+ \backslash I_l^+$. This implies that the distribution identification equation of term $l$ can be disregarded as it fails to provide new information.
Upon gathering all the distribution identification equations that have not been eliminated, we construct an equations system, which encapsulates the accessible information concerning the distribution of $V$. Let these terms be $l^1, ..., l^P$, where $P\in \mathds{Z}^+$ is the number of those terms, the system of equations can be represented as:
Here, $k^p$ corresponds to the treatment associated with term $l^p$, where $p=1,...,P$. In the next step, we will infer the joint distribution $f_V(v)$ with the information provided by these marginal distributions. Assumption (ref) guarantees the inclusion of all dimensions $j=1,...,J$ in the combination of indices ${I_{l^p}^+}$, $p=1,...,P$. Consequently, as stated by Sklar's theorem sklar1959fonctions, there exists a copula $C$ that satisfies the following relation:
Here, $\beta_C$ refers to the copula's coefficients, which may have finite or infinite dimension, and $F_{I_{l}^+}$ is the cumulative distribution of the marginal probability density $f_{I_l^+}$. Furthermore, Assumption (ref) imposes the continuity of $f_V(v)$, which ensures the copula is uniquely determined by Sklar's theorem. Therefore, the distribution $f_V(v)$ is identified across its support $v\in (0,1)^J$.
In order to establish $F_V(v)$ from the copula, we suggest an implementation which consists of the following steps.
Given the intrinsic properties of convergence of maximum likelihood estimator, we have the following theorem.
As implied by Theorem (ref), it is possible to recover the probability distribution of $V$ even in the absence of a full rank leading term.
The precise specification of the threshold function is frequently beyond the grasp of researchers, a phenomenon attributable to a multitude of factors. For instance, the allocation of treatment assignment may be underpinned by a latent group formation process that takes place behind the scenes, such as the influence of social networks. In such circumstances, individuals are implicitly categorised into distinct groups based on unobserved characteristics or shared experiences, resulting in a latent group structure. It is also plausible that the threshold function is kept undisclosed and thus unobservable. For instance, in the context of a training program, the organiser may conceal their exact categorisation criteria from researchers. Under such circumstances, the identification of $Q(Z)$ becomes crucial for further analysis.
To identify the threshold, knowledge of the distribution of the unobserved heterogeneity $V$ is necessary. We make the following assumption about the distribution of $V$:
Assumption (ref) requires that $V$ is dense everywhere in probability measure. I write the probability density function for $V$ as $f_V$. Armed with knowledge about the distribution of $V$, we can pinpoint the threshold function through the conditional probability distribution of treatment assignment. For a fixed treatment $k$, $\Pr(D=k|Z)$ can be directly observed from the sample and is treated as a function of $Z$. The conditional independence of $V$ and $Z$ in Assumption (ref) yields
where the fact $\mathds{1}\{ d_k(v,Q(Z))=1\} = d_k(v,Q(Z))$ is applied in the above equation. By inserting the treatment assignment decomposition from Equation (ref) into the probability distribution, we obtain:
Here, each $\alpha_{lj}(Z)$, $j=1,...,J$ is a function of $Z$ defined as:
Combining Equation (ref) for $k=1,...,T$ yields a system of equations:
Denote the matrix of coefficients in system of Equations (ref) as $\{ c_l^k \}$. The following theorem states the identification of threshold function:
In most cases, the only constraint on the system arises from the completeness condition assumed in Assumption (ref), which implies that the sum of the probabilities for all treatments equals one. However, in some cases, additional constraints may be imposed on the probability distribution of each assigned treatment due to practical considerations of the treatment process. For example, if the treatment probabilities are restricted by external factors such as policy guidelines or operational limitations, this could further reduce the rank of the matrix.
However, obtaining an explicit solution for $\{Q_j(Z)\}_{j=1}^J$ from this equations system through mathematical manipulations like differentiation or matrix operations is challenging, even though the form of $F_V$ is known. This is primarily due to the fact that Equation (ref) does not exhibit a separable form for $Q_j(Z)$ without assuming specific forms for the cumulative distribution function $F_V$. Nevertheless, it is possible to gain insights into the threshold function using a numerical approach.
Before delving into the specifics of the numerical methods employed, we first establish the loss function to be used throughout our numerical experiments. Let us consider a proposed threshold function $\bar{Q}(\cdot)=\left( \bar{Q}_1(\cdot),...,\bar{Q}_J(\cdot) \right)$, under which we can define a corresponding loss function.
The loss function measures the discrepancy between our model's predictions and the actual observed data. For this purpose, we use the sum of the squared differences between the computed probabilities under the proposed threshold function and the observed probabilities. More specifically, the loss function for a single observation $Z$ is given by
Here, $\bar{\alpha}_{lj}(Z)$ is defined as:
The loss function for a given $\bar{Q}$, across all observations of instruments $Z$, can thus be written as the sum of the individual losses:
This loss function, capturing the aggregate discrepancy between our model's predictions under the proposed threshold function and the observed data, will guide our subsequent numerical analyses.
To begin, we propose a parametric approximation of the true threshold function $Q$. This is realized by assuming a parametric form for $\bar{Q}(Z;\beta^Q)$, where $\beta^Q$ denotes the parameters. Formally, we have $\bar{Q}(Z;\beta^Q) = \left(\bar{Q}_1(Z;\beta^Q) ,..., \bar{Q}_J(Z;\beta^Q)\right)$.
The approximation is performed using a set of basis functions $\{q_{jt_q}(\cdot)\}_{t_q=1}^{T_q}$, for each $j=1,...,J$. Here, $T_q \in \mathds{N}^+$ is the number of basis functions used for approximating each $\bar{Q}_j$. For the parameters, we let $\beta^Q_j=(\beta^Q_{j1},...,\beta^Q_{jT_q})\in \mathds{R}^{T_q}$, and $\beta^Q=(\beta^Q_1,...,\beta^Q_J)$ encompass the coefficients of the entire parametrization. Therefore, each $\bar{Q}_j$ is expressed as a linear combination of the basis functions, parameterized by $\beta^Q_{j}$, and can be written as follows:
By specifying $\bar{Q}$ in this way, we re-interpret the loss function as a function of the parameters $\beta^Q$. Consequently, finding an approximation to $Q(\cdot)$ is now equivalent to solving the following optimization problem:
In solving the optimization problem, we can use algorithms like Gradient Descent or Newton's Method. This optimization yields an optimal set of parameters, denoted by $\hat{\beta}^Q$, which provides $\bar{Q}(\cdot;\hat{\beta}^Q)$ as an approximation to the true threshold function $Q(\cdot)$. The following theorem states the convergence of this approximation:
Once the approximation is obtained, it will be used for further analysis in the subsequent sections of this paper.
In this section, I focus on identifying the marginal treatment response $E[Y_k|V=v]$. Since the identification of $Q(Z)$ has been addressed in Section (ref), I will treat the threshold $Q(Z)$ as given and condition on $Q(Z)$ rather than $Z$ in the analysis.
Fixing a treatment $k$, we can observe $E[Y D_k | Q(Z)=q]$ from the sample data, where $q=(q_1,...,q_J) \in [0,1]^J$ is a realization of $Q(Z)$. By the model's construction, $Y D_k=Y_k D_k$. Thus, we can express:
where the last equality follows from the Law of Iterated Expectation.
Using the structure of conditional expectation and the independence of $Y_k$ and $V$ from $Z$, we have:
Since $E[Y_k | V, Q(Z)=q] = E[Y_k | V]$ and $E[d_k(V, Q(Z))|V, Q(Z)=q] = d_k(V, q)$, it follows that:
Using the decomposition of $d_k(V, Q(Z))$ from Equation (ref), we can rewrite $d_k(V, q)$ as a sum of terms $l^1,...,l^N$, where $N \in \mathds{N}^+$ is the number of non-zero terms. Denote $l(q)$ as the set $\{ V: V_j < q_j, \forall j \in l \}$. We then have:
Equation (ref) shows that the expected value $E[Y D_k | Q(Z)=q]$ is a sum of the marginal treatment responses integrated over the regions defined by $d_k(V, q)$.
Unlike Equation (ref), where $f_V(v)$ is independent of the treatment $k$, here $E[Y_k | V=v]$ varies with each treatment $k$. This adds complexity to the identification process because we need to account for the treatment-specific response.
To identify $E[Y_k | V=v]$ from observations, it is crucial to determine whether the leading term of $d_k(V, Q(Z))$ is full rank. This distinction is important and different from the discussion in Sections (ref) and (ref), where the focus was on the presence of at least one treatment with a full rank leading term. Here, we must consider each treatment $k$ separately. If a treatment does not have a full rank leading term, it falls into the category of no full rank leading term.
In this analysis, the variation of $E[Y_k | V=v]$ with different treatments means that we cannot combine equations across different treatments to increase the information available for identification, unlike in Section (ref).
When the treatment $k$ has a full rank leading term, we can follow the method described by lee2018identifying to identify the marginal treatment response within our framework. This approach, similar to the one discussed in Section (ref), relies on the technique of taking derivatives to isolate $E[Y_k|V=v]$ from the summation and integration in Equation (ref).
In Equation (ref), the parameter $q$ determines the integration region through the set $l^n(q)$, for $n=1,...,N$. These regions expand or contract linearly with changes in $q$. Given that $f_V(v)$ is continuous by Assumptions (ref) and (ref), and $E[Y_k|V=v]$, for all $k=1,...,T$, is locally equicontinuous by Assumption (ref), we can ensure the differentiability of Equation (ref) across all dimensions of $q$. Assumption (ref) further guarantees that for any $q \in (0,1)^J$, Equation (ref) is well-defined within an open neighborhood of $q$. Therefore, we can differentiate Equation (ref) with respect to all components of $q$.\footnote{For detailed proof, see Section 6 of lee2018identifying.}
Consider the full rank leading term of treatment $k$, denoted by $\tilde{l}$. When we differentiate $E[Y D_k | Q(Z)=q]$ with respect to all dimensions of $q=(q_1,...,q_J)$, all terms except the one involving $\tilde{l}$ become zero because they do not involve all dimensions of $q$. Thus, we have:
From this derivative, we can identify the marginal treatment response as
almost everywhere over the entire support $v \in (0,1)^J$.
When the treatment $k$ does not have a full rank leading term, identifying the marginal treatment response becomes challenging. Some dimensions of $V$ are not involved in the treatment assignment process, which prevents the direct identification of treatment effects. However, by connecting these dimensions with observations from other treatments, we can still achieve identification, although it will be set identification rather than point identification.
Following the structure of analysis in Section (ref), we can differentiate the dimensions of $v$ associated with each leading term of treatment $k$. This will generate a conditional marginal treatment response for each leading term.
Let the leading terms of treatment $k$ be $l^1, ..., l^{N_k}$, where $N_k \in \mathds{N}^+$ is the number of leading terms for treatment $k$. Fix a leading term $l^i$. Recall the notation introduced in Section (ref). By Equation (ref), differentiating $E[Y D_k | Q(Z)=q]$ with respect to $q_{I^+_{l^i}}$ yields:
where Proposition (ref) ensures that differentiating all elements in a leading term eliminates other terms. The conditional expectation of $Y_k$ given $V$ in the dimensions involved in leading term $l$ is called the conditional marginal treatment response:
where $f_{I^-_{l^i} | I^+_{l^i}}(v_{I^-_{l^i}} | v_{I^+_{l^i}})$ is the conditional probability distribution function of $V_{I^-_{l^i}}$ given $V_{I^+_{l^i}}$, defined as:
where $f_{I^+_{l^i}}$ is the marginal distribution of $V_{I^+_{l^i}}$ defined in Equation (ref). Therefore, Equation (ref) implies:
which allows us to point identify the conditional marginal treatment response $E[Y_k | V_{I^+_{l^i}} = v_{I^+_{l^i}}]$ almost everywhere over its support.
Equation (ref) provides a decomposition of the conditional marginal treatment response by integrating out the dimensions not in the leading term, $I^-_{l^i}$. Combining all the leading terms of treatment $k$, we obtain a system of equations:
almost everywhere for all $v_{I^+_{l^i}} \in [0,1]^{|l^i|}$, $i=1, ..., N_k$.
However, identifying $E[Y_k|V=v]$ from this equation system is challenging, especially when $\bigcup\limits_{i=1}^{N_k} I^+_{l^i} \subsetneqq \mathbf{J}$. This means some dimensions $j\in \mathbf{J}$ are never involved in the treatment determination mechanism for treatment $k$. Nevertheless, by assuming the ranking of the marginal treatment response with respect to the level of treatment, we can achieve set identification for $E[Y_k|V=v]$. This assumption is stated as follows:
This assumption is similar to the monotone treatment response assumption in manski1997monotone. Intuitively, Assumption (ref) means that, conditional on the unobserved heterogeneity, higher treatment levels are expected to generate better outcomes. This assumption is particularly plausible in applications where the treatment variable represents an ordered intensity. For example, in education studies, treatment levels might correspond to the number of years of schooling completed; in labor economics, they might capture the duration or intensity of a training program; in health economics, they may represent different dosage levels of a medical intervention. In such cases, it is often reasonable to expect that higher treatment intensities yield systematically stronger effects on outcomes on average, conditional on unobserved characteristics. \footnote{This ranked treatment assumption can be extended to accommodate other plausible empirical settings. For instance, if higher treatment intensities are expected to produce lower outcomes on average, the assumption remains compatible with the framework after reversing the direction of the inequality. Moreover, even in cases where treatments are not ordinal in a global sense, such as when individuals are assigned to different physical training programs based on multidimensional characteristics like height, endurance, or strength, it may still be reasonable to assume a personalized ranking of treatment effects. That is, for each individual, conditional on $X$ and $V$, there exists a known ordering over treatments based on expected effectiveness. As long as such an individual-specific ordering can be specified over all treatments, the analytical structure developed in this paper remains applicable, with minor adjustments to the integration domain or inequality direction. Hence, the assumption is stated in a common form with minimized loss of generality.}
Note that the ranking of the marginal treatment response here is different from the monotonicity of treatments in previous econometric studies on treatment effects, such as imbens1994identification and heckman2018unordered. In their work, monotonicity indicates that when the instrument changes, there cannot be two individuals moving their treatment assignments in opposite directions. In other words, if the instrument causes one individual to switch from treatment group $k$ to treatment group $k'$, it should not cause another individual to switch from treatment group $k'$ to treatment group $k$. In this paper, the ranking of treatments assumes an order of marginal treatment response across different treatment levels. Monotonicity is not required in my setting, as the flexibility and high-dimensionality of the hyper-rectangle setup naturally allow for individuals to move in opposite directions.
Equipped with Assumption (ref), consider a treatment \(k' < k\), whose leading terms are denoted by \(l'^1, \ldots, l'^{N_{k'}}\), where \(N_{k'} \in \mathds{N}^+\) is the number of leading terms associated with treatment \(k'\). For any given leading term \(l'^{i'}\), Equation (ref) implies that
for almost every \(v_{I^+_{l'^{i'}}} \in [0,1]^{|l'^{i'}|}\), where \(i'=1,\ldots,N_{k'}\). The term \(E[Y_{k'}|V_{I^+_{l'^i}}=v_{I^+_{l'^i}}]\) is point identified from the data, analogous to Equation (ref).
Similarly, for treatments satisfying \(k' > k\), an analogous inequality in the opposite direction holds. For each leading term \(l'^{i'}\) of treatment \(k'\), we have
for almost every \(v_{I^+_{l'^{i'}}} \in [0,1]^{|l'^{i'}|}\), with \(i' = 1, \ldots, N_{k'}\). For ease of exposition, I refer to the relationship in Equation (ref) or (ref) as the inequality constraint between treatment \(k\) and the leading term \(l'^{i'}\) of treatment \(k'\).
Given these constraints, we proceed to analyze the partial identification of marginal treatment responses. I first consider the identification of a single marginal treatment response and then extend the analysis to jointly characterize all marginal treatment responses across treatments.
Consider the identified set of a single marginal treatment response $E[Y_k|V=v]$. It is notable that the identification problem is no longer focusing on a single point $V=q$ but on a function over the interval $[0,1]^J$. The identified set for $E[Y_k|V=v]$ is then given by
which is composed of a continuum of constraints for measurable, locally bounded, and locally equicontinuous function $E[Y_k|V=v]$ on $[0,1]^J$. The following theorem states that $\mathcal{I}^0_k$ is sharp.
Equation (ref) defines a total of $\sum_{k=1}^T N_k$ constraints for the function $E[Y_k|V=v]$. However, not all of these constraints are active. We can simplify the identified set by eliminating slack constraints.
Specifically, consider the following two types of redundancy: First, suppose treatment $k$ has a leading term $l$, and treatment $k'$ has a leading term $l'$, with $l' \subseteq l$. In this case, the inequality constraint between $k$ and $l'$ is redundant. This is because the marginal treatment response can be point-identified conditional on $V_{I^+_l}$; through integration, one can also identify it conditional on $V_{I^+_{l'}}$, rendering the constraint involving $l'$ non-informative.
Second, consider three treatments $k'' < k' < k$, where $k''$ has a leading term $l''$ and $k'$ has a leading term $l'$ such that $l'' \subseteq l'$. I claim that the inequality constraint between $k''$ and $k$ via $l''$ can be eliminated. This follows from combining the inequality between $k''$ and $k'$ via $l''$ and that between $k'$ and $k$ via $l'$. Specifically:
Thus, the constraint between $k''$ and $k$ via $l''$ is implied by other constraints and can be removed. A similar argument applies when $k < k' < k''$ and $l'' \subseteq l'$. Note, however, that this logic does not extend to the cases where $l' \subseteq l''$, or where $k''<k<k'$.
Finally, after eliminating all such redundant inequalities, any remaining equation in (ref) that involves only the leading terms removed in the above step can also be discarded, as they provide no additional identifying information.
Let $\mathcal{I}_k^1$ denote the simplified constraint set after this elimination. From the preceding analysis, it follows that $\mathcal{I}_k^1 = \mathcal{I}_k^0$.
If we are going to jointly identify a set of marginal treatment responses $\{ E[Y_k|V=v] \}_{k\in K}$, where $K$ is a nonempty subset of $\{1,...,T\}$. Note that the joint identification will include internal/incur constraint between different marginal treatment responses, and we need to consider this constraint (interdependence introduced by pooling the marginal treatment responses together). Thus the identified set is given by
which satisfies both the constraint given by $\mathcal I_K^0$ in Equation (ref) as well as the inter-treatment ranking constraints implied by Assumption (ref).
There are also redundant constraints in Equation (ref). First, if $k,k'\in K$, the inequality constraints between treatment $k$ and leading terms of $k'$ are redundant, since they are directly implied by the equality constraint of leading terms of $k'$ and the ranking constraint, and the same for the constraints between treatment $k'$ and leading terms of $k$. Second, we denote $\overline k_K \equiv \min\{\kappa\in K:\kappa> k \}$ and $\underline k_K \equiv \max\{\kappa\in K:\kappa< k \}$, if $k\notin K$, the inequality constraints between any treatment $k'\in K\backslash\{\overline k_K, \underline k_K \}$ and leading terms of $k$ are redundant, since they are implied by the constraint between treatment $\overline k_K$ or $\underline k_K$ with leading terms of $k$ and the ranking constraint.
Similarly, we denote the identified set of marginal treatment responses after eliminating the redundant constraints as $\mathcal I_K^1$, thus $\mathcal I_K^1=\mathcal I_K^0$.
To characterize the identified set for marginal treatment responses $\{E[Y_k|V=v]\}_{k\in K}$, where $K$ may be a singleton or a set with multiple elements, I adopt a sieve-based parametric decomposition of each function $E[Y_k|V=v]$ with respect to $v = (v_1,...,v_J)$. The sieve method transforms the original infinite-dimensional problem into a finite-dimensional one, reducing computational complexity and improving tractability. However, this approximation introduces bias due to the limited expressiveness of finite basis expansions, and may inflate variance if the number of basis functions is chosen inappropriately.\footnote{See chen2007handbook for a comprehensive treatment of sieve estimation.}
where $\{b_{kt}(\cdot)\}_{t=1}^{T_k}$ are known basis functions from $[0,1]^J$ to $\mathds{R}$, $T_k\in \mathds{N}^+$ is the number of basis functions, and $\theta_k \equiv (\theta_{k1},...,\theta_{kT_k}) \in \mathds{R}^{T_k}$ parameterizes the function $E[Y_k|V=v]$ for each $k\in K$.
We then define the feasible set for $\theta_K \equiv \{\theta_k\}_{k\in K}$ as:
This parameterization allows us to impose the identification constraints in $\mathcal{I}^1_K$ directly on the coefficients $\theta_K$, converting the infinite-dimensional constraint system into a finite-dimensional one. To implement this, we sample a finite set of $v$ values from its distribution. These sampled values are used to approximate the constraints in the system, enabling estimation of bounds for $\theta_K$.
Computing the feasible set $\Theta_K$ may be computationally demanding due to the potentially large number of constraints. To improve numerical feasibility and avoid empty identified sets, researchers can introduce slack terms into both equality and inequality constraints, which allow for small tolerances. Monte Carlo simulation methods are particularly useful in this setting. By repeatedly sampling $v$ and solving the associated inequality systems, we can numerically characterize $\Theta_K$.
It is important to note that the use of sieve approximation may result in a loss of accuracy of the identified set $\mathcal{M}_K$ because the true function $E[Y_k|V=v]$ may not lie exactly within the sieve space. This approximation error diminishes as the sieve space becomes richer, but in finite samples, the resulting $\mathcal{M}_K$ only approximates the identified set. Nevertheless, this approach provides a tractable and informative description of the treatment effects of interest.
The final form of the feasible identified set for $\{E[Y_k|V=v]\}_{k\in K}$ is:
Equation (ref) thus describes a feasible approximation of the identified set for marginal treatment responses for treatments $k \in K$, incorporating both model-based restrictions and computational considerations.
Identifying the marginal treatment response is essential, as it provides the foundation for defining a wide of treatment effects. In my multi-valued treatment setting, when restricting attention to comparisons between two treatment levels, one can define treatment effects analogous to those in binary treatment models. As shown in mogstad2018using, these binary-type treatment effects can be expressed in terms of the marginal treatment response. Beyond such pairwise comparisons, the multi-valued treatment structure allows for the definition of more general treatment effects involving multiple treatment levels. Examples include:
These treatment effects share a common structure: they are linear functionals of the marginal treatment responses. Accordingly, their identification reduces to identifying $\{E[Y_k|V=v]\}_{k\in K}$ for a subset $K \subseteq \{1,\dots,T\}$. Letting $G(\{E[Y_k|V=v]\}_{k\in K})$ denote any such treatment effect, the identified set is
This set can be computed using the same procedures introduced in Section (ref). Although these treatment effects may not be point identified, the identified set in Equation (ref) provides informative bounds that can guide empirical evaluation and policy analysis.
Equation (ref) defines the APRTE when a policy changes the treatment assignment mechanism from $(\mathbf{d}, Q)$ to $(\mathbf{d}', Q')$, conditional on $Z$. To evaluate the aggregate welfare impact of such policy changes, I define the Gross Policy Relevant Treatment Effect (GPRTE) based on $N_o$ observed units as:
where $\omega_o \geq 0$ denotes the weight assigned to observation $o$, reflecting policy makers' welfare objectives. Without loss of generality, I normalize $\sum_{o=1}^{N_o} \omega_o = 1$. A common choice is $\omega_o = 1/N_o$.
Such scenarios frequently arise in policymaking contexts. For example, in labor training programs as described in Section (ref), when individuals are assigned to different training contents, a policy reform may replace any discretionary assignment with a talent-based mechanism that depends more systematically on individual-level data. Similarly, in education, students may receive varying levels of instructional support under practices, while a new policy might implement a standardized test-based allocation rule. In both examples, the policy shifts the assignment mechanism and the resulting effect on social outcomes is captured by the GPRTE.
A natural question arises as to whether the proposed policy improves social outcomes. To address this, I develop a framework to test whether the GPRTE is statistically different from zero. For notational convenience, define the individual-level contribution:
where $\mathbf{Z}^{N_o} \equiv (Z^1, \ldots, Z^{N_o})$. Then the GPRTE can be expressed as \[ \Delta W \equiv \sum_{o=1}^{N_o} \omega_o \Delta \mu^o (\mathbf{d}, Q, \mathbf{Z}^{N_o}), \] where the dependence of $\Delta W$ on $\mathbf{d}', Q', \mathbf{d}, Q$, and $\mathbf{Z}^{N_o}$ is suppressed for simplicity.
A key challenge is that the treatment effects $E[Y_k | V = v]$ may not be point identified, as discussed in Section (ref). Therefore, I consider two cases, when the GPRTE is point identified and when it is set identified, and develop appropriate testing methods for each.
I first consider the case in which the GPRTE is point identified. This occurs when all treatment levels involved in the policy change, both before and after implementation, have full-rank leading terms. Under this condition, the marginal treatment responses for the relevant treatments are point identified, and so is the GPRTE defined in Equation (ref).
To assess the impact of the policy, the hypothesis testing problem is specified as:
where $\Delta W$ denotes the GPRTE, representing the aggregate effect of the policy intervention on social outcomes.
The first step is to compute the observed GPRTE from the data. The following algorithm outlines this procedure:
Consider the case where $N_o$ is fixed and $M \rightarrow \infty$. For a given observation $o$, and conditional on $Z^o$, the quantity \[ E[Y_{\bar{D}'^o_m} | V = v_m^o, Z^o] - E[Y_{\bar{D}_m^o} | V = v_m^o, Z^o] = \sum_{k=1}^T (d'_k(v_m^o,Q'(Z^o))-d_k(v_m^o,Q(Z^o))) E[Y_{k}|V=v_m^o] \] is i.i.d.\ across $m = 1, \dots, M$, with bounded mean and variance. This follows because the treatment assignment indicators $d_k(v_m^o, Q(Z^o))$ and $d'_k(v_m^o, Q'(Z^o))$ are binary and mutually exclusive across $k = 1, \dots, T$, and the marginal treatment responses $E[Y_k|V=v]$ are bounded. Hence, their difference is a bounded weighted sum and has finite conditional expectation and variance.
Now, consider the observed GPRTE in Equation (ref). To proceed, we first verify that the conditional expectation of $\Delta W_{N_o}$ equals the true quantity $\Delta W$:
Given the unbiasedness of the observed GPRTE estimator, we can analyze the asymptotic distribution of the test statistic, and establish the hypothesis testing framework when \( M \to \infty \) and \( N_o \) is fixed. For a given observation \( o \), by Central Limit Theorem, the sample mean \( \Delta \mu^o_M \) satisfies: \[ \Delta \mu^o_M \sim N\left(\mu'^o_{E[Y|V]} - \mu^o_{E[Y|V]}, \frac{\sigma_o^2}{M}\right), \] where \( \sigma_o^2 \) is the variance of the individual differences in marginal treatment responses. Since \( \Delta W_{N_o} \) is a weighted sum of \( \Delta \mu^o_M \), it follows that:
To estimate this variance in Equation (ref), we use the sample variance for each observation \( o \): \[ \widehat{\sigma}_o^2 = \frac{1}{M-1} \sum_{m=1}^M \left( \Delta Y_m^o - \Delta \mu^o_M \right)^2, \] where \( \Delta Y_m^o = E[Y_{\bar{D}'^o_m} | V = v_m^o, Z^o] - E[Y_{\bar{D}_m^o} | V = v_m^o, Z^o] \). The estimated variance of \( \Delta W_{N_o} \) is then: \[ \widehat{\mathrm{Var}}(\Delta W_{N_o}) = \sum_{o=1}^{N_o} \omega_o^2 \frac{\widehat{\sigma}_o^2}{M}. \]
Under the null hypothesis \( H_0: \Delta W = 0 \), the test statistic is constructed as: \[ Z_{\text{test}} = \frac{\Delta W_{N_o}}{\sqrt{\widehat{\mathrm{Var}}(\Delta W_{N_o})}}, \] When \( M \to \infty \), the numerator asymptotically follows a normal distribution, and the denominator converges to a constant. By Slutsky's theorem, the test statistic \( Z_{\text{test}} \) asymptotically follows a standard normal distribution: \[ Z_{\text{test}} \sim N(0, 1), \quad \text{under } H_0. \]
With critical value from standard normal distribution, this test evaluates the statistical significance of the GPRTE under the asymptotic framework where \( M \to \infty \) and \( N_o \) is fixed.
Consider the case when the marginal treatment response is not point identified. In this situation, GPRTE is represented as a set rather than a single point estimate. This necessitates a different approach for hypothesis testing.
Our null and alternative hypothesis:
Let \(\Delta \mathcal{W}\) denote the set of possible values for the GPRTE under the policy change from \((\mathbf{d},Q)\) to \((\mathbf{d}',Q')\), we know from Equation (ref),
The region \(\mathcal{I}_K^1\) for each marginal treatment response \(E[Y_k|V=v]\) is defined by a set of inequalities constraints, which determine a continuous set of possible values for \(E[Y_k|V=v]\). The GPRTE, \(\Delta W = \sum_{o=1}^{N_o} \omega_o \Delta \mu^o\), is a weighted sum of \(\Delta \mu^o\), where \(\Delta \mu^o \in \mathcal{I}_k^1\). Since weighted sums of continuous sets are also continuous, \(\Delta \mathcal{W}\) is a connected interval. Inequalities defining \(\mathcal{I}_k^1\) hold with equality at the boundaries, \(\Delta \mathcal{W}\) is a closed. As a consequence, $\Delta\mathcal W$ is a closed interval. I use the following algorithm to compute it.
For each observation \(o\), the lower bounds \(\underline{\delta^o_m}\), \(m = 1, \dots, M\), are i.i.d. because they are functions of independently drawn values \(v_1^o, \dots, v_M^o\). Similarly, the upper bounds \(\overline{\delta^o_m}\) are also i.i.d. across \(m\). However, within the same index \(m\), the pair \((\underline{\delta^o_m}, \overline{\delta^o_m})\) are not independent, as both depend on the same draw \(v_m^o\). Since the outcome variable is bounded, both lower and upper bounds are finite.
As a result, for a given \(o\), the quantities \(\underline{\Delta \mu^o_M}\) and \(\overline{\Delta \mu^o_M}\), which are empirical averages over \(M\) i.i.d. samples, satisfy the conditions of the Central Limit Theorem. Hence, \(\sqrt{M} \, \underline{\Delta \mu^o_M}\) and \(\sqrt{M} \, \overline{\Delta \mu^o_M}\), each converge in distribution to a normal distribution. Consequently, the bounds of GPRTE \(\sqrt{M} \, (\underline{\Delta W_{N_o}}, \overline{\Delta W_{N_o}})\), follows a bivariate normal distribution, which we denote as
To construct the variance-covariance matrix of the joint normal distribution for \( \sqrt{M} (\underline{\Delta W_{N_o}}, \overline{\Delta W_{N_o}}) \), we estimate \( \sigma_{\underline{\Delta W}}, \sigma_{\overline{\Delta W}}, \) and \( \rho_{\underline{\Delta W}, \overline{\Delta W}} \). These components are derived from the variances and covariances of the lower and upper bounds across the observations and weights.
The variance of the lower bound is defined as: \[ \sigma_{\underline{\Delta W}}^2 = \sum_{o=1}^{N_o} \omega_o^2 \frac{\sigma_{\underline{\Delta \mu^o}}^2}{M}, \] where \( \sigma_{\underline{\Delta \mu^o}}^2 \) is the variance of the sample lower bounds for observation \( o \). The unbiased estimator is: \[ \widehat{\sigma}_{\underline{\Delta \mu^o}}^2 = \frac{1}{M-1} \sum_{m=1}^M \left( \underline{\delta^o_m} - \underline{\Delta \mu^o_M} \right)^2, \] where \( \underline{\Delta \mu^o_M} = \frac{1}{M} \sum_{m=1}^M \underline{\delta^o_m} \) is the sample mean of the lower bounds for observation \( o \). Substituting this into the variance formula, the estimator for \( \sigma_{\underline{\Delta W}}^2 \) is: \[ \widehat{\sigma}_{\underline{\Delta W}}^2 = \sum_{o=1}^{N_o} \omega_o^2 \frac{\widehat{\sigma}_{\underline{\Delta \mu^o}}^2}{M}. \]
Similarly, the estimator for the variance of the upper bound $\sigma_{\overline{\Delta W}}^2$ is \[ \widehat{\sigma}_{\overline{\Delta W}}^2 = \sum_{o=1}^{N_o} \omega_o^2 \frac{\widehat{\sigma}_{\overline{\Delta \mu^o}}^2}{M}. \] where \[ \widehat{\sigma}_{\overline{\Delta \mu^o}}^2 = \frac{1}{M-1} \sum_{m=1}^M \left( \overline{\delta^o_m} - \overline{\Delta \mu^o_M} \right)^2, \]
Denote the difference between the upper and lower bounds as \(R_{N_o}\): \[ R_{N_o} = \overline{\Delta W_{N_o}} - \underline{\Delta W_{N_o}}. \] We construct a confidence interval as
where \( \overline{C}_M \) satisfies
and $\alpha$ is the chosen significance level. We state the following theorem:
Therefore, with the given significance level $\alpha$, we can perform the hypothesis test. If $0\notin C^{\Delta W}_{1-\alpha}$, we can reject the null hypothesis that GRPTE equals zero.
This paper develops a hyper-rectangle model to address the challenges of analyzing treatment effects in micro-econometric studies with set-identified parameters. By introducing a framework that leverages the interplay among observed outcomes, treatments, covariates, unobserved heterogeneity, and instrumental variables, the proposed model enhances the identification and estimation of treatment effects under realistic and flexible assumptions.
A key contribution of this study is the identification of the Marginal Treatment Response function and the threshold function $Q(Z)$, which are fundamental for understanding heterogeneous treatment effects. By distinguishing between cases where the threshold function or the distribution of unobserved heterogeneity is known, the model provides a structured approach to derive bounds for treatment effects. This approach offers practical insights into policy-relevant questions by accommodating partial identification and allowing for robust inference.
The model's ability to handle complex treatment assignment mechanisms expands its applicability to a wide range of empirical contexts. It facilitates the estimation of various treatment effect measures, including Average Treatment Effects and Policy Relevant Treatment Effects, while remaining computationally tractable. Moreover, the framework provides tools for hypothesis testing, enabling researchers to draw meaningful conclusions even under partial identification.
Overall, the proposed methodology contributes to the econometric literature by offering a flexible and empirically applicable tool for treatment effect analysis. Future research could extend this framework by exploring dynamic treatment settings, incorporating additional sources of uncertainty, or applying the methodology to large-scale datasets in practice.