Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
123,765 characters · 27 sections · 110 citation commands
Treatment Effects with Targeting Instruments
\onehalfspacing
\thispagestyle{empty}
\setcounter{page}{1}
Much of the literature on the evaluation of treatment effects has concentrated on the paradigmatic “binary/binary” example, in which both treatment and instrument only take two values. Multivalued treatments are common in actual policy implementations, however, as are multivalued instruments. Many different programs aim to help train job seekers for instance, and each of them has its own eligibility rules. Tax and benefit regimes distinguish many categories of taxpayers and eligible recipients. The choice of a college and major has many dimensions too, and responds to a variety of financial help programs and other incentives. Finally, more and more randomized experiments in economics resort to factorial designs\footnote{Factorial:Designs:REStat review recent applications of factorial designs.}.
Existing work on multivalued treatments under selection on observables includes imbens2000role, cattaneo2010efficient, and Ao:et:al among others. As the training, education choice, and tax-benefit examples illustrate, in non-experimental settings multivalued treatments are also subject to selection on unobservables. The use of instruments to evaluate the effects of multivalued treatments under selection on unobservables has received increasing attention in recent literature. In previous work \citep*{LeeSalanie2018}, we analyzed the case when enough continuous instruments are available. Identification is of course more difficult when instruments only take discrete values. We explore in this paper the use of such discrete-valued instruments in order to control for selection bias when evaluating discrete-valued treatments. Our goal is to find plausible conditions on treatment assignment and on the distribution of outcomes under which counterfactual averages and treatment effects are point- or partially identified for various (sometimes composite) complier groups. This distinguishes our paper from the recent contributions of baietal:monotonicityaverage, which focuses on population-wide average outcomes, and of goff:oa, which studies identification without any assumption on outcomes.
In the binary/binary model, the analyst can often take for granted that switching on the binary instrument makes treatment (weakly) more likely for all or no observations. This is satisfied under the local average treatment effect (LATE)-monotonicity assumption LATE1994,vytlacil2002independence,HV2007-handbook. With multiple instrument values and multiple treatments, there may be no natural ordering of instrument or treatment values that would give meaning to the word “monotonicity”. heckmanpinto-pdt defined an “unordered monotonicity” property; various papers have proposed other definitions of (qualified) monotonicity\footnote{See nafjeevanpinto:2022 for a detailed analysis of some of these proposals.}. Even when such a condition holds, there exist several groups of compliers---individuals whose treatment assignment changes with the value of the instrument. Since this may give rise to a multiplicity of cases, existing literature has often added assumptions that reduce this complexity. angrist1995two analyzed two-stage least squares (TSLS) estimation when the treatment takes a finite number of ordered values. Several recent papers have studied the case of binary treatments with multiple instruments. MTW:2020 and Goff analyzed the identifying power of different monotonicity assumptions in this context\footnote{MTW:2024 further apply their framework of monotonicity with multiple instruments to marginal treatment effects heckman2001policy,MIV2005,CHV-AER.}. Others have studied models with binary instruments and multivalued or continuous treatments. Torgovitsky2015, DF2015, Huang:et:al:2009, caetano_escanciano_2020, Feng:2024, and AngristSantosTecchio2025 developed identification results for different models.
In a wide-ranging contribution, heckmanpinto-pdt derived results on partial identification in discrete-instrument, discrete-treatment models; they also showed how additional identifying assumptions, such as unordered monotonicity, can be applied to shrink the identified set of treatment effects for various complier groups. While their results are very general, they are not as transparent as one would like. Our approach to this issue is different: we seek a parsimonious framework within which we can make constructive progress, and that can still be useful in many applications. In order to reduce the complexity of the problem, we start by imposing an additive random-utility model (ARUM) structure, as did HUV2006,HUV2008 and HV2007-handbook-2. Under ARUM, the selection into treatment depends on mean values and additive, observation-specific shocks. Some, but not all, ARUM models satisfy the unordered monotonicity property of heckmanpinto-pdt, which was applied by pinto2021 to the Moving to Opportunity program.
In many applications, a value of the treatment is especially salient; since it often is the no-treatment value $t=0$, we call it the “control”. Under ARUM, each treatment $t$ generates a change in the mean value, relative to the control, that depends on the value $z$ of the instrument. It is natural to speak of an instrument value $z$ {\em targeting\/} a treatment value $t$ when it maximizes this change in mean value. Most of our paper relies on the assumption of {\em strict targeting\/}, which obtains when each instrument only changes the mean values of the treatments it targets. Strict targeting holds for instance in models of imperfect compliance when the cost of non-compliance does not depend on its nature. Some of our results also require {\em one-to-one targeting}, where each non-zero instrument targets one treatment only, and each treatment (apart from the control) is targeted by one instrument only. Finally, we speak of {\em universal targeting\/} when both strict and one-to-one targeting apply.
As we will see, each of these targeting assumptions generates testable implications. These tests can also be used to match instrument values and the treatment values that they target. They only rely on estimates of the generalized propensity scores, which are directly identifiable from the data. In particular, the testable implications of universal targeting that we derive are identical to those of Bai:Tabord-Meehan:2025:arXiv and are therefore sharp.
Our use of “targeting” instruments is similar in spirit to Section 7.3 of HV2007-handbook-2\footnote{See also the recent contribution by Buchinsky:Gertler:Pinto:2023, which uses revealed preference arguments.}. We define it differently and we seek to identify a more general class of treatment effects. The term “targeting” is inspired by the time-honored Targeting Principle\footnote{Early references include tinbergen:econpolicybk and bhagwati:1971.}. Some policies act directly on final outcomes, and others aim to modify choices. Our use of the term “targeting” refers to the latter. Take a Roy model in which workers choose among occupations on the basis of their net utilities; we observe the choice of occupation and the wage in that occupation. A safety regulation that reduces the disutility of labor for (say) construction workers is, in our terminology, an instrument that targets the choice to be a construction worker. Policymakers might also seek to increase average incomes by offering a college credit. While their final aim is to increase wages (an outcome), we would say that the college credit is an instrument that targets the choice to go to college---a treatment variable.
To illustrate the usefulness of our framework, we apply it to the Head Start Impact Study (HSIS), a randomized experiment that sought to evaluate the value added of Head Start preschools. kline2016 revisited the HSIS; they took into account the presence of a substitute treatment (alternative preschools in this case). They found that Head Start was only beneficial for children who would not have attended another preschool program instead. In this study, the instrument is binary: a child is offered admission to Head Start or not. Treatment is ternary, as the child may end up in Head Start, in an alternative preschool, or not be enrolled in preschool. In our language, “no preschool” is the reference treatment. Head Start offers of course target Head Start; and since the instrument is binary, targeting is trivially strict. This makes it an example of universal targeting.
As another example, consider a randomized experiment with imperfect compliance. Individuals are assigned to a treatment arm via instrument (targeting), but may self-select into an alternative treatment (imperfect compliance). When random assignments map uniquely to treatments, the design satisfies our one-to-one targeting assumption. Strict targeting further implies that selecting into treatment is equally costly for all individuals not assigned to it. Consider for instance the interventions reported in Angrist:et:al:2009 and in attanasioetal:colombia:2014 (attanasioetal:colombia:2014,attanasioetal:colombia:2020). These are 4-way factorial randomized experiments: each subject is randomly assigned to a control group, to receive treatment 1, to receive treatment 2, or to receive both treatments. By definition, this is one-to-one targeting. Compliance was very imperfect in Angrist:et:al:2009, and it is described as “high” in the other two papers. If subjects self-selected into treatments on the basis of their expected benefits, then strict targeting is a natural assumption.
Combining ARUM and assumptions on targeting allows us to point-identify the size of some complier groups and the corresponding counterfactual averages and treatment effects on the outcomes, and to partially identify others. We use two examples to demonstrate the identification power and implications of ARUM and targeting. Our first example is a $2 \times T$ model where a binary instrument targets only one of $T \geq 3$ treatment values, as in kline2016. In our second example, three unordered treatment values target three instrument values. This $3 \times 3$ model was also studied by kirkeboen2016\footnote{See also more recent work by Bhuller22, Heinesen:2025, and nos22.}. Unlike them, we do not assume that the data contains information on next-best alternatives. Whereas the $2\times T$ model satisfies unordered monotonicity under our strongest targeting assumptions, the $3\times 3$ model does not\footnote{ It does satisfy the weaker generalized monotonicity assumption of baietal:monotonicityaverage, however.}.
We obtain novel identification results for both examples; they lead to new estimands or bounds for average treatment effects on various groups. Additional identifying assumptions can refine these bounds. Our leading example is what we call {\em positive selection}. This assumes that the average outcome for a given treatment $t$ is larger for some response group than for another. Consider for instance the binary instrument case. It seems natural to assume that the always-takers of a treatment get more from it than compliers who only take it if they are incentivized to do so. Positive selection also obtains under weak assumptions in the generalized Roy model. More generally, let us return to our earlier illustration of a randomized experiment under imperfect compliance. Consider the response group of individuals who would end up in treatment $t^\prime$ both when drawing $z$ and when drawing $z^\prime$. We would expect this response group to have better outcomes under $t^\prime$, on average, than the response group that exhibits perfect compliance to $z$ and $z^\prime$ draws---assuming that these two response groups end up in the same treatment branches for all other instrument values. This falls exactly under our positive selection assumption. It adds identifying power in both of our leading examples.
To show the value of our approach, we apply it to the reanalysis of Head Start by kline2016. We confirm the importance of taking into consideration alternative preschools when evaluating Head Start. Unlike kline2016, we do not rely on parametric selection models. Under a plausible positive selection assumption, our estimates suggest that the large difference between complier groups that they find can only be rationalized under {\em negative\/} selection into Head Start. As a by-product, we provide an upper bound on the welfare effect of expanding access to Head Start. Interestingly, the estimated upper bound turns out to be lower than the point estimate of kline2016; and it yields a lower marginal value for public funds used in expanding access to Head Start.
The paper is organized as follows. (ref) defines our framework. In (ref), we define and discuss the concepts of targeting, one-to-one targeting, and strict targeting. Section (ref) derives their implications for the identification of population shares, counterfactual averages, and the effects of the treatments on various complier groups; it also defines and illustrates positive selection. Finally, we present estimation results for Head Start in (ref). The Appendices contain the proofs of all propositions and lemmata, along with some additional material.
In all of the paper, we denote observations as $i=1,\ldots,n$. Each observation consists of covariates $X_i$, instruments $Z_i$, outcome variables $Y_i$, and treatments $T_i$. We assume that the covariates $X_i$ are exogenous to treatment assignment and outcomes. Since they will not play any role in our identification strategy, we condition on the covariates throughout and we omit them from the notation. Our results should therefore be interpreted as conditional on $X$.
We assume that observations are independent and identically distributed. Random sampling rules out that the treatment status of one observation influences other observations. This further implies that the outcome for a specific observation does not impact the outcomes of other members within the population. In other words, we rely on the Stable Unit Treatment Value Assumption (SUTVA).
We focus in this paper on treatment variables that take discrete values, which we label $t \in \mathcal{T}$. For simplicity, we will call $T=t$ “treatment $t$”. These values do not have to be ordered; e.g., when $t=2$ is available, it does not necessarily indicate “more treatment” than $t=1$. We assume that the only available instruments are discrete-valued, and we label their values as $z \in \mathcal{Z}$. Finally, for any $z\in\mathcal{Z}$ and $t\in \mathcal{T}$, we denote the {\em generalized propensity score\/} by \[ P(t|z) \equiv \Pr(T_i=t\vert Z_i=z). \]
As in most of this literature, we will need an assumption that restricts the heterogeneity in the counterfactual mappings $T_i(z)$, that is, potential treatments. In the binary/binary model, this is most often done by imposing LATE-monotonicity. As is well-known, LATE-monotonicity imposes that (denoting instrument values as $z=0,1$) (i) or (ii) must hold:
With more than two treatment values and/or more than two instrument values, there are many ways to restrict the heterogeneity in treatment assignment. Since treatments may not be ordered in any meaningful way, we cannot apply the results in angrist1995two for instance. MTW:2020 state several versions of monotonicity for a binary treatment model with $\lvert\mathcal{Z}\rvert>2$. They propose a “partial monotonicity” assumption which applies binary LATE-monotonicity component by component. This requires that the instruments be interpretable as combinations of component instruments, which is not necessarily the case here.
To cut through this complexity, we assume from now on that assignment to treatment can be represented by an Additive Random-Utility Model (ARUM), that is by a discrete choice problem with additively separable errors: \[ T_i(z) = \arg\max_{t \in \mathcal{T}} (U_{z}(t) + u_{it}) \] for some real numbers $U_z(t)$ which are common across observations, and random vectors ${(u_{it})}_{t\in \mathcal{T}}$ that are distributed independently of $Z_i$. We do not restrict the codependence of the random variables $u_{it}$.
In a randomized experiment with perfect compliance, we would have $U_z(t^\prime)=-\infty$ and $U_{z^\prime}(t)=-\infty$. With imperfect compliance, these mean values are finite; if for instance $u_{it^\prime}-u_{it}>U_z(t)-U_z(t^\prime)$, individual $i$ will get into treatment $t^\prime$ when drawing $z$ would normally assign her to $t$.
Imposing an ARUM structure will greatly simplify our discussion of treatment assignment. It incorporates a substantial restriction, however. Suppose that observation $i$ has treatment values $t$ under $z$ and $t^\prime$ under $z^\prime$. By the ARUM structure, this implies
Combining these two restrictions implies an “increasing differences” property: \[ U_{z^\prime}(t^\prime) - U_{z^\prime}(t) \geq U_z(t^\prime)-U_z(t). \] This inequality in turn is incompatible with the existence of an observation $j$ that has treatment values $t^\prime$ under $z$ and $t$ under $z^\prime$. Thus we rule out “direct two-way flows”: if a change in the value of an instrument causes an observation to shift from a treatment value $t$ to a treatment value $t^\prime$, it can cause no other observation to switch from $t^\prime$ to $t$. The argument above is a special case of the general discussion in heckmanpinto-pdt; their Theorem {T-3} shows that the treatment assignment models that satisfy unordered monotonicity for each pair of instrument values can be represented as an ARUM. Not all ARUM models satisfy unordered monotonicity, however; unordered monotonicity excludes a more general class of two-way flows. We will illustrate this point on one of our leading examples in (ref).
Assumption (ref) defines the class of models of assignment to treatment that we analyze in this paper.
We will often refer to the $U_z(t)$ as “mean values”. This is only meant to simplify the exposition; it is consistent with, but need not refer to, preferences on the part of the agent. Note that when $\mathcal{T}=\{0,1\}$, Assumption (ref) is just the standard monotonicity assumption, with a threshold-crossing rule \[ T_i(z)={\rm 1\kern-.40em 1}(u_{i0}-u_{i1}\leq U_z(1)-U_z(0)). \] If we add a third treatment value so that $\mathcal{T}=\{0,1,2\}$, the ARUM assumption starts to bite as it excludes direct two-way flows in the treatment model. However, Assumption (ref) is far from sufficient to identify interesting treatment effects in general. In order to better understand what is needed, we now resort to the notion of {\it response-groups\/} of observations, whose members share the same mapping from instruments $z$ to treatments $t$. We first state a general definition\footnote{This is analogous to the definitions in heckmanpinto-pdt.}.
To illustrate, consider the binary instrument/binary treatment case. It has a priori $2^2=4$ response vectors, $R \in \{00, 01, 10, 11\}$ with corresponding response-groups $C_{00}, C_{01}, C_{10}, C_{11}$. In this notation, the first number refers to a treatment value with $z=0$ and the second number with $z=1$. For instance, $C_{01}$ refers to those with $T_i(0) = 0$ and $T_i(1) = 1$, while the composite response-group $C_{\ast 1}$, for which $R(0)=\{0,1\}$, represents the union of $C_{01}$ and $C_{11}$. The LATE-monotonicity assumption implies that either $C_{01}$ or $C_{10}$ is empty.
We start by introducing additional assumptions on the underlying treatment model. We will illustrate these assumptions on three examples: the “binary instrument model” or the “$2\times T$” model; the “$3\times 3$ model”; and a generalized Roy model. We first define them briefly.
When $\lvert\mathcal{T}\rvert=3$, treatment assignment can be represented in the $(u_{i1}-u_{i0},u_{i2}-u_{i0})$ plane. The points of coordinates $P_z=(U_z(0)-U_z(1),U_z(0)-U_z(2))$ play an important role as for a given $z$,
Treatment assignment is illustrated in Figure (ref) for a given $z$, where the origin is in $P_z$. We will make recurrent use of this type of figure.
Finally, our framework also includes multivalued generalized Roy models (see EisenhauerHV:2015).
“Targeting” will be the common thread in our analysis. Just as in general economic discussions a policy measure may target a particular outcome, we will speak of instruments (in the econometric sense) targeting the assignment to a particular treatment.
Under (ref), assignment to treatment is governed by the differences in mean values $(U_z(t)-U_z(\tau))$ and by the differences in unobservables $u_{it}-u_{i\tau}$. Only the former depend on the instrument. From now on, we assume that there is a reference treatment value $t_0$ whose mean utility does not depend on the value of the instrument:
In many applications, $t=0$ is a “no-treatment” value, and instruments only change the mean utilities of the other treatments. For instance, tuition subsidies, investment credits, and invitations to training programs have no effect for those who do not attend college, do not invest, or choose not to train. (ref) seems natural in such cases\footnote{In the generalized Roy model (Example (ref)), it holds if the disutility of occupation 0 does not depend on the values of the instruments.}. For a counter-example, consider a program of unconditional cash transfers with different values $z$, for which we observe the purchases $t$ of several categories of goods the following month. If a household decides to save the transfer ($t=0$), its mean (discounted) utility will still depend on the value of the transfer $z$ that it received\footnote{In that case one could define targeting with the function $\tilde{U}_z(t)= U_z(t)-U_z(0)$. We have not explored the consequences of this alternative definition.}.
Given Assumption (ref), we will say that an instrument value $z$ {\em targets\/} a treatment value $t$ if it maximizes the mean utility $U_z(t) - U_z(0)=U_z(t)$.
Definition (ref) calls for several remarks. First, Assumption (ref) implies that $Z^\ast(0)=\mathcal{Z}$. Therefore $t=0$ is not in $\mathcal{T}^\ast$; the set $\mathcal{T}^\ast$ may exclude other treatment values, however. If a treatment value $t$ is not targeted ($t\not\in \mathcal{T}^\ast$), by definition the function $z\to U_z(t)$ is constant over $z\in\mathcal{Z}$, with value $\bar{U}(t)$. If an instrument value $z$ does not target any treatment ($z\not\in \mathcal{Z}^\ast$), then $U_z(t)<\bar{U}(t)$ for every $t\in\mathcal{T}^\ast$. While non-targeted treatment values ($t\in \mathcal{T}\, \setminus\, \mathcal{T}^\ast$) have mean values that do not respond to changes in the instruments, these mean values may and in general will differ across treatments. The probability that an individual observation takes a treatment $t \in \mathcal{T}\, \setminus\, \mathcal{T}^\ast$ also generally depends on the value of the instrument.
It is important to note here that the values $U_z(t)$ and therefore the targeting maps $Z^\ast$ and $T^\ast$ are not observable; any assumption on targeting instruments and targeted treatments must be a priori and context-dependent. These prior assumptions have consequences that can be tested, however. As a simple example, $z=1$ targets $t=1$ in the $2\times 2$ model under monotonicity; this implies that $P(1\vert 1)\geq P(1\vert 0)$. We will show that targeting always yields linear inequalities on generalized propensity scores in the class of models defined by Assumptions (ref) and (ref).
By definition, the treatment $T_i(z)$ assigned to $i$ under $z$ can only be the top targeted treatment $t^\ast_i(z)$ (if $V^\ast_i(z)>\underline{V}_i(z)$) or the top alternative treatment $\underline{t}_i(z)$ (if $V^\ast_i(z)<\underline{V}_i(z)$). If $z$ is not a targeting instrument, then $T_i(z)$ must be $\underline{t}_i(z)$.
Definition (ref) could be extended to broader settings than our ARUM model. On the other hand, the ARUM structure interacts with our definition of targeting to constrain treatment assignment in interesting ways. Suppose that an instrument value $z$ targets a treatment value $t$: $U_z(t)=\bar{U}(t)$. If some observation $i$ is assigned a treatment value $T_i(z)=t^\prime\neq t$, we must have \[ \bar{U}(t)+u_{it} < U_z(t^\prime)+u_{it^\prime}. \] Now suppose that an instrument value $z^\prime$ targets this $t^\prime$; then $U_z(t^\prime)\leq U_{z^\prime}(t^\prime)=\bar{U}(t^\prime)$, so that
and $T_i(z^\prime)\neq t$. If $t^\prime=0$, we can use Assumption (ref) to get a stronger implication: since $U_z(0)=U_{z^{\prime\prime}}(0)=0$ for all $z^{\prime\prime}$, (ref) holds for all $z^{\prime\prime}$ and $T_i(z^{\prime\prime})\neq t$. We summarize this in (ref).
(ref) (i) implies that the events $T_i(z)=t^\prime$ and $T_i(z^\prime)=t$ are mutually exclusive. Hence, the sum of their probabilities cannot exceed one:
Similarly, (ref) (ii) implies that the event $T_i(z)=0$ is disjoint from the event that $T_i(z'')=t$ for any $z'' \in \mathcal{Z}$. Since the probability of this latter event is at least $\max_{z''} P(t \mid z'')$, we obtain the following implication:
Now suppose that each $z$ consists of a set of (possibly zero or negative) subsidies $S_z(t)$ for treatments $t\in\mathcal{T}$. If there is a no-subsidy regime $z=0$ with $S_0(t)=0$ for all $t$, it seems natural to write the mean value as $U_z(t)=U_0(t)+S_z(t)$. Then for any treatment $t$, the set $Z^\ast(t)$ consists of the instrument values $z$ that subsidize $t$ most heavily. As this illustration suggests, the sets $Z^\ast(t)$ may not be singletons, and they may well intersect. We now introduce a more restrictive definition that rules out these two possibilities.
Under one-to-one targeting, we will often write “$z=t$” if $z$ targets $t$; this is without loss of generality. Let us illustrate these varieties of targeting on Example (ref).
{\bf Example (ref) continued.} Table (ref) shows the values of $U_z(t)$ in the $3\times 3$ model of Example (ref). Suppose that $t=1$ is targeted; choose some $z$ that targets it and relabel it as $z=1$. This means that \[ b\geq \max(a,c) \; \mbox{ and } \; b>\min(a,c). \] If $t=2$ is also targeted by some $z\neq 1$, we relabel this instrument value as $z=2$. This gives \[ f\geq \max(d,e) \; \mbox{ and } \; f>\min(d,e). \] Finally, if targeting is one-to-one we have $b>\max(a,c)$ and $f>\max(d,e)$.
One-to-one targeting implies some useful restrictions on response-groups. Take an observation $i$ and $t\in\mathcal{T}^\ast$. Under one-to-one targeting, treatment $t$ is targeted only by $z=t$. As a result, by (ref), if $T_i(t)=t^\prime$ and $t^\prime\in \mathcal{T}^\ast\setminus \{t\}$ then $T_i(t^\prime)\neq t$; and if $T_i(t)=0$ then $T_i(z)$ can never equal $t$. These restrictions are summarized in the following proposition.
{\bf Example (ref) (continued)} Return to the $3\times 3$ model and to Table (ref). Suppose that both $t=1$ and $t=2$ are targeted. Under the conditions of Proposition (ref), we have $b>\max(a,c)$ and $f>\max(d,e)$.
Since the points $P_z$ have coordinates $(-U_z(1),-U_z(2))$,
This is easily rephrased in terms of the response-vectors of definition (ref). First note that in the $3\times 3$ case, there are a priori $3^3=27$ response-vectors, $R=000$ to $R=222$, with corresponding response-groups $C_{000}$ to $C_{222}$. Groups $C_{ddd}$ are “always-takers”\footnote{Observations in group $C_{000}$ are usually called the “never-takers”. We prefer not to break the symmetry in our notation. We hope this will not cause confusion.} of treatment value $d$. All other groups are “compliers” of some kind, in that their treatment changes under some changes in the instrument. We will also pay special attention to some non-elemental groups. For instance, $C_{0\ast 2}$ will denote the group who is assigned treatment 0 under $z=0$ and treatment 2 under $z=2$, and any treatment under $z=1$. That is, \[ C_{0\ast 2} = C_{002} \operatorname*{\mathsmaller{\bigcup}} C_{012} \operatorname*{\mathsmaller{\bigcup}} C_{022}. \] One-to-one targeting implies the emptiness of four composite groups out of the 27 possible. For any treatment value $\tau$, Proposition (ref)(ii) rules out group $C_{10\tau}$ since this group has $R(1)=0$ and $R(0)=1$. It rules out $C_{\tau 01}$ as $R(1)=0$ and $R(2)=1$. This eliminates the composite groups $C_{10\ast}$ and $C_{\ast 01}$. The same argument applies to composite groups $C_{\ast 20}$ and $C_{2\ast 0}$, which have $R(2)=0$ and $R(1)=2$ or $R(2)=2$.
These four composite groups correspond to 10 elemental groups\footnote{Specifically, they are: $C_{100}, C_{101}, C_{102}, C_{001}, C_{201}, C_{020}, C_{120}, C_{220}, C_{200}$, and $C_{210}$.}. This still leaves us with 17 elemental groups, and potentially complex assignment patterns. Consider for instance Figure (ref). It shows one possible configuration for the $3\times 3$ model; the positions for $P_0, P_1$ and $P_2$ are consistent with one-to-one targeting.
The number of distinct response-groups (ten in this case) and the contorted shape of the $C_{212}$ and $C_{112}$ groups in Figure (ref) point to the difficulties we face in identifying response-groups without further assumptions. Moreover, this is only one possible configuration: other cases exist, which would bring up other response-groups.
heckmanpinto-pdt, pinto2021, and kirkeboen2016 also studied the $3\times 3$ model; they proposed sets of assumptions that identify some treatment effects. The example in heckmanpinto-pdt is rather specific. We show in (ref) how to apply our framework to the Moving to Opportunity experiment studied in pinto2021. The setup in kirkeboen2016 is most similar to ours; we will return to the differences between our approach and theirs in Section (ref). $\qed$
Figure (ref) suggests that if we could make sure that $P_1$ is directly to the left of $P_0$, the shape of $C_{212}$ would become nicer---and group $C_{202}$ would be empty. Bringing $P_2$ directly under $P_0$ would have a similar effect. This translates directly into assumptions on the dependence of the $U_z(t)$ on the instruments: the first one imposes $d=e$ and the second one imposes $a=c$. This can be interpreted as policy regime $z=1$ (resp.\ $z=2$) subsidizing treatment $t=1$ (resp.\ $z=2$) only. To return to the general model, there are applications in which the instruments $z\in Z^\ast(t)$, which maximize $U_z(t)$, do not shift assignment between the other values of the treatment. The following definition is a direct extension of this discussion.
Note that strict targeting only bites if $\mathcal{Z}$ contains at least three instrument values. If $\lvert\mathcal{Z}\rvert=2$ (one binary instrument, as in our Example (ref)) and say $z=1$ targets $t$, then $\mathcal{Z} \setminus Z^\ast(t)$ can only consist of $z=0$ and Assumption (ref) trivially holds.
Suppose for instance that the data comes from a randomized experiment, where the instrument value $z=t$ targets treatment $t$. If compliance is imperfect, an individual will trade off the benefits from switching to a treatment $t^\prime\neq t$ with the costs of the effort required. Strict targeting obtains when the cost of switching to $t^\prime$ does not depend on the value of $t$.
Under strict targeting, turning on instrument $z\in Z^\ast(t)$ promotes treatment $t$ without affecting the mean values $U_z(t^\prime)$ of other treatment values $t^\prime$. This explains our use of the term “strict targeting”. In this ARUM specification, an instrument in $Z^\ast(t)$ plays the same role as a price discount on good $t$ in a model of demand for goods whose mean values only depend on their own prices. In the language of program subsidies, all $z\in Z^\ast(t)$ subsidize $t$ at the same high rate, and all other instrument values offer the same, lower subsidy.
Finally, we should emphasize that one-to-one targeting and strict targeting are logically independent assumptions: neither one implies the other. Consider the $3\times 3$ model of Example (ref) under one-to-one targeting; strict targeting only holds for $t=1$ if $a=c$, and for $t=2$ if $d=e$. On the other hand, the $3\times 3$ model with $b>a=c$ and $e>d=f$ satisfies strict targeting but not one-to-one targeting, as $z=1$ targets both $t=1$ and $t=2$.
Now consider the general model. If a treatment $t$ is strictly targeted, then $U_z(t)$ can only take one of two values: $\bar{U}(t)$ if z targets $t$, and $\underline{U}(t)$ otherwise. By definition, if $t$ is not targeted then the value of $U_z(t)$ does not depend on $z$; we also denote it $\underline{U}(t)$. We will assume in this subsection that all targeted treatments are strictly targeted:
Under full strict targeting, the values of $U_z(t)$ are given in Table (ref). The functions $\underline{V}_i$ and $\underline{t}_i$ that define the top alternative treatment for a given observation are constant over $\mathcal{Z}\setminus\mathcal{Z}^\ast$. The assigned treatment $T_i(z)$ is $\underline{t}_i$ for all such values of $z$, as well as for $z\in\mathcal{Z}^\ast$ if \[ V^\ast_i(z)=\max_{t\in\mathcal{T}^\ast(z)} (\bar{U}_t+u_{it}) < \underline{V}_i=\max_{t\not\in\mathcal{T}^\ast(z)} (\underline{U}_t+u_{it}); \] it is the maximizer $t^\ast_i(z)$ of $V^\ast_i(z)$ otherwise.
Note that in a sense, all instrument values in $\mathcal{Z}\, \setminus\, \mathcal{Z}^\ast$ are equivalent under universal strict targeting. If $z$ and $z^\prime$ are two such values, then both functions $U_{z}$ and $U_{z^\prime}$ equal $\underline{U}$ on all of $\mathcal{T}$ and the counterfactual treatments $T_i(z)$ and $T_i(z^\prime)$ must be $\underline{t}_i$ for any observation $i$.
We now impose one-to-one targeting as well as strict targeting for every targeted treatment value:
Note that (ref) implies (ref).
Under (ref), (ref) becomes (ref) and we have:
Note that the $C(A,t)$ notation is just a shortcut: every $C(A,t)$ is an elemental group, and every elemental group is a $C(A,t)$. If for instance $\vert \mathcal{T} \vert=6$, it is just more convenient to write $C(\{1,3\}, 2)$ than to write $C_{212322}$.
If the set $A_i$ is non-empty and $\underline{t}_i\in A_i$, the observation $i$ is what one could call a {\em strict $A_i$-complier}: when the value of the instrument moves from $\mathcal{Z}\, \setminus \, A_i$ to a $t\in A_i$, observation $i$ switches from its top alternative treatment $\underline{t}_i$ to the top targeted treatment $t$. In the 3-by-3 model with $\mathcal{T}^\ast=\{1,2\}$, there are three groups of strict compliers: $C_{010}=C(\{1\}, 0)$, $C_{002}=C(\{2\},0)$, and $C_{012}=C(\{1, 2\},0)$.
Universal targeting brings us very close to the main identifying assumption in HV2007-handbook-2: the indicator variable ${\rm 1\kern-.40em 1}(Z=t)$ can be used as the $Z^{[t]}$ in their assumption. \citeauthor*{HV2007-handbook-2} use their Assumption B-2a to identify the effect of the preferred treatment $t$ relative to the next-best treatment. Their complier group consists of those individuals who choose treatment $t$ under $Z=z$ and another treatment under $Z=z^\prime$. This can be a very heterogeneous group, as our examples will show. To paraphrase HV2007-handbook-2: the mean effect of treatment $t$ versus the next best option is a weighted average over $t^\prime\in \mathcal{T}\setminus\{t\}$ of the effect of treatment $t$ versus treatment $t^\prime$, conditional on $t^\prime$ being the next best option, weighted by the probability that $t^\prime$ is the next best option. In contrast, we seek a complete characterization of all treatment effects that can be identified under this set of assumptions.
It is easy to show that universal targeting yields two sets of testable implications. First, consider any treatment value $t\in\mathcal{T}$. It is clear from the first row of Table (ref) that the mean utilities $U_z(t)$ are the same for all $z\in\mathcal{Z}\setminus \mathcal{Z}^\ast$. Since $0\not\in\mathcal{Z}^\ast$, it follows that
Next, let $t\in\mathcal{T}^\ast$ be a targeted treatment value and take an instrument $z^\prime$ that targets a different treatment $t^\prime \neq t$. Comparing the rows of Table (ref) reveals the ordering of utilities. When $z=t$, treatment $t$ is boosted; when $z=0$, it is at baseline; when $z=t^\prime$, the rival is boosted (drawing share away from $t$). This implies:
Combining these inequalities implies that for all targeted instrument values $t\in\mathcal{T}^\ast$, the choice probability function $z \mapsto P(t\mid z)$ is strictly maximized at $z=t$.
Moreover, the “encouragement design” assumption of a recent paper by Bai:Tabord-Meehan:2025:arXiv is exactly equivalent to universal targeting when (in their notation) $J_0>0$ and $z=0$ is what they call a “base state”. They show that encouragement design with a base state generates the following testable implications:
and that these implications are sharp: any set of propensity scores that satisfies these inequalities can be rationalized by an ARUM of encouragement design with a base state.
In our framework, combining (ref) for $z\in\mathcal{Z}\setminus\mathcal{Z}^\ast$ and the rightmost inequality in (ref) (strict inequality for rival-targeting $z$) yields (ref). Conversely, the fact that choice probabilities must sum to one ensures that if $P(t \mid z)$ decreases for all non-target treatments, $P(t \mid t)$ must increase, satisfying the leftmost inequality of (ref). Thus, under our assumptions, conditions (ref) and (ref) exhaust the testable implications of universal targeting.
We use the standard counterfactual notation: $T_i(z)$ and $Y_i(t,z)$ denote respectively potential treatments and outcomes. Let ${\rm 1\kern-.40em 1}(A)$ denote the indicator of set $A$. The validity of the instruments requires the usual exclusion restriction together with appropriate independence conditions.
One could impose a stronger condition than (ref)(ii), such as the joint independence of $\{ (Y_i(t), T_i(z)): (t,z) \in \mathcal{T}\times \mathcal{Z}\}$ from $Z_i$. However, we adopt the weaker mean-independence assumption for $Y_i(t)$ in (ref)(ii), as it is sufficient for the results presented in this paper. (ref)(ii) can be viewed as a generalization of Assumption 2 in huber2017sharp, who only consider the case of binary instruments and binary treatments. Under (ref), we define $T_i := T_i(Z_i)$ and $Y_i := Y_i(T_i)$. Throughout the paper, we assume that we observe $(Y_i, T_i, Z_i)$ for each $i$.
To simplify the exposition, we introduce one more element of notation. For any $z\in\mathcal{Z}$ and $t\in \mathcal{T}$, we define the {\em conditional average outcome\/} by \[ \bar{E}_z(t) \equiv \mathbb{E}(Y_i {\rm 1\kern-.40em 1}(T_i=t) \mid Z_i=z). \] For any response-group $C$ and treatment value $t\in\mathcal{T}$, we define the {\em group average outcome\/} as $\mathbb{E} (Y_i(t) \mid i \in C)$. Our goal is to identify the proportions of each response-group, $\Pr(i\in C)$, and their group average outcomes.
While the conditional average outcome $\bar{E}_z(t)$ is directly identified from the data, the group average outcomes are not; they combine with the group probabilities to form the conditional average outcomes. We will repeatedly use the following identity:
If there are $\lvert\mathcal{R}\rvert$ response-groups, the first equation in Lemma (ref) is a system of $(\lvert\mathcal{T}\rvert-1) \times \lvert\mathcal{Z}\rvert$ equations with $(\lvert\mathcal{R}\rvert-1)$ unknowns. Under monotonicity, the binary-binary model $\lvert\mathcal{T}\rvert =\lvert\mathcal{Z}\rvert= 2$ generates the standard LATE case, where the response-groups consist of never-takers ($C_{00}$), compliers ($C_{01}$), and always-takers ($C_{11}$). As is well known, the sizes of these three groups are just identified. On the other hand, even under (ref), one restriction is required to identify the sizes of the response-groups for the $3 \times 3$ model.\footnote{We prove this in (ref).} More generally, the degree of underidentification of group sizes tends to increase exponentially with the number of targeted treatments, like the number of response-groups.
Each of our assumptions on targeting reduces the number of response-groups and therefore the degree of underidentification. Take our strongest assumption: universal targeting. Then the set of response-groups $C=C_R$ such that $R(z)=t$ is as enumerated in Proposition (ref): it consists of
This gives directly the system of identifying equations.
We now introduce an identifying assumption that we call {\em positive selection}. It obtains when a function $h$ of the vector of potential outcomes $\bm{Y}_i\equiv (Y_i(0),\ldots,Y_i(\lvert\mathcal{T}\rvert-1))$ has a lower expectation for a response-group $C$ than for another response-group $C^\prime$:
Two leading examples are (a) $h(\bm{Y}_i)\equiv Y_i(t)$ for some $t$ and (b) $h(\bm{Y}_i)\equiv Y_i(t)-Y_i(t^\prime)$ for some $t\neq t^\prime$. In the $2 \times 2$ model, huber2017sharp consider a variant of form (a), which they call mean dominance. The identifying power of positive selection depends on the context. We will illustrate it using form (a)---that is, a restriction on the level of a potential outcome across groups---in (ref), as well as in our application to Head Start in Section (ref). We explore form (b)---a restriction on treatment effect differences across groups---in (ref).
Before turning to specific applications, let us discuss the characterization of the empirical content of our framework under positive selection. A series of papers\footnote{See balkepearl:97, kitagawa:15, mourifiewan:17, kedagnimourifie:20, and sun:23.} has provided necessary and, in some cases, sufficient conditions for data to be rationalized under an instrument exclusion restriction. Most recently, Bai:Tabord-Meehan:2025:arXiv characterized the sharp testable implications of what they call “encouragement design” under joint independence, that is, when $Z_i$ is independent of $\{(Y_i(t), T_i(z)) : (t,z) \in \mathcal{T} \times \mathcal{Z}\}$. Because we assume only mean independence in (ref) and impose positive selection via (ref), their testable implications involving outcomes $Y_i$ do not directly apply to our framework. However, their sharp restrictions on the generalized propensity scores $P(t\vert z)$ do apply, and we discuss these in our leading examples below. It is worth noting that while they do not impose an ARUM structure, their “encouragement design” condition is satisfied in our context only under our strongest assumption: universal targeting.
In the following sections, we characterize the empirical content for the restrictions given in (ref), together with a positive selection assumption $\mathbb{E}(Y_i(t) \mid i \in C) \leq \mathbb{E}(Y_i(t) \mid i \in C^\prime)$ for given $C$ and $C^\prime$. The primitives of the model are the response-group probabilities $\Pr(i \in C)$ and the group average outcomes $\mathbb{E}( Y_i(t) \mid i \in C)$. The restrictions implied by (ref) generate bilinear equality constraints for $\mathbb{E}( Y_i(t) \mid i \in C)$ and $ \Pr(i \in C)$, while positive selection yields linear inequality constraints for $\mathbb{E}( Y_i(t) \mid i \in C)$. This makes it difficult to obtain simple yet general characterizations. Instead, we present explicit results for specific models in the next two subsections.
In particular, we study the identification region of local average treatment effects, \[ \mathbb{E}[Y_i(t')-Y_i(t)\mid i\in C], \] for specific pairs $(t,t')$ and subgroups $C$.
Recall that with a binary instrument, strict targeting is trivially satisfied.
Under one-to-one targeting, $z=1$ targets only one instrument value, which we call $t=1$; and targeting is universal. (ref) can be applied directly and the group probabilities are just identified in our (ref). Proposition (ref) gives $2(\lvert\mathcal{T}\rvert-1)$ independent equations: for $t\neq 1$, \[ P(t\vert 0)= \Pr(i\in C(\emptyset, t)) +\Pr(i\in C(\{1\}, t)) \; \mbox{ and } \; P(t\vert 1)= \Pr(i\in C(\emptyset, t)). \] Moreover, $C(\emptyset,t)=C_{tt}$ for $t\neq 1$ and $C(\{1\},t)=C_{t1}$ for all $t$.
Note that when $z$ changes from 0 to 1, the only observations that change treatment are in $C_{t1}$ for $t\neq 1$. Since the corresponding $C_{1t}$ group is empty, there are no “two-way flows” and this model satisfies the unordered monotonicity property of heckmanpinto-pdt. Proposition (ref) gives explicit formul\ae\ for the probabilities of all $(2\lvert\mathcal{T}\rvert-1)$ response groups.
Since $\Pr (C_{t1}) \geq 0$, the model has $(\lvertT\rvert-1)$ simple testable predictions: $P(t\vert 0) \geq P(t \vert 1) \mbox{ for } t\neq 1$. While all the response group probabilities are point-identified, only some group average outcomes are point identified without further restrictions, as shown by Proposition (ref).
(ref) shows that we only identify a known convex combination of the $(\lvert\mathcal{T}\rvert-1)$ LATEs\footnote{We use the term “LATEs” for the average treatment effects on the various complier groups. Throughout the remainder of the paper, we assume, as is standard, that probability differences appearing in the denominator of estimands are always nonzero.}. This formula is reminiscent of angrist1995two, which deals with a different model in which treatments are ordered. It is possible to re-derive our identification results in Propositions (ref) and (ref) using the general framework of heckmanpinto-pdt. We provide details in (ref).
So far, we only imposed restrictions on the process by which treatment values are assigned to observations; this is what goff:oa calls an “outcome-agnostic” approach in that it only assumes that the instruments are excluded from the outcome equations. It is possible to bound the average treatment effects in a straightforward manner if we assume that the support of the outcomes is known and bounded. One could instead add restrictions to achieve point identification of average treatment effects for the compliers. Assuming that the ATEs are all equal is one obvious solution. Another one is to assume some degree of homogeneity of group average outcomes. Alternatively, we may consider weaker conditions under which the average treatment effects for the compliers are only partially identified. We explore these ideas below.
Consider the binary instrument model with $T\geq 3$.
\paragraph{Beyond One-to-one Targeting} First note that the probabilities of the response-groups can be identified under weaker restrictions than one-to-one targeting. Suppose for instance that $z=1$ targets all treatment values $t\geq 1$: we have $U_1(t)>U_0(t)$ for all $t\geq 1$. Then the complier groups $C_{t0}$ for $t\geq 1$ must be empty. To see this, suppose that $T_i(0)=t\geq 1$. This implies $U_0(t)+u_{it}>U_0(0)+u_{i0}=u_{i0}$. Adding up these inequalities gives $U_1(t)+u_{it}>u_{i0}$, and $T_i(1)$ cannot be $0$.
All other groups $C_{tt^\prime}$ may exist. This leaves $\vert \mathcal{T} \vert(\vert \mathcal{T} \vert-1)$ unknown group probabilities, which is $\vert \mathcal{T} \vert/2$ times more than the $2(\vert \mathcal{T} \vert-1)$ propensity scores we observe. We need $(\vert \mathcal{T} \vert-1)(\vert \mathcal{T} \vert-2)$ additional constraints to point-identify all group probabilities.
\paragraph{Single-peaked Mean Utilities} Now suppose that mean utilities are “single-peaked” in the sense that the function $t\to U_1(t)-U_0(t)$ is decreasing over $t=1,\ldots,T-1$. This would be a reasonable assumption if $z=1$ makes treatment $t=1$ more attractive and the treatments $t>1$ are ordered by their proximity to $t=1$.
If this holds, then the same argument as above shows that the response groups $C_{tt^\prime}$ must be empty when $t^\prime>t\geq 1$. This eliminates $(\vert \mathcal{T} \vert-1)(\vert \mathcal{T} \vert-2)/2$ response groups; we divided by two the number of additional identification constraints that we need.
\paragraph{Truncated Moment Bounds with Bounded Outcomes}
One alternative approach is to use truncated moment bounds, a method initiated by horowitz1995identification and popularized by lee2009training in the context of randomized experiments with sample attrition. See also huber2017sharp for extensions to the IV setting with binary instruments and binary treatments, and semenova:2025 for refinements of Lee bounds.
\paragraph{Monotone Treatment Responses.} It is sometimes natural to assume that treatments are ordered and that potential outcomes are weakly increasing in the treatment level, that is, $Y_i(t)\ge Y_i(t')$ whenever $t\ge t'$, as in manski1997. A weaker variant imposes monotonicity only in expectation within a response group $C$, $\mathbb{E}[Y_i(t)\mid i\in C]\ge \mathbb{E}[Y_i(t')\mid i\in C]$ for all $t\ge t'$, as in Marx:2024\footnote{See Section 4.2.2 of the job market paper version of Marx:2024, available at \url{https://www.dropbox.com/scl/fi/dq5hnkwwbsjpluuonpcwf/marx-jmp.pdf?rlkey=1zpwmnx5dgtmre5urp5hdzof5&e=4&dl=0}.} . These restrictions differ conceptually from our positive-selection assumptions. Monotone treatment response (MTR) constrains outcomes across treatment levels within a fixed group, whereas positive selection compares average outcomes across response groups at a given treatment level. Since our analysis focuses on unordered treatments, we focus here on selection-based restrictions.
\paragraph{Positive Selection}
The binary instrument model gives a first example of the power of the positive selection defined in Section (ref). Take $\tau\neq 1$ and consider the complier groups $C_{\tau 1}$: they all have $t=1$ when $z=1$, but they shift to it from different treatment values $\tau$ under $z=0$. Depending on the context, there may be a plausible reason to order the corresponding group average outcomes when $t=1$. Suppose for instance that $T=3$, and that
In this $2\times 3$ model under universal targeting, there are five response-groups, as illustrated in (ref). Proposition (ref) shows that the Wald estimator only identifies \[ \alpha_0 \mathbb{E} \left[ Y_i(1) -Y_i(0) \vert i \in C_{01} \right]+ (1-\alpha_0) \mathbb{E} \left[ Y_i(1) -Y_i(2) \vert i \in C_{21} \right], \] where $\alpha_0=(P(0 \vert 0)-P(0 \vert 1))/(P(1\vert 1)-P(1\vert 0))$ is point-identified. (ref) shows that adding inequality (ref) yields bounds on the corresponding LATEs.
The bounds on the local average treatment effects for $C_{01}$ and $C_{21}$ can be estimated directly from sample averages and are sharp in the sense that they incorporate all model restrictions. Specifically, the characterization relies only on two types of constraints: (i) the linear equality constraint in (ref), which corresponds to the special case of the second equality in (ref) from (ref), and (ii) the linear inequality constraint implied by positive selection in (ref). Accordingly, the bounds in (ref) constitute the explicit solutions to the associated linear programming problems. This explicit sharp characterization is possible because the response-group probabilities are point identified, as established in (ref).
Let us now turn to the $3\times 3$ model of Example (ref), where $\mathcal{Z}^\ast=\mathcal{T}^\ast = \{ 1, 2\}$ and $\mathcal{Z}=\mathcal{T} = \{0, 1, 2\}$. We assume universal targeting: for all of our results in this section, we impose Assumptions (ref), (ref), (ref), and (ref); $z=1$ targets $t=1$ and $z=2$ targets $t=2$.
The set $A$ in (ref) can be $\emptyset, \{1\}, \{2\}$, or $\{1,2\}$, with corresponding values of $t$ in $\{0\}, \{0,1\}, \{0,2\}$ or $\{0,1,2\}$ respectively. The set $c(\emptyset,0)$ corresponds to the never-takers $C_{000}$. For $A=\{1\}$ we get $C_{010}$ and $C_{111}$, and for $A=\{2\}$ we get $C_{002}$ and $C_{222}$. Finally, with $A=\{1,2\}$ we have $C_{012}, C_{112}$, and $C_{212}$.
These eight elemental response groups are illustrated in Figure (ref), again with the origin in $P_0$. Comparing Figure (ref) with Figure (ref) shows the identifying power of Assumption (ref). Table (ref) shows which groups take $T_i=t$ when $Z_i=z$.
Unlike the $2\times 3$ model, even under strict one-to-one targeting the $3\times 3$ model does not satisfy unordered monotonicity. One could show it with the matrix algebra in heckmanpinto-pdt.\footnote{See (ref) for details.} It is more straightforward to note that when the instrument value changes from $z=1$ to $z=2$, observations in $C_{010}$ move to treatment value 0, while observations in $C_{002}$ leave treatment 0. This is the definition of a two-way flow, which violates unordered monotonicity. Since the $3\times 3$ model has three instrument values and only two targeted treatments, baietal:monotonicityaverage shows that it satisfies their weaker general monotonicity assumption. As a consequence, the average potential outcomes $\mathbb{E} [Y_i(d)]$ can only be restricted by identification at infinity arguments.
We know that one restriction is missing to point-identify the probabilities of all eight response-groups. The following proposition shows that the probabilities of four of the eight elemental groups are point-identified: two groups of always-takers, and two groups of compliers. The other four probabilities are constrained by three adding-up constraints.
As before, the model has the following testable implications: $P(1 \vert 1) \geq P(1 \vert 0) \geq P(1 \vert 2)$, $P(2 \vert 2) \geq P(2 \vert 0) \geq P(2 \vert 1)$, and $P(0 \vert 0) \geq \max(P(0 \vert 1), P(0 \vert 2))$. It follows from Bai:Tabord-Meehan:2025:arXiv that these implications are sharp.
The following proposition identifies a number of group average outcomes\footnote{Again, these could also be derived using the general framework of heckmanpinto-pdt, even though the unordered monotonicity assumption is not satisfied---see (ref).}.
By itself, (ref) does not allow us to identify an average treatment effect for {\em any\/} (even composite) response-group. Suppose for instance that we want to identify $\mathbb{E} (Y_i(1)-Y_i(0)\vert i\in C)$ for some group $C$. Then $C$ needs to exclude $C_{111}$, $C_{112}$, and $C_{212}$, since $\mathbb{E} (Y_i(0)\vert i\in C^\prime)$ is not identified for any group $C^\prime$ that contains $C_{111}$, $C_{112}$, or $C_{212}$. Since we only know the mean outcome of treatment 1 for groups that contain one of these three elemental groups, the conclusion follows.
Note that if we assumed $\mathbb{E} (Y_i(1)\vert i\in C_{112})=\mathbb{E} (Y_i(1)\vert i\in C_{212})$, then we could combine the two equations in the fourth displayed line of (ref) and the probabilities of $C_{112}$ and $C_{212}$ (which are point-identified by (ref)) to obtain $\mathbb{E} (Y_i(1)\vert i \in C_{010}\operatorname*{\mathsmaller{\bigcup}} C_{012})$. This would point-identify the average effect of treatment 1 vs treatment 0 on this composite complier group $C_{01\ast}$. While this assumption may be overly strong, it seems natural to impose that $Y_i(\tau)$ is on average larger in a response group that has $t=\tau$ for more values of $z$. Assumption (ref) formalizes this intuition in our setting.
(ref) states a form of positive selection into treatment, as defined in Section (ref). Consider (ref) for instance. It says that within the group of “$12$-compliers” $C_{\ast 12}=C_{012}\operatorname*{\mathsmaller{\bigcup}} C_{112}\operatorname*{\mathsmaller{\bigcup}} C_{212}$, those observations with $T(0)=1$ have a larger average counterfactual $Y(1)$ than those with $T(0)=2$. (ref) shows that this gives bounds on the local average treatment effects for $C_{01\ast}$-compliers, with a similar result for (ref) and $C_{0\ast 2}$-compliers.
The lower bounds given in (ref) are sharp relative to the restrictions imposed by the mean-independence of the instruments and the positive selection assumptions. Just as with the bounds derived in (ref), this explicit sharp characterization obtains because the composite response-group probabilities
are point-identified, as established in (ref).
Let us focus on (ref). Given strict one-to-one targeting, $C_{112}$ is defined by \[ \underline{U}(1)-\bar{U}(2)\leq u_{i2}-u_{i1}\leq \underline{U}(1)-\underline{U}(2), \; u_{i1}-u_{i0}\geq -\underline{U}(1). \] $C_{212}$ is defined by \[ \underline{U}(1)-\underline{U}(2)\leq u_{i2}-u_{i1}\leq \bar{U}(1)-\underline{U}(2), \; u_{i2}-u_{i0}\geq -\underline{U}(2). \] To simplify notation, define $\zeta_i=u_{i2}-u_{i1}$ and $\xi_i=u_{i2}-u_{i0}$, so that $u_{i1}-u_{i0}=\xi_i-\zeta_i$. The inequalities above can be rewritten as
Figure (ref) plots these two groups on the $\zeta_i \times \xi_i$ plane. Group $C_{212}$ corresponds to the top-right (infinite) rectangle and group $C_{112}$ is partitioned into the two subgroups: $C_{112}^{(i)}$ is a bottom-left triangle and $C_{112}^{(ii)}$ is a top-left (infinite) rectangle.
Suppose for instance that \[ \mathbb{E}(Y_i(2)\mid u_{i0}, u_{i1},u_{i2})-\mathbb{E}(Y_i(2))=a_0 u_{i0}+a_1 u_{i1}+a_2 u_{i2}, \] where $a_0$, $a_1$, and $a_2$ are constants, and $(u_{i0},u_{i1},u_{i2})$ are jointly normal and mutually uncorrelated with mean 0 and variance 1. Defining $\zeta_i = u_{i2} - u_{i1}$ and $\xi_i = u_{i2} - u_{i0}$, it is easy to see that
If $a_2 \geq \max(a_0,a_1)+\vert a_1-a_0\vert$, then the coefficients of $\zeta_i$ and $\xi_i$ in (ref) are both non-negative. It follows\footnote{See Appendix (ref) for details.} from the geometry of Figure (ref) that $\mathbb{E}(Y_i(2)\mid i\in C_{212}) \geq \mathbb{E}(Y_i(2)\mid i\in C_{112})$. In summary, a sufficiently large value of $a_2$ induces positive selection, generating patterns similar to those of comparative advantage in generalized Roy models.
kirkeboen2016 used a $3\times 3$ model to study the impact of the field of study on later earnings. Their Proposition 2 characterizes what two-stage least squares (TSLS) estimators identify under different sets of assumptions. The least stringent version combines a monotonicity assumption (Assumption 4 in KLM) and condition (iii) in their Proposition 2, which they call “irrelevance and information on next-best alternatives”. “Irrelevance” is a set of exclusion restrictions, while “information on next-best alternatives” assumes the availability of additional data.
While we take quite a different path, our universal targeting assumption turns out to yield exactly the same identifying restrictions as the combination of monotonicity and irrelevance in KLM. We show it in (ref).
This set of assumptions in itself is too weak to give two-stage least squares estimates a simple interpretation. To see this, let $\beta_1$ and $\beta_2$ be the probability limits of the coefficients in a regression of $Y_i$ on the indicator variables ${\rm 1\kern-.40em 1}(T_i=1)$ and ${\rm 1\kern-.40em 1}(T_i=2)$, with instruments $Z_i$. Remember from Table (ref) that under strict one-to-one targeting, five response-groups have $T(1)=1$:
A similar distinction applies to the groups that have $T(2)=2$; it motivates Definition (ref).
The $\beta_1$ and $\beta_2$ coefficients turn out to be weighted averages of the LATEs on these two groups and on the intermediate groups $C_{112}$ and $C_{212}$.
(ref) implies that $\beta_1$ and $\beta_2$ are weighted averages of the four local average treatment effects on the right-hand side of this system of two equations. The weights are functions of the four probabilities on the left-hand side, which are point identified by (ref). However, these weights may be positive or negative. This complicates interpretation further\footnote{MTW:2020 give a set of assumptions under which the weights are positive in a model with multiple binary instruments.}.
\paragraph{Next-best alternatives} Using the additional information on next-best alternatives in KLM amounts, in our notation, to dropping the “intermediate” response-groups $C_{212}$ and $C_{112}$ from the data. Then the system of equations in (ref) becomes diagonal and it yields
where now $\mathcal{C}_1$ reduces to $C_{010}\operatorname*{\mathsmaller{\bigcup}} C_{012}$ and $\mathcal{C}_2$ reduces to $C_{002}\operatorname*{\mathsmaller{\bigcup}} C_{012}$. This is exactly Proposition 2 (iii) of KLM. Alternatively, one may simply assume that the response-groups $C_{212}$ and $C_{112}$ are empty. This is the path taken by Bhuller22\footnote{See their Corollary 5 and Table 1 for details.}.
\paragraph{Positive Selection} Additional information of the type used by KLM often is not available. Moreover, assuming away $C_{112}$ and $C_{212}$ seems rather strong. On the other hand, reasonable assumptions can be used to generate bounds on the local average treatment effects for 1-compliers and 2-compliers. (ref) illustrates this.
Note that the KLM result of the previous paragraph is the limit case where $\mathcal{D}_1=\mathcal{D}_2=0$.
The regularity condition (ref) ensures that the $2 \times 2$ matrix that premultiplies $(\beta_1, \beta_2)^\prime$ in (ref) is invertible\footnote{ It holds if $C_{212}$ and $C_{112}$ have positive probability and either $C_{010}\operatorname*{\mathsmaller{\bigcup}} C_{012}$ or $C_{002}\operatorname*{\mathsmaller{\bigcup}} C_{012}$ has positive probability.}. To interpret the assumptions on signs, suppose that $\mathcal{D}_1$ is positive. Since $\mathcal{C}_1=C_{010}\operatorname*{\mathsmaller{\bigcup}} C_{012}\operatorname*{\mathsmaller{\bigcup}} C_{212}$, the positivity of $\mathcal{D}_1$ states that the average effect of treatment 1 on $C_{010}\operatorname*{\mathsmaller{\bigcup}} C_{012}\operatorname*{\mathsmaller{\bigcup}} C_{212}$ is at least as large as on $C_{112}$. This is a form of positive selection that is in the same spirit as (but different from) (ref). If this form of positive selection holds for both treatments, then the TSLS estimates overestimate the LATEs on the corresponding compliers if $\mathcal{D}>0$, and they underestimate them if $\mathcal{D}<0$.
To summarize, the TSLS estimators in the $3 \times 3$ model are difficult to interpret unless additional information is available and/or some additional assumptions are imposed. If the groups $C_{112}$ and $C_{212}$ are indeed empty, then both the TSLS estimators and those we obtained in (ref) should identify the LATEs on the 1- and 2-compliers. Comparing their values is a useful (if informal) way of testing the assumptions and of exploring further the heterogeneity of the treatment effects.
We now reexamine the \citeapos{kline2016} analysis of the Head Start Impact Study (HSIS) using our framework. We use exactly the same data as they did; we only apply different identifying assumptions\footnote{Kamat:2019, KamatNorrisPecenco2023 and AngristSantosTecchio2025 analyze the HSIS using different identifying approaches.}.
Head Start is a federal program in the US that addresses various factors affecting children's development in low-income families. It provides early childhood education (hereafter “preschool”) and health and nutrition services. HSIS was a longitudinal study conducted from 2002 to 2010 to assess the program's impact on cognitive, social-emotional, and health outcomes. It focused on 84 communities where the demand for Head Start services was larger than the supply. HSIS randomly assigned about 5,000 three and four year old preschool children to either a treatment group which was offered Head Start services, or a control group which received no such offer. Children in either group could also attend other preschool centers if offered a slot.
The structure of HSIS is identical to that of (ref): it is a $2\times 3$ model. The treatments here consist of no preschool ($n$), Head Start ($h$), and other preschool centers ($c$): $\mathcal{T}=\{n, h, c\}$. The instrument is binary, with a control group ($z=0$) and a group that is offered admission to Head Start ($z=1$). The outcome variable is test scores, measured in standard deviations from their mean.
In our notation, $\mathcal{Z}=\{0,1\}$ and $\mathcal{Z}^\ast=\{1\}$; the instrument $z=1$ uniquely and strictly targets the Head Start treatment $h$. In the terminology of this paper, this assignment structure satisfies universal targeting. The set of inequalities in (ref) is empty, as $z=0$ is the only non-targeting instrument value. On the other hand, (ref) yields the following sharp testable implications: \[ P(n\vert 1) < P(n\vert 0) \quad \text{and} \quad P(c\vert 1) < P(c\vert 0). \] In words, the Head Start offer must strictly reduce the choice probabilities of the other two alternatives (no preschool $n$ and other centers $c$). The data is fully consistent with these implications, as evidenced by the proportions of the two complier groups reported in Panel A of (ref) (see also kline2016).
Figure (ref) reproduces Figure (ref) in this setting. In addition to the three always-taker groups $C_{nn}$, $C_{cc}$, and $C_{hh}$, there are two complier groups: $C_{nh}$, and $C_{ch}$. In Sections (ref) and (ref), we focus on the LATEs on the two complier groups $C_{nh}$ and $C_{ch}$. Section (ref) embeds the model into a larger, $3\times 3$ model in order to evaluate the marginal value of the public funds used in Head Start.
Our estimates of the proportions of the two complier groups in the sample use (ref) in (ref); they are shown in Panel A of (ref). As expected, they coincide with those in kline2016.
Panel B of (ref) shows the counterfactual means of test scores for the complier groups, as per (ref). While $\mathbb{E}[ Y_i(n) | i \in C_{nh}]$ is negative, $\mathbb{E}[ Y_i(c) | i \in C_{ch}]$ is above $0.1$ standard deviation---not a negligible value in this field. This suggests that some of the children who enter Head Start would have been at a good preschool otherwise. kline2016 call this pattern the “substitution effect” of Head Start. However, they do not report estimates of $\mathbb{E}[ Y_i(n) | i \in C_{nh}]$ and $\mathbb{E}[ Y_i(c) | i \in C_{ch}]$.
To fully measure the substitution effect, one needs to identify $\mathbb{E} \left[ Y_i(h) | i \in C_{nh} \right]$ and $\mathbb{E} \left[ Y_i(h) | i \in C_{ch} \right]$. However, we know from (ref) that they are only partially identified by
where $\alpha_0=(P(c\vert 0) - P(c\vert 1))/(P(h\vert 1)-P(h\vert 0))$. This is exactly the formula on kline2016: as they point out, the LATE for Head Start is a weighted average of “subLATEs” with weights determined by the proportion of $C_{ch}$ among compliers, which is identified from the data\footnote{Our $\alpha_0$ is denoted $S_c$ in their paper.}.
kline2016 first tried to identify $\mathbb{E}[ Y_i(h) - Y_i(c) | i \in C_{ch}]$ and $\mathbb{E}[ Y_i(h) - Y_i(n) | i \in C_{nh}]$ separately using interactions of the instrument with covariates or experimental sites. They acknowledged the limitations of this approach and resorted to a parametric selection model \`{a} la Heckman1979 instead. They report\footnote{See kline2016.} estimates of the local average treatment effects of $0.370$ for $C_{nh}$ and $-0.093$ for $C_{ch}$, with respective standard errors $0.088$ and $0.154$. The resulting point estimate of the difference is quite large, at $0.463$ standard deviation.
Our (ref) provides an alternative approach to separating the two treatment effects. Given that compliers coming from other preschools ($C_{ch}$) had better test scores than compliers not originally in preschools ($C_{nh}$), it seems reasonable to assume that they also have better test scores under Head Start:
This is a “positive selection” that fits within the framework of (ref). It can be derived in a simple model in which better students benefit more from Head Start and other preschools; and students choose schools as a function of their expected outcome. We show in Appendix (ref) that this model generates positive selection under reasonable assumptions. The pooled cohort estimates in Panel C of (ref) indicate that the upper bound on $\mathbb{E}[Y_i(h)-Y_i(n)\mid i\in C_{nh}]$ is $0.28$, while the lower bound on $\mathbb{E}[Y_i(h)-Y_i(c)\mid i\in C_{ch}]$ is $0.09$.\footnote{Under monotone treatment response, $Y_i(h)\ge Y_i(n)$, the lower bound on $\mathbb{E}[Y_i(h)-Y_i(n)\mid i\in C_{nh}]$ would be zero, yielding a 95% confidence interval $0\le \mathbb{E}[Y_i(h)-Y_i(n)\mid i\in C_{nh}]\le 0.36$.} The difference between these two numbers gives an upper bound of $0.19$ for the difference of these two LATEs, with a standard error of $0.07$. Negative selection (reversing the inequality (ref)) would make $0.19$ a {\em lower\/} bound for the difference of the LATEs. At the same time, it would imply that the lower bound is negative; this is soundly rejected by the data.
Our upper bound of $0.19$ is much lower than the point estimate reported by kline2016. In fact, our 95% and 99% one-sided confidence intervals for \[ \mathbb{E}[ Y_i(h) - Y_i(n) | i \in C_{nh}] - \mathbb{E}[ Y_i(h) - Y_i(c) | i \in C_{ch}] \] are $(-\infty, 0.308)$ and $(-\infty, 0.356)$. We conclude that the $0.463$ estimate in kline2016 may overstate the difference between the two complier groups: it can only be rationalized under negative selection, which is a much less plausible assumption.
kline2016 sought to evaluate the welfare effect of increasing the number of slots in Head Start, as summarized by the marginal value of public funds (MVPF). They note that any expansion of Head Start will vacate some slots at competing preschools, which are oversubscribed. The relaxation of this rationing must be counted as an effect of Head Start expansions. This is what they call “rationed substitutes”\footnote{See kline2016 for details.}.
The children who move from $T_i=n$ to $T_i=c$ when a slot is vacated by a child who moves to Head Start constitute a $C_{nc}$ group that is ruled out by the $2\times 3$ model. These children increase their grades by $Y_i(c)-Y_i(n)$, whose average generates a LATE that we denote $\textrm{LATE}_{nc}$. Equation (9) in kline2016 shows that the value of $\textrm{LATE}_{nc}$ is a crucial input in the computation of the MVPF of a Head Start expansion. Identifying it requires either data on offers to all preschools, which kline2016 do not have\footnote{See footnote 19 in their paper.}, or additional modeling assumptions. They used their parametric selection model to construct an estimate for $\textrm{LATE}_{nc}$. Their estimate of $\textrm{LATE}_{nc} = 0.294$ results in a high MVPF estimate of $2.02$ (see Table IX in their paper).
We take a different approach by embedding the $2\times 3$ model within a $3 \times 3$ model. In this richer model, the instrument can take three values: in addition to the control group ($z=0$) and those offered admission to Head Start ($z=1$), we have a new group that we denote $z=2$. This group receives a direct offer of admission to a competing preschool. Economically, such offers arise when a Head Start expansion vacates slots at competing preschools---what kline2016 call “rationed substitutes”---but formally $z=2$ is simply an additional instrument value requiring only that it satisfies (ref). Note that this maintains strict, one-to-one targeting. Since potential outcomes $Y_i(t)$ only depend on child $i$'s own treatment $t$, SUTVA still holds: we condition on the market equilibrium in which $z=2$ offers exist, consistent with the partial equilibrium framework of kline2016.
Figure (ref) shows the resulting treatment assignment, using tildes to denote the complier groups of the $3\times 3$ model\footnote{Again, it is just Figure (ref) with different notation.}. Using this notation, $\textrm{LATE}_{nc}$ can be written as
where $\tilde{C}_{n*c} = \tilde{C}_{nnc} \operatorname*{\mathsmaller{\bigcup}} \tilde{C}_{nhc}$ is the composite group of $n\to c$ compliers. Comparing Figure (ref) with Figure (ref) shows that the other complier groups of the two models are linked by
Now consider the new group of $n\to c$ compliers. It differs from $\tilde{C}_{chc}$ in that its members will not go to a preschool unless they are offered a slot. We show in Appendix (ref) that the structural model that we used in the binary instrument case predicts the following inequality:
Now consider the composite response-groups $\tilde{C}_{n*c} = \tilde{C}_{nnc} \operatorname*{\mathsmaller{\bigcup}} \tilde{C}_{nhc}$ and $\tilde{C}_{nn\ast}=\tilde{C}_{nnc}\operatorname*{\mathsmaller{\bigcup}} \tilde{C}_{nnn}$. As Figure (ref) shows, they only differ by the substitution of $\tilde{C}_{nnn}$ for $\tilde{C}_{nhc}$. The former never go to Head Start or to another preschool, while the latter are full compliers. Our structural model generates the inequality \[ \mathbb{E}[ Y_i(n) | i \in \tilde{C}_{nnn}] \leq \mathbb{E}[ Y_i(n) | i \in \tilde{C}_{nhc}] \] which implies
Inequalities (ref) and (ref) again are “positive selection” assumptions that fall under our (ref).
Since $\tilde{C}_{nn\ast}$ coincides with $C_{nn}$ and $\tilde{C}_{chc}$ is $C_{ch}$, we already know the values of the right-hand sides of both inequalities. Applying the same logic as in (ref) gives us an upper bound for $\textrm{LATE}_{nc}$:
As the $\textrm{MVPF}$ is an increasing function of $\textrm{LATE}_{nc}$, this gives us in turn an upper bound on its value\footnote{Online Appendix (ref) derives the formula for the MVPF in this model.}. We obtain $\textrm{LATE}_{nc} \leq 0.164$ and $\textrm{MVPF} \leq 1.55$. These upper bounds are noticeably smaller than the point estimates that result from the parametric selection model of kline2016.
We have shown that the idea of targeting is a useful way to analyze models with multivalued treatments and multivalued instruments. The testable implications that we derived suggest a natural three-step procedure to elucidate targeting patterns in the data, supplementing any qualitative information provided by the empirical context.
\paragraph{Step 1: Screening Candidates.} For each instrument-target pair $(z,t)\in\mathcal{Z}\times\mathcal{T}$, test the hypothesis $H_{0,zt}: S_{zt}\le 0$ against $H_{1,zt}: S_{zt}>0$, where \[ S_{z t} = P(0\mid z)+\max_{z''\in\mathcal{Z}, z^{\prime\prime}\neq z} P(t\mid z'')-1. \] Let $\widehat{\mathcal{V}}$ denote the set of surviving pairs: \[ \widehat{\mathcal V} \equiv \{(z,t)\in\mathcal{Z}\times\mathcal{T}: H_{0,zt}\ \text{is not rejected}\}. \] If for a fixed $z$, no pair $(z,t)$ survives in $\widehat{\mathcal V}$, we reject the hypothesis that $z$ targets any treatment in $\mathcal{T}$.
\paragraph{Step 2: Checking Compatibility} For every pair of candidates $(z,t)$ and $(z',t')$ in $\widehat{\mathcal V}$ with $t\neq t'$, we test $H_{0,c}: S_{c} \le 0$, where \[ S_{c} = P(t'\mid z)+P(t\mid z')-1. \] A rejection implies that $(z,t)$ and $(z',t')$ cannot hold simultaneously.
\paragraph{Step 3: Testing for Universal Targeting} The set of valid mappings that survive Steps 1 and 2 may or may not be consistent with universal targeting. One can further use the inequalities in (ref) to formally test it.
As usual, one should be careful about controlling the size of this three-step procedure by applying standard multiple-testing corrections within each step and/or bootstrapping. We leave this for further research.
Our paper only analyzed discrete-valued instruments and treatments. Some of the notions we used would extend naturally to continuous instruments and treatments: the definitions of targeting, one-to-one targeting, and positive selection would translate directly. Strict targeting, on the other hand, is less appealing in a context in which continuous values may denote intensities. Our earlier paper LeeSalanie2018 as well as \citeapos{Mountjoy:2022} can be seen as analyzing continuous-instruments/discrete-treatments models. Extending our analysis to models with continuous treatments is an obvious topic for further research. It would also be interesting to apply the partial identification approach of MST:2018 in our setting. Finally, there has been a surge of recent interest on understanding the properties of OLS and 2SLS estimands when treatment effects vary with the covariates BBMT,Sloc22,GHK:24. We believe that the targeting concept and the identifying assumptions explored in this paper should be relevant in this context and that they merit further investigation.