EconBase
← Back to paper

2SLS with Multiple Treatments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,482 characters · 19 sections · 105 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

2SLS with Multiple Treatments

singlespaceAbstract: We study what two-stage least squares (2SLS) identifies in models with multiple treatments under treatment effect heterogeneity. Two conditions are shown to be necessary and sufficient for the 2SLS to identify positively weighted sums of agent-specific effects of each treatment: average conditional monotonicity and no cross effects. Our identification analysis allows for any number of treatments, any number of continuous or discrete instruments, and the inclusion of covariates. We provide testable implications and present characterizations of choice behavior implied by our identification conditions.
singlespaceKeywords: multiple treatments, monotonicity, instrumental variables, two-stage least squares\\ JEL codes: C36

\thispagestyle{empty}

{\setcounter{page}{1}}

Introduction

In many settings\textemdash e.g., education, career choices, and migration decisions\textemdash estimating the causal effects of a series of different treatments is valuable. For instance, in the context of criminal justice, one might be interested in separately estimating the effects of conviction and incarceration on defendant outcomes Humphries2022,Kamat2023. Identifying treatment effects in settings with multiple treatments using instruments has, however, proven challenging HeckmanEtal2008_Annales. A common approach in the applied literature is to estimate a “multivariate” two-stage least squares (2SLS) regression with indicators for receiving various treatments as multiple endogenous variables and at least as many instruments.\footnote{Examples of studies estimating models with multiple treatments using 2SLS\textemdash either as the main specification or as an extension of a baseline specification with a binary treatment\textemdash include persson2004constitutional,acemoglu2005unbundling,rohlfs2006government,angrist2006instrumental,angrist2009incentives,maestas2015does,muellersmith2015,kline2016evaluating,jaeger2018shift,bombardini2020trade,BhullerEtAl2020,norris2021effects,angrist2022methods,Humphries2022,BhullerSigstad2021,Kamat2023,heinesen2022instrumental.} While such an approach is valid under homogenous treatment effects, it does not generally identify meaningful treatment effects under treatment effect heterogeneity.

To fix ideas about the identification problem, consider a case with three mutually exclusive treatments\textemdash $T\in\left\{ 0,1,2\right\} $\textemdash and a vector of valid instruments $Z$. Let $\beta_{1i}$ and $\beta_{2i}$ be the causal effects of receiving treatments $1$ and $2$ relative to receiving treatment $0$ for agent $i$ on an outcome variable. Then, the 2SLS estimand of the causal effect of receiving treatment $1$ is, in general, a weighted sum of $\beta_{1i}$ and $\beta_{2i}$ across agents, where the weights can be negative. Thus, the estimated effect of treatment $1$ can both put negative weight on the effect of treatment $1$ for some agents and be contaminated by the effect of treatment $2$. In severe cases, the estimated effect of treatment $1$ can be negative even though $\beta_{1i}>0$ for all agents. The existing literature has, however, not clarified which conditions are necessary for the 2SLS estimand of the effect of treatment $1$ to assign proper weights\textemdash non-negative weight on $\beta_{1i}$ and zero weight on $\beta_{2i}$.

In this paper, we present two necessary and sufficient conditions\textemdash besides the standard rank, exclusion, and exogeneity assumptions\textemdash for multivariate 2SLS to assign proper weights: average conditional monotonicity and no cross effects. Our results apply in the general case with $n$ treatments and $m\geq n$ continuous or discrete instruments. We also provide results for settings where exogeneity holds only conditional on covariates\textemdash a common feature in applied work that complicates the analysis of 2SLS already in the binary treatment case sloczynski2020should,blandhol2022tsls. Finally, our results allow the researcher to specify any set of relative treatment effects, not just effects relative to an excluded treatment (treatment $0$).

For expositional ease, we continue with the three-treatment example. When does the 2SLS estimand of receiving treatment $1$ put non-negative weight on $\beta_{1i}$ and zero weight on $\beta_{2i}$? To develop intuition about the required conditions, assume we run 2SLS with the following two instruments: the linear projection of an indicator for receiving treatment $1$ on the instrument vector $Z$, which we refer to as “instrument 1”, and the linear projection of an indicator for receiving treatment $2$ on the instrument vector $Z$, i.e., “instrument 2”. Using instruments 1 and 2\textemdash the predicted treatments from the 2SLS first stage\textemdash as instruments is numerically equivalent to using the full vector $Z$ as instruments. The first condition\textemdash average conditional monotonicity\textemdash requires that, conditional on instrument 2, instrument 1 does not, on average, induce agent $i$ out of treatment $1$.\footnote{This condition generalizes FrandsenEtAl2019's average monotonicity condition to multiple treatments and is substantially weaker than the Imbens-Angrist monotonicity condition imbens1994identification.} The second condition\textemdash no cross effects\textemdash requires that, conditional on instrument 1, instrument 2 does not, on average, induce agent $i$ into or out of treatment $1$. The latter condition is necessary to ensure zero weight on $\beta_{2i}$ and is particular to the case with multiple treatments. We show that this condition is equivalent to assuming homogeneity in agents' relative responses to the instruments: How instrument $1$ affects treatment $2$ compared to how instrument $2$ affects treatment $2$ can not differ between agents.

We derive a set of testable implications of our conditions. First, we show that when outcome variable transformations interacted with a treatment $1$ indicator are regressed on the instruments, the coefficient on instrument $1$ should be non-negative and the coefficient on instrument $2$ should be zero. Second, we show that the same must hold when regressing an indicator for treatment $1$ on the instruments in subsamples that are constructed using pre-determined covariates. Using data from BhullerEtAl2020, we show how these tests can be implemented in practice.

Building upon our general identification results, we consider a prominent special case: 2SLS with $n$ mutually exclusive treatments and $n-1$ mutually exclusive binary instruments. This case is a natural generalization of the canonical case with a binary treatment and a binary instrument to multiple treatments. A typical application is a randomized controlled trial where each agent is randomly assigned to one of $n$ treatments, but compliance is imperfect. In this setting, our two conditions require that each instrument affects exactly one treatment indicator. In particular, there must be a labeling of the instruments such that instrument $k$ moves agents only from the excluded treatment $0$ into treatment $k$. This result gives rise to additional testable implications of our conditions in models with one binary instrument per treatment indicator.

The requirement that each instrument affects exactly one choice margin restricts choice behavior in a particular way. In particular, 2SLS assigns proper weights only when choice behavior can be described by a selection model where the excluded treatment is always the preferred alternative or the next-best alternative. It is thus not sufficient that each instrument influences the utility of only one choice alternative. For instance, an instrument that affects only the utility of receiving treatment $1$ could still affect the take-up of treatment $2$ by inducing agents who would otherwise have selected treatment $2$ to select treatment $1$. Such cross effects are avoided if the excluded treatment is always at least the next-best alternative. To apply 2SLS in this case, the researcher must argue why the excluded treatment is always the best or the next-best alternative. Our results essentially imply that unless researchers can infer next-best alternatives\textemdash as in kirkeboen2016 and the following literature\textemdash 2SLS in models with one binary instrument per treatment indicator does not identify a meaningful causal effect under arbitrary heterogeneous effects.

Until now, we have considered unordered treatment effects\textemdash treatment effects relative to an excluded treatment. Our results, however, also apply to any other relative treatment effects a researcher might seek to estimate through 2SLS. An important case is ordered treatment effects\textemdash the effect of treatment $k$ relative to treatment k-1.\footnote{AngristImbens1995JASA showed the conditions under which 2SLS with the multivalued treatment indicator $T$ as the endogenous variable identifies a convex combination of the effect of treatment $2$ relative to treatment $1$ and the effect of treatment $1$ relative to treatment $0$. In contrast, we seek to determine the conditions under which 2SLS with two binary treatment indicators\textemdash $D_{1}=\mathbf{1}\left[T\geq1\right]$ and $D_{2}=\mathbf{1}\left[T=2\right]$\textemdash \emph{separately }identifies the effect of each of the two treatment margins.} In the ordered case, 2SLS with one binary instrument per treatment indicator assigns proper weights if and only if there exists a labeling of the instruments such that instrument $k$ moves agents only from treatment \emph{k-1} to treatment $k$. This condition also imposes a particular restriction on agents' choice behavior: we show that 2SLS assigns proper weights in such ordered choice models only when agents' preferences can be described as single-peaked over the treatments. When treatments have a natural ordering\textemdash such as years of schooling\textemdash the researcher might be able to make a strong theoretical case in favor of such preferences.

We finally present another special case of ordered choice where our conditions are satisfied: a classical threshold-crossing model applicable when treatment assignment depends on a single index crossing multiple thresholds. For instance, treatments can be grades and the latent index the quality of the student's work, or treatments might be years of prison and the latent index the severity of the committed crime. Suppose the researcher has access to exogenous shocks to these thresholds, for instance through random assignment to judges or graders that agree on ranking but use different cutoffs. Then 2SLS assigns proper weights provided that there is a linear relationship between the predicted treatments\textemdash an easily testable condition.\footnote{More precisely, the conditional expectation of predicted treatment $k$ must be a linear function of predicted treatment $l\neq k$.}

Our paper contributes to a growing literature on the use of instruments to identify causal effects in settings with multiple treatments HeckmanEtal2008_Annales,kline2016evaluating,kirkeboen2016,HeckmanPinto2018,lee2018identifying,Galindo2020Empirical,pinto2022beyond,Humphries2022,heinesen2022instrumental,kamat2023identification or multiple instruments MogstadEtAl2019,Goff2020Vector,mogstad2020policy. Our main contribution is to provide the exact conditions under which 2SLS with multiple treatments assigns proper weights under arbitrary treatment effect heterogeneity. We allow for any number of treatments, any number of continuous or discrete instruments, any definition of treatment indicators, and covariates. Moreover, we show how the conditions can be refuted. By comparison, the existing literature provides only sufficient conditions in the case with three treatments, three instrument values, and no controls BehagheletAl2013,kirkeboen2016.

In the case with one binary instrument per treatment indicator, we show that the extended monotonicity condition provided by BehagheletAl2013 is not only sufficient but also necessary for 2SLS to assign proper weights, after a possible permutation of the instruments. This non-trivial result gives rise to a new testable implication: For 2SLS to assign proper weights in such models, each instrument can only affect one treatment. Furthermore, we show that knowledge of agents' next-best alternatives\textemdash as in kirkeboen2016\textemdash is implicitly assumed whenever estimates from just-identified 2SLS models with multiple treatments are interpreted as a positively weighted sum of individual treatment effects. We thus show that the assumption that next-best alternatives are observed or can be inferred is not only sufficient but also essentially necessary for 2SLS to identify a meaningful causal parameter.\footnote{The only exception being that next-best alternatives need not be observed for always-takers.} We also provide new identification results for ordered treatments. First, we show when 2SLS with multiple ordered treatments identifies separate treatment effects in a standard threshold-crossing model considered in the ordered choice literature (e.g., carneiro2003estimating,cunha2007identification,Heckman2007EconometricII). While Heckman2007EconometricII show that local IV identifies ordered treatment effects in such a model, we show that 2SLS can also identify the effect of each treatment transition under an easily testable linearity condition. A similar result is found concurrently by Humphries2022.\footnote{Humphries2022 also show that if treatment assignment depends on several unobserved latent indices\textemdash instead of just one\textemdash 2SLS does not, in general, assign proper weights.} We also show how the result of BehagheletAl2013 extends to ordered treatment effects. Finally, we show that for 2SLS to assign proper weights in versions of the BehagheletAl2013 model with ordered treatment effects, it must be possible to describe agents' preferences as single-peaked over the treatments.

In contrast to HeckmanPinto2018, who provide general identification results in a setting with multiple treatments and discrete instruments, we focus specifically on the properties of 2SLS\textemdash a standard and well-known estimator common in the applied literature. Other contributions to the literature on the use of instrumental variables to separately identify multiple treatments lee2018identifying,Galindo2020Empirical,lee2020filtered,Mountjoy2019,pinto2022beyond focus on developing new approaches to identification. We do not necessarily recommend 2SLS over these alternative methods. For instance, the method of HeckmanPinto2018 identifies causal effects under strictly weaker assumptions than those required for 2SLS to assign proper weights in models with two binary treatment indicators and three treatments (see Section (ref)). Also, as argued by Heckman2007EconometricII, the weighted average of treatment effects produced by 2SLS in overidentified models is not necessarily an interesting parameter, even when the weights are non-negative. The alternatives to 2SLS, discussed in Section (ref), arguably all target more policy-relevant treatment effects. But given the popularity of 2SLS among practitioners, we still see a need to clarify the exact conditions under which this is a valid approach\textemdash contributing to a recent body of research assessing the robustness of standard estimators to heterogeneous effects (e.g., de2020two,de2020two_several,callaway2020difference,goodman2021difference,sun2020estimating,borusyak2021revisiting,goldsmith2022contamination).

In Section (ref), we develop the exact conditions for the multivariate 2SLS to assign proper weights to agent-specific causal effects and present testable implications. In Section (ref), we consider two special cases. Section (ref) discusses implications when the conditions fail and alternatives to 2SLS. Section (ref) provides an illustration of our conditions for the random judge IV design and tests these conditions using data from BhullerEtAl2020. We conclude in Section (ref). Proofs and additional results are in the Appendix.

Multivariate 2SLS with Heterogeneous Effects

In this section, we develop sufficient and necessary conditions for the multivariate 2SLS to identify a positively weighted sum of individual treatment effects under heterogeneous effects and present testable implications of these conditions.

Definitions

Fix a probability space where an outcome corresponds to a randomly drawn agent $i$.\footnote{All random variables thus correspond to a randomly drawn agent. We omit $i$ subscripts.} Define the following random variables: A multi-valued treatment $T\in\mathcal{T}\equiv\left\{ 0,1,2\right\} $, an outcome $Y\in\mathbb{R}$, and a vector-valued instrument $Z\in\mathcal{Z}\subseteq\mathbb{R}^{m}$ with $m\geq2$.\footnote{We focus on mutually exclusive treatments. Treatments that are not mutually exclusive can always be made into mutually exclusive treatments. For instance, two not mutually exclusive treatments and an excluded treatment can be thought of as four mutually exclusive treatments: Receiving the excluded treatment, receiving only treatment 1, receiving only treatment 2, and receiving both treatments.} For expositional ease, we focus on the case with three treatments and no control variables. In Section (ref), we show that all our results generalize to an arbitrary number of treatments, and in Section (ref), we show how our results extend to the case with control variables.

Define $\mathcal{S}$ as all possible mappings $f:\mathcal{Z}\rightarrow\mathcal{\mathcal{T}}$ from instrument values to treatments\textemdash all possible ways the instrument can affect treatment. Following HeckmanPinto2018, we refer to the elements of $\mathcal{S}$ as response types. The random variable $S\in\mathcal{S}$ describes agents' potential treatment choices: If agent $i$ has $S=s$, then $s\left(z\right)$ for $z\in\mathcal{Z}$ indicates the treatment selected by agent $i$ if $Z$ is set to $z$. The response type of an agent describes how the agent's choice of treatment reacts to changes in the instrument. For example, in the case with a binary treatment and a binary instrument, the possible response types are never-takers ($s\left(0\right)=s\left(1\right)=0$), always-takers ($s\left(0\right)=s\left(1\right)=1$), compliers ($s\left(0\right)=0$, $s\left(1\right)=1$), and\emph{ defiers }($s\left(0\right)=1$, $s\left(1\right)=0$)\emph{.} Similarly, define $Y\left(k\right)$ for $k\in\mathcal{\mathcal{T}}$ as the agent's \emph{potential outcome} when $T$ is set to $k$.

Using 2SLS, a researcher can aim to estimate two out of three relative treatment effects. We focus on two special cases: the unordered and ordered case. In the unordered case, the researcher seeks to estimate the effects of treatment 1 and 2 relative to treatment 0 by estimating 2SLS with treatment indicators $D_{1}^{\text{unordered}}\equiv\mathbf{1}\left[T=1\right]$ and $D_{2}^{\text{unordered}}\equiv\mathbf{1}\left[T=2\right]$.\footnote{The more general case\textemdash where the researcher seeks to estimate all relative treatment effects\textemdash could be analyzed by varying which treatment is considered to be the excluded treatment. In that case, the researcher should discuss and test the conditions in Section (ref) for each choice of excluded treatment. Often, however, the researcher will only be interested in some relative treatment effects.} In that case, the treatment effects of interest are represented by the random vector $\beta^{\text{unordered}}\equiv\left(Y\left(1\right)-Y\left(0\right),Y\left(2\right)-Y\left(0\right)\right)^{T}$. In the ordered case, the researcher is interested in $\beta^{\text{ordered}}\equiv\left(Y\left(1\right)-Y\left(0\right),Y\left(2\right)-Y\left(1\right)\right)^{T}$ and uses treatment indicators $D_{1}^{\text{ordered}}\equiv\mathbf{1}\left[T\geq1\right]$ and $D_{2}^{\text{ordered}}\equiv\mathbf{1}\left[T=2\right]$. In general terms, we let $\beta\equiv\left(\beta_{1},\beta_{2}\right)^{T}$ denote the treatment effects of interest and $D\equiv\left(D_{1},D_{2}\right)^{T}$ the corresponding treatment indicators. Unless otherwise specified, our results hold for any definition of $\beta$. But to ease exposition we focus on the unordered case when interpreting our results. We maintain the following standard IV assumptions throughout:

assumption(Exogeneity and Exclusion). $\left\{ Y\left(0\right),Y\left(1\right),Y\left(2\right),S\right\} \perp Z$
assumption(Rank). $\operatorname{Cov}\left(Z,D\right)$ has full rank.

For a response type $s\in\mathcal{S}$, let $s_{1}$ and $s_{2}$ be the induced mapping between instruments and treatment indicators. For instance, $s_{k}\left(z\right)\equiv\mathbf{1}\left[s\left(z\right)=k\right]$ if we use unordered treatment indicators, and $s_{k}\left(z\right)\equiv\mathbf{1}\left[s\left(z\right)\geq k\right]$ if we use ordered treatment indicators. Define the 2SLS estimand $\beta^{\text{2SLS}}=\left(\beta_{1}^{\text{2SLS}},\beta_{2}^{\text{2SLS}}\right)^{T}$by \[ \beta^{\text{2SLS}}\equiv\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,Y\right) \] where $P$ is the linear projection of $D$ on $Z$ \[ P=\left(P_{1},P{}_{2}\right)^{T}\equiv\operatorname{E}\left[D\right]+\operatorname{Var}\left(Z\right)^{-1}\operatorname{Cov}\left(Z,D\right)\left(Z-\operatorname{E}\left[Z\right]\right). \] We refer to $P$ as the predicted treatments\textemdash the best linear prediction of the treatment indicators given the value of the instruments. Similarly, we refer to $P_{k}$ for $k\in\left\{ 1,2\right\} $ as predicted treatment $k$. Since 2SLS using predicted treatments as instruments is numerically equivalent to using the original instruments, we can think of $P$ as our instruments. In particular, $P$ is a linear transformation of the original instruments $Z$ such that we get one instrument corresponding to each treatment indicator. We will occasionally refer to $P_{1}$ and $P_{2}$ as “instrument $1$” and “instrument $2$”. Let $\tilde{P}_{1}\equiv P_{1}-\frac{\operatorname{Cov}\left(P_{1},P_{2}\right)}{\operatorname{Var}\left(P_{2}\right)}P_{2}$ be the residualized instrument $1$ after netting out any linear association with instrument $2$.

Identification Results

What does multivariate 2SLS identify under Assumptions (ref) and (ref) when treatment effects are heterogeneous across agents? The following proposition expresses the 2SLS estimand as a weighted sum of average treatment effects across response types:\footnote{By symmetry, an analogous expression can be derived for $\beta_{2}^{\text{2SLS}}$.}

propUnder Assumptions (ref) and (ref) \[ \beta_{1}^{\text{2SLS}}=\operatorname{E}\left[w_{1}^{S}\beta_{1}^{S}+w_{2}^{S}\beta_{2}^{S}\right] \] where, for a realization $s\in\mathcal{S}$ of $S$ and $k\in\left\{ 1,2\right\} $ \begin{eqnarray*} w_{k}^{s}\equiv\frac{\operatorname{Cov}\left(\tilde{P}_{1},s_{k}\left(Z\right)\right)}{\operatorname{Var}\left(\tilde{P}_{1}\right)}; & & \beta_{k}^{s}\equiv\operatorname{E}\left[\beta_{k}\mid S=s\right]. \end{eqnarray*}

In the case of a binary treatment and a binary instrument, $\beta^{\text{2SLS}}$ is a weighted sum of average treatment effects for compliers and defiers. Proposition (ref) is a generalization of this result to the case with three treatments and $m$ instruments. In this case, $\beta_{1}^{\text{2SLS}}$ is a weighted sum of the average treatment effects of all response types present in the population. The parameter $\beta_{k}^{s}$ describes the average effects of treatment $k$ for agents of response type $s$. The weight $w_{k}^{s}$ indicates how the average effects of treatment $k$ of agents with response type $s$ contribute to the estimated effect of treatment $1$. Without further restrictions, these weights could be both positive and negative. Thus, in general, both treatment effects for response type $s$ might contribute, either positively or negatively, to the estimated effect of treatment $1$.

In the canonical case of a binary treatment and a binary instrument, identification is ensured when there are no defiers. Proposition (ref) generates similar restrictions on the possible response types in the case of multiple treatments. To become familiar with our notation and the implications of Proposition (ref), consider the following example.

exampleConsider the case of three treatments and two mutually exclusive binary instruments: $\left(Z_{1},Z_{2}\right)\in\left\{ \left(0,0\right),\left(1,0\right),\left(0,1\right)\right\} $. Assume we are interested in unordered treatment effects, $\beta^{\text{unordered}}=\left(Y\left(1\right)-Y\left(0\right),Y\left(2\right)-Y\left(0\right)\right)^{T}$, with corresponding treatment indicators $D_{1}=\mathbf{1}\left[T=1\right]$ and $D_{2}=\mathbf{1}\left[T=2\right]$. One possible response type in this case is defined by \[ s\left(z\right)=\begin{cases} 2 & \text{if }z=\left(0,0\right)\\ 1 & \text{if }z=\left(1,0\right)\\ 2 & \text{if }z=\left(0,1\right) \end{cases}. \] Response type $s$ thus selects treatment 2 unless $Z_{1}$ is turned on. When $Z_{1}=1$, $s$ selects treatment $1$. Assume the best linear predictors of the treatments indicators are $P_{1}=0.1+0.4Z_{1}$ and $P_{2}=0.2+0.5Z_{2}$ and that all instrument values are equally likely. By Proposition (ref), this response type's contribution to $\beta_{1}^{\text{2SLS}}$ would be $w_{1}^{s}\beta_{1}^{s}+w_{2}^{s}\beta_{2}^{s}=2\beta_{1}^{s}-2\beta_{2}^{s}$.\footnote{We have $\tilde{P}_{1}=\frac{1}{4}\left(1.05+2Z_{1}+Z_{2}\right)$, $s_{1}\left(Z\right)=Z_{1}$, and $s_{2}\left(Z\right)=1-Z_{1}$, which gives $w_{1}^{s}=\operatorname{Cov}\left(\tilde{P}_{1},s_{1}\left(Z\right)\right)/\operatorname{Var}\left(\tilde{P}_{1}\right)=\frac{1}{12}/\frac{1}{24}=2$ and $w_{2}^{s}=-2$.} The average effect of treatment $2$ for response type $s$ thus contributes negatively to the estimated effect of treatment $1$. The presence of this response type in the population would be problematic. The response type \[ s\left(z\right)=\begin{cases} 0 & \text{if }z=\left(0,0\right)\\ 1 & \text{if }z=\left(1,0\right)\\ 0 & \text{if }z=\left(0,1\right) \end{cases} \] on the other hand, has weights $w_{1}^{s}=2$ and $w_{2}^{s}=0$. This response type's effect of treatment $2$ does not contribute to the estimated effect of treatment $1$. Two-stage least squares assigns proper weights on this response type's average treatment effects.\footnote{The weights on $\beta_{1}^{s}$ and $\beta_{2}^{s}$ in $\beta_{2}^{\text{2SLS}}$ are both zero in this example.}$\square$

Under homogeneous effects, or homogenous responses to the instruments, the weight 2SLS assigns on the treatment effects of a particular response type is not a cause of concern. By the following corollary of Proposition (ref), the “cross” weights $w_{2}^{s}$ are zero on average and the “own” weights $w_{1}^{s}$ are, on average, positive:

corUnder Assumptions (ref) and (ref), the 2SLS estimand $\beta_{1}^{\text{2SLS}}$ is a weighted sum of $\beta_{1}^{s}$ and $\beta_{2}^{s}$ where the weights on the first sum to one and the weights on the latter sum to zero.

Thus, if $\beta_{1}^{s}=\beta_{1}$ and $\beta_{2}^{s}=\beta_{2}$ for all $s\in\mathcal{S}$, we get $\beta_{1}^{\text{2SLS}}=\operatorname{E}\left[w_{1}^{S}\beta_{1}+w_{2}^{S}\beta_{2}\right]=\beta_{1}$. Also, if the population consists of only one response type $s$, we get $\beta_{1}^{\text{2SLS}}=\beta_{1}^{s}$.\footnote{One can allow for always-takers and never-takers.} But under heterogeneous effects and heterogeneous responses to the instruments, the estimated effect of treatment $1$ might be contaminated by the effect of treatment $2$. This makes interpreting 2SLS estimates hard. Ideally, we would like $w_{1}^{s}$ to be non-negative, and $w_{2}^{s}$ to be zero for all $s$. Only in that case can we interpret the 2SLS estimate of the effect of treatment $1$ as a positively weighted average of the effect of treatment $1$ under heterogeneous effects. Throughout the rest of the paper, we say that the 2SLS estimand of the effect of treatment $1$ assigns proper weights if the following holds:

defnThe 2SLS estimand of the effect of treatment 1 assigns proper weights if for each $s\in\operatorname{supp}\left(S\right)$, the 2SLS estimand $\beta_{1}^{\text{2SLS}}$ places non-negative weight on $\beta_{1}^{s}$ and zero weight on $\beta_{2}^{s}$.

We say that 2SLS assigns proper weights if it assigns proper weights for both treatment effects.\footnote{The requirement that 2SLS assign proper weights is equivalent to 2SLS being weakly causal (blandhol2022tsls)\textemdash providing estimates with the correct sign\textemdash under arbitrary heterogeneous effects (see Section (ref)). Note that we do not require the estimand to assign equal weights to all response types. Such a criterion would essentially rule out the use of 2SLS beyond the case studied in Section (ref): As noted by Heckman2007EconometricII, 2SLS assigns higher weights to response types that are more influenced by the instruments, producing a weighted average that is not necessarily of policy interest. See Section (ref) for alternatives to 2SLS that seek to target more policy-relevant parameters.} When is the weight $w_{1}^{s}$ on $\beta_{1}^{s}$ non-negative for all $s$? The formula for the weight $w_{1}^{s}$ equals the coefficient in a hypothetical linear regression of $s_{1}\left(Z\right)$ on $\tilde{P}_{1}$ across realizations of $Z$.\footnote{Or, equivalently, the coefficient on $P_{1}$ in a hypothetical regression of $s_{1}\left(Z\right)$ on $P_{1}$ and $P_{2}$. Such a regression is “hypothetical” since, in practice, we only observe $s_{1}\left(Z\right)$ for one realization of $Z$.} Thus $\beta_{1}^{\text{2SLS}}$ puts more weight on $\beta_{1}^{s}$ the more $\tilde{P}_{1}$ tends to push response type $s$ into treatment $1$. If $\tilde{P}_{1}$ tends to push the response type out of treatment $1$, the weight $w_{1}^{s}$ will be negative. The condition that ensures that $\beta_{1}^{\text{2SLS}}$ assigns non-negative weights on $\beta_{1}^{s}$ for all $s$ can be written succintly as follows.

assumption(Average Conditional Monotonicity). For all $s\in\mathcal{S}$ \[ \operatorname{E}\left[\tilde{P}_{1}\mid s\left(Z\right)=1\right]\ge\operatorname{E}\left[\tilde{P}_{1}\mid s\left(Z\right)\neq1\right]. \]
corUnder Assumptions (ref) and (ref), the weight on $\beta_{1}^{s}$ in $\beta_{1}^{\text{2SLS}}$ is non-negative for all $s$ if and only if Assumption (ref) holds.

Assumption (ref) requires that, after controlling linearly for predicted treatment $2$, the average predicted treatment $1$ can not be lower at instrument values where a given agent selects treatment $1$ than at other instrument values. Intuitively, Assumption (ref) requires that, controlling linearly for predicted treatment 2, there is a non-negative correlation between predicted treatment $1$ and potential treatment $1$ for each agent across values of the instruments. We refer to the condition as average conditional monotonicity since it generalizes the average monotonicity condition defined by FrandsenEtAl2019 to multiple treatments.\footnote{With one treatment, Assumption (ref) reduces to $\operatorname{E}\left[P_{1}\mid s\left(Z\right)=1\right]\geq\operatorname{E}\left[P_{1}\right]\Leftrightarrow\operatorname{Cov}\left(P_{1},s_{1}\left(Z\right)\right)\geq0$ which coincides with the average monotonicity condition in FrandsenEtAl2019. Under Assumption (ref), Assumption (ref) is equivalent to the correlation between $P_{1}$ and $s_{1}\left(Z\right)$ being non-negative (without having to condition on $P_{2}$).} In particular, the positive relationship between predicted treatment and potential treatment only needs to hold “on average” across realizations of the instruments. Thus, an agent might be a “defier” for some pairs of instrument values as long as she is a “complier” for sufficiently many other pairs. Informally, we can think of Assumption (ref), as requiring the partial effect of $P_{1}$\textemdash instrument $1$\textemdash on treatment $1$ to be, on average, non-negative for all agents.

The weight $w_{2}^{s}$ on $\beta_{2}^{s}$ in $\beta_{1}^{\text{2SLS}}$ equals the coefficient in a hypothetical linear regression of $s_{2}\left(Z\right)$ on $\tilde{P}_{1}$. Intuitively, if $P_{1}$ has a tendency to push certain agents into or out of treatment $2$\textemdash even after controlling linearly for $P_{2}$\textemdash the estimated effect of treatment $1$ will be contaminated by the effect of treatment $2$ on these agents. The following condition is necessary and sufficient to avoid such contamination:

assumption(No Cross Effects). For all $s\in\mathcal{S}$ \[ \operatorname{E}\left[\tilde{P}_{1}\mid s\left(Z\right)=2\right]=\operatorname{E}\left[\tilde{P}_{1}\mid s\left(Z\right)\neq2\right]. \]
corUnder Assumptions (ref) and (ref), the weight on $\beta_{2}^{s}$ in $\beta_{1}^{\text{2SLS}}$ is zero for all $s$ if and only if Assumption (ref) holds.

Assumption (ref) requires that, after controlling linearly for predicted treatment $2$, the average predicted treatment $1$ can not be different at instrument values where a given agent selects treatment $2$ than at other instrument values. In other words, after linearly controlling for predicted treatment $2$, there can be no correlation between predicted treatment $1$ and potential treatment $2$ for any agent. Informally, we can think of Assumption (ref) as requiring there to be no partial effect of instrument 1 on treatment $2$ for any agent. The following restatement of Assumption (ref) provides further intuition:

propAssumption (ref) is equivalent under Assumption (ref) to \[ \operatorname{Cov}\left(P_{1},s_{2}\left(Z\right)\right)=\rho\operatorname{Cov}\left(P_{2},s_{2}\left(Z\right)\right) \] for all $s\in\mathcal{S}$ and a constant $\rho$.

Thus, Assumption (ref) does not require $\operatorname{Cov}\left(P_{1},s_{2}\left(Z\right)\right)=0$\textemdash that take-up of treatment $2$ is (unconditionally) unaffected by instrument $1$. Instead, the condition requires that the extent that instrument $1$ affects treatment $2$ compared to how instrument $2$ affects treatment $2$ is constant across response types. Assumption (ref) thus imposes a certain homogeneity in agents' responses to the instruments. In particular, it is not allowed that some agents react strongly to instrument $1$ but weakly to instrument $2$ in the take-up of treatment $2$, while the opposite is true for other agents. It is important to note that Assumption (ref) is a knife-edge condition\textemdash requiring an exact zero partial effect. Small deviations from Assumption (ref) will, however\textemdash in combination with moderate heterogeneous effects\textemdash lead only to a small asymptotic 2SLS bias (see Section (ref)).

Assumptions (ref) and (ref) are necessary and sufficient conditions for 2SLS to assign proper weights. This is our main result:

corThe 2SLS estimand of the effect of treatment 1 assigns proper weights under Assumptions (ref) and (ref) if and only if Assumptions (ref) and (ref) hold.

Informally, 2SLS assigns proper weights if and only if for all agents and treatments $k$, (i) increasing instrument $k$ tends to weakly increase adoption of treatment $k$ and (ii) conditional on instrument $k$, increases in instrument $l\neq k$ do not tend to push the agent into or out of treatment $k$. In Sections (ref) and (ref), we show how our conditions relate to two natural generalizations of the imbens1994identification's monotonicity condition to multiple treatments. In particular, we show that if either HeckmanPinto2018's unordered monotonicity or Kamat2023's joint monotonicity holds, then 2SLS assigns proper weights under an easily testable linearity condition.

Testable Implications of the Identification Conditions

Assumptions (ref) and (ref) can not be directly assessed since we only observe $s\left(z\right)$ for the observed instrument values. But the assumptions do have testable implications. For simplicity, consider the case with unordered treatment effects $D_{k}=\mathbf{1}\left[T=k\right]$.\footnote{See Section (ref) for how Proposition (ref) generalizes to other treatment indicators such as $D_{k}=\mathbf{1}\left[T\geq k\right]$.} We then have the following result:\footnote{We thank the associate editor for suggesting these testable implications. balke1997bounds and heckman2005structural showed similar testable implications of the binary treatment IV model assumptions. sun2020instrument presents similar testable implications of multiple-treatment IV assumptions under multivalued ordered monotonicity and unordered monotonicity.}

propMaintain Assumptions (ref) and (ref) and let $y\leq y'$. If 2SLS with $D_{k}=\mathbf{1}\left[T=k\right]$ assigns proper weights then \[ \operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,\mathbf{1}_{y\leq Y\leq y'}D\right) \] is a non-negative diagonal matrix.\footnote{In fact, $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,\mathbf{1}_{Y\in\mathcal{B}}D\right)$ must be non-negative diagonal for any set $\mathcal{B}\subset\operatorname{supp}\left(Y\right)$. But by the argument of kitagawa2015test, Lemma B.7, it is sufficient to consider sets of the form $\left\{ \left[y,y'\right]\mid y\leq y'\right\} $.}

For instance, for a binary outcome variable $Y\in\left\{ 0,1\right\} $, Proposition (ref) implies that $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,YD\right)$ and $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,\left(1-Y\right)D\right)$ are non-negative diagonal matrices. That $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,YD\right)$ is non-negative diagonal can be tested by running the following regressions:\footnote{Whether $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,\left(1-Y\right)D\right)$ is non-negative diagonal can be tested in an analogous manner by using $\left(1-Y\right)D_{1}$ and $\left(1-Y\right)D_{2}$ as outcome variables.} \[ YD_{1}=\varphi_{1}+\vartheta_{11}P_{1}+\vartheta_{12}P_{2}+\upsilon_{1} \] \[ YD_{2}=\varphi_{2}+\vartheta_{21}P_{1}+\vartheta_{22}P_{2}+\upsilon_{2} \] and test whether $\vartheta_{11}$ and $\vartheta_{22}$ are non-negative and whether $\vartheta_{12}=\vartheta_{21}=0$. Intuitively, if no-cross-effects is satisfied, $P_{2}$ can not increase nor decrease the share of agents with $T=1$ and $Y\left(1\right)=1$, after controlling for $P_{1}$. For a continuous outcome variable, the implications could be tested using the method of mourifie2017testing\textemdash first transforming the Proposition (ref) condition to conditional moment inequalities and then applying chernozhukov2013intersection. In that case, the relevant conditional moment conditions are that \[ \operatorname{Var}\left(P\right)^{-1}\operatorname{E}\left[\left(P-\operatorname{E}\left[P\right]\right)D\mid Y=y\right] \] is non-negative diagonal for all $y\in\operatorname{supp}\left(Y\right)$.\footnote{For $\mathcal{B}\subset\operatorname{supp}\left(Y\right)$, we have

eqnarray*[eqnarray* omitted — 556 chars of source]

} Note that this test is a joint test of Assumptions (ref), (ref), and (ref). If the test rejects, it could be due to a violation of exogeneity or the exclusion restriction.

Further testable implications can be obtained if the researcher has access to “pre-determined” covariates. In particular, let $X\in\left\{ 0,1\right\} $ be a random variable such that $Z\perp\left(X,S\right)$.\footnote{We have already assumed $Z\perp S$ (Assumption (ref)). The condition $Z\perp\left(X,S\right)$ requires, in addition, that $Z$ is independent of the joint distribution of $X$ and $S$. If the instrument $Z$ is truly random, then any variable $X$ that pre-dates the randomization satisfies $Z\perp\left(X,S\right)$. For instance, in a randomized control trial, $X$ can be any pre-determined characteristic of the individuals in the experiment.} Informally, $X$ is a pre-determined variable not influenced by or correlated with the instrument. We then have the following testable prediction.

propMaintain Assumption (ref), and assume $X\in\left\{ 0,1\right\} $ with $Z\perp\left(X,S\right)$. If Assumptions (ref)\textendash (ref) hold then \[ \operatorname{Var}\left(P\mid X=1\right)^{-1}\operatorname{Cov}\left(P,D\mid X=1\right) \] is a diagonal non-negative matrix.

This prediction can be tested by running the following regressions: \[ D_{1}=\gamma_{1}+\eta_{11}P_{1}+\eta_{12}P_{2}+\varepsilon_{1} \] \[ D_{2}=\gamma_{2}+\eta_{21}P_{1}+\eta_{22}P_{2}+\varepsilon_{2} \] for the subsample $X=1$ and testing whether $\eta_{11}$ and $\eta_{22}$ are non-negative and whether $\eta_{12}=\eta_{21}=0$. The predicted treatments $P$ can be estimated from a linear regression of the treatments on the instruments on the whole sample\textemdash a standard first-stage regression. This test thus assesses the relationship between predicted treatments and selected treatments in subsamples.\footnote{The one-treatment version of this test is commonly applied in the literature (e.g., dobbie2018effects,BhullerEtAl2020) and formally justified by FrandsenEtAl2019.} If Assumptions (ref)\textendash (ref) hold, we should see a positive relationship between treatment $k$ and predicted treatment $k$ and no statistically significant relationship between treatment $k$ and predicted treatment $l\neq k$ in (pre-determined) subsamples.\footnote{If the test is applied on various subsamples, $p$-values need to be adjusted to account for multiple testing. It is only meaningful to apply the test on subsamples. The condition is always mechanically satisfied in the full sample.} Note that failing to reject that $\operatorname{Var}\left(P\mid X=1\right)^{-1}\operatorname{Cov}\left(P,D\mid X=1\right)$ is non-negative diagonal across all observable pre-determined $X$ does not prove that 2SLS assigns proper weights, even in large samples: There might always be unobserved pre-determined characteristics $X$ such that $\operatorname{Var}\left(P\mid X=1\right)^{-1}\operatorname{Cov}\left(P,D\mid X=1\right)$ is not non-negative diagonal.\footnote{Similarly, in randomized control trials, showing that treatment is uncorrelated with observed pre-determined covariates does not prove that treatment is randomly assigned.} Also, note that the test is a joint test of Assumptions (ref)\textendash (ref) and the assumption that $X$ is pre-determined. Thus, if the test rejects, the reason might be that $X$ is not pre-determined.

Finally, note that the Proposition (ref) test can also be applied on the subsample $X=1$ if the researcher is willing to maintain that Assumptions (ref)\textendash (ref) hold conditional on $X$. In that case, the Proposition (ref) test can be seen as a special case of the Proposition (ref) test with $y=-\infty$ and $y'=\infty$.

Special Cases

In this section, we first provide general identification results in a model with binary instruments and derive the implied restrictions on choice behavior under ordered or unordered treatment effects. Then, we provide identification results for a standard threshold-crossing model\textemdash an example of an overidentified model with ordered treatment effects.

One Binary Instrument Per Treatment Indicator

Identification Results

The standard application of 2SLS involves one binary instrument and one binary treatment. In this section, we apply the results in Section (ref) to show how this canonical case generalizes to the case with three possible treatments and three possible values of the instrument.\footnote{The results generalize to $n$ possible treatments and $n$ possible values of the instruments; see Section (ref).} We refer to this case as the “just-identified” case: the number of distinct values of the instruments equals the number of treatments.\footnote{In models with homogeneous treatment effects, having the same number of instrument values as treatments gives “just enough” instruments to ensure identification of all model parameters. Note, however, that this “just-identified” model does not identify all parameters under heterogeneous effects, since the number of parameters is then much larger than the number of treatments.} In particular, assume we have an instrument, $V\in\left\{ 0,1,2\right\} $, from which we create two mutually exclusive binary instruments $Z=\left(Z_{1},Z_{2}\right)^{T}$with $Z_{v}\equiv\mathbf{1}\left[V=v\right]$.\footnote{There are other ways of creating two instruments from $V$. For instance, one could define $Z_{1}=\mathbf{1}\left[V\geq1\right]$ and $Z_{2}=\mathbf{1}\left[V=2\right]$. Such a parameterization would give exactly the same 2SLS estimate. We focus on the case of mutually exclusive binary instruments since it allows for an easier interpretation of our results.} We maintain Assumptions (ref) and (ref). This setting is common in applications. For instance, $Z_{v}$ might be an inducement to take up treatment $v$, as considered by BehagheletAl2013. Under which conditions does the multivariate 2SLS assign proper weights in this setting? It turns out that proper identification is achieved only when each instrument affects only one treatment. For simpler exposition, we represent in this section the response types $s$ as functions $s:\left\{ 0,1,2\right\} \rightarrow\left\{ 0,1,2\right\} $ where $s\left(v\right)$ is the treatment selected by response type $s$ when $V$ is set to $v$. Thus, $s\left(v\right)$ is the potential treatment when $Z_{v}=1$.

prop2SLS assigns proper weights in the above model if and only if there exists a one-to-one mapping $f:\left\{ 0,1,2\right\} \rightarrow\left\{ 0,1,2\right\} $ between instruments and treatments, such that for all $k\in\left\{ 1,2\right\} $ and $s\in\mathcal{S}$ either $s_{k}\left(v\right)=0$ for all $v\in\left\{ 0,1,2\right\} $, $s_{k}\left(v\right)=1$ for all $v\in\left\{ 0,1,2\right\} $, or $s_{k}\left(v\right)=1\Leftrightarrow f\left(v\right)=k$.

In words, for each agent $i$ and treatment $k$, either $i$ never takes up treatment $k$, always takes up treatment $k$, or takes up treatment $k$ if and only if $f\left(V\right)=k$. Each instrument $Z_{v}$ is thus associated with exactly one treatment $D_{f\left(v\right)}$. For ease of exposition, we can thus, without loss of generality, assume that the treatments and instruments are ordered in the same way:

assumptionAssume instrument values are labeled such that $f\left(k\right)=k$ for all $k\in\left\{ 0,1,2\right\} $ where $f$ is the unique mapping defined in Proposition (ref).

To see the implications of Proposition (ref), consider the response types defined in Table (ref). It turns out that, for 2SLS to assign proper weights, the population can not consist of any other response type:

table[table omitted — 574 chars of source]
cor2SLS with $D_{k}=\mathbf{1}\left[T=k\right]$ assigns proper weights under Assumption (ref) if and only if all agents are either never-takers, always-1-takers, always-2-takers, 1-compliers, 2-compliers, or full compliers.

BehagheletAl2013 show that these assumptions are sufficient to ensure that 2SLS assigns proper weights. See also kline2016evaluating for similar conditions in the case of multiple treatments and one instrument.\footnote{Equation 1 in kline2016evaluating, generalized to two instruments, also describes the same response types: $s\left(1\right)\neq s\left(0\right)\Rightarrow s\left(1\right)=1$ and $s\left(2\right)\neq s\left(0\right)\Rightarrow s\left(2\right)=2$. The requirement that the instruments can cause individuals to switch only from “no treatment” (the excluded treatment) into some treatment is shared by rose2021recoding's extensive margin compliers only assumption.} Proposition (ref) shows that these conditions are not only sufficient but also necessary, after a possible permutation of the instruments. These response types are characterized by instrument $k$ not affecting treatment $l\neq k$\textemdash no cross effects. Under the response type restrictions of Table (ref), 2SLS identifies the average treatment effect of treatment $1$ for the combined population of 1-compliers and full compliers and the average treatment effect of treatment $2$ for the combined population of 2-compliers and full compliers.\footnote{Thus, in the just-identified case, whenever 2SLS assigns proper weights, it also assigns equal weight to all complier types. This result does not generalize to the overidentified case.} These treatment effects coincide with the treatment effects identified by the method of HeckmanPinto2018.\footnote{The methods proposed by HeckmanPinto2018 can be applied to identify further parameters of interest. In particular, if we denote always-1-takers by $s_{A1}$, always-2-takers by $s_{A2}$, never-takers by $s_{N}$, 1-compliers by $s_{C1}$, 2-compliers by $s_{C2}$, and full compliers by $s_{F}$, their method allows to identify the counterfactuals $\operatorname{E}\left[Y\left(1\right)\mid S=s_{A1}\right]$, $\operatorname{E}\left[Y\left(2\right)\mid S=s_{A2}\right]$, $\operatorname{E}\left[Y\left(1\right)\mid S\in\left\{ s_{C1},s_{F}\right\} \right]$, $\operatorname{E}\left[Y\left(0\right)\mid S\in\left\{ s_{C1},s_{F}\right\} \right]$, $\operatorname{E}\left[Y\left(2\right)\mid S\in\left\{ s_{C2},s_{F}\right\} \right]$, $\operatorname{E}\left[Y\left(0\right)\mid S\in\left\{ s_{C2},s_{F}\right\} \right]$, $\operatorname{E}\left[Y\left(0\right)\mid S\in\left\{ s_{N},s_{C1}\right\} \right]$, and $\operatorname{E}\left[Y\left(0\right)\mid S\in\left\{ s_{N},s_{C2}\right\} \right]$. Moreover, their method allows for the characterization of always-1-takers, always-2-takers, and the following combined populations: 1-compliers and full compliers, 2-compliers and full compliers, 1-compliers and never-takers, and 2-compliers and never-takers.}

Testing For Proper Weighting in Just-Identified Models

By Proposition (ref), there must exist a permutation of the instruments such that instrument $k$ affects only treatment $k$. This result gives rise to more powerful versions of the Section (ref) testable implications applicable in just-identified models:\footnote{It is straightforward to see that the tests become more powerful when $P$ is replaced by $Z$. For instance, if $Y\left(k\right)=0$ for all $k$, $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,\mathbf{1}_{y\leq Y\leq y'}D\right)$ is always non-negative diagonal but $\operatorname{Var}\left(Z\right)^{-1}\operatorname{Cov}\left(Z,\mathbf{1}_{y\leq Y\leq y'}D\right)$ might not. Conversely, if $\operatorname{Var}\left(Z\right)^{-1}\operatorname{Cov}\left(Z,\mathbf{1}_{y\leq Y\leq y'}D\right)$ is non-negative diagonal for all $y$ and $y$', then this must also be true for $\operatorname{Var}\left(Z\right)^{-1}\operatorname{Cov}\left(Z,D\right)$. Since $P\equiv\operatorname{E}\left[D\right]+\operatorname{Var}\left(Z\right)^{-1}\operatorname{Cov}\left(Z,D\right)\left(Z-\operatorname{E}\left[Z\right]\right)$, we must then have that $\operatorname{Var}\left(P\right)^{-1}\operatorname{Cov}\left(P,\mathbf{1}_{y\leq Y\leq y'}D\right)$ is also non-negative diagonal for all $y$ and $y$'.}

propUnder Assumption (ref), Propositions (ref) and (ref) continue to hold when $P$ is replaced by $Z$.

The stronger version of Proposition (ref) reads as follows: If 2SLS assigns proper weights in just-identified models, there must be a permutation of the instruments such that treatment $k$ is affected only by instrument $k$ in first-stage regressions.\footnote{Typically, the researcher would have a clear hypothesis about which instrument is supposed to affect each treatment, avoiding the need to run the test for all possible permutations.} As in Section (ref), this prediction can be tested by running a first-stage regression\footnote{heinesen2022instrumental show that the coefficients from a first-stage regression on the full sample can be used to partially identify violations of “irrelevance” and “next-best” assumptions invoked in kirkeboen2016. Their Proposition 4 implies that $\eta$ is non-negative diagonal if their monotonicity, irrelevance, and next-best assumptions hold. We show that this test is a valid test not only of their invoked assumptions, but also more generally of whether 2SLS assigns proper weights. Moreover, we show that the test can be applied on subsamples.} \[ D=\gamma+\eta Z+\varepsilon \] in a “pre-determined” subsample and test whether $\eta$ is a non-negative diagonal matrix. This test can also be applied in the full sample. The stronger version of Proposition (ref) can be tested in a similar manner.

Choice-Theoretic Characterization

What does the condition in Proposition (ref) imply about choice behavior? To analyze this, we use a random utility model.\footnote{In a seminal article, vytlacil2002independence showed that the imbens1994identification monotonicity condition is equivalent to assuming that agents' selection into treatment can be described by a random utility model where agents select into treatment when a latent index crosses a threshold. In this section, we seek to provide similar characterizations of the condition in Proposition (ref).} Assume that response type $s$'s indirect utility from choosing treatment $k$ when $V=v$ is $I_{ks}\left(v\right)$ and that $s$ selects treatment $k$ if $I_{ks}\left(v\right)>I_{ls}\left(v\right)$ for all $l\neq k$.\footnote{Since all agents of the same response type have identical behavior, it is without loss of generality to assume that all agents of a response type have the same indirect utility function. We assume response types are never indifferent between treatments.} The implicit assumptions about choice behavior differ according to which treatment effects we seek to estimate. We here consider the cases of ordered and unordered treatment effects.

\paragraph*{Unordered Treatment Effects.}

What are the implicit assumptions on choice behavior when $\beta=\left(Y\left(1\right)-Y\left(0\right),Y\left(2\right)-Y\left(0\right)\right)^{T}$? In this case, instrument $k$ can affect only treatment $k$. Thus, instrument $k$ can not impact the utility of treatment $l\neq k$ in any way that changes choice behavior. Also, for instrument $k$ to impact treatment $k$ without impacting treatment $l$, it must be that instrument $k$ can change treatment status only from treatment $0$ to treatment $k$ and vice versa. Thus, essentially, Proposition (ref) requires that only instrument $k$ affects the indirect utility of treatment $k$, and that the excluded treatment always is at least the next-best alternative:\footnote{lee2020filtered consider a similar random utility model. Their Additive Random Utility Model in combination with strict one-to-one targeting and the assumption that all treatments except $T=0$ are targeted gives $I_{ks}\left(v\right)=u_{ks}+\mu_{k}\mathbf{1}\left[v=k\right]$ for constants $u_{ks}$ and $\mu_{k}>0$. As shown by lee2020filtered, these assumptions are generally not sufficient to point-identify local average treatment effects. Our result shows that\textemdash in the just-identified case\textemdash local average treatment effects can be identified if we additionally assume that all agents have the same next-best alternative.}

prop2SLS with $D_{k}=\mathbf{1}\left[T=k\right]$ assigns proper weights if and only if there exists $u_{ks}\in\mathbb{R}$ and $\mu_{ks}\geq0$ such that $s\left(v\right)=\arg\max_{k}I_{ks}\left(v\right)$ with\footnote{The assumption that $\arg\max_{k}I_{ks}\left(z\right)$ is a singleton implies that preferences are strict.} \[ I_{0s}\left(v\right)=0 \] \[ I_{ks}\left(v\right)=u_{ks}+\mu_{ks}\mathbf{1}\left[v=k\right] \] \[ I_{ks}\left(v\right)>I_{0s}\left(v\right)\Rightarrow I_{ls}\left(v\right)<I_{0s}\left(v\right) \] for all $s\in\mathcal{S}$, $v\in\left\{ 0,1,2\right\} $, $k,l\in\left\{ 1,2\right\} $ with $l\neq k$.

The excluded treatment is, thus, always “in between” the selected treatment and all other treatments in this selection model. Since next-best alternatives are not observed, there are selection models consistent with 2SLS assigning proper weights where other alternatives than the excluded treatment are occasionally next-best alternatives. But when indirect utilities are given by $I_{0s}=0$ and $I_{ks}=u_{ks}+\mu_{ks}\mathbf{1}\left[v=k\right]$, this can only happen when $s$ always selects the same treatment, in which case the identity of the next-best alternative is irrelevant.\footnote{See the proof of Proposition (ref).}

figure[figure omitted — 821 chars of source]

In which settings can a researcher plausibly argue that the excluded treatment is always the best or the next-best alternative? A natural type of setting is when there are three treatments and the excluded treatment is “in the middle”. For example, consider estimating the causal effect of attending “School 1” and “School 2” compared to attending “School 0” on student outcomes. Here, School 0 is the excluded treatment. Assume students are free to choose their preferred school. If School 0 is geographically located in between School 1 and School 2, as depicted in Figure (ref), and students care sufficiently about travel distance, it is plausible that School 0 is the best or the next-best alternative for all students. For instance, student A in Figure (ref), who lives between School 1 and School 0, is unlikely to prefer School 2 over School 0. Similarly, student B in Figure (ref) is unlikely to have School 1 as her preferred school. If we believe no students have School 0 as their least favorite alternative and we have access to random shocks to the utility of attending Schools 1 and 2 multivariate 2SLS can be safely applied.\footnote{Another example where the excluded treatment is always the best or the next-best alternative is when agents are randomly encouraged to take up one treatment and can not select into treatments they are not offered. In that case, one might estimate the causal effects of each treatment in separate 2SLS regressions on the subsamples not receiving each of the treatments. But when control variables are included in the regression, estimating a 2SLS model with multiple treatments can improve precision.}

kirkeboen2016 and a literature that followed exploit knowledge of next-best alternatives for identification in just-identified models. In the Section (ref), we show that the conditions invoked in this literature are not only sufficient for 2SLS to assign proper weights but also essentially necessary.\footnote{The kirkeboen2016 assumptions can be relaxed by allowing for always-takers.} Thus, to apply 2SLS in just-identified models with arbitrary heterogeneous effects, researchers have to either directly observe next-best alternatives or make assumptions about next-best alternatives based on institutional and theoretical arguments.

\paragraph*{Ordered Treatment Effects.}

What are the implicit assumptions on choice behavior when $\beta=\left(Y\left(1\right)-Y\left(0\right),Y\left(2\right)-Y\left(1\right)\right)^{T}$? In that case, instrument $k$ can influence only treatment indicator $D_{k}=\mathbf{1}\left[T\geq k\right]$. Thus, instrument $k$ can not impact relative utilities other than $I_{si}-I_{li}$ for $s\geq k$ and $l<k$ in a way that changes choice behavior. Thus, without loss of generality, we can assume that instrument $k$ increases the utility of treatments $t\geq k$ by the same amount while keeping the utility of treatments $t<k$ constant. But this assumption is not sufficient to prevent instrument $k$ from influencing treatment indicators other than $D_{k}$. In addition, we need that preferences are single-peaked:

defn[Single-Peaked Preferences] Preferences are single-peaked if for all $s\in\mathcal{S}$, $z\in\mathcal{Z}$, and $k,l,r\in\left\{ 1,\dots,n\right\} $ \[ k>l>r\text{ and }I_{ks}\left(v\right)>I_{ls}\left(v\right)\Rightarrow I_{ls}\left(v\right)\geq I_{rs}\left(v\right), \] \[ k<l<r\text{ and }I_{ls}\left(v\right)<I_{rs}\left(v\right)\Rightarrow I_{ks}\left(v\right)\geq I_{ls}\left(v\right). \]
figure[figure omitted — 699 chars of source]

To see this, consider Figure (ref). The solid dots indicate the indirect utilities of agent $i$ over treatments 0\textendash 2 when instrument 1 is turned off ($V=0$). In this case, the agent's preferences has two peaks: one at $k=0$ and another at $k=2$. If instrument $1$ increases the utilities of treatments $k\geq1$ by the same amount, the agent might change the selected treatment from $T=0$ to $T=2$ as indicated by the open dots. This change impacts both $D_{1}$ and $D_{2}$. In Figure (ref), however, where preferences are single-peaked, a homogeneous increase in the utility of treatments $k\geq1$ can never impact any other treatment indicator than $D_{1}$. We thus have the following result.

prop2SLS with $D_{k}=\mathbf{1}\left[T\geq k\right]$ assigns proper weights if and only if there exists $u_{ks}\in\mathbb{R}$ and $\mu_{ks}\geq0$ such that $s\left(v\right)=\arg\max_{k}I_{ks}\left(v\right)$ with \[ I_{0s}=0 \] \[ I_{ks}\left(v\right)=u_{ks}+\mu_{vs}\mathbf{1}\left[k\geq v\right] \] for all $s\in\mathcal{S},$ $v\in\left\{ 0,1,2\right\} $, and $k\in\left\{ 1,2\right\} $ and the preferences are single-peaked.

A Threshold-Crossing Model

In several settings, treatments have a clear ordering, and assignment to treatment could be described by a latent index crossing multiple thresholds. For instance, treatments could be grades and the latent index the quality of the student's work. When estimating returns to education, treatments might be different schools, thresholds the admission criteria, and the latent index the quality of the applicant.\footnote{The model below applies in this setting if applicants agree on the ranking of schools and schools agree on the ranking of candidates.} In the context of criminal justice, treatments might be conviction and incarceration as in Humphries2022 or years of prison as in rose2021does and the latent index the severity of the crime committed. If the researcher has access to random variation in these thresholds, for instance through random assignment to graders or judges who agree on ranking but use different cutoffs, then the identification of the causal effect of each consecutive treatment by 2SLS could be possible.

We here present a simple and easily testable condition under which 2SLS assigns proper weights in such a threshold-crossing model. Our model is a simplified version of the model considered by Heckman2007EconometricII, Section 7.2.\footnote{Heckman2007EconometricII focus on identifying marginal treatment effects and discuss what 2SLS with a multivalued treatment identifies in this model, but do not consider 2SLS with multiple treatments. lee2018identifying consider a similar model with two thresholds where agents are allowed to differ in two latent dimensions.} We assume agents differ by an unobserved latent index $U$ and that treatment depends on whether $U$ crosses certain thresholds.\footnote{This specification is routinely applied to study ordered choices greene2010modeling.} In particular, assume there are thresholds $g_{1}\left(Z\right)<g_{2}\left(Z\right)$ such that \[ T=

cases0 & if U<g_{1}\left(Z\right)\\ 1 & if g_{1}\left(Z\right)\leq U<g_{2}\left(Z\right)\\ 2 & if U\geq g_{2}\left(Z\right)

\] We are interested in the ordered treatment effects $\beta=\left(Y\left(1\right)-Y\left(0\right),Y\left(2\right)-Y\left(1\right)\right)^{T}$, using as treatment indicators $D_{k}=\mathbf{1}\left[U\geq g_{k}\left(Z\right)\right]$ for $k\in\left\{ 1,2\right\} $. Assume the first stage is correctly specified:

assumption(First Stage Correctly Specified.) Assume that $\operatorname{E}\left[D_{k}\mid Z\right]$ is a linear function of $Z$ for all $k$.

This assumption is always true, for instance, when $Z$ is a set of mutually exclusive binary instruments. We maintain Assumptions (ref) and (ref). We then have:\footnote{A similar result is found in a contemporaneous work by Humphries2022 in the context of the random judge IV design. They suggest that the required linearity condition between $P_{1}$ and $P_{2}$ can be relaxed by running a 2SLS specification with $D_{1}$ as a single treatment indicator and $P_{1}$ as the instrument while flexibly controlling for $P_{2}$ (and vice versa). They further show that if treatment assignment depends on several unobserved latent indices\textemdash instead of just one\textemdash 2SLS does not, in general, assign proper weights.}

prop2SLS assigns proper weights in the above model if $E\left[P_{k}\mid P_{l}\right]$ is linear in $P_{l}$ for all $k,l\in\left\{ 1,2\right\} $.

Thus, when treatment can be described by a single index crossing multiple thresholds, and we have access to random shocks to these thresholds, multivariate 2SLS estimands of the effect of crossing each threshold will be a positively weighted sum of individual treatment effects. The required linearity condition between predicted treatments can easily be tested empirically. In Proposition (ref), we show that 2SLS does not assign proper weights if $E\left[P_{k}\mid P_{l}\right]$ is non-linear for one $k$ and $l$ and there is a positive density of agents at all values of $U\in\left[0,1\right]$. A linear relationship between predicted treatments is thus close to necessary for 2SLS to assign proper weights in the above model.

While an exact linear relationship between $P_{1}$ and $P_{2}$ is unrealistic, small deviations from linearity are unlikely to lead to a large 2SLS bias: Deviations from linearity will lead some agents to be pushed out of treatment $2$ by $P_{1}$ and others to be pushed into treatment $2$ by $P_{1}$. If the deviations from linearity are “local”\textemdash small irregularities in an otherwise linear relationship\textemdash the agents pushed into treatment $2$ will have similar $U$s as those pushed out of treatment $2$. If agents with similar $U$s tend to have similar treatment effects, the bias will be small. But large global deviations from linearity have the potential to induce significant 2SLS bias in the presence of heterogeneous effects. If such non-linearities are detected, one solution is to follow the suggestion of Humphries2022: Run 2SLS with one endogenous variable $D_{1}$, $P_{1}$ as instrument, and control non-linearly for $P_{2}$. In Section (ref), we discuss what linearity between $P_{1}$ and $P_{2}$ means in a real application and show how it can be tested using data from BhullerEtAl2020.

If the Conditions Fail

Our identification results show that if either Assumption (ref) or (ref) is violated, then 2SLS does not assign proper weights. If this turns out to be the case, the researcher has two choices: Either impose assumptions on treatment effect heterogeneity or select an alternative estimator.

Assumptions on Treatment Effect Heterogeneity

As we show in Section (ref), 2SLS assigning proper weights is equivalent to 2SLS being weakly causal blandhol2022tsls\textemdash giving estimates with the correct sign\textemdash under arbitrary heterogeneous effects. But correctly signed estimates could also be ensured by imposing assumptions on treatment effect heterogeneity. For instance, if treatment effects do not systematically differ across response types or if the amount of “selection on gains” and the violations of Assumptions (ref) and (ref) are both moderate, 2SLS could still be weakly causal. See Section (ref) for a formal analysis.

Alternative Estimators

The econometrics literature has proposed several ways to identify multiple treatment effects when 2SLS fails. The most general method is provided by HeckmanPinto2018. Their method allows the researcher to learn which treatment effects are identified\textemdash and how they are identified\textemdash for any given restriction on response types. The identification results apply to all settings with discrete-valued instruments\textemdash also in cases where 2SLS is not valid. For instance, if we allow for the presence of the response type $\left(s\left(0\right),s\left(1\right),s\left(2\right)\right)=\left(2,1,2\right)$ in addition to the response types in Table (ref) in the Section (ref) model, 2SLS no longer assigns proper weights, but the method of HeckmanPinto2018 still recovers causal effects. HeckmanPinto2018 and pinto2022beyond discuss how to use revealed preference analysis to restrict the possible response types.\footnote{For instance, pinto2022beyond combines revealed preference analysis with a functional form assumption to identify causal effects in an RCT with multiple treatments and non-compliance.} In a related approach, lee2020filtered show how assuming that certain instrument values target certain treatments can lead to partial identification of treatment effects.

Another important strand of methods to identify multiple treatment effects relies on continuous instruments and the marginal treatment effects framework brought forward by heckman1999local,heckman2005structural. In the case of ordered treatment effects, Heckman2007EconometricII show how separate treatment effects can be identified if treatment is determined by a single latent index crossing multiple thresholds and the researcher has access to shocks to each threshold. Using this method, separate treatment effects can be recovered in the Section (ref) model also when predicted treatments are not linearly related.\footnote{Intuitively, the causal effect of treatment $1$ can be obtained by varying $P_{1}$ while keeping $P_{2}$ fixed. The 2SLS requires $E\left[P_{1}\mid P_{2}\right]$ to be linear since instead of keeping $P_{2}$ fixed it controls linearly for $P_{2}$.} Another advantage of the methods based on marginal treatment effects is that they allow for the calculation of treatment effects that are more policy-relevant than the weighted average produced by 2SLS. In the context of unordered treatments, Heckman2007EconometricII and HeckmanEtal2008_Annales show that analogous assumptions can recover the causal effect of a given treatment versus the next-best alternative.\footnote{A 2SLS version of this approach would be to run 2SLS with a single binarized treatment indicator. The recovered treatment effect\textemdash the effect of receiving a given treatment compared to a mix of alternative treatments\textemdash might sometimes be a parameter of policy interest. HeckmanPinto2018 extend this result to discrete instruments. In particular, they show that the causal effect of a given treatment versus the next-best alternative is identified under unordered monotonicity.} In their framework, causal effects between two specified treatments can be identified by focusing on instrument realizations for which the probability of taking up all other treatments is zero\textemdash if such instrument realizations exist.\footnote{For instance, in the application considered in Section (ref), one could identify the causal effect of incarceration versus conviction by studying cases assigned to judges that never acquit\textemdash if such judges exist. The 2SLS equivalent of this approach is to run 2SLS on the subsample of judges that never acquit.} lee2018identifying show how knowledge of the exact threshold rules can be used to identify the effect of one treatment versus another treatment without relying on such “identification at infinity” arguments.\footnote{kamat2023identification show how treatment effects can be partially identified in a similar model with discrete instruments.} Mountjoy2019 provides a method using continuous instruments where knowledge of the exact threshold rules is not required either, and identification is achieved by two relatively weak assumptions: Partial unordered monotonicity and comparative compliers.\footnote{Mountjoy2019 considers a case with three treatments and two continuous instruments. Partial unordered monotonicity requires that instrument $k$ weakly increases up-take of treatment $k$ and weakly reduces up-take of all other treatments for all agents. Partial unordered monotonicity and our Assumptions (ref) and (ref) do not nest each other. On the one hand, partial unordered monotonicity allows for heterogeneity in how instrument $k$ induces agents out of treatments $l\neq k$ in ways that would violate our no-cross-effect condition. On the other hand, average conditional monotonicity allows monotonicity to be violated for some instrument pairs in ways that would violate partial unordered monotonicity. A main advantage of partial unordered monotonicity, however, is that it has a clear economic interpretation, which Assumptions (ref) and (ref) lack. Comparative compliers requires that those induced from treatment $l$ to treatment $k$ by a marginal increase in instrument $k$ are similar to those induced from treatment $k$ to treatment $l$ by a marginal increase in instrument $l$. This condition is satisfied in a broad class of index models.}

A final alternative to 2SLS is to explicitly model the selection into treatment using the method of heckman1979sample (see, e.g., kline2016evaluating). This approach, however, requires the researcher to make distributional assumptions about unobservables.

Application: The Effects of Incarceration and Conviction

In this section, we show how our results can be applied in practice. Following Humphries2022 and Kamat2023, we consider the identification of the effects of conviction and incarceration on defendant recidivism in a random judge IV design.\footnote{Humphries2022 and Kamat2023 also consider alternative approaches beyond 2SLS.} We first discuss the conditions required for 2SLS to assign proper weights in this setting and then test these conditions using data from BhullerEtAl2020.

When Does 2SLS Assign Proper Weights?

Assume the possible treatments are \[ T\in\left\{ 0,1,2\right\} =\left\{ \text{acquittal},\text{non-incarceration conviction},\text{incarceration}\right\} \] We consider the treatment indicators $D_{1}=\mathbf{1}\left[T\geq1\right]$ (conviction) and $D_{2}=\mathbf{1}\left[T=2\right]$ (incarceration). We thus seek to separately identify the effect of conviction versus acquittal and the effect of incarceration versus conviction. As instruments $Z$, we use randomly assigned judges.\footnote{Formally, $Z$ is a vector of binary judge indicators where $Z_{k}=1$ if the case is assigned judge $k$. In practice, following BhullerEtAl2020, we use as our instruments the leave-one-out incarceration and conviction rates of the assigned judge, calculated across all randomly assigned cases to each judge, excluding the focal case.} The predicted treatments $P_{1}$ and $P_{2}$ then equal the rate at which the randomly assigned judge convicts and incarcerates defendants, respectively.

figure[figure omitted — 1,516 chars of source]

When does 2SLS assign proper weights in this setting? Applying the Section (ref) model, we get that 2SLS assigns proper weights if judges agree on how to rank cases but use different cutoffs for conviction and incarceration and there is a linear relationship between the judges' incarceration and conviction rates.\footnote{In the notation of Section (ref), $U$ would represent the “strength” of the case, $g_{1}\left(Z\right)$ the conviction cutoff for the randomly chosen judge, and $g_{2}\left(Z\right)$ the incarceration cutoff.} In the Figure (ref) example, the relationship between judges' incarceration and conviction rates is linear.\footnote{In Figure (ref), the relationship between $P_{1}$ and $P_{2}$ is exactly linear. As noted in Section (ref), small local deviations from linearity are unlikely to lead to a large 2SLS bias. But large global non-linearities could cause significant 2SLS bias. For instance, it would be problematic if judges with medium conviction rates tend to have a high incarceration rate while judges with low or high conviction rates tend to have a low incarceration rate. We do not find such non-linearities in Figure (ref) for BhullerEtAl2020.}

It is, however, not necessary to assume that judges agree on how to rank cases for 2SLS to assign proper weights.\footnote{The single-index model has been criticized as an excessively restrictive model of judge behavior Humphries2022,Kamat2023. In particular, Humphries2022 rejects the single-index model in their setting by showing that the observable characteristics of those with $T=1$ change when holding $P_{1}$ constant and varying $P_{2}$. Under the single-index model, this should not happen.} An example of a response type not allowed by the single-index model but satisfying Assumptions (ref)\textendash (ref) is given in Figure (ref). Here, there are nine judges with their conviction rate ($P_{1}$) on the $x$-axis and their incarceration rate ($P_{2}$) on the $y$-axis. After controlling linearly for the conviction rate (the blue line), the judges convicting the defendant do not, on average, have a higher nor lower incarceration rate than the judges acquitting the defendant.\footnote{Graphically, the sum of the dotted red lines equals the sum of the green lines.} Thus, by Corollary (ref), the 2SLS estimand of the effect of incarceration places zero weight on this defendant's effect of conviction.\footnote{In Figure (ref), the 2SLS estimand also places a positive weight on the defendant's effect of incarceration: After controlling linearly for the conviction rate, the judges incarcerating the defendant tend to have an above-average incarceration rate. Thus, Assumption (ref) (“average conditional monotonicity”) is satisfied. The 2SLS estimand of the effect of conviction versus acquittal also assigns proper weights.} In fact, Assumptions (ref)\textendash (ref) admit a wide range of response types beyond the Figure (ref) response type. For instance, the Figure (ref) response type also satisfies Assumptions (ref)\textendash (ref).

However, Assumption (ref) (“no cross effects”) is a knife-edge condition that can easily be violated by small deviations from the allowed response types. For instance, the response type in Figure (ref) violates no cross effects.\footnote{Even after controlling linearly for the conviction rate, the judges incarcerating the defendant have a lower-than-average conviction rate. Graphically, the sum of the dotted red lines is higher than the sum of the green lines. The 2SLS estimand of the effect of incarceration will thus place a non-zero (negative) weight on the causal effect of convicting this defendant.} One way to see why “no cross effects” must be violated for either the Figure (ref) response type or the Figure (ref) response type is that the two response types are heterogeneous in their relative responses to the instruments. Instrument $2$\textemdash the judge's incarceration rate\textemdash has the same tendency to push the two response types into conviction.\footnote{The covariance between $P_{2}$ and $s_{1}\left(Z\right)$ is exactly the same.} But instrument $1$\textemdash the judge's conviction rate\textemdash has a higher tendency to push a Figure (ref) agent than a Figure (ref) agent into conviction. Thus, the relative responses to the two instruments are not homogeneous. By Proposition (ref), “no cross effects” is violated for at least one of the response types. Note, however, that if response types do not deviate much from the allowed response types and heterogeneous effects are moderate, the bias would still be small; see Section (ref).

Finally, we note that “average conditional monotonicity” (Assumption (ref)) is satisfied in all of Figures (ref)\textendash (ref). For this condition to be violated, it must be that the relationship between the judge's incarceration rate and whether the agent is incarcerated is flipped for certain cases. The Figure (ref) response type violates average conditional monotonicity: After controlling for the conviction rate, the judges incarcerating the agent have, on average, a lower incarceration rate than the other judges. While unlikely to be widespread, such violations might happen. For example, assume that\textemdash conditional on their conviction rate\textemdash judges with higher incarceration rates are more “conservative”. If conservative judges tend to be less likely to incarcerate defendants for gun law violations, average conditional monotonicity might be violated in gun law cases.

Applying Our Tests

To illustrate the application of the Section (ref) tests, we use data from BhullerEtAl2020. That paper used a sample of 31,428 criminal cases that were quasi-randomly assigned to judges in Norwegian courts from 2005\textendash 2009 and provided an IV estimate of the effect of incarceration on the five-year recidivism rate for defendants. Using exactly the same sample, we consider the causal effects of conviction and incarceration on the same recidivism outcome. Our 2SLS estimates are provided in Table (ref), where Column (1) replicates the baseline single treatment 2SLS estimate in BhullerEtAl2020 (for comparison, see Table 4, Column (3), on p. 1297 in that paper), while Column (2) provides our multivariate 2SLS estimates of conviction and incarceration treatments.\footnote{While not their main focus, BhullerEtAl2020 also considered multivariate 2SLS specifications with alternative non-mutually exclusive treatments (see Appendix Tables B10 and B12 in that paper).}

We now test our conditions using this data. Let $Y$ be an indicator for the defendant's five-year recidivism outcome, as used in the estimation above. The Proposition (ref) test, which we generalize to ordered treatment indicators in Section (ref), requires that $\vartheta_{11},\psi_{11},\vartheta_{22},\psi_{22}\geq0$, $\vartheta_{22}+\vartheta_{12}=\psi_{22}+\psi_{12}=0$, and $\vartheta_{21}=\psi_{21}=0$ in the regressions

eqnarray*[eqnarray* omitted — 386 chars of source]

where $\left\{ \varphi_{1X},\dots,\varphi_{4X}\right\} $ are court-by-year fixed effects.\footnote{In Proposition (ref), we show that the Proposition (ref) test extends to the case with control variables.} The results from these regressions are presented in Table (ref), Panel A.\footnote{Note that there is a mechanical relationship between the coefficients in Columns (1)\textendash (2) and Columns (3)\textendash (4) in Panel A. We get $\psi_{11}=1-\vartheta_{11}$, $\psi_{12}=-1-\vartheta_{12}$, $\psi_{21}=-\vartheta_{21}$, and $\psi_{22}=1-\vartheta_{22}$. The additional tests in Columns (3)\textendash (4) are, however, not redundant: The standard errors are different, and the conditions $\psi_{11}=1-\vartheta_{11}\geq0$ and $\psi_{22}=1-\vartheta_{22}\geq0$ are not tested in Columns (1)\textendash (2).} While we are unable to statistically reject the required conditions at the 5% level, the $\vartheta_{22}+\vartheta_{12}=0$ and $\psi_{22}+\psi_{12}=0$ conditions are close to rejection ($p$-values at 0.12 and 0.08). Also, while we are unable to reject $\psi_{11}\geq0$, this estimate has a negative sign. These results suggest that in a larger sample we could have concluded that either Assumption (ref), Assumption (ref), or Assumption (ref) is violated.

center[center omitted — 14,266 chars of source]

In Table (ref), Panel B, we apply the subsample test of Proposition (ref) on 12 different subsamples based on the defendant's past employment, past incarceration, age, education, parental status, and the predicted likelihood of incarceration based on pre-determined case and defendant characteristics.\footnote{Ideally, one would select the subsamples using machine learning methods as proposed by farbmacher2022instrument. We leave how to implement this in the multiple treatment setting for future research.} In each subsample, we run the regessions \[ D_{1}=\gamma_{1X}+\eta_{11}P_{1}+\eta_{12}P_{2}+\varepsilon_{1} \] \[ D_{2}=\gamma_{2X}+\eta_{21}P_{1}+\eta_{22}P_{2}+\varepsilon_{2} \] where $\gamma_{1X}$ and $\gamma_{2X}$ are court-by-year fixed effects.\footnote{See Section (ref) for how the Proposition (ref) generalizes to the inclusion of controls. Formally, these regressions are joint tests of Assumption (ref)\textendash (ref) and Assumption (ref).} We can not reject $\eta_{11},\eta_{22}\geq0$ or $\eta_{12}=\eta_{21}=0$ in any subsample. With this test, we are unable to reject Assumptions (ref)\textendash (ref).

Finally, in Figure (ref), we assess whether the relationship between incarceration rates and conviction rates among the judges in our data is linear, as required for 2SLS to assign proper weights in the Section (ref) model.\footnote{When the model includes court-by-year effects, such a linear relationship must exist within each court-by-year cell (Proposition (ref)). We only consider the overall relationship in Figure (ref).} While we can statistically reject that the relationship is exactly linear, the deviations from linearity are surprisingly small. Thus, if judges agree on the ranking of cases but use different cutoffs for conviction and incarceration, any bias in 2SLS due to heterogeneous effects is likely to be negligible.

center[center omitted — 1,852 chars of source]
center[center omitted — 34 chars of source]

Conclusion

Two-stage least squares (2SLS) is a common approach to causal inference. We have presented necessary and sufficient conditions for the 2SLS to identify a properly weighted sum of individual treatment effects when there are multiple treatments and arbitrary treatment effect heterogeneity. The conditions require in just-identified models that each instrument only affects one choice margin. In overidentified models, 2SLS identifies ordered treatment effects in a general threshold-crossing model conditional on an easily verifiable linearity condition. Whether 2SLS with multiple treatments should be used depends on the setting. Justifying its use in the presence of heterogeneous effects requires both running systematic empirical tests of the average conditional monotonicity and no-cross-effects conditions and a careful discussion of why the conditions are likely to hold.

Acknowledgments

singlespaceWe thank Gaurab Aryal, Brigham Frandsen, Magne Mogstad, Elie Tamer, the associate editor, and four anonymous referees for useful comments. Manudeep Bhuller gratefully acknowledges support from the European Research Council\textquoteright s Starting Grant Project No. 757279. We thank Statistics Norway, Norwegian Courts Administration and Norwegian Correctional Services for providing access to their data.