EconBase
← Back to paper

On Recoding Ordered Treatments as Binary Indicators

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

54,227 characters · 6 sections · 71 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

On Recoding Ordered Treatments as Binary Indicators

\doublespacing

titlepage{-1cm} \begin{abstract} \singlespacing Researchers using instrumental variables to investigate ordered treatments often recode treatment into an indicator for any exposure. We investigate this estimand under the assumption that the instruments shift compliers from no treatment to some but not from some treatment to more. We show that when there are extensive margin compliers only (EMCO) this estimand captures a weighted average of treatment effects that can be partially unbundled into each complier group's potential outcome means. We also establish an equivalence between EMCO and a two-factor selection model and apply our results to study treatment heterogeneity in the Oregon Health Insurance Experiment. \end{abstract}

A large literature uses instrumental variables to estimate the effects of ordered treatments such as years of education or duration of health insurance coverage angrist1991does, goldin2019health. In these settings, it is common to recode the endogenous variable into a binary indicator for any treatment, e.g., any college or any insurance card1995, finkelstein2012. angrist1995 argued that doing so is a mistake. In the Local Average Treatment Effect (LATE) framework, the estimand recovered by two-stage least squares (2SLS) is the causal effect of any treatment exposure plus bias generated by intensive margin increases in treatment. This linear combination of effects is difficult to map to potential policies and may fall outside the range of treatment effects possible given the support of the outcome.

Instead, angrist1995 advocated for leaving an ordered endogenous variable unchanged. 2SLS then recovers the “Average Causal Response" (ACR), a weighted average of effects of different doses of treatment across compliers. The ACR, however, is also difficult to interpret heckman_etal2006. The complier groups (i.e., populations defined by their set of potential treatments) that generate it are not mutually exclusive. When treatment effects are potentially non-linear in dosage or heterogeneous---as in, for example, health responses to pharmaceuticals or the influence of unemployment duration on reemployment wages---the ACR may also differ in sign and magnitude from the effects of changes in dosage relevant for policy.

This paper investigates an assumption that significantly simplifies 2SLS analysis of ordered treatments. This restriction requires that the instruments induce units to shift from no treatment to some positive quantity but not from some treatment to more. In a study of the effects of health insurance, for example, the instruments must decrease the likelihood of being uninsured but leave the duration of coverage unchanged for individuals who would have obtained insurance regardless. In other words, the restriction requires that there are “extensive margin compliers only" (EMCO). In settings with one-sided non-compliance katz2001moving,kling2007experimental,heller2017thinking---meaning individuals with one value of the instrument all receive no treatment---EMCO holds automatically. More generally, given the frequent use of recoded endogenous variables in applied research aizer2015, Bhuller_etal2018, Norris2020,finkelstein2012,Arteaga2020, we view understanding the assumptions that can justify doing so as important.

Under EMCO, recoding an ordered treatment into an indicator is no longer a mistake. 2SLS estimates of the effect of “any treatment" recover a weighted average of treatment effects for mutually exclusive groups of compliers. Each group is shifted to a different positive quantity of treatment from no treatment. The estimand averages the effects for each group of receiving this quantity vs. no treatment. Weights on each group are easy to recover. Moreover, under EMCO treated means for each complier group are identified, as well as the average of untreated means across all groups, allowing for a partial unbundling of the estimand. While average treatment effects for each complier group are not identified, they can be bounded. This makes it possible to test whether the data are consistent with certain hypotheses, such as that all complier groups or doses have positive treatment effects on average. Bounds can be tightened using shape restrictions motivated by the setting or economic theory, as in recent work on partial identification CSS2018.

The power of EMCO comes from the restrictions it places on choice behavior. In the spirit of vytlacil2002, we show that taken together the standard LATE assumptions and EMCO imply that choices are rationalized by a two-step selection model where units first decide whether to participate in treatment at all and then pick treatment levels. EMCO requires the instrument affect relative utility in the first step and not the second. Importantly, the two-step selection process implies that at least two distinct latent factors govern treatment choices. While the presence of two factors makes marginal treatment effect analysis more complex than the single dimension usually considered heckman2010JEL, two represent a substantial dimension reduction relative to what is generated by LATE alone, which implies selection is governed by possibly as many latent factors as levels of treatment.\footnote{For example, if treatment $D$ falls in $\{0,1,2,\dots,\bar{D}\}$, then LATE alone is equivalent to assuming there are $\bar{D}$ separate selection equations with distinct latent factors vytlacil2006.} This dimension reduction explains why some quantities such as complier means are identified under EMCO, but not otherwise.

In related work, eckhoff2018instrument investigate restrictions that justify “binarizing" an ordered treatment at a given threshold. EMCO is a special case that binarizes treatment around zero.\footnote{While we focus on the extensive margin, analogous results would apply when the instrument shifts individuals exclusively from any given level of treatment to more (e.g., from finishing high school to completing at least some college). eckhoff2018instrument's results, on the other hand, apply when the instrument shifts individuals exclusively from below a given level of treatment to more (e.g., from high school or less to at least some college).} Our results complement eckhoff2018instrument by proving identification of complier means in cases where EMCO makes binarization appropriate. Moreover, we show that EMCO has clear implications for choice modeling that can help guide marginal treatment effect analysis heckman1999,heckman2005 in a setting with ordered or multiple treatments. Our results thus connect the assumptions behind binarizing approaches to the literature on the identification of complier means imbens1997,abadie2003semiparametric and latent factor representations of LATE models vytlacil2002,vytlacil2006, heckman2018unordered. Our results also relate to marshall2016coarsening, who studies binarization when ruling out the exclusion restriction violations due to shifts in treatment above or below the threshold that make binarization problematic. Our assumptions, on the other hand, concern the effect of instruments on treatment but leave the effects of treatment on outcomes unrestricted.

We conclude by examining the plausibility and implications of EMCO in finkelstein2012's analysis of the Oregon Health Insurance Experiment (OHIE), which randomized low-income individuals' access to Medicaid. Details of the experiment make EMCO highly likely to hold in the OHIE. To be eligible to enroll, for example, participants randomized into treatment had to be uninsured for at least six months, which implies non-compliance is one-sided. We use EMCO's identifying power to unpack treatment effects across complier groups. The results reveal clear patterns of intensive-margin adverse selection in the experimental data. Participants induced to remain on Medicaid the longest have the highest levels of healthcare utilization and worst self-reported health.

Setting and notation

Consider a setting with a single binary instrument $Z_i \in \lbrace 0, 1 \rbrace$ and a discrete, ordered treatment $D_i \in \lbrace 0,1,..., \bar{D} \rbrace$. Let $D_i(z)$ denote the treatment status of individual $i$ when $Z_i=z$. Observed treatments are $D_i = D_i(1) Z_i + D_i(0) (1-Z_i)$. Let $Y_i(d)$ denote the potential outcome of interest under treatment status $d$. Observed outcomes are $Y_i = \sum_{d=0}^{\bar{D}} 1(D_i = d) Y_i(d)$. Assume that $Z_i$ satisfies the standard assumptions of the LATE framework imbens1994 and its extension to ordered treatments angrist1995:

ass{(LATE framework)} \begin{align*} & (i) \; \ensuremath{\mathbb{E} \left[ D_i|Z_i=1 \right]} > \ensuremath{\mathbb{E} \left[ D_i|Z_i=0 \right]} \quad (relevance) \\ & (ii) \; (Y_{i}(0), Y_i(1), \dots Y_i(\bar{D}), D_i(1), D_i(0)) \ \rotatebox[origin=c]{90}{$\models$} \ Z_i \quad (exogeneity and exclusion) \\ & (iii) \; D_i(1) \geq D_i(0) \quad \forall i \quad \text{(monotonicity)} \end{align*}

angrist1995 show that under these assumptions the Wald estimand recovers the Average Causal Response (ACR):

align[align omitted — 363 chars of source]

where $\omega_d = \frac{\Pr(D_i(1) \geq d > D_i(0))}{\sum_{k=1}^{\bar{D}} \Pr(D_i(1) \geq k > D_i(0)) }$.

The ACR captures a weighted average of effects of exposure to different “doses" of treatment (i.e., $\ensuremath{\mathbb{E} \left[ Y_i(d)-Y_i(d-1) \right]}$) for potentially overlapping sets of compliers.\footnote{The researcher must sometimes take a stand on the correct discretization of treatment. Since time in schooling, for example, might be most accurately thought of as continuous, considering “months" or “years" of education is necessarily an approximation. In other settings, such as a drug trial where varying doses are administered in discrete units, no approximation is necessary.} While the ACR captures a well-defined causal parameter, it does not correspond to a clear treatment manipulation or policy counterfactual heckman_etal2006. When treatment effects are non-linear or heterogeneous, the ACR may in fact differ in sign and magnitude from policy-relevant causal effects, such as as the average effect of a particular dosage.

Given the difficulty of interpreting the ACR, a common practice in applied research is to simply ignore the ordered nature of the treatment, recode it as binary, and interpret estimates as capturing the effects of “any" treatment. Research on the effects of incarceration, for example, commonly uses an indicator for any prison sentence as the endogenous variable of interest aizer2015, Bhuller_etal2018, Norris2020. angrist1995 showed that doing so may produce a biased estimator of the ACR, while results from eckhoff2018instrument imply that the recoded endogenous variable model recovers a linear combination of effects for those shifted from zero to some treatment and those shifted from some treatment to more:

propositionLet $\beta_{recoded}$ be the Wald estimand when the endogenous variable is $1(D_i > 0)$. Then under Assumption (ref): \begin{align*} \beta_{recoded} &= \underset{Extensive margin}{\underbrace{\ensuremath{\mathbb{E} \left[ Y_{i}(D_i(1)) - Y_{i}(D_i(0)) | D_i(1)>D_i(0) = 0 \right]}}} \ + \\ &\underset{Intensive margin}{\underbrace{\ensuremath{\mathbb{E} \left[ Y_{i}(D_i(1)) - Y_{i}(D_i(0)) | D_i(1)>D_i(0) > 0 \right]}}} \frac{\Pr(D_i(1)>D_i(0)>0)}{\Pr(D_i(1)>D_i(0)=0)} \nonumber \end{align*}

The estimator $\beta_{recoded}$ captures an unintuitive mixture of effects and thus suffers from similar interpretation issues as the ACR.\footnote{angrist1995's result is that $\beta_{\text{recoded}} = \beta_{\text{ACR}} \cdot (1+\kappa)$ where $\kappa \equiv \frac{ \sum_{l=2}^{\bar{D}} \Pr(D_i(1) \geq l > D_i(0)) }{ \Pr(D_i(1) \geq 1 > D_i(0)) }$. The result in Proposition (ref) is a direct implication of Equations (3.7) and (3.8) in eckhoff2018instrument, so we omit a proof. Proposition (ref) also appears in a working paper rose2018does, which includes results from this manuscript and rose_shemtov2019does.} Moreover, because the estimand is a linear combination and not an average, $\beta_{recoded}$ may fall outside the range of treatment effects physically possible given the support of $Y_i$. Intuitively, the primary issue is that the instrument is no longer excludable after recoding. Some individuals' outcomes may change even though their treatment status ($1(D_i > 0)$) does not. For example, when the binarized treatment is any indicator for any prison sentence, the exclusion restriction may be violated if the instrument shifts individuals from no incarceration to some prison sentence and also lengthens prison sentences for those who would have been incarcerated regardless.

These issues can be avoided if one is willing to rule out the existence of problematic complier types, an assumption we call “extensive margin compliers only":

ass{Extensive Margin Compliers Only (EMCO)} \begin{align*} D_i(1) > D_i(0) \Rightarrow D_i(0) = 0 \quad \forall i \end{align*}

Assumption (ref) requires that the instrument only causes some individuals to switch from no treatment to some, and not from some to more. This assumption will be automatically satisfied in any setting with one-sided non-compliance---i.e., where individuals with $Z_i=0$ must have $D_i=0$---but can also hold in more general settings where instruments induce some individuals to take up treatment but do not affect the relative utility of non-zero treatment levels.

EMCO implies the second term in Proposition (ref) disappears and 2SLS using $1(D_i>0)$ as the endogenous variable recovers the average effect of the treatment on individuals shifted from no exposure to some positive amount.

small\begin{align*} \beta_{recoded} &\equiv \frac{ \ensuremath{\mathbb{E} \left[ Y_{i} |Z_i = 1 \right]} - \ensuremath{\mathbb{E} \left[ Y_{i} |Z_i = 0 \right]} }{ \ensuremath{\mathbb{E} \left[ 1(D_i>0)|Z_i = 1 \right]} - \ensuremath{\mathbb{E} \left[ 1(D_i>0)|Z_i = 0 \right]}} \\ &= \ensuremath{\mathbb{E} \left[ Y_{i}(D_i(1)) - Y_{i}(D_i(0)) | D_i(1)>D_i(0) = 0 \right]} \\ &= \sum_{d=1}^{\bar{D}} \omega_d^{r} \ensuremath{\mathbb{E} \left[ Y_{i}(d) - Y_{i}(0) | D_i(1)=d, D_i(0) = 0 \right]} \nonumber \end{align*}

where $\omega_d^{r} = \frac{\Pr(D_i(1)=d, D_i(0) = 0)}{\sum_{l=1}^{\bar{D}} \Pr(D_i(1)=l, D_i(0) = 0)}$.

Our discussion so far has focused on the binary instrument case. In settings with multiple or multi-valued instruments, if EMCO applies to a comparison between two points of support in the value of the instruments, then $\beta_{recoded}$ can be estimated and interpreted treating those two points of support like the binary instrument case. If there are multiple such EMCO-compatible comparisons, the pairwise estimates can be averaged either manually or, as suggested in angrist1995, using 2SLS. Their Theorem 2 shows that 2SLS with an ordered treatment and multiple mutually orthogonal binary instruments recovers a weighted average of ACRs for pairs of comparisons in the support of the instrument. If EMCO holds for each of these comparisons, then 2SLS using the recoded endogenous variable will thus also capture a weighted average of the $\beta_{recoded}$ estimand for each of these pairs.

Our discussion so far has also abstracted from covariates. If Assumptions (ref) and (ref) hold only conditional on some observed $X_i$, then $\beta_{recoded}$ can be estimated for each value in the support of $X_i$ and averaged. Alternatively, a researcher may wish to use 2SLS while controlling for a saturated set of indicators for values of $X_i$ and using the interactions of $Z_i$ and these indicators as instruments. As shown in angrist1995, doing so produces a potentially different weighted average of the $X_i$-conditional estimates of $\beta_{recoded}$. Using linear controls that are not necessarily saturated, on the other hand, requires justifying the implicit parametric structure, as discussed in blandhol2022tsls.

Although EMCO restricts counterfactual choices, it also implies multiple necessary conditions must hold on the joint distribution of the outcome, treatment, and instruments. In the Online Appendix, we provide visual and formal tests of these restrictions that build on the growing literature testing instrument validity kitagawa2015test, Huber_Mellace2015, mourifie2017testing, frandsen2019judging, norris2019examiner. These tests extend results for binary treatments from balkepearl97 and heckman2005 and are based on the observation that outcome densities must be non-negative for all complier groups. They also require that the instrument not decrease the mass of individuals at any positive level of treatment, as is implied by EMCO's restrictions on compliance patterns. Both sets of restrictions can be tested using tools from the moment inequality literature CCK2018, bai2019practical.

EMCO is attractive because when it holds, $\beta_{\text{recoded}}$ has a clear interpretation as a weighted average of causal effects for mutually exclusive complier populations with weights proportional to their population size. $\beta_{ACR}$ also retains a valid causal interpretation if EMCO holds, albeit one less intuitive than $\beta_{recoded}$. Under EMCO, the complier groups summed over in $\beta_{ACR}$ are simply defined by $1\{D_i(1) \geq d, D_i(0) = 0\}$ for each $d \geq 1$, and thus remain potentially overlapping.

Treatment effect non-linearity or heterogeneity can make the magnitude and sign of treatment effects for each complier group in $\beta_{recoded}$ differ. We next show, however, that levels of each complier group's treated potential outcomes are identified and that bounds can be placed on each group's treatment effects.

Advantages of EMCO: Complier means and bounds

The EMCO assumption places strong restrictions on the data generating process (DGP); however, in settings in which it is satisfied, it also provides meaningful advantages. In addition to giving $\beta_{recoded}$ a clear causal interpretation, EMCO identifies other interesting quantities. imbens1997 and abadie2003semiparametric show that when the treatment is binary (i.e., $D_i \in \lbrace 0,1 \rbrace$), compliers' treated and untreated mean potential outcomes are identified by 2SLS regressions using $Y_i D_i$ or $Y_i (1-D_i)$, respectively, as the outcome and $D_i$ or $(1-D_i)$ as the endogenous variable. When the treatment has multiple levels, it is tempting to try to estimate complier means using $Y_i 1(D_i=d)$ as the outcome. However, doing so without assuming EMCO yields a mixture of outcome means for multiple groups. To see why, note that the reduced form effect of $Z_i$ on this outcome is:

align[align omitted — 763 chars of source]

Hence changes in $Y_i1(D_i=d)$ due to $Z_i$ reflect individuals both moving into $D_i=d$ from multiple sources (extensive- and intensive-margin shifts) and moving into higher levels of treatment (intensive-margin shifts). Clearly Equation (ref) when rescaled by the first stage would not yield a meaningful potential outcome mean.

However, EMCO implies that $\Pr(D_i(1)=d>D_i(0)>0)$ and $\Pr(D_i(1)>D_i(0)=d)$ are both zero for $d \geq 1$. Hence treated potential outcome means for each group of “$d$-type" compliers (i.e., individuals with $D_i(1)=d>0=D_i(0)$) are identified, as well as an average of potential outcomes under no treatment. The following proposition formalizes this claim:

propositionIf Assumptions (ref) and (ref) hold, then: \begin{align*} (i) \quad \frac{\ensuremath{\mathbb{E} \left[ Y_i 1(D_i=0)|Z_i=1 \right]}-\ensuremath{\mathbb{E} \left[ Y_i 1(D_i=0)|Z_i=0 \right]}}{\ensuremath{\mathbb{E} \left[ 1(D_i=0)|Z_i=1 \right]}-\ensuremath{\mathbb{E} \left[ 1(D_i=0)|Z_i=0 \right]}} &= \ensuremath{\mathbb{E} \left[ Y_i(0)|D_i(1) > D_i(0)=0 \right]} \\ &= \sum_{d=1}^{\bar{D}} \omega_d^{r} \ensuremath{\mathbb{E} \left[ Y_i(0)|D_i(1)=d>D_i(0)=0 \right]} \nonumber \end{align*} and for any $d > 0$ such that $\ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=1 \right]} - \ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=0 \right]} > 0$: \begin{align*} (ii) \quad \frac{\ensuremath{\mathbb{E} \left[ Y_i 1(D_i=d)|Z_i=1 \right]}-\ensuremath{\mathbb{E} \left[ Y_i 1(D_i=d)|Z_i=0 \right]}}{\ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=1 \right]}-\ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=0 \right]}} &= \ensuremath{\mathbb{E} \left[ Y_i(d)|D_i(1)=d>D_i(0)=0 \right]} \end{align*}

We illustrate how these results can generate additional insights below using data from the Oregon Health Insurance Experiment. Proposition (ref) can be thought of as an extension of the results in imbens1997 and abadie2003semiparametric for the binary case to multi-valued ordered and unordered treatments. In Appendix (ref), we present a more general version of Proposition (ref) that is analogous to Theorem 3.1 in abadie2003semiparametric and implies that functions and means of covariates $X_i$ for each group of $d$-type compliers can also be identified, e.g.:

small\begin{align} \frac{\ensuremath{\mathbb{E} \left[ X_i 1(D_i=d)|Z_i=1 \right]}-\ensuremath{\mathbb{E} \left[ X_i 1(D_i=d)|Z_i=0 \right]}}{\ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=1 \right]}-\ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=0 \right]}} &= \ensuremath{\mathbb{E} \left[ X_i|D_i(1)=d>D_i(0)=0 \right]} \; \forall d > 0 \end{align}

Unfortunately, Proposition (ref) (i) does not identify each $\ensuremath{\mathbb{E} \left[ Y_i(0)|D_i(1)=d>D_i(0)=0 \right]}$ without additional assumptions when $\omega_d^r <1 \ \forall \ d$. Intuitively, when $Z_i$ shifts individuals from $D_i(0) = 0$ to multiple positive levels of treatment, only a weighted average of untreated means across all complier groups is identified. We cannot separately identify the untreated counterfactual for each group of $d$-type compliers because there are more unknowns than equations unless there is only one complier group (which implies $\omega_d^r = 1$ for some $d$). Thus, while treated $d$-type complier means are identified, $d$-type treatment effects are generally not:

align*[align* omitted — 401 chars of source]

Consequently, $\beta_{recoded}$ cannot be fully decomposed into its constituent causal components. The researcher can, however, construct bounds on $d$-type treatment effects.

Specifically, let $Y_d^0$ be the unknown quantity $\ensuremath{\mathbb{E} \left[ Y_i(0)|D_i(1)=d>D_i(0)=0 \right]}$. By the above, $Y_d^d =\ensuremath{\mathbb{E} \left[ Y_i(d)|D_i(1)=d>D_i(0)=0 \right]}$ is point identified for all $d > 0$ where $\ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=1 \right]} - \ensuremath{\mathbb{E} \left[ 1(D_i=d)|Z_i=0 \right]} > 0$. Complier group shares $\omega_d^{r}$ are likewise identified by the ratio of $\Pr(D_i = d|Z_i=1) - \Pr(D_i=d | Z_i =0)$ to $\Pr(D_i > 0|Z_i=1) - \Pr(D_i > 0 | Z_i =0)$. Bounds on $d$-type treatment effects are given by the solution to the linear program:

align[align omitted — 287 chars of source]

where $\mathcal{Y}$ is the support of $Y_i$ and $\mathrm{conv}\left(\mathcal{Y}\right)$ denotes the convex hull of $\mathcal{Y}$.

Given that the unknown quantities $\{Y_d^0\}_{d=1}^{\bar{D}}$ are disciplined by only two sets of restrictions, these bounds are likely to be wide without further assumptions.\footnote{These bounds are not necessarily sharp. As a result of Proposition (ref), for example, the full distribution of $Y_i(0)$ for the population with $D_i(1) > D_i(0)$, denoted $F_0$, is also identified. The unknown means $Y_d^0$ must also be consistent with a distribution of $Y_i(0)$ for each complier type, denoted $F_d^0$, such that $F_0 = \sum_{d=1}^{\bar{D}} \omega_d^r F_d^0$.} Imposing other shape restrictions, such as that treatment effects are decreasing in $d$, can help tighten bounds in this case. Inference can be conducted using methods that are suited to situations where the standard bootstrap fails, such as fang2019inference or hong2020numerical.

Rather than bounding individual complier groups' treatment effects, it may also be interesting to test whether the data are consistent with certain joint hypotheses, such as that average treatment effects for all complier groups are weakly positive. Positive average treatment effects requires that $Y_d^0 \leq Y_d^d$ for all $d > 0$. Hence testing this hypothesis is equivalent to asking whether there exists a set of $Y_d^0$ such that:

align[align omitted — 231 chars of source]

Inference can be conducted by viewing the problem as a shape constrained generalized method of moments problem and applying methods developed by chernozhukov2015constrained.

\FloatBarrier

Implications for choice behavior

vytlacil2006 shows that the LATE framework laid out in Assumption (ref) is equivalent to a selection model defined by the following assumptions. First, treatment choices are governed by $\bar{D}$ selection equations:

flalign1(D_i \geq d) = 1(C^d(Z_i)-V_{i}^d \geq 0), \ for \ d \in \lbrace 1,\dots \bar{D} \rbrace

where $V_i^d$ are random variables and $C^d$ are unknown functions of the instruments satisfying $C^{d-1}(Z_i) - V_{i}^{d-1} \geq C^{d}(Z_i) - V_{i}^d \quad \forall i,d$.\footnote{As in the rest of the paper, for notational convenience we suppress implicit conditioning on observables $X_i$.}$^{,}$\footnote{Footnote 36 in rose_shemtov2019does describes how the notation in Equation (ref) relates to that in vytlacil2006.} Second, the instrument must induce a monotonic treatment response and be relevant, which requires that either $C^{d}(1) \geq C^{d}(0) \ \forall \ d$ and $\exists \ d \ s.t. \ \Pr\left( C^d(1) \geq V_i^d > C^d(0) \right) > 0$ or $C^{d}(1) \leq C^{d}(0) \ \forall \ d$ and $\exists \ d \ s.t. \ \Pr\left( C^d(1) < V_i^d \leq C^d(0) \right) > 0$. Finally, to complete the model the instrument must be independent of both $V_i^d$ and $Y_i(d)$ for all $d$: $\left( V_i^1,\dots,V_i^{\bar{D}}, Y_i(0), Y_i(1),\dots,Y_i(\bar{D}) \right) \ \rotatebox[origin=c]{90}{$\models$} \ Z_i$.

While equivalent to the LATE assumptions, this selection model is difficult to work with due to the $\bar{D}$ dimensions of the unobserved heterogeneity. EMCO restricts this unobserved heterogeneity sharply. In fact, adding EMCO to the LATE framework is equivalent to assuming a two-step decision making process. In the first step, the individual chooses whether to participate or not. In the second step, the individual chooses the level of treatment. A classic example of such behavior is two-stage budgeting deaton1980economics. This simple two-factor “hurdle" model for treatment choices is defined by the following assumption:

ass{Two-factor choice model} \\ Treatment choices are governed by \begin{align*} 1(D_i = d) &= \begin{cases} & 1( \pi_0(Z_i) - U_i^{Ext} \geq 0 ) \ \ if \ \ d= 0 \\ & 1( \pi_0(Z_i) - U_i^{Ext} < 0 ) 1( \pi_{d+1} \leq U_i^{Int} < \pi_{d} ) \ \ if \ \ d > 0 \end{cases} \end{align*} where $(U_i^{Ext}, U_i^{Int}) \ \rotatebox[origin=c]{90}{$\models$} \ Z_i$, $(U_i^{Ext}, U_i^{Int}) \sim F$ with strictly increasing marginal cumulative distribution functions, $\pi_d \geq \pi_{d+1} \ \forall d$, $Pr(\pi_0(1) < U_i^{Ext} \leq \pi_0(0)) > 0$ or $Pr(\pi_0(0) < U_i^{Ext} \leq \pi_0(1)) > 0$, and $(Y_i(0),\dots,Y_i(\bar{D}),U_i^{Ext}, U_i^{Int}) \ \rotatebox[origin=c]{90}{$\models$} \ Z_i$.

Assumption (ref) describes a two-equation system for treatment choices.\footnote{We use $\pi$s as notation for thresholds because $\pi_0(Z_i) = \Pr(D_i = 0 | Z_i)$ when $U_i^{Ext}$ has a marginally uniform distribution over $[0,1]$. Likewise, when $U_i^{Int}$ is marginally uniform over $[0,1]$, $\pi_{d} - \pi_{d+1} = \Pr(D_i = d | Z_i=1, D_i >0)$, as we show in Appendix (ref).} This model includes only two latent factors: $U_i^{Ext}$, which governs the decision of whether or not to participate, and $U_i^{Int}$, which determines the level of participation. Although all threshold functions $\pi_d$ can depend on covariates $X_i$, only $\pi_0(Z_i)$ is a function of $Z_i$. The restriction that $Pr(\pi_0(1) < U_i^{Ext} \leq \pi_0(0)) > 0$ or $Pr(\pi_0(0) < U_i^{Ext} \leq \pi_0(1)) > 0$ delivers monotonicity and relevance, while requiring that $(Y_i(0),\dots,Y_i(\bar{D}),U_i^{Ext}, U_i^{Int}) \ \rotatebox[origin=c]{90}{$\models$} \ Z_i$ guarantees exogeneity and exclusion.

Proposition (ref) formalizes the equivalence between this model and LATE plus EMCO by showing that they jointly impose the same restrictions on behavior:

propositionAssumptions (ref) and (ref) are equivalent to Assumption (ref).

The equivalence in Proposition (ref) is in the sense of vytlacil2002: the selection model in Assumption (ref) satisfies Assumptions (ref) and (ref). But not only that, Assumptions (ref) and (ref) imply that one can always write down a selection model of the type in Assumption (ref) that rationalizes observed and counterfactual choices. Thus, the model in Assumption (ref) imposes the same restrictions on behavior as those imposed by the combination of the LATE framework assumptions and EMCO.

Because Proposition (ref) shows that EMCO is equivalent to invoking a two-factor selection model, an implication is that “single-index" models commonly used to reduce the dimensionality of unobserved heterogeneity in Equation (ref) Dahl2002,heckman_etal2006,rose_shemtov2019does, kowalski2021reconciling are inconsistent with EMCO except in special cases. In particular, Proposition (ref) shows that compatibility requires that if compliers are shifted to some $d > 1$, no individuals can be assigned positive treatment levels below d when $Z=0$. This requirement rules out the presence of always takers at positive levels of treatment below the maximum level of treatment obtained by treated compliers.

propositionConsider the following single-index model of treatment assignment: \begin{align*} 1(D_i = d) = 1(\pi_{d+1}(Z_i) \leq U_i < \pi_d(Z_i)), \quad U_i \sim Uniform[0,1] \end{align*} where $Z_i \ \rotatebox[origin=c]{90}{$\models$} \ U_i$, $\pi_{d+1}(Z_i) \leq \pi_d(Z_i)$, $\pi_{d}(0) \leq \pi_d(1)$, $\pi_0(Z_i) = 1$, and $\pi_{\bar{D}+1}(Z_i) = 0$. Then compatibility with Assumption (ref) requires that if $\pi_d(0) < \pi_d(1)$, then $\pi_{d}(0) = \pi_{d'}(0) \ \forall \ d' \ s.t. \ 1 \leq d' < d$. In addition, compatibility can be tested by examining whether $$\left(\ensuremath{\mathbb{E} \left[ 1(D_i \geq d)|Z_i=1 \right]} - \ensuremath{\mathbb{E} \left[ 1(D_i \geq d)|Z_i=0 \right]}\right)\left(1-\ensuremath{\mathbb{E} \left[ 1(D_i \geq d)|Z_i=0 \right]} - \ensuremath{\mathbb{E} \left[ 1(D_i = 0)|Z_i=0 \right]}\right) = 0$$ holds for all $d > 1$.

The testable implication contained in Proposition (ref) is straightforward. Single-index compatibility with EMCO requires that for levels of treatment greater than one either $\Pr(D_i(1) \geq d) = \Pr(D_i(0) \geq d)$, so that no compliers are shifted to levels of treatment $\geq d$ (implying the first parenthetical term is zero) or $\Pr(1 \leq D_i(0) < d) = 0$ (implying the second parenthetical term is zero).

Instruments that satisfy EMCO-like behavioral restrictions are commonly used in economics. For example, the need for such instruments arises naturally in labor economics when researchers studying wages seek to correct for the choice to work at all in the style of gronau1974wage and heckman1974shadow. mulligan2008selection use these techniques to estimate the influence of the changing composition of women in the labor force on the gender wage gap. Their instrument---a mother's number of children aged zero to six interacted with marital status---must impact the decision to work but not labor supply among mothers already working. Another example comes from card2005estimating, whose strategy to estimate the wages of individuals induced into working by a time-limited earnings subsidy can be justified by an EMCO restriction on the effects of the subsidy on labor supply.

Implications of EMCO for marginal treatment effect analysis

The inconsistency of EMCO with single-index models makes marginal treatment effect (MTE) analysis heckman1999,heckman2005 more complex. While two factors represent a substantial dimension reduction relative to the $\bar{D}$ implied by LATE alone, modeling treatment effect heterogeneity in two dimensions is significantly more challenging than the single dimension considered in the MTE literature heckman2010JEL and in the ordered treatment cases studied in rose_shemtov2019does.

One simplifying assumption that would restore the validity of standard MTE tools is that potential outcomes are not affected by latent factors that govern selection along the intensive margin:

align[align omitted — 142 chars of source]

However, this assumption precludes selection into treatment based on gains and levels along the intensive margin except through correlation between $U_i^{Ext}$ and $U_i^{Int}$.

An alternative approach is to model potential outcomes as functions of both $U_i^{Ext}$ and $U_i^{Int}$. The researcher can then estimate or bound other treatment effects of interest (such as an average treatment effect) consistent with the moments identified under EMCO. Specifically, let $m_d(u_1,u_2) = \ensuremath{\mathbb{E} \left[ Y_i(d) | U_i^{Ext} = u_1, U_i^{Int} = u_2 \right]}$ represent treatment response functions. $d$-type compliers' potential outcome means are given by:

small\begin{align} & \ensuremath{\mathbb{E} \left[ Y_i(d)|D_i(1)=d>0=D_i(0) \right]} = \\ &\quad = \int_{\pi_{d+1}}^{\pi_d} \int_{\pi_0(0)}^{\pi_0(1)}m_d(u_1,u_2)dF(u_1,u_2) / \int_{\pi_{d+1}}^{\pi_d}\int_{\pi_0(0)}^{\pi_0(1)}dF(u_1,u_2) \nonumber \end{align}

The researcher can then pick explicit functional forms for treatment response functions or flexibly approximate them in the style of mogstad2018using and marx2020sharp. Each identified complier mean serves to discipline these functions.

\FloatBarrier

Empirical application: The effects of health insurance

In 2008, a group of low-income adults in Oregon were randomly given the opportunity to apply for Medicaid. finkelstein2012 use this experiment---dubbed the Oregon Health Insurance Experiment (OHIE)---to study the effects of access to Medicaid on health care utilization and financial and physical well-being. They find that insurance increases both primary and emergency care utilization, lowers some health care expenditures, and increases self-reported physical and mental health.

To analyze the experiment, the researchers primarily use 2SLS specifications with “ever on Medicaid" as the endogenous variable. However, they note that the treatment in their setting---duration of Medicaid coverage---is continuous, and argue that coding the treatment as the “`number of months on Medicaid' may be more appropriate than `ever on Medicaid' where the effect of insurance on the outcome is linear in the number of months insured." While a continuous endogenous variable would be appropriate regardless of linearity in the effects of insurance, a binary endogenous variable for “any Medicaid" is also appropriate in this setting because EMCO is highly likely to hold for institutional reasons. The OHIE randomized admission into the Oregon Health Standard Plan, which had been closed to enrollment since 2004. Subjects lotteried into treatment were able to enroll and remain on the plan so long as they were eligible, which required being uninsured for at least six months prior to enrolling and ineligible for other public health insurance programs.\footnote{Candidates also had to be 19–64 years of age, US citizens or legal immigrants, have income under 100 percent of the federal poverty level, and possess assets of less than \$2,000.} Individuals who would have obtained some Medicaid in the control group therefore could not increase their duration of coverage if lotteried into treatment, since they would be ineligible.\footnote{It is possible some subjects would have obtained Medicaid through other means if not lotteried into treatment. However, the data in this case are consistent with EMCO holding nevertheless. Moment inequality tests of EMCO's restrictions (see Appendix (ref) for details) support its applicability to the OHIE.}

\ifdefined\something

figure[figure omitted — 1,615 chars of source]

\else \fi

Tests of EMCO described in the Online Appendix support its applicability to the OHIE. Appendix Figure (ref), for example, shows that individuals lotteried into Medicaid ($Z_i=1$) were much less likely than controls to have zero months of insurance over the follow-up period. They are more likely, however, to have coverage for all positive durations, with particularly large spikes around 6-7 months and 12-13 months. EMCO's restrictions on treatment choices require that there are no decreases in density at any positive level of coverage.\footnote{Formal tests of the restrictions described in the Online Appendix show that we cannot reject at the 5% significance level the null hypothesis that the data are consistent with EMCO ($T_n=3.46$; critical value $3.96$).}

EMCO gives the finkelstein2012 results a simple causal interpretation. The two percentage point increase in hospital admissions due to Medicaid presented in their Table IV, for example, reflects increases relative to how often subjects would have been admitted without any insurance. EMCO also allows the researcher to further decompose treatment effects across complier populations. Table (ref) illustrates this using survey data from finkelstein2012 on self-reported health and health care utilization. Since there are many survey questions relevant for these outcomes, we create standardized indices by averaging normalized answers to seven health questions and four utilization questions (as in Table V and IX in finkelstein2012). Column 1 shows 2SLS estimates of the effect of any Medicaid on these indices. Consistent with the original results, Medicaid increases both utilization and health significantly. The former rises by 0.1 standard deviations, and the latter by 0.2.

Under EMCO, all compliers in the OHIE are shifted from zero months of Medicaid to some positive amount. Column 2 shows the average untreated outcomes for all compliers. Compliers appear to be significantly negatively-selected on health---untreated means are 0.15 standard deviations below the sample average. Utilization is slightly higher than the sample average, but not significantly so. Columns 3-6 then show treated mean outcomes for compliers shifted to 7-12 months of Medicaid, 13-16 months, etc. No individuals were shifted to 1-6 months of coverage. The final row of the table reports the share of all compliers each group comprises. Roughly 33% of compliers, for example, are shifted from zero months to 7-12 months.

These means several interesting patterns. For example, all complier groups with positive density have higher treated health than the average under no insurance. But compliers induced to remain on Medicaid for the longest have the worst treated health. This is consistent with the adverse selection patterns noted in finkelstein2012 and kowalski2021reconciling---only the sickest remain on Medicaid for long.\footnote{For example, comparing OLS to 2SLS estimates, the authors note that “differences suggest that at least within a low-income population, individuals who select health insurance coverage are in poorer health (and therefore demand more medical care) than those who are uninsured, just as standard adverse selection theory would predict."} Compliers who remain on Medicaid for 7-12 months, by contrast, have better than average treated health. The 2SLS estimate in Column 1 is the population share-weighted average of means in Columns 3-5 minus the average in Column 2. As noted above, means under no insurance for each d-type complier group are not identified. The data are consistent with any pattern of treatment effects where the weighted average of complier means under no insurance equals the means in Column 2. It is possible, for example, that treatment effects are negative for compliers induced to stay on Medicaid the longest. Any negative effects for this group, however, would have to be outweighed by positive effects for others.

Utilization follows a similar pattern. Individuals induced to remain on Medicaid for the longest also have the highest levels of treated utilization. Thus the sickest compliers are also the most expensive. Utilization results also suggest that increases in utilization due to Medicaid are not due to one-shot “pent-up" demand, since utilization is increasing in duration on Medicaid. As with health outcomes, the data are consistent with a range of treatment effects on utilization for each complier group.

Taken together, the results show the power of EMCO in a setting where it is institutionally plausible. Not only do the simple 2SLS regressions presented in finkelstein2012 have a coherent causal interpretation, but identification of complier means reveal important insights into selection patterns.

\ifdefined\something

table[table omitted — 3,675 chars of source]

\else \fi

\FloatBarrier

Conclusion

2SLS estimates of the effects of ordered treatments (e.g., years of education) can be difficult to interpret. The estimand is a weighted average of causal effects for overlapping sets of compliers and may differ from the estimand of interest. In some settings, an auxiliary assumption that the instruments only induce units to switch from zero units of treatment to some positive amount (EMCO) may be reasonable. Under EMCO, 2SLS estimates of the effect of a recoded indicator for any treatment capture average treatment effects for mutually exclusive compliers groups shifted into varying units of treatment. Treated means for each complier type can be recovered using standard techniques, as well as the average untreated mean, allowing for a partial decomposition of the estimand. An application to data from the Oregon Health Insurance Experiment reveals clear patterns of adverse selection into Medicaid.

When should researchers invoke EMCO? While there are testable necessary implications of EMCO, in practice the assumption's validity is likely to hinge on the institutional details of the experiment at hand. Along with the standard LATE assumptions, invoking EMCO is equivalent to assuming the data is generated by a two-stage selection process where individuals first decide whether to participate in treatment at all and then pick treatment levels. EMCO requires that the instruments only affect the utility of any participation, but not the relative utility of positive treatment levels. In many settings, such as those with one-sided non-compliance only, it may be clear a priori whether this is reasonable. When EMCO does hold, researchers should invoke it explicitly to justify their choice of models and make use of the additional identifying power it provides.

\singlespacing \onehalfspacing

\FloatBarrier

table[table omitted — 1,554 chars of source]

\FloatBarrier