Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,747 characters · 25 sections · 79 citation commands
Instrument-based estimation of full treatment effects with movers
\begingroup \footnote{ We thank Massimo Anelli, Akanksha Negi, Denni Tommasi, and Dinand Webbink for valuable comments. The authors have no relevant or material financial interests that relate to the research described in this paper. All omissions and errors are our own. } \addtocounter{footnote}{-1} \endgroup
}
\thispagestyle{empty}
\linespread{1.30} \setcounter{page}{1}
This paper develops an instrumental variable (IV) framework that partially identifies the local average treatment effect (LATE) of receiving full treatment compared to no treatment. The effect of the full treatment program is a primary parameter of interest in policy evaluation. However, in many instances, policy evaluations yield estimates of only a subset of the treatment program. For instance, when the effect of college completion may be of interest, the effect of college enrollment is estimated. This is a problem for two main reasons. First, most treatment programs are designed to be received in full and the effects of separate treatment parts may be less informative. Second, there is often non-compliance with randomized treatment assignment and hence identification of treatment effects in an IV framework requires an exclusion restriction. This restriction may be violated by individuals moving in or out of treatment after the start of the program.
Individuals that move in or out of treatment are present in many policy evaluation settings. For instance, leu2017beginning reports that approximately 30% of college students changes major, sugar2021medicaid describe the high prevalence of Medicaid coverage disruptions, often referred to as churning, due to income fluctuations, and heckman2000substitution discuss the incidence of control group substitution and treatment group dropout across various job market training programs. This implies that instead of staying in or out of treatment during the whole program, there are late-adopters missing the first part of treatment or dropouts missing the final part of treatment.
The IV exclusion restriction rules out certain types of movers depending on the definition of the treatment variable. For instance, silliman2022labor, grosz2020returns, burde2013bringing, and hoekstra2009effect take enrollment at the start of an educational program as treatment indicator. In this case, the exclusion restriction does not allow for movers that are induced by the instrument to only take the second part of the treatment. ketel2016returns and zimmerman2014returns use enrollment at the end of an educational program as treatment indicator. In this case, the exclusion restriction does not allow for movers that are induced by the instrument to only take the first part of the treatment.
Researchers are familiar with the challenging restrictions imposed by the exclusion restriction. For instance, silliman2022labor estimate the returns to vocational secondary education enrollment, and discuss how student dropout may violate the exclusion restriction. grosz2020returns discusses how movers in and out of a nursing program affect the exclusion restriction with different definitions of treatment, such as enrolling at the start or ever enrolling, in the nursing program. finkelstein2012oregon suggest that different definitions for the treatment variable of medicaid insurance can be used to provide a lower and upper bound for a LATE using the largest and smallest first stage, respectively.
This paper partially identifies the LATE of receiving full treatment compared to no treatment under a double exclusion restriction. This parameter is referred to as the local average full treatment effect (LAFTE). Our IV framework includes the binary potential treatment status for the first and second part of treatment with only one binary instrument. The instrument may induce individuals to obtain full treatment, or only the first or second part of the treatment. We refer to the latter group of individuals as movers. Within this framework, movers do not violate the IV exclusion restriction and LATEs are only identified under additional assumptions. The proposed double exclusion restriction extends this exclusion restriction --the instrument can only affect the outcome variable through treatment enrollment-- with the condition that the instrument can only affect treatment enrollment in the second part through enrollment in the first part. We provide a procedure for testing necessary conditions of the presence of movers and of the double exclusion restriction.
The partial identification of the LAFTE is achieved with nonparametric sharp bounds. The bounds formalize the intuition that different definitions for the treatment variable can be utilized to identify the effect of receiving a full treatment program, while allowing for movers. We show that the bounds can be tightened using the monotone treatment response and monotone treatment selection assumptions of manski1997monotone and manski2000monotone in addition to the double exclusion restriction. Less informative bounds rely solely on double exclusion and a bounded response.
The double exclusion restriction holds if a delayed treatment response to treatment assignment is absent, as it rules out the movers that are induced by the instrument to only take the second part of treatment. It now follows that the common practice of using IV with enrollment in the first part of treatment as treatment indicator identifies the LATE of taking the first part of the treatment. Our double exclusion restriction is similar to, albeit weaker than, the dynamic exclusion restriction of angrist2022marginal. The dynamic exclusion restriction imposes the additional condition that treatment enrollment in the first part can only affect the outcome variable though enrollment in the second part. This either rules out movers, restricts treatment effect heterogeneity, or a combination of these two. In this case, LAFTE can be point-identified using the IV method of imbens1994identification.
The policy relevance of the LAFTE is illustrated with four empirical applications from different fields of applied economics: a health program by finkelstein2012oregon,finkelstein2016effect, a labour program by wheeler2022linkedin, an educational program by anelli2020returns, and a development program by burde2013bringing. For instance, finkelstein2016effect use a randomized opportunity to apply for Medicaid as an instrument for Medicaid enrollment. They find that half a year of Medicaid enrollment increases the probability of an emergency department visit by 0.088 percentage points for the compliers. However, the short-term impact of Medicaid on health care utilization may be small, and hence policymakers make efforts to reduce Medicaid coverage disruptions sugar2021medicaid. Indeed, our estimated LAFTE shows that continuous Medicaid enrollment for two years increases the probability of an emergency department visit between 0.081 and 0.224 percentage points for the compliers. This suggests that policies encouraging continuous enrollment result in larger effects on health care utilization. We discuss similar policy implications in the other applications.
The applications also underline the empirical applicability of our methods. The four studies all fit our IV framework, with data available on both enrollment in the first and second part of the treatment. By testing the necessary conditions, we find evidence for the presence of movers in all four studies. In the health program, the movers even establish the majority of the compliers, consistent with the high prevalence of Medicaid coverage disruptions. Except for the development program, the necessary conditions for the double exclusion restriction cannot be rejected. In these three applications, the lower bounds on the LAFTEs are statistically significantly different from zero, and two of the three upper bounds are informative on the size of the LAFTE.
There is an extensive literature that bounds LATEs under violations of the exclusion restriction. For instance, flores2013partial establishes partial identification using one treatment variable. Instead of our approach of using treatment indicators for the first and second part of the treatment, mealli2013using use two outcome variables to construct bounds. conley2012plausibly propose Bayesian approaches for identification under relaxed exclusion restrictions. swanson2018partial provide an overview on treatment bounds for settings with a binary instrument, binary treatment, and a binary outcome. heckman2000substitution bound average treatment effects in the presence of control group substitution and treatment group dropping out, without employing instrumental variables.
The identification of causal parameters with multiple treatment parts has been studied in similar settings. First, we study the identification of the LAFTE by indexing the potential outcomes with the full treatment history. Similar potential outcome models are considered in the dynamic treatment literature. For instance, lechner2009sequential identifies treatment effects using conditional independence assumptions and ding2010estimating employ a structural economic model. blackwell2017instrumental identifies the treatment effects for different types of movers between two sequential treatments with one instrument for each treatment. heckman2016dynamic also identify dynamic treatment effects using multiple instruments. In contrast, our identification approach relies on only one instrument and does not require a conditionally exogenous treatment or an underlying structural model.
Second, instead of estimating LATEs, recent difference-in-differences techniques are used to estimate the intention to treat in the presence of movers. For instance, hull2018estimating identifies mover average treatment effects for the individuals that move in or out of treatment. verdier2020average extrapolates treatment effects for movers to stayers: individuals whose treatment status does not change. Specific types of movers are studied in, for example, athey2022design who only allow for late-adopters in a staggered adoption framework.
Third, instead of estimating the LAFTE, causal mediation analysis identifies the direct and indirect effects of enrollment in the first part of treatment. Our proposed double exclusion restriction imposes that all effects of the instrument on the outcome variable go only through enrollment in the first part of treatment. However, the effect of enrollment in the first part on the outcome can either be direct or via treatment enrollment in the second part of the treatment. Hence, enrollment in the second part can be considered a mediator, and huber2019review provides an overview of estimation methods for direct and indirect effects of enrollment in the first part of treatment.
The outline of this paper is as follows. (ref) defines the causal parameter of interest. (ref) discusses our IV framework with movers. (ref) introduces the double exclusion restriction, partial identification of the LAFTE, and testable necessary conditions for the presence of movers and the double exclusion restriction. (ref) applies the proposed methods to four empirical treatment evaluation settings. (ref) concludes.
Assume a setting in which individuals are randomly assigned to treatment, after which full treatment or only a subset of treatment can be received. The binary instrumental variable $Z \in \{0,1\}$ equals one if an individual is randomly assigned to treatment. The treatment indicator $D_t \in \{0,1\}$ equals one if treatment part $t$ is received, where $t=1$ corresponds to, for instance, college enrollment and $t=2$ to college graduation. When compliance with the treatment assignment is not mandatory, the treatment indicators may take different values than the instrument. After assignment, individuals may obtain full treatment $\{D_1=1,D_2=1\}$ or no treatment $\{D_1=0,D_2=0\}$. Moreover, $D_1$ may be different from $D_2$ in settings that allow for late enrollment into treatment $\{D_1=0,D_2=1\}$ and/or dropping out of treatment $\{D_1=1,D_2=0\}$. Finally, let the variable $Y$ be the observed outcome of interest.
A causal parameter of interest to policy makers is the average treatment effect (ATE) of receiving full treatment $\mathbb{E}[Y(1,1)-Y(0,0)]$, where $Y(D_1,D_2)$ is an individual's potential outcome for values of $D_1$ and $D_2$. This ATE is the average difference between receiving both parts of the treatment ($Y(1,1)$) compared to no treatment at all ($Y(0,0)$). The ATE of the first part of treatment $\mathbb{E}[Y(1,d_2)-Y(0,d_2)]$ or the second part $\mathbb{E}[Y(d_1,1)-Y(d_1,0)]$ may be of interest for the evaluation of which treatment period is most effective. However, most treatments are designed to be received in full. For instance, the first part of a training program builds knowledge that prepares for the second part, and the second part builds upon the knowledge obtained in the first. Hence, the average treatment effect of receiving full treatment is often of primary interest to policy makers.
Since we only observe one $Y$ for each individual and individuals may not comply with their assigned treatment, ATEs are in general not identified. In this case it is common to report local ATEs (LATEs) that can be identified by using the random treatment assignment as an instrumental variable for treatment enrollment. The LATE equals the average treatment effect for the individuals who are induced into treatment by the instrument, and is therefore the ATE for a subpopulation.
The standard IV framework introduced by imbens1994identification defines an IV estimand and shows that it identifies a LATE of a treatment variable $D$ on the outcome of interest $Y$ using the potential outcome framework and four assumptions on the instrumental variable $Z$. The IV estimand is defined as
where $\Delta \mathbb{E}[A|Z]=\mathbb{E}[A|Z=1]-\mathbb{E}[A|Z=0]$. The potential outcome framework links the observed treatment indicator $D$ to the potential treatment status $D(Z)$ as $D=D(1)Z+D(0)(1-Z)$, and the observed outcome $Y$ to the potential outcomes $Y(Z,D)$ as $Y=Y(1,1)ZD+Y(1,0)Z(1-D)+Y(0,1)(1-Z)D+Y(0,0)(1-Z)(1-D)$.
The four instrumental variable assumptions are:
Under (ref), the IV estimand in (ref) identifies a LATE if the treatment can be summarized by a single binary variable $D$: $\beta_{IV} = \mathbb{E}[Y(1)-Y(0)|C]$, where $C=\{D(1)-D(0)=1\}$ defines the individuals induced into treatment by the instrument. These individuals are referred to as compliers.
(ref) shows four ways the treatment indicator $D$ can be constructed when $\{Z,D_1,D_2,Y\}$ is observed. The treatment indicator can be defined as enrollment at the start of the treatment ($D_1$), enrollment at the end of the treatment ($D_2$), received full treatment ($D_\land$), or received at least one part of the treatment ($D_\lor$). The third column in (ref) shows that different definitions of the treatment indicator are used in the applied economics literature.
When the instrument induces individuals to obtain only a subset of the treatment, all four definitions for $D$ may result in violations of the exclusion restriction made by (ref).4. These violations arise when $Z$ affects $Y$ while $D$ remains constant. For instance, if the researcher uses $D=D_1$ and the instrument induces some individuals into only the second part of treatment, $Z$ may affect $Y$ through $D_2$ whereas $D_1$ is not affected. Similarly, if the researcher uses $D=D_2$, the exclusion restriction may be violated if the instrument induces some individuals into only the first part of treatment. An instrument inducing individuals from no treatment at all to only part of the treatment, or an instrument inducing individuals from only part of the treatment to full treatment, may run into problems with $D=D_\land$ or $D=D_\lor$ respectively.
This section extends the standard IV framework to include the potential treatment status at the first and second part of the treatment. Within this framework, five different groups of compliers can be distinguished. The four additional groups of compliers cause the violations of the exclusion restriction described above. We show that in the presence of these four groups, LATEs can only be point identified under additional strict assumptions.
We extend the potential outcome model to include two potential treatments $D_1(Z)$ and $D_2(Z,D_1)$ and a potential outcome $Y(Z,D_1,D_2)$. Potential and observed treatments and outcomes are linked as follows,
(ref) replaces (ref) for the potential outcome framework defined above:
These assumptions hold if $Z$ is randomly assigned ((ref).1), if there are individuals that are induced into full treatment by the instrument ((ref).2), and there are no individuals induced to move out of any part of treatment by the instrument ((ref).3). (ref).4 implies that
which shows that $Y$ is a function of the two treatment parts, as defined in (ref).
In contrast, the exclusion restriction in (ref).4 imposes that $Y(z,d_1,d_2)=Y(d)$ for all $z,d_1,d_2$, which only depends on one summary of the treatment $D=d$. (ref).4 is more restrictive as it does not allow $Z$ to affect $Y$ through both treatment parts, whereas (ref).4 does. For instance, with $D=D_1$, (ref).4 imposes that $Y(z,d_1,d_2)=Y(d_1)$ and therefore $Y$ does not depend on the second treatment part.
The potential outcome model in (ref) implies that there are different types of compliers. Since individuals can comply with the instrument in the first part, the second part, or the full treatment, the group of compliers $C$ consists of five different complier groups: $C=\{\{C_1,C_2\},\{C_1,N_2\},\{C_1,A_2\},\{N_1,C_2\},\{A_1,C_2\}\}$. These groups are defined as follows:
where $C_t$, $N_t$, and $A_t$ refer to respectively compliers, never-takers, and always-takers in treatment part $t$. The first group of compliers $\{C_1,C_2\}$ is induced into full treatment by the instrument, and we refer to this group as full compliers. We define the individuals that are induced to obtain only part of the treatment by the instrument as movers. Hence, there are four types of movers: $\{C_1,N_2\}$ and $\{C_1,A_2\}$ are induced into only the first part of treatment, whereas $\{N_1,C_2\}$ and $\{A_1,C_2\}$ are induced into only the second part of treatment.
Our definition of movers concerns individuals who change treatment status after the start of the program depending on the value of the instrument. The two mover types $\{C_1,N_2\}$ and $\{A_1,C_2\}$ dropout from treatment and miss the second part of the program when $Z=1$ and $Z=0$, respectively. The two mover types $\{C_1,A_2\}$ and $\{N_1,C_2\}$ are late-adopters and miss the first part of the program when $Z=0$ and $Z=1$, respectively. Individuals with $\{N_1,A_2\}$ or $\{A_1,N_2\}$ are not included, as they do not affect the LATE identification. Hence, our mover group is different from the more general definition of movers as the set of all individuals with $D_1 \neq D_2$, as used by, for instance, hull2018estimating in difference-in-differences estimation.
Using the potential outcome framework in (ref) and the definition of the complier group $C$, the local average full treatment effect (LAFTE) can now be expressed as $\mathbb{E}[Y(1,1)-Y(0,0)|C]$.
The following lemma shows what the first stages $\Delta\mathbb{E}[D|Z]$ for different definitions of $D$ and the reduced form $\Delta\mathbb{E}[Y|Z]$ identify when the potential outcome framework takes both treatment parts into account.
The proof is deferred to (ref).
(ref) shows that the first stages represent the proportion of different complier types depending on the definition of $D$. Each first stage captures the proportion of the full compliers plus the proportions of two mover types that comply with the corresponding $D$. For instance, with $D=D_1$, the movers $\{C_1,N_2\}$ and $\{C_1,A_2\}$ comply in the first part of treatment.
The reduced form coefficient $\Delta\mathbb{E}[Y|Z]$ equals a weighted sum of LATEs, in which each weight corresponds to the proportion of one of the five complier types. Each LATE corresponds to a different treatment effect for a complier type. The first term of the reduced form in (ref) corresponds to the LAFTE, but only for the full compliers $\{C_1,C_2\}$.
For each of the definitions of $D$ in (ref), the IV estimand in (ref) identifies a weighted average of LATEs plus a bias term. For instance, for $D=D_1$ we have
with $w(G_1,G_2)=\frac{\mathbb{P}[G_1,G_2]}{\mathbb{P}[C_1,C_2]+\mathbb{P}[C_1,N_2]+\mathbb{P}[C_1,A_2]}$, and where $\{G_1,G_2\}$ can represent any of the five complier groups. The bias terms reflect the effect of $D_2$ on $Y$ while keeping $D_1$ constant. It follows that the IV estimand does not have a clear causal interpretation:
The proof follows directly from (ref).
Figure (ref) visualizes the identification problem, where each arrow represents a possible effect among $\{Z,D_1,D_2,Y\}$ under (ref). The instrument induces the full compliers $\{C_1,C_2\}$ into full treatment, but also induces the mover groups into only a subset of the treatment. The figure suggests that the LAFTE can be identified under two different additional assumptions. The first assumption rules out the presence of all mover groups. The second assumes that the treatment effects for movers are identical to their full treatment effect. (ref) and (ref) below formalize this intuition.
The proof follows directly from (ref). (ref) may suit a setting in which, for instance, treatment completion is mandatory after treatment enrollment and treatment completion is impossible without treatment enrollment. It follows that all compliers are full compliers $C=\{C_1,C_2\}$.
The proof is deferred to (ref). (ref) may suit a setting in which, for instance, $\mathbb{E}[Y(1,d_2)-Y(0,d_2)|G_1,G_2]$ does not depend on $d_2$ for the movers $\{G_1,G_2\}$. This setting implies homogeneous treatment effects within mover types.
The LAFTE can also be identified by a combination of assumptions on the presence of certain mover types and the homogeneity of certain treatment effects. This is formalized by the dynamic exclusion model discussed by angrist2022marginal:
where $\delta$, $\pi$, $\alpha$, $\psi$, $\beta$, and $\mu$ are coefficients, $\eta$, $\xi$, and $\varepsilon$ are error terms, and we exclude additional covariates. From (ref) follows that $D_2$ does not depend on $Z$, and consequently the movers $\{N_1,C_2\}$ and $\{A_1,C_2\}$ in (ref) do not exist. Similarly, (ref) shows that $Y$ does not depend on $D_1$. (ref) shows that this is the case if either $\{C_1,N_2\}$ and $\{C_1,A_2\}$ do not exist, or if $Y(0,d_2)=Y(1,d_2)$ for all $d_2$.
This section proposes a partial identification strategy for the LAFTE. We replace the single exclusion restriction in (ref).4, by a double exclusion restriction:
The first part of (ref) is identical to (ref).4, and the second part states that $Z$ must not have a direct effect on $D_2$ other than through $D_1$. These two parts together impose a double exclusion restriction on $Z$.
Figure (ref) shows that the double exclusion allows for $Z$ to have a direct effect on $D_1$, and for $Z$ to have an effect on $D_2$ through $D_1$. Hence, compliers in the second part of treatment have to be compliers in the first part, and the group of compliers now consists of only three types: $C=\{\{C_1,A_2\},\{C_1,N_2\},\{C_1,C_2\}\}$. (ref) holds under (ref), but does not impose (ref), and hence does not exclude all mover types or imposes homogeneous treatment effects.
The double exclusion restriction imposes that $Z$ randomly assigns individuals to the first part of treatment but not the second part of treatment if the first part stays constant. This holds in settings in which a delayed treatment response to treatment assignment can be ruled out. For instance, in the Medicaid experiment analysed by finkelstein2012oregon, individuals with a lottery draw of $Z=1$ could only obtain Medicaid at the start of the treatment, and hence mover type $\{N_1,C_2\}$ is likely to be absent. Since individuals with $Z=0$ had little to no opportunity to obtain medicaid coverage at the start of the treatment, mover type $\{A_1,C_2\}$ is also likely to be absent. In case the first stage estimate for enrollment at the start of treatment is close to one, for instance in a carefully conducted randomized controlled trial duflo2007using,deree2023closing, the double exclusion may also hold.
Since the double exclusion restriction rules out two complier types, $\{A_1,C_2\}$ and $\{N_1,C_2\}$, the first stages and reduced form in (ref) can be simplified.
The proof is deferred to (ref). (ref) shows that under the double exclusion restriction the proportions of all three complier groups are identified.
The identification of the proportions of the complier groups allows for the partial identification of LAFTE under additional assumptions:
The proof is deferred to (ref). Since the bounds in (ref) collapse to the LAFTE for the full compliers in absence of movers, the bounds are sharp.
Assuming that the LATEs $\mathbb{E}[Y(1,1)-Y(1,0)|C_1,N_2]$ and $\mathbb{E}[Y(0,1)-Y(0,0)|C_1,A_2]$ are nonnegative, the difference between the reduced form of the LAFTE and the reduced form in (ref) is positive. Hence, the latter can be used as a lower bound on the LAFTE. The assumption of positive expected treatment effects have been used by, for instance, flores2013partial for the partial identification of LATEs with a single treatment indicator. manski1997monotone introduces the monotone treatment response (MTR) assumption on the individual level instead of in expectation.
The reduced form in (ref) only differs from the reduced form of the LAFTE by the two potential outcomes $\mathbb{E}[Y(0,0)|C_1,A_2]$ and $\mathbb{E}[Y(1,1)|C_1,N_2]$. The upper bound on the LAFTE can be obtained by bounding these potential outcomes using the following assumptions: The potential outcome $\mathbb{E}[Y(0,0)|C_1,A_2]$ is non-negative, and the potential outcome of obtaining full treatment $Y(1,1)$ is smaller for $\{C_1,N_2\}$ than for $\{C_1,C_2\}$ and $\{C_1,A_2\}$. Outcomes can generally be rescaled so that the assumption of a non-negative potential outcome is harmless, for instance when the response has a lower bound. The second assumption is invoked in expectation, and therefore weaker than the monotone treatment selection (MTS) assumption of manski2000monotone.
A combination of assumptions similar to the ones in (ref) have been used for the partial identification of treatment effects by, for instance, molinari2010missing, dehaan2011effect, and kreider2012identifying. In general, partial identification is common in the analysis of treatment effect identification problems. For instance, kreider2007disability, battistin2011misclassified, tommasi2020bounding, and calvi2022late derive bounds on treatment effects when the treatment is misreported.
When the MTR and MTS assumptions are considered too strong, the following bounds can be derived assuming only a bounded response:
The proof is deferred to (ref).
The bounded response allows us to replace $Y(1,0)$ for $\{C_1,N_2\}$ and $Y(0,1)$ for $\{C_1,A_2\}$ with respectively $Y_{\min}$ and $Y_{\max}$ ($Y_{\max}$ and $Y_{\min}$) to construct a lower (upper) bound on the LAFTE. The difference between the bounds equals $(y_{\max}-y_{\min})(\mathbb{P}[C_1,N_2]+\mathbb{P}[C_1,A_2])/(\mathbb{P}[C_1,C_2]+\mathbb{P}[C_1,N_2]+\mathbb{P}[C_1,A_2])$. The bounds equal $\mathbb{E}[Y(1,1)-Y(0,0)|C_1,C_2]$ when no movers are present, but can be wide when the response bounds are wide and the proportion of movers is large.
A complier group that may be of particular interest are the full compliers $\{C_1,C_2\}$. This group is induced to obtain both parts of the treatment by the instrument. However, with only one instrument, it is not possible to affect the full compliers without also affecting the other complier types. Hence, the full compliers are mostly relevant for settings in which the policymaker has an instrument for both the first and second part of treatment. blackwell2017instrumental shows that with two binary instruments, and an additional treatment exclusion restriction that each instrument only affects its own treatment, the LAFTE for the full compliers can be identified.
Our identification strategy for the LAFTE relies on restrictions on the presence of mover types. This section derives testable necessary conditions for these restrictions.
First, (ref) shows that the IV method introduced by imbens1994identification point-identifies the LAFTE when no movers are present. This assumption results in the following necessary conditions.
The proof is deferred to (ref).
The first two necessary conditions (ref) and (ref) in (ref) can be used to test the null-hypothesis that $\mathbb{P}[C_1,N_2]=\mathbb{P}[A_1,C_2]$ and $\mathbb{P}[N_1,C_2]=\mathbb{P}[C_1,A_2]$. A failure to reject this null-hypothesis does not necessarily imply the absence of movers, as it may also indicate mover types with identical proportions. Hence, in case of a failure to reject (ref) and (ref), the necessary conditions (ref) and (ref) can be considered. These conditions are equal to zero if there are no movers or if potential outcomes are homogeneous.
Provided that potential outcomes are heterogeneous, (ref) provides a two-step testing procedure for the presence of movers. If the null-hypothesis that (ref) and (ref) both equal zero is rejected, movers may be present. If this null-hypothesis cannot be rejected, the null-hypothesis that (ref) and (ref) both equal zero has to be tested. In case of a rejection, we still conclude that there may be movers. However, a failure to reject indicates that there are no movers. So if both sets of necessary conditions cannot be rejected, we recommend using standard IV approaches to identify the LAFTE for the full compliers $\{C_1,C_2\}$.
Second, (ref) shows that in the absence of certain mover types, the LAFTE is partially identified. In particular, the double exclusion restriction in (ref) only allows for $\{C_1,A_2\}$ and $\{C_1,N_2\}$. This has sign implications for (ref) and (ref):
The proof follows directly from (ref).
The sign conditions in (ref) are necessary for the absence of the movers ruled out by the double exclusion restriction. Since it could be the case that $\mathbb{P}[C_1,N_2]>\mathbb{P}[A_1,C_2]>0$ and $\mathbb{P}[C_1,A_2]>\mathbb{P}[N_1,C_2]>0$, the conditions are not sufficient for the double exclusion restriction to hold. Hence, a failure to reject the null-hypothesis that the sign restrictions hold need not imply the double exclusion restriction, but a rejection implies that the double exclusion restriction cannot be invoked.
(ref) provides sharp nonparametric bounds for the LAFTE. However, point-identification may be required or in some empirical settings the assumptions may be deemed too strong. For these cases, we discuss three alternative causal objects that can be identified under weaker assumptions, and from which two can be point-identified.
First, (ref) shows that under (ref).1-(ref).3 and (ref), the IV estimand in (ref) with $D=D_1$ identifies the LATE of taking the first part of the treatment:
with $w[G_1,G_2]=\frac{\mathbb{P}[G_1,G_2]}{\mathbb{P}[C_1,C_2]+\mathbb{P}[C_1,N_2]+\mathbb{P}[C_1,A_2]}$.
The weighted average of LATEs in (ref) can be interpreted as a decomposition of the total effect of $D_1$ on $Y$ using $D_2$ as a mediator. The mediation literature, see huber2019review for a review, defines the direct effect as the effect of $D_1$ on $Y$ while the mediator $D_2$ is constant, and the indirect effect as the effect of $D_2$ on on $Y$ while $D_1$ is constant. Hence, the first term in (ref) is the indirect effect, and the remaining terms are direct effects.
(ref) shows that the IV estimand in (ref) is often used in the literature. This paper shows that this is a valid estimator of the effect of $D_1$ under the double exclusion restriction. In addition, we argue that the effect of receiving full treatment is also a causal parameter of interest. Adding $\mathbb{E}[Y(1,1)-Y(1,0)|C_1,N_2]w[C_1,N_2]$ and $\mathbb{E}[Y(0,1)-Y(0,0)|C_1,A_2]w[C_1,A_2]$ to (ref) results in the effect of receiving full treatment. The additional MTR assumptions in (ref) guarantee that these two average treatment effects of receiving the second part of treatment are positive. It follows that the lower bound on the LAFTE equals (ref) and the LAFTE is equal to or larger than the LATE of taking the first part of the treatment.
Second, the double exclusion restriction may be considered too strong, or its necessary conditions may be rejected for the setting at hand. (ref) shows that under only (ref), the IV estimand in (ref) with the multivalued treatment $D=D_1+D_2 \in \{0,1,2\}$ equals
with $w[G_1,G_2]=\frac{\mathbb{P}[G_1,G_2]}{2\mathbb{P}[C_1,C_2]+\mathbb{P}[C_1,N_2]+\mathbb{P}[N_1,C_2]+\mathbb{P}[C_1,A_2]+\mathbb{P}[A_1,C_2]}$. This results in a causal interpretation of a weighted average of causal effects of receiving one part of the treatment, instead of the causal effect of receiving full treatment.
The IV estimand $\beta_{IV}(D_1+D_2)$ is related to the average causal response (ACR) as introduced by angrist1995two. The ACR is also a weighted average of the effects of unit changes in treatment, but considers a multivalued treatment that does not distinguish between $\{D_1=1,D_2=0\}$ and $\{D_1=0,D_2=1\}$. Without late-adopters who miss the first part of the program, the mover types $\{C_1,A_2\}$ and $\{N_1,C_2\}$ are absent and (ref) is equal to the ACR.
Summarising the multivalued treatment by one binary treatment indicator may also lead to violations of the exclusion restriction similar to the ones discussed in Section (ref). andresen2021instrument show that these violations can also be ruled out by restricting treatment effect heterogeneity or mover types, where movers with a multivalued treatment are defined by the individuals induced by the instrument to change treatment status from and to treatment values that are both below or above the binarisation threshold.
The ACR type estimand in (ref) considers shifting from no treatment to the first part and from the first part to full treatment, as separate treatment effects for the full compliers. Therefore, it double counts the full compliers in the denominator of the weighting function $w[G_1,G_2]$. An alternative is to combine the two effects for the full compliers and to interpret it as the full treatment effect. (ref) allows for the partial identification of this alternative weighted average of LATEs:
The proof is based on (ref) and deferred to (ref). The weights in (ref) are nonnegative and add up to one. Hence $\tau$ is a convex combination of LATEs and has a causal interpretation. However, this interpretation may not directly be policy relevant as it measures an average across different treatment effects and different groups.
Intuitively, the treatment indicator with the largest (smallest) first stage estimate may provide a lower (upper) bound on a treatment effect of interest. For instance, finkelstein2012oregon apply this intuition to bound the effects of Medicaid insurance. (ref) shows that this intuition does not apply to $\tau$: Instead of bounding $\tau$ from below, the treatment indicator with the largest first stage is an upper bound on $\tau$.
This section illustrates the empirical relevance of our methods for treatment evaluation. We consider the LAFTE of a health program, a labour program, an educational program, and a development program. Additional details on the empirical specifications are deferred to (ref). This section focuses on the main results following from (ref) and (ref) and (ref). (ref) shows additional empirical results, such as the bounds in (ref), which are in general wide and cannot reject that the LAFTEs are equal to zero.
This application studies the LAFTE of medicaid coverage on health care utilization. Policy makers actively aim to reduce coverage disruptions by continuous enrollment policies. As a recent example, the Families First Coronavirus Recovery Act requires Medicaid programs to keep individuals enrolled for the duration of the public health emergency sugar2021medicaid. To better understand the impact of such policies, the effect of continuous Medicaid coverage across the full study period is of particular interest.
In 2008, a group of uninsured low-income adults in Oregon was randomly given the opportunity to apply for Medicaid. The state opened a waiting list for 10,000 Medicaid spots and subsequently drew names by lottery from the 89,924 individuals that placed themselves on this list. The Medicaid program provided comprehensive benefits with no consumer cost sharing. The monthly enrollment premiums ranged from \$0 to \$20 depending on income.
finkelstein2012oregon use the randomized opportunity to apply for Medicaid as an instrument for Medicaid enrollment and analyze the effects on health care utilization, financial strain, and health. The Oregon Health Insurance Experiment (OHIE) was subsequently used in a series of papers to study the effects of Medicaid on other outcomes. For example, finkelstein2016effect study the impact on emergency department use over time.
The data in finkelstein2016effect contains Medicaid coverage and emergency department use for 24,646 individuals with Portland-area zip codes across four time periods: day 0 to 180, 181 to 360, 361 to 540, and 541 to 720 after lottery notification. Our $D_1$ equals one if an individual was enrolled in Medicaid in the first year after lottery notification, and $D_2$ equals one if an individual was enrolled in the second year after lottery notification. The instrument equals one if an individual was randomly given the opportunity to apply for Medicaid. The binary outcome variable $Y$ equals one if an individual had any emergency department visit during the two years after the lottery notification.
(ref) reports the estimated necessary conditions for the presence of movers and the validity of (ref). Column (1) and (2) in Panel A show, respectively, the results of a regression from $D_{\lor}-D_2$ and $D_{\land}-D_2$ upon $Z$. First, since both estimates are significantly different from zero at the 1% level, we reject that movers are absent. Second, since both estimates are positive, we do not reject the necessary conditions of (ref).
Under the double exclusion restriction, column (1) implies that $\widehat{\mathbb{P}[C_1,N_2]}=0.078$. This group includes movers who drop out of treatment. Since individuals had to recertify eligibility every six months, they could have lost coverage due to, for instance, income fluctuations. Column (2) implies that $\widehat{\mathbb{P}[C_1,A_2]}=0.060$. These movers likely used the opportunity to gain eligibility in the second time period: 14 months after the lottery, Oregon took lottery draws from a new Medicaid waiting list. Column (5) reports the results of a regression from $D_2$ upon $Z$. Under the double exclusion restriction this estimate equals $\widehat{\mathbb{P}[C_1,C_2]}=0.115$, and hence the movers compose a larger proportion of the data than the full compliers.
The double exclusion restriction imposes that $\mathbb{P}[N_1,C_2]=\mathbb{P}[A_1,C_2]=0$. The first mover type is likely absent since individuals with $Z=1$ had to apply for Medicaid within 45 days after the state made them aware of the lottery draw. Individuals that did not apply during this window could not apply after. In general, finkelstein2012oregon describe that the mechanisms to obtain Medicaid coverage are limited to the initial lottery and the lottery 14 months later. They also show that only 2% of the individuals with $Z=0$ obtained Medicaid coverage up unto one year after the lottery. This may explain the absence of the second mover type.
(ref) shows the estimated bounds for the LAFTE from (ref). Panel A identifies a statistically significant LAFTE: The 95% confidence interval of the lower bound, $[0.035,0.128]$, does not include zero. The bounds rely on the assumption that the effect of Medicaid in the second year on emergency department use is non-negative (MTR), and that always takers and compliers of Medicaid in the second year are not less likely to use the emergency department than never takers (MTS). Based on these bounds, we conclude that Medicaid enrollment during both years increases the probability of an emergency department visit between 0.081 and 0.224 percentage points for the individuals induced to enroll by the lottery.
The estimated bounds imply that the LAFTE may be up to three times as large as the LATE of Medicaid enrollment in the first period. Hence, policy makers may expect larger treatment effects in combination with continuous enrollment requirements, which may inform the design of future Medicaid enrollment policies.
This application studies the LAFTE of a LinkedIn training on employment. Job market training programs, such as a training that stimulates LinkedIn usage, are often characterized by high incidence of control group substitution and treatment group dropout heckman2000substitution. Since the programs are carefully designed to receive and follow in full, the treatment effect of LinkedIn usage for a longer period should be of particular interest to policy makers.
wheeler2022linkedin run and evaluate a randomly assigned program that trains work seekers to join and use LinkedIn. Their study sample includes 30 cohorts from existing job readiness training programs in four large South African cities. They randomly assign 15 cohorts to four hours of LinkedIn training during their job readiness training program. The intervention trains participants, among others, to open accounts, build their profiles, and search and apply for jobs. The study examines the effect of the LinkedIn training on a range of outcomes. Although the experiment is not specifically designed to identify the causal effect of LinkedIn usage, one of the analysis uses the random assignment to LinkedIn training as an instrument for LinkedIn usage to study its effect on employment.
The data includes LinkedIn usage and employment outcomes across three time periods: directly, six months, and twelve months after the end of the job readiness training program. Our $D_1$ equals one if a participant has a Linkedin account at the end of the job training and $D_2$ equals one if a participant has a LinkedIn account six months later. The instrument $Z$ equals one if the participant is randomly assigned to the LinkedIn training. The binary outcome variable $Y$ equals one if the participant is employed twelve months after the program ended. The estimation sample includes 988 of the 1,638 participants across the 30 cohorts.
The estimated necessary conditions in column (1) and (2) of Panel B in (ref) are positive, and the latter estimate is also significantly different from zero at the 1% level. Hence, we reject that movers are absent and do not reject the necessary conditions of (ref).
Under the double exclusion restriction, column (1) implies that $\widehat{\mathbb{P}[C_1,N_2]}=0.036$, which is not significantly different from zero. This estimate refers to individuals that were induced to open a LinkedIn account by the LinkedIn training but deleted it after the job training program. Column (2) implies that $\widehat{\mathbb{P}[C_1,A_2]}=0.042$. This suggests that several treated individuals would have also opened a LinkedIn account after the job training program if assigned to the control. Hence, the LinkedIn training only accelerated them to create an account. Column (5) implies that the two groups of movers are small compared to the full compliers.
The double exclusion restriction imposes that $\mathbb{P}[N_1,C_2]=\mathbb{P}[A_1,C_2]=0$. Since the treatment group was incentivized to immediately open a LinkedIn account in the first week of the job training program, the first mover type can be absent. However, if the LinkedIn training succeeds in explaining the benefits of LinkedIn in the long term, both mover types could be present. Hence, even though they are not detected by the necessary conditions, this empirical setting may include movers that violate the double exclusion restriction.
Panel B of (ref) shows a statistically significant LAFTE: The 95% confidence interval of the lower bound, $[0.047,0.305]$, does not include zero. The bounds rely on the assumption that the effect of a LinkedIn account six months after the job training program on employment is non-negative (MTR), and that always takers and compliers of a LinkedIn account six monthts after the job training are not less likely to find a job than never takers (MTS). Based on these bounds, we conclude that having a Linkedin account in both time periods increases the probability to find a job between 0.176 and 0.235 percentage points for individuals induced to open an account due to the LinkedIn training. The estimated bounds imply that the LAFTE is similar to the LATE of the first part of the LinkedIn training, which suggests that having a Linkedin account is most effective shortly after the job training program.
This application studies the LAFTE of a university education on income. School curricula, including the design and sequencing of learning content, are developed to receive in full. Moreover, students enroll in a program with the aim to graduate, which makes the LAFTE a valuable piece of information for their study choice. Hence, for both university educators and students the effect of following a full university program is of interest.
anelli2020returns studies the returns to a selective, expensive, and private university offering business, economics, and law degrees in a large Italian city. Admission to the elite university is based on a uni-dimensional application score. Every year, the number of admitted students is fixed, and the university strictly offers admission to the students with the highest application score. This procedure generates a cutoff in the application score, where students above the cutoff are offered admission. A score above the cutoff is used as an instrument for ever enrolled at the elite university to study the effect of ever enrolled on income.
The main sample is restricted to the elite university applicants between 1995 and 2000. For 645 of these applicants the dataset contains information on the application score, ever being enrolled at the elite university between 1995 and 2005, graduation from any university up unto the year 2005, and yearly income in the year 2005. If an applicant retook the admission test, the score refers to the first observed application.
Our $D_1$ equals one if an applicant was ever enrolled at the elite university between 1995 and 2005. This is the treatment variable in anelli2020returns. Our $D_2$ equals one if an applicant graduated from any university up unto the year 2005. The dataset does not contain information on whether the applicant graduated from the elite university. The instrument $Z$ equals one if the applicant scores above to application score cutoff. The outcome variable $Y$ is the logarithm of taxable income in 2005, measured before taxes and after deductions.
The estimated necessary conditions in column (1) and (2) of Panel C in (ref) are positive and significantly different from zero at the 1% level. Hence, we can reject that movers are absent and do not reject the necessary conditions of (ref).
Under the double exclusion restriction, column (1) implies that $\widehat{\mathbb{P}[C_1,N_2]}=0.047$. This mover type may refer to students who were induced to enroll at the elite university but never graduate. Column (2) implies that $\widehat{\mathbb{P}[C_1,A_2]}=0.537$. Recall that our $D_2$ equals one if a student graduated from any university, not just the elite institution, and so this group is relatively large because students with $Z=0$ may graduate from other universities.
The double exclusion restriction imposes that $\mathbb{P}[N_1,C_2]=\mathbb{P}[A_1,C_2]=0$. The first mover type is absent if university graduation for the never takers of elite university enrollment is not affected by scoring above or below the cutoff. This group of movers could be present if scoring above the cutoff has some positive psychological effect that persists even if a student does not enroll at the elite institution. The large estimate in column (2) of (ref) suggests that such psychological effects are close to zero. The second mover type is likely to be absent since anelli2020returns reports that across all application rounds only 23 students retook the admission test after scoring below the cutoff. Hence the possible number of always takers with enrollment in the elite university is small.
Panel C of (ref) shows a statistically significant LAFTE: the estimated bounds equal 0.532 and 7.545, and the 95% confidence interval of the lower bound, $[0.076, 0.988]$, does not include zero. The bounds rely on the assumption that the effect of university graduation on income is non-negative, and that always takers and compliers of university graduation have higher wages than never takers. The upper bound is large due to the large group of movers $\{C_1,A_2\}$. Based on the bounds, we conclude that the full treatment of enrolling at the elite and graduating at any university increases annual income between 53 and 750 log points for the individuals induced to enroll by scoring above the cutoff. The estimated bounds imply that the LAFTE may be much larger than the LATE of university enrollment. This suggests that consecutive parts of a university degree are complements and students may particularly benefit from fully completing a degree.
This application aims to study the LAFTE of a development program. The implementation of development programs often requires careful design, planning, and coordination with local partners, and can be expensive duflo2007using. Hence, the full treatment effect is of interest to the stakeholders.
burde2013bringing conduct and evaluate the randomized opening of village-based schools in Afghanistan, where primary-school participation rates are low. Their study sample includes 31 villages in rural Afghanistan. In the summer of 2007, they randomly opened schools in 13 of these villages. Village-based schools are public schools that are designed to deliver the official national curriculum to children living in close proximity to the school. They use the random assignment of village-based schools as an instrument for school enrollment to estimate the effect of enrollment on academic performance.
Their household survey data contains information on school enrollment and math and language test scores in two time periods: four months after and eight months after opening the village-based schools. Our $D_1$ equals one if a child was enrolled in school four months after the opening and $D_2$ equals one if a child was enrolled in school eight months after the opening. The binary instrument $Z$ equals one if a child lives in a village that randomly received a school. The outcome variable $Y$ is the standardized test score eight months after the opening of the schools. The final estimation sample includes 1,181 children, of in total 1,490 school-age children across the 31 villages.
Panel D in (ref) shows the estimated necessary conditions. Since the observed $D_2=1$ if $D_1=1$, the variable $D_{\lor}-D_2$ is zero for each individual, and hence the estimates in column (1) and (3) are zero. This indicates the absence of the mover types $\{C_1,N_2\}$ and $\{A_1,C_2\}$. The estimate in column (2) is negative and statistically significant at the 1% level. This estimate implies that $\mathbb{P}[N_1,C_2]>0$, which violates the double exclusion restriction. This mover type may arise if households need time to change their beliefs towards allowing their children to enroll in school. burde2013bringing discuss that the conservative beliefs in Afghanistan may be an impediment for school enrollment, in particular for girls.
(ref) shows the estimates for the parameters that can be identified when the double exclusion restriction is violated: the ACR type estimand in (ref) and the weighted average of LATEs in (ref). To conclude, this application shows that it depends on the empirical context whether the double exclusion restriction holds, and that the testable necessary conditions are able to help researchers detect violations of this assumption.
Policy evaluation has to deal with individuals that only take a subset of a treatment program. This poses the question: Which treatment effect is of interest to the policymaker? For instance, are the outcomes of the individuals receiving the full treatment or any individual that came in contact with the treatment of interest? Does the control group consists of individuals who obtained zero treatment or did not complete the full treatment? This paper defines the effect of full treatment versus no treatment at all as the parameter of interest.
We develop an IV framework in the presence of individuals that are induced by the instrument to obtain only a subset of the treatment, and refer to these individuals as movers. We show that these movers violate the exclusion restriction in the standard IV framework, that necessary conditions on the presence of movers are testable, and that under a double exclusion restriction the average treatment effect of receiving full treatment for the individuals induced into treatment by the instrument is partially identifiable. We refer to this treatment effect as the local average full treatment effect (LAFTE).
We study the LAFTE in empirical applications from four different fields of applied economics. Our methods find evidence for the presence of movers in all four studies. These types of movers align with the intuition provided by the setting at hand. In three of these applications, the necessary conditions for the double exclusion restriction cannot be rejected, and the bounds identify a statistically significant LAFTE. We discuss the potential policy implications within the empirical context of the applications.
We allow the instrument to induce individuals into dropping out or late-adoption of treatment by taking into account two treatment parts. In settings in which the instrument additionally affects dropping out or late-adoption across more treatment parts, the exclusion restriction in our IV framework may also be violated. In case these settings include the observed treatment status across these treatment parts, our framework can be extended to more specific mover types that also take these violations into account.
\linespread{1.00} \addcontentsline{toc}{section}{References}