Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
114,732 characters · 15 sections · 177 citation commands
Sharp Bounds for the Marginal Treatment Effect with Sample Selection
\setstretch{1}
}
\newsavebox{\tablebox} \newlength{\tableboxwidth}
I analyze treatment effects in situations when agents endogenously select into the treatment group and into the observed sample. As a theoretical contribution, I propose pointwise sharp bounds for the marginal treatment effect (MTE) of interest within the always-observed subpopulation under monotonicity assumptions. Moreover, I impose an extra mean dominance assumption to tighten the previous bounds. I further discuss how to identify those bounds when the support of the propensity score is either continuous or discrete. Using these results, I estimate bounds for the MTE of the Job Corps Training Program on hourly wages for the always-employed subpopulation and find that it is decreasing in the likelihood of attending the program within the Non-Hispanic group. For example, the Average Treatment Effect on the Treated is between \$.33 and \$.99 while the Average Treatment Effect on the Untreated is between \$.71 and \$3.00.
\
Keywords: Marginal Treatment Effect, Sample Selection, Partial Identification, Principal Stratification, Program Evaluation, Training Programs.
\
JEL Codes: C31, C35, C36, J38
\doublespacing
In the applied treatment effects literature, there are many problems that face two identification challenges: endogenous selection into treatment and endogenous sample selection. For instance, in Labor Economics, if a researcher wants to evaluate the effect of a job training program on wages, she has to understand why agents choose to enroll in the program and why agents select into her sample by being employed. In this situation, she may combine information on hourly labor earnings (the observable outcome) and employment (sample selection status) to uncover the effect on hourly wages (the outcome of interest). Similar problems appear in Labor Economics when analyzing the college wage premium and scarring effects. In the Health Sciences, a researcher faces the same identification challenges when analyzing the effect of a drug on a health quality index when the drug may save a patient's life. Moreover, in randomized control trials, researchers are concerned with non-compliance and differential attrition rates between treated and control groups. This double selection problem is also present when analyzing the effect of an educational intervention on short- and long-term outcomes and the effect of procedural laws on litigation outcomes.\footnote{Training programs are studied by Heckman1999a, Lee2009 and Chen2015. The college wage premium is analyzed by Altonji1993, Card1999 and Carneiro2011. Scarring effects are discussed by Heckman1980, Farber1993 and Jacobson1993. Some education interventions are studied by Krueger2001, Angrist2006, Angrist2009, Chetty2011 and Dobbie2015. Medical treatments are analyzed by CASS1984, Sexton1984 and Health2004. Litigation outcomes are discussed by Helland2017. RCT with attrition are illustrated by DeMel2013 and Angelucci2015.}
To simultaneously address both idetification challenges, I propose a Generalized Roy Model (Heckman1999) with sample selection in which there is one outcome of interest that is observed only if the individual self-selects into the sample. Under a monotonicity assumption on the sample selection indicator, I decompose the Marginal Treatment Response (MTR) function for the potential observable outcome when treated as a weighted average of (i) the MTR on the outcome of interest for the subpopulation who is always observed and (ii) the Marginal Treatment Effect (MTE) on the observable outcome for the subpopulation who is observed only when treated. Under a bounded (in one direction) support condition, this decomposition is useful because it allows me to propose pointwise sharp bounds for the MTE on the outcome of interest within the always-observed subpopulation ($MTE^{OO}$) as a function of (i) the MTR functions on the observable outcome, (ii) the maximum and (or) minimum of the support of the potential outcome, and (iii) the proportions of always-observed individuals and observed-only-when-treated individuals. I also show that it is impossible to construct bounds without extra assumptions when the support of the potential outcome is the entire real line. After that, I impose an extra mean dominance assumption that compares the always-observed population against the observed-only-when-treated population, tightening the previous bounds. Moreover, under this new assumption, I show that those tighter bounds are also pointwise sharp and derive an informative lower bound even when the support of the potential outcome is the entire real line.
I then proceed to show that those bounds are well-identified. When the support of the propensity score is an interval, the relevant objects are point identified by applying the local instrumental variable approach (LIV, see Heckman1999) to the expectations of the observable outcome and of the selection indicator conditional on the propensity score and the treatment status. However, in many empirical applications, the support of the propensity score is a finite set. In such a context, I can identify bounds for the $MTE^{OO}$ of interest by adapting the nonparametric bounds proposed by Mogstad2017 or the flexible parametric approach suggested by Brinch2017 to encompass a sample selection problem. When using the nonparametric approach, the bounds for the $MTE^{OO}$ of interest are simply an outer set that contains the true $MTE^{OO}$, i.e., they are not pointwise sharp anymore.
Partial identification of the $MTE^{OO}$ of interest is useful for two reasons. First and most importantly, bounds for the $MTE^{OO}$ can be used to shed light on the heterogeneity of treatment effects, allowing the researcher to understand who would benefit and who would lose from a specific treatment, as recently illustrated by Cornelissen2018 and Bhuller2019. This knowledge can be used to optimally design policies that incentivize to agents to take a treatment. Second, bounds for the $MTE^{OO}$ can be used to construct bounds for any treatment effect parameter that is written as a weighted integral of the $MTE^{OO}$. For instance, by taking a weighted average of the pointwise sharp bounds for the $MTE^{OO}$, one can bound the average treatment effect (ATE), the average treatment effect on the treated (ATT), any local average treatment effect (LATE, Imbens1994) and any policy-relevant treatment effect (PRTE, Heckman2001a) within the always-observed subpopulation. Although such bounds may not be sharp for any specific parameter, they are a flexible and easy-to-apply tool for many empirical problems that depend on a varied set of treatment effects.\footnote{As a consequence of this trade-off between flexibility and sharpness, I recommend the use of a specialized tool if the parameter of interest already has specific bounds (e.g., the ITT by Lee2009 and the LATE by Chen2015).}
Finally, I illustrate the usefulness of the proposed bounds for the $MTE^{OO}$ of interest by analyzing the effect of the Job Corps Training Program (JCTP) on hourly wages within the Non-Hispanic always-employed subpopulation. My framework is ideal to analyze this important experiment because it simultaneously addresses the imperfect compliance issue (self-selection into treatment) by focusing on the MTE and the endogenous employment decision (sample selection) by using a partial identification strategy. Although my $MTE^{OO}$ bounds are uninformative when using only the monotonicity assumption, they are tight and positive under a mean dominance assumption, illustrating the identification power of extra assumptions in a context of partial identification. Most interestingly, I find that the bounds of the $MTE^{OO}$ on hourly wages are decreasing in the likelihood of attending the program, implying that the agents who would benefit the most from the JCTP are the least likely to attend it. As a consequence of this result, my estimates suggest that ATU is greater than the ATT for the always-employed subpopulation. Moreover, my bounds for the $LATE^{OO}$ are in line with the estimates of Chen2015 and the effect of the JCTP on employment is positive for every agent according to the test proposed by Machado2018. Finally, as a by-product of my estimation strategy, I also find that the MTE on employment and hourly labor earnings are decreasing in the likelihood of attending the JCTP, a result that is in line with the estimated upper bounds of Chen2017.
I make contributions to three branches of literature: identification of treatment effects using an instrument, identification of treatment effects with sample selection, and the effect of job training programs. They are all vast and only briefly summarized here.
In the literature about treatment effects with an instrument, Imbens1994 show that we can identify the LATE. Heckman1999, Heckman2005 and Heckman2006 define the MTE and explain how to compute any treatment effect as a weighted average of the MTE. However, if the support of the propensity score is not the unit interval, then it is not possible to non-parametrically identify some common treatment effects, such as the ATE, the ATT and the ATU. A parametric solution to this problem is given by Brinch2017, who identify a flexible polynomial function for the MTE whose degree is defined by the cardinality of the propensity score support, while a nonparametric solution is given by Mogstad2017, who use the information contained on IV-like estimands to construct non-parametrically worst- and best- case bounds for policy-relevant treatment effects.\footnote{Other important contributions are made by Manski1990, Manski1997, Manski2000, Heckman2001, Bhattacharya2008, Chesher2010, Chiburis2010, Shaikh2011, Bhattacharya2012, Cornelissen2016, Chen2017, Huber2017, Kowalski2018, Mourifie2018 and Zhou2019.}
I contribute to this literature by extending the non-parametric approach by Mogstad2017 and the flexible parametric approach by Brinch2017 to encompass a sample selection problem. By doing so, I can partially identify the MTE function on the outcome of interest, which, in my framework, is different from the observable outcome.
In the literature about identification of treatment effects with sample selection, the control function approach (Heckman1979, Ahn1993 and Das2003) and the use of auxiliary data (Chen2008) are two classical solutions to this problem. Another approach is to partially identify the parameter of interest by imposing weak monotonicity assumptions. For example, in a seminal paper, Lee2009 imposes that sample selection is monotone on treatment assignment to sharply bound the ITT for the subpopulation of always-observed individuals ($ITT^{OO}$).\footnote{Other relevant contributions are made by Frangakis2002, Blundell2007, Imai2008, Lechner2010, Blanco2013, Mealli2013, Behaghel2015 and Huber2015.}
In the intersection of both literatures, a few authors address the problem of sample selection and endogenous treatment simultaneously. By using two instrumental variables, Fricke2015 and Lee2016 identify different treatment effects. However, since finding a credible instrument for sample selection is challenging in some cases, it is worth developing alternative tools that do not require more than an instrument for selection into treatment. Frolich2014 point identify the LATE by assuming that there is no contemporaneous relationship between the potential outcomes and the sample selection problem. Chen2015 derive bounds for Average Treatment Effect within the always-observed compliers ($LATE^{OO}$) by combining one instrument with a double exclusion restriction with monotonicity assumptions on the sample selection and the selection into treatment problems.\footnote{Other important contributions are made by Huber2014, Steinmayr2014, Blanco2017 and Kedagni2018.}
I contribute to this literature by partially identifying the MTE on the always-observed subsample allowing for a contemporaneous relationship between the potential outcomes and the sample selection problem, and using only one (discrete) instrument combined with a monotonicity assumption. Deriving bounds for the $MTE^{OO}$ is theoretically important because it can unify, in one framework, the bounds for different treatment effects with sample selection. It is also empirically relevant because it allows us to partially identify any treatment effect on the outcome of interest in many empirical problems. For instance, when analyzing the effect of a job training program on wages, it is useful to compare the ATT with the ATU in order to understand whether the workers who would benefit the most from such a policy are actually the ones who receive training.
In the literature about job training programs, Heckman1999a wrote an influential survey paper. In particular, many papers were written about the effects of the Job Corps Training Program (JCTP) after a randomized experiment funded by the U.S. Department of Labor in 1995.\footnote{For example, significant contributions are made by Schochet2001, Schochet2008, Flores-Lagunes2010, Flores2012, Blanco2013, Blanco2013a, Blanco2017 and Chen2017.} Finally, my work is closer to the research done by Lee2009 and Chen2015, who analyze the effect of the Job Corps Training Program on wages by focusing, respectively, on the ITT and the LATE parameters within the always-observed subpopulation. Lee2009 rules out a zero effect after accounting for the loss in labor market experience generated by the extra education acquired by Job Corps participants. Chen2015 find that the $LATE^{OO}$ on hourly wages four years after randomization is between 5.7% and 13.9% for the entire population and between 7.7% and 17.5% for the non-Hispanic population under monotonicity and mean dominance assumptions.
I contribute to this literature by analyzing the MTE on hourly wages within the Non-Hispanic group and formally testing whether this training program has a monotone effect on employment by implementing the test proposed by Machado2018.
This paper proceeds as follows: Section (ref) details the Generalized Roy Model with sample selection; Section (ref) explains how to derive bounds for the $MTE^{OO}$ of interest; Sections (ref) and (ref) discuss identification of the $MTE^{OO}$ bounds when the support of the propensity score is continuous or discrete; and Section (ref) analyzes the effect of the Job Corps Training Program on hourly wages. Finally, Section (ref) concludes.
I begin with the classical potential outcome framework by Rubin1974 and modify it to include a sample selection problem. Let $Z$ be an instrumental variable whose support is given by $\mathcal{Z}$, $X$ be a vector of covariates whose support is given by $\mathcal{X}$, $W \coloneqq \left(X, Z\right)$ be a vector that combines the covariates and the instrument whose support is given by $\mathcal{W} \coloneqq \mathcal{X} \times \mathcal{Z}$, $D$ be a treatment status indicator, $Y_{0}^{*}$ be the potential outcome of interest when the person is not treated, and $Y_{1}^{*}$ be the potential outcome of interest when the person is treated. The outcome variable of interest (e.g., wages) is $Y^{*} \coloneqq D \cdot Y_{1}^{*} + \left(1 - D\right) \cdot Y_{0}^{*}$. Moreover, let $S_{1}$ and $S_{0}$ be potential sample selection indicators when treated and when not treated, and define $S \coloneqq D \cdot S_{1} + \left(1 - D\right) \cdot S_{0}$ as the sample selection indicator (e.g., employment status). Define $Y \coloneqq S \cdot Y^{*}$ as the observable outcome (e.g., labor earnings). I also define $Y_{1} \coloneqq S_{1} \cdot Y_{1}^{*}$ and $Y_{0} \coloneqq S_{0} \cdot Y_{0}^{*}$ as the potential observable outcomes. Observe that, following Lee2009 and Chen2015, my notation implicitly imposes two exclusion restrictions: Z has no direct impact on the potential outcome of interest nor on the sample selection indicator. The second exclusion restriction requires attention in empirical applications. On the one hand, it may be a strong assumption in randomized control trials if sample selection is due to attrition and initial assignment has an effect on the subject's willingness to contact the researchers. On the other hand, it may be a reasonable assumption in many labor market applications, such as the evaluation of a job training program. For instance, in my empirical section, it is plausible that the initial random assignment to the Job Corps Training Program (JCTP) has no impact on future employment status.
I model sample selection and selection into treatment using the Generalized Roy Model Heckman1999. Let $U$ and $V$ be random variables, and $P:\mathcal{W} \rightarrow \mathbb{R}$ and $Q:\left\lbrace 0, 1 \right\rbrace \times \mathcal{X} \rightarrow \mathbb{R}$ be unknown functions. I assume that:
and
As Vytlacil2002 shows, equations (ref) and (ref) are equivalent to assuming monotonicity conditions on the selection-into-treatment problem (Imbens1994) and on the sample selection problem (Lee2009). I stress that both monotonicity assumptions are testable using the tools developed by Machado2018. Note also that, given equation (ref), $S_{0} = \mathbf{1}\left\lbrace Q\left(0, X\right) \geq V \right\rbrace$ and $S_{1} = \mathbf{1}\left\lbrace Q\left(1, X\right) \geq V \right\rbrace$.
The random variables $U$ and $V$ are jointly continuously distributed conditional on $X$ with density $f_{U,V \left\vert X \right.}:\mathbb{R}^{2} \times \mathcal{X} \rightarrow \mathbb{R}$ and cumulative distribution function $F_{U,V \left\vert X \right.}:\mathbb{R}^{2} \times \mathcal{X} \rightarrow \mathbb{R}$. As has been shown in the literature, equations (ref) and (ref) can be rewritten as
where $\tilde{P}\left(W\right) \coloneqq F_{U \left\vert X \right.} \left(P\left(W\right) \left\vert X \right. \right)$, $\tilde{U} \coloneqq F_{U \left\vert X \right.} \left(U \left\vert X \right.\right)$, $\tilde{Q}\left(D, X\right) \coloneqq F_{V \left\vert X \right.} \left(Q\left(D, X\right) \left\vert X \right. \right)$, and $\tilde{V} \coloneqq F_{V \left\vert X \right.} \left(V \left\vert X \right.\right)$. Consequently, the marginal distributions of $\tilde{U}$ and $\tilde{V}$ conditional on $X$ follow the standard uniform distribution. Since this is merely a normalization, I drop the tilde and mantain throughout the paper the normalization that $\left(P\left(w\right), Q\left(d, x\right)\right) \in \left[0, 1\right]^{2}$ for any $\left(x, z, d\right) \in \mathcal{W} \times \left\lbrace 0, 1 \right\rbrace$ and the marginal distributions of $U$ and $V$ conditional on $X$ follow the standard uniform distribution, even though their joint distribution allows for any kind of dependency between those two variables. As a consequence of such normalization, $P\left(w\right)$ represents the propensity score and is equal to $\mathbb{P}\left[\left. D = 1 \right\vert W = w\right]$, while $Q\left(d, x\right)$ is equal to $\mathbb{P}\left[\left. S_{d} = 1 \right\vert X = x\right]$.
Moreover, I assume that:
Assumption (ref) is fairly general. Case 1 covers continuous random variables whose support is convex and bounded below (e.g., wages), while Case 3.a covers continuous variables with bounded convex support (e.g., test scores). Case 3.b encompasses not only binary variables, but also any discrete variable whose support is finite (e.g., years of education). It also includes mixed random variables whose support is not an interval but achieves its maximum and minimum. Case 2 is included for theoretical complementness. Furthermore, Proposition (ref) shows that assumption (ref) is partially necessary to the existence of bounds for the $MTE^{OO}$ of interest in the sense that, if $\underline{y}^{*} = -\infty$ and $\overline{y}^{*} = + \infty$, then it is impossible to bound the marginal treatment effect on the outcome of interest within the always-observed subpopulation without any extra assumptions.
Assumption (ref) goes beyond the monotonicity condition implicitly imposed by equation (ref) by assuming that the direction of the effect of treatment on the sample selection indicator is known and positive, i.e., $Q\left(1, x\right) \geq Q\left(0, x\right)$ for any $x \in \mathcal{X}$. In this sense, it is a standard assumption in the literature.\footnote{Lee2009 and Chen2015 write it in an equivalent way as $S_{1} \geq S_{0}$, while Manski1997 and Manski2000 call it the “monotone treatment response” assumption.} Most importantly, it is also a testable assumption using the tools developed by Machado2018, because, under monotone sample selection (equation (ref)), identification of the sign of the ATE on the selection indicator provides a test for Assumption (ref). However, Assumption (ref) is slightly stronger than what is usually imposed in the literature, because it additionally imposes $Q\left(0, x\right) > 0$ and $Q\left(1, x\right) > Q\left(0, x\right)$ for any $x \in \mathcal{X}$. While the first inequality implies that there is a subpopulation who is always observed, allowing me to properly define my target parameter (the marginal treatment effect on the outcome of interest within the always-observed population, $MTE^{OO}$), the second inequality implies that there is a subpopulation who is observed only when treated, making the problem theoretically interesting by eliminating trivial cases of point identification of the $MTE^{OO}$ as discussed in Proposition (ref). Finally, I emphasize that all my results can be stated and derived with some straightforward changes if I impose $Q\left(0, x\right) > Q\left(1, x\right) > 0$ for any $x \in \mathcal{X}$ instead of Assumption (ref), as is done in Appendix (ref). I also discuss, in Appendix (ref), an agnostic approach to monotonicity in the sample selection problem (equation (ref)) and show, in Appendix (ref), that bounds derived with non-monotone sample selection are uninformative (i.e., equal to $\left(\underline{y}^{*} - \overline{y}^{*}, \overline{y}^{*} - \underline{y}^{*}\right)$) under mild regularity conditions.
In my empirical application, Assumption (ref) imposes that the JCTP has a positive effect on employment for all individuals, which is plausible given the objectives and services provided by this training program. As discussed by Chen2015, the two potential threats against it --- the “lock-in” effect (Ours2004) and an increase in the reservation wage of treated individuals --- are likely to become less relevant in the long run, justifying my focus on the hourly wage after 208 weeks from randomization. Most importantly, this assumption is formally tested by the method developed by Machado2018 and I reject, at the 1%-significance level, the null hypothesis that Assumption (ref) is invalid within the Non-Hispanic group.
Finally, in partial identification contexts, extra assumptions may have a lot of identification power. In the specific case of identifying treatment effects with sample selection, it is common to use mean or stochastic dominance assumptions to tighten the bounds for the parameter of interest (Imai2008, Blanco2013, Huber2015 and Huber2017) and justify them based on the intuitive argument that some population sub-groups have more favorable underlying characteristics than others. In particular, I discuss the identifying power of the following mean dominance assumption\footnote{In appendix (ref), I derive bounds for the MTE of interest when the above inequality holds in the other direction.}:
Unfortunately, this assumption is empirically untestable, implying that its use must be justified for each application based on qualitative or theoretical arguments. In particular, in my empirical application, Assumption (ref) imposes that the marginal treatment response function of wages when treated for the always-employed population is greater than the same object for the employed-only-when-treated population. Intuitively, this assumption imposes that the group with better potential employment outcomes also has, on average, better potential wages, i.e., there is positive selection into employment.
The target parameter, the MTE on the outcome of interest for the subpopulation who is always observed ($MTE^{OO}$), is given by
for any $u \in \left[0, 1\right]$ and any $x \in \mathcal{X}$, and is a natural parameter of interest. In labor market applications where sample selection is due to observing wages only when agents are employed, it is the effect on wages for the subpopulation who is always employed. In medical applications where sample selection is due to the death of a patient, it is the effect on health quality for the subpopulation who survives regardless of treatment status. In the education literature where sample selection is due to students quitting school, it is the effect on test scores for the subpopulation who do not drop out of school regardless of treatment status. In all those cases, the target parameter captures the intensive margin of the treatment effect.\footnote{If the researcher is interested in the extensive margin of the treatment effect, captured by the MTE on the observable outcome ($\mathbb{E}\left[Y_{1} - Y_{0} \left\vert X = x, U = u\right.\right]$) and by the MTE on the selection indicator ($\mathbb{E}\left[S_{1} - S_{0} \left\vert X = x, U = u\right.\right]$), he or she can apply the identification strategies described by Heckman2006, Brinch2017 and Mogstad2017.}
Other possibly interesting parameters are the MTE on the outcome of interest within the subpopulation who is never observed ($\mathbb{E}\left[Y_{1}^{*} - Y_{0}^{*} \left\vert X = x, U = u, S_{0} = 0, S_{1} = 0 \right.\right]$, $MTE^{NN}$), the MTR function under no treatment for the outcome of interest within the subpopulation who is observed only when treated ($\mathbb{E}\left[Y_{0}^{*} \left\vert X = x, U = u, S_{0} = 0, S_{1} = 1 \right.\right]$, $MTR_{0}^{NO}$) and the MTR function under treatment for the outcome of interest within the subpopulation who is observed only when treated ($\mathbb{E}\left[Y_{1}^{*} \left\vert X = x, U = u, S_{0} = 0, S_{1} = 1 \right.\right]$, $MTR_{1}^{NO}$). While the last parameter can be partially identified (Appendix (ref)), the first two parameters are impossible to point identify or bound in an informative way because the outcome of interest ($Y_{0}^{*}$ or $Y_{1}^{*}$) is never observed for the conditioning subpopulations.\footnote{Zhang2008 discuss this identification issue in a deeper way. Moreover, in some applications (e.g., analyzing the impact of a medical treatment on a health quality measure where selection is given by whether the patient is alive), the potential outcome $Y_{d}^{*}$ is not even properly defined when $S_{d} = 0$ for $d \in \left\lbrace 0, 1 \right\rbrace$.} As a consequence, it is not possible to point identify or bound in an informative way the Marginal Treatment Effect for the entire population ($\mathbb{E}\left[Y_{1}^{*} - Y_{0}^{*} \left\vert X = x, U = u\right.\right]$, $MTE$) either. Note also that the subpopulation who is observed only when not treated ($S_{0} = 1$ and $S_{1} = 0$) does not exist by Assumption (ref). Furthermore, observe that the conditioning subpopulations in all the above-mentioned parameters are determined by post-treatment outcomes and, as a consequence, are connected to the statistical literature known as principal stratification (Frangakis2002).
I now focus on the target parameter $\Delta_{Y^{*}}^{OO}\left(x, u\right)$ given by equation (ref). While Subsection (ref) derives bounds for the $MTE^{OO}$ of interest (equation (ref)) using only a monotonicity assumption (Assumptions (ref)-(ref)), Subsection (ref) tightens those bounds by additionally imposing the Mean Dominance Assumption (ref). Finally, Subsection (ref) discusses the empirical relevance of such bounds.
Here, my goal is to derive bounds for $\Delta_{Y^{*}}^{OO}\left(x, u\right)$ under Assumptions (ref)-(ref). Note that the second right-hand term in equation (ref) can be written as\footnote{Appendix (ref) contains a proof of this claim.}
where I define $m_{0}^{Y}\left(x, u\right) \coloneqq \mathbb{E}\left[Y_{0} \left\vert X = x, U = u \right.\right]$ and $m_{0}^{S}\left(x, u\right) \coloneqq \mathbb{E}\left[S_{0} \left\vert X = x, U = u \right.\right]$ as the MTR functions associated with the counterfactual variables $Y_{0}$ and $S_{0}$, respectively. In this section, I assume that all terms in the right-hand side of equation (ref) are point identified, postponing the discussion about their identification to Sections (ref) and (ref).
The first right-hand term in equation (ref) can be written as\footnote{Appendix (ref) contains a proof of this claim.}
where $m_{1}^{Y}\left(x, u\right) \coloneqq \mathbb{E}\left[Y_{1} \left\vert X = x, U = u \right.\right]$ is the MTR function associated with the counterfactual variable $Y_{1}$, $\Delta_{Y}^{NO}\left(x, u\right) \coloneqq \mathbb{E}\left[ Y_{1} - Y_{0} \left\vert X = x, U = u, S_{0} = 0, S_{1} = 1 \right.\right]$ is the MTE on the observable outcome $Y$ for the subpopulation who is observed only when treated, $\Delta_{S}\left(x, u\right) \coloneqq \mathbb{E}\left[S_{1} - S_{0} \left\vert X = x, U = u \right.\right] = m_{1}^{S}\left(x, u\right) - m_{0}^{S}\left(x, u\right)$ is the MTE on the selection indicator, and $m_{1}^{S}\left(x, u\right) \coloneqq \mathbb{E}\left[S_{1} \left\vert X = x, U = u \right.\right]$ is the MTR function associated with the counterfactual variable $S_{1}$. In this section, I also assume that $m_{1}^{Y}\left(x, u\right)$ and $\Delta_{S}\left(x, u\right)$ are point identified, postponing the discussion about their identification to Sections (ref) and (ref).
Although point identification of $\mathbb{E}\left[Y_{1}^{*} \left\vert X = x, U = u, S_{0} = 1, S_{1} = 1 \right.\right]$ is not possible due to the term $\Delta_{Y}^{NO}\left(x, u\right)$ in equation (ref), I can find identifiable bounds for it.\footnote{Appendix (ref) contains a proof of this proposition.}
Note that, even when the support is bounded in only one direction (Assumptions (ref).1 and (ref).2), it is possible to derive lower and upper bounds for $\mathbb{E}\left[Y_{1}^{*} \left\vert X = x, U = u, S_{0} = 1, S_{1} = 1 \right.\right]$.
At this point, it is worth understanding the determinants of the width of those bounds. First, if there is no sample selection problem at all ($\mathbb{P}\left[S_{0} = 1, S_{1} = 1 \left\vert X = x, U = u \right.\right] = 1$, i.e., the always-observed group is the entire population), then $m_{0}^{S}\left(x, u\right) = 1$, $\Delta_{S}\left(x, u\right) = 0$, implying point identification in equation (ref). Second, if there is no problem of differential sample selection with respect to treatment status ($\mathbb{P}\left[S_{0} = 0, S_{1} = 1 \left\vert X = x, U = u \right.\right] = 0$, i.e., the observed-only-when-treated subpopulation has zero mass), then $\Delta_{S}\left(x, u\right) = 0$, once more implying point identification in equation (ref). Both cases are theoretically uninteresting and ruled out by Assumption (ref).
Finally, combining equations (ref) and (ref) and Proposition (ref), I can partially identify the target parameter $\Delta_{Y^{*}}^{OO}\left(x, u\right)$:
Furthermore, I can show that\footnote{The definition of pointwise sharpness used here and in the rest of the paper follows the definition of sharpness given by Canay2017. Moreover, note that, if the functions $m_{0}^{Y}$, $m_{1}^{Y}$, $m_{0}^{S}$ and $\Delta_{S}$ are point identified only in a subset of the unit interval, then pointwise sharpness holds only in that subset.}:
Intuitively, Theorem (ref) says that, for any $\delta\left(\overline{x}, \overline{u}\right) \in \left(\underline{\Delta_{Y^{*}}^{OO}}\left(\overline{x}, \overline{u}\right), \overline{\Delta_{Y^{*}}^{OO}}\left(\overline{x}, \overline{u}\right)\right)$, it is possible to create candidate random variables $\left(\tilde{Y}_{0}^{*}, \tilde{Y}_{1}^{*}, \tilde{U}, \tilde{V}\right)$ that generate the candidate marginal treatment effect $\delta\left(\overline{x}, \overline{u}\right)$ (equation (ref)), satisfy the bounded support condition --- a restriction imposed by my model (Assumption (ref)) and summarized in equation (ref) --- and generate the same distribution of the observable variables --- a restriction imposed by the data and summarized in equation (ref). In other words, the data and the model in Section (ref) do not generate enough restrictions to refute that the true target parameter $\Delta_{Y^{*}}^{OO}\left(\overline{x}, \overline{u}\right)$ is equal to the candidate target parameter $\delta\left(\overline{x}, \overline{u}\right)$.
Moreover, the bounded support condition (Assumption (ref)) is partially necessary to the existence of bounds for the target parameter $\Delta_{Y^{*}}^{OO}\left(\overline{x}, \overline{u}\right)$. When the support is unbounded in both directions (i.e., $\underline{y}^{*} = - \infty$ and $\overline{y}^{*} = + \infty$), then it is impossible to derive bounds for the target parameter $\Delta_{Y^{*}}^{OO}\left(\overline{x}, \overline{u}\right)$ without any extra assumption. Proposition (ref) formalizes this last statement.\footnote{Appendix (ref) contains the proof of this proposition, whose intuition is similar to the one provided for Theorem (ref).}
In other words, when the support of the potential outcome is the entire real line, the data and the model in Section (ref) do not generate enough restrictions to refute that the true target parameter $\Delta_{Y^{*}}^{OO}\left(\overline{x}, \overline{u}\right)$ is equal to an arbitrarily large effect in magnitude. This impossibility result is interesting in light of the previous literature about partial identification of treatment effects with sample selection. In the case of the $ITT^{OO}$ (Lee2009) and the $LATE^{OO}$ (Chen2015), it is possible to construct informative bounds even when the support of the potential outcome is the entire real line. However, when focusing on a specific point of the $MTE^{OO}$ function, it is impossible to construct informative bounds when $\mathcal{Y}^{*} = \mathbb{R}$ due to the local nature of the target parameter.
There is one remark about the results I just derived. Theorem (ref) and Proposition (ref) do not impose any smoothness condition on the joint distribution of $\left(Y_{0}^{*}, Y_{1}^{*}, U, V, Z, X\right)$. In particular, the conditional cumulative distribution functions $F_{V \left\vert X, U \right.}$, $F_{Y_{0}^{*} \left\vert X, U, V \right.}$ and $F_{Y_{1}^{*} \left\vert X, U, V \right.}$ are allowed to be discontinuous functions of U at the point $\overline{u}$. Appendix (ref) states and proves a sharpness result similar to Theorem (ref) and an impossibility result similar to Proposition (ref) when $F_{V \left\vert X, U \right.}$, $F_{Y_{0}^{*} \left\vert X, U, V \right.}$ and $F_{Y_{1}^{*} \left\vert X, U, V \right.}$ must be continuous functions of U.
Here, I use the Mean Dominance Assumption (ref) to tighten the bounds for the target parameter $\Delta_{Y^{*}}^{OO}$ (equation (ref)) given by Corollary (ref). Note that Assumption (ref) implies that $\Delta_{Y}^{NO}\left(x, u\right) \leq \dfrac{m_{1}^{Y}\left(x, u\right)}{m_{1}^{S}\left(x, u\right)} \leq \mathbb{E}\left[Y_{1}^{*} \left\vert X = x, U = u, S_{0} = 1, S_{1} = 1 \right.\right]$ by equations (ref) and (ref). As a consequence, by following the same steps of the proof of corollary (ref), I can derive:
Notice that, under Mean Dominance Assumption (ref), I can increase the lower bounds proposed in Corollary (ref) under Assumption (ref) and provide an informative lower bound even when the support of the outcome of interest is the entire real line, a result in stark contrast with Proposition (ref).\footnote{Appendix (ref) discusses when Corollary (ref) provides bounds that are strictly tighter than the ones provided by Corollary (ref).} These improvements clearly show the identifying power of the Mean Dominance Assumption (ref). Moreover, the phenomenon of obtaining more informative bounds by imposing extra assumptions is common in the partial identification literature, as explained by Tamer2010 and illustrated by Kline2016.
As in Subsection (ref), I assume that $m_{0}^{Y}\left(x, u\right)$, $m_{1}^{Y}\left(x, u\right)$, $m_{0}^{S}\left(x, u\right)$, $m_{1}^{S}\left(x, u\right)$, and $\Delta_{S}\left(x, u\right)$ are point identified, postponing the discussion about their identification to Sections (ref) and (ref).
Now, using the above corollary, I can combine the sharpness and the impossibility results of Subsection (ref) in one single proposition\footnote{Appendix (ref) contains a proof of this proposition, whose intuition is similar to the one provided for Theorem (ref). The only difference is that, now, the function $F_{\tilde{Y}_{0}^{*}, \tilde{Y}_{1}^{*}, \tilde{U}, \tilde{V}, Z, X}$ at $\tilde{U} = \bar{u}$ must also satisfy equation (ref).}:
Note that, in addition to all the restriction imposed by Theorem (ref), the candidate random variables $\left(\tilde{Y}_{0}^{*}, \tilde{Y}_{1}^{*}, \tilde{U}, \tilde{V}\right)$ must also satisfy an extra model restriction (equation (ref)) associated with the Mean Dominance Assumption (ref). Intuitively, Proposition (ref) says that the data (equation (ref)) and the model (equations (ref) and (ref)) do not generate enough restrictions to refute that the true target parameter $\Delta_{Y^{*}}^{OO}\left(\overline{x}, \overline{u}\right)$ is equal to the candidate target parameter $\delta\left(\overline{x}, \overline{u}\right)$ (equation (ref)).
Now, it is worth discussing the empirical relevance of partially identifying the $MTE^{OO}$ of interest. First, bounds for the $MTE^{OO}$ can illuminate the heterogeneity of the treatment effect, allowing the researcher to understand who would benefit and who would lose with a specific treatment. This is important because common parameters (e.g., $ATE^{OO}$, $ATT^{OO}$, $ATU^{OO}$, $LATE^{OO}$) can be positive even when most people lose with a policy if the few winners have very large gains. Moreover, knowing, even partially, the $MTE^{OO}$ function can be useful to optimally design policies that provides incentives to agents to take some treatment. Second, I can use the $MTE^{OO}$ bounds to partially identify any treatment effect that is described as a weighted integral of $\Delta_{Y^{*}}^{OO}\left(x, u\right)$ because
where $\omega(x, \cdot)$ is a known or identifiable weighting function. Even though such bounds may not be sharp for any specific parameter, they are a general and off-the-shelf solution to many empirical problems. As a consequence of this trade-off, I recommend the applied researcher to use a specialized tool if he or she is interested in a parameter that already has specific bounds for it (e.g., $ITT^{OO}$ by Lee2009 and $LATE^{OO}$ by Chen2015). However, I suggest the applied researcher to easily compute a weighted integral of pointwise sharp bounds for the MTE of interest if he or she is interested in parameters without specialized bounds (e.g., ATE, ATT and ATU in the case with imperfect compliance). In other words, facing a trade-off between empirical flexibility and sharpness, the partial identification tool proposed in this paper focus on empirical flexibility while still ensuring pointwise sharpness of the bounds for the MTE of interest.
Tables (ref) and (ref) show some of the treatment effect parameters that can be partially identified using inequality (ref). More examples are given by Heckman2006 and Mogstad2017.
Here, I fix $x \in \mathcal{X}$ and impose that the support of the propensity score, defined by $\mathcal{P}_{x} \coloneqq \left\lbrace P\left(x, z\right): z \in \mathcal{Z} \right\rbrace$, is an interval\footnote{$\mathcal{P}_{x}$ as an interval may be achieved by a continuous instrument $Z$ or by the existence of independent covariates Carneiro2011.}. Then, under Assumptions (ref)-(ref), the MTR functions associated with any variable $A \in \left\lbrace Y, S \right\rbrace$ are point identified by\footnote{Appendix (ref) contains a proof of this claim based on the Local Instrumental Variable (LIV) approach described by Heckman2005.}:
and
for any $p \in \mathcal{P}_{x}$.
Finally, the pointwise sharp bounds for $\Delta_{Y^{*}}^{OO}\left(x, p\right)$ are point identified by combining equations (ref) and (ref), the fact that $\Delta_{S}\left(x, p\right) = m_{1}^{S}\left(x, p\right) - m_{0}^{S}\left(x, p\right)$, and Corollaries (ref) or (ref).
When the support of the propensity score is not an interval, I cannot point identify $m_{0}^{Y}\left(x, u\right)$, $m_{1}^{Y}\left(x, u\right)$, $m_{0}^{S}\left(x, u\right)$, $m_{1}^{S}\left(x, u\right)$, and $\Delta_{S}\left(x, u\right)$ without extra assumptions, implying that I cannot identify the bounds for $\Delta_{Y^{*}}^{OO}\left(x, u\right)$ given by Corollaries (ref) or (ref). There are two solutions for this lack of identification: I can non-parametrically bound those four objects (Mogstad2017) or I can impose flexible parametric assumptions (Brinch2017) to point identify them. While the first approach is discussed in Subsection (ref), the second one is detailed in Subsection (ref).
For any $u \in \left[0, 1\right]$ and $x \in \mathcal{X}$, I can bound $m_{0}^{S}\left(x, u\right)$, $m_{1}^{S}\left(x, u\right)$, $\Delta_{S}\left(x, u\right)$, $m_{0}^{Y}\left(x, u\right)$, $m_{1}^{Y}\left(x, u\right)$ and $\Delta_{Y}\left(x, u\right)$ using the machinery proposed by Mogstad2017. To do so, fix $A \in \left\lbrace S, Y \right\rbrace$ and $d \in \left\lbrace 0, 1 \right\rbrace$ and define the pair of functions $m^{A} \coloneqq \left(m_{0}^{A}, m_{1}^{A}\right)$ and the set of admissible MTR functions $\mathcal{M}^{A} \ni m^{A}$. For example, in the case of a binary function, the admissible set would be $\mathcal{M}^{A} = \left[0,1\right]^{\mathcal{X} \times \left[0, 1\right]} \times \left[0,1\right]^{\mathcal{X} \times \left[0, 1\right]}$ and, in the case of the selection indicator, this set would be further restricted by Assumption (ref) to $$\mathcal{M}^{A} = \left\lbrace \left(m_{0}^{A}, m_{1}^{A}\right) \in \left[0,1\right]^{\mathcal{X} \times \left[0, 1\right]} \times \left[0,1\right]^{\mathcal{X} \times \left[0, 1\right]} \colon m_{1}^{A}\left(x, u\right) \geq m_{0}^{A}\left(x, u\right) \hspace{5pt} \forall \left(x, u\right) \in \mathcal{X} \times \left[0, 1\right] \right\rbrace.$$ Moreover, define the function $\Gamma_{A}^{*} \colon \mathcal{M}^{A} \rightarrow \mathbb{R}$ as:
and observe that $\Gamma_{A}^{*}\left(m^{A}\right) = \Delta_{A}\left(x, u\right)$. Furthermore, define $\mathcal{G}_{A}$ to be a collection of known or identified measurable functions $g_{A} \colon \left\lbrace 0, 1 \right\rbrace \times \mathcal{Z} \rightarrow \mathbb{R}$ whose second moment is finite. For each IV-like specification $g_{A} \in \mathcal{G}_{A}$, define also $\beta_{g_{A}} \coloneqq \mathbb{E}\left[g_{A}\left(D, Z\right) A \left\vert X = x \right.\right]$. According to Mogstad2017, the function $\Gamma_{g_{A}} \colon \mathcal{M}^{A} \rightarrow \mathbb{R}$, defined as
satisfies $\Gamma_{g_{A}}\left(m^{A}\right) = \beta_{g_{A}}$. As a result, $m^{A}$ must lie in the set $\mathcal{M}_{\mathcal{G}_{A}}$ of admissible functions that satisfy the restrictions imposed by the data through the IV-like specifications, where:
Assuming that $\mathcal{M}^{A}$ is convex and $M_{\mathcal{G}_{A}} \neq \emptyset$ for every $A \in \left\lbrace S, Y \right\rbrace$, Mogstad2017 show that:
Based on this result, I can also define bounds for the MTR functions as $$\left(\overline{m_{0}^{A}\left(x, u\right)}, \underline{m_{1}^{A}\left(x, u\right)}\right) \coloneqq \operatorname*{\arg\!\inf}\limits_{\tilde{m}^{A} \in \mathcal{M}_{\mathcal{G}_{A}}} \Gamma_{A}^{*} \left( \tilde{m}^{A} \right) \text{ and } \left(\underline{m_{0}^{A}\left(x, u\right)}, \overline{m_{1}^{A}\left(x, u\right)}\right) \coloneqq \operatorname*{\arg\!\sup}\limits_{\tilde{m}^{A} \in \mathcal{M}_{\mathcal{G}_{A}}} \Gamma_{A}^{*} \left( \tilde{m}^{A} \right),$$ where
As a consequence, I can combine Corollaries (ref) and (ref) and inequalities (ref) and (ref) to provide a non-parametrically identified outer set around $\Delta_{Y^{*}}^{OO}\left(x, u\right)$, that contains the true target parameter $\Delta_{Y^{*}}^{OO}\left(x, u\right)$ by construction. However, the cost of non-parametric partial identification of $m_{0}^{S}\left(x, u\right)$, $m_{1}^{S}\left(x, u\right)$, $\Delta_{S}\left(x, u\right)$, $m_{0}^{Y}\left(x, u\right)$, $m_{1}^{Y}\left(x, u\right)$ and $\Delta_{Y}\left(x, u\right)$ is losing the pointwise sharpness of the bounds around the target parameter $\Delta_{Y^{*}}^{OO}\left(x, u\right)$.
The fully non-parametric approach explained in Subsection (ref) may provide an uninformative outer set (e.g., equal to $\overline{y}^{*} - \underline{y}^{*}$ or $\underline{y}^{*} - \overline{y}^{*}$ when the support of the potential outcome is bounded). In such cases, parametric assumptions on the marginal treatment response functions may buy a lot of identifying power. Although restrictive in principle, parametric assumptions may be flexible enough to provide credible bounds for $\Delta_{Y^{*}}^{OO}\left(x, u\right)$, as illustrated by Brinch2017.
I fix $x \in \mathcal{X}$ and assume that the support of the propensity score $P\left(x, Z\right)$ is discrete and given by $\mathcal{P}_{x} = \left\lbrace p_{x, 1}, \ldots, p_{x,N} \right\rbrace$ for some $N \in \mathbb{N}$. I could directly apply the identification strategy proposed by Brinch2017 by assuming that the MTR functions associated with $Y$ and $S$ are polynomial functions of $U$. However, this assumption is problematic for binary variables, such as the selection indicator $S$. For this reason, I make a small modification to the procedure created by Brinch2017: for $d \in \left\lbrace 0, 1 \right\rbrace$ and $A \in \left\lbrace Y, S \right\rbrace$, the MTR function is given by
for any $u \in \left[0, 1\right]$, where $\Theta_{x}^{A} \subset \mathbb{R}^{2L}$ is a set of feasible parameters, $L \in \left\lbrace 1, \ldots, N \right\rbrace$ is the number of parameters for each treatment group $d$, $\left(\boldsymbol{\theta}_{x, 0}^{A}, \boldsymbol{\theta}_{x, 1}^{A}\right) \in \Theta_{x}^{A}$ is a vector of pseudo-true unknown parameters, and $M^{A} \colon \left[0, 1\right] \times \mathbb{R}^{2L} \rightarrow \mathbb{R}$ is a known function. For instance, in the case of a binary variable, a reasonable choice of $M^{A}$ is the Bernstein Polynomial $\left(M^{A}\left(u, \boldsymbol{\theta}_{x, d}^{A}\right) = \sum_{l = 0}^{L - 1} \theta_{x, d, l}^{A} \cdot {L - 1 \choose l} \cdot u^{l} \cdot \left(1 - u\right)^{L - 1 - l}\right)$ with feasible set $\Theta_{x}^{A} = \left[0, 1\right]^{2L}$. In the case of the selection indicator, the feasible set would be further restricted by Assumption (ref) to $\Theta_{x}^{A} = \left\lbrace \left(\boldsymbol{\tilde{\theta}}_{x, 0}^{A}, \boldsymbol{\tilde{\theta}}_{x, 1}^{A}\right) \in \left[0, 1\right]^{2L} \colon \boldsymbol{\tilde{\theta}}_{x, 1}^{A} \geq \boldsymbol{\tilde{\theta}}_{x, 0}^{A} \right\rbrace$. I stress that the only difference between the Bernstein polynomial model and the simple polynomial model proposed by Brinch2017 is that it is easier to impose feasibility restrictions on the former model.
Back to the parametric model given by equation (ref), I define the parameters $\left(\boldsymbol{\theta}_{x, 0}^{A}, \boldsymbol{\theta}_{x, 1}^{A}\right)$ as pseudo-true parameters in the sense that the parametric model in equation (ref) is an approximation to the true data generating process via the moments $\mathbb{E}\left[A \left\vert X = x, P\left(W\right) = p_{n}, D = d \right.\right]$ for any $d \in \left\lbrace 0, 1 \right\rbrace$ and $n \in \left\lbrace 1, \ldots, N \right\rbrace$. Formally, I define
Note that, to estimate parameters $\left(\boldsymbol{\theta}_{x, 0}^{A}, \boldsymbol{\theta}_{x, 1}^{A}\right)$, I can simply use the sample analogue of equation (ref), i.e., I only have to estimate a constrained OLS regression whose restrictions are given by the set $\Theta_{x}^{A}$. If the model restrictions imposed through the set of feasible parameters $\Theta_{x}^{A}$ are valid and $L = N$, then my parametric model collapses to the model proposed by Brinch2017 and I find that\footnote{Appendix (ref) contains a proof of this claim.}, for any $p_{n} \in \mathcal{P}_{x}$,
I can then combine Corollaries (ref) and (ref) and equations (ref) and (ref) to bound $\Delta_{Y^{*}}^{OO}\left(x, u\right)$.
I focus on analyzing the Marginal Treatment Effect of the Job Corps Training Program (JCTP) on wages for the always-employed subpopulation ($MTE^{OO}$). This program provides free education and vocational training to individuals who are legal residents of the U.S., are between the ages of 16 and 24 and come from a low-income household (Schochet2001 and Lee2009). Besides receiving education and vocational training, the trainees reside in the Job Corps center, that offers meals and a small cash allowance.
In the mid 1990's, the U.S. Department of Labor hired Mathematica Policy Resarch, Inc., to evaluate the JCTP through a randomized experiment. According to Chen2015, eligible people who applied to JCTP for the first time between November 1994 and December 1995 (80,833 applicants) were randomly assigned into a treatment group and a control group. People in the control group (5,977) were embargoed from the program for 3 years, while those in the treatment group (74,856) were allowed to enroll in JC. However, in this randomized control trial, there was non-compliance (selection into treatment) because some individuals in the treated group decided not to participate in the program and some individuals in the control group were able to attend the JCTP even though they were officially embargoed.
To evaluate the JCTP, I start by describing the dataset, providing summary statistics and, most importantly, formally testing the assumptions that the potential treatment status is monotone on the instrument (equation (ref)) and that the potential employment (sample selection status) is positively monotone on the treatment (Assumption (ref)) using the test elaborated by Machado2018. I then estimate and discuss the marginal treatment responses and effects on employment and labor earnings using the parametric tool developed by Brinch2017. Finally, I estimate and discuss the bounds for the $MTE^{OO}$ on wages without and with the mean dominance assumption (Assumption (ref)), given, respectively, by Corollaries (ref) and (ref).
The publicly available National Job Corps Study (NJCS) sample contains 15,386 individuals --- all 5,977 control group individuals and 9,409 randomly selected treatment group individuals. All of them were interviewed at random assignment and at 12, 30 and 48 months after random assignment. Following Lee2009, I only keep individuals with non-missing values for weekly earnings and weekly hours worked for every week after randomization (9,145). Following Chen2015, my instrument ($Z$) is random treatment assignment and my treatment dummy ($D$) is an indicator variable that is equal to one if the individual was ever enrolled in the JCTP during the 208 weeks after random assignment. Since this variable has 51 missing values, the final sample size is 9,094 observations.
The dataset contains information about demographic covariates (sex, age, race, marriage, number of children, years of schooling, criminal behavior, personal income) and pre- and post-treatment labor market outcomes (employment and earnings). Following Chen2015, hourly wages at week 208 are created by dividing weekly earnings by weekly hours worked at that week, implying that a missing wage is equivalent to zero weekly hours worked. I consider the person to be unemployed ($S = 0$) when the wage is missing and to be employed ($S = 1$) when the wage is non-missing. Differently from Lee2009 and Chen2015, who use log hourly wages as their main outcome variable, my outcome of interest ($Y^{*}$) is the level of the hourly wage because Assumption (ref).1 requires that the support $\mathcal{Y}^{*}$ has a finite lower bound. As a consequence, the observable outcome $Y$ is defined as hourly labor earnings. Finally, I use the NJCS design weights in my empirical analysis because some subpopulations were randomized with different, but known, probabilities (Schochet2001).
Considering the results found by Flores-Lagunes2010, who focus on explaining the negative but insignificant effects on employment and labor earnings for the Hispanic subpopulation, I separately analyze two subsamples from the NJCS sample: the Non-Hispanics subsample and the Hispanics subsample. Table (ref) shows descriptive statistics for both subsamples. Note that, as expected, the pre-treatment covariates are, on average, very similar between the groups defined by the random treatment assignment. Consequently, both subsamples maintain the balance of baseline variables. However, when comparing Non-Hispanics and Hispanics, I find numerically small differences with respect to the variables female, never married, has children, ever arrested, has a job at baseline, and had a job.
Table (ref) shows preliminary effects within the Non-Hispanic and the Hispanic subsamples. The first row shows that a large number of individuals did not comply to their treatment assignment. As is expected for any voluntary treatment, a large share of individuals (around 30% for both subsamples) decided not to take the treatment even though they were assigned to the treatment group. There are also some individuals (5% among Non-Hispanics and 3% among Hispanics) who attended the JCTP even though they were embargoed. Moreover, the instrument (treatment assignment) is clearly strong for both subsamples, suggesting that Assumption (ref) is plausible in this context. When analyzing the treatment effects and similarly to the previous literature (e.g., Schochet2008, Flores-Lagunes2010 and Chen2015), we find that the JCTP has a positive and significant effect on Non-Hispanics and a negative but insignificant effect on Hispanics.
This last result, particularly with respect to the employment status, is important for my analysis. Similarly to Lee2009 and Chen2015, I assume that the effect of the treatment on employment (i.e., sample selection) is monotone and positive. However, a negative effect of JCTP on employment is evidence against this assumption as discussed by Flores-Lagunes2010 and Chen2015. For this reason, I formally test Assumption (ref). To do so, I implement the procedure developed by Machado2018, that simultaneously tests instrument exogeneity (Assumption (ref)), monotonicity of treatment take-up on treatment assignment (equation (ref)) and monotonicity of employment on the treatment (equation (ref)). Their procedure also uses this last test as a gate-keeper to test that the effect of the treatment on employment is positive (Assumption (ref)).
In a more detailed way, the test proposed by Machado2018 has three steps. In the first step, the null hypothesis is that the instrument is not exogenous, or treatment take-up is not monotone on treatment assignment, or employment is not monotone on treatment take-up. As a consequence, the alternative hypothesis is that Assumption (ref) and equations (ref) and (ref) hold. In the second step, that is implemented only if the first step rejects its null hypothesis, the second null hypothesis is that the effect of the treatment on employment is non-positive. Consequently, its alternative hypothesis is that Assumptions (ref) and (ref) and equations (ref) and (ref) hold. Finally, in the third step, that is implemented only if the second step does not reject its null hypothesis, the third null hypothesis is that the effect of the treatment on employment is non-negative. Consequently, its alternative hypothesis is that, while Assumption (ref) and equations (ref) and (ref) are valid, Assumption (ref) holds in the opposite direction (see Assumption (ref)).
Table (ref) shows the results of the test described above. Within the Non-Hispanics subsample, steps 1 and 2 reject their null hypotheses at the 1%-significance level, implying that Assumptions (ref) and (ref) and equations (ref) and (ref) are plausible given the data. Consequently, it is reasonable to use Corollary (ref) to bound the $MTE^{OO}$ of the JCTP on wages within the Non-Hispanics subsample. For the Hispanics subsample, step 1 rejects its null hypothesis at the 1%-significance level, while neither step 2 nor step 3 reject their null hypotheses at the 10%-significance level. As a consequence, Assumption (ref) and equations (ref) and (ref) are plausible given the data, but it seems that there is no effect of the treatment on employment, i.e., $S_{1} = S_{0}$ for all individuals. With no differential sample selection for the Hispanic population, point identification of the MTE of interest is trivial as discussed immediately after Proposition (ref). For this reason, I focus my empirical analysis on the Non-Hispanic subsample.
As a preliminary step to estimate the bounds for the $MTE^{OO}$ of the JCTP on hourly wages within the Non-Hispanic subsample, I need to estimate the MTR functions on employment and hourly labor earnings, i.e., I need to estimate the functions $m_{0}^{S}$, $m_{1}^{S}$, $m_{0}^{Y}$, and $m_{1}^{Y}$. To do so, I use the procedure described in Subsection (ref), that adapts the method developed by Brinch2017 to a constrained framework. Specifically, I model the MTR functions of $Y$ and $S$ using Bernstein polynomials with four parameters, i.e., $M^{A}\left(u, \boldsymbol{\theta}_{d}^{A}\right) = \theta_{d, 0}^{A} \cdot \left(1 - u\right) + \theta_{d, 1}^{A} \cdot u$ for any $A \in \left\lbrace Y, S \right\rbrace$ and $d \in \left\lbrace 0, 1 \right\rbrace$ with feasible sets $\Theta^{Y} = \mathbb{R}_{+}^{4}$ and $\Theta^{S} = \left\lbrace \left(\boldsymbol{\theta}_{0}^{S}, \boldsymbol{\theta}_{1}^{S}\right) \in \left[0, 1\right]^{4} \colon \boldsymbol{\theta}_{1}^{S} \geq \boldsymbol{\theta}_{0}^{S} \right\rbrace$. To estimate $\left(\boldsymbol{\theta}_{0}^{A}, \boldsymbol{\theta}_{1}^{A}\right)$. I run the following constrained OLS model:\footnote{Appendix (ref) connects the OLS model (ref) to the minimization problem (ref) when the instrument is binary and there are no covariates. It also provides the explicit formula for the bounds in Corollaries (ref) and (ref) using the parametric model described in Subsection (ref). Appendix (ref) implements a Monte Carlo Simulation that analyzes the coverage rate of confidence intervals around the MTE bounds that are based on the OLS model (ref).}
where $e$ is the error term, $\theta_{0, 0}^{A} = a_{0}^{A} - b_{0}^{A}$, $\theta_{0, 1}^{A} = a_{0}^{A} + b_{0}^{A}$, $\theta_{1, 0}^{A} = a_{1}^{A}$, $\theta_{1, 1}^{A} = a_{1}^{A} + 2 \cdot b_{1}^{A}$ and the constraints on $\left(a_{0}^{A}, b_{0}^{A}, a_{1}^{A}, b_{1}^{A}\right)$ are given by $\Theta^{A}$.
Tabel (ref) reports the point-estimates and 90%-confidence intervals of the parametric models for the MTR functions on employment and hourly labor earnings. Note that the feasibility constraint $\theta_{1,0}^{S} \geq \theta_{0,0}^{S}$ is binding even though Assumption (ref) is plausible according to the test proposed by Machado2018. Moreover, for the upper bound of the 90%-confidence interval, the feasibility constraint $\theta_{1,0}^{S} \leq 1$ is also binding.
It is easier to understand and interpret those estimates using Figure (ref). The solid lines are the point-estimates of the MTR and MTE functions based on the parameters reported in Table (ref). The dotted lines are pointwise 90%-confidence intervals around the estimated functions based on 5,000 bootstrap repetitions. Blue colored lines are associated with treated potential outcomes, while red colored lines are associated with untreated outcomes. In Subfigure (ref), I find that, although the employment probability for the agents who are most likely to attend the JCTP is similar between treated and untreated individuals, the employment probability for the agents who are less likely to attend the JCTP is much higher for treated individuals than for untreated ones. As a consequence, the MTE on employment within the Non-Hispanic subsample (Subfigure (ref)) is increasing in the latent heterogeneity. Similarly, in Subfigure (ref), I find that, although expected hourly labor earnings for the agents who are most likely to attend the JCTP is similar between treated and untreated individuals, expected hourly labor earnings for the agents who are less likely to attend the JCTP is much higher for treated individuals than for untreated ones. As a consequence, the MTE on hourly labor earnings within the Non-Hispanic subsample (Subfigure (ref)) is increasing in the latent heterogeneity. I highlight that the shape of my estimated MTE functions are in line with the results by Chen2017, whose estimated upper bounds also suggest that the ATE on those variables is greater than the ATT.
To partially identify the $MTE^{OO}$ of the JCTP on wages within the Non-Hispanic subsample, I can combine the functions estimated in Subsection (ref) with Corollaries (ref) and (ref). While the first corollary imposes only assumptions that are valid by the experimental design (Assumption (ref)), technical (Assumptions (ref)-(ref)) or testable (Assumptions (ref) and (ref), and equation (ref)), Corollary (ref) additionally uses the Mean Dominance Assumption (ref). This last assumption imposes that the marginal treatment response function of wages when treated for the always-employed population is greater than the same object for the employed-only-when-treated population, implying a positive correlation between potential employment and potential wages, which is supported by standard models of labor supply.\footnote{Chen2015 discuss the connection between the Mean Dominance Assumption (ref) and the Labor Economics literature in a deeper way.}.
Another issue when estimating bounds for a parameter of interest is that there are two ways to construct confidence intervals. The conservative method finds the $\zeta$-confidence intervals around the upper and lower $MTE^{OO}$ bounds and then uses their upper most and lower most bounds to construct a confidence interval that contains the identified region with probability $\zeta$. Since the parameter of interest has to be inside the identified region, this confidence interval contain the parameter of interest with probability at least $\zeta$. An alternative method is proposed by Imbens2004, who directly construct a $\zeta$-confidence interval that contains the parameter of interest. Since they take into account that the parameter of interest has to be inside the identified region by construction, their confidence interval is tighter than the conservative method.
Figure (ref) shows the parametric bounds of the $MTE^{OO}$ on wages using Corollary (ref) (Subfigure (ref)) and using Corollary (ref) (Subfigure (ref)). The solid lines are the point-estimates of the parametric bounds of the MTE on wages, while the dotted lines are pointwise conservative 90%-confidence intervals around the identified region based on 5,000 bootstrap repetitions and the dashed lines are pointwise 90%-confidence intervals of the parameter of interest (Imbens2004) based also on 5,000 bootstraps repetitions.
As a way to understand the magnitude of the effects, I compare the estimated $MTE^{OO}$ bounds against the average observed hourly wage of the Non-Hispanics assigned to the control group, \$7.72. Note that the lower bounds that do not use the mean dominance assumption (Subfigure (ref)) are implausibly negative. Even for the agents who are the most likely to attend the JCTP, the lower bound of the $MTE^{OO}$ on wages (-\$6.51) imply that the JCTP would drive their hourly wages almost to zero. This implausibly negative lower bound is based on the worst-case scenario that unrealistically imposes that the treated potential wage for the always-employed subpopulation is equal to zero.
By imposing the Mean Dominance Assumption (ref), I rule out this extreme case by assuming that there is positive selection into employment. As a consequence, I can increase the lower bound from equation (ref) to equation (ref), narrowing the bounds of the $MTE^{OO}$ on wages (Subfigure (ref)). Under this extra assumption, the $MTE^{OO}$ on wages is significant at the 10%-confidence level for latent heterogeneity values between 0.34 and 0.68 when I use the conservative confidence interval and between 0.35 and 0.73 when I use the confidence interval based on Imbens2004. Most interestingly, the point-estimate of the lower bound of the $MTE^{OO}$ on wages is decreasing in the likelihood of attending the JCTP.
To better understand the magnitude of those effects and compare my results with the previous literature, I summarize the bounds for the $MTE^{OO}$ function using four key parameters --- $ATE^{OO}$, $ATT^{OO}$, $ATU^{OO}$ and $LATE^{OO}$ --- that are described in Tables (ref) and (ref) as integrals of the $MTE^{OO}$ function. Table (ref) reports those bounds in brackets, the 90%-conservative confidence intervals of the identified region in parenthesis and the 90%-confidence intervals of the parameter of interest (Imbens2004) in braces. As expected, the bounds without the mean dominance assumption are wide and uninformative, while, when imposing Assumption (ref), all parameters but the $ATT^{OO}$ are significant at 10% according to both types of confidence intervals.
I stress that my $LATE^{OO}$ estimates represent an effect between 7.51% and 24.74% of the average observed hourly wage of the Non-Hispanics assigned to the control group, which are comparable to the bounds of the $LATE^{OO}$ parameter derived by Chen2015 --- approximately between 7.7% and 17.5% under a similar set of assumptions. The finding that their bounds are tighter than mine for the $LATE^{OO}$ is not surprising because their method leverages all the available information to specifically identify the $LATE^{OO}$ while my tool bounds the $MTE^{OO}$ function and then flexibly bounds the other treatment effects for the always-employed population.
As a consequence of this flexibility, I can partially identify other treatment effects that may be policy-relevant. For example, the $ATE^{OO}$ is bounded between 7.90% and 29.53% of the average observed hourly wage of the Non-Hispanics assigned to the control group. Most interestingly, the $ATT^{OO}$ and the $ATU^{OO}$ are, respectively, bounded between 4.27% and 12.82%, and 9.20% and 38.86%, suggesting that the agents who do not attend the JCTP might be the ones who would benefit the most from it. This result is even stronger when we analyze the confidence intervals around the $ATT^{OO}$ and the $ATU^{OO}$: while the first treatment effect is not significantly different from zero, the second parameter is significantly different from zero. To conclude, I highlight that, even though the upper bound of the treatment effects on wages may be unrealistically large, the magnitude of the lower bounds are similar to the results found by Chen2017 and are reasonable when compared to ITT effects of 16.70% on earnings per week and of 9.87% on hours per week that are shown in Table (ref).
My main theoretical contribution provides pointwise sharp bounds for the MTE of interest within the always-observed subpopulation by imposing a monotonicity assumption that the treatment has a positive impact on sample selection for every agent. Those bounds are tightened by imposing an extra mean dominance assumption that the potential outcome when treated within the always-observed subpopulation is greater than or equal to the same parameter within the observed-only-when-treated subpopulation. Both bounds can be estimated using the LIV approach if the instrument is continuous, using a non-parametric outer set based on the method developed by Mogstad2017, or using a parametric model based on the strategy proposed by Brinch2017. Such bounds are useful to analyze many empirical problems that include endogenous self-selection into treatment and sample selection.
My main empirical findings suggest that the marginal treatment effect of the Job Corps Training Program (JCTP) on employment, hourly labor earnings and hourly wages increases with the latent heterogeneity variable within the Non-Hispanic group. More specifically, while MTEs for the agents who are the most likely to attend the JCTP are very small, the MTEs for the agents who are the least likely to attend the JCTP are considerably large. Economically, this result implies that the agents who are more likely to benefit from the JCTP are not attending it due to some unobserved constraint. A similar result is found by Chen2017, whose empirical evidence suggests that the effects of the JCTP on employment and labor earnings for never-takers are significantly positive. They argue that those agents are not enrolling at the JCTP due to family constraints (lack of childcare services), incomplete information on JCTP's benefits, overconfidence or personal preferences for non-enrollment. A more complete analysis of why agents who would benefit from attending the JCTP are not doing so is beyond the scope of this paper, but is an important question for future research because it may help policy makers to better target the JCTP to the population who would benefit the most from this program.
\singlespace
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{equation}{0}