Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
103,053 characters · 14 sections · 111 citation commands
Identifying Marginal Treatment Effects in the Presence of Sample Selection
\nonstopmode
}
\newsavebox{\tablebox} \newlength{\tableboxwidth}
This article presents identification results for the marginal treatment effect (MTE) when there is sample selection. We show that the MTE is partially identified for individuals who are always observed regardless of treatment, and derive uniformly sharp bounds on this parameter under three increasingly restrictive sets of assumptions. The first result imposes standard MTE assumptions with an unrestricted sample selection mechanism. The second set of conditions imposes monotonicity of the sample selection variable with respect to treatment, considerably shrinking the identified set. Finally, we incorporate a stochastic dominance assumption which tightens the lower bound for the MTE. Our analysis extends to discrete instruments. The results rely on a mixture reformulation of the problem where the mixture weights are identified, extending lee2009training's (lee2009training) trimming procedure to the MTE context. We propose estimators for the bounds derived and use data made available by Deb2006 to empirically illustrate the usefulness of our approach.
\
Keywords: Sample Selection, Instrumental Variable, Marginal Treatment Effect, Partial Identification, Principal Stratification, Program Evaluation, Mixture Models.
JEL Codes: C14, C31, C35.
\
\doublespacing
Many interesting applications in the treatment effects literature involve two simultaneous identification challenges: endogenous selection into treatment and sample selection. For instance, in labor economics, when a researcher wishes to evaluate the effect of a job training program on wages, she has to consider the individuals' decision to enroll in the training program as well as their decision to participate in the labor market. In the health sciences, the same identification challenges appear when analyzing the effect of a drug on well-being as the outcome of interest --- health status --- is observed only for those who survive. Moreover, in randomized control trials (RCTs), non-compliance and endogenous attrition in treated and control groups lead to the same identification concerns. This double selection problem is also present when analyzing the effect of attending college on wages, the effect of an educational intervention on short- and long-term outcomes, and the effect of procedural laws on litigation outcomes.\footnote{Training programs are studied by Heckman1999a, lee2009training and chen2015bounds. The college wage premium is analyzed by Altonji1993, Card1999 and carneiro2011estimating. Some education interventions are studied by Krueger2001, Angrist2006, Angrist2009, Chetty2011 and Dobbie2015. Medical treatments are analyzed by CASS1984, Sexton1984 and Health2004. Litigation outcomes are discussed by Helland2017. RCTs with attrition are illustrated by DeMel2013 and Angelucci2015.}
In this paper, we derive novel uniformly sharp bounds on the marginal treatment effect (MTE) for individuals who would self-select into the sample regardless of their treatment status ($MTE^{OO}$). To do so, we propose identification strategies under increasingly restrictive sets of assumptions that extend the MTE identification to scenarios with endogenous sample selection. Furthermore, the choice of treatment is allowed to be endogenous, and can be related to the sample selection mechanism. To address both identification challenges described above, we analyze a generalized sample selection model in which the realized outcome (e.g., wages) is observed only if the individual self-selects into the sample (employment status), and the treatment choice (training program participation) is observed for all individuals in the data being analyzed.
The $MTE$ and $MTE^{OO}$ provide important and intuitive measures of the treatment effect and its heterogeneity across the population. For example, consider a job training program. Training influences both the likelihood of employment and wages, which are observed only for individuals who are employed. The $MTE$ reflects the returns to training for individuals with different levels of the (latent) cost of attending the program. As a consequence, this parameter sheds light on the heterogeneity of the training program's effects, i.e., to understand who would benefit from taking extra training. This knowledge can be used to design policies focusing on targeting of the program, affordability and services offered. Common parameters evaluated in the literature --- such as the average treatment effect ($ATE$), the average treatment effect on the treated ($ATT$), the average treatment effect on the untreated ($ATU$), and the local average treatment effect ($LATE$) --- may not adequately capture this heterogeneity. For example, they could be positive even when most people are affected adversely by a policy, masking its effects.\footnote{The empirical relevance of the MTE is recently illustrated by carneiro2011estimating, Kline2016_head_start, Arnold2018, Cornelissen2018, Bhuller2019, Humphries2019 and Mountjoy2019 in different contexts.} Similarly, the $MTE^{OO}$ reflects the training program's effects at the intensive margin, i.e., to the group of individuals that are more attached to the labor force, and would be employed even if they had not attended the training program.
Our strategy to partially identify the $MTE^{OO}$ relies on standard assumptions regarding selection into treatment. Those are the same conditions usually imposed for identification of the MTE when there is no sample selection, i.e., there is an exogenous instrument excluded from the outcome determination, treatment choice is monotone in the instrument, and the propensity score is continuous.\footnote{See bjorklund1987, heckman1999, heckman2001, heckman2005structural and Andresen2018 for a detailed discussion.} We add the requirement that the instrument is also excluded from the sample selection mechanism. In the job training literature, a similar assumption is frequently imposed when analyzing RCTs with imperfect compliance, as the interest frequently lies on the effect on labor earnings and employment.
In this paper's first contribution, the partial identification result for the $MTE^{OO}$ leaves sample selection unrestricted, providing very general uniformly sharp bounds on that parameter. Our second result tightens the bounds around the $MTE^{OO}$ by exploiting a “monotonicity in selection” assumption. This condition is similar to the usual LATE monotonicity assumption and requires individuals to be at least as likely to be observed in the sample if they are treated. In the job training program example, this additional assumption imposes that the treatment can induce workers to join the labor force, but not the opposite.
The identified set is further reduced by imposing a stochastic dominance assumption to our third and final set of assumptions. This condition mandates that the subpopulation who self-selects into the sample regardless of the treatment status has higher potential outcomes if treated than the subpopulation who self-selects into the sample only when treated. Intuitively, workers with high attachment to the labor force (that would be employed regardless of receiving job training), would earn higher wages after the extra training than those that would choose to participate in the labor force only if they participated in the training program.
Identification relies on a reformulation of the potential outcomes' conditional probabilities as a mixture between the latent groups of individuals who are “always observed” and “observed only when treated.” This reformulation extends to the MTE case the trimming procedure proposed by Imai2008, lee2009training and chen2015bounds in the context of identifying the ATE and LATE. Crucially, since we are interested in the MTE, the trimming is based on the distribution of the potential outcome conditional on unobserved individual characteristics related to treatment receipt.
The results can be used to construct bounds for any treatment effect parameter that can be written as a weighted average of the $MTE^{OO}$. We derive new weights to obtain sharp bounds on the ATE, the ATT, the ATU, any LATE (imbensangrist94) and any policy-relevant treatment effect (PRTE, Heckman2001a) within the always-observed subpopulation. Differently from the weights for the case without sample selection (heckman2005structural, carneiro2009estimating, and carneiro2011estimating), the new weights must be integrated over the distribution of the latent heterogeneity for the always-observed subpopulation instead of its unconditional distribution. Furthermore, we show these weights are identified under the “monotonicity in selection” assumption if the support of the propensity score is the full unit interval.
Moreover, we propose and discuss nonparametric and parametric estimators for the bounds around the $MTE^{OO}$. The usefulness and feasibility of the parametric procedure is illustrated in an empirical example using the data set organized by Deb2006. We find that the effect of insurance plan choice on ambulatory expenditures depends negatively on the agents' relative cost of choosing managed care plans over fee for service plans. This corroborates the results found by Deb2006, indicating that agents endogenously self-select into managed care plans.
The main identification results are extended to the case where the researcher only has access to multi-valued discrete instruments. In this context, we derive nonparametric sharp bounds on a weighted average of the MTE. This result is conceptually similar to chen2015bounds, who provide an outer set for the LATE parameter when the instrument is binary.\footnote{An additional extension is presented in Appendix (ref), where we derive sharp bounds around a more general object of interest: the distributional marginal treatment effect (DMTE), which captures the effect of the treatment on the outcome's distribution for individuals at the margin for participation. The DMTE can then be used to derive bounds on the quantile version of the marginal treatment effect and many other parameters.}
This paper is organized as follows. Section (ref) presents the structural model and sample selection mechanism considered, followed by a discussion of the identifying assumptions. In Section (ref), we provide the identification results for the MTE bounds in the case of a continuous instrument under each set of assumptions while Section (ref) discusses general aspects of estimation for the bounds. Section (ref) illustrates an estimation procedure and the identifying power of each set of assumptions using data made available by Deb2006. Section (ref) presents identification results for the case of discrete instruments. Section (ref) concludes. The proofs, sharp testable implications of the model, extensions to DMTE, numerical example, proposed estimation methods' details, Monte Carlo simulations for the estimator's performance, economic models illustrating our identifying assumptions, and an alternative specification for our empirical example are presented in the appendix.
Following lee2009training, and chen2015bounds, we consider the generalized sample selection model, described in the potential outcomes framework:
where $Z$ is a vector of observable instrumental variables (e.g., random assignment of cash incentives for participation in a job training program) with support given by $\mathcal Z\subset \mathbb R^{d_{z}}$, $D$ is the treatment status indicator (job training program enrollment). The variable $Y^{*}$ is the possibly censored realized outcome variable (wages) with support $\mathcal Y \subset \mathbb R$, while $Y_{0}^{*}$ and $Y_{1}^{*}$ are the possibly censored potential outcomes when the person is untreated and treated, respectively. Similarly, $S$ is the realized sample selection indicator (employment status), and $S_{0}$ and $S_{1}$ are potential sample selection indicators when individuals are untreated and treated. Finally, $Y$ is the uncensored observed outcome (labor earnings), and $V$ represents unobserved individual characteristics (cognitive or social costs of attending the job training program). The researcher observes only the vector $\left(Y,D,S,Z\right)$, while $Y^{*}_{1}$, $Y^{*}_{0}$, $S_{1}$, $S_{0}$ and $V$ are latent variables.\footnote{For simplicity, we drop exogenous covariates from the model. All results derived in the paper hold conditionally on covariates.} This model is a generalization of that considered in heckman1999, heckman2001, heckman2005structural to the sample selection setting.
The treatment status $D$ is connected to the instrument $Z$ and the unobserved characteristics $V$ through the unknown function $P: \mathcal{Z} \rightarrow \mathbb{R}$. In this model, we assume that the individual receives treatment when its idiosyncratic cost $V$ is less than or equal to a threshold $P(Z)$. This assumption is equivalent to imposing monotonicity of the treatment in the instrument $Z$ imbensangrist94 as shown by vytlacil2002. This setup is similar to the one proposed by heckman2005structural and leads to the definition of the MTE,
for any $p \in \left[0, 1\right]$.
In the setting analyzed here, the task of learning about the MTE is further complicated by the potential for nonrandom sample selection. As pointed out by lee2009training, even in the simpler case of the ATE, point identification is no longer possible even if the treatment is randomly assigned, leading him to derive bounds for the ATE. This paper combines the insights of these literatures to develop sharp bounds for the MTE under sample selection while allowing for treatment to be endogenously determined.
Similarly to the compliance groups defined by imbensangrist94, we define four latent groups based on the potential sample selection indicators. The subpopulations are defined as: always-observed ($S_{0} = 1, S_{1} = 1$), observed-only-when-treated ($S_{0} = 0, S_{1} = 1$), observed-only-when-untreated ($S_{0} = 1, S_{1} = 0$), and never-observed ($S_{0} = 0, S_{1} = 0$).\footnote{Since the conditioning subpopulation is determined by post-treatment outcomes, our work is also connected to the statistical literature known as principal stratification frangakis2002principal, in which the four latent groups would be called strata.} Those subgroups are summarized in Table (ref).
Following Zhang2008 and lee2009training, we focus on the always-observed subpopulation $\left(S_{0} = 1, S_{1} = 1\right)$. Importantly, this subpopulation is the only group whose censored potential outcomes are observed in both treatment arms. For the other three subpopulations, treatment effect parameters are not point-identified or bounded in a non-trivial way without further parametric assumptions, since at least one of the potential outcomes ($Y_{0}^{*}$ or $Y_{1}^{*}$) is never observed.\footnote{In some applications, the potential censored outcome $Y_{d}^{*}$ is not even properly defined when $S_{d} = 0$ for $d \in \left\lbrace 0, 1 \right\rbrace$, e.g., analyzing the impact of a medical treatment on a health quality measure where selection is given by whether the patient is alive} Since our focus is on a fully non-parametric identification strategy, we do not discuss parametric identification of unconditional treatment effect parameters or for treatment effect parameters associated with the latent groups $ON$, $NO$ and $NN$.
Our target parameter is the MTE function for the subpopulation who is always observed ($MTE^{OO} \colon \left[0,1\right] \rightarrow \mathbb{R}$):
for any $p \in \left[0, 1\right]$. Note that this parameter captures the intensive margin of the treatment effect.\footnote{If the researcher is interested in the extensive margin of the treatment effect, captured by the MTE on the observed outcome ($\mathbb{E}\left[Y_{1} - Y_{0} \left\vert V = p\right.\right]$) and by the MTE on the selection indicator ($\mathbb{E}\left[S_{1} - S_{0} \left\vert V = p\right.\right]$), she can apply the identification strategies described by Heckman2006, Brinch2017, Mogstad2017 and Andresen2018.} For example, when evaluating the effect of a job training program on wages, this parameter captures the effect of the training program on the wage of a worker who is employed regardless of her treatment status.\footnote{This conditional parameter is interesting for a policy maker and can offer meaningful information regarding treatment and its targeting. For example, the effect of a training program on individuals' wages reflects their productivity and the overall economic welfare. If a training program only affects overall earnings by attracting more participants to the labor market without any effect on productivity, the program's targeting and overall benefit to society should consider this aspect.} In medical applications where selection is due to the death of a patient, this parameter is the effect on health quality for the subpopulation who survives regardless of treatment status. In the education literature where sample selection is due to students quitting school, this parameter is the effect on test scores for the subpopulation who does not drop out of school in any case. Moreover, the definition of the $MTE^{OO}$ does not depend on the instrument being used, implying that our target parameter is policy-invariant. This feature is an advantage in comparison with the $LATE^{OO}$ chen2015bounds, which is not policy-invariant. Although the $MTE^{OO}$ function shares the policy invariance property with the usual $MTE$ function, it has one important drawback: the definition of $MTE^{OO}$ conditions on a latent group $\left(S_{0} = 1, S_{1} = 1 \right)$ to simultaneously address the selection-into-treatment and sample selection problems.
Analogously to lee2009training, identification of $MTE^{OO}$ is complex because sample selection is nonrandom and possibly impacted by the treatment. To address this issue, we consider three sets of increasingly restrictive assumptions that allow partial identification of the target parameter, shrinking the identified sets as the assumptions are strengthened.\footnote{According to Tamer2010, this approach to identification “characterizes the informational content of various assumptions by providing a menu of estimates, each based on different sets of assumptions, some of which are plausible and some of which are not.” Empirically, this approach is also illustrated by Kline2016.}
Assumptions (ref)-(ref) are sufficient to partially identify the $MTE^{OO}$ function.
Assumption (ref) is a modification of the IV independence assumption to account for sample selection. Instead of assuming that the instrument is independent of the latent heterogeneity and of the potential outcomes only heckman2005structural, we also assume independence of the potential sample selection indicators. Intuitively, we rely on changes in $Z$ shifting treatment status and, hence, sample participation to identify the marginal treatment effect bounds. Such an assumption is common in the empirical literature. A researcher interested in analyzing the impact of a job training program on wages and employment status, usually imposes a similar restriction on the marginal distributions of labor earnings and employment.
Assumption (ref) is important for our bounding strategy of the MTE across values of the latent variable, $V$. In Section (ref), we relax this assumption to allow for discrete instruments, and we show how our methodology can be used to bound the LATE, instead of the MTE.
Assumptions (ref)-(ref) are technical assumptions to ensure that the objects of interest are well-defined and are common in the literature about MTE Heckman2006. Assumption (ref) is crucial for the identification results and requires that there are always-observed individuals for all possible values of the unobserved heterogeneity $V$. This can be restrictive in practice, ruling out the calculation of the MTE bounds for ranges of $V$ in which receipt of treatment determines sample participation heavily. Assumption (ref) can be seen as a normalization if one assumes that the latent variable $V$ is absolutely continuous. Under the same normalization, the image of the function $P: \mathcal{Z} \rightarrow \mathbb{R}$ is contained in the unit interval.
Assumptions (ref)-(ref) form our first set of assumptions required for partial identification of the MTE for the always-observed individuals. Under those assumptions, the function $P(z)$ is identified and is equal to the propensity score $\mathbb P\left[D=1|Z=z\right]$ heckman2005structural. Indeed, $\mathbb P\left[D=1|Z=z\right]=\mathbb P\left[V\leq P(z)|Z=z\right]=\mathbb P\left[V\leq P(z)\right]=P(z)$, where the second equality holds under Assumption (ref) and the last holds under Assumption (ref). This first set of restrictions partially identifies $MTE^{OO}$, as presented in Section (ref).
We also stress that the identified set can be substantially tightened by imposing that the sample selection mechanism is monotone in the treatment.
This monotonicity assumption rules out the existence of the observed-only-when-untreated subpopulation and is commonly used in the sample selection literature (lee2009training, chen2015bounds).\footnote{As in lee2009training, this assumption can be stated as $S_1\geq S_0$ with probability 1. For the sake of simplicity, we assume it to hold for all individuals. Manski1997 and Manski2000 refer to this assumption as the “monotone treatment response” assumption. All results can be stated with some straightforward changes if the inequality in Assumption (ref) holds in the opposite direction.} To obtain some intuition on the mechanisms behind this assumption, consider the job training program example. An individual is employed when her job search skills $\vartheta(D)$, a function of training take-up, are above a threshold $U_{S}$ so that $S=\mathbbm{1}\left\{\vartheta(D) \geq U_S\right\}.$ Additionally, suppose that attending the job training program does not decrease someone's job search skills, i.e., $\vartheta (1) \geq \vartheta (0)$, making it more likely that program's trainees would be observed in the data. In such a case, Assumption (ref) holds. However, if attending the job training program raises the agents' reservation wages or if the lost labor market experience is very costly in terms of job finding, this assumption may not hold.
Assumptions (ref)-(ref) form our second set of identification assumptions and lead to the bounds for $MTE^{OO}$ that are the main result of this paper, presented in Proposition (ref). This second set of assumptions has a testable implication: the treatment positively affects sample selection, i.e., $E[S_{1}-S_{0}|V=p] \geq 0$, implying
In other words, the share of the population for which the outcome is observed rises with $p$. We discuss further testable implications in Section (ref), and formally characterize sharp testable implications arising from those assumptions in Appendix (ref). In the job training example, this testable implication means that the likelihood of employment increases with the probability of attending the training program.
We can further shrink the identified set around the $MTE^{OO}$, by adding Assumption (ref) and completing the third set of identifying assumptions.
This dominance assumption imposes that the always-observed subpopulation has higher potential censored outcomes than the observed-only-when-treated group conditional on $V$. This type of assumption is common in the literature (Imai2008, Blanco2013, Huber2015, and Huber2017) and is intuitively based on the argument that some population sub-groups have more favorable underlying characteristics than others.\footnote{All of our results can be stated if the inequality in Assumption (ref) holds in the opposite direction, as it is the case if larger values of the outcome harms the agent. For example, the researcher might be interested on the effect of a drug on cholesterol levels and the selection is based on whether the patient is alive.} Naturally, the plausibility of Assumption (ref) depends on the empirical context. In some cases this stochastic dominance assumption could be hard to interpret and motivate empirically. However, this challenge can be confronted constructively in a layered policy analysis Manski2011, since we can offer a menu of estimates based on different assumptions, allowing the researcher to understand the continuum of information that we can gather about a specific economic parameter, as advocated by Tamer2010. In Appendix (ref), we provide a simple economic model related to our job training example to illustrate that Assumption (ref) may hold under plausible economic restrictions.
This section presents the main results of this paper, the identification for $MTE^{OO}(p)$ under the three different sets of assumptions described in Section (ref). As stepping stones, Subsection (ref) shows identification of the conditional joint distribution of $\left. \left(Y_{d}^{*}, S_{d}=1\right) \right\vert V$ for any $d \in \left\lbrace 0, 1 \right\rbrace$, while Subsection (ref) shows that the distribution of the potential outcomes can be seen as a mixture of latent groups, an important feature of the model. In the following subsections, we sharply bound the $MTE^{OO}$ under increasingly restrictive assumptions. First, we bound the $MTE^{OO}$ without imposing any assumption on the selection mechanism (Subsection (ref)). We then tighten those bounds by additionally imposing monotone sample selection (Subsection (ref)) and stochastic dominance (Subsection (ref)). For completeness, we also show that the $MTE^{OO}$ is point-identified under the “no selection effect” assumption (Remark (ref)) at the end of Subsection (ref). Finally, in Subsection (ref), we discuss how to sharply bound treatment effect parameters that can be written as weighted averages of the $MTE^{OO}(p)$.
Before we discuss the identification of the $MTE^{OO}$, we point-identify the conditional joint distribution of each potential outcome and sample selection for different levels of individual heterogeneity, $\left. \left(Y_{d}^{*}, S_{d}=1\right) \right\vert V$ for $d \in \left\lbrace 0, 1 \right\rbrace$.
Under Assumptions (ref)-(ref), for any $p \in \text{int }\mathcal{P}$ and any Borel set $A \subseteq \mathcal Y$, we have that
where the second equality follows from Assumption (ref), the third and fourth equalities follow from the Law of Iterated Expectations, and the last equality follows from Assumption (ref). By differentiating each side with respect to $p$, we point-identify the conditional distribution of $\left(Y_{1}^{*}, S_{1}=1\right)$ given $V=p$ as follows:
Similarly, we can show that
Note that since equations (ref) and (ref) reflect probabilities, they generate two testable implications for Assumptions (ref) and (ref):
for all Borel sets $A \subset \mathbb R$ and $p\in (0,1).$ Intuitively, for people with observable characteristics ($Z$) that indicate a higher likelihood of being treated, the share of treated (untreated) individuals that self-select into the sample increases (decreases) for any range of the outcome.\footnote{A formal characterization of the sharp testable implications implied by our model is given in Appendix (ref).}
Similarly to the local IV approach proposed by heckman2005structural, equations (ref) and (ref) can be used to point-identify the MTE on the probability of being observed $\left(\mathbb E[S_{1}-S_{0}|V=p]\right)$, capturing the extensive margin of the treatment. Note that, for $A=\mathcal Y$,
implying that $\mathbb E[S_{1}-S_{0}|V=p]=\frac{\partial \mathbb E[S|P(Z)=p]}{\partial p}$. This effect could be of interest in itself: for example, the researcher may want to evaluate whether a training program increases employment levels. We can also identify the MTE on the observed outcome,
We would like to disentangle the marginal treatment on the observed outcome into the extensive margin and the intensive margin. While the extensive margin is point-identified, we show that the intensive margin ($MTE^{OO}$) is partially identified by considering that the distribution of potential outcomes is a mixture of latent groups.
Fundamental to our identification strategy is recognizing that the observed treated (untreated) group is composed only by $OO$ and $NO$ ($ON$) types, as described in Table (ref). Hence, the conditional distribution $\left. Y_{1}^{*} \right\vert S_1 = 1, V = p$ can be written as the mixture of these latent distributions. For notational simplicity, let $\alpha(p) \equiv \frac{\mathbb{P}\left[OO|V=p\right]}{\mathbb P\left[S_{1}=1|V=p\right]}$ be the share of always-observed individuals among those for which $S_{1}=1$ conditional on $V=p$. Naturally, the remainder, $\frac{\mathbb P\left[NO|V=p\right]}{\mathbb P\left[S_{1}=1|V=p\right]}$, can be described as $1-\alpha(p)$. By the Law of Total Probability, we have that:
As a consequence, $\mathbb E[Y^{*}_{1}|S_{1}=1, V=p]$ is also a mixture of the expectation of $Y^{*}_{1}$ for the always-observed and for observed-only-when-treated given $V=p$,
Similarly, the conditional distribution of $\left. Y_{0}^{*} \right\vert S_{0} = 1, V = p$ is the mixture of $\left. Y_{d}^{*} \right\vert V = p$ for two latent groups, the always-observed and the observed-only-when-untreated group:
where $\beta(p) \equiv \frac{\mathbb{P}\left[OO|V=p\right]}{\mathbb P\left[S_{0}=1|V=p\right]}$.
We exploit these mixture representations to bound the marginal treatment response of the censored treated outcome within the always-observed subpopulation $\left(\mathbb{E}\left[\left. Y^{*}_{1} \right\vert S_{0} = 1, S_{1} = 1, V=p \right]\right)$ by considering the tails of the observed outcomes' distribution for treated individuals. The smallest attainable value of $\mathbb E[Y^{*}_{1}|S_{0} = 1, S_{1} = 1, V=p]$ is obtained when we consider the scenario in which the always-observed individuals are contained entirely in the left tail of mass $\alpha(p)$ of the outcome distribution, i.e., the lowest values of $Y^{*}_{1}$ among the subpopulation $\{S_{1}=1\}$ conditional on $V$ being equal to $p$. Respectively, the largest attainable value of $\mathbb E[Y^{*}_{1}|S_{0} = 1, S_{1} = 1, V=p]$ is obtained in the case that the always-observed individuals would be the right tail of the same distribution, getting the highest values of $Y^{*}_{1}$ on that subpopulation. This is the same intuition behind the trimming procedure suggested by lee2009training and chen2015bounds, but, differently from them, the trimmed distribution is conditional on a specific value for the latent heterogeneity variable. This type of trimming approach, used in the current and above papers, relies on results derived in a more general mixture model by horowitz1995.
Hence, $\mathbb{E}\left[\left. Y^{*}_{1} \right\vert S_0 = 1, S_1 = 1, V=p \right]$ lies within the interval $[LB_{1}(p),UB_{1}(p)],$ where
and $F^{-1}_{Y^{*}_{d}|S_{d}=1,V=p}(\cdot)$ is the quantile function of the distribution of $Y^{*}_{d}$ given $S_{d}=1 \text{ and } V=p$.
Similarly, the conditional distribution of $\left. Y_{0}^{*} \right\vert S_0 = 1, V = p$ can be written as the mixture of $\left. Y_{d}^{*} \right\vert V = p$ for two latent groups, the always-observed and the observed-only-when-untreated group. Analogously to the treated outcome, the marginal treatment response of the untreated outcome within the always-observed subpopulation $\left(\mathbb{E}\left[\left. Y^{*}_{0} \right\vert S_0 = 1, S_1 = 1, V=p \right]\right)$ lies within the interval $[LB_{0}(p),UB_{0}(p)],$ where
Combining the bounds around $\mathbb{E}\left[\left. Y^{*}_{1} \right\vert S_{0} = 1, S_{1} = 1, V=p \right]$ and $\mathbb{E}\left[\left. Y^{*}_{0} \right\vert S_{0} = 1, S_{1} = 1, V=p \right]$, we find that $MTE^{OO}\left(p\right)$ lies within the interval $$[LB_{1}(p) - UB_{0}(p), UB_{1}(p) - LB_{0}(p)].$$
In the next subsections, we investigate the bounds that are generated under the alternative sets of assumptions described in Section (ref). Intuitively, those assumptions impose different restrictions on the possible values of the mixture weights $\left(\alpha(p), \beta(p)\right)$, providing different sets of information about $E[Y^{*}_{1}-Y^{*}_{0}|S_{0} = 1, S_{1} = 1, V=p]$.
Initially, consider the case in which the researcher is only willing to consider Assumptions (ref)-(ref), leaving the sample selection mechanism unrestricted. To learn about $MTE^{OO}$, we need information about the share of always-observed individuals in the total population, $\mathbb{P} \left[S_{0} = 1, S_{1} = 1\right]$ and, hence, the conditional joint distribution of $\left. \left(S_{0}, S_{1}\right) \right\vert V = p$. However, we only have information about the conditional marginal distributions $\left. S_{0} \right\vert V = p$ and $\left. S_{1} \right\vert V = p$ based on equations (ref) and (ref). According to Imai2008 and Mullahy2018, the following Boole-Fréchet bounds are sharp around the share of always-observed individuals:
Combining this information with Equations (ref) and (ref), leads to Lemma (ref).
Note that, since $\Upsilon\left(p\right)$ provides the identified set of possible values for the share of always-observed individuals, we can obtain the equivalent range of possible values for $\alpha(p)$ and $\beta(p)$, the mixture weights described in Subsection (ref). For brevity, let $\mathbb{P} \left[S_{0} = 1, S_{1} = 1 \vert V = p\right]$ take any particular value, $\upsilon \in \Upsilon\left(p\right)$. Define,
Let the bounds in Equations (ref)-(ref), for specific values of $\alpha(p, \upsilon)$ and $\beta(p, \upsilon)$ in the identified set be written as:
Combining the bounds around $\mathbb{E}\left[\left. Y^{*}_{1} \right\vert S_0 = 1, S_1 = 1, V=p \right]$ and $\mathbb{E}\left[\left. Y^{*}_{0} \right\vert S_0 = 1, S_1 = 1, V=p \right]$, we find that $MTE^{OO}\left(p\right)$ lies within the interval $[LB_{1}(p, \upsilon) - UB_{0}(p, \upsilon), UB_{1}(p, \upsilon) - LB_{0}(p, \upsilon)]$ for a particular $\mathbb{P} \left[S_0 = 1, S_1 = 1 \vert V = p\right] = \upsilon$.
To bound the target parameter, we find worst- and best-case scenarios by varying the value $\upsilon$. Explicitly, $MTE^{OO}\left(p\right)$ is partially identified and lies within the interval
Note that $\upsilon$ has a monotone relationship to the mixture weights, which define the trimming points in the bounds. As previously discussed, higher values for $\alpha(p)$ ($\beta(p)$) indicate that a bigger share of the observed treated (untreated) population belongs to the always-observed latent group, thus providing more information and tighter bounds for the parameter of interest. Hence, we only need to focus on the scenario that generates the wider bounds, that is, the smallest admissible $\alpha(p)$ and $\beta(p)$. Let $\upsilon^\ell$ be the lower bound of $\Upsilon(p)$. We have:
Making the same argument to the upper bound, we can rewrite them as,
greatly simplifying our bounds, which need only to be evaluated at the end point of $\Upsilon(p)$.
We can combine these facts with equations (ref), (ref), (ref) and (ref) to propose the first identification result for $MTE^{OO}$, which does not impose meaningful restrictions on the sample selection mechanism.
In this subsection, we introduce monotonicity of sample selection in the treatment (Assumption (ref)), which can considerably shrink the identified set for $MTE^{OO}$. As discussed in Section (ref), under the monotonicity assumption, individuals who self-select into the sample when untreated $\left(S_{0} = 1\right)$ would also be observed if they had been treated, ruling out the subgroup $ON$. In other words, any untreated individuals observed on the sample are members of the always-observed latent subpopulation $\left(S_{0} = 1, S_{1} = 1\right)$. Formally, the following two events are identical: $\left\{S_{0}=1\right\}=\left\{S_{0}=1,S_{1}=1\right\}$, and the mixture weight for the untreated group, $\beta(p)$, equals one.
Consequently, $\mathbb P\left[S_{0}=1,S_{1} = 1|V=p\right]$ is point-identified by equation ((ref)), and we no longer need to rely on the partial identification results in Lemma (ref). Specifically, we have that
This result connects the conditional share of always observed individuals to changes on the conditional mass of observed untreated individuals when the propensity score increases. Looking at the second equality, we find that the conditional probability of being always observed is the difference between the increase in the share of observed treated individuals and the increase in the share of observed individuals when the propensity score is equal to $p$.
Since $\left\{S_{0}=1\right\}=\left\{S_{0}=1,S_{1}=1\right\}$ ($\beta(p)=1$), the distribution of $\left. \left( Y_0^{*}, S_{0}=1,S_{1}=1 \right)\right\vert V$ is equal to the distribution of $\left. \left( Y_0^{*}, S_{0}=1 \right)\right\vert V$, implying that
Note that the right-hand side of equation (ref) is point-identified according to equations (ref) and (ref). Consequently, the expectation $\mathbb E[Y^{*}_{0}|S_0 = 1, S_1 = 1,V=p]$ is also point-identified. Monotonicity also leads to point identification of the mixture weight, $\alpha(p)=\dfrac{\mathbb P\left[S_{0}=1,S_{1}=1|V=p\right]}{\mathbb{P} \left[S_1 = 1 \vert V = p \right]}$ by Equations (ref) and (ref).
Then, under monotonicity, the researcher has to obtain bounds only for the expected potential outcomes under treatment, which still can be written as a mixture of the always-observed and observed-only-when-treated latent subpopulations.
As discussed in Section (ref), the expectation $\mathbb E[Y^{*}_{1}|S_{0}=1,S_{1}=1, V=p]$ lies in the interval $[LB_{1}(p),UB_{1}(p)]$, given in Equations (ref)-(ref). Combining the bounds and identification results in equations (ref), (ref), (ref), (ref) and (ref), the following proposition holds:
In this section, we add the stochastic mean dominance assumption to tighten the identified set for $MTE^{OO}$ under Assumptions (ref)-(ref). Stochastic dominance and equation (ref) imply that
for any $y \in \mathcal{Y}$. As a consequence, the following inequality holds
This tightens the lower bound for $\mathbb E[Y^{*}_{1}|S_{0}=1,S_{1}=1, V=p]$ as we no longer need to focus on the lowest $\alpha(p)$ mass of outcomes as the lower bound, since the stochastic dominance assumption guarantees that the expectation of outcomes for the always observed subpopulation will be larger than the one of the observed treated individuals which mixes $OO$ and $NO$ types. Hence, $\mathbb E[Y^{*}_{1}|S_{0}=1,S_{1}=1, V=p]$ lies within the interval $[LB_{3}(p),UB_{3}(p)],$ where
The upper bound remains unchanged. Naturally, that leads to tighter identified sets relative to the ones in Proposition (ref), which are presented in the following proposition.
For completeness in the identification discussion, note that point identification of $MTE^{OO}$ is achieved if, in addition to assumptions (ref)-(ref), we assume that the always-observed and never-observed subpopulations are the only existing groups, i.e., $S_{0}=S_{1}$. This “no selection effect” assumption imposes that the treatment has no impact on sample selection. In the job training program context, this implies that workers' employment would not be affected by the program.
Under Assumptions (ref)-(ref) and “no selection effect”, the distributions of $\left. Y_{d}^{*} \right\vert S_{1} = 1, V$ for $d \in \left\lbrace 0, 1 \right\rbrace$ are exclusively composed of always-observed individuals ($\alpha(p)=\beta(p)=1$). Then,
where point-identification follows from equations (ref), (ref), (ref) and (ref). The expectations $\mathbb{E}[Y^{*}_{d}|S_0 = 1, S_1 = 1,V=p]$ are also point-identified. Then, $MTE^{OO}\left(p\right) = \frac{\frac{\partial \mathbb E[YS|P(Z)=p]}{\partial p}}{\frac{\partial \mathbb E[SD|P(Z)=p]}{\partial p}}$.
Alternatively, point-identification of the unconditional MTE is achieved if potential sample selection status and outcomes are independent given the unobserved characteristics,$(S_{0},S_{1})\ \rotatebox[origin=c]{90}{$\models$}\ (Y^{*}_{0},Y^{*}_{1})|V$. In this case, the distributions $\mathbb P\left[Y^{*}_{d}\leq y|V=p\right]$ are point-identified from Equations (ref) and (ref), and the unconditional MTE is point-identified as:
The partial identification results for the $MTE^{OO}$ are relevant for a vast array of empirical objectives. First, bounds for the $MTE^{OO}$ can illuminate the treatment effect's heterogeneity, allowing researchers to better understand who would benefit from a specific treatment. This feature is important because common parameters (e.g., $ATE$, $ATT$, $ATU$, and $LATE$ within the always-observed subpopulation) can be positive even when most people are adversely affected by a policy. Moreover, knowing, even partially, the $MTE^{OO}$ function can be useful to design policies that provide incentives to agents to take some treatment.
Second, the $MTE^{OO}$ bounds can be used to partially identify alternative treatment effect parameters that are described as a weighted integral of $MTE^{OO}$ because
where $t \in \left\lbrace 1, 2, 3 \right\rbrace$, $\underline{\Delta}_{t}$ and $\overline{\Delta}_{t}$ are described in Propositions (ref)-(ref), and $\omega(\cdot)$ is a known or identifiable weighting function. Those bounds are sharp, as summarized in Proposition (ref).\footnote{We focus on the case under the monotonicity restriction (Assumptions (ref)-(ref)) for brevity. Similar results hold under our other identifying sets of assumptions.}
Under Assumptions (ref)-(ref), Table (ref) shows the most relevant treatment effect parameters that are partially identified using Proposition (ref). The weights in Table (ref) are derived in Appendix (ref) and are identified if the support of the propensity score is the full unit interval, i.e., $\mathcal{P} = \left[0, 1\right]$.
Importantly, note that, differently from the weights for the case without sample selection (heckman2005structural, carneiro2009estimating, and carneiro2011estimating), these weights must be integrated over the distribution of the latent heterogeneity for the always-observed subpopulation instead of its unconditional distribution.
This section describes the general estimation steps for the bounds proposed in Proposition (ref), while details on two proposed methods are provided in Appendix (ref) and (ref). For brevity, we focus on the bounds identified under monotonicity of sample selection in the treatment (Assumptions (ref)-(ref)), as it is the most relevant (and feasible) case empirically. Estimators for the bounds in Proposition (ref) and Proposition (ref) are natural extensions. Appendix (ref) presents a Monte Carlo Simulation that evaluates the small sample properties of the estimator.
To estimate the bounds in Proposition (ref), it is useful to focus on the building blocks that are the foundation for the identification results. In particular, we need the CDFs:
for any $d \in \left\lbrace 0, 1 \right\rbrace$ and
Consequently, we need to estimate:
Furthermore, the estimation of the propensity score, $P(Z)$, is necessary to obtain the moments of the conditional distribution of the observed outcome.
Each of these components can be estimated by standard approaches, and multiple procedures might be valid depending on the assumptions and model structure the researcher is willing to impose. For example, one could resort to nonparametric methods to estimate $P(Z)$, $ \pi_{0}(p)$, $\pi_{1}(p)$, $\Gamma_{0}(p,y)$, and $\Gamma_{1}(p,y)$, avoiding functional form choices as discussed in Appendix (ref). While appealing, nonparametric estimation can be quite challenging in practice especially when covariates are added, as illustrated in the empirical application. Alternatively, we could estimate the parameters based on parametric functions for $P(Z)$, $\mathbb{P}\left[\left. S = 1, D = d \right\vert P\left(Z\right) = p \right]$, and a partition of the outcome's support $\gamma_{d}(p,k)=\mathbb{P}\left[\left. y_{k-1} \leq Y < y_{k}, S = 1, D = d \right\vert P\left(Z\right) = p \right]$ for $d=\{0,1\}$, $p \in \left[0, 1 \right]$, $k \in \left\lbrace 1,\ldots,K \right\rbrace$. The choice of estimator should be guided by the nature of the problem being studied and the data available to the researcher.
With estimates $\hat{P}(Z)$, $ \hat{\pi}_{0}(p)$, $\hat{\pi}_{1}(p)$, $\hat{\Gamma}_{d}(p,y)$, and $\hat{\gamma}_{d}(p,y)$ at hand, we can estimate $\alpha\left(p\right)$, by its sample analog $\hat{\alpha}\left(p\right) \coloneqq \dfrac{\hat{\pi}_{0}\left(p\right)}{\hat{\pi}_{1}\left(p\right)}.$ Finally, the estimated bounds $LB_{2}(p)$, $UB_{2}(p)$, can be obtained as
where $\hat{f}_{1}(p,k) \coloneqq \frac{\hat{\gamma}_{1}(p,k)}{\hat{\pi}_{1}(p)}$, $\hat{F}_{1}\left(p, y_{k}\right) \coloneqq \sum_{j = 2}^{k} \hat{f}_{1}(p, j)$ and $\overline{y}_{k}$ is the center point of each bin $[y_{k-1}, y_{k}]$ for any $k \in \left\lbrace 2,\ldots,K_{N} \right\rbrace$. Moreover, we can estimate $\mathbb E\left[\tilde{Y}_0|S=1,D=0,P(Z)=p\right]$ in Proposition (ref) using $ \hat{\Xi}_{OO, 0}(p) \coloneqq \sum_{k = 2}^{K_{N}} \overline{y}_{k} \cdot \hat{f}_{0}(p, k)$, where $\hat{f}_{0}(p,k) \coloneqq \frac{\hat{\gamma}_{0}(p,k)}{\hat{\pi}_{0}(p)}$. Naturally, the estimated $MTE^{OO}$ bounds can then be obtained by $\hat{\underline{\Delta}}_{2}\left(p\right) \coloneqq \widehat{LB}_{2}(p)-\hat{\Xi}_{OO, 0}(p)$ and $\hat{\overline{\Delta}}_{2}\left(p\right) \coloneqq \widehat{UB}_{2}(p)-\hat{\Xi}_{OO, 0}(p)$.
We summarize the estimation procedure in the following steps: \setcounter{bean}{0}
To illustrate the empirical usefulness of our partial identification strategy in a concrete application, we analyze the impact of insurance plan choice on ambulatory expenditures using the data set made available by Deb2006 through the Journal of Applied Econometrics' Data Archive.\footnote{The Journal of Applied Econometrics' Data Archive can be accessed at http://qed.econ.queensu.ca/jae/.} We simplify their analyzes in two important dimensions. To enforce a binary treatment, we follow Papadoulos2012 and combine Health Maintenance Organizations (HMO) and Preferred Provider Organization (PPO) in one treatment category (managed care) while the control group contains all individuals who choose fee-for-service (FFS) plans. Second, while Deb2006 analyze ambulatory and hospital expenditures, we focus only on the former since the large share of zero hospital expenditures Deb2006 imply that bounds around the $MTE^{OO}$ would be very wide as explained in Appendix (ref).
In this application, the treatment variable $D$ is equal to one if the person chooses a managed care plan and equal to zero if the person chooses an FFS plan. $Y_{0}^{*}$ and $Y_{1}^{*}$ represent the potential ambulatorial expenditures in dollars, that is only observed if the person seeks care. $S_{0}$ and $S_{1}$ represent the potential indicators for seeking care, i.e., for spending a positive amount of money on ambulatory services. The latent heterogeneity $V$ in our treatment choice model (Equation (ref)) can be interpreted as the individual's relative cost of choosing a managed care plan over FFS plan. Importantly, our treatment effect of interest $\left(MTE^{OO}\left(p\right) \coloneqq \mathbb{E}\left[Y_{1}^{*} - Y_{0}^{*} \left\vert V = p, S_{0} = 1, S_{1} = 1 \right.\right]\right)$ captures the intensive margin of the impact of managed care plans on ambulatory expenditures, that is, the increased intensity in use of services by those individuals that would seek ambulatory care in both insurance scenarios.
We use data from the 1996-2001 waves of the Medical Expenditure Panel Survey (MEPS) made available by Deb2006. The sample is restricted to employed individuals who bought a private insurance plan and whose age is between 21 and 64 years. In this data set, we observe 24 covariates.\footnote{The covariate variables are family size, age, squared age, years of schooling, income, female indicator, interaction term between age and female, African-American indicator, Hispanic indicator, marriage indicator, three geographic region dummies, metropolitan area indicator, three subjective health dummies, physical limitation indicator, number of chronic conditions, injury indicator and four year dummies. The descriptive statistics of all variables can be found in Table 1 in Deb2006.} Moreover, our instruments are spouse's age $\left(Z_{1}\right)$ and spouse's insurance plan type in the previous year $\left(Z_{2}\right)$.\footnote{Differently from Deb2006, we also combine the spouse’s insurance plan choice in a binary variable that is equal to one if the spouse chose a managed care plan and equal to zero if the spouse chose a FFS plan.} The discussion about instruments' validity follows Deb2006. Regarding instrument relevance, the argument is that insurance plans cover an entire family, so spouse's age and lagged choice of the spouse's insurance plan should be a determinant of plan choice. The instruments' independence relies on the assertion that conditional on own age and other individual characteristics, spouse's age should not affect personal medical expenditures directly. Moreover, since lagged spouse's choice was pre-determined, it should not impact personal medical expenditures directly either. Therefore, conditioning on covariates is crucial for the instrument's credibility and can be more clearly dealt by specifying parametric functions for the probabilities underlying the DGP.
We calculate bounds for $MTE^{OO}\left(p, x\right)$ based on parametric estimates for the functions $\mathbb{P}\left[\left. S = 1, D = d \right\vert X = x, P\left(Z\right) = p \right]$, and $\mathbb{P}\left[\left. y_{k-1} \leq Y < y_{k}, S = 1, D = d \right\vert X = x, P\left(Z\right) = p \right]$ for $d=\{0,1\}$, $p \in \left[0, 1 \right]$, $k \in \left\lbrace 1,\ldots,K \right\rbrace$ and covariate values $x$. These probabilities are modeled as logit functions that depend on a linear index of the covariates and a quadratic function of the propensity score, using 20 grid points for the outcome variable. The propensity score is also estimated using a logit model that depends linearly on covariate and instrumental variables. To enforce the common support assumption, we trim the top and bottom 1% of the overlapping estimated propensity score distribution. After estimating those probabilities, we estimate $\alpha\left(p, x\right)$, $\beta\left(p, x\right)$ and the bounds around $MTE^{OO}\left(p, x\right)$ for each covariate value $x$. Then, following Deb2006, we assume that the covariates are independent of $(S_{0}, S_{1},V)$ and average the bounds for $MTE^{OO}\left(p, x\right)$ across the sample using observed covariates values. By doing so, we recover the bounds for the unconditional $MTE^{OO}(p)$.\footnote{By averaging with respect to the observed density of the covariates, we compute bounds around a summary measure of the conditional $MTE$ for the always-observed subgroup: $SCMTE^{OO}\left(p\right) \coloneqq \int \mathbb{E}\left[\left. Y_{1}^{*} - Y_{0}^{*} \right\vert V = p, S_{0} = 1, S_{1} = 1, X = x^{\prime} \right] \, \text{d}F_{X}\left(x^{\prime}\right)$. If $X \rotatebox[origin=c]{90}{$\models$} \left(V = p, S_{0} = 1, S_{1} = 1\right)$ holds, then $SCMTE^{OO}\left(p\right) = MTE^{OO}\left(p\right)$, implying that the summary bounds are valid for the unconditional $MTE$ function for the always-observed subgroup. Importantly, Deb2006 assumed that the covariates are fully exogenous, implying that $X \rotatebox[origin=c]{90}{$\models$} \left(V = p, S_{0} = 1, S_{1} = 1\right)$ holds. carneiro2009estimating assume a similar exogeneity assumption. Alternatively, we can analyze the conditional $MTE^{OO}$ function for pre-specified values of the covariates. These results are available upon request.} For details on this parametric approach, see Appendix (ref).
We estimate bounds around the $MTE^{OO}$ function under three sets of assumptions: (i) no restrictions on the sample selection mechanism (Assumptions (ref)-(ref)), (ii) “monotonicity of sample selection in the treatment” (Assumptions (ref)-(ref)), and (iii) “monotonicity of sample selection in the treatment” and “stochastic dominance” (Assumptions (ref)-(ref)). We are interested in understanding how the different sets of assumptions impact the identified set for $MTE^{OO}$ and interpreting the heterogeneity captured by the MTE.
First, we analyze the $MTE^{OO}$ bounds based only on Assumptions (ref)-(ref) (Subsection (ref)). Subfigures (ref) and (ref) show the estimated proportion of the always-observed subpopulation within the observed-when-treated and observed-when-untreated groups, that is, $\alpha\left(p, \upsilon^\ell\right)$ and $\beta\left(p, \upsilon^\ell\right)$ in Proposition (ref). Importantly, in some regions of the support those proportions are equal to one or zero, implying that $MTE^{OO}$ is point-identified or not-identified, respectively. Subfigure (ref) shows the estimated bounds. For most of the propensity score's support, the bounds without restrictions on the sample selection mechanism are very wide or identification is lost. The estimated bounds do not rule out the possibility of homogeneous treatment effects, i.e., we can still place a constant function inside the bounds in Subfigure (ref).
In order to tighten those bounds, we consider restrictions on the sample selection mechanism. The “monotonicity of sample selection in the treatment” condition (Assumption (ref)) imposes that agents who would spend a positive amount of money on ambulatory services if allocated to a FFS plan would also have positive ambulatory expenditures if allocated to a managed care plan. This assumption's direction is in accordance with the results described by Deb2006, who found that individuals enrolled in HMOs and PPOs are more likely to seek care than FFS enrollees. Note that, under Assumption (ref), $\beta(p)$ is always equal to 1. Subfigure (ref) shows that, under Assumptions (ref)-(ref), the estimated proportion of the always-observed subpopulation within the observed-when-treated group $\left(\alpha\left(p\right)\right)$ is strictly positive everywhere in this example. Consequently, the $MTE^{OO}$ function is at least partially identified for the entire support of the propensity score, as can be seen in the bounds reported Subfigure (ref), illustrating the identifying power of the “monotonicity of sample selection in the treatment” assumption. As a result of tighter bounds, we can rule out the possibility of homogeneous treatment effects, i.e., we cannot fit a constant function inside the bounds in Subfigure (ref). Interestingly, the upper bound under monotonicity is decreasing. Hence, if the $MTE^{OO}$ function followed the pattern from the upper bound, our results would suggest that the agents who are more likely to enroll in a managed care plan are the ones who incur larger additional ambulatory expenses due to their choice of insurance coverage. This interpretation is compatible with individuals taking into account their potential expenditures when selecting their insurance plans, at least among those for which those expenditures will be positive regardless of their plan (always-observed), and provides extra support for the selectivity result found by Deb2006.
In order to further tighten the bounds around the $MTE^{OO}$ function, we impose the stochastic dominance assumption (Assumption (ref)). To interpret this assumption, recall that the “always-observed” ($OO$) individuals are those for whom ambulatory expenditures would be positive regardless of their insurance plans, while the “observed-only-when-treated”($NO$) population encompasses people who would incur expenses only if enrolled in managed care plans. Assumption (ref) compares the (counterfactual) expenditures that would take place if all employees were enrolled in a managed care plan ($Y^{*}_{1}$) between those two groups. Formally, it says that for any particular level of expenditures, say 1000\$, the share of individuals that spend less than 1000\$ will be larger for the $NO$ group compared to the $OO$ type. Alternatively, it implies that the average expenditures among the lowest 25% (or any particular quantile) of spenders, will be smaller for the $NO$ than for the $OO$ subgroup. Intuitively, if everyone had a managed care plan, typical patients who would go to ambulatories only if insured by a managed care plan spend less in services than those that would go regardless of their plan. In Appendix (ref), we provide a simple theoretical framework in which this assumption holds.\footnote{As pointed out by a referee, this assumption is hard to interpret and motivate empirically. However, in a layered policy analysis Manski2011, we offer a menu of estimates based on different assumptions, that may or may not be plausible according to each researcher's own expertise and beliefs Tamer2010.} Importantly, Assumption (ref) has no impact on the information about $\alpha(p)$. Consequently, it only impacts the results by increasing the lower bound for the $MTE^{OO}$, which is much tighter as can be seen in Subfigure (ref), illustrating the identifying power of the stochastic dominance assumption. If we consider Assumption (ref) plausible, we find that the lower bound is positive for most values of the propensity score. This finding reinforces the results in Deb2006, who obtained a positive effect of PPO choice on ambulatory expenditures and a zero effect of HMO choice after controlling for selection. Furthermore, our results also suggest that the choice of managed care plan reduces ambulatory expenditures for individuals who face high latent costs.
This section extends Proposition (ref) to the case with multi-valued discrete instruments, sharply bounding many LATE parameters for the always-observed subpopulation. We focus on the case under the monotonicity restriction (Assumption (ref)-(ref)). Similar results hold under our other identifying sets of assumptions.
In many applications, the only instruments available are discrete, e.g., treatment eligibility, number of children in the household and quarter of birth. In this section, we provide sharp identification results when the instrument is multi-valued discrete, implying the support of the propensity score is finite. The results here can be seen as an extension of chen2015bounds, who provide an outer set for the LATE parameter when the instrument is binary.
Assumption (ref) requires that one can rank the probabilities of receiving treatment for the points of $P(Z)$ that are available, allowing the researcher to partition the $[0,1]$ interval into regions $[p_\ell - p_{\ell-1}]$ for $\ell=2,...,K$. The researcher can only identify an average of the MTE within each region, i.e., a LATE. Naturally, if the instrument has more points of positive mass (providing finer partitions of the probabilities), one could obtain averages of the MTE for more specific ranges of the unobservable characteristic $V$.\footnote{If for some values of $Z$, the probabilities are the same, $p_\ell=p_{\ell-1}$, we cannot refine the partition of the unit interval describing the probabilities and, hence, cannot improve on the detail level of the MTE identified.}
The identification argument is similar to the one presented for the continuous instrument case in Subsection (ref). Under Assumption (ref), we have $p_\ell=\mathbb P\left[V\leq P(z_\ell)\right]$ and $\mathbb P\left[P(z_{\ell-1}) < V\leq P(z_\ell)\right]=p_{\ell}-p_{\ell-1}$. If Assumptions (ref) and (ref) hold, then $P(z_\ell)=p_\ell$. To ease the exposition, we use the shorthand $P \coloneqq P(Z)$.
We have $\mathbb P\left[Y\in A,S=1,D=1|P=p_\ell\right]=\mathbb P\left[Y_1^* \in A, S_1=1,V \leq p_\ell\right]$. Therefore,
which implies that
Similarly, we have
Thus for $A=\mathcal Y$, we can write
We know that $\mathbb P\left[Y_d^* \in A|S_d=1,p_{\ell -1} < V \leq p_\ell\right]=\frac{\mathbb P\left[Y_d^* \in A, S_d=1|p_{\ell -1} < V \leq p_\ell\right]}{\mathbb P\left[S_d=1|p_{\ell -1} < V \leq p_\ell\right]}$ for $d\in \{0,1\}$. Under Assumption (ref), we identify $\mathbb P\left[S_{0}=1,S_{1}=1|p_{\ell -1} < V \leq p_\ell\right]=\mathbb P\left[S_0=1|p_{\ell -1} < V \leq p_\ell\right]$.
To implement the trimming in this setting, we define the discrete case analog of $\alpha(p)$, denoted by $\tilde{\alpha}(p_{\ell-1},p_\ell)$, $$\tilde{\alpha}(p_{\ell-1},p_{\ell}) \coloneqq \frac{\mathbb P\left[S_{0}=1,S_{1}=1|p_{\ell -1} < V \leq p_{\ell}\right]}{\mathbb P\left[S_{1}=1|p_{\ell -1} < V \leq p_{\ell}\right]}=\frac{\mathbb E[S(1-D)|P=p_{\ell -1}] - \mathbb E[S(1-D)|P=p_{\ell}]}{\mathbb E[SD|P=p_{\ell}] - \mathbb E[SD|P=p_{\ell-1}]}.$$ To find bounds around $\mathbb E[Y_{1}^{*}-Y_{0}^{*}|S_{0}=1,S_{1}=1,p_{\ell-1}<V\leq p_{\ell}]$, we follow the steps in Subsection (ref) to derive the following proposition:
As mentioned above, the quantity for which we derive bounds in Proposition (ref) is an average of the $MTE(p)$ evaluated at levels of $p$ in the interval $(p_{\ell-1}, p_{\ell}]$, i.e., we partially identify a LATE. More can be said about the MTE if additional assumptions are made. For example, if we assume that the $MTE$ is flat within each interval, then Proposition (ref) provides sharp bounds on the $MTE$ for the always-observed. This result is related to previous work in which discrete instruments are used to identify $MTE$ in the absence of sample selection. For example, Brinch2017 leverages additional functional structure for identification, while Mogstad2017 provided partial identification results for the $MTE$. An extension of their results to the current framework is an interesting topic for future research.
This paper derives sharp bounds for the marginal treatment effect for the always-observed individuals when there is sample selection. We achieve partial identification results under three increasingly restrictive sets of assumptions. First, we impose standard MTE assumptions without any restrictions to the sample selection mechanism. The second case, which is our main result, imposes monotonicity of the sample selection variable with respect to the treatment, considerably shrinking the identified set. Finally, we consider a strong stochastic dominance assumption which tightens the lower bound for the MTE.
All the results rely on the insight that observed individuals in the sample are a mixture of two possible groups, the ones that would always be observed regardless of treatment status and the ones that would self-select into the sample only when (un)treated. The mixture weights can be identified, leading to a trimming procedure that partially identifies the target parameter, extending Imai2008, lee2009training and chen2015bounds results to the context of MTE. Moreover, we derive testable implications of our identifying assumptions, and provide extensions to bound LATE parameters with multi-valued discrete instruments. A feasible estimator is proposed and implemented in an empirical illustration analyzing the impacts of managed health care options on health related expenditures, following Deb2006 and highlighting the practical relevance of the results.
We thank the Editor Elie Tamer, an Associate Editor, and two anonymous referees for constructive feedback that help improve the quality of the paper. We also thank Joseph Altonji, Nathan Barker, Michael Bates, Ivan Canay, Xiaohong Chen, Xuan Chen, Michael Darden, Nino Doghonadze, John Finlay, Carlos A. Flores, Thomas Fujiwara, Dalia Ghanem, John Eric Humphries, Yuichi Kitamura, Marianne Köhli, Helena Laneuville, Jaewon Lee, Giovanni Mellace, Ismael Mourifi\'e, Yusuke Narita, Pedro Sant'anna, Masayuki Sawada, Azeem Shaikh, Edward Vytlacil, Stephanie Weber, Siuyuat Wong and seminar participants at Iowa State University, University of Iowa, Yale University, UC Davis, UC-Riverside, UNC-Chapel Hill, CEME Conference for Young Econometricians 2019, IZA/CREST Conference on Labor Market Policy Evaluation, Southern Denmark University, Statistics Norway, CMStatistics 2019, the Bristol Econometrics Study Group 2019 and the 42\textsuperscript{nd} Meeting of the Brazilian Econometric Society for helpful discussions, and Seung Jin Cho for excellent research assistance.
\singlespace