EconBase
← Back to paper

Horowitz-Manski-Lee Bounds with Multilayered Sample Selection

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,563 characters · 16 sections · 136 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Horowitz-Manski-Lee bounds with multilayered sample selection

abstractThis paper investigates the causal effect of job training on wage rates in the presence of firm heterogeneity. When training affects the sorting of workers to firms, sample selection is no longer binary but is “multilayered". This paper extends the canonical Heckman (1979) sample selection model -- which assumes selection is binary -- to a setting where it is multilayered. In this setting Lee bounds set identifies a total effect that combines a weighted-average of the causal effect of job training on wage rates across firms with a weighted-average of the contrast in wages between different firms for a fixed level of training. Thus, Lee bounds set identifies a policy-relevant estimand only when firms pay homogeneous wages and/or when job training does not affect worker sorting across firms. We derive analytic expressions for sharp bounds for the causal effect of job training on wage rates at each firm that leverage information on firm-specific wages. We illustrate our partial identification approach with two empirical applications to job training experiments. Our estimates demonstrate that even when conventional Lee bounds are strictly positive, our within-firm bounds can be tight around 0, showing that the canonical Lee bounds may capture only a pure sorting effect of job training. \\ { Keywords: job training, sample selection, unordered treatments} JEL subject classification: C12, C14, C21, and C26.

Introduction

Governments allocate substantial funds to job training programs that are designed to improve worker skills. The United States (U.S.) federal government spends roughly $19$ billion U.S. dollars annually on employment and training programs.\footnote{See the Council of Economic Advisers 2019 report: “Government Employment and Training Programs: Assessing the Evidence on their Performance" (the_council_of_economic_advisers_government_2019 the_council_of_economic_advisers_government_2019).} Federal agencies administer roughly 40 employment and training programs to assist job seekers in gaining employment. Against this backdrop, there is an ongoing debate about the appropriate level of spending on such programs. Advocates argue that they help close the “skills gap” and address worker shortages, while critics argue that they are ineffective and socially wasteful.

At the center of the debate is the long-standing question of whether job training has a causal effect on the labor market outcomes of participants. This question has garnered significant interest from both academics and policymakers and knowing the answer is essential for determining whether to continue spending on training programs. Substantial progress has been made to answer this question as a result of randomized evaluations of job training programs. These evaluations have been the subject of several comprehensive meta-analyzes (see heckman_economics_1999 heckman_economics_1999 and card_active_2010 card_active_2010, card_what_2018).

To date, most training program evaluations have focused on total earnings as the main outcome of analysis. While the earnings impact is surely important for answering some questions, for others it is important to narrow the focus. Earnings naturally reflect both labor supply decisions (employment and hours margin) and wage rates. To better understand whether job training increases worker skills and welfare, standard economic models show that it is important to focus on the latter.\footnote{Labor supply decisions could also be impacted by job training via an increase in human capital. In particular, workers with higher skills are more likely to be offered -- and accept -- better paying jobs.}\textsuperscript{,}\footnote{hendren_unified_2020 measure the willingness to pay for job training using the treatment effect on total earnings. This assumes that all increases in earnings come from returns to human capital (higher wage rate), not from higher levels of labor supply.} However, identifying the causal effect of job training on wage rates is empirically challenging due to the well-known sample selection problem (heckman_sample_1979 heckman_sample_1979). This arises since a researcher only observes wages of the employed, and the likelihood of employment can itself be impacted by job training.

In a seminal contribution, lee_training_2009 showed that the causal effect of job training on wages for the always-employed can partially identify under the imbens_identification_1994 (“IA”) monotonicity assumption. Lee's key insight was to reduce the partial identification problem in this framework to the one considered by horowitz_identification_1995 – the problem of finding sharp bounds for the mean of an unobserved potential outcome that is a component of an observed mixing distribution with set-identified mixing probabilities. IA’s monotonicity condition delivers point identification of both the mixing weight (i.e., the share of always-employed) and the mean of the “untrained” potential outcome for the always-employed. Therefore, to obtain sharp bounds on the causal effect of job training on wages, one needs only to find sharp bounds on the mean of the “trained” potential outcome for the always-employed. Lee's bounding approach has become influential in empirical research.\footnote{In (ref), we describe results from an informal survey we conducted. We counted 56 papers published in `top 5' general interest economic journals -- the American Economic Review, Econometrica, the Journal of Political Economy, the Quarterly Journal of Economics and the Review of Economic Studies -- that cited lee_training_2009 or earlier working paper versions. The purpose of our survey was to determine whether the nature of sample selection in these papers was binary or multilayered. Of these, 42 empirically implemented Lee bounds to address sample selection, and 7 of them featured multilayered selection. To apply Lee bounds in these settings, the researchers collapsed the sample selection problem to a single dimension.}

In developing his approach, Lee focused on the potential for job training to affect labor supply along the extensive margin (work vs. no work). Although it is important to know whether training increases employment, a fundamental question for policymakers is whether job training improves labor market outcomes by increasing job quality. If this is the case, training may increase both the probability individuals are employed, as well as the probability individuals are employed at “better jobs”, both of which can increase earnings. For example, this is a key feature of President Biden's workforce training initiative, the “American Rescue Plan’s Good Jobs Challenge”, which prioritizes job quality and is designed to ensure that workers can access good jobs.

Despite the emphasis on job quality by policymakers, the academic literature on job training programs has mostly ignored firms. In the standard competitive model considered by Lee and most of the training literature, the firm for which an individual works does not matter for wages. However, there is growing empirical evidence that demonstrates the importance of firms for wage determination.\footnote{See, for example, abowd_high_1999, card_workplace_2013, song_firming_2019, bonhomme_distributional_2019 and bonhomme_how_2023.} This research highlights the importance of having a “good job” which can be interpreted as working at a “good firm” that offers a higher wage for all its employees. Thus, an open question is whether job training raises earnings by moving participants to higher-paying firms.

The presence of firm heterogeneity and the potential for worker sorting raises several new questions of interest. First, what estimand do “Lee bounds” partially identify when there is firm heterogeneity in wages? This question cannot be addressed with the canonical heckman_sample_1979 sample selection model since this assumes that sample selection is binary, i.e., job training can increase employment but has no effect on sorting to firms. The first contribution of this paper is to extend the standard sample selection model to a setting where selection is multilayered, and to show that the conventional Lee bounds set identifies (for the always-employed population) a total effect that combines a weighted average of the causal effect of job training on wage rates across firms (we label this the “within-firm effect”) with a weighted average of the contrast in wages between different firms for a fixed level of training (we label this the “sorting effect”).

A natural question is whether it is possible to separate the within-firm wage effect of job training from the sorting effect in the presence of heterogeneous firms. There are several reasons why one would want to separately identify these effects. First, some features of job training programs affect sorting (job search assistance), whereas others affect skill acquisition (classroom and vocational training). Thus, the decomposition could potentially highlight which investments – job search assistance or classroom training – are effective for raising wages and thus improve targeting. Second, the within-firm wage effect is arguably better able to shed light on the causal effect of employer-sponsored job training. Third, for a welfare analysis of job training programs, it is important to focus on the direct wage effects of job training, since labor supply effects have second-order effects on utility (via the envelope theorem) (see hendren_unified_2020 hendren_unified_2020).\footnote{This logic requires that the government is increasing spending on job training by a sufficiently small amount.} Thus, we argue that the within-firm wage effect is the more relevant causal effect for evaluating welfare effects.

The second contribution of this paper is to derive sharp bounds on the within-firm wage effect. Our bounding approach proceeds in two steps. In the first step, we derive sharp closed-form bounds on the response type probabilities.\footnote{The response type represents the pair of firms that an individual would choose to work at if she were externally assigned to the control group or the treatment group, respectively.} In deriving these bounds, we exploit a unique feature of our setting, which is that (unlike in the traditional instrumental variables framework), the exclusion restriction does not hold, since job training can have a direct causal effect on the outcome (wages). We show that this feature implies that the distribution of response types does not depend on the outcome (wage) distribution and allows us to derive closed-form bounds on the distribution of response types.\footnote{While we derive closed-form bounds on the distribution of response types, we show that one can obtain them equivalently using a linear programming approach.} The second step provides closed-form bounds on the treatment effects, as a function of the sharp bounds on the response types derived in the first step. This step involves extending the horowitz_identification_1995 approach (which involves a single-equation mixture model with two components) to our setting which involves two mixture model equations with unknown weights that are interdependent across the equations. Importantly, we show that while this two-step approach provides an easy and tractable way to construct closed-form bounds, it does not entail any loss of information and provides sharp bounds. We also consider a set of additional restrictions on response types and show that they naturally lead to tighter bounds on the treatment effects of interest. Using a numerical illustration, we demonstrate that our bounds can be informative enough to potentially separate cases with a positive within-firm effect from those with no within-firm effect, even when both cases lead to strictly positive Lee bounds.

We consider two empirical applications. Our first application is based on the randomized evaluation of Job Corps following lee_training_2009. We classify firms into observable firm types taking advantage of the fact that in the publicly available survey data there are direct measures of firm amenities, such as the availability of health insurance, paid vacation and retirement, or pension benefits. We show that, on average, firms that offer amenities pay higher wages than firms that do not. Moreover, the wage distribution for amenity-offering firms stochastically dominates the wage distribution for non-amenity-offering, within both the treatment and control groups. We then go on to show that being randomly assigned to Job Corps led individuals to work at firms with better job amenities compared to the control group. This combined evidence suggests that selection is multilayered and motivates our implementation of sharp bounds to evaluate Job Corps. We begin by replicating the findings of lee_training_2009. Our estimates reveal that even in the cases where conventional Lee bounds are strictly positive ($[0.047,0.048]$, 90 weeks after training), our multilayered bounds for the within-firm wage effect (which hold the sorting effect constant) include 0. This suggests that Lee bounds may capture a pure sorting effect of job training rather than a direct human capital effect. Indeed, under additional plausible assumptions on the distribution of response-types, our bounds on the within-firm wage effect become tight but continue to contain 0.

Our second empirical application is based on the WorkAdvance experiment recently examined in katz_why_2022. This is a sectoral employment program that targets high-quality jobs in specific industries with strong labor demand. We classify firms based on whether a firm operates in a WorkAdvance targeted sector. Similar to Job Corps, wages are higher at target sector firms and random assignment to job training increases sorting to these firms. We again document a case where Lee bounds are strictly positive and tight, but our within-firm wage effect includes 0.

The remainder of the paper is organized as follows. Section (ref) considers the multilayered sample selection problem in both a parametric model along the lines of heckman_sample_1979 and a more general treatment effects framework; it defines the key causal estimands of interest. Section (ref) considers the causal interpretation of Lee bounds in the presence of multilayered sample selection and presents the general decomposition. Section 4 derives the sharp bounds of a large class of parameters of interest in the multilayered sample selection model. Section (ref) presents empirical applications that implement the sharp bounds for Job Corps and WorkAdvance. The last Section concludes. The Appendices contain simulations of the model assuming that there are two types of firms, and presents all the proofs for the paper.

Related Literature

Our paper builds on and contributes to the following literature. First, there is a large literature on active labor market programs that is reviewed in heckman_economics_1999 and card_active_2010 (card_active_2010, card_what_2018). Our contribution to this literature is to examine whether, and to what extent, worker sorting to firms affects the wage impacts of job training. To our knowledge, there are only a few empirical analyses that have examined the impact of training on worker sorting to firms. andersson_does_2022 find suggestive evidence of a positive impact of training on firm characteristics, as well as effects on industry of employment. Another related study is katz_why_2022, who evaluate sector-based training programs. Examining evidence from randomized evaluations of programs that combine upfront screening, occupational and soft skills training, wrap-around services, and targeted low-wage workers, katz_why_2022 find substantial and persistent earnings gains after training. With regard to mechanisms, they interpret the earnings gain as driven in part by the sorting of workers to higher-paying industries and occupations. They do not, however, provide a framework for isolating the sorting effect as a causal mechanism. Finally, schochet_does_2008 evaluate the impact of Job Corps on the sorting of workers to jobs with different amenities, such as the availability of health insurance, and retirement or pension benefits, and report positive impacts. However, they do not disentangle the effects on these characteristics for those who would be employed in any case from a selection effect that comes from the impact of Job Corps on employment.

Second, our paper relates to the literature that has documented firm heterogeneity in wages. Firms have been shown to be important for wage inequality (abowd_high_1999 abowd_high_1999), the cyclicality of wages and early career progression (card_workplace_2013 card_workplace_2013), the earnings losses of displaced workers (lachowska_sources_2020 lachowska_sources_2020; schmieder_costs_2023 schmieder_costs_2023), and gender (card_bargaining_2016 card_bargaining_2016) and racial wage gaps (gerard_assortative_2021 gerard_assortative_2021). Our contribution to this literature is to examine the role of firms in understanding the wage effect of job training.

Third, our paper relates to econometric approaches that address the sample selection problem. The heckman_sample_1979 sample selection model has been extended in various dimensions. First, a series of papers, including gallant_semi-nonparametric_1987, newey_semiparametric_1990, and ahn_semiparametric_1993, propose estimation and inference methods that relax the normality assumption imposed by heckman_sample_1979; see li_nonparametric_2007 (li_nonparametric_2007, Chapter 10) for a review of such extensions. Second, lee_training_2009 extends heckman_sample_1979 by relaxing the exclusion restriction of instrumental variables and derives bounds on the parameters of interest. honore_selection_2020 study a semiparametric version of Lee's model. Additionally, semenova_generalized_2020 and olma_nonparametric_2021 propose various approaches for inference on Lee's bounds conditional on (potentially continuous) covariates. To our knowledge, this paper represents the first attempt to extend the seminal heckman_sample_1979 sample selection model to multilayered settings. Although the focus of our paper is primarily on firms, which we consider to be the main layer of interest, our analysis can be extended in various directions. For example, one could consider occupation as a layer and examine the returns to occupation while controlling for sorting, similar to the approach taken by gottschalk_taking_2014.

Finally, one can view the firm as a “mediator” in the context of the literature on mediation analysis (see, for example, robins_identifiability_1992 and pearl_direct_2001). Traditionally, most of this literature ignores sample selection where the outcome is not observed at some mediator values. For example, recently kwon2024testing developed a method to test for the presence of a mediator but abstracts from sample selection. A rare exception is zuo_mediation_2022, who consider identification of direct and indirect effects within a mediation analysis framework, when both the outcome and the mediator are missing. They focus on point identification under various assumptions, including the abstract and non-falsifiable assumption of completeness; however, many of their assumptions do not apply to our setting.\footnote{For an in-depth review of completeness, see dhaultfoeuille_completeness_2011, and also canay_testability_2013.} For instance, in our setting, whether the wage is observed can depend on wage rate offered by firms (the mediator), and this situation is ruled out by their assumptions. Our paper therefore complements zuo_mediation_2022 by establishing partial identification of the direct and indirect effects without imposing completeness. Our maintained assumptions are transparent and directly apply to the primitives of our model. Moreover, our approach accommodates an endogenous mediator and allows the outcome to be missing non-randomly, even conditional on covariates.

Analytical Framework

Multilayered Sample Selection: A parametric model

Since Heckman's seminal work in 1979, the sample selection model has been formalized as follows:

eqnarray[eqnarray omitted — 239 chars of source]

where $D$ captures the binary sample selection model (i.e., $D$ is equal to $1$ if employed and $0$ if not) and $Y$ represents the outcome which is observed only when $D$ is equal to $1$ (i.e., $Y$ is the observed wage when employed). The latent variables in the model are denoted as $(U, V)$, $X$ is a vector of observed exogenous covariates, and $Z \in \{0,1\}$ is a binary variable that respects the conditional independence assumption: $(U,V) \perp Z|X$ (i.e., job training, $Z$, is randomly assigned). In the Heckman sample selection model, identification of the parameters in the outcome equation requires at least one variable that is independent of the latent variables but is excluded from the outcome equation. In model ((ref), (ref)) this is equivalent to assuming $\alpha$ to be equal to $0$, implying that $Z$ is excluded from the outcome equation. In this context, $Z$ becomes a valid instrumental variable to consistently estimate $\beta$, satisfying both the independence and exclusion restriction conditions.

However, as highlighted in lee_training_2009 and recognized more generally, the exclusion restriction can be violated in certain cases implying that $\alpha \neq 0$. Moreover, in such cases, $\alpha$ is potentially a parameter of primary interest. For instance, in the job training example, participating in training could boost an individual's human capital and directly affect their wage rate. Consequently, $\alpha$ can be interpreted as the causal effect of job training on the wage rate, which is the key parameter studied by lee_training_2009. The primary methodological contribution of Lee is to provide a method that allows researchers to partially identify $\alpha$ in the presence of sample selection (i.e., when training can affect labor supply through $\eta$).

A key assumption in the parametric model above is that the sample selection problem is binary: individuals are either employed or unemployed. Yet if job training affects not only whether an individual works but also which firm they work at, the sample selection problem becomes multilayered. We now generalize the seminal Heckman sample selection model ((ref), (ref)) to allow for a richer model of labor supply where individuals choose layers, i.e., firms. We refer to this extended model as the “parametric multilayered selection model":

eqnarray[eqnarray omitted — 474 chars of source]

where $\eta_{0} Z + \theta_0 X^{'}+V_{0}=0$. In this model, each layer ($D$) represents a distinct firm, with corresponding parameters $\alpha_d$, $\beta_d$, and latent variable $U_d$. Expected utility for a given firm $d$ is given by $\eta_{d} Z + \theta_d X' + V_{d}$. The utility of the outside option (i.e., unemployment) is $\eta_{0} Z + \theta_0 X' + V_{0}=0$. The worker selects the firm with the highest expected utility.

In the parametric multilayered sample selection model, $\alpha_d$ is the causal effect of job training on the wage rate at firm $d$. We refer to this causal effect as the “within-firm effect” for layer $d$. The vector $(\eta_1,...,\eta_K)$ includes parameters that reflect the causal effect of job training on firm choice. In the next section, we demonstrate that Lee bounds combine both the within-firm effects $(\alpha_1,...,\alpha_K)$ with the sorting effects summarized by $(\eta_1,...,\eta_K)$. As discussed in the Introduction, it is interesting to separately identify these causal effects. For example, it is straightforward to show that the change in utility associated with a small increase in spending on job training is captured by the mechanical increase in wages. The sorting effect has only a second-order effect on worker utility due to the envelope theorem.

Multilayered Sample Selection: Generalized version using the potential outcome model

Let $(\Omega,\mathcal F, P)$ be a probability space, where we interpret $\Omega$ as the population of interest, and $\omega\in \Omega$ as a generic individual in the population. Let $Y_{z,d}(\omega)$ be the potential outcome (i.e., potential wage) if agent $\omega$ is externally assigned to the treatment group $z \in \{0,1\}$ (i.e., job training) and to a specific layer $d \in \{0,1,...,K\}$, where $d=0$ denotes the layer for which the outcome is not observed (i.e., $Y_{z,0}(\omega)$ is not observed).\footnote{In our empirical application, we will assume that the layer corresponds to a firm's type, where the type is constructed based on a firm's observable characteristics.} $D_z$ denotes the potential layer that the individual selects if externally assigned to the treatment group $z \in \{0,1\}$. We denote the realized outcome by $Y \in \mathcal Y \subseteq \mathbb R$ and the realized layer by $D \in \{0,1,...,K\}$. Finally, let $Z$ be the assigned treatment group and let $X$ be a vector of covariates. We assume that $(D,Z,X)$ is observed for every individual, but the realized outcome $Y$ is only observed if $D\neq 0$ (i.e., if the individual is employed). This implies the following model:\footnote{Strictly speaking, equation ((ref)) implies that the outcome $Y_{z,d}(\omega)=0$ when $D=0$, but it should really be interpreted as $Y_{z,d}(\omega)$ is unobserved.}

eqnarray[eqnarray omitted — 126 chars of source]

along with the following conditional independence assumption.

assumption[Conditional Random Assignment] Individuals are randomly assigned to a treatment group. $\left\{(Y_{z,d}, D_z): d \in \{0,1,...,K\}, z \in \{0,1\} \right\} \perp Z |X.$

The outcome equation ((ref)) collapses to the outcome model of lee_training_2009 when there is no heterogeneity between the potential layered outcomes, that is, $Y_{z,d}=Y_{z}$ for $d \in \{1,...,K\}$. In this case, we have:

multline*[multline* omitted — 174 chars of source]

What causal interpretation should be given to Lee bounds in the presence of multilayer sample selection, where $Y_{z,d}\neq Y_{z,d'}$, for $d,d' \in \{1,...,K\}$? Depending on the researcher's interest, various causal estimands of interests could be defined. Before defining our parameters of interest, we show that there is a link between the causal effects in our multilayered framework and those typically considered in the mediation analysis literature. We use this link to characterize our key estimands below.

Direct and indirect effects in presence of sample selection

In our model, particularly in equation ((ref)), there is a notable connection to the literature on mediation analysis, as discussed by pearl_direct_2001 and others. The graphical representation of the outcome equation in our model takes the form:\footnote{For the sake of clarity, this graph simplifies the discussion by omitting sample selection.}

figure[figure omitted — 1,643 chars of source]

In the context of mediation analysis, where $Z$ represents the randomized treatment, $D$ is conceptualized as the “mediator", and $Y$ denotes the outcome, our model allows the treatment (job training) to influence the outcome through two channels: a direct channel and an indirect channel that passes through the mediator. $W$ is a vector of latent unobserved variables often called “confounding variables" which simultaneously affect $D$ and $Y$, making $D$ an endogenous variable. In the parametric model, $W\equiv (U_1,...,U_K,V_1,...,V_K)$. In our framework, the mediator corresponds to the firm where the individual would be employed if they were externally assigned to job training.

In the mediation analysis literature, two categories of causal estimands have garnered attention: “direct effects" and “indirect effects". Focusing on the former, two types of “direct effects" have been conceptualized. First, the control direct effect (CDE) is defined as:

eqnarray[eqnarray omitted — 74 chars of source]

This captures the causal effect of job training on earnings within a specific firm $d$ when the firm is held fixed. It is equivalent to the within-firm wage effect for layer $d$. In lee_training_2009's (lee_training_2009) terminology, the CDE corresponds to the causal impact of job training on the wage rate illustrated by the curved arrow in Figure (ref). This parameter is the primary focus of lee_training_2009.

The CDE can vary significantly across firms ($d$), reflecting the potentially heterogeneous impact of job training on wages for different firms. CDEs are useful when the policymaker is primarily interested in the impact of job training on wages at a specific firm. More generally, policymakers may also be interested in understanding the overall impact of training on wages at firms that workers naturally choose when they receive training. The second type of direct effect -- the “natural direct effect" (NDE) -- is introduced to capture this notion:

eqnarray[eqnarray omitted — 140 chars of source]

where $Y_{z,D_{z'}} \equiv\sum_{d=0}^{K} Y_{z,d}1\{D_{z'}=d\}$ for $z,z' \in \{0,1\}$, and $d \in \{0,...,K\}$. The expression $Y_{1,D_1}(\omega)-Y_{0,D_1}(\omega)$ represents the causal impact of job training on wages for the specific firm that the worker $\omega$ would have selected if she had been externally assigned to receive job training. The NDE is essentially the average of these individual effects.

Turning to indirect effects, the “natural indirect effect" (NIE) is defined as:

eqnarray[eqnarray omitted — 202 chars of source]

The term $Y_{0,d}-Y_{0,d'}$ represents the wage contrast between firms $d$ and $d'$ in the absence of job training. However, rather than specifying the pair of firms $(d, d')$, we can examine this wage difference at the “natural representative" firms $D_1$ and $D_0$, resulting in $Y_{0,D_1}-Y_{0,D_0}$. The indirect effect aims to capture the causal impact of job training on the outcome purely due to the change in firms, and can therefore be seen as the influence of job training transitioning through a change in the firm. As highlighted by pearl_causal_2009, the empirical relevance of the indirect effect estimand is controversial and questionable. Implementing an intervention that would suppress the direct effect of $Z$ on $Y$ while allowing the indirect channel through $D$ is not realistic. However, it remains a key parameter in the mediation analysis literature.

In the context of sample selection, the outcome is observable only when $D \neq 0$. We can categorize the population into four major groups: $\Omega=\{\omega: D_0(\omega)=0, D_1(\omega)=0\} \cup \{\omega: D_0(\omega)>0, D_1(\omega)=0\} \cup \{\omega: D_0(\omega)=0, D_1(\omega)>0\} \cup \{\omega: D_0(\omega)>0, D_1(\omega)>0\}$. For the first group we never observe outcomes, regardless of training status. For the second group only outcomes under job training are never observed, whereas for the third group only outcomes when not assigned to job training are never observed. If we are unwilling to assume that the outcome is missing at random (or selection on observables only) or impose parametric assumptions, the observed data cannot provide information on the causal effect of job training for individuals belonging to those groups. We refrain from imposing such stringent restrictions and focus solely on the causal effects for the final group, the subpopulation $\{\omega: D_0(\omega)>0, D_1(\omega)>0\}$.

It is useful to further divide our population of interest $\{\omega: D_0(\omega)>0, D_1(\omega)>0\}$ into finer groups, which we label response types, i.e. $\{\omega: D_0(\omega)>0, D_1(\omega)>0\}=\cup_{\{d,d'\in \{1,..., K\}\}}\{\omega \in \Omega: D_1(\omega)=d, D_0(\omega)=d'\}$.\footnote{See heckman_unordered_2018 for a more detailed discussion on the advantages of such a partition.} Response types are defined by the pair of firms that individual $\omega$ would choose to work for if externally assigned to the control group or the treatment group. Formally, the response type is defined as the random variable $T=(D_0, D_1)$ and $\mathcal T$ represents its support.

We now introduce two pivotal parameters, the Local Controlled Direct Effect (LCDE) and the Local Controlled Indirect Effect (LCIE):

eqnarray[eqnarray omitted — 129 chars of source]

and

eqnarray[eqnarray omitted — 135 chars of source]

The individual CDEs may vary between individuals, i.e. $Y_{0,d}(\omega)-Y_{0, d'}(\omega) \neq Y_{0,d}(\omega')-Y_{0, d'}(\omega')$ for $\omega \neq \omega'$. By considering the LCDE, we allow for heterogeneity of the CDE across response types. Note that in certain instances, a specific LCDE may be more policy-relevant than the CDE itself. Both the LCDE and CDE exhibit their own policy relevance, akin to the extensive debate in the instrumental variable (IV) literature regarding the empirical relevance between the average treatment effect (ATE) and local ATE (LATE). This analogy extends to the LCIE.

We now show that the “sample selection" versions of the CDE, NDE, and NIE can be perceived as a weighted average of $\text{LCDE}(d|t)$ or $\text{LCIE}(z,d, d'|t)$. By “sample selection” version, we mean that the effects are defined conditionally on being always employed, i.e. $\{\omega: D_0(\omega)>0, D_1(\omega)>0\}$.

multline*[multline* omitted — 468 chars of source]

As illustrated, $\text{LCDE}(d|t)$ or $\text{LCIE}(z,d, d'|t)$ represent more primitive parameters compared to CDE, NDE, and NIE. This paper will focus in particular on identifying the $\text{LCDE}(d|t)$. It should be noted that, in the absence of individual heterogeneity, whenever $Y_{z,d}(\omega)=Y_{z,d}(\omega')$ for $\omega \neq \omega'$, we have $\text{LCDE}(d|t)=\text{LCDE}(d)=\alpha_d$, as in the parametric version of the model. Moreover, since the outcome is never observed when $D=0$, hereafter we use the following notation $Y_{z,D_{z'}} \equiv\sum_{d=1}^{K} Y_{z,d}1\{D_{z'}=d\}$ for $z,z' \in \{0,1\}$.

remarkIn lee_training_2009, training assignment $Z$ is a randomly assigned treatment that fails to satisfy the exclusion restriction required to address sample selection using standard methods. In our more general setting, while $Z$ remains a treatment of interest, it also plays the role of being an instrument for $D$. In some settings, the impact of $D$ on the outcome may be of independent interest, and our results also apply to such settings. Moreover, since our results apply even when there is no sample selection (i.e. $Y$ is always observed, $P(D = 0) = 0$), our approach generalizes the IV model to settings where the instrument does not satisfy the exclusion restriction.

The Causal Interpretation of Lee's bounds in the presence of Multilayered Sample Selection

First, notice that the generalized multilayered selection model, i.e., equations ((ref), (ref)) implies the following:

eqnarray[eqnarray omitted — 129 chars of source]

Additionally, Lee imposes the following monotonicity assumption:

assumption[Conditional Lee's Monotonicity Assumption] We impose the following restriction: $\mathbb P\left[1\{D_1>0\}\geq 1\{D_0>0\}\right|X]=1$ a.s.

This assumption means that being assigned to the treatment group can never lower employment, and this applies uniformly for all agents in the population. In our general framework with multilayered sample selection, this assumption requires that all agents are more likely to join an employment layer when assigned to treatment. Since Lee's monotonicity assumption is only required to hold conditional on $X$, it can be modified to allow the direction of monotonicity to vary across different values of $X$. Such a modification does not present a challenge for our identification analysis, which holds $X$ fixed throughout, but inference methods need to be adapted to accommodate such an assumption, especially when $X$ is continuous. For further details on such adaptations, see sloczynski_when_2020 and semenova_generalized_2020, which provide inference methods that are valid under such assumptions.

All remaining analysis, results, and assumptions should be understood as implicitly conditioning on $X = x$ for some value $x$ of the vector of observed covariates, $X$, which will generally be suppressed in notation.

lemma[Lee Bounds] Under Assumptions (ref) and (ref) Lee bounds set identifies the following estimand $\mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0]$: \begin{eqnarray} \theta^{\ell} \leq \mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0] \leq \overline{\theta}^{\ell} \end{eqnarray} where \begin{enumerate} • For continuous outcome: \begin{eqnarray} &&\theta^{\ell}\equiv \mathbb E[Y|D>0,Z=1, Y\leq F^{-1}_{Y|D>0,Z=1}(p)]-\mathbb E[Y|D>0,Z=0],\\ &&\overline{\theta}^{\ell} \equiv \mathbb E[Y|D>0,Z=1, Y\geq F^{-1}_{Y|D>0,Z=1}(1-p)]-\mathbb E[Y|D>0,Z=0], \end{eqnarray} • For binary outcome: \begin{eqnarray} &&\theta^{\ell}\equiv \max\left\{0, 1-\frac{1}{p}P[Y=0|D>0,Z=1]\right\}-\mathbb E[Y|D>0,Z=0] ,\\ &&\overline{\theta}^{\ell} \equiv \min\left\{1, \frac{1}{p}\mathbb P[Y=1|D>0,Z=1]\right\}-\mathbb E[Y|D>0,Z=0], \end{eqnarray} \end{enumerate} with $F_W^{-1}(u)\equiv \inf\{w \in \mathbb R: \mathbb P(W\leq w) \geq u\}$ for $u \in [0,1]$ and $p\equiv\frac{\mathbb P(D>0|Z=0)}{\mathbb P(D>0|Z=1)}$.

Lemma (ref) shows that in the presence of heterogeneous firms, Lee's identification approach bounds the estimand $\mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0]$. What is the causal interpretation of this estimand? The following lemma sheds light on this.

lemma[Decomposition] Assuming the generalized multilayered sample selection model, we have the following decomposition: \begin{itemize} • General decomposition: \begin{multline} \mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0]\\=\underbrace{ \sum_{d=1}^{K}\sum_{d'=1}^{K}LCDE(d|d',d)\times \mathbb P[T=(d',d)|D_0>0,D_1>0]}_{\mathbb E[Y_{1,D_1}-Y_{0,D_1}|D_0>0,D_1>0]} \\ + \underbrace{\sum_{d=1}^{K}\sum_{d'=1:d\neq d'}^{K}LCIE(0,d,d'|d',d)\times \mathbb P[T=(d',d)|D_0>0,D_1>0]}_{\mathbb E[Y_{0,D_1}-Y_{0,D_0}|D_0>0,D_1>0]} \end{multline} • No mediation effect (No firm-specific wage rate, i.e., $Y_{z,d}=Y_{z\bullet}$) or no sorting across firms, i.e., $\mathbb P[T=(d',d)|D_0>0,D_1>0]=0$ for $d \neq d'$. \begin{multline} \mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0]\\=\sum_{d=1}^{K}\sum_{d'=1}^{K}\mathbb E[Y_{1\bullet}-Y_{0\bullet}|D_1=d,D_0=d']\mathbb P(D_1=d,D_0=d'|D_0>0,D_1>0)\\=\mathbb E[Y_{1\bullet}-Y_{0\bullet}|D_0>0,D_1>0] \end{multline} • No direct effect (i.e., $Y_{z,d}=Y_{\bullet d}$). \begin{multline} \mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0]\\=\sum_{d=1}^{K}\sum_{d'=1:d\neq d'}^{K}\mathbb E[Y_{\bullet d}-Y_{\bullet d'}| D_0=d',D_1=d]\times \mathbb P[D_0=d',D_1=d|D_0>0,D_1>0]\\=\mathbb E[Y_{\bullet D_1}-Y_{\bullet D_0}|D_0>0,D_1>0] \end{multline} \end{itemize}

Lemma (ref) (i) shows that in the presence of firm heterogeneity, Lee's partial identification approach establishes bounds for a total effect. This total effect combines the sample selection version of the NDE and the NIE (i.e., conditional on $D_0>0$ and $D_1>0$), with each possessing distinct interpretations. Importantly, as discussed above, the NDE and NIE do not hold the mediator (firm) $D$ fixed. The NDE is an average of causal effects of job training at each firm weighted by fraction of workers choosing that firm under job training. The NIE is an average of causal effects of firm on wages (in the no-job training counterfactual scenario) weighted by the share of response types choosing those firms. Thus, without additional assumptions, this approach does not allow one to separately identify the CDEs (the within-firm wage effects) from the labor supply effects or sorting effects that transit through $D$. Lemma (ref) (ii) shows that when there are no mediation effects (or no heterogeneity in wages between firms), i.e., $Y_{z,d}=Y_{z\bullet}$ as assumed in lee_training_2009 and illustrated in Figure (ref),

figure[figure omitted — 1,418 chars of source]

the NIE vanishes while the NDE reduces to the CDE which is the target parameter in Lee's framework.\footnote{Notice that Figures (ref) and (ref) are drawn for the subpopulation of always observed, i.e. $\{\omega: D_0(\omega)>0, D_1(\omega)>0\}$.} Finally, Lemma (ref) (iii) reveals that in the absence of a direct effect of job training on wages (i.e., $Y_{z,d}=Y_{\bullet d}$) as depicted in Figure (ref), Lee bounds capture the effect of job training on wages coming exclusively from the sorting of individuals into different firms.

figure[figure omitted — 1,365 chars of source]

This shows that, in general, interpreting Lee bounds as informative about the direct wage effect of job training is problematic, unless there is clear empirical evidence of the absence of mediation effects. Unfortunately, Lee's approach does not provide a means to assess this. Given these challenges, the next section introduces an alternative partial identification approach that is designed to overcome these limitations and aims to partially identify the true causal impact of job training on wage rates.

Sharp bounds in the multilayered sample selection model

In this section, we develop a partial identification strategy to recover the parameters $\text{LCDE}(d|t)$ and $\text{LCIE}(z,d, d'|t)$ that will allow us to isolate the within-firm effect of job training from the sorting effect. Under Assumption (ref), the response type $T$ is independent of $Z$. Assumption (ref) restricts the response type support. For example, under Assumption (ref), $\mathbb P[T=(d,0)]=0$ for $d \in \{1,...,K\}$. We denote by $f_{Y_{z,d}|D,Z}(y|d',z')$ the conditional density of $Y_{z,d}$ given $\{D=d',Z=z'\}$ and assume that it is absolutely continuous with respect to a dominating measure $\mu$ on $Y_{z,d}$. We note that $f_{Y_{z,d},D|Z}(y,d|z)\equiv f_{Y_{z,d}|D,Z}(y|d,z)\mathbb P(D=d|Z=z)$. For $d,d' \in \{1,...,K\}$ and $z \in \{0,1\}$, and any $y \in \mathcal Y$ we have the following:

multline[multline omitted — 176 chars of source]

where the first equality holds under Assumption (ref).

More precisely, under Assumption (ref), the following system of equations characterizes the empirical content of the multilayered sample selection model.

eqnarray[eqnarray omitted — 205 chars of source]

and this holds for any $d,d' \in \{1,...,K\}$ and $y \in \mathcal Y$. The left-hand side of equations ((ref)) and ((ref)) are observed while the individual types, i.e., $\mathbb P[T=(d,d')]$ and the conditional potential outcome distributions, i.e., $f_{Y_{z,d}|T}(y|d',d)$ on the right-hand side of the equations are unknown. For a given $d$, the number of unknown quantities $2K+1 + 2(K+1)|\mathcal Y|$ is larger than the number of equations $2 |\mathcal Y|$. We therefore have an under-determined system of linear equations with unknown coefficients. As such, it is only possible to set identify these parameters. The identified set of unknown parameters could naturally shrink if the researcher is willing to impose additional assumptions, such as Assumption (ref). For example, under Assumption (ref), $\mathbb P[T=(d,0)]=0$ for $d \in \{1,...,K\}$ which implies that $\mathbb P[T=(d,0)] f_{Y_{z,d}|T}(y|d,0)=0$ for $d \in \{1,...,K\}$. Consequently, for a fixed $d$, this leads to a reduction of $|\mathcal Y|+1$ in the total number of unknown parameters while keeping the number of equations fixed. As a result, the system of equations becomes more tightly constrained. When the support of $Y$, i.e., $\mathcal Y$, is finite, the system of equations ((ref))-((ref)) could be solved using a linear programming method, with the drawback that the linear programming approach does not provide intuition about the source of identification power.\footnote{ If the researcher is interested in analyzing a discrete outcome and wishes to explore this avenue further, she could employ the inferential method developed by fang_inference_2023.} More importantly, the linear programming approach can no longer be used when $Y$ is continuous, as is the case in our empirical application. To address this issue, we develop a two-step identification approach. The first step provides sharp bounds on the response types. This step involves only the distribution on $(D, Z)$ that has finite support in our framework and can therefore be solved using a linear programming approach, since it does not involve $Y$ which could have continuous support. The second step provides closed-form bounds on the treatment effects of interest, as functions of the sharp bounds on the response types computed in the first step. We show that these two steps provide sharp bounds on our parameters of interest.

Step 1: Sharp bounds on response type probabilities

In this step, we focus on the partial identification of the distribution of response types. Integrating equations ((ref)) and ((ref)) over $\mathcal Y$, we obtain the following system of equations, for each $d$:

eqnarray[eqnarray omitted — 155 chars of source]

In general, in the standard IV model, the distribution of response types depends on the full joint distribution of the observed data $(Y, D, Z)$, not just on the distribution of $(D,Z)$.\footnote{This has been pointed out by huber_sharp_2017 and is also implicit in the results of kitagawa_identification_2021. See Theorem 3 in vayalinkal_sharp_2024 for a result characterizing the relationship between outcome distributions and the identified set of response-type probabilities.} This complexity occurs because in the IV framework, the exclusion restriction is imposed, i.e., $Y_{z,d}=Y_{\bullet d}$. In fact, when this restriction is imposed, the response-type conditional density of $Y_{\bullet d}$ appears in both equations ((ref)) and ((ref)), and integrating each equation separately can lead to a loss of information on the response-type probabilities (leading to non-sharp bounds). However, in the absence of the exclusion restriction, each response-type conditional density $f_{Y_{z,d}|T}$ in the system of equations ((ref)) and ((ref)) only appears in one equation, so the integration step can be performed without losing any information on response-type probabilities. Therefore, we show that in our model, the response-type probabilities are entirely characterized by the distribution of $(D,Z)$ which justifies proceeding in two steps. We will say that a vector $\mathbf{v}$ satisfies ((ref), (ref)) if $\left(\mathbb P[T=(d,d')] : d,d' \in \{0,...,K\}\right) = \mathbf{v}$ is a solution to ((ref), (ref)) for all $d$.

lemmaConsider the model ((ref), (ref)). Under Assumption (ref), the (sharp) identified set of response-type probabilities is the set of non-negative vectors that satisfy ((ref), (ref)).

The researcher may also seek to apply additional restrictions on the distribution of response types, including, but not limited to, Assumption (ref). For example, one could assume that there are more upward switchers than downward switchers or more stayers than downward switchers. We consider such restrictions as a possible auxiliary assumption. We designate the set of linear constraints that can be applied to the response types as $\mathcal R_T$.

assumption[Restriction on response types] Consider that the layers (i.e., firms) are ordered. \begin{enumerate} • [Strong Monotonicity] $1\{D_1(\omega)=d\}\geq 1\{D_0(\omega)=d'\}$ for $d\geq d'$, or equivalently $\mathbb P[T=(d',d)]=0$ for $d\geq d'$. • [More upward switchers than downward switchers] $\mathbb P[T=(d,d')] \geq \mathbb P[T=(d',d)]$ for $d\geq d'$. • [More stayers than downward switchers] $\mathbb P[T=(d,d)] \geq \mathbb P[T=(d',d)]$ for $d\geq d'$. \end{enumerate}

A notable aspect of Assumptions (ref) and (ref) is that these restrictions can seamlessly integrate into equations ((ref))-((ref)) as supplementary linear constraints. Consequently, the process of recovering response types that conform to all these behavioral restrictions simplifies to a feasible linear programming problem. For example, researchers may choose $\mathcal R_T = \{\text{Assumption } \ref{Ass:LM}\}$, $\mathcal R_T = \{\text{Assumption } \ref{Ass:LM}$, $ \text{Assumption } \ref{Ass:Type-rest}(i)\}$, or $\mathcal R_T = \{\text{Assumption } \ref{Ass:LM}, \text{Assumption } \ref{Ass:Type-rest}\}$.

As mentioned in Lemma (ref), equations ((ref), (ref)) sharply characterize the restrictions on the distribution of $T$ imposed by the model ((ref), (ref)). Therefore, the identified set for response type probabilities under model ((ref), (ref)), Assumptions (ref), and responses type restrictions $\mathcal R_T$, denoted $\Theta_{I}(\mathcal R_T)$, is simply the set of non-negative vectors that jointly satisfy both $\mathcal R_T$ and ((ref), (ref)), i.e., $$\Theta_{I}(\mathcal R_T)\equiv \left\{ \mathbf{v} \in \mathbb{R}_{\geq 0}^{(K+1)^2}:

array[array omitted — 165 chars of source]

\right\}.$$

We also have the following result:

lemmaConsider the model ((ref), (ref)). Assumption (ref) and $\mathcal R_T$ are jointly rejected by the data if and only if $\Theta_{I}(\mathcal R_T)=\emptyset$.

Lemma (ref) has significant practical implications. It shows that assessing the validity of model assumptions does not depend on knowledge of the outcome distribution, and therefore, can be reduced to testing the feasibility of a linear program. This property greatly simplifies the implementation of a falsification test for our model.\footnote{For example, implementing a (sharp) test of our model is much simpler than the sharp tests for instrument validity and monotonicity proposed by kitagawaTestInstrumentValidity2015 and mourifieTestingLocalAverage2017.} In other words, once we can find a distribution of type that aligns with the assumptions of the model and the observed data on $(D,Z)$, it is always possible to find a corresponding distribution of potential outcomes $f_{Y_{z,d}|T}(y|d',d)$ that would rationalize the observed joint distribution of $(Y,D,Z)$.

For simplicity, we introduce the shorthand notation, $p_{d,d'}\equiv \mathbb P[T=(d,d')]$, and $\gamma^z_{d,d'}\equiv\frac{p_{d,d'}}{\mathbb P(D=d|Z=z)}.$ When $\Theta_{I}(\mathcal R_T)\neq \emptyset$, let $\underline{p}^{ r}_{d,d'}$ denote the infimum over all values for $p_{d,d'}$ induced by distributions that belong to $\Theta_{I}(\mathcal R_T)$. Subsequently, we can define: $\underline{\gamma}^{z,r}_{d,d'}=\frac{\underline{p}^r_{d,d'}}{\mathbb P(D=d|Z=z)}.$ The “$r$" superscript is used to emphasize the dependence of $\underline{p}^{ r}_{d,d'}$ and $\underline{\gamma}^{z,r}_{d,d'}$ on $\mathcal R_T$, the set of assumptions imposed.

For all the potential choices of $\mathcal R_T$ considered above, $\Theta_{I}(\mathcal R_T)$ is the set of nonnegative solutions to a linear system. Therefore, $\underline{\gamma}^{z,r}_{d,d'}$ for $d,d' \in \{0,...,K\}$ and $z \in \{0,1\}$ can be obtained as a solution to a linear program. Since the linear system of interest here is generally small, it is also possible to obtain an analytic solution for $\underline{p}^r_{d,d'}$ (and therefore for $\underline{\gamma}^{z,r}_{d,d'}$) via Fourier-Motzkin elimination. The details of both the computational and analytical approaches are presented in Appendix (ref).

Step 2: Sharp bounds on the treatment effects

As evident from equations ((ref)) and ((ref)), the observed conditional distribution of earnings, $F_{Y|D,Z}(y|d,z)$, can be expressed as a finite mixture of the conditional potential outcome distributions given the response types, $F_{Y_{z,d}|T}(y|l,l')$. More precisely, we have:

eqnarray[eqnarray omitted — 193 chars of source]

The unknowns in this mixture are the weights, $\gamma^z_{d,d'}$ for $z \in \{0,1\}$, and $d,d' \in \{0,...,T\}$. In lee_training_2009, Assumption (ref) implies that the weights are point identified, and establishing identification reduces to recovering the mean average of the mixture components. However, in our scenario, the mixture weights are not point identified under Assumption (ref). Even when we strengthen Assumption (ref) with Assumption (ref), point identification is still not achieved. This underidentification issue arises primarily due to the presence of numerous unobserved types stemming from the multiple layers, i.e., firms. Nevertheless, we can still derive informative bounds on these weights, as elaborated in the previous subsection.

horowitz_identification_1995 proposed sharp bounds on the distribution of mixture components in a single-equation mixture model with two components, where the weights are unknown but researchers possess non-trivial bounds for these weights, and cross_regressions_2002 extended these results to single-equation models with many components. However, the empirical content of our model is characterized by a set of systems of mixture equations, one system for each $d \in \left\{1, \dots, K\right\}$, each with up to $(K+1)^2$ components. Importantly, in our setting, the weights are unknown and shared between these systems, introducing a cross-equation dependence that is not present in horowitz_identification_1995 or cross_regressions_2002. We derive the identified set for the weights and then extend the approaches of horowitz_identification_1995 and cross_regressions_2002 to this more general case, deriving closed-form bounds on our key parameters of interest.

Moreover, Lee demonstrated that for continuous outcomes, the bounds proposed by horowitz_identification_1995 can be equivalently expressed as the mean of a truncated distribution. We extend Lee’s results by introducing a generalized truncated mean representation that applies regardless of the outcome distribution, whether it is continuous, discrete or mixed.

Hereafter, to simplify our notation and enhance readability, we introduce the following notation. For any $d$ and $z$, define: $\underline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\gamma;d,z)\equiv \mathbb E[F^{-1}_{Y|D=d,Z=z}(U) | U \leq \gamma]$, and $\overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\gamma;d,z)\equiv \mathbb E[F^{-1}_{Y|D=d,Z=z}(U) | U \geq 1-\gamma]$, for some $U\sim \text{Uniform}(0,1)$. We define $y_L$ as the lower bound of the support of $Y$ and $y_U$ as the upper bound.\footnote{Note that these bounds need not be finite.}

Before stating the main result, we note the following. When the outcome is continuously distributed $\underline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\gamma;d,z)$ is exactly equal to the truncated mean used in lee_training_2009 i.e., $$\mathbb E[F^{-1}_{Y|D=d,Z=z}(U) | U \leq \gamma]=\mathbb E[Y|D=d,Z=z, Y\leq F^{-1}_{Y|D=d,Z=z}(\gamma)].$$ When the outcome is binary we have$$\mathbb E[F^{-1}_{Y|D=d,Z=z}(U) | U \leq \gamma]=\max\left\{0, 1-\frac{1}{\gamma}P[Y=0|D=d,Z=z]\right\}.$$ This formulation provides a general truncation formula that applies to any type of outcomes, continuous, discrete, or mixed.

theoremSuppose that Assumptions (ref), and restrictions $\mathcal R_T$ hold. Whenever $\Theta_{I}(\mathcal R_T)\neq \emptyset$, then the following bounds are pointwise sharp: \begin{enumerate} • Local Controlled Direct Effect (LCDE): \begin{eqnarray} && \mathbb E_{F^{-1}_{Y|D,Z}}(\gamma^{1,r}_{d,d};d,1) -\overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\gamma^{0,r}_{d,d};d,0) \leq LCDE(d|d,d) \leq \overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\gamma^{1,r}_{d,d};d,1) - \mathbb E_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{0,r}_{d,d};d,0), \nonumber\\ \nonumber \\ && \underline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{1,r}_{d',d};d,1) -y_U \leq \text{LCDE}(d|d',d) \leq \overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{1,r}_{d',d};d,1) - y_L, \text{ for } d\neq d', \nonumber\\ \nonumber\\ && y_L -\overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{0,r}_{d,d'};d,0) \leq \text{LCDE}(d|d,d') \leq y_U - \underline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{0,r}_{d,d'};d,0) \text{ for } d\neq d'. \nonumber \end{eqnarray} • Local Controlled Indirect Effect (LCIE). \begin{eqnarray} && \underline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{1,r}_{l,d};d,1) -y_U \leq \text{LCIE}(1,d, d'|l,d) \leq \overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{1,r}_{l,d};d,1) - y_L, \text{ for } d\neq d'\text{ and any } l, \nonumber\\ \nonumber\\ && \underline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{0,r}_{d,l};d,0) -y_U \leq \text{LCIE}(0,d, d'|d,l) \leq \overline{\mathbb E}_{F^{-1}_{Y|D,Z}}(\underline{\gamma}^{0,r}_{d,l};d,0) - y_L, \text{ for } d\neq d', \text{ and any } l.\nonumber \end{eqnarray} • Aggregate LCDE: \begin{multline} \inf_{\substack{\left\{(p_{d,d}: d \in \{1,...,K\})\right.\\\left. \in \Theta_I(\mathcal R_T)\vphantom{(p_{d,d}: d \in \{1,...,K\})}\right\}}}\sum_{d=l}^{l'} \frac{p_{d,d}}{\sum _{d'=l}^{l'}p_{d',d'}}\left[ \underline{\mathbb E}_{F^{-1}_{Y|D,Z}}\left(\frac{p_{d,d}}{\mathbb P(D=d|Z=1)};d,1 \right) -\overline{\mathbb E}_{F^{-1}_{Y|D,Z}}\left(\frac{p_{d,d}}{\mathbb P(D=d|Z=0)};d,0 \right)\right] \\ \leq \sum _{d=l}^{l'} \frac{p_{d,d}}{\sum _{d'=l}^{l'}p_{d',d'}}\text{LCDE}(d|d,d) \leq \\ \sup_{\substack{\left\{(p_{d,d}: d \in \{1,...,K\})\right.\\\left. \in \Theta_I(\mathcal R_T)\vphantom{(p_{d,d}: d \in \{1,...,K\})}\right\}}} \sum _{d=l}^{l'} \frac{p_{d,d}}{\sum _{d'=l}^{l'}p_{d',d'}}\left[ \overline{\mathbb E}_{F^{-1}_{Y|D,Z}}\left(\frac{p_{d,d}}{\mathbb P(D=d|Z=1)};d,1 \right) -\underline{\mathbb E}_{F^{-1}_{Y|D,Z}}\left(\frac{p_{d,d}}{\mathbb P(D=d|Z=0)};d,0 \right)\right].\nonumber \end{multline} \end{enumerate}

The derivation of the bounds in Theorem (ref) comes from extending the horowitz_identification_1995 bounding approach summarized in Lemma (ref) in Appendix (ref). However, demonstrating their sharpness presents a considerably more complex challenge. This involves showing that solving equations ((ref)) to ((ref)) for all $y \in \mathcal{Y}$ and $d \in \{1,...,K\}$ while imposing the restrictions defined in $\mathcal{R}_T$ consistently produces the same information as the closed-form bounds presented in Theorem (ref). As explained above, the absence of the IV exclusion restrictions facilitates this result.

Theorem (ref) (i) indicates that, without additional assumptions on the potential outcome distributions, the derived bounds can best determine the direction (sign) of the within-firm effect at layer $d$ solely for individuals who remain with firm $d$ under any treatment assignment $Z$, i.e., $\mathbb E[Y_{1,d}-Y_{0,d}|T=(d,d)]$. This finding is somewhat intuitive given that these “stayers" are equivalent to the so-called “always-employed" in Lee's model, where firm heterogeneity is not taken into account. On the other hand, the bounds for those who switch firms due to treatment (“switchers”), such as $\mathbb E[Y_{1,d}-Y_{0,d}|T=(d, d')]$ and $\mathbb E[Y_{1,d}-Y_{0,d}|T=(d',d)]$ for $d\neq d'$, always include $0$. This is because the observed data $(Y,D,Z)$ do not reveal any information on the following unobserved counterfactuals $\mathbb E[Y_{0,d}|T=(d', d)]$ and $\mathbb E[Y_{1,d}|T=(d, d')]$.

Similarly, Theorem (ref) (ii) reveals that in the absence of additional restrictions on the potential outcome distributions, it is impossible to identify the sign of $\text{LCIE}(z,d,d'|t)=\mathbb E[Y_{z,d}-Y_{z,d'}|T=t]$. This underscores the inherent challenges of identifying some specific treatment effects without imposing further assumptions on the outcome distributions.

Finally, Theorem (ref) (iii) presents the closed form bounds that correspond to the weighted average of $\text{LCDE}(d|d,d)$, $\sum _{d=l}^{l'} \frac{p_{d,d}}{\sum _{d=l}^{l'}p_{d,d}}\text{LCDE}(d|d,d)$. These bounds are sharp and take into account the interdependence between equations ((ref)) and ((ref)).

remarkTo derive bounds on the aggregate LCDE, defined as $\sum_{d=l}^{l'} \frac{p_{d,d}}{\sum_{d'=l}^{l'} p_{d',d'}} \text{LCDE}(d \mid d,d),$ one might be tempted to adopt a na\"ive approach by taking a weighted average of the pointwise sharp bounds derived in Theorem (ref) (i). However, this approach not only fails to provide sharp bounds on the aggregate quantity but may also yield invalid bounds. We discuss this in more detail in Appendix (ref). The difference between the sharp bounds and na\"ive “bounds” is illustrated using simulations ((ref)) in (ref), and using our empirical applications ((ref)) in (ref).

To better illustrate our theoretical results, in Appendix (ref), we examine the special case with two firm types, which follows our empirical applications below. We derive closed-form expressions for the bounds on response-type probabilities. Plugging the resulting expression(s) into the bounds given in Theorem 1 (i)-(ii) yields analytic expressions for the bounds on the LCDEs and LCIEs. Through a numerical illustration, we show that our bounds can be sufficiently informative to distinguish cases with a positive within-firm effect from those with no within-firm effect, even when both yield strictly positive Lee bounds.

Inference

Inference on the causal parameters considered in (ref) can often be performed using existing methods for inference on parameters bounded by truncated conditional expectations. The three sets of bounds in (ref) are known functions of the $(D,Z)$ conditional expectations of $Y$ truncated above or below at particular quantiles. Inference in such settings is complicated by the need to estimate two nuisance parameters: the conditional quantile functions themselves and the truncation quantile level.

In parts (i) and (ii) of (ref), the truncation quantile levels are of the form $\underline{\gamma}^{z,r}_{d,d'}=\frac{\underline{p}^r_{d,d'}}{\mathbb P(D=d|Z=z)}$ and these can often be estimated at a faster rate than the second step (truncated conditional expectation). Moreover, each bound involves either only one truncated conditional expectation, or truncated conditional expectations that can be independently estimated. Therefore, in such cases, inference on the parameters considered in (ref)(i)-(ii) can be performed by plugging in the estimators of lee_training_2009, semenova_generalized_2020, or olma_nonparametric_2021 for the truncated conditional expectation(s) and adapting the inference approaches discussed there (subject to the regularity conditions outlined there).\footnote{However, $\underline{\gamma}^{z,r}_{d,d'}$ can often be only directionally differentiable with respect to the propensity score vector, as is evident in the closed form expressions for the 2 firms type case provided in (ref). In such cases, inference procedures that depend on the bootstrap will need to be adapted to remain valid (see fangInferenceDirectionallyDifferentiable2019 for details).} These approaches allow for conditioning on, and aggregating over the covariates $X$, which can lead to much tighter bounds than the unconditional case, as noted by lee_training_2009. The approach proposed in lee_training_2009 applies when $X$ is finitely supported, whereas the approaches developed by semenova_generalized_2020 and olma_nonparametric_2021 allow $X$ to be continuous.

Inference on the “aggregate” parameters considered in part (iii) of (ref) is more complicated since the truncation quantile level to be estimated depends on the solution to a higher-dimensional optimization problem involving the outcome distributions. We defer the development of estimation and inference methods for this case to future work.

Empirical Applications: Job Corps Study and WorkAdvance RCT

Application \#1: Job Corps Study

Job Corps is the largest residential career training program in the U.S. It is free for participants and targets disadvantaged people aged 16 to 24 with the aim of helping these people become more responsible, employable, and productive citizens (johnson_national_1999 johnson_national_1999). Most participants live at a local Job Corps center and complete 440 hours of academic instruction and 700 hours of vocational training. Job Corps also provides job search assistance upon participant completion of the program. The typical participant completes the program over a span of 30 weeks. Job Corps has trained more than two million individuals since its inception under the Economic Opportunity Act of 1964. The program trains over 60,000 enrollees per year, at roughly $130$ Job Corps centers nationwide, with an estimated cost of 34,301 USD per enrollee and 57,312 USD per graduate (liu_estimating_2020 liu_estimating_2020).

During the mid- to late 1990s, the U.S. Department of Labor funded a randomized evaluation of Job Corps, which was implemented by Mathematica Policy Research, Inc. Existing evaluations of the Job Corps Study include schochet_national_2001, schochet_does_2008, lee_training_2009, and blanco_effects_2013. The Job Corps Study randomized 80,883 eligible individuals who applied to Job Corps for the first time between November 1994 and December 1995 into two groups: (i) 5,977 individuals into the control group who were embargoed from participating in Job Corps for three years and (ii) 74,906 individuals into the treatment group. Of the 74,906 individuals assigned to treatment, 9,409 individuals were randomly selected for data collection and all control individuals were selected for data collection. The final sample therefore consists of 15,386 participants who were interviewed at the time of random assignment and then subsequently 12, 30, and 48 months after random assignment.

We use the publicly available data from the National Job Corps Study (schochet_national_2003 schochet_national_2003). We impose two sample restrictions to address missing values due to interview non-response and sample attrition over time. The first sample restriction, which follows lee_training_2009, is to keep individuals who have nonmissing values for weekly earnings and hours for every week following random assignment. Restricting the sample in this manner decreases the sample size to 9,145 individuals (= 3,599 control units + 5,546 treated units).

The second sample restriction is to keep individuals who have non-missing values of health insurance for the weeks of interest (90, 135, 180 and 208). This comes from our classification of firm type based on the provision of health insurance.\footnote{Another potentially useful classification of firm type is based on industry since Job Corps targets certain sectors. However, industry is not observed in the public use of the Job Corps data. We consider this type of classification in the next empirical application.} This is motivated by dey_flinn_2005 who consider a setting where workers and firms bargain over wages and health insurance and find that, on average, better productivity matches lead to higher wages and the provision of health insurance.\footnote{At the time of the National Job Corps study, there were no legal requirements for firms to provide any of these amenities and, conditional on firm provision, federal law generally prohibited discriminatory provision across workers (united_states_equal_employment_opportunity_commission_federal_2009 united_states_equal_employment_opportunity_commission_federal_2009). The relevant federal laws at the time of the National Job Corps Study included the following. Title VII of the Civil Rights Act of 1964; the Age Discrimination in Employment Act of 1967; Title I and Title V of the Americans with Disabilities Act of 1990 (united_states_equal_employment_opportunity_commission_federal_2009 united_states_equal_employment_opportunity_commission_federal_2009).} This restriction results in a final sample size of 6,403 individuals (= 2,540 control units + 3,863 treated units).

The three key variables of interest are employment, hourly wage, and provision of health insurance for employed individuals. We follow lee_training_2009 by defining employment in a week based on whether an individual has positive earnings in that week and defining the hourly wage in a week by dividing weekly earnings by weekly hours worked. In Appendix C, we present summary statistics for our final sample and demonstrate that, on average, firms that offer these amenities pay higher wages than firms that do not. We also show that the wage distribution for firms that offer amenities stochastically dominates the wage distribution for firms that do not, in both the treatment and control groups. Finally, we show that being randomly assigned to Job Corps led individuals to work at firms with better job amenities compared to the control group. This combined evidence suggests that sample selection is multilayered and motivates the implementation of our sharp bounds to these data.

Multilayered Bounds for Job Corps Study

As a first step, we replicate the bounds reported in lee_training_2009 which, as discussed above, target the parameter $\mathbb E[Y_{1,D_1}-Y_{0,D_0}|D_0>0,D_1>0]$.\footnote{In this section, certain empirical results only presented are for week 90 which follows the preferred specification in lee_training_2009. In these cases, tables and figures for other weeks of interest (135, 180 and 208) are presented in (ref)} In (ref) we report Lee's bounds for weeks 90, 135, 180 and 208 along with the trimming proportion $p\equiv\mathbb P(AE)$, e.g., the share of the always-employed among individuals receiving job training. Lee focuses on week 90 which produces the tightest bounds $[0.0468, 0.0484]$.\footnote{As we detail in (ref), these are Lee's bounds when treating $ln(\text{hourly wage})$ as a continuous variable, as we do throughout this paper. lee_training_2009 uses vingtiles of $ln(\text{hourly wage})$ that produce bounds $[0.0423,0.0428]$.}

We now consider the scenario with two firm types, denoted as $L$ and $H$. Firms are classified based on the provision of health insurance with $H$ denoting firms that offer health insurance and $L$ denoting firms that do not.\footnote{Results for classifying firms based on the provision of pension/retirement benefits and paid vacation are presented in (ref). } As in Section 4, we always impose Assumptions 1-2 and then sequentially impose the restrictions in (ref): (i) $p_{H,H} \geq p_{H,L}$ (more stayers than downward switchers) (ii) $\displaystyle\min_t \mathbb P[T=t] = p_{H,L}$ (more upward switchers than downward switchers) and (iii) $p_{H,L} = 0$ (strong monotonicity).

(ref) presents the estimated propensity scores for each week of interest from the National Job Corps Study.\footnote{Our sample restriction to keep observations with non-missing amenity values for the weeks of interest drops only employed individuals. This restriction mechanically reduces the propensity scores. To ensure comparability with lee_training_2009, we rescale our estimated propensity scores so that the probabilities of employment by treatment status are the same as those reported in lee_training_2009.} As expected, in all weeks $\mathbb P[D>0|Z=1] > \mathbb P[D>0|Z=0]$ the treated individuals are more likely to be employed. The table also shows that $\mathbb P[D=H|D>0,Z=1] > \mathbb P[D=H|D>0,Z=0]$ so that individuals who receive Job Corps training have a higher propensity to be employed in firms that offer health insurance than individuals who do not receive Job Corps training, conditional on employment, consistent with the evidence presented in Appendix (ref).

table[table omitted — 412 chars of source]

Using the week 90 propensity scores from (ref), (ref) presents the identified set for $\left(p_{L,L},p_{H,H}\right)$. Naturally, incorporating additional restrictions on the response types leads to a tightening of the identified sets. Having characterized the identified set of response-type probabilities, we now present our multilayered bounds.

figure[figure omitted — 735 chars of source]

Recall that under (ref), the support of possible response types is as follows:

align*[align* omitted — 165 chars of source]

The always-employed (AE) definition used in lee_training_2009 therefore combines four different response types: $\{D_0>0, D_1>0\}=\left\{\left(L, L\right), \left(H, H\right),\left(L, H\right),\left(H, L\right)\right\}\equiv AE$. We focus on the bounds for stayers, defined as the response types $(H,H)$ and $(L,L)$.

(ref) presents our multilayered bounds for $\mathbb E[Y_{1,H}-Y_{0,H}|T=\left(H, H\right) ]$ and $\mathbb E[Y_{1,L}-Y_{0,L}|T=\left(L, L\right) ]$, along with our aggregate bounds, for weeks 90, 135, 180 and 208. We illustrate the bounds in the baseline case (only (ref) is imposed) and also when we sequentially impose the following restrictions: (i) more stayers than downward switchers, (ii) more upward switchers than downward switchers, and (iii) strong monotonicity. As discussed above, the Lee bounds for week 90 are $[0.0468, 0.0484]$. Focusing on the type $H$ firms (firms that offer health insurance), our estimates indicate $\mathbb E[Y_{1,H}-Y_{0,H}|T=\left(H, H\right) ] \in [-2.1415,2.3907]$. Assuming more stayers than downward switchers tightens these bounds to $[-0.4214, 0.5020]$. Moreover, assuming that $(H,L)$ is the smallest response type, further tightens them to $[-0.0023,0.0754]$. Finally, assuming strong monotonicity narrows these bounds to $[-0.0018,0.0750]$.

We find a similar pattern of results for the $L$-type bounds. These patterns persist across all weeks of interest and also when classifying firms based on the provision of alternative amenities (paid vacation and retirement/pension benefits) as shown in Appendix (ref). (ref) reports our estimated bounds across weeks.\footnote{Appendix (ref) provide our estimated within-firm-type bounds when classifying firms based on the provision of paid vacation and retirement/pension benefits, respectively.} The aggregate multilayered bounds are reported in Appendix (ref).\footnote{Appendix (ref) provide our estimated aggregate bounds when classifying firms based on the provision of paid vacation and retirement/pension benefits, respectively.}

table[table omitted — 1,365 chars of source]

Thus, while the conventional Lee bounds are strictly positive, our bounds for the within-firm effects include 0 even under strong assumptions on the response types. This suggests that Lee bounds may capture a pure sorting response to job training rather than a direct wage effect.

figure[figure omitted — 845 chars of source]

Application \#2: WorkAdvance RCT

As a second empirical application, we evaluate MDRC's WorkAdvance sectoral employment program. WorkAdvance trains disadvantaged adults with the goal of matching them to high-quality jobs in specific industries with strong labor demand. The MDRC WorkAdvance demonstration featured four different community-based providers in three different locations (New York City, Tulsa, and Northeast Ohio). In this section, we apply our bounds to the Madison Strategies RCT in Tulsa, Oklahoma, and the Towards Employment RCT in Northeast Ohio.\footnote{Our primary dataset is state-level administrative Unemployment Insurance (UI) data, obtained via a confidential data use agreement with MDRC. The administrative UI data for Oklahoma (Madison Strategies RCT) and Ohio (Towards Employment RCT) provide the two-digit North American Industry Classification System (NAICS) code for the industry in which the participant worked, while the New York data (Per Scholas and St. Nicks Alliance RCTs) do not. Therefore, we focus on the Oklahoma and Ohio evaluations.}\textsuperscript{,}\footnote{For more details on the WorkAdvance RCT and an evaluation of the impacts, see katz_why_2022. For Madison Strategies, the reported impact is $12.4$ percent on earnings two years after the program. For Towards Employment, the impact is $14$ percent.}

The Madison Strategies RCT targeted high-quality jobs in Transportation and Manufacturing. The duration of the program ranged from 4 to 32 weeks. The curriculum focused on pre-employment career readiness and sector-specific skills training. It also involved sector-specific placement services, post-employment retention, and advancement services. Upon completion of the program, participants received the appropriate certification.

The Madison Strategies RCT enrollment period was from June 2011 to June 2013. It targeted low-income adults who met the following skill requirements: (i) tested at the eighth grade level; (ii) passed a behavioral assessment; (iii) passed mechanical aptitude and manual dexterity exams; (iv) had a driver's license. The control group was not eligible to receive WorkAdvance services. Eligible potential participants (697 units) were randomly assigned to treatment (353 units) and control (344 units).

The Towards Employment RCT targeted jobs in Health Care and Manufacturing. The target participants were low-income adults who: (i) tested at the sixth to tenth grade level; (ii) passed background and drug tests. The curriculum, program duration, and enrollment period were the same as Madison Strategies. Eligible potential participants (698 units) were randomly assigned to treatment (349 units) and control (349 units).

The key variables for our analysis of both RCTs are quarterly wages along with employment status and sector. We define quarterly wages as quarterly earnings subject to UI and follow both lee_training_2009's and our evaluation of the Job Corps RCT in defining quarterly employment status based on whether an individual has positive earnings in a quarter.\footnote{To interpret our estimates as reflecting wage effects, this requires that job training does not affect hours of work.} We assume that $d=H$ if the firm of employment is in the target sector and $d=L$ if the firm is not. All of our results focus on 8 quarters post-random assignment.

Appendix (ref) show summary statistics at baseline, as well earnings and employment outcomes 8 quarters post-randomization, for Madison Strategies and Towards Employment, respectively. For Madison Strategies, the employment rates 8 quarters post randomization are quite similar between the treatment group ($0.67$) and the control group ($0.66$). However, the composition of employment by target sector differs meaningfully by treatment status: among the treatment group, $44$ percent are employed in the target sector whereas, among the control group, only $31$ percent are. Towards Employment increased overall employment from $0.62$ to $0.69$ and the share of employment in the target sector from $0.47$ to $0.51$. (ref) shows the propensity score estimates for both Madison Strategies and Towards Employment which further illustrate the sorting effects of both RCTs. As was the case for the Job Corps study, this combined evidence suggests multilayered sample selection and again motivates the use of our bounds.

Appendix (ref) shows that average wages are higher in the target sector and Appendix (ref) (Madison Strategies) and Appendix (ref) (Towards Employment) show that the wage distributions in the target sector stochastically dominate the wage distributions in non-targeted sectors, conditional on treatment status.

table[table omitted — 373 chars of source]
figure[figure omitted — 558 chars of source]
figure[figure omitted — 558 chars of source]

(ref) and (ref) show the identified sets for $\left(p_{L,L},p_{H,H}\right)$ for Madison Strategies and Towards Employment, respectively. Although sorting appears to be stronger in the WorkAdvance RCT than in the Job Corps study, $p^*_{H,H}$ and $p^*_{L,L}$ are fairly similar. This is driven by differences in $p_{0,0}$, which is a point identified by $P(D=0|Z=1)$. In Job Corps, $p_{0,0}=0.54$. By comparison, $p_{0,0}=0.33$ in Madison Strategies and $p_{0,0}=0.31$ in Towards Employment. Thus, the reason for the similarities in $p^*_{H,H}$ and $p^*_{L,L}$ is driven by much stronger sorting (larger $p_{LH}$) in WorkAdvance. All else being equal, this attenuates $p^*_{d,d}$ and increases the bounds.

table[table omitted — 1,021 chars of source]
figure[figure omitted — 917 chars of source]
figure[figure omitted — 910 chars of source]

(ref) present Lee bounds, our multilayered bounds for $\mathbb E[Y_{1,H}-Y_{0,H}|T=\left(H, H\right) ]$ and $\mathbb E[Y_{1,L}-Y_{0,L}|T=\left(L, L\right) ]$, along with our aggregate bounds, for Madison Strategies and Towards Employment, respectively. As before, we present our multilayered bounds in the baseline case (only (ref) are imposed) and then sequentially impose the following restrictions: (i) more stayers than downward switchers, (ii) more upward switchers than downward switchers, and (iii) strong monotonicity. Even under the most restrictive assumption of strong monotonicity, the within-firm bounds include 0 for both sites, suggesting that the impact of WorkAdvance on wages is primarily driven by sorting to the target sector. Notably, this contrasts with Lee bounds in the case of Madison Strategies, which are positive and tight. (ref) presents the within-firm-type multilayered bounds for Madison Strategies (top panel) and Towards Employment (bottom panel). Appendix (ref) present the aggregate multilayered bounds and Lee bounds, respectively, for both sites.

Conclusion

This paper develops a new methodology to partially identify the causal effect of job training on wages in the presence of multilayered sample selection. We define new treatment effects that operate within and between firms and provide a new identification approach that extends the horowitz_identification_1995 bounds. As a proof of concept, we show how to empirically implement these bounds by considering applications to the Job Corps Study and the WorkAdvance RCT.

Although we consider our approach in the context of job training where a layer corresponds to a firm, we view it as naturally extending to other settings. In particular, it applies to any setting where sample selection is multilayered. As an example, consider a setting where a researcher is interested in estimating the causal effect of a tuition subsidy on labor market outcomes.\footnote{bettinger_long-run_2019 evaluate the impact of California's state-based financial aid on long-run earnings.} The subsidy may have an effect on the type of institution that an individual enrolls in and graduates from. If earnings depend on institutional quality, part of the earnings effects of the subsidy could reflect the value-added of institutions that are affected by the subsidy.

Although our framework has focused mainly on nonparametric (partial) identification, we are currently reexamining the classic parametric sample selection approach of heckman_sample_1979 in the context of multilayered sample selection, as well as its semiparametric version discussed in honore_selection_2020. By imposing additional structure on the unobservables, this approach has the potential to significantly tighten the bounds and may achieve point identification of causal effects.