EconBase
← Back to paper

Probability of Causation with Sample Selection: A Reanalysis of the Impacts of Jóvenes en Acción on Formality

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

59,244 characters · 10 sections · 55 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Probability of Causation with Sample Selection: A Reanalysis of the Impacts of Jóvenes en Acción on Formality

\def\spacingset#1{ {#1}} \spacingset{1}

\if00 \fi

\if10 {

center[center omitted — 148 chars of source]

} \fi

abstractThis paper identifies the probability of causation when there is sample selection. We show that the probability of causation is partially identified for individuals who are always observed regardless of treatment status and derive sharp bounds under three increasingly restrictive sets of assumptions. The first set imposes an exogenous treatment and a monotone sample selection mechanism. To tighten these bounds, the second set also imposes the monotone treatment response assumption, while the third set additionally imposes a stochastic dominance assumption. Finally, we use experimental data from the Colombian job training program Jóvenes en Acción to empirically illustrate our approach's usefulness. We find that, among always-employed women, at least 10.2% and at most 13.4% transitioned to the formal labor market because of the program. However, our 90%-confidence region does not reject the null hypothesis that the lower bound is equal to zero.

{\it Keywords:} Probability of Causation, Sample Selection, Partial Identification, Job Training Programs.

\spacingset{1.8}

Introduction

\newsavebox{\tablebox} \newlength{\tableboxwidth}

{ Many policy evaluation questions involve two simultaneous identification challenges: the causal parameter of interest depends on the joint distribution of potential outcomes {Heckman1997,Pearl1999,Tian2000,Jun2019,Cinelli2021}, and sample selection is present {Lee2009,Chen2015,Bartalotti2021}. For example, when evaluating the effects of job training programs {Heckman1999a,Attanasio2011,Attanasio2017,Blanco2018}, the researcher may be interested in learning to what extent the transition from informal to formal employment can be attributed to the policy. Still, she only observes formality status among those who are employed. This double identification challenge also arises when researchers analyze the effects of a political campaign on agents' opinions {DellaVigna2007,DellaVigna2010} if agents may not reply to the researchers' survey.}

In this paper, we derive novel sharp bounds around the probability of causation parameter Pearl1999,Tian2000,Jun2019,Cinelli2021 for individuals who self-select into the sample regardless of their treatment assignment. The probability of causation parameter summarizes one crucial aspect of the effects of treatments on binary outcomes: the proportion of individuals who benefit from being treated within the subgroup who would, counterfactually, experience a negative untreated outcome. Thus, our target parameter helps researchers gauge to what extent the transition from one state to another can be attributed to the treatment in a relevant latent sub-population.

Our partial identification strategies are based on three increasingly restrictive sets of assumptions. They extend the identification of probabilities of causation to scenarios with endogenous sample selection. In our model, treatment effects can be related to the sample selection mechanism even though treatment take-up is exogenous. We also discuss when our assumptions have identification power and how to test them through necessary observable conditions.

Our first identification result relies on a monotone sample selection mechanism. This condition imposes that treatment has a non-negative effect on the sample selection indicator for all individuals. In the job training example, this restriction implies that the treatment can move workers into employment but never out of employment.

Our second result further assumes a monotone treatment response to tighten the identified bounds. This condition imposes that treatment has a non-negative effect on the potential outcomes for all individuals. In the job training example, this restriction implies that the treatment can move workers into formal jobs but never into informal jobs.

Our final result additionally relies on a stochastic dominance assumption to further reduce the identified set. This condition imposes that the sub-population that self-selects into the sample regardless of the treatment status has higher treated potential outcomes than the sub-population that self-selects into the sample only when treated. In the job training example, this restriction implies that the agents who are always employed are more likely to have a formal job if treated than those who are employed only when treated.

Additionally, we propose parametric estimators for all these bounds. We also combine the precision-corrected bounds proposed by Chernozhukov2013 with a Bonferroni-style correction to derive confidence regions that contain the identified region with a pre-specified confidence level.

To empirically illustrate the usefulness of our approach, we provide bounds for the probability of causation of an intensive training program: Jóvenes en Acción. This program aimed to improve the labor market prospects and, in particular, the quality of jobs held by disadvantaged youths in seven large cities in Colombia. It offered in-classroom intensive training in occupational skills to qualify unemployed individuals for locally demanded jobs. Additionally, it focused on socioemotional development and offered on-the-job internships with formal employers.

Previous research Attanasio2011,Attanasio2017 finds that this program positively affects employment and unconditional formality. However, less is known about whether the program achieves its goal of improving job quality conditioning on having a job. We study its effects on the job quality margin by considering the share of women that transitioned to the formal labor market because they participated in the training program. We find that incorporating selection and bounding the probability of causation leads to a pessimistic view of the program’s impacts. More precisely, we find that at most 13.4% of the always-employed women switched their formality status because they were assigned to the Jóvenes en Acción training program. Moreover, our 90%-confidence region includes the zero, implying that we cannot reject the null hypothesis that our target parameter's lower bound is equal to zero.

Concerning its theoretical contribution, our work is inserted in two research areas: identification of probabilities of causation and identification in the presence of sample selection.

Heckman1997 motivate the focus on a parameter closely connected to the probability of causation based on the political economy of policy evaluation. They argue that a program would only be adopted in a democracy if it benefited most people in the population. {They either make strong probabilistic assumptions or impose model restrictions on treatment take-up decisions} to point-identify this parameter, while we focus entirely on partial identification strategies based on a menu of easily interpretable assumptions.

Pearl1999 and Tian2000 discuss how to interpret and partially identify probabilities of causation in a single population where agents are always observed. Cinelli2021 extend their work by combining experimental results from multiple trials to extrapolate probabilities of causation from one population to a different population. Moreover, Jun2019 extend their work by considering endogenous selection into treatment.

We extend the work by Pearl1999 and Tian2000 in a different direction. We identify probabilities of causation when the agents' realized outcomes may not be observed due to endogenous sample selection. To do so, we combine the tools developed in the literature about probabilities of causation with the trimming bounds developed in the sample selection literature Horowitz1995,Lee2009,Chen2015,Bartalotti2021.

Concerning its empirical contribution, our work is inserted in the literature about job training programs. Attanasio2011 and Attanasio2017 analyze the average treatment effect (ATE) of Jóvenes en Acción on short and long-term outcomes associated with labor force attachment. We extend their work by analyzing a treatment effect parameter that focuses on job quality instead of labor force attachment. Importantly, Blanco2018 also analyze the impact of a job training program on job quality using partial identification strategies. However, we focus on different contexts (Job Corps v. Jóvenes en Acción) and different target parameters (Quantile Treatment Effects v. Probabilities of Causation).

{This paper is organized as follows. Section (ref) presents our structural model, sample selection mechanism, and identifying assumptions. It also discusses the testable restrictions imposed by our model. Section (ref) describes our main identification results, while Section (ref) proposes a parametric estimator for our bounds and discusses an inferential method for the identified region. Moreover, Section (ref) discusses the results of our empirical application. In the end, Section (ref) concludes.

Moreover, we also have an online appendix with additional details and results. Appendix (ref) presents the proofs of all our identification results, while Appendix (ref) intuitively explains them using a numerical example. Moreover, Appendix (ref) brings a detailed discussion about the testable restrictions of our identifying assumptions, while Appendix (ref) compares our target parameter against other causal parameters. Furthermore, Appendix (ref) detailedly explains our estimator and inferential method. Finally, Appendix (ref) presents additional empirical results.}

Analytical Framework

We aim to identify the probability of causation Pearl1999,Tian2000,Jun2019,Cinelli2021 within the always-observed subsample. To do so, we consider the generalized sample selection model Lee2009, described in the potential outcomes framework:

equation[equation omitted — 188 chars of source]

where $D$ is the treatment status indicator (in our application, being selected to enroll in the Jóvenes in Acción training program). The variable $Y^{*}$ is the possibly censored realized outcome variable (indicator for whether the agent has a formal or informal job) with support $\mathcal Y = \left\lbrace 0, 1 \right\rbrace$, while $Y_{0}^{*}$ and $Y_{1}^{*}$ are the possibly censored potential outcomes when the person is untreated and treated, respectively. Similarly, $S$ is the realized sample selection indicator (indicator for whether the agent holds a job), and $S_{0}$ and $S_{1}$ are potential sample selection indicators when individuals are untreated and treated. Moreover, $Y$ is the uncensored observed outcome. {Finally, $X$ is a set of exogenous covariates (indicator variables for each course-city pair in the Jóvenes in Acción training program) whose support is denoted by $\mathcal{X}$. The researcher observes only the vector $\left(Y,D,S,X\right)$,} while $Y^{*}_{1}$, $Y^{*}_{0}$, $S_{1}$ and $S_{0}$ are latent variables.

In the setting analyzed here, learning about the probability of causation Pearl1999,Tian2000,Jun2019,Cinelli2021 is further complicated by the potential for nonrandom sample selection. As pointed out by Lee2009, even in the simpler case of the average treatment effect (ATE), point identification is no longer possible, leading him to derive bounds for the ATE.

This paper combines the insights of these literatures to develop sharp bounds for the probability of causation under sample selection. To do so, we define four latent groups based on the potential sample selection indicators. The sub-populations are defined as: always-observed ($S_{0} = 1, S_{1} = 1$), observed-only-when-treated ($S_{0} = 0, S_{1} = 1$), observed-only-when-untreated ($S_{0} = 1, S_{1} = 0$), and never-observed ($S_{0} = 0, S_{1} = 0$). They are denoted by $OO$, $NO$, $ON$ and $NN$ respectively.

Following Zhang2008 and Lee2009, we focus on the always-observed sub-population $\left(S_{0} = 1, S_{1} = 1\right)$. Importantly, this sub-population is the only group with censored potential outcomes observed in both treatment arms. For the other three sub-populations, treatment effect parameters are not point-identified or bounded in a non-trivial way without further parametric assumptions because at least one of the potential outcomes ($Y_{0}^{*}$ or $Y_{1}^{*}$) is never observed. Since we focus on a fully non-parametric identification strategy, we do not discuss parametric identification of unconditional treatment effect parameters or treatment effect parameters associated with the latent groups $ON$, $NO$ and $NN$.

Our target parameter is the probability of causation within the sub-population that is always observed:

equation[equation omitted — 138 chars of source]

and depends on the joint distribution of potential outcomes $\left(Y_{0}^{*}, Y_{1}^{*}\right)$.

The unconditional probability of causation $\left(\mathbb{P}\left[\left. Y_{1}^{*} = 1 \right\vert Y_{0}^{*} = 0\right]\right)$ captures, within the sub-population whose untreated potential outcome is equal to zero, the share whose treated potential outcome is equal to one. Intuitively, it measures the share of agents who benefited from the treatment within the subgroup with a negative untreated outcome. In our empirical application, the unconditional probability of causation captures, within the population with an informal job if untreated, the share of workers with a formal job if treated. (In Appendix (ref), we compare the probability of causation parameter against other treatment effect parameters frequently discussed in the literature. In particular, we discuss the concepts of “persuasion effect” proposed by Jun2019, of “distribution of gains at selected base state values” and “probability of employed with treatment, not employed without treatment” proposed by Heckman1997, and of the average treatment effect.)

Our target parameter in Equation (ref) focuses on the probability of causation for the always-observed latent group. In our empirical application, our target parameter captures, within the population who is employed regardless of treatment status and has an informal job if untreated, the share of workers with a formal job if treated. Intuitively, we focus on the population who is always employed and found a job of higher observable quality because they were assigned to the Jóvenes in Acción training program.

Analogously to Heckman1997, Jun2019 and Cinelli2021, identification of $\theta^{OO}$ is complicated because it depends on the joint distribution of the potential outcomes $\left(Y_{0}^{*}, Y_{1}^{*}\right)$ while, even in a randomized controlled trial, we can only identify the marginal distributions of the potential outcomes. Analogously to Lee2009, identification of $\theta^{OO}$ is complex because sample selection is nonrandom and possibly impacted by the treatment.

To simultaneously address these issues, we follow a layered policy analysis approach Manski2011 and consider three sets of assumptions to partially identify our target parameter. The identified set {weakly} shrinks when stronger assumptions are used. Assumptions (ref)-(ref) are sufficient {to derive sharp bounds around $\theta^{OO}$.}

assumption[Random Assignment] {Treatment $D$ is randomly assigned after conditioning on the covariates, i.e., $\left. D \protect\mathpalette{\protect\independenT}{\perp} (Y^{*}_{0},Y^{*}_{1},S_{0},S_{1}) \right\vert X$.}

Assumption (ref) modifies the standard independence assumption Imbens2009a to account for sample selection. Instead of assuming that the treatment variable is independent of the potential outcomes only, we also assume independence between the treatment variable and the potential sample selection indicators similarly to Lee2009. In our empirical application, it holds conditionally on course indicators because the possibility of enrolling in the Jóvenes in Acción training program was randomly allocated within oversubscribed courses.

assumption[Positive Mass] {Both treatment groups and the always-observed sub-population who chooses $Y_{0}^{*} = 0$ exist after conditioning on the covariates, i.e., $0 < \mathbb{P}\left[\left. D = 1 \right\vert X = x\right] < 1$ and $\mathbb{P}\left[\left. Y_{0}^{*} = 0, S_{0} = 1, S_{1} = 1 \right\vert X = x \right] > 0$ for every value $x \in \mathcal{X}$.}

Assumption (ref) is crucial for the identification results because it ensures that our sub-population of interest exists. In our empirical application, it requires that oversubscribed courses are the only ones to exist and that there are always-employed individuals who have an informal job when untreated for every course-city pair.

assumption[Monotone Sample Selection] Treatment has a non-negative effect on the sample selection indicator for all individuals, i.e., $S_{1} \geq S_{0}$.

Assumption (ref) is a monotonicity restriction that rules out the existence of the observed-only-when-untreated sub-population and is commonly used in the literature about sample selection Lee2009, Chen2015,Bartalotti2021. In our empirical application, it imposes that the Jóvenes in Acción training program can only move agents into employment. {This assumption is plausible if the training program improves the workers' social skills, boosting their performance in job interviews. However, this assumption is implausible if the training program stimulates them to pursue further education.}

Assumptions (ref)-(ref) form our first set of assumptions required to {derive sharp bounds around} the probability of causation within the always-observed individuals. Importantly, this set of assumptions has a testable implication, as discussed in Lemma (ref).

Even though these assumptions are sufficient to {derive sharp bounds around} $\theta^{OO}$, the identified set {may} be substantially tightened by additionally imposing that the treatment can only increase the possibly censored potential outcome.

assumption[Monotone Treatment Response] Treatment has a non-negative effect on the censored outcome variable for all individuals, i.e., $Y_{1}^{*} \geq Y_{0}^{*}$.

Assumption (ref) is a monotonicity restriction common in the partial identification literature Manski1997,Manski2000,Jun2019. In our empirical application, it imposes that the Jóvenes in Acción training program can only move agents from informal jobs to formal ones. {This assumption is plausible if the training program increases the workers' productivity. However, this assumption is implausible if the training program stimulates them to open their own informal firms.}

Assumptions (ref)-(ref) form our second set of assumptions required to {derive sharp bounds around} the probability of causation within the always-observed individuals. Importantly, this set of assumptions has an extra testable implication, as discussed in Proposition (ref).

We {may} further shrink the identified set around $\theta^{OO}$ by adding Assumption (ref) and completing our final set of identifying assumptions.

assumption[Stochastic Dominance] {After conditioning on the covariates, the treated counterfactual for the always-observed group stochastically dominates the treated counterfactual for the observed-only-when-treated group, i.e., $$\mathbb{P}\left[\left. Y_{1}^{*} = 1 \right\vert S_{0} = 1, S_{1} = 1, X = x \right] \geq \mathbb{P}\left[\left. Y_{1}^{*} = 1 \right\vert S_{0} =0, S_{1} = 1, X = x \right]$$ for every value $x \in \mathcal{X}$.}

Assumption (ref) is a stochastic dominance restriction that imposes that the always-observed sub-population has higher potential treated outcomes than the observed-only-when-treated group. This type of assumption is common in the literature Imai2008,Blanco2013,Huber2015,Huber2017,Bartalotti2021 and is intuitively based on the argument that some sub-groups have more favorable underlying characteristics than others. In our empirical application, it imposes that the always-employed sub-population has higher potential formality when treated than the employed-only-when-treated sub-population. {This assumption is plausible if individuals with better employment status are more likely to have better (i.e., formal) jobs because they are more productive or skillful. However, this assumption will be invalid if always-employed individuals have jobs because they are willing to accept any working opportunity, even if it is an informal job.}

Testable Restrictions

This subsection discusses testable restrictions implied by the assumptions described in Section (ref).

First, the testable restriction implied by Assumptions (ref)-(ref) was already derived by Lee2009. We state it here for completeness.

lemmaUnder Assumptions (ref)-(ref), the following inequality holds: \begin{equation*} \mathbb{P}\left[\left. S = 1 \right\vert D = 1, X\right] - \mathbb{P}\left[\left. S = 1 \right\vert D = 0, X\right] \geq 0. \end{equation*}

Second, we derive a set of testable restrictions implied by Assumptions (ref)-(ref) as detailed in Proposition (ref). Its proof is in Appendix (ref).

propositionUnder Assumptions (ref)-(ref), the following inequalities hold: \begin{align} \mathbb{P}\left[\left. S = 1 \right\vert D = 1, X\right] - \mathbb{P}\left[\left. S = 1 \right\vert D = 0, X\right] & \geq 0, \\ \mathbb{P}\left[\left. Y = 1 \right\vert D = 1, X \right] - \mathbb{P}\left[\left. Y = 1 \right\vert D = 0, X\right] & \geq 0. \end{align}

Intuitively, the monotonicity of the sample selection indicator and the censored potential outcome implies that treatment positively affects the uncensored potential outcome.

These restrictions can be easily tested using two one-sided tests of mean differences. {In Appendix (ref), we discuss the relationship between these testable restrictions and the bounds proposed in Section (ref).}

Identification Results

{

In this section, we partially identify the probability of causation within the always-observed sub-population (Equation (ref)). To do so, we start by identifying the conditional probability of causation within the always-observed sub-population, $$\theta^{OO}\left(x\right) \coloneqq \mathbb{P}\left[\left. Y_{1}^{*} = 1 \right\vert Y_{0}^{*} = 0, S_{0} = 1, S_{1} = 1, X = x\right],$$ and, then, integrate over the distribution of the covariates for the always-observed sub-population with a zero untreated potential outcome, $ \left. X \right\vert Y_{0}^{*} = 0, S_{0} = 1, S_{1} = 1,$ to identify our target parameter $\theta^{OO}$ (Equation (ref)).

First, we identify $\theta^{OO}\left(x\right)$ under our three sets of assumptions and discuss the identifying power of our assumptions.

Combining Assumptions (ref)-(ref), we derive sharp bounds around the conditional probability of causation within the always-observed sub-population as detailed in Proposition (ref). Its proof is in Appendix (ref).

propositionUnder Assumptions (ref)-(ref), the conditional probability of causation is partially identified for the always-observed subgroup, i.e., \begin{equation*} LB_{1}\left(x\right) \leq \theta^{OO}\left(x\right) \leq UB_{1}\left(x\right), \end{equation*} where $$LB_{1}\left(x\right) \coloneqq \max\left\lbrace \dfrac{ \left[B\left(x\right) - \left(1 - A\left(x\right) \right)\right] \cdot \left[A\left(x\right)\right]^{-1} + C\left(x\right) - 1}{C\left(x\right)} , 0 \right\rbrace,$$ $$UB_{1}\left(x\right) \coloneqq \min \left\lbrace \dfrac{B\left(x\right) \cdot \left[A\left(x\right)\right]^{-1}}{C\left(x\right)} , 1 \right\rbrace,$$ $A\left(x\right) \coloneqq \dfrac{\mathbb{P}\left[\left.S = 1 \right\vert D = 0, X = x \right]}{\mathbb{P}\left[\left.S = 1 \right\vert D = 1, X = x \right]},$ $B\left(x\right) \coloneqq \mathbb{P}\left[\left.Y = 1 \right\vert S = 1, D = 1, X = x \right],$ and $C\left(x\right) \coloneqq \mathbb{P}\left[\left.Y = 0 \right\vert S = 1, D = 0, X = x \right]$ for every value $x \in \mathcal{X}$. Moreover, these bounds are sharp.

{

Corollary (ref) describes when Assumptions (ref)-(ref) have identifying power, i.e., the identified set in Proposition (ref) is strictly smaller than the unit interval. Its proof is in Appendix (ref).

corollaryIf Assumptions (ref)-(ref) hold and \begin{align} & \mathbb{P}\left[\left. Y_{0}^{*} = 0, S_{0} = 1\right\vert X = x\right] \nonumber \\ & > \max \left\lbrace \mathbb{P}\left[\left. Y_{1}^{*} = 0, S_{1} = 1 \right\vert X = x\right], \mathbb{P}\left[\left. Y_{1}^{*} = 1, S_{1} = 1 \right\vert X = x\right] \right\rbrace \end{align} for every value $x \in \mathcal{X}$, then $LB_{1}\left(x\right) > 0$ and $UB_{1}\left(x\right) < 1$.

Intuitively, Assumptions (ref)-(ref) have identifying power if the group who is informally employed when untreated is sufficiently large.

}

In practice, the bounds in Proposition (ref) may be {wide} even though they are sharp. To derive tighter bounds, researchers can add increasingly stronger assumptions. Even though the credibility of these assumptions depends on their empirical contexts, applied researchers frequently have some prior about the direction of the treatment effect. Using this prior, the researcher can impose the monotone treatment response condition.

Formally, combining Assumptions (ref)-(ref), we derive sharp bounds around $\theta^{OO}\left(x\right)$ as detailed in Proposition (ref). Its proof is in Appendix (ref).

propositionUnder Assumptions (ref)-(ref), the conditional probability of causation is partially identified for the always-observed subgroup, i.e., \begin{equation*} LB_{1}\left(x\right) \leq \theta^{OO}\left(x\right) \leq UB_{2}\left(x\right), \end{equation*} where $$UB_{2}\left(x\right) \coloneqq \min \left\lbrace \dfrac{B\left(x\right) \cdot \left[A\left(x\right)\right]^{-1} + C\left(x\right) - 1}{C\left(x\right)} , 1 \right\rbrace$$ for every value $x \in \mathcal{X}$. Moreover, these bounds are sharp.

{

Corollary (ref) describes when Assumption (ref) has additional identifying power, i.e., the identified set in Proposition (ref) is strictly smaller than the identified set in Proposition (ref). Its proof is in Appendix (ref).

corollaryIf Assumptions (ref)-(ref) hold, Inequality (ref) holds, and \begin{equation} \mathbb{P}\left[\left. Y_{0}^{*} = 1, Y_{1}^{*} = 1 \right\vert S_{0} = 1, S_{1} = 1, X = x\right] > 0 \end{equation} for every value $x \in \mathcal{X}$, then $LB_{1}\left(x\right) > 0$ and $UB_{2}\left(x\right) < UB_{1}\left(x\right) < 1$.

Note that the identifying power of Assumption (ref) is illustrated by a strictly smaller upper bound in Proposition (ref) in comparison with Proposition (ref). Intuitively, Assumption (ref) has additional identifying power if some always-employed individuals have a formal job regardless of their treatment status. }

To achieve even tighter bounds, researchers can impose the stochastic dominance condition. Formally, combining Assumptions (ref)-(ref), we derive sharp bounds around the conditional probability of causation within the always-observed sub-population as detailed in Proposition (ref). Its proof is in Appendix (ref).

propositionUnder Assumptions (ref)-(ref), the conditional probability of causation is partially identified for the always-observed subgroup, i.e., \begin{equation*} LB_{3}\left(x\right) \leq \theta^{OO}\left(x\right) \leq UB_{2}\left(x\right), \end{equation*} where $$LB_{3}\left(x\right) \coloneqq \max\left\lbrace \dfrac{B\left(x\right) + C\left(x\right) - 1}{C\left(x\right)} , 0 \right\rbrace$$ for every value $x \in \mathcal{X}$. Moreover, these bounds are sharp.

{

Corollary (ref) describes when Assumption (ref) has additional identifying power, i.e., the identified set in Proposition (ref) is strictly smaller than the identified set in Proposition (ref). Its proof is in Appendix (ref).

corollaryIf Assumptions (ref)-(ref) hold, Inequalities (ref) and (ref) hold, $\mathbb{P}\left[\left. S_{0} = 0, S_{1} = 1 \right\vert X = x\right] > 0$ and $\mathbb{P}\left[\left. Y_{0}^{*} = 0, Y_{1}^{*} = 0 \right\vert S_{1} = 1, X = x\right] > 0$ for every value $x \in \mathcal{X}$, then $LB_{3}\left(x\right) > LB_{1}\left(x\right) > 0$ and $UB_{2}\left(x\right) < UB_{1}\left(x\right) < 1$.

Note that the identifying power of Assumption (ref) is illustrated by a strictly larger lower bound in Proposition (ref) in comparison with Proposition (ref). Intuitively, Assumption (ref) has additional identifying power if there are employed-only-when-treated individuals and if some employed-when-treated individuals never have a formal job.

}

Second, we identify the distribution of the covariates for the always-observed sub-population with a zero untreated potential outcome, $ \left. X \right\vert Y_{0}^{*} = 0, S_{0} = 1, S_{1} = 1,$ in Lemma (ref). For ease of notation, we assume that all covariates $X$ are discrete, as in our empirical application. This lemma's proof is in Appendix (ref).

lemmaUnder Assumptions (ref)-(ref), the distribution of the covariates for the always-observed sub-population with a zero untreated potential outcome is point identified, i.e., \begin{align*} \omega\left(x\right) & \coloneqq \mathbb{P}\left[\left. X = x \right\vert Y_{0}^{*} = 0, S_{0} = 1, S_{1} = 1\right] \\ & = \dfrac{\mathbb{P}\left[\left. Y = 0, S = 1 \right\vert D = 0, X = x \right] \cdot \mathbb{P}\left[X = x\right]}{\sum_{x^{\prime} \in \mathcal{X}} \mathbb{P}\left[\left. Y = 0, S = 1 \right\vert D = 0, X = x^{\prime} \right] \cdot \mathbb{P}\left[X = x^{\prime}\right]} \end{align*} for every $x \in \mathcal{X}$.

Finally, we can combine Propositions (ref)-(ref) and Lemma (ref) to partially identify our target parameter $\theta^{OO}$ (Equation (ref)) as detailed in Corollary (ref).

corollaryThe probability of causation is partially identified for the always-observed subgroup, i.e., $$\sum_{x \in \mathcal{X}} LB_{1}\left(x\right) \cdot \omega\left(x\right) \leq \theta^{OO} \leq \sum_{x \in \mathcal{X}} UB_{1}\left(x\right) \cdot \omega\left(x\right)$$ under Assumptions (ref)-(ref), $$\sum_{x \in \mathcal{X}} LB_{1}\left(x\right) \cdot \omega\left(x\right) \leq \theta^{OO} \leq \sum_{x \in \mathcal{X}} UB_{2}\left(x\right) \cdot \omega\left(x\right)$$ under Assumptions (ref)-(ref), and $$\sum_{x \in \mathcal{X}} LB_{3}\left(x\right) \cdot \omega\left(x\right) \leq \theta^{OO} \leq \sum_{x \in \mathcal{X}} UB_{2}\left(x\right) \cdot \omega\left(x\right)$$ under Assumptions (ref)-(ref).

Furthermore, in Appendix (ref), we illustrate this section's results with a numerical example that captures the intuition behind them.

}

Estimation and Inference

This section is divided in two parts. In the first part, we discuss how to estimate the bounds proposed in Section (ref). In the second part, we propose estimators for 90%-confidence regions that contain the identified sets described in Corollary (ref).

Importantly, in Section (ref), we do not discuss how to conduct inference around the target parameter in Equation (ref). Our choice of conducting inference around the target parameter's identified region may have a cost in terms of statistical power and may explain our null results in Section (ref). However, our chosen procedure has the advantage of being simpler and more intuitive.

Estimation

{ In this section, we propose estimators for the bounds described in Propositions (ref)-(ref) and Corollary (ref), and the weights in Lemma (ref). To do so, we need to estimate $\mathbb{P}\left[\left. S = 1\right\vert D = d, X = x \right]$, $\mathbb{P}\left[\left. Y = y\right\vert S = 1, D = d, X = x \right]$, $\mathbb{P}\left[\left. Y = 0, S = 1\right\vert D = 0, X = x \right]$ and $\mathbb{P}\left[X = x\right]$ for any $y \in \left\lbrace 0,1 \right\rbrace$, $d \in \left\lbrace 0,1 \right\rbrace$ and $x \in \mathcal{X}$.

We estimate these objects parametrically using maximum likelihood estimators. {To simplify our notation, we follow our empirical application and impose that the covariates $X$ are stratum (course-city pair) fixed effects (417 strata). Moreover, to ensure that the first part of Assumption (ref) holds, we delete non-oversubscribed strata (327 strata remain). Finally, to estimate $B\left(x\right)$ and $C\left(x\right)$, we delete strata without post-treatment employed individuals (246 strata remain).}

Let $\lambda\left(\cdot\right)$ be a link function, such as the logistic link function or the normal link function. Our parametric regression models are given by:

enumerate$\mathbb{P}\left[\left. S = 1\right\vert D = d, X = x \right] = \lambda\left(\alpha_{0} + \alpha_{1} \cdot d + \alpha_{x}\right)$, • $\mathbb{P}\left[\left. Y = 1\right\vert S = 1, D = d, X = x \right] = \lambda\left(\beta_{0} + \beta_{1} \cdot d + \beta_{x}\right)$, where we only use the employed subsample to estimate $\beta_{0}$, $\beta_{1}$ and $\beta_{x}$, and • $\mathbb{P}\left[\left. W = 1 \right\vert D = d, X = x \right] = \lambda\left(\gamma_{0} + \gamma_{1} \cdot d + \gamma_{x}\right)$, where $W \coloneqq \mathbf{1}\left\lbrace Y = 0, S = 1 \right\rbrace$.

Denoting our coefficients' estimators with the hat notation, the bounds in Propositions (ref)-(ref) can be estimated using the following objects:

enumerate$\hat{A}\left(x\right) = \dfrac{\lambda\left(\hat{\alpha}_{0} + \hat{\alpha}_{x}\right)}{\lambda\left(\hat{\alpha}_{0} + \hat{\alpha}_{1} + \hat{\alpha}_{x}\right)}$, • $\hat{B}\left(x\right) = \lambda\left(\hat{\beta}_{0} + \hat{\beta}_{1} + \hat{\beta}_{x}\right)$, and • $\hat{C}\left(x\right) = 1 - \lambda\left(\hat{\beta}_{0} + \hat{\beta}_{x}\right)$.

Furthermore, the weights in Lemma (ref) can be estimated by $$\hat{\omega}\left(x\right) = \dfrac{\lambda\left(\hat{\gamma}_{0} + \hat{\gamma}_{x}\right) \cdot \sum_{i = 1}^{N} \mathbf{1}\left\lbrace X_{i} = x \right\rbrace}{\sum_{x^{\prime} \in \mathcal{X}} \lambda\left(\hat{\gamma}_{0} + \hat{\gamma}_{x^{\prime}}\right) \cdot \sum_{i = 1}^{N} \mathbf{1}\left\lbrace X_{i} = x^{\prime} \right\rbrace}.$$

In Appendix (ref), we present the full formulas of our estimators for the bounds in Propositions (ref)-(ref) and Corollary (ref).

We must also test the restrictions in Proposition (ref). The first restriction is equivalent to testing the null hypothesis that $\alpha_{1} \geq 0$. The second restriction is equivalent to testing the null hypothesis that $\delta_{1} \geq 0$ in the following model: $$\mathbb{P}\left[\left. Y = 1\right\vert D = d, X = x \right] = \lambda\left(\delta_{0} + \delta_{1} \cdot d + \delta_{x}\right).$$ To control size appropriately, we use a Bonferroni correction for the p-values of both tests. {When using either a Probit Model or a Logit Model for the link function $\lambda\left(\cdot\right)$, we find Bonferroni corrected p-values equal to 1.00 for $H_{0}: \alpha_{1} \geq 0$ and $H_{0}: \delta_{1} \geq 0$.} These results suggest, based on Proposition (ref), that our identifying assumptions are not refuted.

}

Inference

In this section, we suggest possible estimators for 90%-confidence regions that contain the identified sets described in Corollary (ref). To fix ideas, we will focus on the bounds under Assumptions (ref)-(ref), but all the ideas here extend to the bounds under our other sets of assumptions.

Imposing Assumptions (ref)-(ref), we have that $\theta^{OO} \in \left[\sum_{x \in \mathcal{X}} LB_{3}\left(x\right) \cdot \omega\left(x\right), \sum_{x \in \mathcal{X}} UB_{2}\left(x\right) \cdot \omega\left(x\right) \right]$ and $\theta^{OO}\left(x\right) \in \left[ LB_{3}\left(x\right), UB_{2}\left(x\right) \right]$ for any $x \in \mathcal{X}$. We want to find random sets $\widehat{Q}_{N}\left(x\right)$ and $\widehat{R}_{N}$ such that

equation[equation omitted — 183 chars of source]

for any $x \in \mathcal{X}$ and

equation[equation omitted — 270 chars of source]

where $N$ is the sample size, $p_{Q} \in \left(\sfrac{1}{2}, 1\right)$ and $p = 0.9$.

The $p_{Q}$-confidence region $\widehat{Q}_{N}\left(x\right)$ is given by the precision-corrected estimator proposed by Chernozhukov2013. The $p$-confidence region $\widehat{R}_{N}$ is given by a set that combines the precision-corrected estimator proposed by Chernozhukov2013 with a Bonferroni-style correction.

For any $x \in \mathcal{X}$, let $\widehat{Q}_{N}\left(x\right) \coloneqq \left[ \widehat{LB}_{3,N}^{CLR}\left(x,\sfrac{\left(1 + p_{Q}\right)}{2}\right), \widehat{UB}^{CLR}_{2,N}\left(x,\sfrac{\left(1 + p_{Q}\right)}{2}\right) \right]$, where $\widehat{LB}_{3,N}^{CLR}\left(x,\sfrac{\left(1 + p_{Q}\right)}{2}\right)$ and $\widehat{UB}^{CLR}_{2,N}\left(x,\sfrac{\left(1 + p_{Q}\right)}{2}\right)$ are the precision-corrected estimators proposed by Chernozhukov2013 for the bounds $LB_{3}\left(x\right)$ and $UB_{2}\left(x\right)$. These estimators satisfy $$\mathbb{P}\left[ \widehat{LB}_{3,N}^{CLR}\left(x,\sfrac{\left(1 + p_{Q}\right)}{2}\right) \leq LB_{3}\left(x\right) \right] \geq \dfrac{1 + p_{Q}}{2} - o\left(1\right)$$ and $$\mathbb{P}\left[ UB_{2}\left(x\right) \leq \widehat{UB}^{CLR}_{2,N}\left(x,\sfrac{\left(1 + p_{Q}\right)}{2}\right) \right] \geq \dfrac{1 + p_{Q}}{2} - o\left(1\right),$$ implying that Equation (ref) holds. We formally prove this result in Appendix (ref).

Now, we define

equation[equation omitted — 341 chars of source]

This choice of estimator for a feasible $p-$confidence region is inspired by the unfeasible set given by

equation[equation omitted — 337 chars of source]

which assumes we know the true population weights ${\omega}\left(\cdot\right)$ instead of using the estimated weights $\hat{\omega}\left(\cdot\right)$ proposed in Section (ref).

In Appendix (ref), we show that the unfeasible set ${R}_{N}$ is a valid $p$-confidence region around the identified set $\left[\sum_{x \in \mathcal{X}} LB_{3}\left(x\right) \cdot \omega\left(x\right), \sum_{x \in \mathcal{X}} UB_{2}\left(x\right) \cdot \omega\left(x\right) \right]$. In particular, a Bonferroni-style correction implies that $p = 90\%$ if $p_{Q} = 99.96\%$. Additionally, if our goal was to derive half-median unbiased estimators, we could use $p_{Q} = 99.8\%$.

Appendix (ref) also contain details on how to implement the precision-corrected estimators proposed by Chernozhukov2013. This appendix relies heavily on the work done by Flores2013, who intuitively explain the method proposed by Chernozhukov2013.

As a caveat, we highlight that we do not show that the feasible set $\widehat{R}_{N}$ is a valid $p$-confidence region around the identified set. We believe that taking into consideration the uncertainty behind the estimation of $\omega\left(\cdot\right)$ is beyond the scope of this paper and emphasize that a rigorous treatment of feasible inference around the identified set $\left[\sum_{x \in \mathcal{X}} LB_{3}\left(x\right) \cdot \omega\left(x\right), \sum_{x \in \mathcal{X}} UB_{2}\left(x\right) \cdot \omega\left(x\right) \right]$ is an interesting area for future work.

Despite the absence of a formal proof, Appendix (ref) describes a Monte Carlo Simulation that illustrates the finite sample properties of the feasible inference procedure proposed in this section. We find that, in our simulated data-generating process, $\widehat{R}_{N}$ covers the identified set more frequently than its nominal confidence level of 90%. This result suggests that using the feasible set $\widehat{R}_{N}$ in place of the unfeasible set ${R}_{N}$ may work appropriately, suggesting the importance of developing formal results related to this inference procedure in the future.

Empirical Application: Transition into Formality in the Jóvenes in Acción Training Program

{

Our empirical application uses experimental data on a large job training program called Jóvenes en Acción, implemented in Colombia's seven largest cities between 2002 and 2005. The program’s main goals were to increase the labor market attachment and the quality of jobs that disadvantaged young individuals (between 18 and 25 years old) held. To this end, Jóvenes en Acción combined three main components: (i) three months of classroom training on occupational-specific skills in private training centers, with an additional focus on building “soft” skills, such as proactive behavior, resourcefulness, openness to feedback and teamwork; (ii) three months of on-the-job training provided by legally registered companies in the form of an unpaid internship; (iii) elaboration of a project of life, orienting youth towards a positive visualization of their abilities and work perspectives.

An additional key feature of Jóvenes en Acción was that the payment structure of training centers incentivized them to help their trainees complete the program and secure jobs after the program. Specifically, training centers received a large fraction of their payment conditional on the student completing the course and obtaining an internship. More importantly, they were awarded an additional bonus if the firm hired the trainee on a formal contract. This tight incentive structure and curricula encompassing a large set of potentially productive skills allows one to consider Jóvenes en Acción as an intensive program with high potential to improve the employability and the quality of jobs held by its beneficiaries.

The short-run experimental effects of the program have been described in Attanasio2011 and point to improvements along the employability and job quality margins. We follow Attanasio2011 and Attanasio2017 in analyzing effects separately by gender, focusing on women since there was a significant differential sample selection into employment in this sub-sample in the short run. Specifically, women selected to participate in Jóvenes en Acción were 6.1 percentage points (or 9.6%) more likely to be employed between 13 and 15 months after exiting the program according to Attanasio2011. Moreover, they also document that women selected to participate in Jóvenes en Acción were 7.1 percentage points (or 36%) more likely to be formally employed approximately one year after exiting the program.

Differently from Attanasio2011, we are interested in learning more about the effects of Jóvenes en Acción on job quality after accounting for sample selection. Distinguishing between effects on the job quality margin that would occur irrespective of the movements towards employment is important to understand better whether the program led to more favorable labor market outcomes. We focus on formality, which, in most developing countries, is strongly associated with employer compliance with labor market statutes (minimum wage and firing regulations), higher productivity and pay, and social security contributions Meghir2015wages, Attanasio2017.

We use our partial identification results to learn about the share of women who became formal because they were selected to participate in the program. As explained in Section (ref), our target parameter is the probability of causation for the latent group that would be employed regardless of treatment assignment. We compute bounds around this probability of causation by considering assignment to the program as the treatment indicator, employment (either in the formal or the informal sector) as the selection indicator, and an indicator that equals one if the person has a formal job and zero if the person has an informal job as our variable of interest.

We start by providing descriptive statistics on the size of our latent groups of interest, i.e., the share of the female population who would be employed regardless of being assigned to the Jóvenes en Acción training program and, within this group, the share of women who would have an informal job if they were assigned to the control group. Since both objects are point-identified under Assumptions (ref)-(ref), we focus on our first set of assumptions when estimating them. We find that 71.9% of the women are always-employed using either a Probit or Logit model as the link function $\lambda\left(\cdot\right)$. Within this subgroup, we also estimate the probability of having an informal job when untreated as 49.7% using either a Probit or Logit model as the link function $\lambda\left(\cdot\right)$. Thus, our latent group of interest represents a non-negligible share (approximately 35.7%) of the program's pool of potential female participants.

Our main results are presented in Figure (ref). The intervals in this figure represent estimated lower and upper bounds on the probability of causation for the always-employed women (Corollary (ref)) using data from the job training program Jóvenes en Acción and the estimator proposed in Section (ref). The black estimated intervals are based on Assumptions (ref)-(ref). The dark gray estimated intervals are based on Assumptions (ref)-(ref). The light gray estimated intervals are based on Assumptions (ref)-(ref). Subfigure (ref) uses a Probit Model as the link function $\lambda\left(\cdot\right)$ while Subfigure (ref) uses a Logit Model. The dots represent the lower and upper bounds of 90%-confidence regions around the identified sets. These confidence regions are based on the inferential method proposed by Chernozhukov2013 and explained in Section (ref). Since the bounds with a Probit or a Logit link function are very similar, we focus our discussion on the former.

figure[figure omitted — 1,720 chars of source]

We start by presenting the bounds on the probability of causation for the always-employed women (Corollary (ref)) under Assumptions (ref)-(ref). In this case, we only impose, beyond the random assignment and positive mass assumptions, that participation in the program does not deter employment (monotone sample selection).

Assumption (ref) is plausible in the Jóvenes in Acción context. First, the training program's focus on “soft skills” is likely to boost the workers' performance in job interviews, improving their employment prospects. Second, as discussed in Section (ref), the test proposed in Lemma (ref) does not reject the null hypothesis that is implied by Assumptions (ref)-(ref).

We find that the estimated bounds are very wide. They imply that our estimates are consistent with a large variety of values for the probability of causation for the always-employed women ($[6.6\%, 41.8\%]$). It implies that the Jóvenes in Acción training program formalized, at least, 6.6% of the women who are always-employed and would have an informal job if untreated. Moreover, the 90%-confidence region includes the zero, implying that we cannot reject the null hypothesis that our target parameter's lower bound is equal to zero.

To tighten the estimated intervals, we now discuss the bounds obtained by additionally imposing Assumption (ref). In this case, we assume that participation in the program can only move agents from informal jobs to formal ones.

Assumption (ref) is plausible in the Jóvenes in Acción context. First, the program's occupational-specific classes and on-the-job training are likely to increase the workers' productivity, helping them find better (i.e., formal) jobs. Second, training centers are incentivized to help their trainees secure a formal job in the firm where they interned. Furthermore, as discussed in Section (ref), the test proposed in Proposition (ref) does not reject the null hypotheses that are implied by Assumptions (ref)-(ref).

We find that imposing a monotone treatment response decreases the upper bound substantially. The dark gray interval in Figure (ref) suggests that Jóvenes en Acción formalized at most 13.4% of the women who are always-employed and would have an informal job if untreated. Furthermore, the upper bound of the 90%-confidence region decreases to 29.4%.

To further tighten the estimated intervals, we discuss the bounds obtained by additionally imposing Assumption (ref). In this case, we assume that the always-employed sub-population has higher potential formality when treated than the employed-only-when-treated sub-population. This assumption is plausible because individuals with better employment status are more likely to be more skillful, increasing their chances of having a better (i.e., formal) job.

We find that imposing this stochastic dominance assumption increases the lower bound. The light gray interval in Figure (ref) suggests that Jóvenes en Acción formalized at least 10.2% of the women who are always-employed and would have an informal job if untreated. Importantly, the 90%-confidence region includes zero, implying that we cannot reject the null hypothesis that the lower bound of the probability of causation for the always-employed women is zero.

Finally, in Appendix (ref), we present additional results focusing on the heterogeneity generated by different course-city pairs.

}

Conclusion

This paper partially identifies the probability of causation for the always-observed subgroup when sample selection occurs. This parameter is important for researchers aiming to describe treatment effects in a way that is relevant to policy-makers. Intuitively, it describes the share of the population induced by the treatment to switch from a negative to a positive state. We derive sharp bounds around this parameter under three increasingly restrictive sets of assumptions.

To illustrate the usefulness of our partial identification strategy, we use experimental data from the Colombian job training program Jóvenes en Acción. { Contradicting the positive effects on the share of women employed in the formal labor market Attanasio2011, we find that incorporating selection and bounding the probability of causation leads to a pessimistic view of the program’s impacts. More precisely, we find that at most 13.4% of the always-employed women switched their formality status because they were assigned to the Jóvenes en Acción training program. Moreover, even our tightest 90%-confidence region includes zero, implying that we cannot reject the null hypothesis that our lower bound is equal to zero.}

{Beyond the analysis of job training programs, our partial identification strategy can be useful for researchers interested in assessing the impacts of interventions in the presence of sample selection. For example, when analyzing the effects of a political campaign {DellaVigna2007,DellaVigna2010}, the researcher may be interested in identifying the share of the population who supports policy A when treated, given that they would support policy B if untreated. In this case, the researcher only observes the agents' opinions if they reply to a survey. This double identification challenge also arises when researchers consider the effects of health interventions on health quality {Health2004} if agents may pass away, or the effects of educational interventions on learning {Angrist2006,Chetty2011,Dobbie2015} if there is selection into test-taking.}

Acknowledgment

We thank Donald Andrews, Xiaohong Chen, Fernanda Estevan, Bruno Ferman, Sergio Firpo, John Eric Humphries, Helena Laneuville, Guilherme Lichand, Yusuke Narita, Cormac O'Dea, Giovanni Di Pietra, Rudi Rocha, Edward Vytlacil, Siu Yuat Wong, and seminar participants at Yale University, EPGE Brazilian School of Economics and Finance, Sao Paulo School of Economics, Federal University of Paraiba and State University of New York (Albany) for helpful suggestions. We thank Joana Getlinger for providing excellent research assistance.