Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
130,518 characters · 19 sections · 88 citation commands
Difference-in-Differences with Sample Selection
\baselineskip 20pt \setlength\abovedisplayskip{5pt} \setlength\belowdisplayskip{5pt}
\begingroup \footnote{$^a$Department of Econometrics and Business Statistics, Monash University, Australia. $^b$ Central Bank of Sri Lanka. $^c$IZA, Germany. Emails: [email removed], [email removed], [email removed], [email removed]. We thank Alyssa Carlson, \`{A}ureo de Paula, Desire Kedagni, Pedro Sant'Anna, Rami Tabri, Quentin Brummet, Valentin Verdier, Vitor Possebom, Martin Huber, Giovanni Mellace, Philip Heiler, four anonymous referees, and participants at seminars at the University of Melbourne, 2025 IAAE Annual Meetings, and the 2025 Econometric Society World Congress for their useful suggestions and comments.} \addtocounter{footnote}{-1} \endgroup
{\it Keywords: Sample selection, Partial identification, Difference-in-differences, Panel data, Heterogeneous treatment effects}
{\em JEL Classifications: C14, C31, C33}
Nonrandom sample selection is a pervasive challenge in empirical research that can compromise the validity of usual causal inference approaches. In many studies, the outcome of interest may only be observed for a non-random subset of the population due to issues such as attrition, survey non-response, and measurement error.\footnote{This is frequently observed in policy evaluation studies that employ a panel survey of individuals before and after a program is implemented (such as cash transfer, poverty alleviation, or a training subsidy program) to evaluate its impact holzer1993training,bobonis2011impact,asadullah2016evaluating. Frequently, the follow-up survey will face the problem of non-ignorable attrition.} This can pose a significant problem for difference-in-differences (DiD) methods whose popularity in empirical research has grown over time goldsmith2024tracking. In this paper, we address the challenges posed by endogenous sample selection in a DiD setting and propose a partial identification strategy for average treatment effects on the treated based on a latent subpopulation structure under alternative sets of identifying assumptions.
We first demonstrate that na\"{\i}vely applying DiD to units observed in both pre- and post-treatment periods without accounting for sample selection does not identify a meaningful causal parameter; either the overall ATT or an adequately weighted average of treatment effects. When the goal is to recover the overall ATT, the canonical DiD estimand will generally be biased unless one imposes restrictive assumptions on untreated outcome trends and treatment effect heterogeneity that are unlikely to be true when sample selection is endogenous. Interestingly, even if the selection mechanism is independent of treatment assignment, the bias in na\"{\i}ve DiD does not disappear unless selection is exogenous to the outcome of interest.\footnote{In Section (ref), we demonstrate this through a simulation where the na\"{\i}ve DiD estimate exhibits an upward bias, even when the selection and treatment assignment mechanisms are independent.}
Our first contribution is to propose a partial identification strategy for the average treatment effect on the treated (ATT) for individuals belonging to the latent group whose outcome would be observed regardless of their treatment state ($\tau_{OOO}$, with OOO referring to being “always-observed” in the pre-treatment period and in both the untreated and treated post-treatment counterfactual states, respectively) under different assumptions on the sample selection mechanism. The ATT for the always-observed latent group is of interest both as a component of the overall ATT and on its own. As discussed in bartalotti2023identifying, $\tau_{OOO}$ can be seen as a measure of the effect of treatment on the intensive margin. Substantively, it captures the effect of the treatment for a stable subpopulation that can be identified based on data observed in both time periods.\footnote{For example, the OOO group could refer to (a) workers with stable labor force attachment in the context of a training program; (b) employees who would stay with a firm irrespective of whether they are offered a work-from-home option, (c) patients who would continue to receive medical care or (d) survive throughout the study period.} Our approach combines the trimming procedure of lee2009 to address the identification challenges posed by endogenous sample selection and treatment assignment.
Identification of $\tau_{OOO}$ relies on parallel trends in outcomes (PTO) between the treated and untreated units within the same latent group, and a no-anticipation assumption. The trimming procedure requires knowledge of the latent groups' proportions imai2008sharp, lee2009, semenova2025generalized, which we acquire by considering combinations of two assumptions that govern the relationship between selection and treatment assignment: (i) “ignorability of treatment in potential selection” (IS) and (ii) “monotonicity of selection” (MS) in treatment. IS requires that, conditional on being observed in the pre-treatment period, the probability of being observed post-treatment in each counterfactual state (treated or untreated) is independent of the treatment received. In turn, “monotonicity of selection” (MS) requires that treatment has an increasing effect on the probability of selection for all units in the post-treatment period.\footnote{MS is commonly used in the sample selection literature chen2015bounds, huber2015sharp and is similar to the LATE monotonicity condition imbens1994identification.}
We derive alternative bounds for the ATT of the OOO group under different sets of assumptions, both with and without imposing MS. Without MS, our first result establishes partial identification of the proportions of the “always-observed” latent group among the treated (untreated) under no anticipation in selection and IS for potential selection if untreated (treated). Then, $\tau_{OOO}$ can be partially identified if IS holds for both counterfactual treatment states, producing bounds that are general and allow flexible post-treatment selection patterns, including cases in which treatment induces individuals to enter or exit the sample. Our second result tightens the $\tau_{OOO}$ bounds by imposing MS which allows us to relax the IS assumption for one of the two counterfactual treatment states\footnote{Under positive (negative) MS, we only constrain potential selection in the untreated (treated) state, leaving selection in the treated (untreated) state unrestricted.} and point-identifies the latent group proportions. Following lee2009, we establish the sharpness of both these bounds.
The second contribution is to extend the partial identification results for $\tau_{OOO}$ to include pre-treatment covariates, whereby we relax the unconditional PTO and IS assumptions to their conditional analogues. These two assumptions, along with alternative MS assumptions, allow us to identify the ATT for the always-observed within each covariate subpopulation. The resulting conditional bounds are then aggregated to obtain the identified set for the unconditional ATT for the OOO group. Importantly, we also consider the relaxed monotonicity framework of semenova2025generalized, which permits groups with different observed characteristics to exhibit different monotonicity directions. This provides an intermediate case between the two baseline scenarios of no monotonicity and global monotonicity (discussed in Section (ref)). Formal results and a detailed discussion are deferred to Appendix (ref).
The third contribution of this paper is to extend the bounding approach to identify the ATT for additional latent groups on whom the researcher has limited information compared to the “always-observed” type. Similar to the OOO group, these latent subpopulations are defined based on their observed selection status in the pre-treatment period and their counterfactual selection behavior in the post-treatment period under both treated and untreated states. We combine basic support restrictions and cross-group mean dominance assumptions, that impose economically intuitive rankings on potential outcome means across latent types, to obtain informative identified sets for the group-specific ATTs. These parameters are policy-relevant in many empirical settings and can reveal meaningful treatment effect heterogeneity across different selection types. For example, in evaluating job training programs, policymakers may care about effects for NOO (units that were unemployed before treatment but employed post-treatment whether or not they receive training) or NNO (units who were unemployed before treatment but will be employed post-treatment if given training but not otherwise). For workplace policy evaluations (such as work-from-home (WFH) policies), firms may be interested in the effects of WFH for the ONO group (employees who are observed prior to implementation of the policy and will leave the company if WFH is not provided but would stay if it's offered). The shares of these latent groups in the population are identified based on MS and a strengthening of IS to a joint independence assumption that conditions on pre-treatment selection.
Finally, we show that the overall ATT can be expressed as a weighted average of latent-group-specific ATTs. We obtain bounds on this parameter by combining the identified sets for each group along with appropriate population weights. This delivers partial identification for the overall population ATT, which is the typical estimand in DiD studies.
We illustrate our approach with two empirical applications. The first evaluates the effect of the National Supported Work training program on the Aid to Families with Dependent Children sample of women lalonde1986evaluating using the dataset from calonico2017women. Here, we consider the sample selection problem arising from unemployment. For the second application, we evaluate the effect of a WFH policy on employee performance, considering the sample selection problem arising from employee attrition bloom2015does.
While the approach developed here focuses on a two-period panel, we also explore extensions to repeated cross-sections sant2026difference,abadie2005semiparametric,finkelstein2002effect, meyer1990workers and staggered adoption DiD settings with multi-period panels callaway2021difference. With repeated cross sections, since observations cannot be tracked across time, it becomes impossible to distinguish between individuals observed in both pre- and post-treatment periods from those observed in only one period. We develop identification results for the always-observed latent group by relying on the assumption of no compositional changes sant2026difference. Additionally, we discuss how the current identification argument for the always-observed group can be adapted to multi-period staggered adoption settings by considering $2\times 2$ pre- and post-treatment comparisons for cohorts first treated in a specific period. These extensions are provided in appendices (ref) and (ref), respectively.
This paper contributes to both the DiD and the sample selection literatures in causal inference. In panel data settings, the traditional sample selection literature has primarily focused on parametric or semiparametric models to achieve point identification of treatment effects wooldridge1995selection,kyriazidou1997estimation,rochina1999new,semykina2010estimating. lechner2016difference study the implications of panel non-response in the outcome on the parallel trends assumption by comparing ordinary least squares and fixed effects estimates through simulations and applications.\footnote{They conclude that deviation of ordinary least squares and fixed effects estimation indicates nonignorable attrition.}
Our paper directly relates to a more recent nonparametric, instrument-free strand that pursues partial identification of treatment effects using principal stratification frangakis2002principal. Within this framework, zhang2003estimation, zhang2008evaluating, lee2009, and chen2015bounds derive bounds for the average treatment effect among always-observed units. honore2020selection build on the trimming logic of lee2009 and obtain tighter bounds by imposing additional structure on the selection model. In contrast, we follow lee2009 in targeting treatment effects for latent principal strata and extend this framework to a DiD setting. Subsequent work has explored extensions of this framework along other directions. bartalotti2023identifying extend this approach to marginal treatment effects for the always-observed group, while huber2015sharp derive bounds for additional latent subpopulations that go beyond the always-observed. All these papers focus on the cross-section setting. In contrast, we incorporate pre-treatment information about selection and outcomes via IS and PTO to additionally account for the endogeneity of treatment with respect to outcome and selection. Our paper also builds on semenova2025generalized, who incorporates pre-treatment covariates to relax the global monotonicity assumption used in lee2009 to a weaker conditional monotonicity assumption. viviens2025difference proposes an alternative relaxation of global monotonicity in settings where multiple sources of sample selection are observed and allows the direction of monotonicity to vary across sources.\footnote{viviens2025difference provides an example in which students' outcomes are missing because they (i) dropped out of college or (ii) graduated from college. Obtaining a scholarship (treatment) would affect each source of missingness in a different monotone direction.}
A closely related work is ghanem2024correcting, which studies attrition using the changes-in-changes (CiC) approach of athey2006identification and point identifies treatment effects. Their approach relies on the assumption that the distribution of unobservables affecting outcomes remains stable over time within each treatment-response subgroup and that potential outcomes are strictly monotone in unobserved heterogeneity. We impose weaker restrictions on outcome dynamics and instead leverage restrictions on the selection mechanism to deliver bounds that are robust to a wider class of outcome heterogeneity and more general forms of sample selection beyond just follow-up non-response. Our approach does not require monotonicity between outcomes and unobservables and also delivers bounds for group-specific treatment effects across latent selection types, in addition to bounds for the overall ATT. viviens2025difference also studies partial identification of the average and quantile treatment effects in a CiC framework that could be specialized to DiD under the more restrictive conditions needed for CiC. His approach imposes an absorbing-state restriction on missingness under which units not observed at baseline remain permanently unobserved. This restriction limits viviens2025difference's analysis to just four latent groups. His identification results for the always-observed subgroup, including the mixing proportions and bounds, coincide with ours for the monotonicity case. In contrast, we treat baseline non-observability as an integral part of the selection problem, allowing for a richer principal-strata structure that enables the study of additional latent groups. We also incorporate covariates and develop extensions beyond the canonical two-by-two framework.
Concurrently\footnote{We were only made aware after finishing our first draft paper that shin2024difference also independently studies the same setting.}, shin2024difference also studies missing outcomes in a DiD framework and partially identifies the ATT for the always-observed group using the trimming procedure by zhang2003estimation and lee2009. Similar to our approach, she considers identification with and without monotonicity of selection. Once Shin’s implicit conditioning on baseline observability is made explicit, her selection assumptions are equivalent to our IS assumption and the identified mixing proportions under each scenario are also equivalent. This implies that both approaches produce the same bounds for the always-observed group. The key difference lies in scope. Our framework allows for baseline non-observability, explicitly models all principal strata, and develops identification results for additional latent groups (ONO, NON, and NOO). In that sense, Shin (2024) can be viewed as a special case of our more general framework. We additionally extend identification for the always-observed group to settings with covariates and relaxed conditional monotonicity assumption in the spirit of semenova2025generalized. We further extend the framework to repeated cross-section and staggered adoption settings. These features are not considered in shin2024difference. Instead, the latter pursues point identification of the overall ATT using instrumental variables, whereas we bound the overall ATT without instruments.
The remainder of this article is organized as follows. Section (ref) introduces the principal stratification framework and identifying assumptions. Section (ref) discusses what na\"{\i}ve DiD identifies when sample selection is ignored. Section (ref) develops the identification strategy for the always-observed latent subgroup and derives ATT bounds for this group both with and without the MS assumption. It also presents the corresponding identification extension that incorporates covariates. Section (ref) presents results for three additional latent groups under outcome mean dominance assumptions, while Section (ref) discusses estimation and inference of the proposed bounds. Section (ref) presents simulation evidence under a range of data-generating processes. Section (ref) provides two empirical illustrations, and Section (ref) concludes. Additional identification results, extensions, and technical proofs are collected in the appendices.
Consider a setting with two time periods denoted by $t=0, 1$. Treatment, $D_{t}$, is available only at period $t=1$, such that $D_{0}=0$ for everyone and $D_1 \equiv D$. For each unit, let $Y_{t}^{\ast}(0)$ and $Y_{t}^{\ast}(1)$ be two continuous latent potential outcomes and $Y_{t}^{\ast} = Y_{t}^{\ast}(0)\cdot(1-D) + Y_{t}^{\ast}(1)\cdot D$ be the realized outcome, which is only observed for a non-random subset of the population. To formalize this, let $S_{t}(0)$ and $S_{t}(1)$ be two potential binary selection indicators such that
and the researcher observes the data vector $(Y_{t},S_{t},D)$ where
and $S_{t}\in {\{1,0}\}$ is the realized selection indicator, which equals one if the outcome for a unit is observed in period `$t$' and zero otherwise. For example, those with $S_{t}(0)=0$ and $S_{t}(1)=1$ are individuals for whom the outcome would not be observed if they are untreated but would be observed if treated.
Assumption (ref) formalizes the no-anticipation assumptions on selection and potential outcomes in the pre-treatment period. It states that there can be no anticipatory effects of the treatment assignment on sample selection or on the latent potential outcomes at baseline. This is plausible in situations where the treatment is not announced in advance, thereby discouraging individuals from basing their decision to be observed in the sample on whether they will receive the treatment in the future.
We consider the principal stratification framework introduced by frangakis2002principal to divide the population into latent subgroups based on the potential sample selection indicators in both periods. This results in sixteen groups, which can be reduced to the eight groups presented in Table (ref) since $S_{0}(0) = S_{0}(1)$ (Assumption (ref)).\footnote{Following the nomenclature used in lee2009, huber2015sharp and bartalotti2023identifying, we use “O” and “N” to denote observed and not observed, respectively.} Let $G$ denote the principal strata or latent group to which a unit belongs, with `g' denoting the group denomination.
Following lee2009, we define our target parameter to be the ATT for the subpopulation that is always observed, denoted by OOO, and indicates that selection equals one in all periods and under both counterfactual treatment states. Formally,
For $\tau_{OOO}$, we require a less restrictive version of parallel trends in outcomes that applies only to the OOO group.
Assumption (ref) states that the changes in the potential outcomes for always-observed (OOO) individuals, in the absence of treatment, would have been the same across the treatment and control groups. This assumption is weaker than requiring parallel trends for the full population of treated and control units, since that also includes other latent types beyond the OOO group. If additional pre-treatment periods are available, one can estimate a placebo DiD using only pre-treatment data. As shown in Appendix (ref), a non-zero estimate may reflect violations of parallel trends within the OOO or ONO group, cross-group trend differentials across latent groups, or any joint combination thereof. Therefore, such a test cannot isolate or falsify the plausibility of parallel trends for the always-observed group alone.\footnote{We thank an anonymous referee for suggesting this approach.}
While not required for the most general results in Section (ref), we consider a monotonicity assumption that is widely used in the literature lee2009,huber2015sharp,chen2015bounds,bartalotti2023identifying, which requires that the treatment affect sample selection in only one direction.
Without loss of generality, Assumption (ref) assumes that treatment increases the probability of selection or has a non-decreasing effect on sample selection for all individuals. Positive selection implies that there are no individuals whose outcome is observed only when untreated. For example, attending the job training program cannot decrease any individual's employment probability and, thus, does not decrease his/her chance of being observed. Assumption (ref) rules out the strata NON and OON, that is, individuals that would be observed in period one if untreated but not if treated.\footnote{Symmetric results can be derived under negative monotonicity, which are discussed in Appendix (ref).} We also assume that our setup has no spillovers and hidden treatment variations. In other words, we assume that the stable unit treatment value assumption holds.
Finally, we impose the following restriction on the relationship between the potential selection mechanism and treatment assignment.
Each part of Assumption (ref) imposes that the counterfactual proportion of individuals observed in the post-treatment period among those observed in the pre-treatment period be the same across the two treatment groups. As we discuss in Section (ref), under positive (negative) monotonicity, $\tau_{OOO}$ is partially identified under (ref)(a) ((ref)(b)), which restricts only the selection behavior in the untreated (treated) counterfactual.
Assumption 4 is plausible when the unobservables affecting sample selection in the post-treatment period are not systematically related to factors affecting the decision to select into treatment, once we condition on baseline observability. For example, in the WFH application, Assumption 4 is plausible when eligibility is determined based on factors that are not related to employees’ latent retention risk, conditional on being observed at baseline. It is less plausible when WFH is granted selectively based on managerial judgments of burnout risk, outside offers, or other unobserved predictors of attrition.
Note that this assumption only focuses on post-treatment selection behavior of individuals observed in the pre-treatment period ($S_0=1$) but not those who are unobserved in the pre-treatment period ($S_0=0$). Assumption (ref) can also be viewed as selection on lagged outcomes, and its connection to parallel-trends-type restrictions has been noted in ding2019bracketing. However, parallel trends for binary outcomes are known to be very restrictive (see marx2024parallel and ghanem2022selection).
Although Assumption (ref) is inherently untestable, since the relevant counterfactual selection rates across treatment arms (for example, $\mathbbm{P}(S_{1}(1)=1\mid D=0,S_{0}=1)$ and $\mathbbm{P}(S_{1}(0)=1\mid D=1,S_{0}=1)$) are not simultaneously observed, its plausibility can be partially assessed by examining whether pre-treatment selection rates are similar across treatment and control groups. Any significant pre-treatment differences may indicate that the assumption may be less tenable in the post-treatment period. Formally, the null hypothesis corresponding to Assumption 4(a) can be written as
Both probabilities in (ref) are identified from observed data under no-anticipation for selection in the pre-treatment period. A separate empirical test for Assumption 4(b) is not feasible because, under the same no anticipation in selection assumption,
As a result, testing equality of treated counterfactual selection probabilities across treatment groups produces the same restriction as in (ref). Hence, Assumptions 4(a) and 4(b) imply a testable implication that is identical in the pre-treatment period.
Our notion of sample selection is different from the idea of compositional changes that is discussed in hong2013measuring and sant2026difference. In those papers, selection arises from comparing independently drawn random samples from different time periods (pre- and post-treatment). This can lead to compositional changes, i.e., the joint distribution of covariates and treatment assignment can vary over time due to sample randomness, creating a “non-stationarity” problem that, if unaddressed, produces incorrect estimands for the ATT of interest due to the heterogeneity of the treatment effect. On the other hand, we consider a panel data setting where the same individuals are sampled in both periods. Therefore, sampling for $Y_{i0}$ and $Y_{i1}$ in our framework is stationary and have no compositional changes. Any compositional changes in the observed outcomes can only arise because of outcomes being non-randomly observed due to endogenous selection, which we explicitly model.
Before we delve into identification of $\tau_{OOO}$, it's useful to understand what a na\"{\i}ve DiD estimand that ignores the problem of sample selection identifies. We denote this as $\tau_{\textup{DiDs}}$. With panel data, $\tau_{\textup{DiDs}}$ compares average outcomes over time between the treated and control groups for individuals that are observed in both periods $(S_{0}=1, S_{1}=1)$. Lemma (ref) shows that $\tau_{\textup{DiDs}}$ is biased for the overall ATT, defined as $\tau \equiv \mathbbm{E}[Y_1^*(1) - Y_1^*(0) |D = 1]$.
The proof is presented in Appendix (ref). The first part of Lemma (ref) shows that $\tau_{\textup{DiDs}}$ is biased for the overall ATT. The bias components are given by \( \Delta_0 \) and \(\Delta_{\text{Het}}\). The term \( \Delta_0 \) represents differential trends in untreated potential outcomes among treated and untreated groups that are observed in both periods. The second bias term \( \Delta_{\text{Het}} \) captures treatment effect heterogeneity in the ATTs across the $S_0=s_0, S_1=s_1$ subpopulations. It becomes clear from this bias expression that assuming trends in untreated potential outcomes to be parallel between the treated and untreated i.e. \(\mathbbm{E}[Y_{1}^\ast(0)-Y_{0}^\ast(0)|D=1, S_0=1, S_1=1]=\mathbbm{E} [Y_{1}^\ast(0)-Y_{0}^\ast(0)|D=0, S_0=1, S_1=1] \), would eliminate $\Delta_0$. However, this would implicitly assume that selection is exogenous with respect to the untreated potential outcome trends, which is misguided in the current framework of endogenous sample selection. Even then, one would still be left with the bias term $\Delta_{\text{Het}}$.
The second part of Lemma (ref) establishes that the na\"{\i}ve DiD estimand could also be represented as a weighted average of ATTs for the always-observed (OOO) and the observed-only-when-treated (ONO) subgroups, with weights given by their respective proportions in the treated population, $p_{OOO1}$ and $1-p_{OOO1}$, plus additional terms representing selection bias. These terms include (i) the difference in trends in the absence of treatment between the treated and untreated ONO group and (ii) the cross-group differential in untreated potential outcomes trends across the different latent groups (i.e. ONO, OOO, OON) in the untreated subpopulation. If we assume parallel trends for the ONO group, the first bias term disappears, and we are left with heterogeneity in untreated potential outcome trends between the three latent groups. Naturally, this implies that if there is no heterogeneity in trends between these groups, i.e. $\mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=0, OOO] = \mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=0, ONO] = \mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=0, OON]$, then $\tau_{\textup{DiDs}}$ has a causal interpretation as a weighted average of group-specific ATTs.
The bias of the na\"{\i}ve DiD simplifies when one imposes positive monotonicity (Assumption (ref)). In this case, the OON latent stratum disappears from the $D=0$ group and $p_{OOO0}=1$. Just as before, the bias now arises from (i) the difference in trends in the absence of treatment between the treated and untreated ONO group ($\mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=1, ONO] - \mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=0, ONO]$) and (ii) the difference in untreated potential outcome trends between the OOO and ONO groups ($\mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=0, ONO] - \mathbbm{E}[Y_1^\ast(0)-Y_0^\ast(0)|D=0, OOO]$). This implies that even if we assume parallel trends for the ONO group, the DiD estimand will still be biased on account of differential trends between the untreated OOO and ONO groups. The extent of this bias depends on the share of ONO among those assigned treatment and the heterogeneity in average trends in the untreated potential outcomes between the ONO and OOO. A larger share of OOO will cause $\tau_{\textup{DiDs}}$ to be primarily influenced by $\tau_{OOO}$, thereby reducing the impact of the cross-group differences in trends, whereas a larger share of ONO will cause $\tau_{\textup{DiDs}}$ to be primarily influenced by $\tau_{ONO}$ while amplifying the cross-group difference in trends. Notice that even if the selection mechanism is completely independent of the treatment assignment process, na\"{\i}ve DiD would still be biased since selection might still be endogenous to the outcome of interest. If $S_1(0)$ is exogenous to the trends in the untreated potential outcome, then the bias in na\"{\i}ve DiD would disappear and it would give us $p_{OOO1}\cdot \tau_{OOO}+(1-p_{OOO1})\cdot \tau_{ONO}$.
It is important to note that a value of $p_{OOO1}$ close to one is informative but not sufficient for assessing the overall validity of the na\"{\i}ve DiD estimand in the presence of endogenous sample selection. This is because $p_{OOO1}\rightarrow 1$ indicates that the treated units observed in both periods are composed almost entirely of the OOO units, which eliminates selection bias for this group. As shown in Lemma (ref)(2b), under positive monotonicity, this is enough to ensure that the na\"{\i}ve DiD estimand recovers $\tau_{OOO}$, because the comparison group is also composed solely of OOO units. However, without monotonicity, the untreated group observed in both periods may still contain a mixture of OOO and OON units, and cross-group differences in untreated potential outcome trends can still generate bias even if $p_{OOO1} \approx 1$.
This section presents alternative conditions under which we can identify the parameter of interest, $\tau_{OOO}$. First, we discuss identifying the difference in the expected potential outcomes for latent groups. The identification problem arises from the fact that we do not observe the latent group membership directly since we either observe $S_{1}(0)$ or $S_{1}(1)$, but never both.
It is useful to note that $\tau_{OOO}$ could be identified by a hypothetical DiD estimand for members of the OOO latent group.
Although $\mathbbm{E}[Y_{1}^{\ast}(d)-Y_{0}^{\ast}(d)|D=d, OOO]$ cannot be generally point identified for $d=\{0,1\}$, it can be partially identified under different combinations of the monotonicity and selection mechanism assumptions. The plausibility of the assumptions required for partial identification depends on the empirical context. We approach this constructively by obtaining bounds for $\tau_{OOO}$ under less informative assumptions that might be valid on a larger range of empirical settings and then moving towards more restrictive assumptions that could be more informative for the parameter of interest. This allows a layered policy analysis manski2011, offering various estimates based on different assumptions so that the researcher can explore the information gathered about the parameter of interest by each restriction, as advocated in tamer2010partial.
Following the literature, we take advantage of the representation of observed subgroups of individuals as mixtures of latent groups, as shown in Table (ref) lee2009,chen2015bounds,huber2015sharp,bartalotti2023identifying. The relationship between observed and latent groups partially identifies $\mathbbm{E}[Y_{1}^{\ast}(d)-Y_{0}^{\ast}(d)|D=d, OOO]$, which we can use to recover $\tau_{OOO}$.
For instance, consider the group of treated individuals for whom the outcome is observed in both periods $(D=1, S_{0}=1, S_{1}=1)$. Table (ref) shows that their observed average outcome reflects a mixture of the potential outcomes for the OOO and ONO latent groups with mixing probabilities corresponding to their relative proportions. Then, {
} Now, consider the group of control individuals for whom the outcome is observed in both periods $(D=0, S_{0}=1, S_{1}=1)$. Their observed average outcome is a mixture of the potential outcomes for the OOO and OON latent groups, {
} We use these mixture representations to bound the expected change in potential outcomes within the always-observed subpopulation by looking at the observed outcomes’ distribution for treated individuals. Specifically, the lower bound for $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=1,OOO]$ is obtained by considering the worst-case scenario in which the OOO group comprises of individuals with the lowest values of $Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)$ among the subpopulation of treated individuals that has been observed in both periods. This corresponds to the left tail of mass $p_{OOO1}$ of the distribution of changes in outcomes between the pre- and post-treatment periods for treated individuals. The upper bound analogously assumes that the OOO group lies in the right tail of the same distribution, with the highest values of $Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)$. The intuition is similar to the trimming procedure suggested by lee2009 and others, where we assume that the OOO group corresponds to the set of treated individuals who either had the lowest observed changes in outcome between the pre- and post-treatment periods (yielding the lower bound) or experienced the highest observed changes (giving us the upper bound). Hence, $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=1,OOO]$ lies within the interval $[LB_{OOO1},UB_{OOO1}]$ where,
Similarly, the conditional distribution $Y_{1}-Y_{0}$ for the untreated individuals observed in both time periods can be trimmed to obtain the bounds for $\mathbbm{E}[Y_{1}^{\ast}(0)-Y_{0}^{\ast}(0)|D=0,OOO]$, which lies within the interval $[LB_{OOO0},UB_{OOO0}]$,
In equations (ref)-(ref) above, $F_{\Delta Y|dss'}^{-1}(.)$ is the quantile function of the distribution of the variable $\Delta Y \equiv Y_1-Y_0$ given $D=d,S_{0}=s,S_{1}=s'$.
Combining the bounds for $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=1,OOO]$ and $\mathbbm{E}[Y_{1}^{\ast}(0)-Y_{0}^{\ast}(0)|D=0,OOO]$ we find that the parameter of interest $\tau_{OOO}$ is in the interval
The fundamental aspect of identifying the target parameter is what can be learned about the weights, $p_{OOO0}$ and $p_{OOO1}$. Since we are interested in the always-observed group, a higher share of OOO among the treated individuals for which we have complete data implies that the observed sample provides more information about the changes in outcome for that group. In the extreme case, $p_{OOO1}\rightarrow 1$ and $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=1,OOO]$ is point identified as the observed sample reflects only the OOO type. In the opposite case, $p_{OOO1}\rightarrow 0$ and the observed sample would be uninformative about the always-observed group.
To this end, we consider alternative assumptions that impose different restrictions on the admissible values of the latent mixing proportions, $p_{OOO0}$ and $p_{OOO1}$, thereby yielding more information about $\tau_{OOO}$.
Initially, consider the case where the researcher is unwilling to assume monotonicity in selection (Assumption (ref)). We are interested in the unobserved share of always-observed individuals, $\mathbbm{P}[S_0(0)=1, S_1(0)=1, S_1(1)=1, D=d]$. The share of units observed in both periods among each treatment group is informative about the mixing proportions. For the treated group,
And for untreated observations,
The first equality in the equations above formalize the intuition that we can identify the marginal conditional proportions $\mathbbm{P}[S_{0}=1, S_{1}(d)=1| D=d]$ from observed data. It is useful to express {
} where $\mathbbm{P}[S_0=1|D=d]$ is directly observed in the data whereas $\mathbbm{P}[S_{1}(0)=1, S_{1}(1)=1| D=d, S_{0}=1]$ can be partially identified using Fr\'echet bounds imai2008sharp as follows: {
}
Note that $\mathbbm{P}[S_{1}(0)=1|D=0, S_{0}=1]$ and $\mathbbm{P}[S_{1}(1)=1|D=1, S_{0}=1]$ are directly identified from the observed data. We consider assumptions restricting the relationship between the selection mechanism and treatment assignment to identify their counterfactual counterparts, $\mathbbm{P}[S_{1}(0)=1|D=1, S_{0}=1]$ and $\mathbbm{P}[S_{1}(1)=1|D=0, S_{0}=1]$. Since the identification of $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=1,OOO]$ depends only on $p_{OOO1}$ and equivalently that of $\mathbbm{E}[Y_{1}^{\ast}(0)-Y_{0}^{\ast}(0)|D=0,OOO]$ solely on $p_{OOO0}$, we consider the assumptions for each term separately.
By combining Assumption (ref) and the information in equations ((ref)) and ((ref)), we can identify the missing counterfactual probabilities through the observed proportions for the treated and untreated groups, leading to Lemma (ref).
The proof of Lemma (ref) can be found in Appendix (ref). The restrictive nature of assuming both parts of Assumption (ref) becomes clear as the identified set for $\mathbbm{P}[S_{1}(0)=1, S_{1}(1)=1|D=d, S_{0}=1]$ is the same for both treated and control groups in that case, reflecting that the probability of being always-observed is independent of treatment under that assumption. This simplifies the identification of the mixing weights and is similar to scenarios in which the treatment is exogenous lee2009, or an instrument is available for selection and treatment bartalotti2023identifying. However, the weights will still differ between treated and untreated groups, which remains a challenge for identification in this setting that is not yet addressed in the previous literature.
Lemma (ref) can be used to obtain the range of possible values that $p_{OOO1}$ and $p_{OOO0}$ can take. For any value $v_d$ in the identified set for $\mathbbm{P}[S_{1}(0)=1, S_{1}(1)=1|D=d, S_{0}=1]$, the $p_{gd}$ associated with it is given by $p_{gd}(v_d)=\frac{v_d}{\mathbbm{P}[S_{1}=1| D=d, S_{0}=1]}$. As previously discussed, higher values for $p_{OOO1}$ and $p_{OOO0}$ indicate that a larger share of the observed - treated and untreated, respectively - population belongs to the always-observed latent groups, thus providing more information and tighter bounds for the target parameters. Hence, we only need to focus on the scenario that generates the wider bounds, that is, the smallest $p_{OOO1}(v_1)$ and $p_{OOO0}(v_0)$ bartalotti2023identifying. Since $v_d$ has a monotone relationship to the mixture weights, the relevant case is obtained at the lower bound of each of the identified sets for $\mathbbm{P}[S_{1}(0)=1, S_{1}(1)=1|D=d, S_{0}=1]$ described in Lemma (ref), which we call $v^{l}_{d}$ for $d=0,1$.
Evaluating equations ((ref))-((ref)) at the least favorable values for $p_{OOO1}(v_1)$ yields,
Similarly, for the bounds for $\mathbbm{E}[Y_{1}^{\ast}(0)-Y_{0}^{\ast}(0)|D=0,OOO]$ based on equations ((ref))-((ref)), evaluated at the smallest admissible value for $p_{OOO0}(v_0)$ yields,
Note that when $p_{OOOd}(v_d^l)=0$, the trimming regions become empty. This occurs when $\mathbbm{P}[S_{1}=1\mid D=0,S_{0}=1]+\mathbbm{P}[S_{1}=1\mid D=1,S_{0}=1]\le 1$ which implies that $F^{-1}_{\Delta Y| d11}(p_{OOOd}(v_d^l)) = -\infty$ and $F^{-1}_{\Delta Y| d11}(1 - p_{OOOd}(v_d^l)) = +\infty$ . Consequently, the trimmed expectations based on the empty sets $\{\Delta Y\le -\infty\}$ in (ref) and (ref) and $\{\Delta Y> +\infty\}$ in (ref) and (ref), will be undefined. To avoid this issue, we follow the literature and assume that the true mixing proportion is strictly positive semenova2025generalized,shin2024difference,lee2009.\footnote{If the estimated proportions are zero even though $p_{OOOd}>0$, this indicates that one cannot rule out the presence of OOO units in either the treated or control groups in the sample. In such cases, one would need additional information about the relative shares of OOO individuals within each group to achieve identification. A natural next step is to impose the monotonicity restriction in Assumption (ref). In this case, the proportions are point-identified, and the corresponding lower and upper bounds for $\tau_{OOO}$ are well-defined.} Combining the results above, we now propose partial identification of $\tau_{OOO}$.
Proof of Theorem (ref) can be found in Appendix (ref). The partial identification results in Theorem (ref) allow somewhat flexible patterns of potential selection into the sample. All latent group types are possible, and treatment is allowed to induce individuals to join or leave the sample in the post-treatment period since monotonicity in selection is not assumed. Nevertheless, to achieve identification, we imposed substantial restrictions on the relationship between the selection mechanism and treatment assignment through assumptions (ref)(a) and (ref)(b).
In specific applications, monotonicity in sample selection may be a plausible assumption. In the previous section, we saw that with just (ref)(a) or (ref)(b), we can partially identify $p_{OOO0}$ and $p_{OOO1}$, respectively. It is worth investigating how much leverage monotonicity alone has in terms of bounding the target parameter, $\tau_{OOO}$.
Proof can be found in Appendix (ref).
Positive monotonicity rules out the NON and OON strata. Hence, all untreated individuals observed in both periods are from the “always-observed” latent group and $p_{OOO0}=1$. Therefore, $\mathbbm{E}[Y_{1}^{\ast}(0)-Y_{0}^{\ast}(0)|D=0,OOO]$ is point identified by $\mathbbm{E}[Y_{1}-Y_{0}|D=0,S_{0}=1,S_{1}=1]$.
On the other treatment arm, individuals observed in both periods are still a mixture of OOO and ONO types. However, monotonicity guarantees that, $\mathbbm{P}[S_1(0)=1, S_1(1)=1|D=1, S_0=1] = \mathbbm{P}[S_1(0)=1|D=1, S_0=1]$ which means that we can focus on the values $p_{OOO1}$ can take over all possible $\mathbbm{P}[S_1(0)=1|D=1, S_0=1]$. Under positive monotonicity, the probability of selection in the treated counterfactual is always higher than the probability of selection in the untreated counterfactual, and $$0<\mathbbm{P}[S_1(0)=1|D=1, S_0=1]\leq \mathbbm{P}[S_1(1)=1|D=1, S_0=1]=\mathbbm{P}[S_1=1|D=1, S_0=1].$$ Even though monotonicity significantly constraints the possible values that $\mathbbm{P}[S_1(0)=1|D=1, S_0=1]$ can take, this information does not help us in learning about the proportion $p_{OOO1}=\frac{\mathbbm{P}[S_1(0)=1|D=1, S_0=1]}{\mathbbm{P}[S_1=1|D=1, S_0=1]}$, as it can still take any value in the unit interval.
To be able to partially identify $p_{OOO1}$ and $\tau_{OOO}$ we need to complement monotonicity with restrictions on $\mathbbm{P}[S_1(0)=1|D=1, S_0=1]$ that shrink its possible range to the interior of $[0, \mathbbm{P}[S_1=1|D=1, S_0=1]]$. A natural choice is to consider Assumption (ref)(a), which point identifies $p_{OOO1}$ by assuming $\mathbbm{P}[S_1(0)=1|D=1, S_0=1]=\mathbbm{P}[S_1(0)=1|D=0, S_0=1]$, as we show in Section (ref).
Alternatively, one can use a weaker version of this ignorability assumption, say, $\mathbbm{P}[S_{1}(0)=1|D=0, S_{0}=1]\leq \mathbbm{P}[S_{1}(0)=1|D=1, S_{0}=1]$. Intuitively, this condition requires that the probability of selection into the sample in the absence of treatment be at least as strong for the treated group as observed in the untreated group, allowing for “stronger trends” among the treated. This puts a floor on the lowest value possible for $p_{OOO1}\in\left[\frac{\mathbbm{P}[S_1=1|D=0, S_0=1]}{\mathbbm{P}[S_1=1|D=1, S_0=1]},1\right]$, which can then be used to construct identified sets for $\tau_{OOO}$ in a similar way to that described in Theorem (ref). However, the least favorable bounds in this case do not improve over those derived using Assumption (ref)(a).
As discussed in Section (ref), positive monotonicity rules out latent groups NON and OON, and $p_{OOO0}=1$ and $p_{OOO1} = \frac{\mathbbm{P}[S_1(0)=1|D=1, S_0=1]}{\mathbbm{P}[S_1=1|D=1, S_0=1]}$ (Lemma (ref)). Since $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(0)|D=0,OOO]$ is point identified in that case, there is no need for assumption (ref)(b).\footnote{In the case of negative monotonicity, latent groups NNO and ONO are ruled out. This results in $p_{OOO1}=1$, point identification for $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(0)|D=1,OOO]$. The identification results under this scenario are discussed in Appendix (ref).}
As suggested in the previous section, we can obtain point identification of $p_{OOO1}$ by combining positive monotonicity in selection and Assumption (ref)(a). Then, $p_{OOO0}=1$ and $p_{OOO1}=\frac{\mathbbm{P}[S_{1}=1|S_{0}=1,D=0]}{\mathbbm{P}[S_{1}=1|S_{0}=1,D=1]}$.
With point identified $p_{OOO0}$ and $p_{OOO1}$ we propose the partial identification of $\tau_{OOO}$.
The proof of Theorem (ref) is given in Appendix (ref). The identified set for $\tau_{OOO}$ under the assumptions of Theorem (ref) is more informative since, by construction, point identification of $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=0,OOO]$ tightens the overall bounds for $\tau_{OOO}$. Similarly, the proportion of the always-observed among the treated, $p_{OOO1}=\frac{\mathbbm{P}[S_{1}=1|S_{0}=1,D=0]}{\mathbbm{P}[S_{1}=1|S_{0}=1,D=1]}$, is the upper bound for $p_{OOO1}$ obtained under the conditions for Lemma (ref). Since higher shares of always-observed individuals imply more informative identified sets about that group, monotonicity leads to tighter bounds for $\tau_{OOO}$ as well.
Suppose that the researcher observes a vector of pre-treatment covariates $X$. We extend the partial identification results for the OOO group to explicitly incorporate covariates in the analysis. Specifically, we impose conditional versions of the parallel trends assumption, which is standard in the conditional DiD literature abadie2005semiparametric,sant2020doubly,caetano2024difference, along with a conditional ignorability assumption which assumes independence between treatment and selection conditional on pre-treatment characteristics and observability. These assumptions allow identification of the ATT within each covariate subpopulation, denoted as $\tau_{OOO}(X)$.
We establish results under two cases: (i) no monotonicity and (ii) a relaxed form of monotonicity following the framework in semenova2025generalized. In the latter case, we allow for covariate-specific monotonicity in selection patterns. This is accomplished by partitioning the covariate space into regions of positive monotonicity, negative monotonicity, and no monotonicity.
For each case of identification, the lower and upper bounds for the conditional ATT are derived using the same trimming logic as applied earlier to each subpopulation of $X$. Bounds for the unconditional ATT, $\tau_{OOO}$, are then obtained by aggregating conditional bounds over the distribution of covariates for the always-observed units in the treated group. The lower bound is then given by:
Similar results follow for the upper bound, where,
Detailed exposition of the partial identification arguments, along with formal statements of the results for both cases, is provided in Appendix (ref). In addition, we also provide moment-based representations of the bounds for the unconditional ATT for OOO in Appendix Section (ref). Estimation of the conditional bounds proceeds by discretizing the covariate space and then estimating the cell-specific bounds in the case of no monotonicity and classifying individuals into positive, negative, or no monotonicity regions before estimating the bounds in each region, for the case of relaxed monotonicity. This is discussed in Appendix Section (ref). Results for the two empirical applications incorporating covariates are presented in Appendix Section (ref). A joint test of monotonicity is also provided in Section (ref).
So far, the discussion has focused on identifying $\tau_{OOO}$, the ATT for the always-observed group, which often accounts for a large proportion of the population in many applications. However, in specific applications, policymakers may also be interested in identifying the treatment effect for other latent groups. For example, in evaluating the effects of a training program on earnings, policymakers are interested in the impacts on those unemployed before treatment (e.g., NOO and NNO latent groups). In other cases, the ONO latent group might be of interest. For instance, when considering the impact of working from home (WFH) on employee performance, the company's management may be interested in the effect on the productivity of employees who leave the company if WFH is not provided, but would stay if WFH is provided (i.e. ONO latent group).
This section studies the identification of $\tau_{g}$, the ATT for latent group $g$, for $g\in \{ONO, NOO ,$ $ NNO\}$.\footnote{As discussed in section (ref), positive MS rules out NON and OON latent groups. Furthermore, there is no information on ONN and NNN groups in either treatment arm in the post-treatment period. We therefore focus our attention on partially identifying the remaining groups.} Since less information is available for these groups relative to the OOO group, we introduce additional cross-group mean dominance assumptions to obtain informative bounds. These assumptions are admissible for many empirical situations. We also restrict the support of the potential outcomes to be bounded such that $Y_t^\ast(0),Y_t^\ast(1)\in y=[Y_t^{LB}, Y_t^{UB}]$, where $-\infty<Y_t^{LB}<Y_t^{UB}<\infty$ for $t=0,1$.
To consider $\tau_{ONO}$, $\tau_{NOO}$, and $\tau_{NNO}$, we extend the within-group potential outcomes parallel trends provided in Assumption (ref) to include these groups.
Next, we introduce cross-group mean dominance assumptions to aid the identification of the ATT for these latent groups.
Our mean dominance assumptions compare the untreated potential outcomes of a latent group in a specific treatment arm $D=d$ at a given point in time to its closest latent counterpart. Specifically, the inequality restrictions posit that selection into being observed is (weakly) positively correlated with the untreated potential outcomes for group $D=d$ at time $t$. For instance, invoking Assumption 6(a) helps obtain a tighter upper bound on the counterfactual expectation $\mathbbm{E}[Y_{1}^{\ast}(0)\mid D=0, \text{ONO}]$, thereby narrowing the overall bounds for $\tau_{ONO}$. This is achieved by comparing the ONO group to the OOO group, whose untreated potential outcome mean in the post-treatment period can be point identified using $\mathbbm{E}[Y_1 \mid D=0, S_0=1, S_1=1]$, under positive MS. Similar arguments apply to the other latent groups, where the mean dominance assumptions help to refine the theoretical bounds on counterfactual means. Because each latent group's selection behavior varies across treatment states and time periods, the mean dominance assumptions are stratum-specific.
In the context of the job training example, all these assumptions imply that individuals with higher attachment to the labor force or those less prone to be unemployed in some period/treatment scenario have better wages on average than peers with lower attachment in similar situations (time period, treatment counterfactual, etc.). As the always-observed group will be employed irrespective of training, assuming their potential wages to be higher than those of the other groups is reasonable. The justifiability of these assumptions depends on the empirical context, and researchers need to consider them carefully.
To identify bounds for $\tau_{ONO}$, $\tau_{NOO}$, and $\tau_{NNO}$, we introduce a stronger version of IS (Assumption (ref)). It imposes independence on the joint counterfactual selection distribution rather than only relating to the marginal distributions.
Assumption (ref) states that conditional on the initial period selection status, the joint counterfactual selection mechanism is independent of treatment assignment. This is a stronger assumption than its marginal version in Assumption (ref), and is a sufficient condition for the latter. Under this assumption, the observed selection probabilities conditional on initial period selection and treatment enable us to identify all latent group proportions.\footnote{See Lemma (ref) and its proof in Appendix (ref).}
To derive the ATT bounds for these latent groups, decompose $\tau_{g}$ as follows,
The treatment effect for the ONO group ($\tau_{ONO}$) can be further decomposed using Equation ((ref)) as,
As explained in Section (ref), we can use the group of treated individuals for whom the outcome is observed in both periods to partially identify $\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{0}^{\ast}(1)|D=1, ONO]$. Similarly, we can use the group of untreated individuals for whom the outcome is observed in the first period only ($D = 0, S_0 = 1, S_1 = 0$) to partially identify $\mathbbm{E}[Y_{0}^{\ast}(0)|D=0, ONO]$. Identification of $\mathbbm{E}[Y_{1}^{\ast}(0)|D=0, ONO]$ combines the theoretical upper and lower bound of the outcome distribution huber2015sharp and the mean dominance Assumption (ref)(a).
Proof of Theorem (ref) is given in the Appendix (ref).
The ATT for NNO group ($\tau_{NNO}$) also can be further decomposed using equation ((ref)) as,
We can use the group of treated individuals for whom the outcome is not observed in the pre-treatment period but observed in the post-treatment period ($D = 1, S_0 = 0, S_1 = 1$) to partially identify $\mathbbm{E}[Y_{1}^{\ast}(1)|D=1, NNO]$. The other terms, $\mathbbm{E}[Y_{0}^{\ast}(1)|D=1, NNO]$, $\mathbbm{E}[Y_{0}^{\ast}(0)|D=0, NNO]$ and $\mathbbm{E}[Y_{1}^{\ast}(0)|D=0, NNO]$ can be partially identified by imposing the theoretical upper and lower bounds of the respective outcome distributions huber2015sharp and tighten these by imposing outcome mean dominance assumptions (ref)b.(i) and (ref)b.(ii), respectively.
Proof of Theorem (ref) is given in the Appendix (ref).
The ATT for NOO group ($\tau_{NOO}$) can be decomposed using equation ((ref)) as follows,
We can use the group of treated individuals for whom the outcome is not observed in the pre-treatment period but observed in the post-treatment period ($D = 1, S_0 = 0, S_1 = 1$) to partially identify $\mathbbm{E}[Y_{1}^{\ast}(1)|D=1, NOO]$. The term, $\mathbbm{E}[Y_{1}^{\ast}(0)|D=0, NOO]$, can be point identified using $\mathbbm{E}[Y_{1}|D=0,S_{0}=0,S_{1}=1]$ which considers the untreated individuals not observed in the pre-treatment period but observed in the post-treatment period ($D=0, S_0=0, S_1=1$). Under positive monotonicity, this observed group is composed exclusively of the NOO latent subgroup. The remaining terms, $\mathbbm{E}[Y_{0}^{\ast}(1)|D=1, NOO]$ and $\mathbbm{E}[Y_{0}^{\ast}(0)|D=0, NOO]$, can be partially identified by imposing the theoretical upper and lower bounds of the respective outcome distributions huber2015sharp where we tighten them by imposing outcome mean dominance Assumption (ref)(c).
Proof of Theorem (ref) is given in the Appendix (ref).
After partially identifying the ATT for each of these latent groups, we provide a way to partially identify the overall ATT, $\tau$, which relies on combining the identified sets of the ATT of these different latent groups with their appropriate population weights.
Proof can be found in Appendix (ref). Theorem (ref) provides a useful identification result about a parameter that is the typical target of a DiD analysis. The expressions for the lower and upper bounds give a detailed and transparent description of the sources of heterogeneity and information affecting the overall ATT. This can help researchers gauge how important each group is. For example, the weights assigned to each identified set and the width of the corresponding group-specific interval together indicate the relative influence of each group on the overall bounds.
The estimation of the bounds defined in Theorem (ref) and Theorem (ref) are based on the sample analogues of the population counterparts. To estimate the bounds defined in Theorem (ref) we first have to estimate the mixing proportions $p_{OOO1}(v^{l}_{1})$ and $p_{OOO0}(v^{l}_{0})$. Formally, we have,
where,
With these estimated mixing proportions, the bounds for $\tau_{OOO}$ under Theorem (ref) can be estimated as follows,
where $\hat{y}_{\hat{p}_{OOO0}(v^{l}_{0})}$ and $\hat{y}_{1-\hat{p}_{OOO0}(v^{l}_{0})}$ are $\hat{p}_{OOO0}(v^{l}_{0})$-th and $(1-\hat{p}_{OOO0}(v^{l}_{0}))$-th quantile of the conditional distribution $Y_{1}-Y_{0}$ for the untreated individuals observed in both time periods. Similarly, $\hat{y}_{\hat{p}_{OOO1}(v^{l}_{1})}$ and $\hat{y}_{1-\hat{p}_{OOO1}(v^{l}_{1})}$ are $\hat{p}_{OOO1}(v^{l}_{1})$-th and $(1-\hat{p}_{OOO1}(v^{l}_{1}))$-th quantile of the conditional distribution $Y_{1}-Y_{0}$ for the treated individuals observed in both time periods. In general, the relevant $q$-th quantile of the conditional distribution $Y_{1}-Y_{0}$ for the treated individuals observed in both time periods is calculated as,
The bounds for $\tau_{OOO}$ under Theorem (ref) can be estimated similarly. First, estimate the required mixing proportion $p_{OOO1}$ as follows,
Next, the estimated versions of $LB_{OOO1}$ and $UB_{OOO1}$ can be obtained as,
where $\hat{y}_{\hat{p}_{OOO1}}$ and $\hat{y}_{1-\hat{p}_{OOO1}}$ are $\hat{p}_{OOO1}$-th and $(1-\hat{p}_{OOO1})$-th quantile of the conditional distribution $Y_{1}-Y_{0}$ for the treated individuals observed in both time periods. Next, $\mathbbm{E}[Y_{1}-Y_{0}|D=0,S_{0}=1,S_{1}=1]$ (denote as $E_{OOO0}$ for notational ease) will be estimated using its sample analogues as,
Finally, the bounds for $\tau_{OOO}$ defined in Theorem (ref) can be estimated as,
The bounds for other latent groups defined in Theorems (ref), (ref), and (ref) can be estimated similarly. The estimation steps are detailed in Appendix (ref).
These sample analogue estimators of the bounding functions are functions of conditional probabilities, means, and trimmed means, and include non-smooth functions of auxiliary parameters for which we use plug-in estimates. Fortunately, when the true mixing proportions are strictly positive,\footnote{When the true mixing proportion is zero, the asymptotic behavior of the bounds can be characterized by letting the trimming proportion shrink to zero sufficiently slowly such that, asymptotically, there are enough observations in the trimming region to guarantee that the expectation is well defined. See andrews2013inference and references therein for similar approaches.} $\sqrt{n}$-consistency and asymptotic normality of these estimators follow from results in chen2003estimation.\footnote{We thank the associate editor for indicating that the proposed estimators are encompassed by the results in chen2003estimation. In the supplementary materials Section (ref), we verify that the requisite conditions for applying chen2003estimation hold, and we derive the resulting asymptotic distributions and their properties.} Hence, we can obtain two types of confidence intervals by applying standard inference procedures. Following lee2009 and huber2014treatment, let $\widehat{LB}_{\tau_g}$ and $ \widehat{UB}_{\tau_g}$ be the estimated bounds for a specific latent group $g$ using the estimation method discussed above and $\hat{\sigma}_{LB_{\tau_g}}$ and $\hat{\sigma}_{{UB}_{\tau_g}}$ denote their respective standard deviations, obtained through bootstrap huber2014treatment,chen2015bounds.\footnote{See bugni2010bootstrap, bugni2015specification, and andrews2024misspecified, among others, for inferential methods for partially identified models that solve moment inequalities.} Then, we can compute the first confidence interval as,
which will contain the true bounds with at least 95% probability. The second option for confidence intervals is based on imbens2004confidence. These confidence intervals are focused on covering the true treatment effect with 95% probability, which are calculated as $[\widehat{LB}_{\tau_g}-C_{n}\cdot \frac{\hat{\sigma}_{LB_{\tau_g}}}{\sqrt{n}},\widehat{UB}_{\tau_g}+C_{n}\cdot \frac{\hat{\sigma}_{UB_{\tau_g}}}{\sqrt{n}}]$ where $C_{n}$ satisfies
and are uniformly valid.\footnote{stoye2009more highlights that the refined confidence interval from Imbens and Manski achieves uniform coverage only under the assumption that the estimator of the identified set length ($\widehat{UB}_{\tau_g}-\widehat{LB}_{\tau_g}$) is super-efficient near point-identification. He establishes a weaker sufficient condition of super-efficiency as being joint-normality of the lower and upper bound estimators along with the bounds being ordered (by construction). Both of these conditions hold in our case. stoye2009more proposes a refinement with and without super-efficiency, and stoye2020simple extends it to partial identification of a pseudo true parameter.}
This section presents simulation evidence of the bias induced by sample selection on the standard DiD estimates. It further demonstrates the feasibility of the identification and estimation procedures for partially identifying the ATTs of the different latent groups ($\tau_{OOO}$, $\tau_{ONO}$, $\tau_{NOO}$ and $\tau_{NNO}$) proposed above. The main data generating process (DGP) used in the simulations is as follows,
where $\bigl(a_i,\;c_i,\;u_{i0},\;v_{i0},\;u_{i1},\;v_{i1}\bigr)'$ is drawn jointly from the six-dimensional truncated multivariate normal distribution, $\mathcal{N}_6\bigl(0,\Sigma_6\bigr)$, whose support is restricted to the hypercube ${[-M,M]^6}$. The covariance matrix $\Sigma_6$ has unit marginal variances, $\operatorname{Cov}(a_i,c_i)=\rho_{ac}$, $\operatorname{Cov}(u_{it},v_{it})=\rho_{u_t v_t}$ for $t=0,1$, and all remaining covariances are set to zero. The variable $c_i$ plays the role of time-invariant unobserved heterogeneity. Finally, $b_i$ and $w_{i1}$ are independent standard normal random variables. Latent group $g$ is defined by the tuple $\bigl(S_{i0},S_{i1}(0),S_{i1}(1)\bigr)\in\{NNN,NNO,NOO,ONN,\\ ONO,OOO\}$. The treatment effects $\tau^{g_i}$ and time effects $t^{g_i}_t$ are both group-specific, with
We set $M=5$, $\zeta=1.5$, $\rho_{ac}=0.7$, $\rho_{u_0v_0}=0.7$, $\rho_{u_1v_1}=0.6$, $\tau_g=(-1,1,3,1,4,5)'$ , $t_0=(0,2,3,1,4,5)'$ and $t_1=(1,3,4,2,5,6)'$. The observed data are $\left\{Y_{it},S_{it},D_{i}\right\}_{i=1}^n$ where $Y^\ast_{i1}=Y^\ast_{i1}(0)\cdot (1-D_i)+Y^\ast_{i1}(1)\cdot D_i$, $S_{i1}=S_{i1}(0)\cdot (1-D_i)+S_{i1}(1)\cdot D_i$, and $Y_{it}=S_{it}\cdot Y^\ast_{it}$, for each $t=\{0,1\}$.
This DGP satisfies Assumptions (ref), (ref), (ref), (ref), (ref), and (ref). The true overall ATT is $\tau=2.8504$ and the latent group-specific ATTs are given by: $\tau_{OOO}=5$, $\tau_{ONO}=4$, $\tau_{ONN}=1$, $\tau_{NOO}=3$, $\tau_{NNO}=1$ and $\tau_{NNN}=-1$.
Under this DGP, the true mixing probability is $p_{OOO1}$ = 0.7052 and the true numerical bounds for $\tau_{OOO}$ are given by $[LB_{\tau_{OOO}}, UB_{\tau_{OOO}}] = [3.6829,5.1969]$. Notice that these contain the true ATT for the OOO group. The mathematical derivation of these numerical bounds can be found in Appendix (ref).
We draw random samples of 500, 1000, and 2000 observations from this DGP and compute the empirical distribution of the estimated bounds over 10,000 simulation draws (replications). The bounds for $\tau_{OOO}$ are estimated under two sets of assumptions. The first corresponds to the bounds given in Theorem (ref), which we refer to as the without-monotonicity scenario. The second corresponds to the bounds presented in Theorem (ref), which we refer to as the with-monotonicity scenario. Figure (ref) plots the distribution of the estimated mixing probability, $\hat{p}_{OOO1}$, for the with-monotonicity scenario, which is centred around the black vertical line, representing the true value of the mixing probability. The average estimated bounds for $\tau_{OOO}$ under the two sets of assumptions are given in Table (ref). These are seen to contain the true ATT for OOO of 5. As we discuss in Section (ref), the na\"{\i}ve DiD estimate exhibits an upward bias of around 55.4% in this relatively simple DGP, which satisfies the full independence assumption (ref)$(Joint)$. The 95% CI reflects the coverage probability of the true interval, which hovers around 93%, whereas the Imbens and Manski (IM) 95% CI is the coverage probability of the true parameter, which hovers around 99%. Simulation results for the remaining latent groups (ONO, NOO, and NNO) are provided in Appendix (ref).
\paragraph{Mixing Proportion Approaching One:} We also examine how the bounds for $\tau_{OOO}$ behave under monotonicity (Theorem (ref)) as we increase the mixing proportion to approach 1. To do this, we modify the main DGP given above by varying the parameter $\zeta$ between 0.1 and 3, which generates mixing proportions ranging from 0.9601 to 0.6685, respectively. The results for this case can be found in Appendix Table (ref). It becomes clear from these results that as the mixing proportion approaches 1, the estimated bounds get tighter and contract towards the true value of $\tau_{OOO}=5$.
\paragraph{Violation of Monotonicity:} To further investigate the behavior of the proposed bounds for $\tau_{OOO}$ under the no-monotonicity setting of Theorem (ref), we modify the DGP by drawing $\zeta \in \{-1.5, 0, 1.5\}$ at random. This generates positive selection for some units, negative selection for others, and no selection for the rest, such that monotonicity is violated on average for the overall sample. The resulting bounds are reported in Appendix Table (ref) and contain the true ATT for the OOO group. These are seen to be wider compared to the bounds obtained under monotonicity (Table (ref)). In conclusion, these findings validate our theoretical results on the relationship between proposed no-monotonicity and monotonicity bounds.
\paragraph{Strength of Assumption (ref):} To assess how different versions of Assumption (ref) affect identification, we modify the DGP so that Assumptions (ref)(a) and (ref)(b) hold, while both monotonicity and Assumption (ref) fail. We then estimate the without-monotonicity bounds for $\tau_{OOO}$ based on Theorem (ref). This case is discussed in Appendix section (ref) and the results are reported in Appendix Table (ref). We find that these bounds are noticeably wider than the without-monotonicity bounds obtained under Assumption (ref) (Table (ref)).
We conduct a similar exercise by modifying the DGP so that Assumption (ref)(a) and monotonicity hold together. We then estimate the with-monotonicity bounds given in Theorem (ref). The estimated bounds provided in Appendix Table (ref) are again wider than the with-monotonicity bounds estimated under Assumption (ref) that are reported in Table (ref). These findings illustrate how strengthening Assumption (ref) narrows the identified set for $\tau_{OOO}$ and produces more informative bounds.
In this section, we illustrate our partial identification approach with two empirical applications. First, we revisit the 1970's experiment described in lalonde1986evaluating by using the Aid to Families with Dependent Children (AFDC) sample of women from the National Supported Work (NSW) training program. Here, we consider the sample selection problem arising from unemployment (reported zero earnings). For the second application, we consider the study by bloom2015does in which an experiment evaluates the effectiveness of a company's working from home policies. In this application, sample selection bias arises from employee attrition. Although both applications were conducted as randomized experiments, outcomes are only observed conditional on being selected into the sample. If selection is endogenous to the outcomes of interest in both the pre- and post-treatment periods, then random assignment does not necessarily ensure that treated and control units remain comparable, especially if individuals enter or exit the sample differently across groups and periods.
NSW was a temporary employment program implemented in the United States between 1975-1979. It was designed to help individuals from disadvantaged populations find stable employment by offering them structured work experience and counselling in a sheltered environment. The program targeted four disadvantaged socio-economic groups, with qualified applicants being randomly assigned to training. In our empirical analysis, we only consider the AFDC sub-sample of women originally studied in calonico2017women, consisting of 1185 individuals out of which 600 received training. We apply the proposed approach to account for sample selection arising from unobserved earnings.
We treat zero earnings as unobserved wages due to individuals' inability to find suitable employment. In this case, $D=1$ if an individual is assigned to receive training and zero otherwise. $Y_{t}^\ast(0)$ and $Y_{t}^\ast(1)$ are potential earnings of an individual and $S_{t}(1)$ and $S_{t}(0)$ are potential indicators for being employed or not. Our treatment effect of interest, $ \tau_{OOO}=\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{1}^{\ast}(0)|D=1, S_{0}=1, S_{1}(0)=1, S_{1}(1)=1]$, captures the effect of training on earnings for the latent group that is “always employed”. Policy makers often care about this latent group as it reflects the returns to training for workers who would remain employed regardless of training, as they are highly attached to the labor force. Many policies aim to improve outcomes among continuously employed participants. Therefore, this parameter is informative for evaluating the impact of a program within a committed, consistently engaged population, reflecting an important component of the treatment effect on the intensive margin. During the pre-treatment period, 73.3% of the treated and 74.7% of the control samples reported zero earnings. The follow-up survey indicates 45.0% of the treated and 46.2% of the control group individuals reported zero earnings (see Table (ref)).\footnote{Tables (ref) and (ref) present the covariates used in the analysis along with means for the observed (unobserved) samples for both treatment groups.}
We estimate bounds for $\tau_{OOO}$ under two sets of assumptions. The first considers Assumptions (ref), (ref), (ref)(a), (ref)(b), and (ref) which we refer to as the without-monotonicity scenario and the second under Assumptions (ref), (ref), (ref), (ref)(a), and (ref) which we refer to as the with-monotonicity scenario. Intuitively, Assumption (ref) requires that women's pre-treatment wage offers and the ability/willingness to secure a job (i.e., be employed) are unaffected by their eventual assignment into the training program. For example, if we consider $S(0)$ and $S(1)$ to reflect the decision to work as determined by the interplay of reservation wages and wage offers, the no anticipation assumption rules out scenarios in which, upon learning of their assignment to the training program, women immediately increase their reservation wages or firms increase their wage offers to those slated to participate in the program. As mentioned in calonico2017women, this assumption is likely to hold for earnings measured in 1975, before assignment to the program began. Assumption (ref) will be plausible if, in the absence of treatment, latent wage trends do not differ between the treated and untreated samples of women. Importantly, to partially identify $\tau_{OOO}$, this condition only needs to hold for the always-observed group. In our setting, these are likely women who have higher attachment to the labor force. Given the program's focus on low-income women with dependents and recent unemployment spells, this group is more homogeneous in their employment patterns and wage trends, making the parallel trends assumption credible. Similarly, the ignorability of potential selection (Assumption (ref)) requires that the counterfactual employment behavior of women in the post-treatment period be similar across training assignment groups conditional on pre-treatment unemployment. This is plausible if differences are driven primarily by pre-existing employment patterns rather than differential assignment to training, and rules out that women assigned to the training program would have been, on average, more likely to become unemployed than their counterparts not assigned for training. For example, it rules out the possibility that, in the absence of treatment, reservation wages for the $D=1$ and $D=0$ groups would respond differently to shocks or unobservables, such as having a young child in the household in the post-training period.
In this application, we assume monotonicity operates in the positive direction, implying that receiving training increases the probability of employment and, hence, the likelihood of observing employed individuals in the treated sample compared to the control sample. The results under each set of assumptions are presented in Table (ref). The bounds derived without Assumption (ref) are wide and uninformative. Imposing monotonicity substantially tightens the bounds, demonstrating the assumption's informational content. We cannot rule out $\tau_{OOO}$ between \$1,404 and \$1,718. The na\"{\i}ve DiD estimates suggest an increase of around \$1,613 dollars in annual earnings due to training.\footnote{As discussed in Appendix (ref), the na\"{\i}ve DiD is always within the identified set for $\tau_{OOO}$ based on the trimming approach discussed in Theorems (ref)-(ref).}
While it's interesting to characterize the effects of training for the always-observed group of women, who have relatively high attachment to the labor force and whose estimated ATT could serve as a useful benchmark when treatment effect heterogeneity is limited, it's natural to wonder about the effects for other groups. This is especially true since the NSW demonstration explicitly targeted more vulnerable populations with weaker labor market attachment. Specifically, policymakers may be more interested in estimating the ATT for those i) unemployed before training who will be employed post-treatment irrespective of training (i.e. $\tau_{NOO}$), ii) employed only if they are given training (i.e. $\tau_{NNO}$), or iii) people employed before training who will only be employed post-treatment if they are given training (i.e. $\tau_{ONO}$). The estimated bounds corresponding to these ATTs are presented in Table (ref) and use the results in Theorems (ref)-(ref), which impose assumptions (ref), (ref), (ref), (ref), (ref) and (ref).
The last column of Table (ref) reports the estimated population shares corresponding to each latent group type. The proportion of NOO is estimated to be large (37.4%), which aligns with the selection criteria for participation in NSW. This also matches with the data where we observe a high fraction of individuals reporting zero earnings in the pre-treatment period and a sizable increase in post-treatment employment among both the treated and control groups. Meanwhile, the share of the NNO latent group reflects the relatively small extensive margin effect of the training program among the individuals who are initially unemployed, implying that training induces only a small increase in employment. The share of OOO is important, at 16.5%, but also highlights the fact that we can only learn about the impact of the policy with more certainty for a small part of the population. The limited information available for the NOO, NNO, and ONO latent groups is reflected in the wide estimated ATT bounds, which in all cases encompass both the estimated set for $\tau_{OOO}$ and zero. Therefore, we cannot rule out either the case of homogeneous treatment effects relative to OOO or the possibility of a nil effect of the training program for these groups. The overall ATT lies in the interval $\tau \in [-13.5636,15.4139]$. These bounds are estimated based on the results in Theorem (ref). One should also note that, even under these assumptions, bounds for the overall ATT are not very informative given the wide identified sets for the ATT of each latent group and the limited importance of the always-observed group in the population at large.
In this section, we revisit the results of an experiment at Ctrip, a 16,000-employee, NASDAQ-listed Chinese travel agency bloom2015does. This was carried out to evaluate the effectiveness of working from home (WFH) policies on employee performance. The experiment was conducted from January 2010 to August 2011 among eligible employees of the airfare and hotel departments at the company's Shanghai call centre, who volunteered to participate. Out of the employees who volunteered, only 49.5% (249) met the eligibility criteria set by the company. Employees with even-numbered birthdays were assigned to the treated group, where they were allowed to WFH and those with odd birthdays were assigned to the control group with no option to WFH. Accordingly, 52.6% (131) and 47.4% (118) were allocated to the treatment and control groups, respectively. We use average individual weekly performance z-scores, a combination of different key performance indicators standardized based on each job type, to evaluate employee performance.\footnote{See bloom2015does for a detailed description of the experiment and data collection process.} We will illustrate how our identification strategy can be used to account for selection bias due to employee attrition in the experimental period.\footnote{Table (ref) and (ref) in Appendix (ref) reports descriptive statistics for the observed covariates.}
In this application, $D=1$ if an employee is eligible to WFH and zero otherwise. $Y_{t}^\ast(0)$ and $Y_{t}^\ast(1)$ are potential average individual weekly performance z-scores of employees under the two treatment scenarios. $S_{t}(1)$ and $S_{t}(0)$ are potential indicators for attrition, where $S_{t}=1$ is the realized sample participation, indicating the particular employee stayed with the company. Our primary focus is the identification of $\tau_{OOO}=\mathbbm{E}[Y_{1}^{\ast}(1)-Y_{1}^{\ast}(0)|D=1, S_{0}=1, S_{1}(0)=1, S_{1}(1)=1]$ which captures the ATT of WFH eligibility on employee performance for the subgroup of employees who will stay with the company irrespective of WFH or not.\footnote{Since the treatment variable $D=1$ indicates eligibility to the program, we could interpret $\tau_{OOO}$ as the intent-to-treat effect of the WFH program on worker's performance for the always-observed group.} This could be interpreted as the “intensive” margin effects of the WFH policy, that is, changes in performance for employees that would have been retained even in the absence of WFH.
Table (ref) shows the attrition rates during the experimental period. We observe an attrition rate for the control group (34.8%) that is more than double that for the treated group (16.0%).
We estimate two sets of bounds for $\tau_{OOO}$. The without-monotonicity case imposes assumptions (ref), (ref), (ref)(a), (ref)(b), and (ref), and the with-monotonicity scenario under assumptions (ref), (ref), (ref), (ref)(a), and (ref). Intuitively, Assumption (ref) requires that employee performance and the decision to remain with the firm before the WFH policy is introduced are unaffected by the assignment to remote work or not. For example, it rules out scenarios where workers who would have otherwise left the firm in the pre-treatment period decide to stay in order to benefit from the WFH scheme. Given the random assignment to WFH based on birth date, it is unlikely that workers could anticipate their assignments and alter pre-treatment quitting behavior. However, we cannot rule out the possibility that workers who were considering quitting might have been more likely to volunteer for the program and delay their decisions until after assignment to the treatment group. This would not violate the no anticipation assumption and is naturally incorporated by the ignorability of potential selection assumption, which only requires the same post-treatment quitting behavior between the two WFH assignment groups, conditional on being observed in the pre-treatment period. For parallel trends on latent outcomes (assumption (ref)), it is plausible to expect that, in the absence of WFH policy, performance trends for employees would have followed similar paths across treatment assignment groups. This is because the always-observed group comprises of workers with stable attachment to the firm who continue to perform the same roles within the same organizational environment and are evaluated under the same performance criteria. As a result, differences in performance trends are less likely to reflect systematic differences in underlying productivity growth, making parallel trends a reasonable assumption. For the second case, we assume positive monotonicity, implying that WFH employees are at least as likely to stay in the company during the experimental period as those who do not WFH. This can be justified as employees assigned to the WFH treatment can choose whether or not to take advantage of the policies, and are likely better off due to increased convenience and reduced commuting costs if they decide to participate.
The results are presented in Table (ref). The na\"{\i}ve DiD implies that the overall performance of the treated group is 0.1759 standard deviations higher than what would have been in the absence of the WFH treatment.\footnote{As discussed in Appendix (ref), the na\"{\i}ve DiD is always within the identified set for $\tau_{OOO}$ based on the trimming approach discussed in theorems (ref)-(ref).} The more flexible bounds estimated without assuming monotonicity are wide, but rule out a decline in standardized performance larger than 0.33 std. deviations as well as increases above 0.67 std. deviations. These can be tightened by imposing Assumption (ref), leading to an identified set ranging between 0.0057 and 0.3806 standard deviations improved performance for the effect of being assigned to WFH among employees who stay with the company regardless of being eligible for the policy.
In contrast to the previous application, all workers in this setting are observed in the initial period, which reduces the number of possible latent types. However, the share of always-observed workers among treated individuals observed in both periods is at most 77.7%.
Furthermore, as presented in Table (ref), under monotonicity, the OOO group is estimated to represent close to 65.3% of the overall population. The ONO latent group is an interesting subtype and represents those workers who would have left the company if WFH were not provided but would stay otherwise. Using results in Theorem (ref), we can partially identify the ATT for this group by imposing Assumptions (ref), (ref), (ref), (ref), (ref) and (ref)(a). As can be seen in Table (ref) the ONO encompasses 18.7% of the population. The estimated identified set for $\tau_{ONO}$ includes zero and a wide range of values. Therefore, we cannot determine the direction of the effect of WFH eligibility (if any) on workers' performance for this latent group without additional information.
In this particular application, since there is no sample selection in the pre-treatment period and we assume monotonicity, there are only three latent groups, OOO, ONO and ONN. We have estimated bounds for the ATT for both the OOO and ONO groups, which together account for roughly 84.0% of the total population. The remaining 16.0% corresponds to workers who would have exited the company regardless of their WFH eligibility status. The overall ATT lies in the interval $\tau \in [-0.9413,1.7086]$. These bounds are estimated based on the results in Theorem (ref).
In this article, we address the challenge of endogenous sample selection within the difference-in-differences (DiD) framework. We demonstrate that the standard DiD estimates can be biased when sample selection is ignored, even under the strong assumption of independence between the selection and treatment assignment mechanisms. We propose methods for partial identification of average treatment effects for different latent treated subpopulations by building on the trimming procedure of lee2009.
For the OOO-group, identification relies on the insight that individuals observed in both periods are a mixture of two possible latent groups. The mixture proportions are identified under alternative sets of assumptions on the selection and treatment assignment mechanisms. Specifically, we consider scenarios with and without MS. When MS is not imposed, the latent strata proportions are identified under the IS assumption, which assumes counterfactual selection probabilities to be equal between the treated and untreated units, conditional on being observed in the baseline period. In cases where monotonicity holds, we achieve tighter bounds on the ATT for the OOO by imposing ignorability for selection in one direction only. We also discuss identification of ATT bounds for the OOO group when covariates are observed by considering both no monotonicity and a relaxed version of monotonicity. We also extend the basic identification argument (without covariates) to settings in which the researcher has access only to repeated cross-sectional data and to staggered treatment adoption across multiple time periods.
Additionally, we present the identified sets for the ATT of other empirically relevant latent groups, such as ONO, NON, and NOO, under MS and outcome mean dominance assumptions. Combining these group-specific identified sets with weights also allows us to partially identify the overall ATT in the population. We illustrate our results through two empirical illustrations: 1) bounding the effects of a job training program and 2) bounding the effects of a work-from-home policy on employee performance. These applications highlight the practical relevance of the proposed bounds in two different empirical settings.
\singlespacing