Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Estimating the Intensive Margin Effect in Panel Data Settings
abstractMany policies operate through two different channels: the extensive margin (e.g., the decision to participate) and the intensive margin (e.g., the intensity of the response among participants). This paper develops a novel identification strategy to estimate the intensive margin effect in panel data settings. I adapt the Horowitz-Manski-Lee bounds to the Changes-in-Changes framework to partially identify both the average and quantile intensive margin treatment effects. Additionally, I explore how to leverage multiple sources of sample selection to relax the monotonicity assumption in the original Horowitz-Manski-Lee bounds, which may be of independent interest. Alongside the identification strategy, I present estimators and inference results. I illustrate the relevance of the proposed methodology by analyzing a job training program in Colombia.
JEL-Code: C14, C21, C23.
Keywords: Intensive Margin, Difference-in-differences, Changes-in-Changes, Sample selection, Panel Data, Repeated Cross-Sections, Principal stratification, Partial identification.
\thispagestyle{empty}
\setcounter{page}{1}
spacing{1.5}
\section{Introduction}
A central concern in economics is understanding the mechanisms through which policies operate. Aggregate policy effects often mask distinct channels through which individuals are affected.
A prominent instance arises in policies that impact the extensive margin (e.g., the decision to participate) and the intensive margin (e.g., the intensity of the response among participants).
A canonical example of such policies is job training programs, which affect employment (extensive margin) and also wages for those employed (intensive margin) heckman_economics_1999, ham_effect_1996,angrist_estimation_2001.\footnote{Additional examples include, but are not limited to, education policies that affect enrolment and performance for those enrolled krueger_experimental_1999,cornwell_student_2005; carbon taxes that affect firms' entry and exit and emissions for active firms greenstone_effects_2012; weather shocks that affect the decision to migrate and income for those who do not migrate groger_internal_2016. In the medical literature, this phenomenon is known as truncation due to death zhang_estimation_2003,ding_identifiability_2011,rubin_causal_2006.}
Distinguishing between these margins is essential for both welfare analysis and policy design.
The main contribution of this paper is to develop an identification strategy for estimating the intensive margin effect in panel data settings. The existing literature has proposed identification and estimation strategies for intensive margin effects in settings where the treatment can be assumed uncounfounded (e.g., experimental settings).\footnote{This literature was pioneered by zhang_estimation_2003, imai_sharp_2008, lee_training_2009, and zhang_likelihood-based_2009, who introduced a principled approach to partially identify the intensive margin effect. Subsequent research has extended this literature in multiple directions, including sharpening bounds with covariates semenova_generalized_2025, long_sharpening_2013, grilli_nonparametric_2008,samii_generalizing_2025, ding_identifiability_2011, instruments mattei_identification_2014, post-censoring outcomes yang_using_2016, or a structural model for selection honore_selection_2020, incorporating longitudinal data comment_survivor_2025, grossi_bayesian_2023, considering noncompliance chen_bounds_2015, blanco_bounds_2020, targeting different estimands huber_sharp_2015,bartalotti_identifying_2023, and exploring non-binary treatment regimes lee_lee_2025. However, all these papers rely on the assumption that treatment is unconfounded.} This assumption is often implausible in panel data settings, where treatment is confounded with unobservable characteristics liu_practical_2024,arkhangelsky_causal_2024,ghanem_selection_2025. This paper fills this gap by adapting intensive-margin estimands and identification strategies to panel data settings in which treatment is not randomly assigned.
I formalize the intensive margin effect using principal stratification analysis frangakis_principal_2002. This framework
characterizes each unit by its potential participation decision under both treatment arms, thereby partitioning the sample into four distinct subpopulations or principal strata. For example, in the job training case, those who are always employed, those employed only if trained, those employed only if not trained, and those who are always unemployed. The extensive margin effect refers to the treatment effect on participation (employment in the job training example).
The intensive margin effect is defined as the treatment effect on the response
for the subset of units that would participate under both treatment arms, and hence for whom there is no effect on participation, i.e., no extensive margin effect. In the job training case, these are the workers who are employed regardless of whether they receive the training or not. For these workers, once participation is defined in terms of employment, any remaining treatment effect operates through the intensive margin, such as changes in wages.
The identification strategy developed in this paper focuses on the Always-Observed principal stratum, i.e., the units with observed outcomes regardless of their treatment status.\footnote{The methodology developed in this paper can also be applied in settings where outcomes are observed for all units, but we are interested in isolating the intensive margin effect. Consider the `bad control' example in angrist_mostly_2009, where we are interested in the effect of education on earnings. Part of this effect operates through occupational choice. In this example, we may be interested in the effect of education on earnings net of occupational changes. In this case, we would estimate the effect on the subpopulation of units that always have a white-collar occupation, regardless of their education level.} The effect on these units can be interpreted as the policy's intensive margin effect. To achieve identification, I extend
the Howorowitz-Manski-Lee bounds horowitz_identification_1995, lee_training_2009 to the Changes-in-Changes (CiC) model athey_identification_2006. The identification strategy relies on
the assumption
that all the unobserved confounders affecting the outcome for the Always-Observed units can be captured in a single latent index (for example, ability in the case of wages). The distribution of this index may differ for the Always-Observed units in the treatment and control groups, but it is assumed to be constant over time within each group (that is, the distribution of ability for workers who are always employed can be different in the treatment and control groups as long as this distribution does not vary over time). Combined with information on the principal strata proportions, this assumption delivers partial identification of two causal estimands: (i) the Average Treatment Effect on the Treated for the Always-Observed stratum ($\text{ATT}_{\text{AO}}$) and (ii) the Quantile Treatment Effect on the Treated for the same stratum ($\text{QTT}_{\text{AO}}$).
Another contribution of this paper is the estimation of the principal strata proportions. I also employ the CiC approach to model participation. Combined with a monotonicity assumption, as in the original Horowitz-Manski-Lee bounds, this method delivers point identification of the principal strata proportions. The required monotonicity assumption implies that treatment can affect participation only in `one direction' for all the units, ruling out the existence of one of the strata (e.g., units observed only if treated).\footnote{This assumption is analogous to the monotonicity assumption in IV settings angrist_identification_1996 that rules out the existence of defiers.} In the job training context, this assumption implies that all workers employed without training would also be employed if trained. I relax this assumption when the data include multiple sources of participation. For instance, wages may be unobserved for both unemployed workers and workers who decide to migrate. I assume each source is monotonic while allowing the direction of monotonicity to differ across sources. For example, the training may reduce the probability of unemployment but increase the likelihood of migration for all units. This extension enables all four principal strata to exist while still achieving point identification of their proportions. This novel extension of the original Horowitz-Manski-Lee bounds could also be of interest in settings where treatment is unconfounded.
I illustrate the proposed methodology by analyzing a job training program in Colombia. Using data from attanasio_subsidizing_2011, I estimate the intensive margin effect of the training on salaried earnings. I contrast these findings with those obtained using a `naive' approach,
in which the CiC estimator is applied to units with observed wages.
The `naive' approach suggests that the training increased wages by 13% for trained units. In contrast, using the methodology presented in this paper, I fail to reject the null that the training's intensive margin effect was zero. When examining the $\text{QTT}_{\text{AO}}$, I find that the estimated bounds vary across the outcome distribution, suggesting heterogeneous intensive margin effects with potential distributional implications.
To date, the literature on panel data
has focused on correcting selection bias and targeting the ATT heckman_sample_1979,hausman_attrition_1979,fitzgerald_analysis_1998,ghanem_correcting_2024,bhattacharya_inference_2008, semykina_estimating_2010,carlson_addressing_2024. While these approaches recover causal effects under sample selection, their estimand cannot be interpreted as the intensive margin effect.\footnote{This estimand identifies the effect for all treated units and thus integrates over all the principal strata, mixing intensive and extensive margin effects.} More
recently, there has been growing attention to estimating the intensive margin effect in panel data settings. In parallel with this work, shin_difference--differences_2024, rathnayake_difference--differences_2024 extend Horowitz-Manski-Lee bounds to a Difference-in-Differences (DiD) setting and target one of the estimands proposed in this paper (the ATT on the Always-Observed principal stratum). The methodology proposed by this paper can also be applied to the DiD design, as illustrated in Appendix (ref).
A key distinction is that both shin_difference--differences_2024, rathnayake_difference--differences_2024 identify the principal strata proportions under a parallel trends assumption in participation, whereas I identify them using the Changes-in-Changes model.
An advantage of my approach is that it guarantees estimated probabilities that lie in the unit interval.
Moreover, this paper is the first to establish partial identification of distributional effects, the $\text{QTT}_{\text{AO}}$.
This estimand is particularly relevant when policymakers are interested not only in the average impact of a policy but also in its distributional effects. By shifting the focus from means to distributions, the CiC framework enables the exploration of heterogeneity of treatment effects along the outcome distribution. Finally, I consider several practical extensions—including covariates, binary outcomes, and repeated cross sections—which further enhance the applicability of the proposed framework.
The remainder of this paper is organized as follows. Section (ref) introduces the notation and set-up. Section (ref) introduces the identification strategy to estimate causal effects for the Always-Observed units. Section (ref) derives the estimation and inference results. Section (ref) applies the methodology to data from a job training program from attanasio_subsidizing_2011. Section (ref) concludes. The Appendix sections describe the extensions of the main results ((ref)), prove them ((ref)) and provide supplementary material ((ref)).
\section{Notation and set up}
For simplicity, I consider the two-groups two-periods case. Consider a random sample of $N$ units observed during two periods, $t \in \{1,2\}$. Let $w_{it}$ denote the treatment of unit $i$ at time $t$. No one is treated in the pre-treatment period, $w_{i1} = 0$ $\forall i$. Some units are treated in the post-treatment period, $w_{i2}\in\{0,1\}$. Let $G$ denote a treatment group indicator, with $G_{i} = 1$ if unit $i$ is treated in the second period and zero otherwise, $G_{i} \equiv w_{i2}$. Let $Y_{it}$ be the outcome of unit $i$ at time $t$. I adopt the potential outcome framework and assume the Stable Unit Treatment Value Assumption (SUTVA) rubin_randomization_1980, rubin_application_1990. This assumption implies that there is no interference between units and there are no hidden versions of treatment, allowing me to relate the observed $Y_{it}$ to the potential outcomes of unit $i$, $Y_{it}(w_{i1},w_{i2})$:
\begin{equation*}
Y_{it} = Y_{it}(0,1)G_{i} + Y_{it}(0,0)(1 - G_{i}).
\end{equation*}
The outcome for some sampled units may not be observed. Let $S_{it}$ be a binary selection indicator equal to 1 if the outcome of unit $i$ is observed at time $t$ and zero otherwise. A crucial aspect of this setup is the recognition that the selection indicator, $S$, is also a post-treatment variable. Consequently, I must also relate the observed selection outcome, $S_{it}$, to the potential selection outcomes:
\begin{equation*}
S_{it} = S_{it}(0,1)G_{i} + S_{it}(0,0)(1 - G_{i}).
\end{equation*}
Throughout the paper, I will maintain the assumption of no treatment anticipation for both the outcome and the selection, a standard assumption in the literature arkhangelsky_causal_2024.
\begin{assumption}
No treatment anticipation
\begin{align*}
S_{i1}(0,w_{i2}) = S_{i1}(0,w_{i2}^{\prime}) =S_{i1} \quad \forall w_{i2},w_{i2}^{\prime}\forall i ,\\
Y_{i1}(0,w_{i2}) = Y_{i1}(0,w_{i2}^{\prime}) = Y_{i1} \quad \forall w_{i2},w_{i2}^{\prime}\forall i .
\end{align*}
\end{assumption}
Assumption (ref) implies that the treatment does not affect the outcome and the selection in the pre-treatment period. Since no unit is treated in the first period, it is sufficient to index the second-period potential outcomes with $w_{i2}$:
\begin{align*}
S_{i2} = S_{i2}(1)G_{i} + S_{i2}(0)(1-G_{i}), \\
Y_{i2} = Y_{i2}(1) G_{i} + Y_{i2}(0)(1-G_{i}).
\end{align*}
Since selection is a post-treatment variable, the population can be partitioned into subpopulations defined by joint values of the potential selection indicator in the post-treatment period, i.e., $(S_{i2}(1), S_{i2}(0))$. This approach is known as principal stratification frangakis_principal_2002. Every unit belongs to one of the four principal strata. Table (ref) presents the principal strata, where $V_{i}$ denotes the principal stratum of unit $i$. Because potential outcomes are not affected by the treatment, the principal strata exist prior to the treatment assignment. While the observed selection status is a post-treatment outcome, the principal strata act as pre-treatment covariates.\footnote{In the example discussed in footnote (ref), occupational choice is a post-treatment variable and deemed a bad control angrist_mostly_2009, rosenbaum_consequences_1984. Nevertheless, principal strata can be defined in terms of the potential occupational choices under each treatment arm. The treatment determines which of these potential choices is realized, but does not determine to which principal stratum the unit belongs.}
\begin{table}[H]
\caption{Principal Strata.}
\begin{tabular}{|cc|c|c|}
\hline
$\boldsymbol{S_{i2}(0)}$ & $\boldsymbol{S_{i2}(1)}$ & Description & $\boldsymbol{V_{i}}$ \\ \hline
1&1 & Always Observed & AO \\
0&0 & Never Observed & NO \\
1&0 & Observed only in Control & OC \\
0&1 & Observed only in Treatment & OT \\ \hline
\end{tabular}
\begin{minipage}{0.81\linewidth \setstretch{0.75}}
{\scriptsize Notes: Principal Strata defined by the joint value of the potential selection indicators, $S(0)$ and $S(1)$. For example, consider the case where the treatment is a job training program aiming to increase workers' wages. The AO stratum comprises workers who are always employed, both with and without training. The NO stratum consists of workers who are always unemployed, regardless of whether they are treated or not. The OC stratum refers to workers who would find a job without the training but would be unemployed if they were treated. Conversely, the OT stratum includes workers who are unemployed in the absence of treatment but find a job if they receive it. }
\end{minipage}
\end{table}
To facilitate the exposition of the proposed methodology, I impose the following assumption on missingness:
\begin{assumption}
Missingness as an absorbing state
\begin{equation*}
S_{i1} = 0 \implies S_{i2} = 0.
\end{equation*}
\end{assumption}
Assumption (ref) implies that once a unit’s outcome is unobserved, it remains unobserved in all subsequent periods. Thus, all the units with $S_{i1} = 0$ will belong to the Never-Observed stratum at $t=2$. In contrast, units with $S_{i1} = 1$ can transition to any stratum in the second period. Beyond this, Assumptions (ref) and (ref) do not impose additional constraints on the proportion of the strata in the population, which may vary across groups, nor on the distribution of potential outcomes $Y(0)$ and $Y(1)$ across strata.
Assumption (ref) is plausible in settings with censoring due to death or similar events, such as clinical trials where patients cannot be observed after death, or educational contexts where students who drop out or graduate do not subsequently re-enroll.
In other settings, however, researchers may be reluctant to impose this assumption, as outcomes that are unobserved in the pre-treatment period may be observed at a later date.
This assumption imposes testable restrictions on the data, since $S_{i1}$ and $S_{i2}$ are observed for all units.
I impose Assumption (ref) to simplify the exposition of the methodology, but it is not required for the identification of the causal effects considered in this paper. Appendix (ref) presents the identification results of this paper without this assumption.
It is also important to note that this assumption does not correspond to restricting the sample to those observed in the first period. Allowing for missingness in the pre-treatment period is crucial for imputing the missing potential selection outcomes, which is central to the methodology developed in this paper. Restricting the sample to units with observed outcomes at baseline would discard valuable information contained in baseline selection. At the same time, as discussed in Appendix (ref), under Assumption (ref), the identification strategy in the main text targets causal effects only for units with observed outcomes at baseline. Nonetheless, this strategy can be extended to target causal effects regardless of whether the units were observed in the pre-treatment period, as shown in Appendix (ref).
Finally, I formalize the random sampling assumption in Assumption (ref).
\begin{assumption}
Random Sampling. $\{Y_{i1},Y_{i2}, S_{i1}, S_{i2},G_{i}\}_{i=1}^{N}$ are independent and identically distributed with joint distribution $P_{Y,G,S}$, and supp$(Y_{it}) \subset \mathbb{R}\cup\{*\}$, supp$(S_{it}) = \{0,1\}$, and supp$(G_{i}) = \{0,1\}$.
\end{assumption}
Assumption (ref) implies that the units are sampled from a larger population. Even when the outcome $Y_{it}$ is not observed for some units, the researcher knows that those units were included in the sample. I use the convention that $Y_{it} = *$ for units with $S_{it} = 0$. For example, wages are only observed for employed workers, but researchers also observe which workers are unemployed; quality of life is only observed for survivors, but researchers know which patients died; and academic performance is only observed for enrolled students, but researchers know which students dropped out. Accordingly, the expectations throughout the paper are meant to capture the sampling-based uncertainty abadie_sampling-based_2020.
\section{Identification of intensive margin effects}
This section presents a novel identification strategy to estimate the intensive margin effect in longitudinal settings.
Unlike previous literature on intensive margin effects,\footnote{See footnote (ref).} I do not assume treatment unconfoundedness. Instead, I allow the treatment to be confounded with unobservable unit-level characteristics. However, I impose restrictions on these unobservables' distribution over time. In what follows, I propose two causal estimands and discuss their identification.
\begin{estimand}
Average Treatment Effect on the Treated Always-Observed units ($\text{ATT}_{\text{AO}}$)
\begin{equation*}
ATT_{AO} = \mathbb{E}[Y_{i2}(1) - Y_{i2}(0) \mid G_{i} = 1, V_{i} = AO]
\end{equation*}
\end{estimand}
Estimand (ref) is the ATT for the subpopulation of Always-Observed units. This estimand captures the average effect for treated units whose outcome is observed under both treatment arms, and therefore can be interpreted as the policy's intensive margin effect. This estimand is especially relevant in settings where the researcher is interested in the causal effect net of any compositional changes.
In many applications where treatment effects are likely heterogeneous, policymakers may be interested not only in the average treatment effect but also in its distributional impact. For instance, consider the case of a job training program. A positive $\text{ATT}_{\text{AO}}$ indicates that, on average, the training increases wages for the workers employed both with and without the program. However, this average positive effect might be driven primarily by workers at the lower end of the wage distribution or, conversely, by those at the top. These contrasting scenarios carry significant policy implications, as the distributional effects of the training differ substantially. To fill this gap, I propose a new estimand that targets the treatment effect at various quantiles of the outcome distribution for the Always-Observed treated units.
\begin{estimand}
Quantile Treatment Effect on the Treated Always-Observed units ($\text{QTT}_{\text{AO}}$)
\begin{align*}
QTT_{AO}(q) &= Q_{Y_{2}(1) \mid G = 1, V_ = AO}^(q)-Q_{Y_{2}(0) \mid G = 1, V_ = AO}^(q), \\
Q_{Y}^(q) &= \inf\{y : F_{Y}(y) \geq q\}, \\
& q \in[0,1]
\end{align*}
where $F_{Y_{t}(w) \mid G = 1, V_{} = AO}$ denotes the cumulative distribution of the Potential Outcome $Y_{t}(w)$ at period $t$ for units that belong to the the treatment group, $G=1$, and to the Always-Observed stratum, $V = AO \equiv S_{2}(1) =S_{2}(0) = 1$.
\end{estimand}
Estimand (ref) contrasts the quantile $q$ of the treated potential outcome distribution for AO units in the treatment group with the same quantile of the control potential outcome distribution for these units. For example, $\text{QTT}_{\text{AO}}(0.5)$ captures the treatment effect at the median of the outcome distribution. This estimand captures a distributional effect: it does not require any rank-invariance assumption across potential outcomes and should not be confused with the median treatment effect imbens_causal_2015.
Both Estimands (ref) and (ref) are principal causal effects, in that they define causal effects for a subpopulation characterized by a principal stratum. Because principal stratum membership depends on joint potential outcomes, which are never simultaneously observed, this subpopulation is not identifiable without additional assumptions. Analogously to the IV setting, where individual compliance types cannot be directly observed, it may not be possible to know which units belong to the Always-Observed stratum. This feature may be viewed as a limitation of the proposed methodology. Nevertheless, because these estimands admit a clear interpretation as intensive margin effects, they are policy-relevant and practically meaningful for evaluating interventions and informing scale-up decisions. Finally, as in every principal stratification analysis, it is advisable to characterize the units belonging to the Always-Observed principal stratum mealli_refreshing_2012.
An alternative to targeting principal causal effects is given by bias-correction methods, that attempt to recover the ATT in the presence of sample selection. This estimand can be interpreted as “the treatment effect if all the treated units were observed” and therefore cannot be interpreted as the intensive-margin effect without further assumptions. Moreover, bias-correction methods have several limitations. First, they assume that outcomes are always well-defined, even if they are not directly observed zhang_likelihood-based_2009. In the job training example, bias-correction methods assume the existence of a `wage offer', allowing us to consider what the wage of an unemployed worker is. Second, they believe that the selection status is manipulable mealli_refreshing_2012, meaning that employment status could, in principle, be directly manipulated by the researcher.\footnote{When selection status is manipulable, one can define potential outcomes that can never be observed in the data and are `a priori counterfactuals', such that the potential wage without training for a worker who is only employed if trained. Another example of this type of potential outcome is given by the IV setting, where we can hypothesize about the potential outcome under treatment for a never-taker. See mealli_refreshing_2012 for a discussion.} Third, bias-correction methods often rely on exclusion restriction assumptions, requiring variables that affect employment but not wages lee_training_2009 and place restrictions on how potential outcomes are distributed across principal strata.
Next, I outline the assumptions required for partial identification of these estimands. First, I discuss the partial identification of both the $\text{ATT}_{\text{AO}}$ and the $\text{QTT}_{\text{AO}}(q)$ in a Changes-in-Changes setting. I then examine the identification of the principal strata proportions, which constitutes a key component of the proposed methodology, under alternative sets of assumptions.
\subsection{Identification of causal effects}
The existing literature estimating the effects for the Always-Observed units has relied on the assumption that treatment is unconfounded. However, this assumption is often unrealistic in causal models for panel data, where treatment assignment is typically non-random and likely correlated with unobserved factors ghanem_selection_2025, liu_practical_2024. To address this limitation, I reformulate the Horowitz-Manski-Lee bounds horowitz_identification_1995,lee_training_2009 by replacing the unconfoundedness assumption with the Changes-in-Changes (CiC) model as in athey_identification_2006, allowing the treatment assignment to be confounded with unobservable characteristics that affect the potential outcomes.
The choice of the CiC outcome modeling is meant to address several limitations in the canonical DiD research design. The canonical DiD may be sensitive to functional form roth_when_2023 and implicitly relies on additivity and linearity assumptions on the potential outcomes. To address these issues, athey_identification_2006 generalize the DiD identification strategy to a nonlinear and scale-free model, Changes-in-Changes. In the CiC model, individual and time `fixed effects' can influence the outcome in more flexible ways beyond the linear and additive fashion typically imposed in DiD. The CiC setup also offers additional advantages. First, shifting the focus from means to distributions enables the identification of distributional treatment effects, such as the $\text{QTT}_{\text{AO}}$. Second, it avoids extrapolation outside the outcome's support, ensuring that the counterfactual outcomes $Y(0)$ for treated units lie within the support of $Y$. On the other hand, because the method reconstructs the full distribution of $Y(0)$ for treated units, it requires assumptions that are not needed in the DiD framework, where identification concerns only the first moment of this distribution. As a result, the canonical DiD is not nested as a particular case of the CiC research design (see roth_when_2023 for an example). Nevertheless, the methodology developed in this paper can be adapted to the canonical DiD setting, as shown in Appendix (ref).
\begin{assumption}
Outcome model for Always-Observed units \\
\begin{equation*}
(Y_{it}(0) \mid V_{i} = AO )\quad =\quad m(\mathcal{U}_{it}, t),
\end{equation*}
where $m(u,t)$ is strictly increasing in $u$ $\forall t$, and $\mathcal{U}_{it}$ is an unobservable scalar for unit $i$ at time $t$ with constant distribution over time within groups,
\begin{equation*}
\mathcal{U}_{i1} \mid G_{i}, V_{i} = AO \sim \mathcal{U}_{i2} \mid G_{i}, V_{i} = AO.
\end{equation*}
\end{assumption}
Assumption (ref) is the principal counterpart to assumptions 3.1-3.4 in athey_identification_2006. It models the untreated potential outcome as a function of an unobservable individual characteristic, $\mathcal{U}_{it}$. This model embeds several restrictions on the unobservable, its distribution, and the outcome function $m(\cdot)$. These assumptions do not restrict the data in any way and, thus, are not testable.
This model assumes that all the individual characteristics affecting the outcome can be captured in a single index, $\mathcal{U}_{it}$. The function $m(\cdot)$ being strictly increasing implies that units with higher values of the unobservable also have higher outcomes. This assumption is intuitive when the unobservable is interpreted as an individual characteristic such as ability or productivity. It holds by construction in additively separable models, such as the two-way fixed effects model. However, it also accommodates a broad range of non-additive nonlinear outcome functions.
The distribution of the unobservable may differ between the treatment and control groups. However, it remains constant over time within each group. This assumption is crucial, as it allows interpreting changes in the outcome in the control group as changes in the function $m(\cdot)$. Once the trend in $m(\cdot)$ is estimated, the missing potential outcome for the treatment group can be identified.
\begin{example}
(Two-way fixed effects model). Assume a separable additive model. Let $m(\mathcal{U}_{it},t) = \mathcal{U}_{it} + \lambda_{t}$, and assume $\mathcal{U}_{it} = \alpha_{i} + \varepsilon_{it}$, where $\varepsilon_{it}$ is random noise. Then, the untreated potential outcome for AO units can be expressed as:
\begin{equation*}
Y_{it}(0) = \alpha_{i} + \lambda_{t} + \varepsilon_{it}.
\end{equation*}
This example illustrates that the standard two-way fixed effect model can be seen as a special case of the CiC outcome model.
\end{example}
\begin{proposition}
Let $Y_{it}(1)$ and $Y_{it}(0)$ be continuous with compact
support. Furthermore, let the support of $Y_{it}(0)$ for the treatment group be contained in the support of $Y_{it}(0)$ for the control group. Then, if Assumptions (ref), (ref), (ref) and (ref) hold, then $\Lambda^{LB}(q)$ and $\Lambda^{UB}(q)$ are
lower and upper bounds for the Quantile Treatment Effect on the Treated Always-Observed units ($\text{QTT}_{\text{AO}}(q)$), where:
\begin{align*}
\Lambda^{LB}(q) &= Q_{Y_2 \mid G=1, S_{2} = 1}(q\pi_1) - Q_{Y_{2} \mid G=0, S_2 = 1}\left(F_{Y_{1}\mid G=0, S_2 = 1}\left(Q_{Y_{1}\mid G=1,S_2 = 1}(q\pi_1 + 1 - \pi_1)\right) + 1 - \pi_0 \right), \\
\Lambda^{UB}(q) &= Q_{Y_2 \mid G=1, S_{2} = 1}(q\pi_1 + 1 - \pi_1) - Q_{Y_{2} \mid G=0,S_2 = 1}\left(F_{Y_{1}\mid G=0, S_2 = 1}\left(Q_{Y_{1}\mid G=1,S=1}(q\pi_1)\right)-(1-\pi_0)\right), \\
\pi_{1} & = Pr(S_{i2}(0) = 1 \mid G_{i} = 1, S_{i2}(1) = 1),\\
\pi_{0} & = Pr(S_{i2}(1) = 1 \mid G_{i} = 0, S_{i2}(0) = 1), \\
F_{Y}(y) & := Pr(Y \leq y) \\
Q_{Y}(q) & := \inf\{y: F_{Y}(y) \geq q\}
\end{align*}
provided that
\begin{align*}
F_{Y_{1}\mid G=0, S_2 = 1}\left(Q_{Y_{1}\mid G=1,S_2 = 1}(q\pi_1 + 1 - \pi_1)\right) &\leq \pi_0 \\
F_{Y_1 \mid G=0, S_2 = 1}\left(Q_{Y_1\mid G=1, S_2 = 0} (q\pi_1)\right) &\geq 1 - \pi_0.
\end{align*}
If $F_{Y_{1}\mid G=0, S_2 = 1}\left(Q_{Y_{1}\mid G=1,S_2 = 1}(q\pi_1 + 1 - \pi_1)\right) > \pi_0$, then
\begin{equation*}
\Lambda^{LB}(q) = Q_{Y_2 \mid G=1, S_{2} = 1}(q\pi_1) - Q_{Y_{2} \mid G=0, S_2 = 1}\left(1\right).
\end{equation*}
If $F_{Y_1 \mid G=0, S_2 = 1}\left(Q_{Y_1\mid G=1, S_2 = 0} (q\pi_1)\right) < 1 - \pi_0$, then
\begin{equation*}
\Lambda^{UB}(q) = Q_{Y_2 \mid G=1, S_{2} = 1}(q\pi_1 + 1 - \pi_1) - Q_{Y_{2} \mid G=0,S_2 = 1}\left(0\right)
\end{equation*}
Proof: See Appendix (ref).
\end{proposition}
\begin{remark}
The two conditions, $$F_{Y_{1}\mid G=0, S_2 = 1}\left(Q_{Y_{1}\mid G=1,S_2 = 1}(q\pi_1 + 1 - \pi_1)\right) \leq \pi_0$$ and $$F_{Y_1 \mid G=0, S_2 = 1}\left(Q_{Y_1\mid G=1, S_2 = 0} (q\pi_1)\right) \geq 1 - \pi_0,$$ ensure that extrapolation is not required outside the support of $Y_{2} \mid G=0$. When these conditions do not hold, the lower bound for $Q_{Y_2(0) \mid G=1, V=AO}(q)$ is given by the infimum of the support of $Y_2 \mid G=0$, and the upper bound by its supremum. This implies that, for extreme quantiles, the bounds on $\text{QTT}_{\text{AO}}(q)$ may be uninformative. Moreover, the smaller the proportion of Always-Observed units, $\pi_1$ and $\pi_0$, the larger the region where these bounds may be uninformative. Note that when $\pi_0 = 1$, both conditions are always satisfied.
\end{remark}
\begin{remark}
Once the distributions of $Y_{2}(1)_{ \mid G = 1, V_{} = AO}$ and $Y_{2}(0)_{ \mid G = 1, V_{} = AO}$ are partially identified, the bounds for the $\text{ATT}_{\text{AO}}$ follow immediately:
$$\int_{0}^{1} \Lambda^{LB}(q)dq \leq \text{ATT}_{\text{AO}} \leq \int_{0}^{1} \Lambda^{UB}(q)dq.$$
\end{remark}
\begin{remark}
Proposition (ref) builds on continuity and support assumptions. These results can be extended to cases where the support of $Y_{it} \mid G_i=1$ does not fully overlap with that of $Y_{it}\mid G_{i} = 0$; see athey_identification_2006 for a discussion. Appendix (ref) explores the identification for discrete outcomes. However, the identification is unsuitable when $Y$ is mixed or has non-compact support.
\end{remark}
\begin{remark}
When the OT and OC strata are empty, it follows that $\pi_{1} = \pi_{0} = 1$, $\Lambda^{LB} = \Lambda^{UB}$, and the $\text{QTT}_{\text{AO}}(q)$ is point identified. The intuition behind this result is that all the units observed in the post-treatment period belong to the AO stratum. In this scenario, $S$ becomes an indicator of membership in the AO stratum.
\end{remark}
Proposition (ref) describes a trimming procedure to bound the $\text{QTT}_{\text{AO}}(q)$. It builds on horowitz_identification_1995, who derive bounds for the components of a mixture distribution when the mixture proportions are known—here, the proportions of Always-Observed units in each group. I combine this with results from athey_identification_2006, who derive the counterfactual distribution of $Y_{2}(0)$ for treated units. Once the distributions of $Y_{2}(1)$ and $Y_{2}(0)$ for the treated AO units are bounded, the identification of the $\text{QTT}_{\text{AO}}(q)$ and the $\text{ATT}_{\text{AO}}$ follows.
The intuition behind Proposition (ref) is as follows. Suppose that the proportion of AO units among treated units with observed outcomes, $\pi_1$, is known. For example, if $\pi_1 = 0.9$, the observed distribution for treated units trimmed at the 90% percentile constitutes a lower bound for the real latent distribution for these 90% AO treated units. Similarly, the top 90% of the observed distribution constitutes an upper bound. Analogously, the distribution of outcomes for control AO units can be partially identified using the corresponding proportion, $\pi_0$.
Once the outcome distributions for AO units are partially identified, they can be combined with the results from athey_identification_2006 to partially identify the $\text{QTT}_{\text{AO}}(q)$. Under the CiC model for the AO units, any change in the control group’s outcome distribution must come from changes in the outcome mapping $m(\cdot)$, as the distribution of $\mathcal{U}_{it}$ is assumed to be stable over time within groups. This lets me use the control group to recover how $m(\cdot)$ shifts over time. I can then apply the same shift to the pre-treatment distribution of $\mathcal{U}_{it}$ in the treatment group, inferred from their pre-treatment distribution of the outcome, to construct the counterfactual post-treatment distribution. Figure (ref) provides graphical intuition behind Proposition (ref).
\begin{figure}[h]
\caption{Graphical intuition on Proposition (ref).}
\vskip\baselineskip
\raisebox{0.1cm}{$\Lambda^{LB}$}
\fbox{
\begin{minipage}{0.98\textwidth}
\begin{minipage}{0.48\textwidth}
\end{minipage}
\begin{minipage}{0.48\textwidth}
\end{minipage}
\end{minipage}
}
\vskip\baselineskip
\raisebox{0.1cm}{$\Lambda^{UB}$}
\fbox{
\begin{minipage}{0.98\textwidth}
\begin{minipage}{0.48\textwidth}
\end{minipage}
\begin{minipage}{0.48\textwidth}
\end{minipage}
\end{minipage}
}
\vskip\baselineskip
\begin{minipage}{0.99\linewidth \setstretch{0.75}}
{\scriptsize Notes: This figure plots a hypothetical distribution of $Y_2 \mid S_{2} =1$. The distributions on the left correspond to the Treated group ($G=1$) and the one on the right to the Control group ($G=0$). The shaded areas correspond to the proportion of strata in the given group: black corresponds to the AO units ($\pi_{1}$ on the left and $\pi_{0}$ on the right), green corresponds to the OT stratum and blue to the OC. The vertical black lines denote where a given quantile $q$ lies in the observed distribution. The vertical red lines denotes where the same quantile $q$ lies in the trimmed distributions, used in Proposition (ref). This figure also illustrates Remark (ref): as $\pi_{1}$ and $\pi_{0}$ go to 1, the black shadowed areas expand, and the trimmed (red) quantiles converge to those of the entire distribution (black).
}
\end{minipage}
\end{figure}
\subsection{Identification of principal strata proportions}
A pivotal factor of Proposition (ref) is the proportion of AO units in each group, $\pi_{1}$ and $\pi_{0}$. Next, I impose additional assumptions on the selection mechanism that enable the identification of these proportions.
\begin{assumption}
Monotonicity
\begin{align*}
S_{i2}(1) \geq S_{i2}(0) \quad & \forall i \quad Positive Monotonicity \\
& \text{or} \\
S_{i2}(1) \leq S_{i2}(0) \quad & \forall i \quad \text{Negative Monotonicity}
\end{align*}
\end{assumption}
Assumption (ref) implies that treatment can affect selection only in ‘one direction’ for all the units and rules out one of the four principal strata. Positive monotonicity excludes the OC stratum. This implies that the control units observed in both periods consist solely of Always-Observed units: $S_{i2}(1) \geq S_{i2}(0) \implies Pr(S_{i2}(1) = 0\mid S_{i2}(0) = 1) = 0 \implies \pi_{0} = 1$. Conversely, negative monotonicity rules out the OT stratum, assuming that all the treated units observed in both periods belong to the AO stratum. While restrictive, this assumption is standard in the literature and fundamental to ensure point identification of principal strata proportions. It is analogous to the monotonicity assumption that rules out the existence of defiers in IV settings. In section (ref), I propose a novel methodology to relax this assumption when different sources of selection are available in the data.
The preference for positive or negative monotonicity depends on the context. For example, consider a researcher estimating the effect of job training on wages. It is reasonable to assume that a worker employed without training would also be employed with it. This assumption makes positive monotonicity the preferred assumption. Conversely, take the example of a researcher studying the effect of increased tuition fees on university students' performance. In this case, sample selection arises as some students drop out. Under the assumption that higher education costs increase dropouts, it is reasonable to conclude that if a student remains enrolled despite the fee increase, they would also stay enrolled without it. In this case, negative monotonicity is preferred.
When positive (negative) monotonicity is assumed, $\pi_{0}$ ($\pi_{1}$) is equal to one, and only the treatment (control) group needs to be trimmed. I adopt the CiC athey_identification_2006 methodology to estimate the proportion of AO units in this group. This identification strategy assumes that an unknown function of individual unobservables determines the potential selection outcomes. This approach has some key advantages over the canonical DiD framework when applied to selection: it allows the identification of the missing potential selection outcomes for both groups. This is crucial when trimming the control group under negative monotonicity. Furthermore, it ensures that all the estimated probabilities lie in the unit interval.
\begin{assumption}
Selection Model \\
Under Positive Monotonicity:
\begin{equation*}
\begin{aligned}
S_{it}(0) = h^{0}(U_{it},t) .
\end{aligned}
\end{equation*}
Under Negative Monotonicity:
\begin{equation*}
\\
S_{it}(1) = h^{1}(U_{it},t),
\end{equation*}
where $U_{it}$ is an unobservable scalar for unit $i$ at time $t$ and $h^{w}(u,t)$ is non decreasing function in $u$ $\forall$ $t \in\{1,2\}$, $w = \{0,1\}$. \\
The unobservable $U_{}$ is continuously distributed and has the same compact support in both groups. Its distribution is constant over time within groups
\begin{equation*}
U_{i1} \mid G_{i} \sim U_{i2} \mid G_{i} .
\end{equation*}
Additionally, given the realized selection outcome, the distribution of $U$ is independent of the group in a given time period.
\begin{equation*}
U_{it} \perp \! \! \! \perp G_{i} \mid S_{it}
\end{equation*}
\end{assumption}
Assumption (ref) models the selection mechanism as a function of an unobservable individual characteristic, $U_{it}$. Depending on the monotonicity direction in Assumption (ref), a model for $S_{it}(0)$ or $S_{it}(1)$ needs to be assumed. Under Positive (Negative) Monotonicity, the control (treatment) group is composed solely of AO units, and the treatment (control) group is a mixture of AO and OT (OC) units, and therefore $\mathbb{E}[S_{i2}(0)\mid G_{i} =1, S_2 = 1]$ $\left(\mathbb{E}[S_{i2}(1)\mid G_{i} = 0, S_2 = 1]\right)$ needs to be imputed. This selection model shares significant similarities with the outcome model in Assumption (ref). The main components of both models are the unobservable scalars, $U$ in the case of selection, and $\mathcal{U}$ in the case of the outcome. While they may be independent, the potential correlation between the unobserved factors influencing selection and those determining the outcome is central to the non-ignorable sample selection problem.
There are three noteworthy differences between the outcome model in Assumption (ref) and the selection model in Assumption (ref), albeit both share the general intuition behind a CiC model. First, depending on the direction of the monotonicity assumption, imputing the missing potential selection outcomes may be necessary for the treated or control units. This requires specifying functions for both potential selection outcomes, $S(1)$ in case negative monotonicity holds, and $S(0)$ in case positive monotonicity holds. In contrast, only the missing potential outcome $Y(0)$ for treated units needs to be imputed. Thus, modeling $Y_{it}(1)$ is not required. Second, the unknown function $m(\cdot)$ is assumed to be strictly increasing in the unobservable. Since $S$ is binary, the equivalent assumption cannot be made for $h^{0}(\cdot)$ and $h^{1}(\cdot)$. Nevertheless, the non-decreasing assumption ensures that observed units will not have lower values of the unobservable than the units that leave the sample. Third, Assumption (ref) states that $ U_{it} \perp \! \! \! \perp G_{i} \mid S_{it}$. While the corresponding assumption, $ \mathcal{U}_{it} \perp \! \! \! \perp G_{i} \mid Y_{i}, V_{i} = AO$, is not explicitly made, it is embedded in Assumption (ref). Since the function $m(u,t)$ is strictly increasing in $u$, it follows that $\mathcal{U}_{it} = m^{-1}(Y_{it}, t)$. Thus, when conditioning on $Y$ and $t$, the distribution of $\mathcal{U}$ becomes degenerate and, therefore, identical for both groups. As a result, the conditional independence assumption is trivially satisfied.
\begin{lemma}
If Assumptions (ref), (ref), (ref), (ref), and (ref) hold, then the missing potential selection outcomes are given by:
Under Positive Monotonicity,
\begin{align*}
\mathbb{E}[S_{i2}(0) \mid G_{i} = 1] = \mathbb{E}[S_{i1} \mid G_{i} = 1]\frac{\mathbb{E}[S_{i2} \mid G_{i} =0]}{\mathbb{E}[S_{i1} \mid G_{i} = 0]}.
\end{align*}
Under Negative Monotonicity,
\begin{equation*}
\mathbb{E}[S_{i2}(1) \mid G_{i} = 0] = \mathbb{E}[S_{i1} \mid G_{i} = 0] \frac{\mathbb{E}[S_{i2} \mid G_{i} = 1]}{\mathbb{E}[S_{i1}\mid G_{i} = 1]}.
\end{equation*}
\end{lemma}
\begin{remark}
Both $ \mathbb{E}[S_{i2}(0) \mid G_{i} = 1]$ and $\mathbb{E}[S_{i2}(1) \mid G_{i} = 0]$ will always lie in the unit interval. They will always be well-defined, except for the case when $\mathbb{E}[S_{i1} \mid G=g]=0$, which would mean that no unit is observed in the pre-treatment period in group $g$.
\end{remark}
Lemma (ref) is equivalent to Theorem 4.2 in athey_identification_2006, and the proof can be found in athey_identification_2006. Under the CiC model for selection, the missing selection outcome in the post-treatment period for group $g$ is given by their baseline selection in the pre-treatment period scaled by the proportional change for the other group. Table (ref) provides a numerical example for illustration purposes.
\begin{table}[H]
\caption{Numerical example of Lemma (ref)}
\begin{tabular}{|c|c|c|c|}
\hline
Group ($G$) & $\mathbb{E}[S_{1}]$ & $\mathbb{E}[S_{2}(G)]$ & $\mathbb{E}[S_{2}(1-G)]$ \\ \hline
Treatment $(G=1)$ & 0.9 & 0.55 & 0.45 \\
Control $(G = 0)$ & 0.7 & 0.35 & 0.43 \\ \hline
\end{tabular}
\begin{minipage}{0.81\linewidth \setstretch{0.75}}
{\scriptsize Notes: This table illustrates Lemma (ref) with a numerical example. Columns 2 and 3 display the observed mean in the selection indicator in each of the groups. Column 2 corresponds to the pre-treatment period and column 3 to the post-treatment period. The fourth column presents the missing potential selection outcome for each of the groups, imputed using Lemma (ref).}
\end{minipage}
\end{table}
\begin{remark}
The treatment effect on selection, $\mathbb{E}[S_{i2}(1) - S_{i2}(0) \mid G_{i} = g]$ will have the same sign for both groups $g \in \{0,1\}$. This is consistent with the monotonicity assumption. Furthermore, these two effects will be the same if the expected selection in the pre-treatment period is the same across groups, i.e., $\mathbb{E}[S_{i1} \mid G_{i}= 1] = \mathbb{E}[S_{i1} \mid G_{i} = 0]$. This aligns with the selection model in Assumption (ref) as
\begin{equation*}
\mathbb{E}[S_{i1} \mid G_{i}= 1] = \mathbb{E}[S_{i1} \mid G_{i} = 0] \iff U_{it} \perp\!\!\!\!\perp G_{i} \implies \mathbb{E}[S_{i2}(1) - S_{i2}(0) ] \perp\!\!\!\!\perp G_{i}
\end{equation*}
Proof: See Appendix (ref)
\end{remark}
\begin{proposition}
Under Assumptions (ref), (ref), (ref), (ref), and (ref), the proportion of Always-Observed units in the treatment group ($\pi_{1}$) and control group ($\pi_{0}$) are identified as follows:
\begin{itemize}
• If Positive Monotonicity holds:
\begin{align*}
\pi_{0} &= 1 \\
\pi_{1} & = \frac{\mathbb{E}[S_{i2}(0) \mid G_{i} = 1]}{\mathbb{E}[S_{i2} \mid G_{i} = 1]} = \frac{\mathbb{E}[S_{i1} \mid G_{i} = 1]}{\mathbb{E}[S_{i2} \mid G_{i} = 1]}\frac{\mathbb{E}[S_{i2} \mid G_{i} = 0]}{\mathbb{E}[S_{i1} \mid G_{i} = 0]} \in [0 ,1]
\end{align*}
• If Negative Monotonicity holds:
\begin{align*}
\pi_{0} &= \frac{\mathbb{E}[S_{i2}(1) \mid G_{i} = 0]}{\mathbb{E}[S_{i2} \mid G_{i} = 0]} = \frac{\mathbb{E}[S_{i1}\mid G_{i} = 0]}{\mathbb{E}[S_{i2} \mid G_{i} = 0 ]}\frac{\mathbb{E}[S_{i2} \mid G_{i} = 1]}{\mathbb{E}[S_{i1}\mid G_{i} = 1]} \in [0,1] \\
\pi_{1} &= 1
\end{align*}
\end{itemize}
Proof: See Appendix (ref).
\end{proposition}
\subsection{Relaxing the monotonicity assumption}
Assumption (ref) can be restrictive, as it requires the treatment to affect selection in only one direction for all the units, even when selection may respond to treatment through multiple channels. Similarly, Assumption (ref) implies that all unobservables affecting selection can be summarized in a single scalar, ruling out different unobservables affecting selection in distinct ways. When multiple sources of selection are observed, the monotonicity assumption can be relaxed by allowing each source to exhibit monotonicity independently, potentially with different signs. For instance, consider a scholarship program designed to improve academic performance. Such a scholarship may affect selection through two distinct mechanisms: reducing dropout rates (positive monotonicity) while increasing graduation rates (negative monotonicity).
Suppose that there are $J$ different sources of sample selection, where $s_{it}^{j}$ is a binary variable equal to 0 if unit $i$'s outcome is not observed because of source $j$ and 1 otherwise. In the previous example, $J=2$, with $s_{it}^{1}$ equal to 0 if the student $i$ drops out at period $t$, and $s_{it}^{2} = 0$ if the student graduates. By definition, these sources are mutually exclusive, meaning that when a unit leaves the sample, it is due to a specific source $j$. For instance, students may leave the sample either because they graduated or because they dropped out before completing the degree, but both events cannot happen simultaneously. This framework allows the expression of the generic selection indicator as the product of the different sources of selection: $S_{it} = \prod_{j=1}^{J}s_{it}^{j}$, where the mutual exclusivity of the sources implies that $\sum_{j = 1}^{J}s_{it}^{j} \in \{J-1, J\}$ $ \forall i,t$. This notation extends to potential selection indicators as well: $S_{it}(w) = \prod_{j=1}^{J}s_{it}^{j}(w)$.
The four strata defined in Table (ref) exist for each source $j$. I can define analogously a variable $V_{i}^{j}$ indicating the principal stratum of unit $i$ defined by the joint values of the potential selection indicator of the source $j$, $(s_{i2}^{j}(1), s_{i2}^{j}(0))$.
\begin{assumption}
Source-specific monotonicity \\
For any source $j\in\{1,..., J \}$:
\begin{align*}
s_{i2}^{j}(1) \geq s_{i2}^{j}(0) \quad & \forall i \quad \text{Positive Monotonicity} \\
& \text{or} \\
s_{i2}^{j}(1) \leq s_{i2}^{j}(0) \quad & \forall i \quad \text{Negative Monotonicity}
\end{align*}
\end{assumption}
Assumption (ref) is analogous to Assumption (ref) defined for each specific source of selection. It rules out the existence of one stratum for each source $j$ ($V_{i}^{j} = OC$ if positive monotonicity is assumed and $V_{i}^{j} = OT$ if negative monotonicity is assumed). Similarly, I can formulate an analogous, source-specific Assumption (ref):
\begin{assumption}
Source-specific Selection Model \\
For any source $j\in\{1,...,J\}$:\\
Under Positive Monotonicity,
\begin{equation*}
\begin{aligned}
s_{it}^{j}(0) = h_j^{0}(u_{it}^{j},t) .
\end{aligned}
\end{equation*}
Under Negative Monotonicity,
\begin{equation*}
s_{it}^{j}(1) = h_j^{1}(u_{it}^{j},t),
\end{equation*}
where $u_{it}^{j}$ is an unobservable scalar for unit $i$ at time $t$ and $h_j^{w}(u,t)$ is non decreasing function in $u$ $\forall$ $t \in\{1,2\}$, $w = \{0,1\}$. \\
The unobservable $u_{}^{j}$ is continuously distributed and has the same compact support in both groups. Its distribution is constant over time within groups
\begin{equation*}
u_{i1}^{j} \mid G_{i} \sim u_{i2}^{j} \mid G_{i} .
\end{equation*}
Additionally, given the realized selection outcome, the distribution of $u^{j}$ is independent of the group in a given time period.
\begin{equation*}
u^{j}_{it} \perp \! \! \! \perp G_{i} \mid s_{it}^{j}
\end{equation*}
\end{assumption}
Assumption (ref) states that unobservables affecting selection can be summarized by a single scalar that affects the eventual selection in a monotonic way. Assumption (ref) relaxes this assumption by allowing for multiple unobservable scalars that may affect selection in different directions. Going back to the college example, one can think of two distinct unobservables affecting selection: ability and motivation. Higher ability increases the likelihood of graduation, whereas lower motivation increases the likelihood of dropping out. Because students with high ability and high motivation may leave the sample through graduation, while those with low ability and low motivation may leave through dropout, it is difficult to summarize these distinct selection forces using a single index.
Let $\mathcal{X}_{j}$ denote the set of units that belong to stratum $X$ according to source $j$. For instance, $\mathcal{AO}_{j} = \{i : V_{i}^{j} = AO\} \equiv \{i:s_{i2}^{j}(1) = 1, s_{i2}^{j}(0) = 1\}$.
\begin{assumption}
No intersection of OC and OT strata. \\
\begin{equation*}
\mathcal{OT}_{j} \cap \mathcal{OC}_{k} = \varnothing \quad \forall j,k \in \{1,...,J\}
\end{equation*}
\end{assumption}
Assumption (ref) rules out the scenario where a unit belongs to the OT stratum defined by source $j$ and to the OC stratum defined by a different source. Consequently, units that belong to the OC or OT strata according to source $j$ can only belong to the AO stratum according to all the other sources. Formally, $s_{i2}^{j}(1) \neq s_{i2}^{j}(0) \implies s_{i2}^{k}(1) = s_{i2}^{k}(0) = 1$ $\forall k \neq j$.
For instance, consider a student who drops out in the control group but remains enrolled if treated ($s_{i2}^{1}(0) = 0,s_{i2}^{1}(1) = 1$). Since sources are mutually exclusive and $s_{i2}^{1}(0) = 0$, it must be true that this unit does not graduate in the control group, $s_{i2}^{2}(0) = 1$. Assumption (ref) rules out the possibility that this student, who drops out in the control group, graduates in the treatment group, i.e., assumes that $s_{i2}^{2}(1) = 1$. Hence, given that the student belongs to the OT stratum defined by the joint values of the dropout source, $(s_{i2}^{1}(0) , s_{i2}^{1}(1) ) = (0, 1)$, they must belong to the AO stratum defined by the joint values of the graduation source, $(s_{i2}^{2}(1),s_{i2}^{2}(0)) = (1,1)$.
\begin{remark}
If sources are mutually exclusive, and Assumption (ref) holds, then Assumptions (ref) and (ref) also hold, with all the sources of selection having the same sign of monotonicity. On the other hand, when Assumptions (ref) and (ref) hold, then:
\begin{itemize}
• If all sources exhibit the same sign of monotonicity, then Assumption (ref) also holds.
• If some sources exhibit positive monotonicity while others exhibit negative monotonicity, then Assumption (ref) is violated.
\end{itemize}
\end{remark}
While I acknowledge that Assumptions (ref) and (ref) are restrictive, they relax Assumption (ref) as noted in Remark (ref). This relaxation, which allows the existence of the four principal strata, comes at no cost, as the proportion of AO units in each group remains point-identified. This section can therefore be viewed as a generalization of Section (ref): when Assumptions (ref) and (ref) hold, then Assumptions (ref), (ref), and (ref) also hold; when Assumptions (ref) and (ref) fail, then Assumptions (ref), (ref), and (ref) may still hold. A significant limitation, though, is that different sources must exist and be identifiable within the data, which is not always guaranteed. Intuitively, if selection is driven by multiple channels operating in different directions, the assumptions in Section (ref) are violated. If these channels manifest themselves in distinct, observable sources of selection, the assumptions in this section may still be satisfied. If not, these assumptions are violated as well.
\begin{lemma}
If assumptions (ref) and (ref) hold, the four principal strata defined by the joint values of $(S_{i2}(1),S_{i2}(0))$ that partition the population are given by:
\begin{align*}
\mathcal{AO} = \bigcap\limits_{j} \mathcal{AO}_{j} \quad ; \quad
\mathcal{NO} = \bigcup\limits_{j} \mathcal{NO}_{j} \quad ; \quad
\mathcal{OT} = \bigcup\limits_{j} \mathcal{OT}_{j} \quad ; \quad
\mathcal{OC} = \bigcup\limits_{j} \mathcal{OC}_{j}
\end{align*}
Proof: See Appendix (ref).
\end{lemma}
Lemma (ref) states that to belong to the AO stratum, a unit must belong to the AO stratum defined by the potential selection indicators of all sources. Conversely, if a unit belongs to the NO, OT, or OC stratum for any source, it will also belong to the corresponding principal stratum defined by the overall selection indicator $S$. Table (ref) provides an example with two different sources of selection.
\begin{proposition}
Let $J^{+}$ denote the set of sources for which positive monotonicity holds and $J^{-}$ the set of sources for which negative monotonicity holds. Under Assumptions (ref), (ref), (ref), (ref), (ref), and (ref), the proportion of Always-Observed units in the control group ($\pi_{0}$) and the treatment group ($\pi_{1}$) are identified as follows:
\begin{align*}
\pi_{0} = \frac{1}{\mathbb{E}[S_{i2} \mid G_{i} = 0]}\left(1 - \sum_{j \in J^{+}}\left(1 - \mathbb{E}[s_{i2}^{j} \mid G_{i} = 0]\right) - \sum_{j \in J^{-}}\left(1 - \mathbb{E}[s_{i1}^{j} \mid G_{i} = 0]\frac{\mathbb{E}[s_{i2}^{j} \mid G_{i} = 1]}{\mathbb{E}[s_{i1}^{j} \mid G_{i} = 1]}\right)\right) \\
\pi_{1} = \frac{1}{\mathbb{E}[S_{i2} \mid G_{i} = 1]}\left(1 - \sum_{j \in J^{+}}\left(1 - \mathbb{E}[s_{i1}^{j} \mid G_{i} = 1]\frac{\mathbb{E}[s_{i2}^{j} \mid G_{i} = 0]}{\mathbb{E}[s_{i1}^{j} \mid G_{i} = 0]} \right) - \sum_{j \in J^{-}}\left(1 - \mathbb{E}[s_{i2}^{j} \mid G_{i} = 1]\right)\right)
\end{align*}
Proof: See Appendix (ref).
\end{proposition}
Proposition (ref) identifies the proportion of always observed in both the treatment and control groups. By Assumption (ref), some sources exhibit positive monotonicity, while others exhibit negative monotonicity. Following the same logic as in Proposition (ref), the former can identify the AO units in the treatment group, and the latter can identify the AO units in the control group. Appendix (ref) shows how Proposition (ref) is a specific case of Proposition (ref) with only one source of selection, $J = 1$.
\section{Estimation and inference}
This section proposes estimators for the bounds of the $\text{ATT}_{\text{AO}}$ and $\text{QTT}_{\text{AO}}(q)$. I show that these estimators are consistent and asymptotically normal and discuss how to construct confidence intervals for the estimated bounds. The estimators are constructed using the sample analogs of the objects in Propositions (ref) and (ref). For simplicity and without loss of generality, this section considers the case with two different sources of sample selection. For the first one, $j =1$, positive monotonicity holds. For the second one, $j = 2$, negative monotonicity holds. All the summations in this section are over the entire sample.
\subsection{Estimation}
Let $N$ denote the sample size, $N_{0}$ the number of units in the control group, and $N_{1}$ the number of units in the treatment group. The estimators of the proportions of Always-Observed units defined in Proposition (ref) are given by:
\begin{align}
\hat \pi_{0} &= \frac{1}{\frac{\sum_{i}S_{i2}(1-G_{i})}{N_{0}}}\left(1 - \left(1 - \frac{\sum_{i}s_{i2}^{1}(1 - G_{i})}{N_{0}}\right) - \left(1 - \frac{\sum_{i}s_{i1}^{2}(1-G_{i})}{N_{0}}\frac{\sum_{i}s_{i2}^{2}G_{i}}{\sum_{i}s_{i1}^{2}G_{i}}\right)\right) \\
\hat \pi_{1} & = \frac{1}{\frac{\sum_{i}S_{i2}G_{i}}{N_{1}}}\left(1 - \left(1 - \frac{\sum_{i}s_{i1}^{1}G_{i}}{N_{1}}\frac{\sum_{i}s_{i2}^{1}(1-G_{i})}{\sum_{i}s_{i1}^{1}(1-G_{i})}\right) - \left(1 - \frac{\sum_{i}s_{i2}^{2}G_{i}}{N_{1}}\right)\right)
\end{align}
where I have substituted the expectations in Proposition (ref) with their sample analogs.
Given the estimated proportions, $\hat \pi_{0}$ and $\hat\pi_{1}$, the estimators for the bounds of the $\text{QTT}_{\text{AO}}(q)$ defined in Proposition (ref) are given by:
\begin{align}
\widehat{ \Lambda^{LB}}(q) &= \widehat Q_{Y_{2} \mid G = 1, S_{2} = 1}^(q\hat\pi_{1}) - \widehat Q_{Y_{2} \mid G = 0, S_{2} = 1}^\left(\hat q_{LB}^{*}\right), \\
\hat{ q}_{LB}^{*} &= \min\left\{\widehat F_{Y_{1} \mid G = 0, S_{2} = 1}^\left(\widehat Q_{Y_{1} \mid G=1, S_{2} = 1}^(q\hat\pi_{1}+1 - \hat\pi_{1}) \right) + 1 - \hat\pi_{0} ,1\right\}, \\
\widehat{\Lambda^{UB}}(q)& = \widehat Q_{Y_{2} \mid G = 1 , S_{2} = 1}^(q\hat\pi_{1}+1-\hat\pi_{1}) -\widehat Q_{Y_{2} \mid G = 0, S_{2} = 1}^\left( \hat q_{UB}^{*}\right), \\
\hat q_{UB}^{*} &= \max\left\{\widehat F_{Y_{1} \mid G = 0, S_{2} = 1}^\left(\widehat Q_{Y_{1} \mid G=1, S_{2} = 1}^(q\hat \pi_{1}) \right)-(1 - \hat \pi_{0}), 0\right\},
\end{align}
where $\widehat F_{Y}$ denotes the empirical distribution of $Y$. For instance,
\begin{align*}
\widehat F_{Y_{1} \mid G = 0, S_{2} = 1}(y) &= \frac{\sum_{i}S_{i2}(1-G_{i})\mathbbm{1}(Y_{i1}\leq y)}{\sum_{i}S_{i2}(1-G_{i})} \\
\widehat Q_{Y_{2} \mid G=0, S_{2} = 1}^(q) &= \inf\left\{y: \widehat F_{Y_{2} \mid G=0, S_{2} = 1}(y) \geq q\right\}.
\end{align*}
The bounds for the $\text{ATT}_{\text{AO}}$ can be estimated under the CiC framework using the estimated bounds for the $\text{QTT}_{\text{AO}}$, as stated in Remark (ref):
\begin{equation}
\widehat{ATT}_{AO} \in \left[\int_{0}^{1} \widehat{ \Lambda^{LB}}(q) dq \quad , \quad \int_{0}^{1}\widehat{\Lambda^{UB}}(q) dq \right]
\end{equation}
\subsection{Asymptotic Normality}
The next propositions establish the consistency and asymptotic normality of the proposed estimators. These results are indispensable for conducting valid inference on the estimated bounds.
\begin{proposition}
Asymptotic normality of the bounds of the $\text{QTT}_{\text{AO}}$
\begin{align}
\sqrt{n}(\widehat{\Lambda^{LB}}(q) - {\Lambda^{LB}}(q)) \xrightarrow{d} \mathcal{N}\left(0,\varsigma_{LB}^{2}\right) \\
\sqrt{n}(\widehat{\Lambda^{UB}}(q) - {\Lambda^{UB}}(q)) \xrightarrow{d} \mathcal{N}\left(0,\varsigma_{UB}^{2}\right)
\end{align}
Provided that $$q \in (0,1),$$ $$F_{Y_{1}\mid G=0, S_2 = 1}\left(Q_{Y_{1}\mid G=1,S_2 = 1}(q\pi_1 + 1 - \pi_1)\right) +1 - \pi_{0} < 1,$$ and $$F_{Y_1 \mid G=0, S_2 = 1}\left(Q_{Y_1\mid G=1, S_2 = 0} (q\pi_1)\right) -( 1 - \pi_0) > 0.$$ \\
Proof: See Appendix (ref)
\end{proposition}
Under Positive Monotonicity, $\pi_{0} = 1 $ and the bounds are asymptotically normal for any $q \in (0,1)$. When $\pi_{0} \neq 1$, there might be some extreme quantiles in which the estimators in equations ((ref)-(ref)) are the sample maximum/minimum of $Y_{2} \mid G=0$. In these cases, the estimator for $Q_{Y_{2}(0) \mid G_{i} = 1, S_2 = 1}(q)$ follows a extreme value distribution.
\subsection{Confidence Intervals}
Provided that the estimators of the bounds are asymptotically normal, imbens_confidence_2004 provide an expression to construct confidence intervals around a partially identified parameter:
\begin{align}
CI_{\alpha} = \left[\widehat{\Lambda^{LB}} - Z_{\alpha} \frac{\hat \sigma_{LB}}{\sqrt{n}} \quad ,\quad \widehat{ \Lambda^{UB}} + Z_{\alpha} \frac{\hat\sigma_{UB}}{\sqrt{n}}\right]
\end{align}
where $Z_{\alpha}$ is such that:
\begin{equation*}
\Phi\left(Z_{\alpha} + \sqrt{n}\frac{\widehat{\Lambda^{UB}} - \widehat{\Lambda^{LB}} }{\max\{\hat\sigma_{UB},\hat\sigma_{LB}\}}\right) - \Phi\left(-Z_{\alpha}\right) = \alpha,
\end{equation*}
where $\Phi(\cdot)$ is the c.d.f. of the standard normal distribution. The interval described in equation ((ref)) contains the set $[\Lambda^{LB}, \Lambda^{UB}]$ at least $\alpha\%$ of the times. Intuitively, when the bounds are far away from each other and/or the standard errors of the bounds are small, the confidence interval is constructed by adding or subtracting the one-side critical value associated with a confidence level $\alpha$ (e.g., 1.645 for a 95% confidence interval) times the standard error of the bound. This critical value increases when the bounds are close to each other or have large standard errors. In the extreme case where $\widehat{ \Lambda^{LB}} = \widehat{\Lambda^{UB}}$ and the $\text{ATT}_{\text{AO}}$ is point identified, $Z_{0.95} \approx 1.96$ and equation ((ref)) becomes the standard confidence interval constructed around a point-identified parameter.
\section{Empirical application}
This section illustrates the methodology using a job training program, \textit{Jóvenes en Acción,} implemented in Colombia between 2002 and 2005. The program provided three months of in-classroom training and three months of on-the-job training to disadvantaged youth to improve their labor market outcomes. The program has been previously studied by attanasio_subsidizing_2011, who found a positive effect on wages and the probability of having paid employment for women. For men, they find a positive effect on job quality but no significant impact on earnings. In a follow-up study, attanasio_vocational_2017 examined the program's long-term effects, finding similar and persistent impacts on earnings and job quality for both men and women. More recently, possebom_probability_2024 revisited these results using principal stratification analysis. They partially identify the effect of the training on the probability of being employed in the formal sector for the women who are employed, regardless of their treatment status. For this stratum, they cannot reject the null that the program had no impact on formal employment.
In this paper, I use data from attanasio_subsidizing_2011, which consists of survey responses from 3,955 individuals. The participants were interviewed twice: initially in January 2005, before the training, and then between August and October 2006, after the training. Using the baseline data, I construct a panel dataset.
Further details on data collection and program implementation are available in attanasio_subsidizing_2011.
I analyze the effect of the training on salaried wages. This outcome is observed only among those with paid employment, as unemployed and self-employed workers do not receive salaried wages. Of the 3,955 individuals in the sample, 1453 have paid employment in the pre-treatment period. Among these, 20% were no longer employed in salaried positions during the post-treatment period, and 19% could not be reached in the post-treatment survey. As a result, I observe salaried wages in both periods for 888 individuals. Table (ref) presents these numbers separately for each treatment group.
I can define two different sources of sample selection in my sample: transitioning to unemployment/self-employment and survey non-response. For those with missing salaried earnings, I observe the specific source of selection in the data. I assume negative monotonicity for unemployment/self-employment and positive monotonicity for survey non-response.\footnote{
The negative monotonicity assumption implies that all workers who did not transition into unemployment or self-employment after receiving training would also have retained their salaried positions in the absence of treatment. In the context of disadvantaged youth in Colombia, many salaried jobs are informal. As a result, training may reduce the probability of holding a salaried position by increasing treated workers’ reservation wages, enhancing their entrepreneurial skills and likelihood of self-employment, or raising aspirations that lead to longer job-search durations in the short run following program completion. By contrast, the positive monotonicity assumption implies that all control units that were reachable in the post-treatment period would also have responded to the survey had they been treated. This assumption is plausible if program participation increases respondents’ commitment to survey participation and if attrition is correlated with migration, which the treatment may reduce.
}
Under these assumptions, I estimate the proportion of units that have paid employment in both treatment arms among my final sample of 888 workers. That is, I estimate the proportions of Always-Observed units, $\pi_{1}$ and $\pi_{0}$ as defined in Proposition (ref), using the estimators described in equations ((ref)) and ((ref)). I estimate $\hat \pi_{1} = 0.93$ and $\hat \pi_{0} = 0.96$. In other words, of the 501 units observed in both periods in the treatment group and the 387 units in the control group, 93% of the treated units and 96% of the control group units would have also been observed in the opposite group.
Using these estimated proportions, I partially identify the $\text{ATT}_{\text{AO}}$. Table (ref) reports the estimates of the bounds with and without the inclusion of covariates.\footnote{The covariate adjustment is described in Appendix (ref).} For comparison, columns 3 and 4 present results from a complete-case analysis, which discards observations with missing outcomes and applies the Changes-in-Changes estimator to the subsample of units with observed outcomes.
\begin{table}[H]
\caption{Estimates of the bounds for the $\text{ATT}_{\text{AO}}$}
\begin{center}
\begin{tabular}{r*{8}{c}}
\toprule
Outcome & \multicolumn{8}{c}{Log of salaried earnings} \\
\cmidrule(lr){2-9}
Estimand & \multicolumn{4}{c}{$\text{ATT}_{\text{AO}}$ } & \multicolumn{4}{c}{Complete Case} \\
\cmidrule(lr){2-5} \cmidrule(lr){6-9}
Estimate & [-0.11 ,& 0.429] & [-0.095 ,& 0.319] & \multicolumn{2}{c}{0.129} & \multicolumn{2}{c}{0.123} \\
95% CI & (-0.169 ,& 0.527) & (-0.158 ,& 0.416) & (0.037 ,& 0.22) & (0.031 ,& 0.215) \\
Covariates & \multicolumn{2}{c}{No} &\multicolumn{2}{c}{Yes} & \multicolumn{2}{c}{No} & \multicolumn{2}{c}{Yes} \\
N & \multicolumn{2}{c}{888}&\multicolumn{2}{c}{888}&\multicolumn{2}{c}{888}&\multicolumn{2}{c}{888} \\
\bottomrule
\end{tabular}
\end{center}
\begin{minipage}{1\linewidth \setstretch{0.75} }
{\scriptsize Notes: Columns 1 and 2 present the estimates of the bounds for the $\text{ATT}_{\text{AO}}$ using the estimated proportions $\hat \pi_{1} = 0.93$ and $\hat \pi_{0} = 0.96$. Column 1 reports the bounds estimated without covariates with the estimators presented in Section (ref), while Column 2 includes education as a covariate as explained in Appendix (ref). Columns 3 and 4 show estimates using the complete case analysis without and with covariates respectively. The 95% confidence intervals are computed using equation ((ref)).
}
\end{minipage}
\end{table}
Table (ref) illustrates the importance of the methodology developed in this paper. Disregarding the problem of sample selection and estimating the ATT using units with observed outcomes suggests that the training had a positive and significant effect on salaried earnings, increasing them by 12% for treated units. However, this increase may be driven by differences in the distribution of outcomes across strata. Instead, I estimate the intensive-margin effect of the training by targeting treated units with paid employment regardless of their treatment status. The bounds for this effect include 0. Accordingly, I cannot reject the null hypothesis that the program had no impact on wages at any significance level.
In this context, it is relevant to examine the distributional effects. Suppose the treatment effects are heterogeneous along the outcome distribution. In that case, the bounds for the average effect may include zero even if the effect is positive and significant in some parts of the distribution. For instance, consider two different training programs. The first only benefits workers at the top of the salaried earnings distribution. The other only benefits workers at the bottom. These two policies may generate the same bounds for the $\text{ATT}_{\text{AO}}$. However, their policy implications change drastically: while the first program increases wage dispersion, the second helps reduce earnings inequality. Figure (ref) plots the $\text{QTT}_{\text{AO}}(q)$ for different quantiles $q$.
\begin{figure}[H]
\caption{$\text{QTT}_{\text{AO}}(q)$. Outcome: log of salaried earnings}
\begin{minipage}{0.99\linewidth \setstretch{0.75}}
{\scriptsize Notes: Quantile Treatment effect on the Treated Always Observed units ($\text{QTT}_{\text{AO}}(q)$) for different values of $q$. The number of units used to compute these bounds is 888.}
\end{minipage}
\end{figure}
Figure (ref) shows that the $\text{QTT}_{\text{AO}}(q)$ bounds vary across different quantiles, reflecting heterogeneity in the treatment effects. The 95% confidence interval includes zero for all the quantiles. However, their widths and ranges differ substantially across the outcome distribution. The bounds are vast and not very informative at the lower quantiles. However, as $q$ approaches the median, the bounds tighten and the lower bound approaches or even exceeds zero. A precise zero effect is estimated between the median and the third quartile. Finally, a moderate but positive effect is estimated around the eighth decile, before the bounds widen again towards the upper tail of the distribution.
\section{Conclusion and extensions}
This paper develops a novel identification strategy for estimating the intensive margin effect in panel data settings. I propose two new estimands, the $\text{ATT}_{\text{AO}}$ and the $\text{QTT}_{\text{AO}}(q)$, which can be interpreted as the average and the quantile intensive-margin effects, respectively. I explore partial identification of these causal quantities within a CiC framework, which allows the treatment to be confounded with unobservable unit-level characteristics.
Appendix (ref) extends the methodology presented in the main text in different directions. Appendix (ref) replaces the Changes-in-Changes outcome model with the canonical Difference-in-Differences framework. Appendix (ref) revisits the methodology without Assumption (ref), which treats missingness as an absorbing state. Appendix (ref) studies how to incorporate covariates into the identification and estimation strategies. Appendix (ref) extends the identification to repeated cross-sectional data. Appendix (ref) explores the identification when the outcome is discrete. Yet several extensions remain for future research.
The intensive margin effects are instrumental for policy evaluation. Beyond estimating the intensive margin effect, I encourage practitioners to examine the extensive margin effect by characterizing all the principal strata, including those in the Observed-if-Control and Observed-if-Treated strata. Furthermore, the two estimands proposed in this paper focus on the treatment group. Since the treatment is not unconfounded, its effect will likely differ across the treatment and control groups. However, the CiC framework can be extended to impute the missing potential outcome for control units, thereby enabling the identification of the average treatment effect for the control group. As a result, the proposed methodology can accommodate the identification of two additional estimands, the $\text{ATE}_{\text{AO}}$ and the $\text{QTE}_{\text{AO}}(q)$, enhancing the external validity of the conclusions drawn from the analysis.
This paper focuses on the two-group, two-period case. However, many applications involve multiple pre- and post-treatment periods. Incorporating this information is essential for both assessing the identification assumptions and gaining a deeper understanding of the policy's impacts. Moreover, many policies are implemented in a staggered fashion. While the proposed methodology can be smoothly extended to the multiple groups, multiple periods setting through the estimation and aggregation of pairwise comparisons athey_identification_2006,sun_estimating_2021, a more ambitious and valuable direction for research would be to integrate it with doubly robust estimators that combine outcome modeling with propensity score weighting arkhangelsky_doubly_2022, santanna_doubly_2020. Another promising direction for future research is the extension to time-varying estimands, as proposed by comment_survivor_2025,lin_longitudinal_2008, which may be particularly policy-relevant in dynamic settings.
One limitation of the proposed methodology is that the derived bounds may be uninformative. In general, lower proportions of Always-Observed units lead to wider bounds, and even when this proportion is close to one, certain outcome distributions may still generate wide identification regions. Adopting a Bayesian perspective may help infer where the target estimand is most likely to lie within the bounds. Another approach is to tighten the bounds by incorporating covariates, as discussed in Appendix (ref). Under treatment unconfoundedness, covariate-adjusted bounds are weakly tighter than unadjusted bounds grilli_nonparametric_2008, mealli_using_2013, lee_training_2009, long_sharpening_2013. However, in the setting studied here, where treatment is confounded and covariate distributions can differ substantially across groups, there is no guarantee that covariate adjustment will tighten the bounds. Determining whether, how, and which covariates should be included in panel data settings remains an important topic for future research.
Finally, this paper relaxes the monotonicity assumptions while maintaining the point identification of principal strata proportions. This extension may also be of interest in settings under unconfoundedness. Nevertheless, it relies on multiple sources of sample selection. Furthermore, it maintains the monotonicity assumption for each available source. If the researcher does not observe multiple sources of selection, believes that these multiple sources do not capture all the unobservables that may affect selection in different directions, and/or does not want to rely on monotonicity of any form, shin_difference--differences_2024,rathnayake_difference--differences_2024 provide identification of the $\text{ATT}_{\text{AO}}$ without monotonicity for the DiD case. However, they replace the monotonicity assumption with the assumption that the average treatment effect on selection is the same in the treatment and control groups\footnote{In the case of rathnayake_difference--differences_2024, conditional on the pre-treatment selection outcome.}, as in settings where treatment is unconfounded. An interesting approach to address the monotonicity assumption is the automated method proposed by duarte_automated_2024, which can be used to assess the plausibility of the assumption and derive sharp bounds that relax or even remove it.
spacing{1}
{
\nocite{*}
\newrefcontext[sorting=nyt]
\printbibliography
}