EconBase
← Back to paper

Sharp Bounds and Inference in Sample Selection Models with Treatment Endogeneity

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

95,783 characters

Sharp Bounds and Inference in Sample Selection Models with Treatment Endogeneity


	\begin{titlepage}
		\title{Sharp Bounds and Inference in Sample Selection Models with Treatment Endogeneity \thanks{ \scriptsize We would like to thank Otavio Bartalotti, Giovanni Mellace, Vitor Possebom, Vira Semenova, and the participants at the IAAE 2025 and USC-UCLA Mini Workshop for helpful comments and discussion. Phillip Heiler would also like to thank the Danish National Research Foundation (DNRF186) and the Independent Research Fund Denmark for support (DFF 3099-00077B).  All remaining errors are ours.}}
		\author{
   Yingying Dong\thanks{{\scriptsize University of California Irvine, Department of Economics, Irvine, CA, 92697, USA.
 email: [email removed]}} \quad
  Phillip Heiler\thanks{{\scriptsize Aarhus University. Department of Economics and Business Economics, Aarhus Center for Econometrics (ACE), TrygFonden's Centre for Child Research, Universitetsbyen 51, 8000 Aarhus C, Denmark. email: [email removed]}} \quad
      }

		\date{\underline{This version}: \today}
		\maketitle
		\thispagestyle{empty}

		\begin{abstract} \singlespacing	\small
    This paper provides partial identification and inference for treatment effects in nonparametric sample selection models with endogenous treatment and (weak) sample selection monotonicity. Outcomes are observed only for a non-randomly selected subsample and treatment is endogenous because of noncompliance with assignment. The proposed bounds for intensive margin treatment effects among compliers are sharp and tighter than those of \cite{chen2015bounds}. For inference, we develop semiparametrically efficient orthogonal moments and a debiased machine learning procedure that permits valid root-$n$ inference under high-dimensional covariates and/or flexible functional forms. Simulation results indicate good finite sample performance. Applications to Job Corps and the Oregon Health Insurance Experiment show that the method can deliver substantially tighter effect bounds and confidence intervals than existing alternatives.
		\end{abstract}
		\noindent \textbf{Keywords:} Debiased/double machine learning; Lee bounds;  Partial identification; Principal strata; Noncompliance; Sample selection \\
		\textbf{JEL classification}: C13, C14, C21
	\end{titlepage}

\setcounter{page}{1}
	\newpage



\section{Introduction}

This paper provides sharp bounds and inference for the causal effect of a binary treatment in settings with nonrandom sample selection and endogenous treatment take-up -- two prominent features of many empirical applications. Our paper can be viewed from two complementary perspectives: From the perspective of the sample selection literature, we extend
intensive margin bounds from reduced-form effects of treatment assignment to complier treatment effects in the presence of treatment take-up endogeneity. From the perspective of the instrumental variable/local average treatment effect (IV/LATE) literature, we extend the standard LATE framework to allow for sample selection, or partial observability of the outcome. Thus, this paper targets the intensive margin analogue of LATE for always-selected compliers. The framework nests the widely used \cite{lee2009training}
bounds
as a special case. Identification is obtained under standard  IV/LATE assumptions together with weak sample selection monotonicity for treatment compliers. We also develop semiparametrically efficient orthogonal moments and debiased machine learning (DML) estimators that allow valid root-$n$ inference on the treatment effect and its sharp bounds under high-dimensional covariates and flexible functional forms.

Partial identification of causal effects under sample selection or partial observability has been extensively studied.
Much of the existing work, including \cite{lee2009training}, derives bounds for the effect of a randomly assigned instrument without accounting for endogenous treatment take-up, and is therefore limited to intention-to-treat (ITT) effects when assignment and treatment receipt differ, see, e.g., \cite{horowitz2000nonparametric}, \cite{zhang2003estimation}, \cite{imai2008sharp}, \cite{huber2015sharp}, \cite{heiler2024heterogeneous}, \cite{heiler2024treatmentevaluationintensiveextensive}, \cite{sun2024partiallyidentifiedheterogeneoustreatment}, or \cite{lee2025leeboundscontinuoustreatment}.\footnote{\cite{horowitz2000nonparametric} propose nonparametric bounds for randomized experiments with missing covariate and outcome data. \cite{zhang2003estimation} develop bounds for the survivor average causal effect, targeting individuals who are always selected regardless of treatment status using principal stratification \citep{frangakis2002principal}. They provide both assumption-free bounds and bounds under sample selection monotonicity and stochastic dominance assumptions. Their monotonicity based bounds are now regularly referred to as ``Lee bounds'' or ``Zhang-Rubin-Lee bounds'' \citep{andersen2023guide}. \cite{imai2008sharp} proves the sharpness of the \cite{zhang2003estimation} bounds and extends them to quantile treatment effects. \cite{huber2015sharp} derive bounds for additional subpopulations, such as individuals selected only under treatment or only under control, and for the observed subpopulation. \cite{heiler2024heterogeneous} provides heterogeneous treatment effect bounds. \cite{heiler2024treatmentevaluationintensiveextensive} provide identification for principal strata and extensive and intensive margin bounds. \cite{sun2024partiallyidentifiedheterogeneoustreatment} consider bounds for type-specific potential outcome means, assuming exogenous treatment assignment and using an excluded variable that shifts selection but not potential outcomes. \cite{lee2025leeboundscontinuoustreatment} consider intensive margin bounds for continuous treatments.}

In practice, however, realized treatment take-up is often endogenous in both experimental and quasi-experimental settings as assignment, eligibility, or encouragement may shift participation without perfectly determining treatment. For instance, in our empirical applications a nontrivial fraction of individuals do not comply with the assignment. In the National Job Corps (JC) Study, 26.2\% of the individuals provided with access to training did not enroll in JC, while 4.4\% of individuals assigned to the control group eventually enrolled in JC after randomization \citep{schochet2008does}. In the Oregon Health Insurance Experiment (OHIE), only 30\% of those eligible to apply for Medicaid successfully enrolled, and a small share of controls also enrolled \citep{finkelstein2012oregon}. Noncompliance of these magnitudes can substantially affect treatment effect bounds.

Without sample selection, ITT and complier treatment effect differ only by a scaling factor, the share of compliers \citep{angrist1996identification}. We show that this simple relationship breaks down under sample selection: Sharp treatment effect bounds cannot be obtained by merely scaling sharp bounds of ITT effects via primitive probabilities. Our analysis therefore fills an important gap by delivering identification and inference for sharp treatment effect bounds in the more realistic case where both endogenous sample selection and noncompliance (implying treatment endogeneity) are present.


Other papers have leveraged IV/LATE-type assumptions for partial identification of various parameters in sample selection models:
\cite{lechner2010partial} derive bounds for mean and quantile treatment effects for observable subpopulations, such as the ``treated and selected''.
\cite{christelis2019partial} provide bounds on potential outcome distributions under selected samples and a monotone IV assumption \citep{manski1997monotone}.

Within this literature, the closest contribution to ours is Chen and Flores (2015; CF).
To our  knowledge, CF is the only paper studying bounds in a setting with both sample selection and noncompliance that also provides statistical inference. Unlike the CF bounds, our bounds are sharp. Related identification contributions include \cite{imai2007identification}\footnote{The derivation of bounds in \cite{imai2007identification} is based on a principal-strata approach that is complementary to ours. However, the corresponding bound expressions imported from \cite{imai2008sharp} appear to contain errors in the expressions for $Q_0$ and $Q_1$ in RESULT 1 on page 4. Hence, the resulting bounds cannot be reconciled with ours. Moreover,
\cite{imai2007identification} does not develop estimation or inference, and
sharpness is not formally established for the noncompliance-and-selection
parameter. Sharpness is only indirectly suggested by the arguments in
\cite{imai2008sharp}, which is formulated for ITT bounds under different
assumptions. By contrast, we provide corrected expressions for the bounds, prove
their sharpness, compare them with CF bounds, and develop a full estimation and
inference framework that extends to settings with both discrete and continuous
covariates and weak sample selection monotonicity.} and \cite{bartalotti2023identifying}.\footnote{\cite{bartalotti2023identifying}
study treatment-effect bounds under sample selection within a latent-index marginal treatment effect (MTE)
framework. Their Proposition 5 derives sharp bounds for always-observed LATEs
with multi-valued discrete instruments. In the binary-instrument, no-covariate
case, their latent-index interval $p_0<V\le p_1$ coincides with our complier
stratum as defined in Section \ref{sec:bounds}. Under the same strong sample selection monotonicity, their Proposition 5 bounds are algebraically equivalent
to our basic sharp bounds without covariates in Section \ref{sec:bounds}. They informally discuss plug-in estimation for MTE bounds without covariates, but do not develop accompanying inference guarantees for the discrete-instrument always-observed LATE bounds in Proposition 5.}
Relative to all three papers, we
contribute along multiple dimensions: First, we allow for the inclusion of covariates, thereby accommodating unconfounded rather than independently assigned instruments, which is often more realistic in quasi-experimental settings.
Even when the instrument is independently assigned, incorporating covariates can improve efficiency, much as in standard average treatment effect (ATE) estimation. Second, we relax their individual-level strong sample selection monotonicity to weak, i.e.,~covariate-dependent, monotonicity. This is empirically relevant: \cite{semenova2025generalized} shows that strong sample selection monotonicity can be rejected in JC. We also find substantial violations in the OHIE. Third, we provide a simple covariate-profiling procedure for the targeted complier population at the intensive margin. This helps to assess both the external validity of the estimates and the substantive relevance of the subpopulation to which they apply, paralleling complier profiling approaches from the LATE literature \citep{abadie2003semiparametric,angrist2004treatmenteffectheterogeneity,singh2023doublerobustnessforcomplier}. Fourth, we propose a semiparametric DML procedure that
delivers valid root-$n$ inference using generic nonparametric or machine learning methods for the nuisance components, allowing for potentially high-dimensional covariates and/or unknown functional forms. In strong contrast to \cite{chen2015bounds}, our estimators are asymptotically linear, and thus confidence intervals have the standard ``estimate $\pm$ standard error $\times$ critical value'' form and do not rely on alternative concepts such as half-median-unbiasedness \citep{chernozhukov2013intersection}.


Recent research has generalized estimation and inference of Lee-type bounds to high-dimensional or otherwise flexible settings using semiparametric DML methodology.
\cite{semenova2025generalized} extends \cite{lee2009training} intensive margin bounds to multiple outcomes and provides orthogonal moments for debiased estimation. \cite{heiler2024heterogeneous} develops a DML procedure to obtain corresponding heterogeneous treatment effect bounds and inference under local misspecification. \cite{heiler2024treatmentevaluationintensiveextensive} provide refined DML methods for outer identification regions for intensive and extensive margin effects that admit regular inference in a larger class of DGPs. These contributions all focus on inference for (conditional) ITT effects.
We instead construct semiparametrically efficient orthogonal moment functions for the sharp bounds on the treatment effect for compliers at the intensive margin. They nest several leading moments from the literature, including for ITT effect bounds (Semenova 2025, under perfect compliance), LATE (Fr\"olich 2007\nocite{frolich2007nonparametric}, in the absence of sample selection), and ATE (Hahn 1998\nocite{hahn1998role}, under perfect compliance and no sample selection).

A crucial difference to the literature on DML inference on Lee-type ITT bounds is that our sharp bounds require trimming based on a population whose quantiles are identified only \textit{indirectly} via inversion of a compliance weighted cumulative distribution in the sense of \cite{abadie2003semiparametric}. The latter by itself is a linear, but not necessarily convex, combination of two conditional CDFs. As a result, debiasing relies on an implicit function characterization. We leverage this representation to impose learning rate requirements only on primary, reduced-form type conditional CDFs and conditional means instead of the counterfactual conditional quantiles used for trimming the relevant observed strata distributions. This also has the advantage of making it straightforward in practice to guarantee non-crossing properties for the conditional quantiles involved, e.g., via inversion of isotonic (distributional) regression \citep{henzi2021isotonic}. Our implementation also follows this inversion strategy. Monte Carlo simulations suggest good coverage and power properties of the inference method in finite samples.

We apply and compare our method to evaluate the effects of Job Corps on earnings and Medicaid on healthcare utilization. Under the strong sample selection monotonicity assumption, our sharp bounds tighten the existing \cite{chen2015bounds} bounds by 68.4\% (JC) and 31.1\% to 69.4\% (OHIE, depending on outcome).
For example, the sharp bounds without covariates for the intensive margin complier effect of JC on hourly wages are now $[0.019,0.067]$ compared to $[-0.022,0.130]$ for CF.
Similar reductions apply to the 95\%-confidence intervals for the effects. Their widths are reduced by 46.7\% (JC) and 25.8\% to 36.9\% (OHIE, depending on outcome).
Under weak sample selection monotonicity, our sharp DML bounds still tend to be contained by the restrictive CF bounds for most outcomes and, despite relying on weaker assumptions, continue to yield shorter confidence intervals of around 7.9\% (JC) and 7.3\% to 42.3\% (OHIE, depending on outcome).

The rest of the paper is organized as follows: Section \ref{sec:bounds} derives the basic bounds without covariates and compares them with Lee and CF bounds. Section \ref{sec:bounds_covariates} extends these bounds to incorporate covariates. Section \ref{sec:inference} develops estimation and inference.
Section \ref{sec:profiling} presents the profiling method for the targeted complier population at the intensive margin. Section \ref{sec:JC} provides the JC application.
Section \ref{sec:conclusion} concludes. The OHIE application, Monte Carlo simulations as well as proofs and supplementary derivations are in the Supplementary Appendix.

\section{Basic Bounds Without Covariates} \label{sec:bounds}

\subsection{Model and Identification of Basic Bounds}

\label{subsec:baseline_bounds}

Let $D\in\{0,1\}$ denote a binary treatment and $Z\in\{0,1\}$ a binary instrument. Let $S\in\{0,1\}$ be a sample selection indicator for a continuous outcome $Y\in\mathcal{Y}\subset\mathbb{R}$ that is observed only if $S=1$. The observed data are independent draws of $(SY,S,D,Z)$.
For $d,z\in\{0,1\}$, let $Y_{d,z}$ and $S_{d,z}$ denote, respectively, the potential outcome and selection indicator that would be observed if treatment and the instrument were set exogenously to $(d,z)$. Let $D_{z}$ denote the potential treatment if $Z$ were set to $z$. We impose the following IV assumptions:

\begin{assumption}[IV]
\textbf{\label{ass:IV}}
\end{assumption}

\ref{ass:IV}.1 \textit{(Exclusion) For $d=0,1$, $Y_{d,1}=Y_{d,0}=:Y_{d}$ and $S_{d,1}=S_{d,0}=:S_{d}$.}

\ref{ass:IV}.2 \textit{(Independence) $(Y_{1},Y_{0},S_{1},S_{0},D_{1},D_{0})\perp Z$.}

\ref{ass:IV}.3 \textit{(Strong Treatment Response Monotonicity) $P(D_{1}\geq D_{0})=1$.}

\ref{ass:IV}.4 \textit{(First Stage) $P(D_{1}\ne D_{0})>0$.}

\ref{ass:IV}.5 \textit{(Non-trivial Assignment) $P(Z=1)\in(0,1)$.}

\medskip

These assumptions are the standard IV/LATE assumptions \citep{imbens1994identification,angrist1996identification} with the added requirement that instrument $Z$ is excluded from directly causing selection (2.1.1) and is allocated independently of potential selection (2.1.2). This matches classic encouragement or eligibility designs with selected samples where it is credible that the instrument causes selection and outcome only indirectly through the treatment.


Within this framework, types and causal effects can be decomposed into extensive and intensive margins. Point identification of the extensive margin effect is possible only for \textit{compliers} ${D_{1}>D_{0}}$. In particular, the extensive-margin effect for compliers is
\begin{equation}
\theta_{E}:=E[S_{1}-S_{0}| D_{1}>D_{0}],
\end{equation}
which under Assumption \ref{ass:IV} is point identified by the usual Wald ratio
\begin{equation}
\theta_{E}=\frac{E[S| Z=1]-E[S| Z=0]}{E[D| Z=1]-E[D| Z=0]}. \label{eq:thetaE}
\end{equation}
Our main target parameter is the average causal effect for compliers at the intensive margin, i.e.,~the compliers $D_1 > D_0$ who are selected regardless of treatment response $S_1 = S_0 = 1$.
We refer to this parameter as the \textit{always-selected} or \textit{survivor local average treatment effect} (SLATE):
\begin{equation}
\theta_{SLATE}:=E[Y_{1}-Y_{0}| S_{1}=S_{0}=1,D_{1}>D_{0}]. \label{eq:SLATE_def}
\end{equation}

\medskip
\par{\textbf{Remark 1.}}
\textit{Without noncompliance, $\theta_{SLATE}$ reduces to the intensive margin/always-selected ATE in \cite{lee2009training}. Without sample selection, $\theta_{SLATE}$  reduces to the LATE \citep{imbens1994identification}. Without both, it collapses to the ATE. $\theta_{SLATE}$ is the natural intensive margin analogue of the LATE: It focuses on always-selected compliers, whose treatment status is changed by the instrument while their selection status is not. It is therefore a policy effect on the level of the outcome for a principal stratum, uncontaminated by selection into the observed sample or noncompliance. The policy relevance of $\theta_{SLATE}$ has to be discussed on a case-by-case basis. We note that, in what follows, our assumptions will be able to
identify the share of always-selected compliers as well as their observable characteristics. This can be used to compare their features with those of the overall population or other relevant groups, thereby assessing how substantive the parameter is in a given empirical setting. We provide all details in Section \ref{sec:profiling} and empirical examples in Section \ref{sec:JC} and Appendix \ref{sec:OHIE}.}
\medskip

Under selected samples $\theta_{SLATE}$ is only partially identified because the joint selection status $(S_{1},S_{0})$ is not observed. In particular, even for compliers who are selected only under treatment $(S_{1}=1,S_{0}=0)$ or only under control $(S_{1}=0,S_{0}=1)$, either $Y_{0}$ or $Y_{1}$ is never observed, so their treatment effects are not identified without additional (untestable) assumptions.
To make progress, we first state some standard IV identities under Assumption \ref{ass:IV}:
\begin{lemma}
\label{lem:IV_identities}
Suppose Assumption \ref{ass:IV} holds. Then, for any integrable function $g: \mathcal{Y}\rightarrow \mathbb{R}$,
{\begin{align}
P(D_{1}>D_{0})&=E[D| Z=1]-E[D| Z=0], \label{eq:PrC}\\
P(S_{1}=1,D_{1}>D_{0})&=E[DS| Z=1]-E[DS| Z=0], \label{eq:PrS1C}\\
E[g(Y_{1})| S_{1}=1,D_{1}>D_{0}]&=\frac{E[DSg(Y)| Z=1]-E[DSg(Y)| Z=0]}{E[DS| Z=1]-E[DS| Z=0]}. \label{eq:EY1S1C}
\end{align}}
Moreover, replacing $D$ with $(D-1)$ in \eqref{eq:PrS1C} and \eqref{eq:EY1S1C} identifies $P(S_{0}=1,D_{1}>D_{0})$ and $E[g(Y_{0})| S_{0}=1,D_{1}>D_{0}]$, respectively.
\end{lemma}
In particular, taking $g(Y)=\mathbbm{1}(Y\leq y)$ yields conditional CDFs $F_{Y_{1}| S_{1}=1,D_{1}>D_{0}}(y)$ and $F_{Y_{0}| S_{0}=1,D_{1}>D_{0}}(y)$, while taking $g(Y)=Y$ yields the corresponding conditional means.

Next, we introduce the complier strata that underlie the SLATE parameter. Among compliers ($D_{1}>D_{0}$), define

\begin{alignat}{2}
ac&:=\{S_{0}=S_{1}=1,\;D_{1}>D_{0}\} &\quad\text{always-selected compliers} \\
dc&:=\{S_{1}<S_{0},\;D_{1}>D_{0}\} &\quad\text{treatment-de-selected compliers} \\
cc&:=\{S_{1}>S_{0},\;D_{1}>D_{0}\} &\quad\text{treatment-selected compliers}
\end{alignat}
Rewriting \eqref{eq:SLATE_def} in terms of these strata yields

\begin{equation}
\theta_{SLATE}=E[Y_{1}| ac]-E[Y_{0}| ac].
\end{equation}
The set $\{S_{0}=1,D_{1}>D_{0}\}$ combines $ac$ and $dc$, while $\{S_{1}=1,D_{1}>D_{0}\}$ combines $ac$ and $cc$. Lemma \ref{lem:IV_identities} allows us to identify $P(S_{0}=1,D_{1}>D_{0})$ and $P(S_{1}=1,D_{1}>D_{0})$, which yield two linear restrictions on the shares of three unknown types ($ac$, $dc$, $cc$). To fully identify these shares, we further impose a strong sample selection monotonicity condition for the compliers:

\begin{assumption}[\textbf{Strong Sample Selection Monotonicity for Compliers}]
Either $P(S_{1}\geq S_{0}| D_{1}>D_{0})=1$ or $P(S_{1}\leq S_{0}| D_{1}>D_{0})=1$. \label{ass:monotoneS}
\end{assumption}
Assumption \ref{ass:monotoneS} imposes a common direction of treatment-induced sample selection within the complier stratum $D_{1} > D_{0}$. In particular, the treatment is assumed to either weakly increase or weakly decrease selection for all compliers. It cannot move some compliers into the sample and others out of it, effectively ruling out either $dc$ or $cc$ types.\footnote{This restriction is implied, for example, by an additively separable selection equation, see, e.g., \cite{vytlacil2002independence} for treatment selection or \cite{heiler2024treatmentevaluationintensiveextensive} for sample selection. However, these models impose monotonicity for the full population and are thus stronger than Assumption \ref{ass:monotoneS}.} Note that no analogous restriction is needed for strata with $D_{1} = D_{0}$ (always-takers and never-takers) as their treatment status does not vary with the instrument. In particular, never-takers satisfy $D_0=D_1=0$, so their observed selection status is governed only by $S_0$; always-takers satisfy $D_0=D_1=1$, so their observed selection status is governed only by $S_1$. The counterfactual selection status under the unrealized treatment therefore does not enter the identification argument.

We now provide identification for the case where $P(S_{1}\geq S_{0}| D_{1}>D_{0})=1$. The case where $P(S_{1}\leq S_{0}| D_{1}>D_{0})=1$ can be handled analogously.
Under Assumption \ref{ass:monotoneS}, the stratum ${S_{0}=1,D_{1}>D_{0}}$ consists only of always-selected compliers ($ac$), while ${S_{1}=1,D_{1}>D_{0}}$ combines $ac$ and $cc$. Lemma \ref{lem:IV_identities} then implies that

\begin{equation}
\beta_{0}:=E[Y_{0}| ac]=E[Y_{0}| S_{0}=1,D_{1}>D_{0}]
\end{equation}
is point identified by taking $g(Y)=Y$ and replacing $D$ with $(D-1)$ in \eqref{eq:EY1S1C}.
In contrast, $E[Y_{1}| ac]$ is not point identified. However, Lemma \ref{lem:IV_identities} and Assumption \ref{ass:monotoneS} allow us to identify the mixture distribution of $Y_{1}$ for $ac$ and $cc$,

\begin{equation}
F_{Y_{1}| ac\cup cc}(y):=F_{Y_{1}| S_{1}=1,D_{1}>D_{0}}(y),
\end{equation}
as well as the frequencies of $ac$ and $cc$ within this mixture. Let

\begin{equation}
\pi_{ac}:=P(S_{0}=S_{1}=1,D_{1}>D_{0}),\quad \pi_{cc}:=P(S_{1}>S_{0},D_{1}>D_{0}).
\end{equation}
Then, under Assumption \ref{ass:monotoneS},

\begin{equation}
\pi_{ac}=P(S_{0}=1,D_{1}>D_{0}),\qquad \pi_{ac}+\pi_{cc}=P(S_{1}=1,D_{1}>D_{0}),
\end{equation}
and both probabilities are identified. Moreover, define

\begin{equation}
p:=\frac{\pi_{ac}}{\pi_{ac}+\pi_{cc}}.
\end{equation}
$p$ is equal to the fraction of always-selected compliers in  mixture ${S_{1}=1,D_{1}>D_{0}}$. Since we only observe this mixture, sharp bounds are obtained by assigning $ac$ to the ``best'' or ``worst'' location within the mixture distribution of $Y_{1}$. Let $F_{1}(y):=F_{Y_{1}| S_{1}=1,D_{1}>D_{0}}(y)$ and $Q_{1}(u):=\inf\{y\in\mathcal{Y}:F_{1}(y)\geq u\}$ denote the identified CDF and quantile function of this mixture.

To simplify exposition for the remainder of this section, we maintain the following regularity condition: the relevant identified distributions entering the trimming
formulas have continuous, strictly increasing CDFs. This rules out mass points at
the relevant trimming cutoffs and allows the bounds to be written as ordinary
trimmed means. This condition is not needed for identification per se.\footnote{A weaker local regularity condition would suffice: for each result below, it is enough that the corresponding identified CDF be continuous and strictly increasing in a neighborhood of the relevant trimming threshold(s). Without this regularity, the same conclusions also hold once the bounds are written in exact fractional-trimming (quantile-integral)
form.}

The lower and upper bounds for $E[Y_{1}| ac]$ are
\begin{align}
\beta_{L,1} &:=E[Y_{1}| Y_{1}\leq Q_{1}(p),ac \cup cc]=E[Y_{1}| Y_{1}\leq Q_{1}(p),S_{1}=1,D_{1}>D_{0}], \label{eq:betaL1} \\
\beta_{U,1} &:=E[Y_{1}| Y_{1}\geq Q_{1}(1-p),ac \cup  cc]=E[Y_{1}| Y_{1}\geq Q_{1}(1-p),S_{1}=1,D_{1}>D_{0}], \label{eq:betaU1}
\end{align}
i.e., the trimmed means obtained by placing all $ac$ individuals in the bottom $p$ fraction (for the lower bound) or the top $p$ fraction (for the upper bound) of the mixture. We obtain the following proposition:

\begin{proposition}[Sharpness]
\label{prop:sharp_bounds}
Suppose Assumptions \ref{ass:IV}-- \ref{ass:monotoneS} hold and $\pi_{ac}>0$. Then,

\begin{equation*}
\beta_{L,1}\leq E[Y_{1}| ac]\leq \beta_{U,1}.
\end{equation*}
Moreover, $\beta_{L,1}$ and $\beta_{U,1}$ are sharp: They are, respectively, the largest lower bound and the smallest upper bound for $E[Y_{1}| ac]$ that are consistent with the observed data and Assumptions \ref{ass:IV}--\ref{ass:monotoneS}. Any other valid bounds under these assumptions contain $[\beta_{L,1},\beta_{U,1}]$.
\end{proposition}

Combining the point-identified $\beta_{0}$ with these bounds yields sharp bounds for $\theta_{SLATE}$:

\begin{equation}
\beta_{L}:=\beta_{L,1}-\beta_{0},\qquad \beta_{U}:=\beta_{U,1}-\beta_{0}. \label{eq:theta_bounds}
\end{equation}
Using Assumption \ref{ass:monotoneS} and Lemma \ref{lem:IV_identities} with $g(Y)=Y\mathbbm{1}(Y\leq Q_{1}(p))$ and $g(Y)=Y\mathbbm{1}(Y\geq Q_{1}(1-p))$, all components of $\beta_{L}$ and $\beta_{U}$, namely $\pi_{ac}$, $p$, $F_{1}$, $Q_{1}$, and the trimmed means are functionals of the observed distribution of $(SY,S,D,Z)$. In particular,

{\begin{align}
\beta_{L}=\frac{1}{\pi_{ac}}\left\{
\begin{array}{c}
E\left[DSY\mathbbm{1}(Y\leq Q_{1}(p))+(1-D)SY| Z=1\right]\\
-\,E\left[DSY\mathbbm{1}(Y\leq Q_{1}(p))+(1-D)SY| Z=0\right]
\end{array}
\right\}, \label{eq:betaL} \\
\beta_{U}=\frac{1}{\pi_{ac}}\left\{
\begin{array}{c}
E\left[DSY\mathbbm{1}(Y\geq Q_{1}(1-p))+(1-D)SY| Z=1\right]\\
 -\,E\left[DSY\mathbbm{1}(Y\geq Q_{1}(1-p))+(1-D)SY| Z=0\right]
\end{array}
\right\}, \label{eq:betaU}
\end{align}}
where

\begin{align}
    p &=\frac{E\left[ \left( D-1\right) S| Z=1\right] -E\left[ \left(
D-1\right) S| Z=0\right] }{E\left[ DS| Z=1\right] -E\left[ DS| Z=0 \right] }, \\
Q_{1}(u ) &= \inf \{y \in \mathcal{Y}:F_{1}(y)\geq u \} \text{ for any } u \in \left( 0,1\right), \\
F_{1}(y) &=
\frac{E\left[ \mathbbm{1}\{Y\leq y\}DS| Z=1\right] -E\left[ \mathbbm{1}
\{Y\leq y\}DS| Z=0\right] }{E\left[ DS| Z=1\right] -E\left[ DS| Z=0
\right] }.
\end{align}
We next compare our bounds with two benchmarks: the ITT bounds of \cite{lee2009training} and the CF bounds, which target $\theta_{SLATE}$ under essentially identical identification assumptions.




\subsection{Comparison with Existing Bounds}
\label{subsec:Bound_comparison}
\subsubsection{Comparison with \cite{lee2009training} Bounds}
\label{subsec:Lee_comparison}

Without sample selection, the ITT effect
\begin{equation}
\tau_{Y}:=E[Y| Z=1]-E[Y| Z=0]
\end{equation}
and the average treatment effect for compliers,
\begin{equation}
\theta_{LATE} := \frac{\tau_{Y}}{\tau_{D}},\qquad \tau_{D}:=E[D| Z=1]-E[D| Z=0], \label{eq:LATE1}
\end{equation}
differ only by the scaling factor $\tau_{D}$, which corresponds to the share of compliers. That is, in the standard IV point identification setting, ITT can be converted into the complier average treatment effect by dividing by $\tau_{D}$. This does not generalize to treatment effect bounds under sample selection.

In particular, in the presence of sample selection, \cite{lee2009training} derives sharp bounds for the ITT effect among the always-selected assuming strong sample selection monotonicity. These ``Lee bounds'' can be viewed as a special case of our basic bounds under perfect compliance, i.e., when treatment equals instrument, $D=Z$. We only consider the lower bound. The analysis for the upper bound is analogous and omitted for brevity. Setting $D=Z$ in Equation \eqref{eq:betaL}, we obtain the bound

{\begin{align}
\beta_{L}^{Lee} &:=E[Y| Y\leq Q_{1}^{Lee}(p^{Lee}),S=1,Z=1]-E[Y| S=1,Z=0] \notag \\
&= \beta_{L,1}^{Lee}-\beta_{0}^{Lee}, \label{eq:Lee_lower_def}
\end{align}}
where $Q_{1}^{Lee}(u):=\inf \{y\in \mathcal{Y}: F_{Y|S=1,Z=1}(y)\ge u \}$  is the $u$-quantile of $\left(Y|S=1,Z=1\right)$, and $p^{Lee}:=E[S| Z=0]/E[S| Z=1]$ is the trimming fraction.

To relate \eqref{eq:Lee_lower_def} to our SLATE bounds, it is useful to decompose the selected population into latent types. In addition to the complier types introduced in Section~\ref{subsec:baseline_bounds}, let

\begin{align}
an &:=\{S_{0}=1,D_{0}=D_{1}=0\}\quad\text{always-selected never-takers}, \\
aa &:=\{S_{1}=1,D_{0}=D_{1}=1\}\quad\text{always-selected always-takers},
\end{align}
and denote their shares by $\pi_{an}$ and $\pi_{aa}$ respectively. Under Assumptions \ref{ass:IV}--\ref{ass:monotoneS}, the selected group in each arm of the experiment is a mixture of these types. We show in Appendix~\ref{app:Lee_decomposition} that

\begin{align}
F_{Y| S=1,Z=1}(y)&=\omega F_{Y_{0}| an}(y)+(1-\omega)F_{Y_{1}| ac\cup cc\cup aa}(y), \label{eq:FY_SZ1_mix}\\
F_{Y| S=1,Z=0}(y)&=\delta F_{Y_{0}| ac\cup an}(y)+(1-\delta)F_{Y_{1}| aa}(y), \label{eq:FY_SZ0_mix}
\end{align}
for suitable mixture weights $\omega,\delta\in(0,1)$ depending only on type shares. Let $q :=Q^{Lee}_1(p^{Lee})$ denote the Lee trimming threshold. The trimmed mean in the first term of \eqref{eq:Lee_lower_def} can be written as
\begin{align}
\beta^{Lee}_{L,1}
&:= E\!\left[Y \mid Y \le q,\, S=1,\, Z=1\right] \notag \\
&= w(q)\,E\!\left[Y_0 \mid Y_0 \le q,\, an\right]
+ \bigl(1-w(q)\bigr)\,E\!\left[Y_1 \mid Y_1 \le q,\, ac\cup cc\cup aa\right],
\end{align}
where $w\left( q\right)=\tfrac{\omega F_{Y_0\mid an}(q)}{p^{Lee}}$ is the effective $an$ stratum weight from trimming applied after mixing. The second term of \eqref{eq:Lee_lower_def} can be written as
{ \begin{align}
\beta _{0}^{Lee} &:=E\left[ Y|S=1,Z=0\right]  \notag \\
&=E\left[ Y_{0}|ac \cup an\right] \delta +E\left[ Y_{1}|aa\right] \left( 1-\delta \right) .
\end{align}}
Thus, $\beta_{L}^{Lee}$ is a difference of two mixtures that involve not only treatment compliers ($ac$, $cc$), but also always-takers and never-takers ($aa$, $an$). In contrast, our lower bound for SLATE,

\begin{equation}
\beta_{L}=E[Y_{1}| Y_{1}\leq Q_{1}(p),ac \cup cc]-E[Y_{0}| ac],
\label{eq:sharp_basic_lower}
\end{equation}
depends only on the treatment complier types ($ac$ and $cc$), with $p$ equal to the fraction of always-selected compliers within that complier mixture.

Except in knife-edge cases where the shares of always-selected always-takers and never-takers vanish ($\pi_{aa}=\pi_{an}=0$), the $Y_{1}$ term in $\beta_{L}^{Lee}$ does not reduce to $E[Y_{1}| Y_{1}\leq Q_{1}(p),ac \cup cc]$, and the $Y_{0}$ term does not reduce to $E[Y_{0}| ac]$. Consequently, \cite{lee2009training} bounds cannot be transformed into sharp bounds for $\theta_{SLATE}$ by simple rescaling via primitive probabilities. We summarize this result formally in the following proposition.
\begin{proposition}[Lee bounds do not rescale to SLATE bounds]
\label{prop:Lee_not_scaling}
Suppose Assumptions~\ref{ass:IV}--\ref{ass:monotoneS} hold and $\pi_{ac}>0$.
If $\pi_{an}+\pi_{aa}>0$, then, for generic joint distributions of $(Y_0,Y_1)$
across latent types,
\begin{align*}
\beta_L^{Lee}\neq \beta_L.
\end{align*}
Moreover, there does not exist a scalar function $c$, depending only on the
primitive probabilities (i.e.~the principal-strata shares and instrument probabilities), such that
\begin{align*}
\beta_L = c\,\beta_L^{Lee}
\end{align*}
uniformly over DGPs.
\end{proposition}



Proposition \ref{prop:Lee_not_scaling} implies that \cite{lee2009training} ITT bounds cannot be converted into sharp bounds for SLATE by a simple rescaling analogous to $\tau_{Y}/\tau_{D}$ in the standard IV case without sample selection.


\subsubsection{Comparison with \cite{chen2015bounds} Bounds}\label{subsec:CF_comparison}

CF study partial identification of $\theta_{SLATE}$ under essentially identical model assumptions. We show that our bounds are
weakly tighter uniformly over DGPs, and strictly tighter on a nonempty subset of DGPs than the CF bounds. For brevity, we focus on the lower bounds. The upper bounds can be treated analogously.

CF propose two lower bounds for $E[Y_{1}| ac]$, based on trimming the distribution of $Y$ among selected treated units, $Y| DSZ=1$, equivalent to $Y_1|ac\cup cc\cup aa$ under Assumptions \ref{ass:IV}--\ref{ass:monotoneS}, and then taking the maximum of the two. In contrast, our sharp bounds trim within the smaller identified mixture $Y_1|ac\cup cc$. Let $Q_{1}^{CF}(u)$ denote the $u$-quantile of $Y| DSZ=1$. Their basic lower bound is

{\begin{align}
\beta_{L}^{mix}&:=E\left[Y| DSZ=1,Y\leq Q_{1}^{CF}(p^{CF}_1)\right]-\beta_{0} \nonumber\\
&=\beta_{L,1}^{mix}-\beta_{0},
\label{eq:CF_basic_lower}
\end{align}}

where $p^{CF}_1:=\pi_{ac}/(\pi_{ac}+\pi_{cc}+\pi_{aa})$ and $\beta_{0}:=E[Y_{0}| ac]$. Their alternative lower bound uses additional information from the always-taker stratum ($aa$),

{\begin{align}
\beta_{L}^{adj}&:=\left\{
\begin{array}{l}
\left(1+\dfrac{\pi_{aa}}{\pi_{ac}}\right)
E\left[Y\mid DSZ=1,Y\leq Q_{1}^{CF}(p^{CF}_2)\right]\\[6pt]
-\dfrac{\pi_{aa}}{\pi_{ac}}E\left[Y\mid DS(1-Z)=1\right]
\end{array}
\right\}-\beta_{0}\nonumber\\
&=\beta_{L,1}^{adj}-\beta_{0}, \label{eq:CF_alt_lower}
\end{align}}
where $p^{CF}_2:=(\pi_{ac}+\pi_{aa})/(\pi_{ac}+\pi_{cc}+\pi_{aa})$ and $E[Y| DS(1-Z)=1]=E[Y_{1}| aa]$ is point identified. The CF lower bound is then
\begin{equation}
\beta_{L}^{CF}:=\max\{\beta_{L}^{mix},\beta_{L}^{adj}\}.
\end{equation}

Under Assumptions \ref{ass:IV}--\ref{ass:monotoneS}, the conditional distribution $Y| DSZ=1$ coincides with $Y_{1}$ for the mixture of types $ac$, $cc$, and $aa$:

\begin{equation}
F_{Y| DSZ=1}(y)=F_{Y_{1}| ac\cup cc\cup aa}(y),
\end{equation}
so both $\beta_{L}^{mix}$ and $\beta_{L}^{adj}$ are obtained by trimming this mixture distribution and then reweighting to isolate $ac$ and $aa$. CF show that $\beta_{L}^{mix}$ corresponds to the worst-case lower bound when all $ac$ individuals are placed in the bottom $p^{CF}_1$-
fraction of $Y_{1}| ac\cup cc\cup aa$, while $\beta_{L}^{adj}$ corresponds to the worst case when $ac$ and $aa$ together occupy the bottom $p^{CF}_2$-
fraction of that distribution, using the fact that $E[Y_{1}| aa]$ is identified.

In contrast, our lower bound $\beta_{L}$ uses the smallest stratum that is point identified, namely $Y_{1}| ac\cup cc$. As shown in Lemma \ref{lem:F_Y1_ac_cc_identification} in Appendix \ref{app:CF_dominance}, the distribution $F_{Y_{1}| ac\cup cc}$ can be recovered from the two observable selected distributions

\begin{equation}
F_{Y| DSZ=1}(y)=F_{Y_{1}| ac\cup cc\cup aa}(y),\qquad F_{Y| DS(1-Z)=1}(y)=F_{Y_{1}| aa}(y).
\end{equation}
We then construct the sharp lower bound

\begin{equation}
\beta_{L}=E\left[Y_{1}| Y_{1}\leq Q_{1}(p),ac \cup cc\right]-\beta_{0},\qquad p:=\frac{\pi_{ac}}{\pi_{ac}+\pi_{cc}},
\end{equation}
by trimming $Y_{1}$ only within this complier mixture.

Intuitively, $\beta_{L}$ assumes, in a worst-case fashion, that all $ac$ individuals are located below all $cc$ individuals in the distribution of $Y_{1}$ among compliers, but it imposes no restrictions on the relative ranking of $ac$ versus $aa$. By contrast, the CF bounds are based on worst-case restrictions on the joint ranking of $ac$, $cc$, and $aa$ within $Y_{1}| ac\cup cc\cup aa$. This makes them conservative. In particular, we obtain the following proposition:



\begin{proposition}[Dominance over CF bounds]
\label{prop:CF_dominance_main}
Suppose Assumptions \ref{ass:IV}--\ref{ass:monotoneS} hold and $\pi_{ac}>0$. Then,
\begin{align*}
\beta_{L,1}\geq \beta_{L,1}^{CF}\geq \beta_{L,1}^{mix} \quad \text{ and } \quad
\beta_{U,1}\leq \beta_{U,1}^{CF}\leq \beta_{U,1}^{mix},
\end{align*}
where $\beta_{L,1}$ is defined in \eqref{eq:betaL1}, $\beta_{L,1}^{mix}$ in
\eqref{eq:CF_basic_lower}, and
$\beta_{L,1}^{CF}:=\max\{\beta_{L,1}^{mix},\beta_{L,1}^{adj}\}$,
with $\beta_{L,1}^{adj}$ defined in \eqref{eq:CF_alt_lower}. The upper-bound
counterparts $\beta_{U,1}$, $\beta_{U,1}^{mix}$, $\beta_{U,1}^{adj}$, and
$\beta_{U,1}^{CF}$ are defined analogously. Hence, imposing either of the two restrictions used by CF cannot tighten our sharp lower or upper bound
for $E[Y_1\mid ac]$.

\noindent
Moreover, these inequalities are strict on a nonempty class of DGPs:
\begin{align*}
\beta_{L,1}>\beta_{L,1}^{CF}
&\iff
\pi_{aa}>0 \quad\text{and}\quad 0<F_{Y_1\mid aa}\!\bigl(Q_1(p)\bigr)<1,  \\
\beta_{U,1}<\beta_{U,1}^{CF}
&\iff
\pi_{aa}>0 \quad\text{and}\quad0<F_{Y_1\mid aa}\!\bigl(Q_1(1-p)\bigr)<1.
\end{align*}
\end{proposition}




























\section{Bounds with Covariates}
\label{sec:bounds_covariates}

We now extend the unconditional baseline bounds in Section~\ref{subsec:baseline_bounds} to incorporate predetermined covariates.
Incorporating such covariates serves three purposes: First, it allows for an unconfounded instead of a fully randomly assigned instrument. Second, even under fully randomized assignment, it can improve efficiency by conditioning on relevant covariates when constructing trimmed means. Third, it allows us to relax strong sample selection monotonicity. In particular, we permit the direction of sample selection for treatment compliers to vary with covariates. For simplicity, we maintain the strong treatment response monotonicity assumption \ref{ass:IV}.3.
This no-defiers restriction is natural in eligibility and encouragement
designs such as JC and OHIE: departures from perfect compliance are largely one-sided, arising mainly from incomplete take-up among those assigned to treatment rather than from substantial treatment receipt among controls. Relaxing this assumption is possible via further covariate partitioning analogous to the sample selection case in what follows.

Let $X\in\mathcal{X}\subset\mathbb{R}^{d_{x}}$ denote a vector of predetermined covariates. The observed data are $(SY,S,D,Z,X)$. We impose the following conditional IV assumptions:

\begin{assumption}[Conditional IV]
\label{ass:conditional_IV}
For $d=0,1$:

\ref{ass:conditional_IV}.1 (Exclusion) $Y_{d,1}=Y_{d,0}=:Y_{d}$, $S_{d,1}=S_{d,0}=:S_{d}$.

\ref{ass:conditional_IV}.2 (Independence) $(Y_{1},Y_{0},S_{1},S_{0},D_{1},D_{0})\perp Z| X$.

\ref{ass:conditional_IV}.3 (Strong Treatment Response Monotonicity) $P(D_{1}\geq D_{0})=1$.

\ref{ass:conditional_IV}.4  (First Stage) $P(D_1 \ne D_0 \mid X=x)>0$ for $P_X$-almost every $x \in \mathcal{X}$

\ref{ass:conditional_IV}.5
(Non-trivial Assignment)
$P(Z=1\mid X=x)\in(0,1)$ for $P_X$-almost every $x\in\mathcal{X}$.

\end{assumption}

These assumptions are standard extensions of the conditional IV/LATE assumptions \citep{frolich2007nonparametric} generalizing Assumption \ref{ass:IV} with instrument $Z$ again being excluded from directly causing selection and being independently allocated of potential selection given covariates.

It is important to note that Assumptions \ref{ass:conditional_IV}.4 and \ref{ass:conditional_IV}.5 are not
intended to impose substantive restrictions beyond their unconditional counterparts
Assumptions \ref{ass:IV}.4 and \ref{ass:IV}.5. Rather, they should be interpreted
as support restrictions for the conditional analysis: if the conditional first-stage
or non-trivial assignment condition fails on a set of positive $P_X$-probability,
$\mathcal{X}$ should be understood as restricted to the support on which these
conditions hold.\footnote{Failure of the first stage and failure of non-trivial assignment have slightly different interpretations. If $P(D_1\ne D_0\mid X=x)=0$, there are no compliers at that $x$. so complier-based conditional objects are not meaningful there. If $P(Z=1\mid X=x)\in\{0,1\}$, there is no within-$x$ variation in the instrument, so the conditional IV comparisons are not identified there.}

Under Assumption~\ref{ass:conditional_IV}, the identities in Lemma~\ref{lem:IV_identities} hold pointwise in $x$.
Next, we relax strong sample selection monotonicity to a weaker conditional version that allows the direction of monotonicity to vary with covariates.


\begin{assumption}[Weak Sample Selection Monotonicity for Compliers]\label{ass:conditional_monotoneS}
Either \\ \noindent $P( S_1 \geq S_0 | D_{1}>D_{0},X=x)=1$ or
    $P( S_1 \leq S_0 | D_{1}>D_{0}, X=x)=1.$
\end{assumption}
Assumption \ref{ass:conditional_monotoneS} yields subsets of the covariate space where treatment compliers can have either a weakly positive or negative selection responses to treatment. At their intersection, treatment does not affect selection. Without loss of generality, we now define a distinct partitioning where this intersection is combined with the weakly positive responders. Formally, let \begin{align*}
    \mathcal{X}^0 &=\left\{ x \in \mathcal{X}: P(S_1=S_0|D_{1}>D_{0},X=x)=1 \right\}
\end{align*}
and partitions
\begin{align}
    \mathcal{X}^+ &= \{ x \in \mathcal{X}: P(S_1 \geq S_0|D_{1}>D_{0},X=x)=1,  \notag \\ &\qquad  P(S_1>S_0|D_{1}>D_{0},X=x)>0 \} \cup  \mathcal{X}^0,  \\
    \mathcal{X}^- &=\{  x \in \mathcal{X}:   P(S_1\leq S_0|D_{1}>D_{0},X=x)=1,  \notag \\ &\qquad P(S_1<S_0|D_{1}>D_{0},X=x)>0 \}.
\end{align}







We note that the presence of $\mathcal{X}^0$ is harmless for identification. For regular semiparametric inference, however, it will be necessary that $P(\mathcal{X}^0) = 0$ \citep{heiler2024treatmentevaluationintensiveextensive}. We return to this in Section \ref{sec:inference}.

\medskip
\par{\textbf{Remark 2.}}
\textit{Assumption \ref{ass:conditional_monotoneS} weakens strong sample selection monotonicity by allowing the sign of the treatment effect on selection to vary with observed covariates. Thus, treatment may increase selection for some values in $\mathcal{X}$ and decrease it for others. The restriction is nevertheless substantive: conditional on $X=x$ and complier status, residual heterogeneity in the selection response must be one-sided. In particular, the assumption rules out the coexistence of treatment-induced entry and exit among compliers with the same covariate values.
Its credibility therefore depends on whether $\mathcal{X}$ is rich enough to absorb the economically relevant heterogeneity in the direction of selection. It is more plausible when the main determinants of the sign of the selection response are observed, such as baseline characteristics governing participation, employment, survival, etc. It is less plausible when unobserved factors can generate opposing responses within the same covariate cell, see \cite{heiler2024treatmentevaluationintensiveextensive} and \cite{semenova2025generalized} for additional discussion.
}
\medskip

Let $\pi_{ac}(x):=P(ac| X=x)$ denote the conditional share of always-selected compliers, and let $\pi_{ac} := E[\pi_{ac}(X)]$ denote their unconditional share. Further let $\mathcal X_{ac}:=\{x\in\mathcal X:\pi_{ac}(x)>0\}$.\footnote{For identification of the unconditional bounds, the substantive requirement is $\pi_{ac}>0$. $\pi_{ac}(x)$ may be zero on subsets of the covariate support as they receive zero weight in numerator and denominator of the unconditional bounds.
As a convention for display, we treat $\theta_{SLATE}(x)=0$ for $x\in\mathcal X\setminus\mathcal X_{ac}$.}
For any $x \in \mathcal X_{ac}$, define the conditional SLATE parameter

\begin{equation}
\theta_{SLATE}(x):=E[Y_{1}-Y_{0}| ac,X=x].
\end{equation}
Let
\begin{equation}
\lambda_{d}(x):=P(S_{d}=1,D_{1}>D_{0}| X=x),\quad d=0,1,
\end{equation}
and write $\pi_{cc}(x):=P(cc| X=x)$ and $\pi_{dc}(x):=P(dc| X=x)$ for the conditional shares of $cc$ and $dc$, respectively. Under Assumption~\ref{ass:conditional_monotoneS},
\begin{align}
\lambda_{0}(x) &=\pi_{ac}(x),\quad \lambda_{1}(x)=\pi_{ac}(x)+\pi_{cc}(x)\quad \text{for }x\in\mathcal{X}^{+}, \\
\lambda_{1}(x) &=\pi_{ac}(x),\quad \lambda_{0}(x)=\pi_{ac}(x)+\pi_{dc}(x)\quad \text{for }x\in\mathcal{X}^{-}.
\end{align}
This yields the ratio
\begin{equation}
p(x):=\frac{\lambda_{0}(x)}{\lambda_{1}(x)}.
\end{equation}
Thus $0 \le p(x) \leq 1$ on $\mathcal{X}^+_{ac}:=\mathcal{X}^+\cap\mathcal{X}_{ac}$ and $p(x) > 1$ on $\mathcal{X}^-_{ac}:=\mathcal{X}^-\cap\mathcal{X}_{ac}$. For $d=0,1$ and $x\in \mathcal{X}_{ac}$, define the conditional quantile
\begin{equation}
Q_{d}(u,x):=\inf\left\{y\in\mathcal{Y}:F_{Y_{d}| S_{d}=1,D_{1}>D_{0},X=x}(y)\geq u\right\}.
\end{equation}

To simplify exposition, for the remainder of this section we again maintain that the
relevant identified conditional CDF entering each trimming formula is continuous
and strictly increasing in a neighborhood of the corresponding trimming
threshold(s). This rules out mass points at the trimming thresholds and allows the conditional sharp bounds below to be written as
ordinary trimmed means.\footnote{For $x\in \mathcal{X}^+_{ac}$, the relevant CDF
is $F_{Y_{1}| S_{1}=1,D_{1}>D_{0},X=x}(\cdot)$ and the relevant thresholds are $Q_1(p(x),x)$ and
$Q_1(1-p(x),x)$. For $x\in \mathcal{X}^-_{ac}$, the relevant CDF is $F_{Y_{0}| S_{0}=1,D_{1}>D_{0},X=x}(\cdot)$ and the
relevant thresholds are $Q_0(1-1/p(x),x)$ and $Q_0(1/p(x),x)$. Without this
regularity, the same bounds can be written in exact fractional-trimming form.}
Then Proposition~\ref{prop:sharp_bounds} can be applied pointwise in $x$. This yields the following sharp lower and upper bounds for $\theta_{SLATE}(x)$:
{\begin{align}
\beta_{L}^{+}(x)&:=E\left[Y_{1}| S_{1}=1,D_{1}>D_{0},Y_{1}\leq Q_{1}(p(x),x),X=x\right] \notag \\
&\quad - E\left[Y_{0}| S_{0}=1,D_{1}>D_{0},X=x\right], \label{eq:betaL_cond_plus}\\
\beta_{U}^{+}(x)&:=E\left[Y_{1}| S_{1}=1,D_{1}>D_{0},Y_{1}\geq Q_{1}(1-p(x),x),X=x\right] \notag \\
&\quad -E\left[Y_{0}| S_{0}=1,D_{1}>D_{0},X=x\right], \label{eq:betaU_cond_plus}
\end{align}} for $x\in\mathcal{X}^{+}_{ac}$, and

{\begin{align}
\beta_{L}^{-}(x)&:=E\left[Y_{1}| S_{1}=1,D_{1}>D_{0},X=x\right] \notag \\
&\quad -E\left[Y_{0}| S_{0}=1,D_{1}>D_{0},Y_{0}\geq Q_{0}(1-1/p(x),x),X=x\right], \label{eq:betaL_cond_minus}\\
\beta_{U}^{-}(x)&:=E\left[Y_{1}| S_{1}=1,D_{1}>D_{0},X=x\right] \notag \\
&\quad -E\left[Y_{0}| S_{0}=1,D_{1}>D_{0},Y_{0}\leq Q_{0}(1/p(x),x),X=x\right], \label{eq:betaU_cond_minus}
\end{align}}
for $x\in\mathcal{X}^{-}_{ac}$. As a notational convention, set $\beta_B^\pm(x)=0$ for $x\in \mathcal{X}\setminus \mathcal X_{ac}$, $B\in\{L,U\}$. These values are used only in the reduced-forms $\beta_B^\pm(x)\pi_{ac}(x)$ entering the unconditional bounds. We note that when treatment does not cause selection, conditional effects bounds collapse to a point, i.e.~for any $x \in  (\mathcal{X}^0 \cap \mathcal{X}_{ac}) \subseteq \mathcal{X}^+_{ac}$ we have that \begin{align}
\beta_{L}^{+}(x)& = \beta_{U}^{+}(x) =E\left[Y_{1}| S_{1}=1,D_{1}>D_{0},X=x\right] - E\left[Y_{0}| S_{0}=1,D_{1}>D_{0},X=x\right].  \label{eq:betaU_cond_plus}
\end{align}
Now let
\begin{equation}
\mathbbm{1}^{+}(x):=\mathbbm{1}(x\in\mathcal{X}^{+}),\qquad \mathbbm{1}^{-}(x):=\mathbbm{1}(x\in\mathcal{X}^{-}), \label{eq:indicators1-def}
\end{equation}
and define for any $B \in \{L,U\}$
\begin{equation}
\beta_{B}(x):=\mathbbm{1}^{+}(x)\beta_{B}^{+}(x)+\mathbbm{1}^{-}(x)\beta_{B}^{-}(x).
\end{equation}
The unconditional SLATE parameter can be written as
\begin{equation}
\theta_{SLATE}=\frac{E[\theta_{SLATE}(X)\pi_{ac}(X)]}{\pi_{ac}},
\end{equation}
The corresponding sharp lower and upper bounds are
\begin{equation}
\beta_{L}=\frac{E[\beta_{L}(X)\pi_{ac}(X)]}{\pi_{ac}},\qquad
\beta_{U}=\frac{E[\beta_{U}(X)\pi_{ac}(X)]}{\pi_{ac}}.
\end{equation}

To make the link with observed data explicit, we next provide the estimands of the denominator and numerator moment functions in $\beta_{L}$ and $\beta_{U}$. Define the instrument propensity score $e(x):= P(Z=1| X=x)$, and the inverse probability weight

{\begin{equation}
W(Z,X):=\frac{Z}{e(X)}-\frac{1-Z}{1-e(X)}. \label{eq:ipw}
\end{equation}}
Extending Lemma~\ref{lem:IV_identities} to hold conditional on $X$, one can show that

{\begin{align}
\pi_{ac}&=E\left[\mathbbm{1}^{+}(X)\lambda_{0}(X)+\mathbbm{1}^{-}(X)\lambda_{1}(X)\right] \notag \\
&=E\left[\min\{\lambda_{0}(X),\lambda_{1}(X)\}\right]. \label{eq:Pr_ac_cov_uncond}
\end{align}}
where now

{\begin{align}
\lambda_{0}(x)&= E\left[W(Z,X)(D-1)S| X=x\right],\\
\lambda_{1}(x)&= E\left[W(Z,X)DS| X=x\right],\\
\mathbbm{1}^{+}(x)&=\mathbbm{1} (\lambda_{0}(x)\le \lambda_{1}(x)),\\
\mathbbm{1}^{-}(x)&=\mathbbm{1} (\lambda_{0}(x) > \lambda_{1}(x)).
\end{align}}
Thus $\pi_{ac}$ is identified. In addition, for any $y$ and $x \in \mathcal{X}_{ac}$,

{\begin{align}
F_{1}(y| x)&:=F_{Y_{1}| S_{1}=1,D_{1}>D_{0},X=x}(y)
=\frac{E\left[W(Z,X)DS\mathbbm{1}(Y\leq y)| X=x\right]}{E\left[W(Z,X)DS| X=x\right]}, \label{eq:F1_cond_main}\\
F_{0}(y| x)&:=F_{Y_{0}| S_{0}=1,D_{1}>D_{0},X=x}(y)
=\frac{E\left[W(Z,X)(D-1)S\mathbbm{1}(Y\leq y)| X=x\right]}{E\left[W(Z,X)(D-1)S| X=x\right]}. \label{eq:F0_cond_main}
\end{align}}
The conditional quantiles $Q_{d}(u,x)$ used for trimming in the conditional bounds in Equations  \eqref{eq:betaL_cond_plus} -- \eqref{eq:betaU_cond_minus} are defined as the generalized inverses of these CDFs and are thus identified. Using these identities, the following proposition collects the explicit estimands for the unconditional bounds $\beta_{L}$ and $\beta_{U}$.

\begin{proposition}
\label{prop:betaL_betaU_covariates}
Suppose Assumptions \ref{ass:conditional_IV}--\ref{ass:conditional_monotoneS} hold and $\pi_{ac}>0$. Then the sharp lower and upper bounds for $\theta_{SLATE}$ are identified as
{\begin{equation*}
\beta_{L}=\frac{E[\beta_{L}(X)\pi_{ac}(X)]}{\pi_{ac}},\qquad
\beta_{U}=\frac{E[\beta_{U}(X)\pi_{ac}(X)]}{\pi_{ac}},
\end{equation*}}
where $\pi_{ac}$ is given by Equation \eqref{eq:Pr_ac_cov_uncond} and
\begin{align*}
\beta_{L}(x) &:=\mathbbm{1}^{+}(x)\beta_{L}^{+}(x)+\mathbbm{1}^{-}(x)\beta_{L}^{-}(x), \\
\beta_{U}(x) &:=\mathbbm{1}^{+}(x)\beta_{U}^{+}(x)+\mathbbm{1}^{-}(x)\beta_{U}^{-}(x),
\end{align*}
with
{\small
\begin{align}
\beta_{L}^{+}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY\mathbbm{1}(Y\leq Q_{1}(p(x),x))+(1-D)SY\right)| X=x\right], \label{eq:num-plus-lower}\\
\beta_{L}^{-}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY+(1-D)SY\mathbbm{1}(Y\geq Q_{0}(1-1/p(x),x))\right)| X=x\right], \label{eq:num-minus-lower}\\
\beta_{U}^{+}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY\mathbbm{1}(Y\geq Q_{1}(1-p(x),x))+(1-D)SY\right)| X=x\right], \label{eq:num-plus-upper}\\
\beta_{U}^{-}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY+(1-D)SY\mathbbm{1}(Y\leq Q_{0}(1/p(x),x))\right)| X=x\right],\label{eq:num-minus-upper}
\end{align}}
on $\mathcal{X}_{ac}$, and $\beta_B^\pm(x)\pi_{ac}(x)=0$ by convention on $\mathcal X\setminus\mathcal X_{ac}$. $\lambda_d(x)$ and $W(z,x)$ are point identified on $\mathcal{X}$ and conditional quantiles $Q_d(\cdot,x)$ are point identified on $\mathcal X_{ac}$.
\end{proposition}

In summary, $(\beta_{L},\beta_{U})$ are functionals of the observed law of $(SY,S,D,Z,X)$ and a collection of nuisance functions (instrument propensity score, conditional means, and (inverted) conditional distribution functions). In the next section, we derive orthogonal moment conditions and semiparametric influence functions for these functionals, which will allow us to construct semiparametrically efficient debiased machine learning estimators for $(\beta_{L},\beta_{U})$ and confidence intervals for $\theta_{SLATE}$.

\section{Estimation and Inference}
\label{sec:inference}
\subsection{Nuisance Functions and Target Parameter}

The previous section shows that for any $B\in\{L,U\}$, the unconditional bound can be written as
\begin{equation}
\beta_{B}=\frac{E[\beta_{B}(X)\pi_{ac}(X)]}{\pi_{ac}}, \label{eq:betaB_cov_repeat}
\end{equation}
with $\beta_{B}(X)$ defined in Proposition \ref{prop:betaL_betaU_covariates} and $\pi_{ac}$ in \eqref{eq:Pr_ac_cov_uncond}. For notational convenience, we decompose the numerator of \eqref{eq:betaB_cov_repeat} as
\begin{equation}
E[\beta_{B}(X)\pi_{ac}(X)]=N_{0}^{+}+N_{0}^{-}+N_{B,1}^{+}+N_{B,1}^{-}, \label{eq:N0-NB-all1}
\end{equation}
where, for $B\in\{L,U\}$ and $\ast\in\{+,-\}$, we define
\begin{equation}
N_{0}^{\ast}:=E[\beta_{0}^{\ast}(X)\pi_{ac}(X)], \qquad
N_{B,1}^{\ast}:=E[\beta_{B,1}^{\ast}(X)\pi_{ac}(X)].
\end{equation}

Table \ref{tab:nuisance} contains all primary nuisance functions as well as all derived quantities that are used for the remainder of this section. We collect the required primary nuisance functions in vector
\begin{align}
\eta(x):=(e(x),\{r(z,x),m(z,x),\mu(z,x),\nu(z,x),F(\cdot|  d,z, x),G_{L}(\cdot| d,z, x),G_{U}(\cdot| d,z, x)\}_{d,z}).
\end{align}

\begin{table}[!h]
\caption{Primary and Derived Nuisance Functions for Bounds $(\beta_{L},\beta_{U})$}
\label{tab:nuisance}
\centering{\footnotesize
\begin{tabular}{c|c} \hline \\[-0.5ex]
Primary Nuisance Parameter in $\eta$ & Definition \\[1ex]\hline
&\\
$e(x)$ & $P(Z=1| X=x)$\\
$r(z,x)$ & $E[DS| Z=z,X=x]$\\
$m(z,x)$ & $E[(1-D)S| Z=z,X=x]$\\
$\mu(z,x)$ & $E[(1-D)SY| Z=z,X=x]$\\
$\nu(z,x)$ & $E[DSY| Z=z,X=x]$\\
$F(y| d,z,x)$ & $E[\mathbbm{1}(Y\leq y)| D=d,S=1,Z=z,X=x]$\\
$G_{L}(y| d,z,x)$ & $E[Y\mathbbm{1}(Y \leq y)| D=d,S=1,Z=z,X=x]$\\
$G_{U}(y| d,z,x)$ & $E[Y\mathbbm{1}(Y > y)| D=d,S=1,Z=z,X=x]$\\\hline
&\\[-0.5ex]
Derived Nuisance Parameter or Variable & Definition \\[1ex]\hline
&\\[-0.5ex]
$\lambda_{0}(x)$ & $-(m(1,x)-m(0,x))$\\
$\lambda_{1}(x)$ & $r(1,x)-r(0,x)$\\
$\pi_{ac}(x)$ & $\min\{\lambda_{0}(x),\lambda_{1}(x)\}$\\
$p(x)$ & $\dfrac{\lambda_{0}(x)}{\lambda_{1}(x)}$\\
$\mathbbm{1}^+(x)$ & $\mathbbm{1}(p(x) \leq 1)$ \\
$\mathbbm{1}^-(x)$ & $\mathbbm{1}(p(x) > 1)$ \\
$W(z,x)$ & $\dfrac{z}{e(x)}-\dfrac{1-z}{1-e(x)}$\\
$F_{1}(y| x)$
& $\dfrac{F(y| 1,1,x)r(1,x)-F(y| 1,0,x)r(0,x)}{r(1,x)-r(0,x)}$\\
$F_{0}(y| x)$
& $\dfrac{F(y| 0,1,x)m(1,x)-F(y| 0,0,x)m(0,x)}{m(1,x)-m(0,x)}$\\
$Q_{1}(u,x)$ & $\inf\{y\in\mathcal{Y}:F_{1}(y| x)\geq u\}$\\
$Q_{0}(u,x)$ & $\inf\{y\in\mathcal{Y}:F_{0}(y| x)\geq u\}$\\
$\psi_{L}^{+}(z,x)$ & $G_{L}(Q_{1}(p(X),X)| 1,z,x)$\\
$\psi_{U}^{+}(z,x)$ & $G_{U}(Q_{1}(1-p(X),X)| 1,z,x)$\\
$\psi_{L}^{-}(z,x)$ & $G_{L}(Q_{0}(1-1/p(X),X)| 0,z,x)$\\
$\psi_{U}^{-}(z,x)$ & $G_{U}(Q_{0}(1/p(X),X)| 0,z,x)$\\
&\\ \hline
\end{tabular}
}
\end{table}



\subsection{Semiparametric Influence Functions and Nuisance Functions} \label{sec:moment-functions1}
For any $B \in \{L,U\}$, we treat $\pi_{ac}$ in \eqref{eq:Pr_ac_cov_uncond} as well as the four components in \eqref{eq:N0-NB-all1} as target functionals and derive their semiparametric influence functions/efficient orthogonal moments. They are then combined to yield the efficient influence function and estimator for the respective $\beta_B$.
Denote the influence function operator $\mathbbm{IF}: \Theta \rightarrow L^2(P,\mathbb{R}^q)$ as the operator that, for a parameter $\theta:\mathcal{P}\rightarrow \mathbb{R}^q$, returns its influence function, i.e.,~the standard Riesz-representer of the pathwise derivative \citep{bickel1993efficient, kennedy2024semiparametric}.\footnote{The influence function is defined via the functional derivative of the target parameter with respect to perturbations along the tangent space evaluated at the true probability distribution. For example, if we observe iid data $X\sim \mathcal{P}$ and the target functional is $E[X]$, then for $\mathbbm{IF}(E[X]) = X - E[X]$. For a regular parameter under the nonparametric model $\mathcal{P}$, this is the unique semiparametric efficient influence function, see, e.g., \cite{hines2022demystifying}, \cite{kennedy2024semiparametric} or \cite{heiler2026heterogeneity} for additional examples.}

We now present the influence function for the target parameters $\beta_B$ for $B \in \{L,U\}$. We suppress dependence on nuisances and data in what follows whenever it does not cause confusion. By linearity of the $\mathbbm{IF}$ operator, we have that  \begin{align}
     \mathbbm{IF}(\beta_B)
     &= \frac{1}{\pi_{ac}}\left[ \mathbbm{IF}(N_0^+)  + \mathbbm{IF}(N_0^-) + \mathbbm{IF}(N_{B,1}^+) + \mathbbm{IF}(N_{B,1}^-) - \beta_B \mathbbm{IF}(\pi_{ac})\right] \label{eq:IF_betaB_general1}
\end{align}


\begin{table}[!h]
\caption{Influence Functions of Components} \label{tab:IF-components}
\centering
{\footnotesize
\begin{tabular}{>{\raggedright\arraybackslash}p{0.12\textwidth}|
                >{\centering\arraybackslash}p{0.82\textwidth}}
\hline \\[-1.5ex]
Component & Influence Function $\mathbbm{IF}$ \\ \hline \\[-0.5ex]

$\pi_{ac}$ &
$\begin{aligned}[t]
& -\mathbbm{1}^{+}(X)\Big[ W(Z,X)((1-D)S-m(Z,X)) + m(1,X)-m(0,X) \Big] \\
& + \mathbbm{1}^{-}(X)\Big[ W(Z,X)(DS-r(Z,X)) + r(1,X)-r(0,X) \Big] - \pi_{ac}
\end{aligned}$ \\[8ex]

$N_{0}^{+}$ &
$\mathbbm{1}^{+}(X)\Big\{ W(Z,X)((1-D)SY-\mu(Z,X)) + \mu(1,X)-\mu(0,X) \Big\}
- N_{0}^{+}$ \\[4ex]

$N_{L,1}^{+}$ &
$\begin{aligned}[t]
& \mathbbm{1}^{+}(X)\Big\{
- Q_{1}(p(X),X) W(Z,X)\Big[ (1-D)S-m(Z,X) \\
&\quad + (DS-r(Z,X))F(Q_{1}(p(X),X)|1,Z,X) \\
&\quad + DS(\mathbbm{1}(Y\le Q_{1}(p(X),X)) - F(Q_{1}(p(X),X)|1,Z,X)) \Big] \\
&\quad + W(Z,X)\Big[ DSY\mathbbm{1}(Y\le Q_{1}(p(X),X)) - \psi_L^{+}(Z,X)r(Z,X) \Big] \\
&\quad + \psi_L^{+}(1,X)r(1,X) - \psi_L^{+}(0,X)r(0,X)
\Big\}
- N_{L,1}^{+}
\end{aligned}$ \\[24ex]

$N_{U,1}^{+}$ &
$\begin{aligned}[t]
& \mathbbm{1}^{+}(X)\Big\{
- Q_{1}(1-p(X),X) W(Z,X)\Big[ (1-D)S-m(Z,X) \\
&\quad + (DS-r(Z,X))(1-F(Q_{1}(1-p(X),X)|1,Z,X)) \\
&\quad + DS(\mathbbm{1}(Y>Q_{1}(1-p(X),X)) - (1-F(Q_{1}(1-p(X),X)|1,Z,X))) \Big] \\
&\quad + W(Z,X)\Big[ DSY\mathbbm{1}(Y>Q_{1}(1-p(X),X)) - \psi_U^{+}(Z,X)r(Z,X) \Big] \\
&\quad + \psi_U^{+}(1,X)r(1,X) - \psi_U^{+}(0,X)r(0,X)
\Big\}
- N_{U,1}^{+}
\end{aligned}$ \\[24ex]

$N_{0}^{-}$ &
$\mathbbm{1}^{-}(X)\Big\{ W(Z,X)(DSY-\nu(Z,X)) + \nu(1,X)-\nu(0,X) \Big\}
- N_{0}^{-}$ \\[4ex]

$N_{L,1}^{-}$ &
$\begin{aligned}[t]
& \mathbbm{1}^{-}(X)\Big\{
- Q_{0}(1-\tfrac{1}{p(X)},X) W(Z,X)\Big[ DS-r(Z,X) \\
&\quad + ((1-D)S-m(Z,X))(1-F(Q_{0}(1-\tfrac{1}{p(X)},X)|0,Z,X)) \\
&\quad + (1-D)S(\mathbbm{1}(Y\ge Q_{0}(1-\tfrac{1}{p(X)},X)) - (1-F(Q_{0}(1-\tfrac{1}{p(X)},X)|0,Z,X))) \Big] \\
&\quad + W(Z,X)\Big[ (1-D)SY\mathbbm{1}(Y\ge Q_{0}(1-\tfrac{1}{p(X)},X)) - \psi_L^{-}(Z,X)m(Z,X) \Big] \\
&\quad + \psi_L^{-}(1,X)m(1,X) - \psi_L^{-}(0,X)m(0,X)
\Big\}
- N_{L,1}^{-}
\end{aligned}$ \\[24ex]

$N_{U,1}^{-}$ &
$\begin{aligned}[t]
& \mathbbm{1}^{-}(X)\Big\{
- Q_{0}(\tfrac{1}{p(X)},X) W(Z,X)\Big[ DS-r(Z,X) \\
&\quad + ((1-D)S-m(Z,X))F(Q_{0}(\tfrac{1}{p(X)},X)|0,Z,X) \\
&\quad + (1-D)S(\mathbbm{1}(Y\le Q_{0}(\tfrac{1}{p(X)},X)) - F(Q_{0}(\tfrac{1}{p(X)},X)|0,Z,X)) \Big] \\
&\quad + W(Z,X)\Big[ (1-D)SY\mathbbm{1}(Y\le Q_{0}(\tfrac{1}{p(X)},X)) - \psi_U^{-}(Z,X)m(Z,X) \Big] \\
&\quad + \psi_U^{-}(1,X)m(1,X) - \psi_U^{-}(0,X)m(0,X)
\Big\}
- N_{U,1}^{-}
\end{aligned}$ \\[24ex]

\hline
\end{tabular}
}
\end{table}


The influence functions of the components are in Table \ref{tab:IF-components}.
Putting upper and lower bounds together then yields \begin{align}
    \mathbbm{IF}(\beta) = \begin{pmatrix}
        \mathbbm{IF}(\beta_L) \\
        \mathbbm{IF}(\beta_U)
    \end{pmatrix}
\end{align}
The influence functions $\mathbbm{IF}(\beta_B)$ for $B \in \{L,U\}$ nest the existing literature on LATE and sample selection bounds. In particular, we obtain the following proposition: \begin{proposition} \label{prop:nesting-of-IF}
 For any $B \in \{L,U\}$, the efficient influence function in \eqref{eq:IF_betaB_general1} under (i) perfect compliance, (ii) no sample selection or (iii) both collapse to their respective efficient influence functions for Lee bounds, LATE, or ATE:
\begin{enumerate}[label=(\roman*), leftmargin=*]
    \item If $P(D = Z) = 1$, then $\mathbbm{IF}(\beta_B) = \mathbbm{IF}(\beta_B^{Lee})$ a.s.
    \item If $P(S = 1) = 1$, then $\mathbbm{IF}(\beta_B) = \mathbbm{IF}(\theta_{LATE})$ a.s.
    \item If $P(S=1, D=Z) = 1$, then $\mathbbm{IF}(\beta_B) = \mathbbm{IF}(\theta_{ATE})$ a.s.
\end{enumerate}

\end{proposition}
Practically, this means that, as $P(D = Z) \rightarrow 1$ or $P(S = 1) \rightarrow 1$, our influence functions will get closer in probability to the  semiparametrically efficient influence functions for Lee-bounds on the intensive margin ATE \citep{heiler2024treatmentevaluationintensiveextensive,semenova2025generalized} or the LATE \citep{frolich2007nonparametric} respectively. If both apply, the functions approach the efficient influence of the ATE \citep{hahn1998role}.

\subsection{Finite Sample Implementation} \label{sec:estimation1}
We now outline the steps to estimate bounds and effect confidence intervals and provide more details regarding nuisance function estimation. For a generic random variable $X$ we denote $E_n[X] = \frac{1}{n}\sum_i^n X_i$ in what follows.  Assume we have independent data $O_i = (S_iY_i, S_i, D_i, Z_i, X_i)'$ for $i=1,\dots,n$. Algorithm \ref{alg:estimation} contains a step-by-step explanation of how to obtain bounds and inference. Estimation of bounds is fairly standard within the DML framework for composite ratio parameters. In particular, we estimate all their components separately with cross-fitted nuisances to obtain an estimate for the eventual bounds and their respective influence functions. The latter then yield the full variance-covariance matrix estimates that can be used for standard-normal-based inference on bounds or, more importantly, confidence intervals for the effect using refined critical values \citep{imbens2004confidence,stoye2020simple}.
All of these have at least $(1-\alpha)$ asymptotic coverage under some regularity conditions on the conditional distribution and suitable rate conditions on the nuisances, see Section \ref{sec:largesample1} for more details.

\begin{algorithm}[!h]
  \caption{Estimation and Inference} \label{alg:estimation}
  \begin{algorithmic}[1]
   \Require Data $\{O_i\}_{i=1}^n$, number of folds $K$.
   \State Randomly partition data into $K$ disjoint folds $I_1,\dots,I_K$ of approximately equal size.
    \For{each $f \in \{1,\dots,K\}$}
    \State estimate primary nuisance parameters using data in $I_f^c$.
    \State evaluate the nuisances on $I_f$.
    \EndFor

    \For{each $N \in \{N_0^+,N_0^-,N_{L,1}^+,N_{L,1}^-,N_{U,1}^+,N_{U,1}^-,\pi_{ac}\}$}
    \State for all $i=1,\dots,n$ plug in all nuisances $\hat{\eta} $ into the orthogonal moments $\mathbbm{IF}(\cdot)$.
      \State obtain $\hat{N}$ by setting the sample mean of the orthogonal moments equal to zero $$E_n[\mathbbm{IF}(\hat{N},\hat{\eta})] = 0$$
    \EndFor
    \State
    \Return the bound estimators as $$\hat{\beta}_B = \frac{\hat{N}_0^+ + \hat{N}_0^- + \hat{N}_{B,1}^+ + \hat{N}_{B,1}^-}{\hat{\pi}_{ac}}$$
    \State   \Return standard errors via estimated composite moment: $$\hat{\sigma}_{B,n} = \sqrt{\frac{E_n[\mathbbm{IF}(\hat{\beta}_B,\hat{\eta})^2]}{n}}$$
    where the composite moment is given by $$\mathbbm{IF}(\hat{\beta}_B,\hat{\eta}) = \frac{1}{\hat{\pi}_{ac}}\left[ \mathbbm{IF}(\hat{N}_0^+,\hat{\eta})  + \mathbbm{IF}(\hat N_0^-,\hat{\eta}) + \mathbbm{IF}(\hat N_{B,1}^+,\hat{\eta}) + \mathbbm{IF}(\hat N_{B,1}^-,\hat{\eta}) - \hat\beta_B \mathbbm{IF}(\hat\pi_{ac},\hat{\eta})\right]$$
    \State \Return $(1-\alpha)$ effect confidence intervals as
    \begin{align*}
        CI_{1-\alpha}(\theta_{SLATE}) = \left[\hat\beta_L - c_{L,\alpha}\hat\sigma_{L,n},\  \hat\beta_U + c_{U,\alpha}\hat\sigma_{U,n}\right]
    \end{align*}
    where refined critical values $c_{L,\alpha},c_{U,\alpha} \leq z_{1-\alpha/2}$ can be chosen according to \cite{imbens2004confidence} or \cite{stoye2020simple}. Using $z_{1-\alpha/2}$ is valid for inference on the bounds.
  \end{algorithmic}
\end{algorithm}

We now discuss primary nuisance function estimation and how to obtain derived nuisances as defined in Table \ref{tab:nuisance}. Our high-level assumptions in Section \ref{sec:largesample1} match these primitive objects that are conditional means/probabilities, conditional CDFs, and trimmed conditional means.
For primary nuisance functions that are simple conditional means or probabilities, $e,r,m,\mu, \nu$, a plethora of off-the-shelf nonparametric and machine learning methods such as neural networks, forests or high-dimensional sparse parametric models are available.
The derived nuisances $\lambda_0,\lambda_1, \pi_{ac}, p, \mathbbm{1}^+, \mathbbm{1}^-, W$ only depend on the primary and can be obtained via simple plug-in versions.

The functional parameters $F$ and $G_B$ require special attention. In particular, we suggest to estimate the primary conditional CDF $F$ via machine learning analogues of distributional regression \citep{foresi1995conditional,klein2024distributional}. We also recommend direct imposition or post-processing, e.g., via isotonic regression \citep{henzi2021isotonic} that further refine these estimates by enforcing nondecreasing estimated conditional CDFs. Together with $r$ and $m$, these yield plug-in versions of $F_0$ and $F_1$. The $G_B$ components can be similarly obtained and refined as a sequence of regressions with outcome variable equal to trimming indicator times actual outcome.

To obtain the derived conditional quantiles $Q_0$ and $Q_1$, inversion of the previously obtained conditional CDFs, $F_0$ and $F_1$, can be used. These yield the trimming indicators that are required to evaluate the conditional expectation models $G_B$ at their respective trimming quantiles to obtain derived nuisances $\psi_B$.
The inversion-based approach ensures algebraic compatibility between the different nuisance quantities. Thus, we use this approach in all of our simulations and applications in Section \ref{sec:JC} as well as Appendix \ref{sec:simulation} and \ref{sec:OHIE}.


\subsection{Large Sample Properties} \label{sec:largesample1}

We now present additional technical assumptions and the resulting large sample properties of the efficient influence function based estimators of the causal effect bounds.
Denote the true nuisances as $\eta(x) =: \eta \in \mathcal{H}$ where $\mathcal{H}$ is a convex subset of a suitably normed vector space. Denote $\mathcal{H}_n \subset \mathcal{H}$ the realization set of the estimated nuisance quantities $\hat{\eta}(x) =: \hat{\eta}$, i.e.,~the set containing estimated nuisances with probability $1-u_n$ where $u_n = o(1)$. All nuisances are cross-fitted according to Definition 3.2 in \cite{chernozhukov2018double}, see Algorithm \ref{alg:estimation}.

For the remainder, write for generic nuisance $h=h(x)$ its estimation error $\Delta h=\hat h-h$. Denote $\|\cdot\|_2$ as $L^2(P)$ norm and $\|\cdot\|_\infty$ as the uniform norm.
By abuse of notation, if the object depends on $Z=z$ and/or $D=d$, we suppress dependence and take all norms to be uniform over these finite dimension as well, e.g., \begin{align}
    ||r||_p = \sup_{z,d \in \{0,1\}}||r(z,x)||_p  =\sup_{z \in \{0,1\}}||r(z,x)||_p
    = \sup_{z \in \{0,1\}} \left(\int r^p(x,z)dP(x)\right)^{1/p}
\end{align}
and equivalently for other nuisances. For a given $x$, we also denote
$\|\cdot\|_{\infty,\mathcal{N}_x}$ as the uniform norm over a neighborhood $\mathcal{N}_x$. In particular, for a generic object $A(\cdot|x)$,
\begin{align}
||\hat{A}(y|x) - A(y|x)||_{\infty,\mathcal{N}_x} &= \sup_{y \in \mathcal{N}_x}|\hat{A}(y|x) - A(y|x)|.
\end{align}
We denote the supremum over these neighborhoods as \begin{align}
||\hat{A}(y|x) - A(y|x)||_{\infty,\mathcal{N}} &= \sup_x\sup_{y \in \mathcal{N}_x}|\hat{A}(y|x) - A(y|x)|.
\end{align} Moreover, we write shorthand \begin{align}
 ||\Delta G ||_p &= ||\Delta G_L||_{p,\mathcal{N}} + ||\Delta G_U||_{p,\mathcal{N}}
\end{align} and $a_n \lesssim b_n$ and $a_n \lesssim_P b_n$, whenever $a_n = O(b_n)$ or $a_n = O_p(b_n)$ respectively. If not stated differently, the following assumptions are all uniformly over $n$.





\noindent {\textbf{Assumption A (Regularity, Overlap, and Learning Rates)}}
\textit{
\begin{enumerate}[itemsep=0pt] \singlespacing
\item[A.1] (Moments) The conditional potential outcome moments are bounded, i.e.,~for some $m > 0$, \begin{align*}\sup_{x\in\mathcal{X},d\in\{0,1\}}E[|Y(d)|^{2+m}|X=x] \lesssim 1.\end{align*}
\item[A.2] (Eigenvalues) The variance-covariance matrix of the influence function $E[\mathbbm{IF}(\beta)\mathbbm{IF}(\beta)']$ has finite eigenvalues bounded away from zero.
\item[A.3] (No Point Mass at Trimming Points) The conditional distributions are continuous at the trimming points. For $z \in \{0,1\}$, and $x$ on the relevant support,\footnote{Here and in A.4, the relevant support is $\mathcal X^+_{ac}$ for conditions involving $(Q_1(p(x),x)$ or $Q_1(1-p(x),x)$, and $\mathcal X^-_{ac}$ for
conditions involving $Q_0(1-1/p(x),x)$ or $Q_0(1/p(x),x)$.}\begin{align*}
    &P(Y=Q_1(p(x),x)| DS=1,Z=z,X=x) = 0,\\
    &P(Y=Q_1(1-p(x),x)| DS=1,Z=z,X=x) = 0,\\
&P(Y=Q_0(1-1/p(x),x)| (1-D)S=1,Z=z,X=x) = 0, \\
&P(Y=Q_0(1/p(x),x)| (1-D)S=1,Z=z,X=x) = 0.
\end{align*}
\item[A.4] (Bounded Mixture Outcome Density and Local Lipschitz Trimmed Means) Let $f_1(\cdot|x)$ and $f_0(\cdot|x)$ denote the respective densities of $F_1(\cdot|x)$ and $F_0(\cdot|x)$. Assume they are bounded at the trimming thresholds and the trimmed conditional means are locally Lipschitz, i.e.,~there exist a $C>0$ and constants $0<f_{\min}\le f_{\max}<\infty$, $L_G<\infty$ such that (i) for all $x$ on the relevant support:
\begin{align*}
&f_1(Q_1(p(x),x)\mid x) \in [f_{\min},f_{\max}], \\
&f_1(Q_1(1-p(x),x)\mid x) \in [f_{\min},f_{\max}], \\
&f_0(Q_0(1-1/p(x),x)\mid x) \in [f_{\min},f_{\max}], \\
&f_0(Q_0(1/p(x),x)\mid x) \in [f_{\min},f_{\max}],
\end{align*}
and (ii) uniformly in $(z,x)$, for $|u|\le C$,
\begin{align*}
&|G_L(Q_1(p(x),x) + u|1,z,x) - G_L(Q_1(p(x),x)|1,z,x)| \leq L_G |u|, \\
&|G_U(Q_1(1-p(x),x) + u|1,z,x) - G_U(Q_1(1-p(x),x)|1,z,x)| \leq L_G |u|, \\
&|G_L(Q_0(1-1/p(x),x) + u|0,z,x) - G_L(Q_0(1-1/p(x),x)|0,z,x)| \leq L_G |u|, \\
&|G_U(Q_0(1/p(x),x) + u|0,z,x) - G_U(Q_0(1/p(x),x)|0,z,x)| \leq L_G |u|.
\end{align*}
\item[A.5] (Margin Condition) The distribution of the positive and negative monotonicity type is well-behaved around the margin of indifference, i.e.,~there exist $C_M<\infty$ and $\kappa>0$ such that
\[
P\!\left(\big|-[m(1,X)-m(0,X)]-[r(1,X)-r(0,X)]\big|\le t\right)\le C_M\,t^\kappa\quad\text{for all }t>0.
\]
\item[A.6] (Strong Overlap) There are comparable units across instrument levels and within the differently treated and selected populations. Moreover, always-selected complier probabilities are bounded away from
zero on their relevant supports, i.e.,~there exists some $\underline{c} \in (0,1)$ such that \begin{align*}
    \underline{c} < \inf_{z,x}\{r(z,x),m(z,x),e(x)\} \leq \sup_{z,x}\{r(z,x),m(z,x),e(x)\} < 1-\underline{c}.
\end{align*}
and
 \begin{align*}
\inf_{x\in\mathcal X_{ac}^+}-[m(1,x) - m(0,x)]> \underline{c},\qquad
\inf_{x\in\mathcal X_{ac}^-}[r(1,x) - r(0,x)]> \underline{c}.
\end{align*}
\item[A.7] (Machine Learning Bias) Let $u_n = o(1)$. For all folds, the nuisance parameters obtained via cross-fitting belong to a shrinking neighborhood $\mathcal{H}_n$ around $\eta$ with probability of at least $1-u_n$, such that, uniformly over the neighborhood, the nuisance functions are consistent
\begin{align*}
    \|\Delta\mu\|_2\ + \|\Delta\nu\|_2\ + \|\Delta e\|_2\ + \|\Delta m\|_{\infty}\ + \|\Delta r\|_{\infty} + ||\Delta G||_{\infty,\mathcal{N}} + ||\Delta F ||_{\infty,\mathcal{N}}  &= o(1),
\end{align*}
and obey convergence rates \begin{align*}
 &(\|\Delta\mu\|_2\ + \|\Delta\nu\|_2) \|\Delta e\|_2\ + (||\Delta m||_{\infty} + ||\Delta r||_{\infty})^{\kappa + 1}  +(||\Delta e ||_2 + ||\Delta m ||_2 + ||\Delta r ||_2) \times \\ &\bigg[||\Delta e ||_{\infty} + ||\Delta m ||_{\infty} + ||\Delta r ||_{\infty} + ||\Delta G||_{\infty,\mathcal{N}} + ||\Delta F ||_{\infty,\mathcal{N}} \bigg] = o(n^{-1/2}).
\end{align*}
\end{enumerate}}

Assumption A.1 and A.2 are simple regularity conditions that rule out heavy tails and degenerate DGPs where bounds are close or equal to a point or otherwise degenerate.

Assumption A.3 rules out point masses at the trimming thresholds in the observed selected outcome distributions used to construct the reduced-form nuisance functions. This ensures that the trimming indicators are unambiguous at the cutoff and that weak and strict inequality conventions coincide at the relevant thresholds.\footnote{The theory can be extended to mass points as discussed in Section \ref{sec:bounds}. However, in contrast to identification where mass points are harmless once one uses fractional trimming \citep{huber2015sharp,kitagawa2021theidentificationregion}, inference can be affected as the functional may be nonregular. This problem arises only in boundary cases where the trimming probability coincides exactly with the edge of a mass point. For DGPs in which the cutoff lies in the interior of a mass point, fractional trimming yields a regular functional, so standard root-$n$ semiparametric inference can proceed using the corresponding influence function.}

Assumption A.4(i) is complementary to A.3 and provides primitive density regularity that guarantees local invertibility and differentiability. It controls error propagation through the quantile mapping (Bahadur-type representation). A.4(ii) adds local smoothness to the trimmed mean functions, ensuring uniform control over the nuisance functions evaluated close to the trimming points.

Assumption A.5 is a margin condition that controls the mass of observations near the boundary between positive and negative selection monotonicity types among compliers. In particular, it rules out excessive concentration of probability that would make correct classification difficult. In the case of a bounded density, A.5 holds with $\kappa = 1$, see, e.g.,~\cite{audibert2007fast} or \cite{heiler2024treatmentevaluationintensiveextensive} for related assumptions and discussion. The larger $\kappa$, the less demanding the convergence requirements in A.7 for learning the conditional joint probability of being selected and in treatment/control status. It implies the necessary regularity condition $P(\mathcal{X}^0) = 0$.

Assumption A.6 assures that there are comparable units for instrument and the treatment within the selected group and a relevant share of always-selected compliers. Strong overlap is imposed to obtain a finite variance bound and avoid irregular identification \citep{khan2010irregular,HEILER2021valid}.

Assumption A.7 imposes the standard DML requirement that all nuisances are consistent and converge to their truth sufficiently fast for the second-order remainder of the orthogonal von Mises expansion to be \(o(n^{-1/2})\), but it does so in a relatively weak form tailored to our target parameter and estimation procedure and its primitives.  In particular, in contrast to much of the Lee bounds-type literature, we do not assume global uniform convergence rates for the estimated quantiles or trimmed mean functionals. Instead, it only imposes local sup-norm control of the estimated CDFs
\(F\) and trimmed means \(G_B\) on neighborhoods relevant for trimming, with the regularity of quantiles and trimmed means derived from these local conditions. The first product term in A.7 matches the usual \(L^2\)-rate requirement in \cite{chernozhukov2018double}, while the additional terms capture the non-smooth features of monotonicity type classification and trimming indicators. This is related to rate conditions in \cite{heiler2024heterogeneous} and \cite{semenova2025generalized} for generalized Lee bounds as well as \cite{heiler2024treatmentevaluationintensiveextensive} for more general intensive and extensive margin treatment effects, but expressed directly in terms of CDF errors
rather than through separate uniform convergence assumptions on the associated quantile-based
functionals. Examples for global uniform rates of nonparametric and machine learning estimation of conditional CDFs can be found in, e.g., \cite{xie2023uniform} or \cite{cattaneo2025uniformestimationinferencenonparametric} respectively.

We obtain the following Theorem: \begin{theorem} \label{thm_asyN1}
    Under Assumptions \ref{ass:conditional_IV}, \ref{ass:conditional_monotoneS}, and A.1--A.7, the estimated bounds are jointly asymptotically normal and semiparametrically efficient, i.e.\begin{align*}
    \sqrt{n}(\hat{\beta} - \beta) \overset{d}{\rightarrow} \mathcal{N}\big(0,E[\mathbbm{IF}(\beta)\mathbbm{IF}(\beta)']\big).
\end{align*}
\end{theorem}
Theorem \ref{thm_asyN1} can be used directly for inference on the bounds using the usual standard normal critical values. Importantly, the assumptions imply that the identified set always has a non-empty interior in the population. Thus, Theorem \ref{thm_asyN1} is sufficient to construct tighter confidence intervals for the effect $\theta_{SLATE}$ directly using \cite{imbens2004confidence} or \cite{stoye2020simple} critical values. We use the latter in our empirical applications.

\section{Always-selected Complier Profiling} \label{sec:profiling}
 $\theta_{SLATE}$ is the average treatment effect for the target population $ac$, the always-selected compliers. While individual members of this group are not identified and the group may differ from observed subpopulations or the overall population in both observed and unobserved characteristics, its observable features can nevertheless be characterized, much like complier profiling for the LATE.\footnote{See, for example, \cite{abadie2003semiparametric}, \cite{angrist2004treatmenteffectheterogeneity}, and \cite{singh2023doublerobustnessforcomplier} for various approaches to complier profiling.} This makes it possible to compare always-selected compliers with other populations of interest and thereby better assess the external validity and substantive importance of the effect bounds.

Observable $ac$ characteristics can be identified using the same ingredient, $\pi_{ac}(x)$, as the SLATE bounds in Section \ref{sec:bounds_covariates}. Debiased estimation and inference follow analogously from our influence functions. In particular, for any integrable $g:\mathcal{X}\rightarrow \mathbb{R}$, consider target parameter \begin{align}
    E[g(X)|ac] = \frac{E[g(X)\pi_{ac}(X)]}{E[\pi_{ac}(X)]}.
\end{align}
This is the average $g(x)$ in the population of always-selected compliers. Its influence function is given by
\begin{align}
    \mathbbm{IF}(E[g(X)|ac]) &= \frac{1}{\pi_{ac}}\big(\left(\mathbbm{IF}(\pi_{ac}) + \pi_{ac}\right)\left(g(X) - E[g(X)|ac]\right)\big), \label{eq:IFg(X)}
\end{align}
where nuisances and influence functions of the components can be found in Section \ref{sec:moment-functions1}. The corresponding estimator can be obtained by solving the empirical analogue as in Algorithm \ref{alg:estimation}.
It is root-$n$ consistent and asymptotically normal analogously to Theorem \ref{thm_asyN1}. Moreover, influence function \eqref{eq:IFg(X)} can directly be used for statistically valid comparisons between mean covariates $g(x)$ of $ac$ and other populations such as unconditional, treated or assigned units. We provide some specific examples in Section \ref{sec:JC} and Appendix \ref{sec:OHIE}.


\section{Empirical Study I: Job Corps Revisited}\label{sec:JC}
\subsection{Data and Methods}
In this section, we re-evaluate the earnings effect of participating in JC, a large US federally funded training program providing free academic education, vocational training and employment assistance to disadvantaged youth. We make use of the National Job Corps Study by Mathematica Policy Research. This experiment implemented stratified randomized assignment of applicants, incorporating over 15,400 individuals between ages 16 and 24. Multiple outcomes such as earnings and job status were gathered at various points after assignment.

Our data and main variables are identical to \cite{chen2015bounds}. In particular, we use a subset of 9,090 units (3,599 control and 5,491 treated) with non-missing work hours, earnings, and participation information. The outcome is log hourly wages which is only observed for the employed. Assignment is given by the original randomization. Treatment is defined as eventual JC participation within the 208 weeks of the evaluation period. We additionally make use of
socio-economic pre-assignment covariates including job and earnings history, education, parental background and more, matching the variables used in \cite{lee2009training} for ITT bounds. The list of covariates along with sample summary statistics are provided in Table \ref{tab:jc_balance}.

The data are suitable for our method: Assignment was based on stratified randomization justifying conditional independence. Exclusion is credible as any earnings effects likely require actual training and not just assignment. Importantly, there was a significant amount of noncompliance with the randomized treatment assignment. In particular, only 73.8\% of individuals assigned to the treatment group ever participated in JC. Additionally, a small fraction (4.4\%) of individuals assigned to the control group also ended up participating.\footnote{Among the 4.4\%, 1.2\% of controls enrolled in JC before the end of the
embargo, while 3.2\% enrolled afterward.}

We evaluate the always-selected complier effect $\theta_{SLATE}$ at week $208$ after assignment using (i) \cite{chen2015bounds} bounds and (ii) sharp-basic bounds -- both assuming strong sample selection monotonicity -- as well as (iii) sharp DML bounds with covariates under weak monotonicity. Implementation of (i) follows \cite{chen2015bounds}. (ii) uses simple sample analogues of \eqref{eq:betaL} and \eqref{eq:betaU}. (iii) is obtained via the procedure in Section \ref{sec:estimation1}, see Appendix \ref{app:supp-empirical} for additional details. We also provide a profiling analysis of always-selected compliers using the method from Section \ref{sec:profiling}.

\subsection{Results: Bounds and Shares}
\begin{table}[!htbp]
\caption{Bounds for the Effect of Job Corps on Log Wages}
\label{tab:jc_bounds2}\centering
{\footnotesize
\begin{tabular}{@{}lccc}
\toprule & CF & Sharp-basic & Sharp-DML \\
\midrule Estimates & $[-0.022,\ 0.130]$ & $[0.019,\ 0.067]$ & $[-0.013,\ 0.124]$
\\
Standard Errors & --- & $(0.023,\ 0.022)$ & $(0.023,\ 0.022)$ \\
95\% Confidence Interval & $[-0.061,\ 0.168]$ & $[-0.018,\ 0.104]$ & $[-0.051,\ 0.160]$ \\[1ex]
Share of positive sample selection $^{a}$ & \multicolumn{2}{c}{$1 \text{ (assumed)}$} & \multicolumn{1}{l}{$0.933\ (0.004)$}  \\
Share of $ac$ $^{b}$ & \multicolumn{2}{c}{$0.391\ (0.010)$} & \multicolumn{1}{l}{$0.393\ (0.010)$} \\
\bottomrule &  &  &
\end{tabular}
}
{\footnotesize
\begin{minipage}{0.88\linewidth}
\footnotesize
\textit{Notes:} $N=9,090$. All calculations use design weights.
\textbf{CF} refers to the \cite{chen2015bounds} bounds under strong sample selection monotonicity without covariates using half-median-unbiased estimates with 95\% confidence intervals following \cite{chernozhukov2013intersection}.
CF reports no standard errors as inference relies on the CLR projection.
\textbf{Sharp-basic} refers to the sharp bounds without covariates under strong sample selection monotonicity.
\textbf{Sharp-DML} refers to the sharp bounds estimated by DML under weak sample selection monotonicity.
Standard errors for the two bounds are in parentheses as
$(\widehat{SE}_{\mathrm{lower}},\,\widehat{SE}_{\mathrm{upper}})$,
and the 95\% CI for $\theta_{SLATE} = E[Y_{1}-Y_{0}\mid ac]$ is calculated using \cite{stoye2020simple}.
\\
$^{a}$ Estimated share of positive sample selection~$= E[\mathbb{I}^{+}(X)]$, the fraction of observations
in the positive sample selection class (Sharp-DML only, CF and Sharp-basic assume this fraction
equals~1).\\
$^{b}$ Estimated share of always-selected compliers $\pi_{ac}$.
\end{minipage}
}
\end{table}





Table \ref{tab:jc_bounds2} reports the evaluation results. It delivers four main messages. First, the target population is empirically
relevant: the estimated share of always-selected compliers is about 39\% across specifications, so $\theta_{SLATE}$ pertains to a sizable subpopulation. Second, the data do not support imposing strong sample selection monotonicity: the estimated share in the positive sample selection region is 93.3\% and
strong sample selection monotonicity is rejected ($p<0.01$).\footnote{Negative employment effects at week 208 are economically plausible because
employment responses to Job Corps are heterogeneous across applicants. Consistent with this, \cite{semenova2025generalized} shows that positive conditional employment effects do not arise for all applicants
even four years after random assignment and documents subgroups with significantly
negative employment effects at later horizons.} Thus,
Sharp-DML appears to be the most credible specification.
Third, under the strong monotonicity benchmark without covariates, the sharp bounds and effect confidence interval are substantially tighter than CF, by 68.4\% and 46.7\%, respectively. Moreover, the estimated basic bounds no longer include zero and the corresponding confidence interval rules out negative effects below  -1.8\%.
Fourth, estimates using Sharp-DML are broadly comparable to CF despite relying only on weak sample selection monotonicity. Both rule out negative effects beyond -5.1\% and -6.1\% respectively. However, despite imposing less restrictive assumptions, Sharp-DML delivers tighter bounds and confidence intervals than CF by 9.8\% and 7.9\% respectively, highlighting the relevance of both sharpness and efficient inclusion of covariates.


The results also speak to the importance of the no-scaling result in Proposition \ref{prop:Lee_not_scaling}. In particular, naively scaling basic ITT bounds $[-0.019,0.093]$ \citep{lee2009training} with the JC compliance probability of 69.4\% would suggest SLATE bounds of $[-0.027,  0.134]$ which vastly exceed our Sharp-basic bounds of $[0.019, 0.067]$.




\subsection{Always-selected Complier Profiling}
\begin{table}[!h]
\centering
\caption{Job Corps Covariates: Always-Selected Compliers vs.\ Full Sample}
\label{tab:jc_profiling}
\footnotesize
\begin{tabular}{@{}l l l l@{}}
\toprule
Covariate & $ac$ & Full sample & ~\quad ~Difference \\
\midrule
Female                        & $0.400$~$(0.014)$ & $0.443$~$(0.005)$ & 
  \makebox[3.2em][r]{$-0$}.$044$~$(0.013)^{***}$
 \\
Age (in yrs.) at baseline     & $18.45$~$(0.058)$ & $18.44$~$(0.023)$ & 
  \makebox[3.2em][r]{$0$}.$013$~$(0.053)$
 \\
Black, non-Hispanic           & $0.433$~$(0.014)$ & $0.500$~$(0.005)$ & 
  \makebox[3.2em][r]{$-0$}.$067$~$(0.013)^{***}$
 \\
Hispanic                      & $0.194$~$(0.010)$ & $0.172$~$(0.004)$ & 
  \makebox[3.2em][r]{$0$}.$022$~$(0.009)^{**}$
 \\
Other race/ethnicity          & $0.067$~$(0.007)$ & $0.071$~$(0.003)$ & 
  \makebox[3.2em][r]{$-0$}.$004$~$(0.007)$
 \\
Married                       & $0.015$~$(0.004)$ & $0.022$~$(0.002)$ & 
  \makebox[3.2em][r]{$-0$}.$007$~$(0.004)^{*}$
 \\
Living together               & $0.033$~$(0.006)$ & $0.041$~$(0.002)$ & 
  \makebox[3.2em][r]{$-0$}.$008$~$(0.005)$
 \\
Separated                     & $0.019$~$(0.004)$ & $0.023$~$(0.002)$ & 
  \makebox[3.2em][r]{$-0$}.$004$~$(0.004)$
 \\
Has children                  & $0.162$~$(0.011)$ & $0.204$~$(0.004)$ & 
  \makebox[3.2em][r]{$-0$}.$042$~$(0.010)^{***}$
 \\
Number of children            & $0.214$~$(0.017)$ & $0.291$~$(0.007)$ & 
  \makebox[3.2em][r]{$-0$}.$077$~$(0.017)^{***}$
 \\
Education (in yrs.)           & $10.23$~$(0.042)$ & $10.12$~$(0.016)$ & 
  \makebox[3.2em][r]{$0$}.$111$~$(0.039)^{***}$
 \\
Mother's education            & $11.54$~$(0.064)$ & $11.48$~$(0.024)$ & 
  \makebox[3.2em][r]{$0$}.$061$~$(0.059)$
 \\
Father's education            & $11.54$~$(0.060)$ & $11.45$~$(0.023)$ & 
  \makebox[3.2em][r]{$0$}.$096$~$(0.054)^{*}$
 \\
Ever arrested                 & $0.235$~$(0.011)$ & $0.248$~$(0.005)$ & 
  \makebox[3.2em][r]{$-0$}.$013$~$(0.011)$
 \\

Household income: & & & \\
\quad\quad\quad\quad$[\$ 3,000, \$ 6,000)$  & $0.189$~$(0.009)$ & $0.208$~$(0.003)$ & 
  \makebox[3.2em][r]{$-0$}.$019$~$(0.008)^{**}$
 \\
\quad\quad\quad\quad$[\$ 6,000, \$ 9,000)$  & $0.122$~$(0.007)$ & $0.116$~$(0.003)$ & 
  \makebox[3.2em][r]{$0$}.$006$~$(0.006)$
 \\
\quad\quad\quad\quad$[\$ 9,000, \$ 18,000)$ & $0.258$~$(0.010)$ & $0.244$~$(0.004)$ & 
  \makebox[3.2em][r]{$0$}.$014$~$(0.009)$
 \\
\quad\quad\quad\quad$\geq \$ 18,000$        & $0.198$~$(0.009)$ & $0.179$~$(0.003)$ & 
  \makebox[3.2em][r]{$0$}.$019$~$(0.008)^{**}$
 \\

Personal income: & & & \\
\quad\quad\quad\quad$[\$ 3,000, \$ 6,000)$  & $0.130$~$(0.009)$ & $0.131$~$(0.003)$ & 
  \makebox[3.2em][r]{$-0$}.$001$~$(0.008)$
 \\
\quad\quad\quad\quad$[\$ 6,000, \$ 9,000)$  & $0.066$~$(0.006)$ & $0.051$~$(0.002)$ & 
  \makebox[3.2em][r]{$0$}.$015$~$(0.005)^{***}$
 \\
\quad\quad\quad\quad$\geq \$ 9,000$         & $0.038$~$(0.005)$ & $0.033$~$(0.002)$ & 
  \makebox[3.2em][r]{$0$}.$004$~$(0.005)$
 \\

At baseline: & & & \\
\quad\quad\quad\quad Has job                          & $0.242$~$(0.011)$ & $0.195$~$(0.004)$ & 
  \makebox[3.2em][r]{$0$}.$047$~$(0.010)^{***}$
 \\
\quad\quad\quad\quad Months worked, previous year     & $4.377$~$(0.119)$ & $3.566$~$(0.045)$ & 
  \makebox[3.2em][r]{$0$}.$810$~$(0.109)^{***}$
 \\
\quad\quad\quad\quad Had a job, previous year         & $0.716$~$(0.013)$ & $0.629$~$(0.005)$ & 
  \makebox[3.2em][r]{$0$}.$087$~$(0.012)^{***}$
 \\
\quad\quad\quad\quad Earnings, previous year          & $3470$~$(137.0)$  & $2870$~$(58.00)$  & \multicolumn{1}{r}{$600.0~(111.0)^{***}$}  \\
\quad\quad\quad\quad Weekly hours, most recent job    & $24.35$~$(0.560)$ & $21.42$~$(0.220)$ & 
  \makebox[3.2em][r]{$2$}.$930$~$(0.520)^{***}$
 \\
\quad\quad\quad\quad Weekly earnings, most recent job & $121.8$~$(5.800)$ & $107.6$~$(2.900)$ & \multicolumn{1}{r}{$14.10~(3.800)^{***}$} \\
\bottomrule
\end{tabular}
\par\medskip
\begin{minipage}{0.95\linewidth}
\footnotesize
\textit{Notes:} $N=9,090$. The table reports estimated means for the always-selected complier ($ac$) subpopulation and for the full sample, along with their difference. Standard errors are in parentheses. All estimates use design weights and are based on the weak-monotonicity DML specification with GRF learners. Joint $\chi^{2}$ tests for the null that all differences within a category are jointly zero:
Race/ethnicity ($p < 0.001$);
Marital status ($p = 0.084$);
Household income ($p = 0.011$);
Personal income ($p = 0.027$).
$^{*}$~$p<0.1$; $^{**}$~$p<0.05$; $^{***}$~$p<0.01$.
\end{minipage}
\end{table}





We now conduct the $ac$ profiling analysis as discussed in Section \ref{sec:profiling}.
Table \ref{tab:jc_profiling} contains the baseline characteristics of always-selected compliers, $ac$, and those of the full sample. Demographic differences are modest but systematic: On average, $ac$ individuals are less likely to be female (-4.4 pp) and Black (-6.7 pp), slightly more likely to be Hispanic (+2.2 pp) and have fewer children (-4.2 pp in incidence; -0.08 in number). Own education is marginally higher (+0.11 years), with parental education and age being similar across groups.  $ac$ individuals are slightly less concentrated in the lowest household income bracket and more represented in the highest (1.9 pp each). The most pronounced differences arise in pre-assignment labor market outcomes: $ac$ individuals are significantly more likely to be employed at baseline (+4.7 pp), more likely to have worked in the previous year (+8.7 pp), and exhibit higher labor supply and earnings, including +0.81 months worked, +\$600 annual earnings, +2.9 weekly hours, and +\$14 weekly earnings.

Taken together, these differences indicate that $ac$ subpopulation is positively selected, in particular on baseline labor-market attachment. The stronger pre-program employment and earnings profiles suggest that $ac$ may be better positioned to translate JC participation into earnings gains, but also imply more favorable counterfactual trajectories in the absence of treatment. At the same time, prior evidence on JC shows that impacts operate through channels such as increased GED, increased vocational credential attainment and reduced criminal involvement, which are not confined to the most labor-market-ready participants \citep{schochet2008does,flores2012estimating}. Thus, while the observable composition of the $ac$ group points to potential attenuation when extrapolating our estimates for $\theta_{SLATE}$ to the full sample, baseline differences alone do not fully pin down the direction or magnitude of the extrapolation.

\section{Concluding Remarks}\label{sec:conclusion}


For both identification and semiparametric inference, this paper provides a synthesis of treatment evaluation under sample selection as well as noncompliance building on Lee-type bounds and the LATE framework.
Given the analytic form of our population bounds, it is
relatively simple to incorporate additional restrictions, such as stochastic dominance of $ac$ potential outcome distribution over that of $cc$ to further tighten bounds \citep{zhang2003estimation,huber2015sharp,heiler2024treatmentevaluationintensiveextensive}. Moreover, given the simple ratio-of-linear-moment structure, our influence function components can be directly leveraged for estimation and inference on heterogeneous group-specific $ac$ effect bounds by combining them with nonparametric projection methods as in \cite{heiler2024heterogeneous}.

\newpage


\addcontentsline{toc}{section}{References}
	 {\setstretch{1}
	\bibliography{DH1}
	}


\subsection*{Declaration of generative AI and AI-assisted technologies in the manuscript preparation process}

During the preparation of this work, the authors used OpenAI's ChatGPT to assist with mathematical proofs, coding, formatting of tables and outputs, and spelling. The authors reviewed and edited the content as needed and take full responsibility for the content of the article.

\newpage