EconBase
← Back to paper

Sharp Bounds and Inference in Sample Selection Models with Treatment Endogeneity

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

95,843 characters · 19 sections · 101 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Sharp Bounds and Inference in Sample Selection Models with Treatment Endogeneity

titlepage\thispagestyle{empty} \begin{abstract} \singlespacing This paper provides partial identification and inference for treatment effects in nonparametric sample selection models with endogenous treatment and (weak) sample selection monotonicity. Outcomes are observed only for a non-randomly selected subsample and treatment is endogenous because of noncompliance with assignment. The proposed bounds for intensive margin treatment effects among compliers are sharp and tighter than those of chen2015bounds. For inference, we develop semiparametrically efficient orthogonal moments and a debiased machine learning procedure that permits valid root-$n$ inference under high-dimensional covariates and/or flexible functional forms. Simulation results indicate good finite sample performance. Applications to Job Corps and the Oregon Health Insurance Experiment show that the method can deliver substantially tighter effect bounds and confidence intervals than existing alternatives. \end{abstract} Keywords: Debiased/double machine learning; Lee bounds; Partial identification; Principal strata; Noncompliance; Sample selection \\ JEL classification: C13, C14, C21

\setcounter{page}{1}

Introduction

This paper provides sharp bounds and inference for the causal effect of a binary treatment in settings with nonrandom sample selection and endogenous treatment take-up -- two prominent features of many empirical applications. Our paper can be viewed from two complementary perspectives: From the perspective of the sample selection literature, we extend intensive margin bounds from reduced-form effects of treatment assignment to complier treatment effects in the presence of treatment take-up endogeneity. From the perspective of the instrumental variable/local average treatment effect (IV/LATE) literature, we extend the standard LATE framework to allow for sample selection, or partial observability of the outcome. Thus, this paper targets the intensive margin analogue of LATE for always-selected compliers. The framework nests the widely used lee2009training bounds as a special case. Identification is obtained under standard IV/LATE assumptions together with weak sample selection monotonicity for treatment compliers. We also develop semiparametrically efficient orthogonal moments and debiased machine learning (DML) estimators that allow valid root-$n$ inference on the treatment effect and its sharp bounds under high-dimensional covariates and flexible functional forms.

Partial identification of causal effects under sample selection or partial observability has been extensively studied. Much of the existing work, including lee2009training, derives bounds for the effect of a randomly assigned instrument without accounting for endogenous treatment take-up, and is therefore limited to intention-to-treat (ITT) effects when assignment and treatment receipt differ, see, e.g., horowitz2000nonparametric, zhang2003estimation, imai2008sharp, huber2015sharp, heiler2024heterogeneous, heiler2024treatmentevaluationintensiveextensive, sun2024partiallyidentifiedheterogeneoustreatment, or lee2025leeboundscontinuoustreatment.\footnote{horowitz2000nonparametric propose nonparametric bounds for randomized experiments with missing covariate and outcome data. zhang2003estimation develop bounds for the survivor average causal effect, targeting individuals who are always selected regardless of treatment status using principal stratification frangakis2002principal. They provide both assumption-free bounds and bounds under sample selection monotonicity and stochastic dominance assumptions. Their monotonicity based bounds are now regularly referred to as “Lee bounds” or “Zhang-Rubin-Lee bounds” andersen2023guide. imai2008sharp proves the sharpness of the zhang2003estimation bounds and extends them to quantile treatment effects. huber2015sharp derive bounds for additional subpopulations, such as individuals selected only under treatment or only under control, and for the observed subpopulation. heiler2024heterogeneous provides heterogeneous treatment effect bounds. heiler2024treatmentevaluationintensiveextensive provide identification for principal strata and extensive and intensive margin bounds. sun2024partiallyidentifiedheterogeneoustreatment consider bounds for type-specific potential outcome means, assuming exogenous treatment assignment and using an excluded variable that shifts selection but not potential outcomes. lee2025leeboundscontinuoustreatment consider intensive margin bounds for continuous treatments.}

In practice, however, realized treatment take-up is often endogenous in both experimental and quasi-experimental settings as assignment, eligibility, or encouragement may shift participation without perfectly determining treatment. For instance, in our empirical applications a nontrivial fraction of individuals do not comply with the assignment. In the National Job Corps (JC) Study, 26.2% of the individuals provided with access to training did not enroll in JC, while 4.4% of individuals assigned to the control group eventually enrolled in JC after randomization schochet2008does. In the Oregon Health Insurance Experiment (OHIE), only 30% of those eligible to apply for Medicaid successfully enrolled, and a small share of controls also enrolled finkelstein2012oregon. Noncompliance of these magnitudes can substantially affect treatment effect bounds.

Without sample selection, ITT and complier treatment effect differ only by a scaling factor, the share of compliers angrist1996identification. We show that this simple relationship breaks down under sample selection: Sharp treatment effect bounds cannot be obtained by merely scaling sharp bounds of ITT effects via primitive probabilities. Our analysis therefore fills an important gap by delivering identification and inference for sharp treatment effect bounds in the more realistic case where both endogenous sample selection and noncompliance (implying treatment endogeneity) are present.

Other papers have leveraged IV/LATE-type assumptions for partial identification of various parameters in sample selection models: lechner2010partial derive bounds for mean and quantile treatment effects for observable subpopulations, such as the “treated and selected”. christelis2019partial provide bounds on potential outcome distributions under selected samples and a monotone IV assumption manski1997monotone.

Within this literature, the closest contribution to ours is Chen and Flores (2015; CF). To our knowledge, CF is the only paper studying bounds in a setting with both sample selection and noncompliance that also provides statistical inference. Unlike the CF bounds, our bounds are sharp. Related identification contributions include imai2007identification\footnote{The derivation of bounds in imai2007identification is based on a principal-strata approach that is complementary to ours. However, the corresponding bound expressions imported from imai2008sharp appear to contain errors in the expressions for $Q_0$ and $Q_1$ in RESULT 1 on page 4. Hence, the resulting bounds cannot be reconciled with ours. Moreover, imai2007identification does not develop estimation or inference, and sharpness is not formally established for the noncompliance-and-selection parameter. Sharpness is only indirectly suggested by the arguments in imai2008sharp, which is formulated for ITT bounds under different assumptions. By contrast, we provide corrected expressions for the bounds, prove their sharpness, compare them with CF bounds, and develop a full estimation and inference framework that extends to settings with both discrete and continuous covariates and weak sample selection monotonicity.} and bartalotti2023identifying.\footnote{bartalotti2023identifying study treatment-effect bounds under sample selection within a latent-index marginal treatment effect (MTE) framework. Their Proposition 5 derives sharp bounds for always-observed LATEs with multi-valued discrete instruments. In the binary-instrument, no-covariate case, their latent-index interval $p_0<V\le p_1$ coincides with our complier stratum as defined in Section (ref). Under the same strong sample selection monotonicity, their Proposition 5 bounds are algebraically equivalent to our basic sharp bounds without covariates in Section (ref). They informally discuss plug-in estimation for MTE bounds without covariates, but do not develop accompanying inference guarantees for the discrete-instrument always-observed LATE bounds in Proposition 5.} Relative to all three papers, we contribute along multiple dimensions: First, we allow for the inclusion of covariates, thereby accommodating unconfounded rather than independently assigned instruments, which is often more realistic in quasi-experimental settings. Even when the instrument is independently assigned, incorporating covariates can improve efficiency, much as in standard average treatment effect (ATE) estimation. Second, we relax their individual-level strong sample selection monotonicity to weak, i.e., covariate-dependent, monotonicity. This is empirically relevant: semenova2025generalized shows that strong sample selection monotonicity can be rejected in JC. We also find substantial violations in the OHIE. Third, we provide a simple covariate-profiling procedure for the targeted complier population at the intensive margin. This helps to assess both the external validity of the estimates and the substantive relevance of the subpopulation to which they apply, paralleling complier profiling approaches from the LATE literature abadie2003semiparametric,angrist2004treatmenteffectheterogeneity,singh2023doublerobustnessforcomplier. Fourth, we propose a semiparametric DML procedure that delivers valid root-$n$ inference using generic nonparametric or machine learning methods for the nuisance components, allowing for potentially high-dimensional covariates and/or unknown functional forms. In strong contrast to chen2015bounds, our estimators are asymptotically linear, and thus confidence intervals have the standard “estimate $\pm$ standard error $\times$ critical value” form and do not rely on alternative concepts such as half-median-unbiasedness chernozhukov2013intersection.

Recent research has generalized estimation and inference of Lee-type bounds to high-dimensional or otherwise flexible settings using semiparametric DML methodology. semenova2025generalized extends lee2009training intensive margin bounds to multiple outcomes and provides orthogonal moments for debiased estimation. heiler2024heterogeneous develops a DML procedure to obtain corresponding heterogeneous treatment effect bounds and inference under local misspecification. heiler2024treatmentevaluationintensiveextensive provide refined DML methods for outer identification regions for intensive and extensive margin effects that admit regular inference in a larger class of DGPs. These contributions all focus on inference for (conditional) ITT effects. We instead construct semiparametrically efficient orthogonal moment functions for the sharp bounds on the treatment effect for compliers at the intensive margin. They nest several leading moments from the literature, including for ITT effect bounds (Semenova 2025, under perfect compliance), LATE (Fr\"olich 2007\nocite{frolich2007nonparametric}, in the absence of sample selection), and ATE (Hahn 1998\nocite{hahn1998role}, under perfect compliance and no sample selection).

A crucial difference to the literature on DML inference on Lee-type ITT bounds is that our sharp bounds require trimming based on a population whose quantiles are identified only indirectly via inversion of a compliance weighted cumulative distribution in the sense of abadie2003semiparametric. The latter by itself is a linear, but not necessarily convex, combination of two conditional CDFs. As a result, debiasing relies on an implicit function characterization. We leverage this representation to impose learning rate requirements only on primary, reduced-form type conditional CDFs and conditional means instead of the counterfactual conditional quantiles used for trimming the relevant observed strata distributions. This also has the advantage of making it straightforward in practice to guarantee non-crossing properties for the conditional quantiles involved, e.g., via inversion of isotonic (distributional) regression henzi2021isotonic. Our implementation also follows this inversion strategy. Monte Carlo simulations suggest good coverage and power properties of the inference method in finite samples.

We apply and compare our method to evaluate the effects of Job Corps on earnings and Medicaid on healthcare utilization. Under the strong sample selection monotonicity assumption, our sharp bounds tighten the existing chen2015bounds bounds by 68.4% (JC) and 31.1% to 69.4% (OHIE, depending on outcome). For example, the sharp bounds without covariates for the intensive margin complier effect of JC on hourly wages are now $[0.019,0.067]$ compared to $[-0.022,0.130]$ for CF. Similar reductions apply to the 95%-confidence intervals for the effects. Their widths are reduced by 46.7% (JC) and 25.8% to 36.9% (OHIE, depending on outcome). Under weak sample selection monotonicity, our sharp DML bounds still tend to be contained by the restrictive CF bounds for most outcomes and, despite relying on weaker assumptions, continue to yield shorter confidence intervals of around 7.9% (JC) and 7.3% to 42.3% (OHIE, depending on outcome).

The rest of the paper is organized as follows: Section (ref) derives the basic bounds without covariates and compares them with Lee and CF bounds. Section (ref) extends these bounds to incorporate covariates. Section (ref) develops estimation and inference. Section (ref) presents the profiling method for the targeted complier population at the intensive margin. Section (ref) provides the JC application. Section (ref) concludes. The OHIE application, Monte Carlo simulations as well as proofs and supplementary derivations are in the Supplementary Appendix.

Basic Bounds Without Covariates

Model and Identification of Basic Bounds

Let $D\in\{0,1\}$ denote a binary treatment and $Z\in\{0,1\}$ a binary instrument. Let $S\in\{0,1\}$ be a sample selection indicator for a continuous outcome $Y\in\mathcal{Y}\subset\mathbb{R}$ that is observed only if $S=1$. The observed data are independent draws of $(SY,S,D,Z)$. For $d,z\in\{0,1\}$, let $Y_{d,z}$ and $S_{d,z}$ denote, respectively, the potential outcome and selection indicator that would be observed if treatment and the instrument were set exogenously to $(d,z)$. Let $D_{z}$ denote the potential treatment if $Z$ were set to $z$. We impose the following IV assumptions:

assumption[IV]

(ref).1 (Exclusion) For $d=0,1$, $Y_{d,1}=Y_{d,0}=:Y_{d}$ and $S_{d,1}=S_{d,0}=:S_{d}$.

(ref).2 (Independence) $(Y_{1},Y_{0},S_{1},S_{0},D_{1},D_{0})\perp Z$.

(ref).3 (Strong Treatment Response Monotonicity) $P(D_{1}\geq D_{0})=1$.

(ref).4 (First Stage) $P(D_{1}\ne D_{0})>0$.

(ref).5 (Non-trivial Assignment) $P(Z=1)\in(0,1)$.

These assumptions are the standard IV/LATE assumptions imbens1994identification,angrist1996identification with the added requirement that instrument $Z$ is excluded from directly causing selection (2.1.1) and is allocated independently of potential selection (2.1.2). This matches classic encouragement or eligibility designs with selected samples where it is credible that the instrument causes selection and outcome only indirectly through the treatment.

Within this framework, types and causal effects can be decomposed into extensive and intensive margins. Point identification of the extensive margin effect is possible only for compliers ${D_{1}>D_{0}}$. In particular, the extensive-margin effect for compliers is

equation[equation omitted — 56 chars of source]

which under Assumption (ref) is point identified by the usual Wald ratio

equation[equation omitted — 93 chars of source]

Our main target parameter is the average causal effect for compliers at the intensive margin, i.e., the compliers $D_1 > D_0$ who are selected regardless of treatment response $S_1 = S_0 = 1$. We refer to this parameter as the always-selected or survivor local average treatment effect (SLATE):

equation[equation omitted — 95 chars of source]

{Remark 1.} Without noncompliance, $\theta_{SLATE}$ reduces to the intensive margin/always-selected ATE in lee2009training. Without sample selection, $\theta_{SLATE}$ reduces to the LATE imbens1994identification. Without both, it collapses to the ATE. $\theta_{SLATE}$ is the natural intensive margin analogue of the LATE: It focuses on always-selected compliers, whose treatment status is changed by the instrument while their selection status is not. It is therefore a policy effect on the level of the outcome for a principal stratum, uncontaminated by selection into the observed sample or noncompliance. The policy relevance of $\theta_{SLATE}$ has to be discussed on a case-by-case basis. We note that, in what follows, our assumptions will be able to identify the share of always-selected compliers as well as their observable characteristics. This can be used to compare their features with those of the overall population or other relevant groups, thereby assessing how substantive the parameter is in a given empirical setting. We provide all details in Section (ref) and empirical examples in Section (ref) and Appendix (ref).

Under selected samples $\theta_{SLATE}$ is only partially identified because the joint selection status $(S_{1},S_{0})$ is not observed. In particular, even for compliers who are selected only under treatment $(S_{1}=1,S_{0}=0)$ or only under control $(S_{1}=0,S_{0}=1)$, either $Y_{0}$ or $Y_{1}$ is never observed, so their treatment effects are not identified without additional (untestable) assumptions. To make progress, we first state some standard IV identities under Assumption (ref):

lemmaSuppose Assumption (ref) holds. Then, for any integrable function $g: \mathcal{Y}\rightarrow \mathbb{R}$, {\begin{align} P(D_{1}>D_{0})&=E[D| Z=1]-E[D| Z=0], \\ P(S_{1}=1,D_{1}>D_{0})&=E[DS| Z=1]-E[DS| Z=0], \\ E[g(Y_{1})| S_{1}=1,D_{1}>D_{0}]&=\frac{E[DSg(Y)| Z=1]-E[DSg(Y)| Z=0]}{E[DS| Z=1]-E[DS| Z=0]}. \end{align}} Moreover, replacing $D$ with $(D-1)$ in (ref) and (ref) identifies $P(S_{0}=1,D_{1}>D_{0})$ and $E[g(Y_{0})| S_{0}=1,D_{1}>D_{0}]$, respectively.

In particular, taking $g(Y)=\mathbbm{1}(Y\leq y)$ yields conditional CDFs $F_{Y_{1}| S_{1}=1,D_{1}>D_{0}}(y)$ and $F_{Y_{0}| S_{0}=1,D_{1}>D_{0}}(y)$, while taking $g(Y)=Y$ yields the corresponding conditional means.

Next, we introduce the complier strata that underlie the SLATE parameter. Among compliers ($D_{1}>D_{0}$), define

alignat{2} ac&:=\{S_{0}=S_{1}=1,\;D_{1}>D_{0}\} &\quadalways-selected compliers \\ dc&:=\{S_{1}<S_{0},\;D_{1}>D_{0}\} &\quadtreatment-de-selected compliers \\ cc&:=\{S_{1}>S_{0},\;D_{1}>D_{0}\} &\quadtreatment-selected compliers

Rewriting (ref) in terms of these strata yields

equation[equation omitted — 57 chars of source]

The set $\{S_{0}=1,D_{1}>D_{0}\}$ combines $ac$ and $dc$, while $\{S_{1}=1,D_{1}>D_{0}\}$ combines $ac$ and $cc$. Lemma (ref) allows us to identify $P(S_{0}=1,D_{1}>D_{0})$ and $P(S_{1}=1,D_{1}>D_{0})$, which yield two linear restrictions on the shares of three unknown types ($ac$, $dc$, $cc$). To fully identify these shares, we further impose a strong sample selection monotonicity condition for the compliers:

assumption[Strong Sample Selection Monotonicity for Compliers] Either $P(S_{1}\geq S_{0}| D_{1}>D_{0})=1$ or $P(S_{1}\leq S_{0}| D_{1}>D_{0})=1$.

Assumption (ref) imposes a common direction of treatment-induced sample selection within the complier stratum $D_{1} > D_{0}$. In particular, the treatment is assumed to either weakly increase or weakly decrease selection for all compliers. It cannot move some compliers into the sample and others out of it, effectively ruling out either $dc$ or $cc$ types.\footnote{This restriction is implied, for example, by an additively separable selection equation, see, e.g., vytlacil2002independence for treatment selection or heiler2024treatmentevaluationintensiveextensive for sample selection. However, these models impose monotonicity for the full population and are thus stronger than Assumption (ref).} Note that no analogous restriction is needed for strata with $D_{1} = D_{0}$ (always-takers and never-takers) as their treatment status does not vary with the instrument. In particular, never-takers satisfy $D_0=D_1=0$, so their observed selection status is governed only by $S_0$; always-takers satisfy $D_0=D_1=1$, so their observed selection status is governed only by $S_1$. The counterfactual selection status under the unrealized treatment therefore does not enter the identification argument.

We now provide identification for the case where $P(S_{1}\geq S_{0}| D_{1}>D_{0})=1$. The case where $P(S_{1}\leq S_{0}| D_{1}>D_{0})=1$ can be handled analogously. Under Assumption (ref), the stratum ${S_{0}=1,D_{1}>D_{0}}$ consists only of always-selected compliers ($ac$), while ${S_{1}=1,D_{1}>D_{0}}$ combines $ac$ and $cc$. Lemma (ref) then implies that

equation[equation omitted — 69 chars of source]

is point identified by taking $g(Y)=Y$ and replacing $D$ with $(D-1)$ in (ref). In contrast, $E[Y_{1}| ac]$ is not point identified. However, Lemma (ref) and Assumption (ref) allow us to identify the mixture distribution of $Y_{1}$ for $ac$ and $cc$,

equation[equation omitted — 75 chars of source]

as well as the frequencies of $ac$ and $cc$ within this mixture. Let

equation[equation omitted — 98 chars of source]

Then, under Assumption (ref),

equation[equation omitted — 96 chars of source]

and both probabilities are identified. Moreover, define

equation[equation omitted — 54 chars of source]

$p$ is equal to the fraction of always-selected compliers in mixture ${S_{1}=1,D_{1}>D_{0}}$. Since we only observe this mixture, sharp bounds are obtained by assigning $ac$ to the “best” or “worst” location within the mixture distribution of $Y_{1}$. Let $F_{1}(y):=F_{Y_{1}| S_{1}=1,D_{1}>D_{0}}(y)$ and $Q_{1}(u):=\inf\{y\in\mathcal{Y}:F_{1}(y)\geq u\}$ denote the identified CDF and quantile function of this mixture.

To simplify exposition for the remainder of this section, we maintain the following regularity condition: the relevant identified distributions entering the trimming formulas have continuous, strictly increasing CDFs. This rules out mass points at the relevant trimming cutoffs and allows the bounds to be written as ordinary trimmed means. This condition is not needed for identification per se.\footnote{A weaker local regularity condition would suffice: for each result below, it is enough that the corresponding identified CDF be continuous and strictly increasing in a neighborhood of the relevant trimming threshold(s). Without this regularity, the same conclusions also hold once the bounds are written in exact fractional-trimming (quantile-integral) form.}

The lower and upper bounds for $E[Y_{1}| ac]$ are

align[align omitted — 266 chars of source]

i.e., the trimmed means obtained by placing all $ac$ individuals in the bottom $p$ fraction (for the lower bound) or the top $p$ fraction (for the upper bound) of the mixture. We obtain the following proposition:

proposition[Sharpness] Suppose Assumptions (ref)-- (ref) hold and $\pi_{ac}>0$. Then, \begin{equation*} \beta_{L,1}\leq E[Y_{1}| ac]\leq \beta_{U,1}. \end{equation*} Moreover, $\beta_{L,1}$ and $\beta_{U,1}$ are sharp: They are, respectively, the largest lower bound and the smallest upper bound for $E[Y_{1}| ac]$ that are consistent with the observed data and Assumptions (ref)--(ref). Any other valid bounds under these assumptions contain $[\beta_{L,1},\beta_{U,1}]$.

Combining the point-identified $\beta_{0}$ with these bounds yields sharp bounds for $\theta_{SLATE}$:

equation[equation omitted — 113 chars of source]

Using Assumption (ref) and Lemma (ref) with $g(Y)=Y\mathbbm{1}(Y\leq Q_{1}(p))$ and $g(Y)=Y\mathbbm{1}(Y\geq Q_{1}(1-p))$, all components of $\beta_{L}$ and $\beta_{U}$, namely $\pi_{ac}$, $p$, $F_{1}$, $Q_{1}$, and the trimmed means are functionals of the observed distribution of $(SY,S,D,Z)$. In particular,

{

align[align omitted — 446 chars of source]

} where

align[align omitted — 427 chars of source]

We next compare our bounds with two benchmarks: the ITT bounds of lee2009training and the CF bounds, which target $\theta_{SLATE}$ under essentially identical identification assumptions.

Comparison with Existing Bounds

Comparison with lee2009training Bounds

Without sample selection, the ITT effect

equation[equation omitted — 45 chars of source]

and the average treatment effect for compliers,

equation[equation omitted — 113 chars of source]

differ only by the scaling factor $\tau_{D}$, which corresponds to the share of compliers. That is, in the standard IV point identification setting, ITT can be converted into the complier average treatment effect by dividing by $\tau_{D}$. This does not generalize to treatment effect bounds under sample selection.

In particular, in the presence of sample selection, lee2009training derives sharp bounds for the ITT effect among the always-selected assuming strong sample selection monotonicity. These “Lee bounds” can be viewed as a special case of our basic bounds under perfect compliance, i.e., when treatment equals instrument, $D=Z$. We only consider the lower bound. The analysis for the upper bound is analogous and omitted for brevity. Setting $D=Z$ in Equation (ref), we obtain the bound

{

align[align omitted — 159 chars of source]

} where $Q_{1}^{Lee}(u):=\inf \{y\in \mathcal{Y}: F_{Y|S=1,Z=1}(y)\ge u \}$ is the $u$-quantile of $\left(Y|S=1,Z=1\right)$, and $p^{Lee}:=E[S| Z=0]/E[S| Z=1]$ is the trimming fraction.

To relate (ref) to our SLATE bounds, it is useful to decompose the selected population into latent types. In addition to the complier types introduced in Section (ref), let

align[align omitted — 162 chars of source]

and denote their shares by $\pi_{an}$ and $\pi_{aa}$ respectively. Under Assumptions (ref)--(ref), the selected group in each arm of the experiment is a mixture of these types. We show in Appendix (ref) that

align[align omitted — 221 chars of source]

for suitable mixture weights $\omega,\delta\in(0,1)$ depending only on type shares. Let $q :=Q^{Lee}_1(p^{Lee})$ denote the Lee trimming threshold. The trimmed mean in the first term of (ref) can be written as

align[align omitted — 217 chars of source]

where $w\left( q\right)=\tfrac{\omega F_{Y_0\mid an}(q)}{p^{Lee}}$ is the effective $an$ stratum weight from trimming applied after mixing. The second term of (ref) can be written as {

align[align omitted — 160 chars of source]

} Thus, $\beta_{L}^{Lee}$ is a difference of two mixtures that involve not only treatment compliers ($ac$, $cc$), but also always-takers and never-takers ($aa$, $an$). In contrast, our lower bound for SLATE,

equation[equation omitted — 108 chars of source]

depends only on the treatment complier types ($ac$ and $cc$), with $p$ equal to the fraction of always-selected compliers within that complier mixture.

Except in knife-edge cases where the shares of always-selected always-takers and never-takers vanish ($\pi_{aa}=\pi_{an}=0$), the $Y_{1}$ term in $\beta_{L}^{Lee}$ does not reduce to $E[Y_{1}| Y_{1}\leq Q_{1}(p),ac \cup cc]$, and the $Y_{0}$ term does not reduce to $E[Y_{0}| ac]$. Consequently, lee2009training bounds cannot be transformed into sharp bounds for $\theta_{SLATE}$ by simple rescaling via primitive probabilities. We summarize this result formally in the following proposition.

proposition[Lee bounds do not rescale to SLATE bounds] Suppose Assumptions (ref)--(ref) hold and $\pi_{ac}>0$. If $\pi_{an}+\pi_{aa}>0$, then, for generic joint distributions of $(Y_0,Y_1)$ across latent types, \begin{align*} \beta_L^{Lee}\neq \beta_L. \end{align*} Moreover, there does not exist a scalar function $c$, depending only on the primitive probabilities (i.e. the principal-strata shares and instrument probabilities), such that \begin{align*} \beta_L = c\,\beta_L^{Lee} \end{align*} uniformly over DGPs.

Proposition (ref) implies that lee2009training ITT bounds cannot be converted into sharp bounds for SLATE by a simple rescaling analogous to $\tau_{Y}/\tau_{D}$ in the standard IV case without sample selection.

Comparison with chen2015bounds Bounds

CF study partial identification of $\theta_{SLATE}$ under essentially identical model assumptions. We show that our bounds are weakly tighter uniformly over DGPs, and strictly tighter on a nonempty subset of DGPs than the CF bounds. For brevity, we focus on the lower bounds. The upper bounds can be treated analogously.

CF propose two lower bounds for $E[Y_{1}| ac]$, based on trimming the distribution of $Y$ among selected treated units, $Y| DSZ=1$, equivalent to $Y_1|ac\cup cc\cup aa$ under Assumptions (ref)--(ref), and then taking the maximum of the two. In contrast, our sharp bounds trim within the smaller identified mixture $Y_1|ac\cup cc$. Let $Q_{1}^{CF}(u)$ denote the $u$-quantile of $Y| DSZ=1$. Their basic lower bound is

{

align[align omitted — 159 chars of source]

}

where $p^{CF}_1:=\pi_{ac}/(\pi_{ac}+\pi_{cc}+\pi_{aa})$ and $\beta_{0}:=E[Y_{0}| ac]$. Their alternative lower bound uses additional information from the always-taker stratum ($aa$),

{

align[align omitted — 311 chars of source]

} where $p^{CF}_2:=(\pi_{ac}+\pi_{aa})/(\pi_{ac}+\pi_{cc}+\pi_{aa})$ and $E[Y| DS(1-Z)=1]=E[Y_{1}| aa]$ is point identified. The CF lower bound is then

equation[equation omitted — 72 chars of source]

Under Assumptions (ref)--(ref), the conditional distribution $Y| DSZ=1$ coincides with $Y_{1}$ for the mixture of types $ac$, $cc$, and $aa$:

equation[equation omitted — 63 chars of source]

so both $\beta_{L}^{mix}$ and $\beta_{L}^{adj}$ are obtained by trimming this mixture distribution and then reweighting to isolate $ac$ and $aa$. CF show that $\beta_{L}^{mix}$ corresponds to the worst-case lower bound when all $ac$ individuals are placed in the bottom $p^{CF}_1$- fraction of $Y_{1}| ac\cup cc\cup aa$, while $\beta_{L}^{adj}$ corresponds to the worst case when $ac$ and $aa$ together occupy the bottom $p^{CF}_2$- fraction of that distribution, using the fact that $E[Y_{1}| aa]$ is identified.

In contrast, our lower bound $\beta_{L}$ uses the smallest stratum that is point identified, namely $Y_{1}| ac\cup cc$. As shown in Lemma (ref) in Appendix (ref), the distribution $F_{Y_{1}| ac\cup cc}$ can be recovered from the two observable selected distributions

equation[equation omitted — 107 chars of source]

We then construct the sharp lower bound

equation[equation omitted — 132 chars of source]

by trimming $Y_{1}$ only within this complier mixture.

Intuitively, $\beta_{L}$ assumes, in a worst-case fashion, that all $ac$ individuals are located below all $cc$ individuals in the distribution of $Y_{1}$ among compliers, but it imposes no restrictions on the relative ranking of $ac$ versus $aa$. By contrast, the CF bounds are based on worst-case restrictions on the joint ranking of $ac$, $cc$, and $aa$ within $Y_{1}| ac\cup cc\cup aa$. This makes them conservative. In particular, we obtain the following proposition:

proposition[Dominance over CF bounds] Suppose Assumptions (ref)--(ref) hold and $\pi_{ac}>0$. Then, \begin{align*} \beta_{L,1}\geq \beta_{L,1}^{CF}\geq \beta_{L,1}^{mix} \quad and \quad \beta_{U,1}\leq \beta_{U,1}^{CF}\leq \beta_{U,1}^{mix}, \end{align*} where $\beta_{L,1}$ is defined in (ref), $\beta_{L,1}^{mix}$ in (ref), and $\beta_{L,1}^{CF}:=\max\{\beta_{L,1}^{mix},\beta_{L,1}^{adj}\}$, with $\beta_{L,1}^{adj}$ defined in (ref). The upper-bound counterparts $\beta_{U,1}$, $\beta_{U,1}^{mix}$, $\beta_{U,1}^{adj}$, and $\beta_{U,1}^{CF}$ are defined analogously. Hence, imposing either of the two restrictions used by CF cannot tighten our sharp lower or upper bound for $E[Y_1\mid ac]$. Moreover, these inequalities are strict on a nonempty class of DGPs: \begin{align*} \beta_{L,1}>\beta_{L,1}^{CF} &\iff \pi_{aa}>0 \quadand\quad 0<F_{Y_1\mid aa}\!\bigl(Q_1(p)\bigr)<1, \\ \beta_{U,1}<\beta_{U,1}^{CF} &\iff \pi_{aa}>0 \quadand\quad0<F_{Y_1\mid aa}\!\bigl(Q_1(1-p)\bigr)<1. \end{align*}

Bounds with Covariates

We now extend the unconditional baseline bounds in Section (ref) to incorporate predetermined covariates. Incorporating such covariates serves three purposes: First, it allows for an unconfounded instead of a fully randomly assigned instrument. Second, even under fully randomized assignment, it can improve efficiency by conditioning on relevant covariates when constructing trimmed means. Third, it allows us to relax strong sample selection monotonicity. In particular, we permit the direction of sample selection for treatment compliers to vary with covariates. For simplicity, we maintain the strong treatment response monotonicity assumption (ref).3. This no-defiers restriction is natural in eligibility and encouragement designs such as JC and OHIE: departures from perfect compliance are largely one-sided, arising mainly from incomplete take-up among those assigned to treatment rather than from substantial treatment receipt among controls. Relaxing this assumption is possible via further covariate partitioning analogous to the sample selection case in what follows.

Let $X\in\mathcal{X}\subset\mathbb{R}^{d_{x}}$ denote a vector of predetermined covariates. The observed data are $(SY,S,D,Z,X)$. We impose the following conditional IV assumptions:

assumption[Conditional IV] For $d=0,1$: (ref).1 (Exclusion) $Y_{d,1}=Y_{d,0}=:Y_{d}$, $S_{d,1}=S_{d,0}=:S_{d}$. (ref).2 (Independence) $(Y_{1},Y_{0},S_{1},S_{0},D_{1},D_{0})\perp Z| X$. (ref).3 (Strong Treatment Response Monotonicity) $P(D_{1}\geq D_{0})=1$. (ref).4 (First Stage) $P(D_1 \ne D_0 \mid X=x)>0$ for $P_X$-almost every $x \in \mathcal{X}$ (ref).5 (Non-trivial Assignment) $P(Z=1\mid X=x)\in(0,1)$ for $P_X$-almost every $x\in\mathcal{X}$.

These assumptions are standard extensions of the conditional IV/LATE assumptions frolich2007nonparametric generalizing Assumption (ref) with instrument $Z$ again being excluded from directly causing selection and being independently allocated of potential selection given covariates.

It is important to note that Assumptions (ref).4 and (ref).5 are not intended to impose substantive restrictions beyond their unconditional counterparts Assumptions (ref).4 and (ref).5. Rather, they should be interpreted as support restrictions for the conditional analysis: if the conditional first-stage or non-trivial assignment condition fails on a set of positive $P_X$-probability, $\mathcal{X}$ should be understood as restricted to the support on which these conditions hold.\footnote{Failure of the first stage and failure of non-trivial assignment have slightly different interpretations. If $P(D_1\ne D_0\mid X=x)=0$, there are no compliers at that $x$. so complier-based conditional objects are not meaningful there. If $P(Z=1\mid X=x)\in\{0,1\}$, there is no within-$x$ variation in the instrument, so the conditional IV comparisons are not identified there.}

Under Assumption (ref), the identities in Lemma (ref) hold pointwise in $x$. Next, we relax strong sample selection monotonicity to a weaker conditional version that allows the direction of monotonicity to vary with covariates.

assumption[Weak Sample Selection Monotonicity for Compliers] Either \\ $P( S_1 \geq S_0 | D_{1}>D_{0},X=x)=1$ or $P( S_1 \leq S_0 | D_{1}>D_{0}, X=x)=1.$

Assumption (ref) yields subsets of the covariate space where treatment compliers can have either a weakly positive or negative selection responses to treatment. At their intersection, treatment does not affect selection. Without loss of generality, we now define a distinct partitioning where this intersection is combined with the weakly positive responders. Formally, let

align*[align* omitted — 98 chars of source]

and partitions

align[align omitted — 298 chars of source]

We note that the presence of $\mathcal{X}^0$ is harmless for identification. For regular semiparametric inference, however, it will be necessary that $P(\mathcal{X}^0) = 0$ heiler2024treatmentevaluationintensiveextensive. We return to this in Section (ref).

{Remark 2.} Assumption (ref) weakens strong sample selection monotonicity by allowing the sign of the treatment effect on selection to vary with observed covariates. Thus, treatment may increase selection for some values in $\mathcal{X}$ and decrease it for others. The restriction is nevertheless substantive: conditional on $X=x$ and complier status, residual heterogeneity in the selection response must be one-sided. In particular, the assumption rules out the coexistence of treatment-induced entry and exit among compliers with the same covariate values. Its credibility therefore depends on whether $\mathcal{X}$ is rich enough to absorb the economically relevant heterogeneity in the direction of selection. It is more plausible when the main determinants of the sign of the selection response are observed, such as baseline characteristics governing participation, employment, survival, etc. It is less plausible when unobserved factors can generate opposing responses within the same covariate cell, see heiler2024treatmentevaluationintensiveextensive and semenova2025generalized for additional discussion.

Let $\pi_{ac}(x):=P(ac| X=x)$ denote the conditional share of always-selected compliers, and let $\pi_{ac} := E[\pi_{ac}(X)]$ denote their unconditional share. Further let $\mathcal X_{ac}:=\{x\in\mathcal X:\pi_{ac}(x)>0\}$.\footnote{For identification of the unconditional bounds, the substantive requirement is $\pi_{ac}>0$. $\pi_{ac}(x)$ may be zero on subsets of the covariate support as they receive zero weight in numerator and denominator of the unconditional bounds. As a convention for display, we treat $\theta_{SLATE}(x)=0$ for $x\in\mathcal X\setminus\mathcal X_{ac}$.} For any $x \in \mathcal X_{ac}$, define the conditional SLATE parameter

equation[equation omitted — 58 chars of source]

Let

equation[equation omitted — 72 chars of source]

and write $\pi_{cc}(x):=P(cc| X=x)$ and $\pi_{dc}(x):=P(dc| X=x)$ for the conditional shares of $cc$ and $dc$, respectively. Under Assumption (ref),

align[align omitted — 237 chars of source]

This yields the ratio

equation[equation omitted — 60 chars of source]

Thus $0 \le p(x) \leq 1$ on $\mathcal{X}^+_{ac}:=\mathcal{X}^+\cap\mathcal{X}_{ac}$ and $p(x) > 1$ on $\mathcal{X}^-_{ac}:=\mathcal{X}^-\cap\mathcal{X}_{ac}$. For $d=0,1$ and $x\in \mathcal{X}_{ac}$, define the conditional quantile

equation[equation omitted — 107 chars of source]

To simplify exposition, for the remainder of this section we again maintain that the relevant identified conditional CDF entering each trimming formula is continuous and strictly increasing in a neighborhood of the corresponding trimming threshold(s). This rules out mass points at the trimming thresholds and allows the conditional sharp bounds below to be written as ordinary trimmed means.\footnote{For $x\in \mathcal{X}^+_{ac}$, the relevant CDF is $F_{Y_{1}| S_{1}=1,D_{1}>D_{0},X=x}(\cdot)$ and the relevant thresholds are $Q_1(p(x),x)$ and $Q_1(1-p(x),x)$. For $x\in \mathcal{X}^-_{ac}$, the relevant CDF is $F_{Y_{0}| S_{0}=1,D_{1}>D_{0},X=x}(\cdot)$ and the relevant thresholds are $Q_0(1-1/p(x),x)$ and $Q_0(1/p(x),x)$. Without this regularity, the same bounds can be written in exact fractional-trimming form.} Then Proposition (ref) can be applied pointwise in $x$. This yields the following sharp lower and upper bounds for $\theta_{SLATE}(x)$: {

align[align omitted — 375 chars of source]

} for $x\in\mathcal{X}^{+}_{ac}$, and

{

align[align omitted — 380 chars of source]

} for $x\in\mathcal{X}^{-}_{ac}$. As a notational convention, set $\beta_B^\pm(x)=0$ for $x\in \mathcal{X}\setminus \mathcal X_{ac}$, $B\in\{L,U\}$. These values are used only in the reduced-forms $\beta_B^\pm(x)\pi_{ac}(x)$ entering the unconditional bounds. We note that when treatment does not cause selection, conditional effects bounds collapse to a point, i.e. for any $x \in (\mathcal{X}^0 \cap \mathcal{X}_{ac}) \subseteq \mathcal{X}^+_{ac}$ we have that

align[align omitted — 171 chars of source]

Now let

equation[equation omitted — 156 chars of source]

and define for any $B \in \{L,U\}$

equation[equation omitted — 100 chars of source]

The unconditional SLATE parameter can be written as

equation[equation omitted — 80 chars of source]

The corresponding sharp lower and upper bounds are

equation[equation omitted — 131 chars of source]

To make the link with observed data explicit, we next provide the estimands of the denominator and numerator moment functions in $\beta_{L}$ and $\beta_{U}$. Define the instrument propensity score $e(x):= P(Z=1| X=x)$, and the inverse probability weight

{

equation[equation omitted — 73 chars of source]

} Extending Lemma (ref) to hold conditional on $X$, one can show that

{

align[align omitted — 195 chars of source]

} where now

{

align[align omitted — 249 chars of source]

} Thus $\pi_{ac}$ is identified. In addition, for any $y$ and $x \in \mathcal{X}_{ac}$,

{

align[align omitted — 346 chars of source]

} The conditional quantiles $Q_{d}(u,x)$ used for trimming in the conditional bounds in Equations (ref) -- (ref) are defined as the generalized inverses of these CDFs and are thus identified. Using these identities, the following proposition collects the explicit estimands for the unconditional bounds $\beta_{L}$ and $\beta_{U}$.

propositionSuppose Assumptions (ref)--(ref) hold and $\pi_{ac}>0$. Then the sharp lower and upper bounds for $\theta_{SLATE}$ are identified as {\begin{equation*} \beta_{L}=\frac{E[\beta_{L}(X)\pi_{ac}(X)]}{\pi_{ac}},\qquad \beta_{U}=\frac{E[\beta_{U}(X)\pi_{ac}(X)]}{\pi_{ac}}, \end{equation*}} where $\pi_{ac}$ is given by Equation (ref) and \begin{align*} \beta_{L}(x) &:=\mathbbm{1}^{+}(x)\beta_{L}^{+}(x)+\mathbbm{1}^{-}(x)\beta_{L}^{-}(x), \\ \beta_{U}(x) &:=\mathbbm{1}^{+}(x)\beta_{U}^{+}(x)+\mathbbm{1}^{-}(x)\beta_{U}^{-}(x), \end{align*} with { \begin{align} \beta_{L}^{+}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY\mathbbm{1}(Y\leq Q_{1}(p(x),x))+(1-D)SY\right)| X=x\right], \\ \beta_{L}^{-}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY+(1-D)SY\mathbbm{1}(Y\geq Q_{0}(1-1/p(x),x))\right)| X=x\right], \\ \beta_{U}^{+}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY\mathbbm{1}(Y\geq Q_{1}(1-p(x),x))+(1-D)SY\right)| X=x\right], \\ \beta_{U}^{-}(x)\pi_{ac}(x)&:=E\left[W(Z,X)\left(DSY+(1-D)SY\mathbbm{1}(Y\leq Q_{0}(1/p(x),x))\right)| X=x\right], \end{align}} on $\mathcal{X}_{ac}$, and $\beta_B^\pm(x)\pi_{ac}(x)=0$ by convention on $\mathcal X\setminus\mathcal X_{ac}$. $\lambda_d(x)$ and $W(z,x)$ are point identified on $\mathcal{X}$ and conditional quantiles $Q_d(\cdot,x)$ are point identified on $\mathcal X_{ac}$.

In summary, $(\beta_{L},\beta_{U})$ are functionals of the observed law of $(SY,S,D,Z,X)$ and a collection of nuisance functions (instrument propensity score, conditional means, and (inverted) conditional distribution functions). In the next section, we derive orthogonal moment conditions and semiparametric influence functions for these functionals, which will allow us to construct semiparametrically efficient debiased machine learning estimators for $(\beta_{L},\beta_{U})$ and confidence intervals for $\theta_{SLATE}$.

Estimation and Inference

Nuisance Functions and Target Parameter

The previous section shows that for any $B\in\{L,U\}$, the unconditional bound can be written as

equation[equation omitted — 98 chars of source]

with $\beta_{B}(X)$ defined in Proposition (ref) and $\pi_{ac}$ in (ref). For notational convenience, we decompose the numerator of (ref) as

equation[equation omitted — 109 chars of source]

where, for $B\in\{L,U\}$ and $\ast\in\{+,-\}$, we define

equation[equation omitted — 124 chars of source]

Table (ref) contains all primary nuisance functions as well as all derived quantities that are used for the remainder of this section. We collect the required primary nuisance functions in vector

align[align omitted — 131 chars of source]
table[table omitted — 1,568 chars of source]

Semiparametric Influence Functions and Nuisance Functions

For any $B \in \{L,U\}$, we treat $\pi_{ac}$ in (ref) as well as the four components in (ref) as target functionals and derive their semiparametric influence functions/efficient orthogonal moments. They are then combined to yield the efficient influence function and estimator for the respective $\beta_B$. Denote the influence function operator $\mathbbm{IF}: \Theta \rightarrow L^2(P,\mathbb{R}^q)$ as the operator that, for a parameter $\theta:\mathcal{P}\rightarrow \mathbb{R}^q$, returns its influence function, i.e., the standard Riesz-representer of the pathwise derivative bickel1993efficient, kennedy2024semiparametric.\footnote{The influence function is defined via the functional derivative of the target parameter with respect to perturbations along the tangent space evaluated at the true probability distribution. For example, if we observe iid data $X\sim \mathcal{P}$ and the target functional is $E[X]$, then for $\mathbbm{IF}(E[X]) = X - E[X]$. For a regular parameter under the nonparametric model $\mathcal{P}$, this is the unique semiparametric efficient influence function, see, e.g., hines2022demystifying, kennedy2024semiparametric or heiler2026heterogeneity for additional examples.}

We now present the influence function for the target parameters $\beta_B$ for $B \in \{L,U\}$. We suppress dependence on nuisances and data in what follows whenever it does not cause confusion. By linearity of the $\mathbbm{IF}$ operator, we have that

align[align omitted — 236 chars of source]
table[table omitted — 2,631 chars of source]

The influence functions of the components are in Table (ref). Putting upper and lower bounds together then yields

align[align omitted — 135 chars of source]

The influence functions $\mathbbm{IF}(\beta_B)$ for $B \in \{L,U\}$ nest the existing literature on LATE and sample selection bounds. In particular, we obtain the following proposition:

propositionFor any $B \in \{L,U\}$, the efficient influence function in (ref) under (i) perfect compliance, (ii) no sample selection or (iii) both collapse to their respective efficient influence functions for Lee bounds, LATE, or ATE: \begin{enumerate}[label=(\roman*), leftmargin=*] • If $P(D = Z) = 1$, then $\mathbbm{IF}(\beta_B) = \mathbbm{IF}(\beta_B^{Lee})$ a.s. • If $P(S = 1) = 1$, then $\mathbbm{IF}(\beta_B) = \mathbbm{IF}(\theta_{LATE})$ a.s. • If $P(S=1, D=Z) = 1$, then $\mathbbm{IF}(\beta_B) = \mathbbm{IF}(\theta_{ATE})$ a.s. \end{enumerate}

Practically, this means that, as $P(D = Z) \rightarrow 1$ or $P(S = 1) \rightarrow 1$, our influence functions will get closer in probability to the semiparametrically efficient influence functions for Lee-bounds on the intensive margin ATE heiler2024treatmentevaluationintensiveextensive,semenova2025generalized or the LATE frolich2007nonparametric respectively. If both apply, the functions approach the efficient influence of the ATE hahn1998role.

Finite Sample Implementation

We now outline the steps to estimate bounds and effect confidence intervals and provide more details regarding nuisance function estimation. For a generic random variable $X$ we denote $E_n[X] = \frac{1}{n}\sum_i^n X_i$ in what follows. Assume we have independent data $O_i = (S_iY_i, S_i, D_i, Z_i, X_i)'$ for $i=1,\dots,n$. Algorithm (ref) contains a step-by-step explanation of how to obtain bounds and inference. Estimation of bounds is fairly standard within the DML framework for composite ratio parameters. In particular, we estimate all their components separately with cross-fitted nuisances to obtain an estimate for the eventual bounds and their respective influence functions. The latter then yield the full variance-covariance matrix estimates that can be used for standard-normal-based inference on bounds or, more importantly, confidence intervals for the effect using refined critical values imbens2004confidence,stoye2020simple. All of these have at least $(1-\alpha)$ asymptotic coverage under some regularity conditions on the conditional distribution and suitable rate conditions on the nuisances, see Section (ref) for more details.

algorithm[algorithm omitted — 1,926 chars of source]

We now discuss primary nuisance function estimation and how to obtain derived nuisances as defined in Table (ref). Our high-level assumptions in Section (ref) match these primitive objects that are conditional means/probabilities, conditional CDFs, and trimmed conditional means. For primary nuisance functions that are simple conditional means or probabilities, $e,r,m,\mu, \nu$, a plethora of off-the-shelf nonparametric and machine learning methods such as neural networks, forests or high-dimensional sparse parametric models are available. The derived nuisances $\lambda_0,\lambda_1, \pi_{ac}, p, \mathbbm{1}^+, \mathbbm{1}^-, W$ only depend on the primary and can be obtained via simple plug-in versions.

The functional parameters $F$ and $G_B$ require special attention. In particular, we suggest to estimate the primary conditional CDF $F$ via machine learning analogues of distributional regression foresi1995conditional,klein2024distributional. We also recommend direct imposition or post-processing, e.g., via isotonic regression henzi2021isotonic that further refine these estimates by enforcing nondecreasing estimated conditional CDFs. Together with $r$ and $m$, these yield plug-in versions of $F_0$ and $F_1$. The $G_B$ components can be similarly obtained and refined as a sequence of regressions with outcome variable equal to trimming indicator times actual outcome.

To obtain the derived conditional quantiles $Q_0$ and $Q_1$, inversion of the previously obtained conditional CDFs, $F_0$ and $F_1$, can be used. These yield the trimming indicators that are required to evaluate the conditional expectation models $G_B$ at their respective trimming quantiles to obtain derived nuisances $\psi_B$. The inversion-based approach ensures algebraic compatibility between the different nuisance quantities. Thus, we use this approach in all of our simulations and applications in Section (ref) as well as Appendix (ref) and (ref).

Large Sample Properties

We now present additional technical assumptions and the resulting large sample properties of the efficient influence function based estimators of the causal effect bounds. Denote the true nuisances as $\eta(x) =: \eta \in \mathcal{H}$ where $\mathcal{H}$ is a convex subset of a suitably normed vector space. Denote $\mathcal{H}_n \subset \mathcal{H}$ the realization set of the estimated nuisance quantities $\hat{\eta}(x) =: \hat{\eta}$, i.e., the set containing estimated nuisances with probability $1-u_n$ where $u_n = o(1)$. All nuisances are cross-fitted according to Definition 3.2 in chernozhukov2018double, see Algorithm (ref).

For the remainder, write for generic nuisance $h=h(x)$ its estimation error $\Delta h=\hat h-h$. Denote $\|\cdot\|_2$ as $L^2(P)$ norm and $\|\cdot\|_\infty$ as the uniform norm. By abuse of notation, if the object depends on $Z=z$ and/or $D=d$, we suppress dependence and take all norms to be uniform over these finite dimension as well, e.g.,

align[align omitted — 161 chars of source]

and equivalently for other nuisances. For a given $x$, we also denote $\|\cdot\|_{\infty,\mathcal{N}_x}$ as the uniform norm over a neighborhood $\mathcal{N}_x$. In particular, for a generic object $A(\cdot|x)$,

align[align omitted — 115 chars of source]

We denote the supremum over these neighborhoods as

align[align omitted — 119 chars of source]

Moreover, we write shorthand

align[align omitted — 96 chars of source]

and $a_n \lesssim b_n$ and $a_n \lesssim_P b_n$, whenever $a_n = O(b_n)$ or $a_n = O_p(b_n)$ respectively. If not stated differently, the following assumptions are all uniformly over $n$.

{Assumption A (Regularity, Overlap, and Learning Rates)}

enumerate[itemsep=0pt] \singlespacing • (Moments) The conditional potential outcome moments are bounded, i.e., for some $m > 0$, \begin{align*}\sup_{x\in\mathcal{X},d\in\{0,1\}}E[|Y(d)|^{2+m}|X=x] \lesssim 1.\end{align*} • (Eigenvalues) The variance-covariance matrix of the influence function $E[\mathbbm{IF}(\beta)\mathbbm{IF}(\beta)']$ has finite eigenvalues bounded away from zero. • (No Point Mass at Trimming Points) The conditional distributions are continuous at the trimming points. For $z \in \{0,1\}$, and $x$ on the relevant support,\footnote{Here and in A.4, the relevant support is $\mathcal X^+_{ac}$ for conditions involving $(Q_1(p(x),x)$ or $Q_1(1-p(x),x)$, and $\mathcal X^-_{ac}$ for conditions involving $Q_0(1-1/p(x),x)$ or $Q_0(1/p(x),x)$.}\begin{align*} &P(Y=Q_1(p(x),x)| DS=1,Z=z,X=x) = 0,\\ &P(Y=Q_1(1-p(x),x)| DS=1,Z=z,X=x) = 0,\\ &P(Y=Q_0(1-1/p(x),x)| (1-D)S=1,Z=z,X=x) = 0, \\ &P(Y=Q_0(1/p(x),x)| (1-D)S=1,Z=z,X=x) = 0. \end{align*} • (Bounded Mixture Outcome Density and Local Lipschitz Trimmed Means) Let $f_1(\cdot|x)$ and $f_0(\cdot|x)$ denote the respective densities of $F_1(\cdot|x)$ and $F_0(\cdot|x)$. Assume they are bounded at the trimming thresholds and the trimmed conditional means are locally Lipschitz, i.e., there exist a $C>0$ and constants $0<f_{\min}\le f_{\max}<\infty$, $L_G<\infty$ such that (i) for all $x$ on the relevant support: \begin{align*} &f_1(Q_1(p(x),x)\mid x) \in [f_{\min},f_{\max}], \\ &f_1(Q_1(1-p(x),x)\mid x) \in [f_{\min},f_{\max}], \\ &f_0(Q_0(1-1/p(x),x)\mid x) \in [f_{\min},f_{\max}], \\ &f_0(Q_0(1/p(x),x)\mid x) \in [f_{\min},f_{\max}], \end{align*} and (ii) uniformly in $(z,x)$, for $|u|\le C$, \begin{align*} &|G_L(Q_1(p(x),x) + u|1,z,x) - G_L(Q_1(p(x),x)|1,z,x)| \leq L_G |u|, \\ &|G_U(Q_1(1-p(x),x) + u|1,z,x) - G_U(Q_1(1-p(x),x)|1,z,x)| \leq L_G |u|, \\ &|G_L(Q_0(1-1/p(x),x) + u|0,z,x) - G_L(Q_0(1-1/p(x),x)|0,z,x)| \leq L_G |u|, \\ &|G_U(Q_0(1/p(x),x) + u|0,z,x) - G_U(Q_0(1/p(x),x)|0,z,x)| \leq L_G |u|. \end{align*} • (Margin Condition) The distribution of the positive and negative monotonicity type is well-behaved around the margin of indifference, i.e., there exist $C_M<\infty$ and $\kappa>0$ such that \[ P\!\left(\big|-[m(1,X)-m(0,X)]-[r(1,X)-r(0,X)]\big|\le t\right)\le C_M\,t^\kappa\quad\text{for all }t>0. \] • (Strong Overlap) There are comparable units across instrument levels and within the differently treated and selected populations. Moreover, always-selected complier probabilities are bounded away from zero on their relevant supports, i.e., there exists some $\underline{c} \in (0,1)$ such that \begin{align*} c < \inf_{z,x}\{r(z,x),m(z,x),e(x)\} \leq \sup_{z,x}\{r(z,x),m(z,x),e(x)\} < 1-c. \end{align*} and \begin{align*} \inf_{x\in\mathcal X_{ac}^+}-[m(1,x) - m(0,x)]> c,\qquad \inf_{x\in\mathcal X_{ac}^-}[r(1,x) - r(0,x)]> c. \end{align*} • (Machine Learning Bias) Let $u_n = o(1)$. For all folds, the nuisance parameters obtained via cross-fitting belong to a shrinking neighborhood $\mathcal{H}_n$ around $\eta$ with probability of at least $1-u_n$, such that, uniformly over the neighborhood, the nuisance functions are consistent \begin{align*} \|\Delta\mu\|_2\ + \|\Delta\nu\|_2\ + \|\Delta e\|_2\ + \|\Delta m\|_{\infty}\ + \|\Delta r\|_{\infty} + ||\Delta G||_{\infty,\mathcal{N}} + ||\Delta F ||_{\infty,\mathcal{N}} &= o(1), \end{align*} and obey convergence rates \begin{align*} &(\|\Delta\mu\|_2\ + \|\Delta\nu\|_2) \|\Delta e\|_2\ + (||\Delta m||_{\infty} + ||\Delta r||_{\infty})^{\kappa + 1} +(||\Delta e ||_2 + ||\Delta m ||_2 + ||\Delta r ||_2) \times \\ &\bigg[||\Delta e ||_{\infty} + ||\Delta m ||_{\infty} + ||\Delta r ||_{\infty} + ||\Delta G||_{\infty,\mathcal{N}} + ||\Delta F ||_{\infty,\mathcal{N}} \bigg] = o(n^{-1/2}). \end{align*}

}

Assumption A.1 and A.2 are simple regularity conditions that rule out heavy tails and degenerate DGPs where bounds are close or equal to a point or otherwise degenerate.

Assumption A.3 rules out point masses at the trimming thresholds in the observed selected outcome distributions used to construct the reduced-form nuisance functions. This ensures that the trimming indicators are unambiguous at the cutoff and that weak and strict inequality conventions coincide at the relevant thresholds.\footnote{The theory can be extended to mass points as discussed in Section (ref). However, in contrast to identification where mass points are harmless once one uses fractional trimming huber2015sharp,kitagawa2021theidentificationregion, inference can be affected as the functional may be nonregular. This problem arises only in boundary cases where the trimming probability coincides exactly with the edge of a mass point. For DGPs in which the cutoff lies in the interior of a mass point, fractional trimming yields a regular functional, so standard root-$n$ semiparametric inference can proceed using the corresponding influence function.}

Assumption A.4(i) is complementary to A.3 and provides primitive density regularity that guarantees local invertibility and differentiability. It controls error propagation through the quantile mapping (Bahadur-type representation). A.4(ii) adds local smoothness to the trimmed mean functions, ensuring uniform control over the nuisance functions evaluated close to the trimming points.

Assumption A.5 is a margin condition that controls the mass of observations near the boundary between positive and negative selection monotonicity types among compliers. In particular, it rules out excessive concentration of probability that would make correct classification difficult. In the case of a bounded density, A.5 holds with $\kappa = 1$, see, e.g., audibert2007fast or heiler2024treatmentevaluationintensiveextensive for related assumptions and discussion. The larger $\kappa$, the less demanding the convergence requirements in A.7 for learning the conditional joint probability of being selected and in treatment/control status. It implies the necessary regularity condition $P(\mathcal{X}^0) = 0$.

Assumption A.6 assures that there are comparable units for instrument and the treatment within the selected group and a relevant share of always-selected compliers. Strong overlap is imposed to obtain a finite variance bound and avoid irregular identification khan2010irregular,HEILER2021valid.

Assumption A.7 imposes the standard DML requirement that all nuisances are consistent and converge to their truth sufficiently fast for the second-order remainder of the orthogonal von Mises expansion to be \(o(n^{-1/2})\), but it does so in a relatively weak form tailored to our target parameter and estimation procedure and its primitives. In particular, in contrast to much of the Lee bounds-type literature, we do not assume global uniform convergence rates for the estimated quantiles or trimmed mean functionals. Instead, it only imposes local sup-norm control of the estimated CDFs \(F\) and trimmed means \(G_B\) on neighborhoods relevant for trimming, with the regularity of quantiles and trimmed means derived from these local conditions. The first product term in A.7 matches the usual \(L^2\)-rate requirement in chernozhukov2018double, while the additional terms capture the non-smooth features of monotonicity type classification and trimming indicators. This is related to rate conditions in heiler2024heterogeneous and semenova2025generalized for generalized Lee bounds as well as heiler2024treatmentevaluationintensiveextensive for more general intensive and extensive margin treatment effects, but expressed directly in terms of CDF errors rather than through separate uniform convergence assumptions on the associated quantile-based functionals. Examples for global uniform rates of nonparametric and machine learning estimation of conditional CDFs can be found in, e.g., xie2023uniform or cattaneo2025uniformestimationinferencenonparametric respectively.

We obtain the following Theorem:

theoremUnder Assumptions (ref), (ref), and A.1--A.7, the estimated bounds are jointly asymptotically normal and semiparametrically efficient, i.e.\begin{align*} \sqrt{n}(\hat{\beta} - \beta) \overset{d}{\rightarrow} \mathcal{N}\big(0,E[\mathbbm{IF}(\beta)\mathbbm{IF}(\beta)']\big). \end{align*}

Theorem (ref) can be used directly for inference on the bounds using the usual standard normal critical values. Importantly, the assumptions imply that the identified set always has a non-empty interior in the population. Thus, Theorem (ref) is sufficient to construct tighter confidence intervals for the effect $\theta_{SLATE}$ directly using imbens2004confidence or stoye2020simple critical values. We use the latter in our empirical applications.

Always-selected Complier Profiling

$\theta_{SLATE}$ is the average treatment effect for the target population $ac$, the always-selected compliers. While individual members of this group are not identified and the group may differ from observed subpopulations or the overall population in both observed and unobserved characteristics, its observable features can nevertheless be characterized, much like complier profiling for the LATE.\footnote{See, for example, abadie2003semiparametric, angrist2004treatmenteffectheterogeneity, and singh2023doublerobustnessforcomplier for various approaches to complier profiling.} This makes it possible to compare always-selected compliers with other populations of interest and thereby better assess the external validity and substantive importance of the effect bounds.

Observable $ac$ characteristics can be identified using the same ingredient, $\pi_{ac}(x)$, as the SLATE bounds in Section (ref). Debiased estimation and inference follow analogously from our influence functions. In particular, for any integrable $g:\mathcal{X}\rightarrow \mathbb{R}$, consider target parameter

align[align omitted — 72 chars of source]

This is the average $g(x)$ in the population of always-selected compliers. Its influence function is given by

align[align omitted — 168 chars of source]

where nuisances and influence functions of the components can be found in Section (ref). The corresponding estimator can be obtained by solving the empirical analogue as in Algorithm (ref). It is root-$n$ consistent and asymptotically normal analogously to Theorem (ref). Moreover, influence function (ref) can directly be used for statistically valid comparisons between mean covariates $g(x)$ of $ac$ and other populations such as unconditional, treated or assigned units. We provide some specific examples in Section (ref) and Appendix (ref).

Empirical Study I: Job Corps Revisited

Data and Methods

In this section, we re-evaluate the earnings effect of participating in JC, a large US federally funded training program providing free academic education, vocational training and employment assistance to disadvantaged youth. We make use of the National Job Corps Study by Mathematica Policy Research. This experiment implemented stratified randomized assignment of applicants, incorporating over 15,400 individuals between ages 16 and 24. Multiple outcomes such as earnings and job status were gathered at various points after assignment.

Our data and main variables are identical to chen2015bounds. In particular, we use a subset of 9,090 units (3,599 control and 5,491 treated) with non-missing work hours, earnings, and participation information. The outcome is log hourly wages which is only observed for the employed. Assignment is given by the original randomization. Treatment is defined as eventual JC participation within the 208 weeks of the evaluation period. We additionally make use of socio-economic pre-assignment covariates including job and earnings history, education, parental background and more, matching the variables used in lee2009training for ITT bounds. The list of covariates along with sample summary statistics are provided in Table (ref).

The data are suitable for our method: Assignment was based on stratified randomization justifying conditional independence. Exclusion is credible as any earnings effects likely require actual training and not just assignment. Importantly, there was a significant amount of noncompliance with the randomized treatment assignment. In particular, only 73.8% of individuals assigned to the treatment group ever participated in JC. Additionally, a small fraction (4.4%) of individuals assigned to the control group also ended up participating.\footnote{Among the 4.4%, 1.2% of controls enrolled in JC before the end of the embargo, while 3.2% enrolled afterward.}

We evaluate the always-selected complier effect $\theta_{SLATE}$ at week $208$ after assignment using (i) chen2015bounds bounds and (ii) sharp-basic bounds -- both assuming strong sample selection monotonicity -- as well as (iii) sharp DML bounds with covariates under weak monotonicity. Implementation of (i) follows chen2015bounds. (ii) uses simple sample analogues of (ref) and (ref). (iii) is obtained via the procedure in Section (ref), see Appendix (ref) for additional details. We also provide a profiling analysis of always-selected compliers using the method from Section (ref).

Results: Bounds and Shares

table[table omitted — 1,881 chars of source]

Table (ref) reports the evaluation results. It delivers four main messages. First, the target population is empirically relevant: the estimated share of always-selected compliers is about 39% across specifications, so $\theta_{SLATE}$ pertains to a sizable subpopulation. Second, the data do not support imposing strong sample selection monotonicity: the estimated share in the positive sample selection region is 93.3% and strong sample selection monotonicity is rejected ($p<0.01$).\footnote{Negative employment effects at week 208 are economically plausible because employment responses to Job Corps are heterogeneous across applicants. Consistent with this, semenova2025generalized shows that positive conditional employment effects do not arise for all applicants even four years after random assignment and documents subgroups with significantly negative employment effects at later horizons.} Thus, Sharp-DML appears to be the most credible specification. Third, under the strong monotonicity benchmark without covariates, the sharp bounds and effect confidence interval are substantially tighter than CF, by 68.4% and 46.7%, respectively. Moreover, the estimated basic bounds no longer include zero and the corresponding confidence interval rules out negative effects below -1.8%. Fourth, estimates using Sharp-DML are broadly comparable to CF despite relying only on weak sample selection monotonicity. Both rule out negative effects beyond -5.1% and -6.1% respectively. However, despite imposing less restrictive assumptions, Sharp-DML delivers tighter bounds and confidence intervals than CF by 9.8% and 7.9% respectively, highlighting the relevance of both sharpness and efficient inclusion of covariates.

The results also speak to the importance of the no-scaling result in Proposition (ref). In particular, naively scaling basic ITT bounds $[-0.019,0.093]$ lee2009training with the JC compliance probability of 69.4% would suggest SLATE bounds of $[-0.027, 0.134]$ which vastly exceed our Sharp-basic bounds of $[0.019, 0.067]$.

Always-selected Complier Profiling

table[table omitted — 4,553 chars of source]

We now conduct the $ac$ profiling analysis as discussed in Section (ref). Table (ref) contains the baseline characteristics of always-selected compliers, $ac$, and those of the full sample. Demographic differences are modest but systematic: On average, $ac$ individuals are less likely to be female (-4.4 pp) and Black (-6.7 pp), slightly more likely to be Hispanic (+2.2 pp) and have fewer children (-4.2 pp in incidence; -0.08 in number). Own education is marginally higher (+0.11 years), with parental education and age being similar across groups. $ac$ individuals are slightly less concentrated in the lowest household income bracket and more represented in the highest (1.9 pp each). The most pronounced differences arise in pre-assignment labor market outcomes: $ac$ individuals are significantly more likely to be employed at baseline (+4.7 pp), more likely to have worked in the previous year (+8.7 pp), and exhibit higher labor supply and earnings, including +0.81 months worked, +\$600 annual earnings, +2.9 weekly hours, and +\$14 weekly earnings.

Taken together, these differences indicate that $ac$ subpopulation is positively selected, in particular on baseline labor-market attachment. The stronger pre-program employment and earnings profiles suggest that $ac$ may be better positioned to translate JC participation into earnings gains, but also imply more favorable counterfactual trajectories in the absence of treatment. At the same time, prior evidence on JC shows that impacts operate through channels such as increased GED, increased vocational credential attainment and reduced criminal involvement, which are not confined to the most labor-market-ready participants schochet2008does,flores2012estimating. Thus, while the observable composition of the $ac$ group points to potential attenuation when extrapolating our estimates for $\theta_{SLATE}$ to the full sample, baseline differences alone do not fully pin down the direction or magnitude of the extrapolation.

Concluding Remarks

For both identification and semiparametric inference, this paper provides a synthesis of treatment evaluation under sample selection as well as noncompliance building on Lee-type bounds and the LATE framework. Given the analytic form of our population bounds, it is relatively simple to incorporate additional restrictions, such as stochastic dominance of $ac$ potential outcome distribution over that of $cc$ to further tighten bounds zhang2003estimation,huber2015sharp,heiler2024treatmentevaluationintensiveextensive. Moreover, given the simple ratio-of-linear-moment structure, our influence function components can be directly leveraged for estimation and inference on heterogeneous group-specific $ac$ effect bounds by combining them with nonparametric projection methods as in heiler2024heterogeneous.

\addcontentsline{toc}{section}{References} {\setstretch{1} }

Declaration of generative AI and AI-assisted technologies in the manuscript preparation process

During the preparation of this work, the authors used OpenAI's ChatGPT to assist with mathematical proofs, coding, formatting of tables and outputs, and spelling. The authors reviewed and edited the content as needed and take full responsibility for the content of the article.