Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Conformalized Lee Inference: Distribution-Free Individual Treatment Effect Intervals under Monotone Sample Selection
center[center omitted — 13 chars of source]
abstractEmpirical studies often observe outcomes only for selected units, and treatment may change who is observed. This paper studies prediction in randomized studies with one-sided selection. Standard prediction intervals can fail because treated selected observations are not the same group as selected controls. The paper asks how to predict missing treated outcomes and individual treatment effects for always-observed units. The proposed conformalized Lee procedure uses treated selected observations to train and check any prediction rule, then adjusts the cutoff using the observed treatment-control selection gap. For selected controls, the missing treated-outcome interval is shifted by the observed untreated outcome to produce an individual treatment-effect interval. The method provides reliable coverage without requiring the prediction rule to be correctly specified. The key result shows that the proposed adjustment uses the exact amount of uncertainty implied by the monotone selection logic of lee2009training. In simulations, ordinary conformal prediction demonstrates a lower coverage rate under selection-induced distribution shift, while the Lee-adjusted methods achieve the desired coverage rate. The results show that the proposed selection correction method can support reliable counterfactual prediction, while retaining practical implementation with modern prediction tools.
JEL Classification: C14, C21, C24, C53.
Keywords: Conformal prediction; Lee bounds; sample selection; individual treatment effects; counterfactual prediction; partial identification; distribution-free inference.
Introduction
commentEmpirical studies often observe outcomes only for units who are selected into the sample. This creates a fundamental problem when treatment itself affects selection. In that case, the treated-selected and control-selected groups may represent different latent populations, so comparisons among observed units combine treatment effects with changes in who is observed. The monotone-selection framework of lee2009training addresses this problem by assuming that treatment can only increase the chance of being selected, not decrease it. Under this assumption, selected-control units come from the always-selected population, while treated-selected units may include both always-selected units and additional marginal units induced into selection by treatment.
This paper develops “conformalized Lee inference”, a distribution-free method for constructing prediction intervals under monotone sample selection. The goal is not only to bound an average treatment effect, but to provide distribution-free predictive intervals for counterfactual outcomes of randomly drawn always-selected units. For selected controls, the untreated outcome is observed and monotonicity implies membership in the always-selected stratum. Therefore, a predictive interval for the missing treated potential outcome can be translated into a predictive interval for that unit’s realized treatment effect by subtracting the observed untreated outcome. The resulting guarantee is marginal over the selected-control population, rather than conditional on a fixed individual’s realized covariates and untreated outcome.
The main challenge is that the data available for calibration are not drawn from the same distribution as the target counterfactual population. Treated-selected observations come from a mixture of always-selected and marginal-in units. By contrast, the target population consists only of always-selected units. The share of always-selected units within the treated-selected group is identified from the selection rates in the treatment and control arms. This share plays the same role as the trimming proportion in the original Lee bounds.
The paper first characterizes the set of possible distributions for the always-selected treated outcomes. Because the always-selected population is nested inside the treated-selected population, the target distribution cannot place probability on events that never occur among treated-selected observations. More strongly, it cannot overweight any event by more than the inverse of the always-selected share. This gives a likelihood-ratio type restriction: the unknown target law must be dominated by the observable treated-selected law, with the amount of possible upweighting limited by the selection-rate ratio.
This restriction is the distributional analogue of the trimming argument of lee2009training. Lee bounds trim the treated-selected outcome distribution to account for the unknown identity of the always-selected units. Here, the same logic is applied to the full joint distribution of covariates and treated potential outcomes. Rather than identifying a single target distribution, the model identifies a class of distributions that are all consistent with the observed treated-selected data and the monotone-selection assumptions.
The paper then shows that this class is sharp. In other words, every distribution in the class can arise from some data-generating process that satisfies the maintained assumptions and matches the same observed selection rates and treated-selected distribution. Therefore, the class is not a conservative approximation. It is exactly the set of target counterfactual laws that cannot be ruled out without imposing additional assumptions. This sharpness result is important because it determines the level of robustness required from any valid prediction procedure.
Building on this identification result, the paper proposes a split-conformal prediction method. A prediction model is trained using treated-selected observations, and a separate treated-selected calibration sample is used to compute nonconformity scores. Ordinary split conformal prediction would calibrate these scores as if the future target unit were drawn from the same distribution as the calibration sample. That is not valid here, because the calibration observations come from the treated-selected mixture, while the target counterfactual unit comes from the always-selected population.
The likelihood ratio restriction provides the needed adjustment. Since the always-selected distribution can concentrate more heavily than the treated-selected distribution on high-score events, the conformal cutoff must be more conservative than the usual one. Specifically, the procedure calibrates the treated-selected score distribution at a higher quantile, using the selection-rate ratio to inflate the required calibration level. Intuitively, the method asks for the score threshold to make the treated-selected tail small enough that, even after the worst allowable upweighting, the target counterfactual tail remains below the desired error level.
The resulting procedure is simple. Estimate the share of always-selected units from the treatment and control selection rates. Train any predictive model using treated-selected observations. Compute calibration scores on held-out treated-selected observations. Then choose a conformal cutoff that accounts for the identified amount of selection-induced contamination. The prediction set for a new always-selected unit contains all treated outcomes whose nonconformity scores fall below this cutoff. The validity of the procedure does not depend on correct model specification; the prediction model may be arbitrary, and coverage is obtained from the conformal calibration step together with the monotone-selection structure.
The same construction also yields intervals for individual treatment effects among selected-controls. For such units, the untreated outcome is observed, and monotonicity implies that they belong to the always-selected stratum. Therefore, a valid interval for the missing treated potential outcome can be translated directly into a valid interval for the individual treatment effect by subtracting the observed untreated outcome. This gives a distribution-free way to move from Lee-style selection correction to marginal predictive inference for counterfactual outcomes and realized treatment effects among selected controls.
Finally, the paper shows that the conformal threshold is minimax optimal over the reduced-information class of possible target distributions. The key point is that the worst-case always-selected distribution can place as much mass as allowed on the upper tail of the nonconformity score. As a result, the adjusted conformal quantile is not merely a sufficient conservative choice. It is the smallest population-level cutoff that guarantees uniform coverage over all target laws consistent with the monotone-selection model. Any smaller cutoff would fail for some data-generating process that is observationally indistinguishable under the maintained assumptions.
This perspective distinguishes conformalized Lee inference from sensitivity analysis approaches to individual treatment effect prediction. In sensitivity analysis methods, the analyst typically chooses a sensitivity parameter that governs how far the target distribution may deviate from the calibration distribution (see, e.g., rosenbaum1987sensitivity, yadlowsky2022bounds, jin2023sensitivity, yin2024conformal, dorn2023sharp, dorn2025doubly). Here, the robustness parameter is not chosen by the analyst. It is identified from the observed selection rates. Thus, the method converts the classical Lee selection model into a distribution-free conformal prediction procedure whose coverage guarantee and optimality are both driven by the identified structure of the sample-selection problem.
The selection-induced shift considered in this paper is generally more complex than
ordinary covariate shift. Under covariate shift, the marginal distribution of
$X$ may change while the conditional outcome law given $X$ remains fixed.
Here, principal-stratum membership may be related to both $X$ and $Y(1)$. Thus, the target law $Q$ may differ from the calibration law $P$ in both its covariate
distribution and its conditional outcome distribution. The likelihood ratio
restriction is therefore imposed on the joint law of $(X,Y(1))$. The method
does not require $Q(Y(1)\mid X)=P(Y(1)\mid X)$, and should be understood as robust prediction under a partially identified
joint distribution shift.
Introduction
Empirical studies often observe outcomes only for units who are selected into the sample. This creates a fundamental problem when treatment itself affects selection. In a randomized treatment setting, selected treated and selected control units need not represent the same latent population: comparisons among observed outcomes may combine treatment effects with changes in who is observed. The monotone-selection framework of lee2009training addresses this problem by assuming that treatment can only increase selection. Under this restriction, selected controls are always-selected units, while treated-selected units are a mixture of always-selected units and marginal-in units induced into selection by treatment.
This paper develops conformalized Lee inference, a distribution-free method for constructing prediction intervals under monotone sample selection. The primary target is a prediction set for the treated potential outcome of a future draw from the always-selected population. For selected controls, the untreated outcome is observed and monotonicity implies membership in the always-selected stratum. Hence, a prediction set for the missing treated potential outcome can be translated into a prediction interval for the realized individual treatment effect by subtracting the observed untreated outcome. The resulting guarantee is marginal over the selected-control population, not conditional on a fixed individual's covariates and untreated outcome.
The central difficulty is that the calibration data are not drawn from the target law. Treated-selected observations come from a mixture of always-selected and marginal-in units, whereas the target law contains only always-selected units. Let \(P\) denote the observed treated-selected law of \((X,Y(1))\), and let \(Q_0\) denote the corresponding always-selected target law. If the selection rate is \(p_d=\Pr(S(d)=1)\) for $d \in \{0,1\}$, then monotonicity identifies
\[
\pi=\frac{p_0}{p_1},
\]
the fraction of treated-selected units who are always-selected. The identity of those units is not observed, so \(Q_0\) is not point identified.
The first contribution is to characterize the resulting distributional identification problem. The always-selected target law must be dominated by the treated-selected law and cannot overweight any event by more than \(1/\pi\):
\[
Q_0 \ll P,
\qquad
0 \leq \frac{dQ_0}{dP} \leq \frac{1}{\pi}.
\]
This is the distributional analogue of Lee's trimming argument. Lee bounds trim the treated-selected outcome distribution because the always-selected treated units are unidentified. Here, the same nesting logic is applied to the joint law of covariates and treated potential outcomes. The paper shows that this ambiguity set is sharp relative to the reduced observables \((P,p_0,p_1)\): every law satisfying the likelihood-ratio restriction can arise under random assignment, monotone selection, and the same observed selection rates.
The second contribution is to turn this identification result into a conformal prediction procedure. Ordinary split conformal prediction would calibrate treated-selected scores as if the future target unit were drawn from the same law as the calibration sample. That is generally invalid here. The Lee ambiguity set gives the necessary correction. Since the target law may place up to \(1/\pi\) times more probability on high-score events than the treated-selected law, the method calibrates at the higher treated-selected score quantile \(1-\alpha\pi\), rather than at the usual \(1-\alpha\) quantile. Equivalently, the treated-selected score tail is made small enough that even the worst admissible always-selected target law has tail probability at most \(\alpha\).
The resulting algorithm is simple. It trains any prediction rule using treated-selected observations, computes held-out nonconformity scores from treated-selected calibration observations, and chooses the Lee-adjusted conformal cutoff. The coverage result is finite-sample and distribution-free for arbitrary prediction rules and score functions. When the always-selected share \(\pi\) is unknown, the paper distinguishes plug-in estimation from lower-bound implementation. If the value used in calibration is no larger than the true \(\pi\), the procedure remains conservative; one-sided lower confidence bounds therefore provide formal high-probability coverage guarantees, while plug-in estimates offer a less conservative empirical implementation.
The third contribution is a minimax interpretation of the cutoff. Over the sharp Lee ambiguity set, the worst-case target law places as much mass as allowed on the upper tail of the nonconformity score. Therefore, the \(1-\alpha\pi\) population score quantile is not merely a sufficient conservative adjustment. It is the smallest score threshold that guarantees uniform coverage over all target laws consistent with the reduced monotone-selection information. Any smaller threshold fails for some observationally equivalent data-generating process.
The simulations compare naive conformal prediction, oracle Lee adjustment, plug-in Lee adjustment, and lower-bound Lee procedures across benign, tail-selection, and smooth-selection designs. Naive conformal prediction is valid when the calibration and target laws coincide, but it under-covers when always-selected units are more likely to have large nonconformity scores. The Lee-adjusted procedures restore coverage close to the target in the shifted designs, while Clopper--Pearson and Hoeffding lower-bound versions are conservative, as expected.
The rest of the paper proceeds as follows. Section 2 introduces the setup, principal strata, target population, and Lee ambiguity set, and establishes sharp identification under reduced observables. Section 3 develops the split-conformal Lee procedure, proves finite-sample coverage, derives individual treatment-effect intervals for selected controls, and establishes minimax optimality of the cutoff. Section 4 reports simulation evidence. The appendix gives proofs, practical lower-bound constructions for \(\pi\), and the full observed law identification result.
comment\subsection{Related Work}
\paragraph{Sample selection.}
This paper contributes to the literature on treatment effect inference under sample selection (see, e.g., lee2009training, honore2020selection, honore2024sample, lee2024lee, kurisu2026lee). The monotone selection framework of lee2009training provides a restriction for randomized or conditionally randomized settings. Specifically, when treatment weakly increases selection, selected controls are drawn from the always-selected population, while treated-selected observations combine always-selected and marginal-in units. lee2009training uses this structure to obtain sharp bounds on average treatment effects by trimming the treated-selected outcome distribution. The present paper keeps the monotonicity structure but changes the inferential target. Rather than bounding only an average treatment effect, it constructs distribution-free prediction intervals for counterfactual treated outcomes and individual treatment effects in the always-selected stratum.
Recent work has extended Lee-type selection correction in covariate-rich settings (see, e.g., heiler2024heterogeneous, semenova2025generalized, dong2026sharp). For instance, generalized Lee bounds of semenova2025generalized allow the direction of selection to vary with pre-treatment covariates and develop modern inference methods for low- and high-dimensional settings. This is especially relevant for the full data version of the present paper, where the selected-control covariate distribution identifies the always-selected covariate marginal and the conditional Lee restriction becomes $dQ_x / dP_x \le 1 / \pi(x)$. Thus, the paper can be positioned as complementary to generalized Lee bound methods: it uses the same selection logic to obtain distribution-free predictive inference rather than average effect bounds.
\paragraph{Principal stratification.}
The paper is also closely related to principal stratification. frangakis2002principal introduce principal stratification as a framework for defining causal effects within latent subpopulations determined by post-treatment variables. zhang2003estimation develop this logic for outcomes truncated by death. The present setting has the same structure, with selection replacing survival. Under monotone selection, the population decomposes into always-selected, marginal-in, and never-selected strata, and the target population is the always-selected stratum. The main difference is that this paper uses the principal-strata structure not only to define the causal target, but also to derive the distributional restriction needed for valid conformal calibration.
\paragraph{Contamination models.}
The identification result connects Lee bounds to partial identification and contamination models (see, e.g., huber1981robust). horowitz1995identification study identification under contaminated and corrupted sampling, while Manski’s work on partial identification emphasizes set-valued identification of probability distributions when the available data do not support point identification. In the present paper, the treated-selected law can be written as $P = \pi Q + (1-\pi)R$, where $Q$ is the always-selected target law and $R$ is an unrestricted marginal-in distribution. Monotone selection and random assignment identify $\pi = p_0 / p_1$, which yields the sharp likelihood ratio ambiguity set. This is the distributional analogue of trimming logic introduced by lee2009training.
\paragraph{Conformal prediction.}
On the prediction side, the paper builds on conformal prediction, beginning with the exchangeability based framework of vovk2005algorithmic and the regression focused split-conformal methods developed by lei2018distribution. Standard split conformal prediction is not directly applicable here because the calibration scores are computed from treated-selected observations drawn from $P$, while the target counterfactual law is $Q$. This connects the paper to conformal prediction under distribution shift, including weighted conformal prediction under covariate shift and more general conformal methods beyond exchangeability. The distinction is that in the present setting the likelihood ratio is not known pointwise. Instead, monotone selection identifies a sharp upper bound $dQ / dP \le 1 / \pi$, which leads to the adjusted $(1 - \alpha\pi)$-quantile cutoff.
Finally, the paper contributes to conformal causal inference. lei2021conformal develop conformal methods for counterfactuals and individual treatment effects, while chernozhukov2021exact use conformal ideas for counterfactual and synthetic control inference. The closest comparison is jin2023sensitivity, who propose robust conformal sensitivity analysis for individual treatment effects. The present paper differs because the robustness class is not chosen by the analyst as a sensitivity model. Instead, the Lee monotonicity assumption and the observed treatment-control selection rates identify the ambiguity set and therefore determine the conformal cutoff. This yields both a distribution-free coverage guarantee and a minimax interpretation tied directly to the sample-selection model.
Related Work
This paper connects four literatures: treatment-effect inference under sample selection, principal stratification and partial identification, contamination models, and conformal prediction under distribution shift and causal inference.
\paragraph{Sample selection and Lee bounds.}
This paper contributes to the literature on treatment-effect inference under sample selection (see, e.g., lee2009training, honore2020selection, honore2024sample). The closest starting point is lee2009training, who shows that monotone selection can deliver sharp bounds on average treatment effects by trimming the treated-selected outcome distribution. The present paper keeps the same monotone-selection logic but changes the inferential target. Rather than bounding only an average treatment effect, it constructs distribution-free prediction sets for counterfactual treated outcomes and individual treatment effects in the always-selected population.
Recent work extends Lee-type selection correction in richer settings, including covariate-rich selection models, generalized Lee bounds, continuous treatments, random objects, and treatment endogeneity (see, e.g., heiler2024heterogeneous, semenova2025generalized, lee2024lee, kurisu2026lee, dong2026sharp). These papers focus primarily on bounding average or distributional causal parameters. The present paper is complementary: it uses Lee-type nesting to derive a sharp ambiguity set for the target counterfactual law and then converts that ambiguity set into a conformal prediction rule. The appendix's full observed law result is especially related to covariate-rich Lee methods, since it shows that the selected-control covariate distribution identifies the target covariate marginal and that the conditional treated-outcome law satisfies the local restriction
\[
Q_x \ll P_x, \qquad 0 \leq \frac{dQ_x}{dP_x} \leq \frac{1}{\pi(x)}.
\]
The main procedure, however, deliberately works with the reduced information \((P,p_0,p_1)\), yielding a simple and distribution-free robust calibration rule.
\paragraph{Principal stratification and partial identification.}
The paper is also related to principal stratification. frangakis2002principal introduce principal stratification as a framework for defining causal effects within latent subpopulations determined by post-treatment variables, and zhang2003estimation apply this logic to outcomes truncated by death. The present setting has the same structure, with sample selection replacing survival. Under monotone selection, selected controls reveal the always-selected stratum, while treated-selected observations combine always-selected and marginal-in units. The contribution here is to use this principal-strata structure not only to define the causal target, but also to derive the distributional restriction required for valid conformal calibration.
The identification result is also connected to partial identification and contamination models (see, e.g., huber1981robust, horowitz1995identification, manski2003partial). The observed treated-selected law can be written as
\[
P=\pi Q+(1-\pi)R,
\]
where \(Q\) is the always-selected target law and \(R\) is an unrestricted marginal-in law. Because the identity of always-selected treated units is unobserved, \(Q\) is not point identified. Monotone selection and random assignment identify the mixture share \(\pi=p_0/p_1\), which yields the sharp likelihood-ratio ambiguity set. This is the distributional analogue of Lee's trimming argument.
\paragraph{Conformal prediction under distribution shift.}
On the prediction side, the paper builds on conformal prediction under exchangeability and split conformal prediction for regression (see, vovk2005algorithmic, lei2018distribution). Standard split conformal prediction is not directly valid here because calibration scores are computed from treated-selected observations drawn from \(P\), while the target counterfactual law is an unknown \(Q\) in the Lee ambiguity set.
This connects the paper to conformal prediction under distribution shift, including weighted conformal prediction under covariate shift and conformal prediction beyond exchangeability (see, tibshirani2019conformal, barber2023conformal). The distinction is that, in the present setting, the pointwise density ratio is not known. Instead, Lee monotonicity identifies a sharp upper bound on the joint density ratio, \(dQ/dP\leq 1/\pi\). Moreover, the shift is not only a change in the covariate distribution: principal-stratum membership may be related to both \(X\) and \(Y(1)\), so the target law may differ from the calibration law in both its covariate marginal and its conditional treated-outcome distribution. The conformalized Lee cutoff is therefore a robust conformal threshold for a partially identified joint distribution shift.
\paragraph{Conformal causal inference and sensitivity analysis.}
The paper also contributes to conformal causal inference. lei2021conformal develop conformal methods for counterfactuals and individual treatment effects, and chernozhukov2021exact use conformal ideas for counterfactual and synthetic control inference. The closest comparison is robust conformal sensitivity analysis for individual treatment effects (see, e.g., jin2023sensitivity, yin2024conformal). Related sensitivity-analysis approaches study robustness to deviations such as unobserved confounding or analyst-specified departures from identifying assumptions (see, e.g., rosenbaum1987sensitivity, yadlowsky2022bounds, dorn2023sharp, dorn2025doubly).
The present paper differs because the robustness class is not chosen by the analyst as a sensitivity model. It is identified by Lee monotonicity and the observed treatment-control selection rates. This delivers a closed-form adjustment from the usual \((1-\alpha)\) conformal cutoff to the \((1-\alpha\pi)\) treated-selected score quantile. The paper further shows that this cutoff is minimax optimal over the reduced-information Lee ambiguity set: any smaller population threshold fails for some target law that is observationally equivalent under the maintained monotone-selection assumptions.
Setup, Principal Strata, and Predictive Target
We consider a randomized treatment setting with sample selection. For each unit $i=1,...,n$, let $D_i \in \{0,1 \}$ denote treatment assignment, let $X_i \in \mathcal{X} \subseteq \mathbb{R}^d$ denote predetermined covariates, and let $Y_i(1), Y_i(0)$ be the treated and untreated potential outcomes. Outcomes are observed only when the unit is selected into the sample. Let $S_i(1), S_i(0) \in \{0,1\}$ denote the potential selection indicators under treatment and control. The realized selection indicator is $S_i = D_i S_i(1) + (1 - D_i)S_i(0)$, and the observed outcome is $Y_i = D_i Y_i(1) + (1 - D_i) Y_i(0)$ when $S_i = 1$. Thus, the observed data are $\{(D_i,X_i,S_i, S_i Y_i)\}_{i=1}^n$.
The central challenge is that the treated-selected sample is not necessarily the same population as the always-selected target population. Under monotone selection, treatment may induce additional units to become selected. Therefore, treated-selected observations contain both always-selected units and marginal-in units, while selected-controls contain only always-selected units.
We impose the following assumptions.
assumption[Random assignment]
\[
(Y(1),Y(0),S(1),S(0),X) \perp \!\!\! \perp D.
\label{ass:independ}
\]
assumption[Monotone Selection]
\[
S(1) \geq S(0) \quad \text{almost surely.}
\label{ass:monosel}
\]
Assumption (ref) identifies potential-outcome and potential-selection distributions from treatment arms. Assumption (ref) rules out units who would be selected under control but not under treatment. Hence, the population can be partitioned into three principal strata:$$
\mathcal{A}=\{S(0)=S(1)=1\},
\qquad
\mathcal{M}=\{S(0)=0,S(1)=1\},
\qquad
\mathcal{N}=\{S(0)=S(1)=0\}.
$$
where $\mathcal{A}$ is the always-selected stratum, $\mathcal{M}$ is the marginal-in stratum induced into selection by treatment, and $\mathcal{N}$ is the never-selected stratum. Because of monotonicity, every selected-control unit belongs to $\mathcal{A}$, whereas a treated-selected unit may belong either to $\mathcal{A}$ or to $\mathcal{M}$.
The object of interest in this paper is the distribution-free prediction problem for individual treatment effects in the always-selected stratum. For an always-selected unit, the individual treatment effect is $$\tau_i = Y_i(1) - Y_i(0), \quad i \in \mathcal{A}.$$
For selected-controls, $Y(0)$ is observed and monotonicity implies membership in $\mathcal{A}$. Therefore, prediction intervals for the missing treated potential outcome $Y(1)$ among always-selected units can be translated directly into prediction intervals for $\tau_i$.
Let $p_d = \Pr(S(d) = 1)$ for $d \in \{0,1\}$. Under Assumption (ref), these are identified by $p_d = \Pr(S = 1 \mid D=d)$ for $d \in \{0,1\}$. Furthermore, under Assumption (ref), $p_1 \ge p_0$. Define $\pi = p_0 / p_1 = \Pr( \mathcal{A} \mid S(1) = 1 )$. Then, $\pi$ is the fraction of treated-selected units who are always-selected. Equivalently, $1 - \pi$ is the fraction of treated-selected units who are marginal-in.
assumption[Positivity]
There exist $\eta_0>0$ and $\eta_1>0$ such that $\eta_0 <p_0 < 1-\eta_0$ and $\eta_1 < p_1 <1-\eta_1$.
This assumption ensures that the selected populations under treatment and control are nondegenerate and that $\pi = p_0 / p_1$ is well-defined.
commentWe observe i.i.d. samples $\{(D_i,X_i',S_i, S_i\cdot Y_i)\}_{i=1}^n$, where $D_i \in \{0,1\}$ is a binary treatment indicator, $X_i \in \mathcal{X} \subseteq \mathbbm{R}^d$ is a vector of predetermined covariates, $S_i \coloneqq S(1) D + S(0)(1-D)$ is the sample selection status where $S_i(1)$ and $S_i(0)$ are potential selection indicators, and $S_i \cdot Y_i \coloneqq S_i \cdot (Y_i(1) D_i + Y_i(0)(1-D_i))$ with $Y_i \in \mathbbm{R}$ is the outcome we observe where $Y_i(1)$ and $Y_i(0)$ are the real-valued treated and untreated potential outcomes, respectively.
The parameter of interest in lee2009training is the average treatment effect for always-observed stratum:
\[
\tau_{\text{ATE}} \coloneqq \operatorname*{\mathbb{E}}[ Y(1) - Y(0) \mid S(0)=S(1)=1 ],
\]
Furthermore, lee2009training introduce the following assumptions to derive the bounds for $\tau_{\text{ATE}}$.
\begin{assumption}[Independence]
\[
(Y(1),Y(0),S(1),S(0),X) \perp \!\!\! \perp D.
\label{ass:independ}
\]
\end{assumption}
\begin{assumption}[Monotone Selection]
\[
S(1) \geq S(0) \quad \text{almost surely.}
\label{ass:monosel}
\]
\end{assumption}
Assumption (ref) states that there is no unobserved confounder. Under Assumption (ref), the treatment weakly increases the selection. Thus, there are three principal strata:
\[
\begin{aligned}
\mathcal{A} &\coloneqq \{ S(0) = S(1) = 1 \}, \\
\mathcal{M} &\coloneqq \{ S(0)=0, S(1) = 1 \}, \\
\mathcal{N} &\coloneqq \{ S(0) = S(1) = 0 \}.
\end{aligned}
\]
where $\mathcal{A}$ is a stratum of “always-selected” (or always-takers / inframarginal), $\mathcal{M}$ is “marginal-in” (or compliers for selection), and $\mathcal{N}$ is “never-selected”. Assumption (ref) requires that any sample in the control group $\{D=0, S=1\}$ must satisfy $(S(0)=1, S(1)=1)$, and so it belongs to $\mathcal{A}$. On the contrary, the treated-selected observations $\{ D=1, S = 1 \}$ may belong to either $\mathcal{A}$ or $\mathcal{M}$.
We hope to learn about the individual treatment effect (ITE) for always-selected stratum:
\[
\tau_i \coloneqq Y_i(1) - Y_i(0), \quad i \in \mathcal{A},
\]
with bounds obtained by trimming the distribution of the treated-selected outcome by the known proportion of excess individuals induced to be selected. Let $\pi$ be the weight of always-selected stratum identified by selection rates:
\[
\pi \coloneqq \Pr (\mathcal{A} \mid D=1, S=1) = \Pr(S(0)=1 \mid S(1)=1 ) = \frac{\Pr(S=1 \mid D=0)}{\Pr (S=1 \mid D=1)} = \frac{p_0}{p_1}.
\]
where $p_0 \coloneqq \Pr(S=1 \mid D=0)$ and $p_1 \coloneqq \Pr(S=1 \mid D=1)$.
\begin{assumption}[Positivity]
There exist $\eta_0>0$ and $\eta_1>0$ such that $\eta_0 <p_0 < 1-\eta_0$ and $\eta_1 < p_1 <1-\eta_1$.
\end{assumption}
Assumption (ref) requires that $\pi=p_0/p_1$ to be well-defined. Because $p_1\geq p_0$ under Assumption (ref), $\pi \in (0,1]$.
Furthermore, define two conditional joint distributions of $(X, Y(1))$:
\[
P \coloneqq \mathcal{L}((X,Y(1)) \mid S(1)=1) \quad \text{and} \quad Q \coloneqq \mathcal{L}((X,Y(1)) \mid S(0)=1),
\]
where $P$ is the observable distribution for treated-selected observations and $Q$ is our target counterfactual distribution for the always-selected principal stratum. Under Assumption (ref), $P=\mathcal{L}( X,Y \mid D=1, S=1 )$. Furthermore, under Assumptions (ref) and (ref), the target population is strictly a subpopulation of the treated-selected population. This nesting structure allows us to formally characterize the identification class for $Q$. The distribution of $P$ is a mixture of always-selected and marginal-in distributions:
\begin{equation}
P = \pi Q + (1-\pi) R,
\end{equation}
where $R$ is the distribution among compliers under treatment, which is unknown and remains unrestricted. The mixture representation implies the contaminated sampling model for the treated-selected sample (see, e.g., huber1981robust, horowitz1995identification).
For any measurable set $B \subseteq \mathcal{X} \times \mathcal{Y}$, a sharp event-probability bound follows immediately:
\[
\max \Bigl\{ 0, \frac{P_{1,\text{obs}}(B) - (1-\pi)}{\pi} \Bigr\} \leq P_{1, \mathcal{A}}(B) \leq \min \Bigl\{ 1, \frac{P_{1,\text{obs}}(B) }{\pi} \Bigr\}.
\]
Therefore, $P_{1, \mathcal{A}} \ll P_{1, \text{obs}}$ and the Radon-Nikodym derivative satisfies the one-sided bound:
\[
0 \leq \frac{d P_{1, \mathcal{A}}}{dP_{1, \text{obs}}}(x,y) \leq \frac{1}{\pi}.
\]
The always-selected treated distribution is an unknown subdistribution of the treated-selected distribution with maximal upweighting $1/\pi$. Thus $P_{1, \mathcal{A}}$ is within bounded distance from $P_{1, \text{obs}}$ for some fixed upper and lower bounds 0 and $1/\pi$. This is the prediction analogue of the worst-case trimming of lee2009training.
Lee-Type Distributional Identification
To state the identification problem, Let $Z = (X,Y(1)) \in \mathcal{Z} = \mathcal{X} \times \mathcal{Y}$, where $(\mathcal{Z}, \mathcal{B})$ is a standard Borel space. Let $\mathcal{P}(\mathcal{Z}, \mathcal{B})$ denote the set of probability measure on $(\mathcal{Z}, \mathcal{B})$. The observed treated-selected distribution is $$P = \mathcal{L}( Z \mid S(1)=1).$$ By random assignment, this law is identified from the treated-selected observations:
$$
P = \mathcal{L}( X,Y \mid D=1, S=1).
$$
The target law is the true distribution of treated potential outcomes for always-selected units:
$$
Q_0 = \mathcal{L}( Z \mid S(0)=1) = \mathcal{L}( Z \mid \mathcal{A}).
$$
The key problem is that $Q_0$ is not observed. Among treated-selected units, only a fraction $\pi$ are always-selected; the remaining fraction $1 - \pi$ are marginal-in. Hence the observable treated-selected law can be written as
$$
P= \pi Q_0 + (1 - \pi) R,
$$ where
$$
R = \mathcal{L} (Z \mid S(0)=0, S(1)=1)
$$ is the distribution of $(X,Y(1))$ among marginal-in units. Because $R$ is unrestricted, the data do not point identify $Q_0$. Instead, they identify a class of possible target laws. The mixture representation implies the contaminated sampling model for the treated-selected sample (see, e.g., huber1981robust, horowitz1995identification).
This mixture representation is the distributional version of Lee’s trimming argument. Lee bounds trim the treated-selected outcome distribution to account for the unknown identity of always-selected units. Here, the same logic is applied to the joint law of $(X,Y(1))$, because the later conformal procedure calibrates nonconformity scores using treated-selected observations drawn from $P$, while the desired prediction guarantee concerns a future always-selected unit drawn from $Q_0$. Furthermore, the nesting $\{ S(0)=1 \} \subseteq \{ S(1)=1 \}$ implies that $Q_0$ must be dominated by $P$. More precisely, the always-selected law cannot put mass on events that have zero probability under the treated-selected law, and it cannot upweight any event by more than $1 / \pi$.
commentThe previous discussion shows that treated-selected observations contain two latent principal strata. Under monotone selection, $\{ S(0)=1 \} \subseteq \{S(1) = 1\}$. Thus, every always-selected unit is included among the treated-selected population, but the treated-selected population may also contain marginal-in units. Consequently, the observed treated-selected law is not the target law itself; it is a mixture of the target always-selected law and an unrestricted marginal-in law.
Let $Z = (X, Y(1))$, and define $P = \mathcal{L}(Z \mid S(1) = 1)$ and $Q = \mathcal{L}(Z \mid S(0) = 1)$. Here, $P$ is dentified from treated-selected observations, since by random assignment $P = \mathcal{L}( (X,Y) \mid D=1, S=1 )$, whereas $Q$ is the latent distribution of treated potential outcomes for the always-selected stratum. The selection-rate ratio $\pi = p_0 / p_1$ is the fraction of treated-selected units who are always-selected.
This nesting structure implies that P can be written as $P = \pi Q + (1 - \pi)R$, where $R$ is the joint distribution of $(X,Y(1))$ among marginal-in units. Since $R$ is unrestricted, the data do not identify $Q$ pointwise. However, the mixture representation does impose a sharp domination restriction: the target law $Q$ must be absolutely continuous with respect to the observed treated-selected law $P$, and no event can receive more than $1/ \pi$ times its probability under $P$.
This restriction is the distributional analogue of Lee’s trimming logic. Lee bounds remove a fraction $1 - \pi$ of the treated-selected outcome distribution to isolate what could be the always-selected component. Here, instead of trimming only outcome means or quantiles, we characterize the full set of joint laws for $(X,Y(1))$ that could correspond to the always-selected stratum.
For $P \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$ and $\pi \in (0,1]$, we define the Lee ambiguity set.
definition[Lee ambiguity set]
\[
\mathcal{Q}(P, \pi) = \{ Q \in \mathcal{P}(\mathcal{Z}, \mathcal{B}): Q \ll P, \quad 0 \le \frac{dQ}{dP} \le \frac{1}{\pi} \quad P\text{-almost surely.} \}
\]
Proposition (ref) formalizes the necessity of this class under Assumptions (ref)--(ref). It shows that the true always-selected treated distribution $Q_0$ must lie in the resulting identification class $\mathcal{Q}(P, \pi)$. Lee ambiguity set can therefore be interpreted as the identified set of possible target distributions under selection-induced distribution shift. The treated-selected law $P$ is the source or calibration distribution, while each
$Q\in\mathcal Q(P,\pi)$ is a possible always-selected target distribution. This is the first step toward conformalized Lee inference. Later sections will construct prediction sets that are valid uniformly over all target laws satisfying this likelihood-ratio bound.
commentTo construct distribution-free prediction sets for counterfactuals under monotone selection, we must formalize the relationship between the observable data distribution and the latent target distribution.
prop[Lee-type likelihood ratio domination]
Suppose Assumptions (ref)--(ref) hold. Then, the true always-selected target law $Q_0$ is absolutely continuous with respect to the treated-selected law $P$ and, its Radon-Nikodym derivative satisfies
$$
\quad 0 \le \frac{dQ_0}{dP}(z) \le \frac{1}{\pi}\quad P\textnormal{-almost surely.}
$$
Equivalently, $Q_0 \in \mathcal{Q}(P, \pi)$.
Moreover, if $\pi < 1$, then there exists a probability law $R_{\mathrm res} \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$ such that
$$
P = \pi Q_0 + (1-\pi)R_{\mathrm res}.
$$
If $\pi = 1$, then $Q_0 = P$.
comment\begin{proof}[Proof of Proposition (ref)]
First, we show that $Q_0 \ll P$. Let $B \in \mathcal B$ satisfy $P(B)=0$. Then under Assumption (ref),
$$
0 = P(B) = \Pr \big( (X,Y(1)) \in B \mid S(1)=1 \big) \implies \Pr \big( (X,Y(1)) \in B, S(1)=1 \big) = 0.
$$
By Assumption (ref), $\{ S(0)=1 \} \subseteq \{S(1)=1\}$ almost surely. Thus,
$$
\Pr \big( (X,Y(1)) \in B, S(0)=1 \big) \le \Pr \big( (X,Y(1)) \in B, S(1)=1 \big) = 0.
$$
Dividing by $\Pr(S(0)=1) = p_0 > 0$ yields $$
\frac{\Pr\big((X,Y(1))\in B,\ S(0)=1 \big)}{\Pr(S(0)=1)}
=\Pr\big((X,Y(1))\in B\mid S(0)=1\big)
=Q_0(B) = 0.$$
Therefore, $Q_0$ is absolutely continuous with respect to $P$.
Second, we construct the Radon-Nikodym derivative and bound it. Because $(\mathcal{Z}, \mathcal{B})$ is standard Borel, regular conditional probabilities exist. Let $B \in \mathcal{B}$. By Assumption (ref), $\operatorname*{\mathbbm{1}} \{S(0)=1 \} = \operatorname*{\mathbbm{1}} \{S(0) = S(1)=1 \}$. Thus,
\[
\begin{aligned}
&\Pr \big( (X,Y(1)) \in B , S(0)=1 \big) \\
&= E\big[ \operatorname*{\mathbbm{1}}\{ (X,Y(1)) \in B \} \operatorname*{\mathbbm{1}}\{ S(0)=S(1)=1 \} \big] \\
&= E\big[ \operatorname*{\mathbbm{1}}\{(X,Y(1))\in B \} \operatorname*{\mathbbm{1}}\{ S(1)=1\} \; E[\operatorname*{\mathbbm{1}}\{S(0)=1\} \mid (X,Y(1),S(1)) ] \big].
\end{aligned}
\]
Hence,
$$
\Pr \big( (X,Y(1)) \in B , S(0)=1 \big) = E\big[ \operatorname*{\mathbbm{1}}\{(X,Y(1))\in B \} \operatorname*{\mathbbm{1}}\{ S(1)=1\} \; r(X,Y(1)) \big].
$$
where
$$
r(X,Y(1)) = E[\operatorname*{\mathbbm{1}}\{S(0)=1\} \mid X,Y(1),S(1)=1 ].
$$
Furthermore, conditioning on $S(1)=1$ yields
\begin{equation}
\Pr\big((X,Y(1))\in B,\ S(0)=1 \big)
=\Pr(S(1)=1) E \big[ \operatorname*{\mathbbm{1}} \{(X,Y(1))\in B\}\,r(X,Y(1))\mid S(1)=1 \big],
\end{equation}
Under $P = \mathcal{L} ( X,Y(1) \mid S(1)=1 ) $, we have
\begin{equation}
E \big[ \operatorname*{\mathbbm{1}} \{(X,Y(1))\in B\}\,r(X,Y(1))\mid S(1)=1 \big] = \int_B r(z) P(dz).
\end{equation}
Therefore, combining ((ref)) and ((ref)) yields
\[
\begin{aligned}
Q_0(B)
&= \Pr\big( (X,Y(1)) \in B \mid S(0)=1 \big) \\
&= \frac{1}{p_0}\Pr\big( (X,Y(1)) \in B, S(0) = 1 \big) \\
&= \frac{p_1}{p_0}\int_B r(z)\,P(dz) \\
&=\frac{1}{\pi}\int_B r(z)\,P(dz).
\end{aligned}
\]
This identifies a Radon-Nikodym derivative:
$$
\frac{dQ_0}{dP}(z)=\frac{r(z)}{\pi} \quad P\text{-almost surely.}
$$
Since $r(z) \in [0,1]$, we obtain
$$
0 \le \frac{dQ_0}{dP}(z) \le \frac{1}{\pi} \quad P\text{-almost surely.}
$$
This proves the likelihood domination and $Q_0 \in \mathcal{Q}(P,\pi)$.
Finally, we derive the mixture representation. If $\pi=1$, the bound $0 \le dQ_0/dP \le 1$ and $\int dQ_0/dP \, dP =1$ imply $dQ_0/dP =1$ $P$-almost surely. Thus, $Q_0=P$, and the mixture statement is trivial.
Suppose $\pi < 1$. For $B \in \mathcal{B}$, define the finite measure $\mu$ by
$$
\mu(B) = P(B) - \pi Q_0(B).
$$
Using $w = dQ_0/dP$,
$$
\mu(B) = \int_B (1 - \pi w(z)) P(dz).
$$
Since $w \le 1/\pi$, we have $1 - \pi w \ge 0$ $P$-almost surely. Thus, $\mu$ is nonnegative. Also,
$$
\mu(\mathcal{Z}) = P(\mathcal{Z}) - \pi Q_0(\mathcal{Z}) = 1 - \pi.
$$
Define
$$
R_{\mathrm res}(B) = \frac{\mu(B)}{1-\pi} = \frac{P(B) - \pi Q_0(B)}{1 - \pi}.
$$
Then, $R_{\mathrm res}$ is a probability measure and $P = \pi Q_0 + (1 - \pi)R_{\mathrm res}$.
\end{proof}
Sharpness of the Lee Ambiguity Set Relative to Reduced Observables
Proposition (ref) shows that monotone selection imposes a likelihood ratio restriction on the always-selected treated distribution. Equivalently, the target counterfactual law $Q_0$ must belong to $\mathcal{Q}(P, \pi)$. This result gives an outer bound on what the always-selected distribution can be, given the observable treated-selected law $P$ and the selection-rate ratio $\pi = p_0 / p_1$. However, $\mathcal{Q}(P,\pi)$ may still contain distributions that satisfy the likelihood ratio inequality but cannot be generated by any data-generating process consistent with monotone selection and the observed selection rates $(p_0,p_1)$.
The purpose of this subsection is to rule out that possibility. We show that the Lee ambiguity set is sharp relative to the reduced information $(P, p_0, p_1)$: every probability law $Q \in \mathcal{Q}(P, \pi)$ can be rationalized as the treated potential outcome distribution of always-selected units under some data-generating process satisfying Assumptions (ref)--(ref) and matching the same observables $(P, p_0, p_1)$. Thus, the likelihood ratio restriction from Proposition (ref) is not merely necessary but also sufficient.
The argument is constructive. Given any candidate $\widetilde{Q} \in \mathcal{Q}(P,\pi)$ and Borel measurable $B \in \mathcal{B}$, define the residual law
$$
R(B) = \frac{P(B) - \pi \widetilde{Q}(B)}{1-\pi},
$$
when $\pi < 1$. The bound $0 \le d\widetilde{Q} / dP \le 1 / \pi$ guarantees that $R$ is a valid probability measure. Hence, the observed treated-selected distribution $P$ can be decomposed as
$$
P = \pi \widetilde{Q} + (1 - \pi) R.
$$
This mixture representation has a direct principal-strata interpretation: $\widetilde{Q}$ is assigned to the always-selected stratum, while $R$ is assigned to the marginal-in stratum. Because the marginal-in distribution is unrestricted under Assumptions (ref)--(ref), the observable law cannot distinguish among different choices of $\widetilde{Q}$ inside $\mathcal{Q}(P, \pi)$.
Proposition (ref) formalizes this construction and establishes that $\mathcal{Q}(P, \pi)$ is exactly the identified region for the always-selected treated distribution. Consequently, any inference procedure that is uniformly valid under Assumptions (ref)--(ref) must be valid uniformly over all $\widetilde Q \in \mathcal{Q}(P, \pi)$. Conversely, no uniformly tighter distributional restriction is available without imposing additional assumptions.
commentWhile Proposition (ref) establishes that the true counterfactual distribution $Q$ reside within the identified set of probability laws $\mathcal{Q}(P,\pi)$, this alone provides necessary restrictions and does not guarantee that the class is as tight as possible. To ensure that our inferential procedure is not overly conservative, we must prove that the bounds implied by $\mathcal{Q}(P,\pi)$ are sharp. Proposition (ref) achieves this by demonstrating observational equivalence: every $Q \in \mathcal{Q}(P,\pi)$ is attainable under some DGP satisfying Assumptions (ref) and (ref) and matching the same observables $(P,p_0,p_1)$, i.e., an identified region exhausts all the information in the model under the maintained assumption. Thus, any inference that is uniformly valid under Assumptions (ref) and (ref) are robust to all $Q \in \mathcal{Q}(P,\pi)$.
prop[Sharpness of $\mathcal{Q}(P,\pi)$ relative to reduced observables]
Let $(\mathcal{Z}, \mathcal{B})$ be a standard Borel space. Let $P \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$, $0 < p_0 \le p_1 < 1$, and $\pi = p_0/p_1$. For any candidate law $\widetilde{Q} \in \mathcal{Q}(P,\pi)$, there exists a data-generating process for $(X, Y(1),Y(0),S(1),S(0),D)$ satisfying Assumptions (ref)--(ref) such that
$$
\Pr(S(0)=1)=p_0, \quad \Pr(S(1)=1)=p_1,
$$
$$
\mathcal{L}( X,Y \mid D=1, S=1) = P, \;\; \text{and} \;\; \mathcal{L} (X,Y(1) \mid S(0)=1) = \widetilde{Q}.
$$
Consequently, conditional on the reduced observables $(P,p_0,p_1)$, the sharp identification region for the always-selected treated law $Q_0$ equals $\mathcal{Q}(P,\pi)$.
comment\begin{proof}[Proof of Proposition (ref)]
The proof of sharpness relies on a constructive mixture argument that assigns the observed data to latent principal strata.
\paragraph{Case 1.} $\pi < 1$
Fix the observed treated-selected law $P$ and any valid candidate target $\widetilde{Q} \in \mathcal{Q}(P,\pi)$. For $B \in \mathcal{B}$, define a residual measure to represent the induced subpopulation: $$R(B) = \frac{P(B) - \pi \widetilde{Q}(B)}{1-\pi}.$$
We want to show that $R$ is a probability measure. Since $\widetilde{Q} \ll P$, let $w = d\tilde{Q}/dP$. Then $$P(B) - \pi \widetilde{Q}(B) = \int_B (1 -\pi w)dP.$$
Because $w \le 1/\pi$, the integrand $1- \pi w \ge 0$ $P$-almost surely. Thus $P - \pi \widetilde{Q}$ is nonnegative. Also,
$$
(P - \pi \widetilde{Q})(\mathcal{Z}) = 1 - \pi.
$$
Therefore, $R$ is a well-defined probability measure and
\begin{equation}
P = \pi \widetilde{Q} + (1-\pi)R.
\end{equation}
We can then construct a latent population comprised of three types $T \in \{\mathcal{A}, \mathcal{M}, \mathcal{N}\}$ representing “always-observed”, “marginal-ins” , and “never-observed”, and match $(p_0,p_1)$. By assigning the population shares, we match the selection probabilities: $$\Pr(T=\mathcal{A}) = p_0, \quad \Pr(T=\mathcal{M}) = p_1 - p_0, \quad \Pr(T=\mathcal{N}) = 1 - p_1.$$
Define the potential selection indicators
\[
(S(0),S(1)) =
\begin{cases}
(1,1), \quad T=\mathcal{A}, \\
(0,1), \quad T=\mathcal{M}, \\
(0,0), \quad T=\mathcal{N}.
\end{cases}
\]
Then $S(1) \ge S(0)$ almost surely and Assumption (ref) holds. Also,
\begin{equation}
\Pr(S(0)=1) = p_0, \quad \Pr(S(1)=1) = p_1.
\end{equation}
Next, we assign $(X,Y(1))$ conditional on principal stratum defined by three types. We introduce the following lemma.
\begin{lemma}[Randomization]
Let $(\mathcal{Z},\mathcal{B})$ be the standard Borel and $\nu$ be any probability measure on $(\mathcal{Z},\mathcal{B})$. Then, there exists a Borel measurable map $g_{\nu}:[0,1] \to \mathcal{Z}$ such that for $U \sim \textnormal{Unif}(0,1)$, $g_{\nu}(U) \sim \nu$.
\end{lemma}
Lemma (ref) allows us to generate random elements with prescribed laws using a single uniform seed. Let $U \sim \text{Unif}(0,1)$ be independent of $T$. By Lemma (ref), there exist measurable maps $g_{\widetilde{Q}}$ and $g_R$ such that $g_{\widetilde{Q}}(U) \sim \widetilde{Q}$ and $g_R(U) \sim R$. Define
\[
Z = (X,Y(1)) \sim
\begin{cases}
g_{\widetilde{Q}}(U), &T = \mathcal{A}, \\
g_R(U), &T = \mathcal{M} \\
\text{arbitrary}, &T = \mathcal{N},
\end{cases}
\]
and define $Y(0)$ arbitrary because it is irrelevant for the equalities to be matched here. Let $D \sim \text{Bernoulli}(1/2)$ be independent of $(T,U)$. Then $D \perp (X,Y(0),Y(1),S(0),S(1))$ and Assumption (ref) holds.
Finally, we verify the matched observables. By ((ref)) and Assumption (ref), $$\Pr(S=1 \mid D=1) = \Pr(S(1) = 1) = p_1, \quad \Pr(S=1 \mid D=0) = \Pr(S(0) = 1) = p_0.$$
Moreover, conditional on $\{ D=1, S=1 \}$, $Y = Y(1)$. Thus,
$$
\mathcal{L}\big( X,Y \mid D=1, S=1 \big) = \mathcal{L}( Z \mid S(1)=1 ).
$$
Under our construction, $S(1)=1$ if and only if $T \in \{ \mathcal{A}, \mathcal{M} \}$. Hence,
\[
\begin{aligned}
\mathcal{L}( Z \mid S(1)=1 )
&= \frac{\Pr(T=\mathcal{A})}{\Pr(S(1)=1)} \mathcal{L}( Z \mid T=\mathcal{A} ) + \frac{\Pr(T=\mathcal{M})}{\Pr(S(1)=1)} \mathcal{L}( Z \mid T=\mathcal{M} ) \\
&= \pi\widetilde{Q} + (1-\pi) R,
\end{aligned}
\]
which equals $P$ by ((ref)). Finally, $S(0)=1$ if and only if $T = \mathcal{A}$. Hence,
$$
\mathcal{L} \big( X,Y(1) \mid S(0)=1 \big) = \mathcal{L}( Z \mid T=\mathcal{A}) = \widetilde{Q}.
$$
This completes the construction and proves sharpness for $\pi < 1$.
\paragraph{Case 2.} $\pi=1$
If $\pi =1$, $\widetilde{Q} \in \mathcal{Q}(P,\pi)$ implies $\widetilde{Q}=P$. Writing $d\widetilde{Q}/dP=w$, we have $0 \le w \le 1$ and $\int w dP=1$. Thus, $w=1$ $P$-almost surely.
We construct the data-generating process as follows. Draw $D\sim\mathrm{Bernoulli}(\rho)$ independently of $(X,Y(1),Y(0),S(1),S(0),T),$ where $\rho\in(0,1)$ may be chosen to match the observed treatment assignment probability. Let $T \in \{\mathcal{A}, \mathcal{N}\}$ with $\Pr(T = \mathcal{A}) = p_1 (=p_0)$, and set
\[
(S(0),S(1)) =
\begin{cases}
(1,1), \quad T=\mathcal{A}, \\
(0,0), \quad T=\mathcal{N}.
\end{cases}
\]
Thus, Assumption (ref) holds. On $\{T = \mathcal{A}\}$, draw $Z = (X,Y(1)) \sim P$, and on $\{ T = \mathcal{N}\}$, define $Z$ arbitrarily. Let $Y(0)$ be arbitrary everywhere. Then $$\Pr(S(0)=1) = p_0, \quad \Pr(S(1)=1) = p_1, \quad \mathcal{L}( Z \mid S(1)=1 ) =P, \quad \mathcal{L}( Z \mid S(0)=1 ) =P = \widetilde{Q}.$$
Finally, $(Y(0), Y(1), S(0), S(1) , X) \perp \!\!\! \perp D$ by construction and therefore Assumption (ref) holds.
\end{proof}
The sharpness result in this subsection is also the identification-level counterpart of the sharpness discussion in jin2023sensitivity. Their analysis emphasizes that validity alone is not enough, since one can always construct arbitrarily conservative prediction sets; sharpness requires calibrating against the actual worst-case score distribution over the identified class of target laws. Here, Proposition (ref) establishes that $\mathcal{Q}(P, \pi)$ is not merely a convenient outer bound: every $Q \in \mathcal{Q}(P, \pi)$ is attainable under some data-generating process satisfying monotone selection and matching the reduced observables $(P, p_0, p_1)$. Consequently, the worst-case CDF over $\mathcal{Q}(P, \pi)$ is the relevant object for any procedure that uses only this reduced information. This parallels the robust-prediction sharpness criterion of jin2023sensitivity, with the Lee ambiguity set $\mathcal{Q}(P, \pi)$ replacing their generic likelihood ratio class.
A useful qualification is that this sharpness is relative to the reduced information $(P, p_0, p_1)$. Proposition (ref) establishes sharpness relative to the reduced information $(P, p_0, p_1)$. The full observed law contains additional information. In particular, because selected-controls are always-selected under monotonicity, the selected-control covariate distribution identifies the covariate marginal of the always-selected target population: $$\mathcal{L}(X \mid \mathcal{A}) = \mathcal{L}(X \mid D=0, S=1).$$ Therefore, the reduced-information class $\mathcal{Q}(P, \pi)$ may be conservative relative to the full observed data. This mirrors the distinction in jin2023sensitivity between sharpness for a general robust prediction problem and sharpness for counterfactual inference, where additional structure in the observed law can shrink the identification set.
Define $s_d(x) = \Pr(S(d) =1 \mid X=x)$ for $d \in \{0,1\}$ and $\pi(x) = s_0(x)/s_1(x)$. Under the random assignment, $s_d(x) = \Pr(S =1 \mid D=d, X=x)$ for $d \in \{0,1\}$. Furthermore, define $Q_X = \mathcal{L}(X \mid \mathcal{A})$ and $H_X = \mathcal{L}(X \mid D=0, S=1)$, where $Q_X$ is the marginal distribution of $X$ under $Q$ and $H_X$ is the observed selected-control covariate distribution. The full-data sharp class is obtained by intersecting $\mathcal{Q}(P, \pi)$ with this identified covariate marginal restriction: $$\mathcal{Q}_{\text{full}}= \{ Q \in \mathcal{Q}(P,\pi) : Q_X = H_X \}.$$
Equivalently, define conditional distributions $P_x = \mathcal{L}(Y \mid D=1, S=1, X=x)$, $Q_x = \mathcal{L}(Y(1) \mid X=x, \mathcal{A})$. Then, the full-data sharp class can be written as $$
\mathcal{Q}_{\text{full}}= \{ Q(dx,dy) = H_X(dx) Q_x(dy): Q_x \ll P_x, \; 0 \le \frac{dQ_x}{dP_x} \le \frac{1}{\pi(x)} \; H_X\text{-a.e.} \}.$$
Detailed discussion on the full observed law identification is provided in the \hyperref[sec:appendix]{Appendix}.
Distribution-free Counterfactual Intervals for Always-selected
The goal of this section is to construct prediction intervals for the unobserved treated potential outcome $Y(1)$ and marginal ITE among always-selected units. The central prediction problem is cross-population nonexchangeability because the available calibration sample is not drawn from the target law. Conditional on the fitted score, the calibration observations are exchangeable among themselves because they are drawn from $P$. The target observation, however, is drawn from an unknown $Q$. Thus, it is not generally
exchangeable with the calibration observations. Applying ordinary split-conformal prediction without modification would therefore control the score tail under \(P\), but not necessarily under the target law \(Q\).
The identification results above provide the structure needed to restore a distribution-free guarantee. Under monotone selection, the target law satisfies $Q \in \mathcal{Q}(P, \pi)$, where $\pi = p_0 / p_1$ is the share of always-selected units among treated-selected units. Consequently, for any measurable event $A$, $Q(A) \le P(A) / \pi$. This likelihood ratio domination converts the counterfactual coverage problem into a robust calibration problem.
The method does not restore exchangeability between $P$ and $Q$. Instead, it combines a distribution-shift transfer inequality with an ordinary exchangeable-rank argument under $P$. For the score-tail event
$$
A_t=\{(x,y):V(x,y)>t\},
$$
the likelihood ratio restriction implies
$$
Q(A_t)\le\frac1\pi P(A_t).
$$
Thus, target miscoverage is at most $\alpha$ whenever the corresponding
treated-selected tail probability is at most $\alpha\pi$. Therefore, to guarantee $Q(A_t) \le \alpha$ uniformly over all $Q \in \mathcal{Q}(P,\pi)$, it is sufficient to choose $t$ so that $P(A_t) \le \alpha \pi$. The conformalized Lee procedure estimates this $(1 - \alpha \pi)-$quantile of the score distribution under the observable treated-selected law $P$.
Split-Conformal Algorithm
We now describe the split-conformal implementation. Let $I_1 = \{i: D_i=1, S_i=1\}$ and $I_0 = \{i: D_i=0, S_i=1\}$. Split $I_1$ into a training fold $I_{\mathrm train}$ and a calibration fold $I_{\mathrm cal}$. The training fold is used to fit any predictive model for $Y$ given $X$, using only treated-selected observations. Given the trained model, let $V : \mathcal{X} \times \mathcal{Y} \rightarrow \mathbb{R}$ be the resulting nonconformity score. Let $\mathcal{H}$ be the sigma-field generated by the training fold $I_{\mathrm train}$, selection sample $\{ (D_i,S_i) \}_{i=1}^n$, and random split process. Conditional on $\mathcal H$, the calibration scores $V_i=V(X_i, Y_i)$, $i \in I_{\mathrm cal}$, are i.i.d. draws from the score distribution induced by $P$.
Let $m = |I_{\text{cal}}|$ and write order statistics $V_{(1)} \le \cdots \le V_{(m)}$ with $V_{(m+1)} = +\infty$. Define the split-conformal index $k = \lceil (m+1)(1-\alpha\pi) \rceil$ and the split-conformal threshold $\hat{t} = V_{(k)}$. The counterfactual prediction set for the treated potential outcome of an always-selected unit with covariates $x$ is $\widehat{C}_1(x) = \{y: V (x,y) \leq \hat{t} \}$. When $\pi=1$, the target law $Q$ coincides with the observed treated-selected law $P$, and the procedure reduces to ordinary split conformal prediction. When $\pi < 1$, the threshold is more conservative: the method calibrates to the $(1-\alpha \pi)$-quantile rather than the usual $(1-\alpha)$-quantile, reflecting the fact that the always-selected distribution may concentrate more heavily in the upper tail of the score distribution than the observed treated-selected mixture.
algorithm[algorithm omitted — 1,213 chars of source]
Conditional on the sigma-field $\mathcal{H}$, let $Z_i = \{(X_i, Y_i)\}_{i=1}^m$ be i.i.d. calibration observations from $P = \mathcal{L}((X,Y(1)) \mid S(1)=1)$, and $Z_{m+1} = (X_{m+1}, Y_{m+1})$ be an independent test point drawn from $Q$. Theorem (ref) shows that this modification delivers finite-sample, distribution-free marginal coverage uniformly over all counterfactual laws $Q \in \mathcal{Q}(P,\pi)$.
commentTo construct the Conformalized Lee Inference (CLI) procedure, we adopt a split-conformal framework. The sample is partitioned into a training fold and an independent calibration fold. The training fold is used to fit a predictive model for the outcome $Y$ conditional on covariates $X$, utilizing only the treated-selected observations $\{i : D_i = 1, S_i = 1\}$.
Let $\mathcal{I}_1 \coloneqq \{i: D_i=1, S_i=1\}$ and $\mathcal{I}_0 \coloneqq \{i: D_i=0, S_i=1\}$. Suppose we have a predictive model. Let $V : \mathcal{X} \times \mathbb{R} \to \mathbb{R}$ denote a generic nonconformity score function evaluated using the trained model. $V(x,y)$ measures how well $y$ conforms to the prediction of the model given $x$. Then we split $\mathcal{I}_1$ into train/calibration folds $\mathcal{I}_{\text{train}}$ and $\mathcal{I}_{\text{calib}}$ with $n = | \mathcal{I}_{\text{calib}} |$.
Define nonconformity scores $V_i = V(X_i,Y_i)$ for $i \in \mathcal{I}_{\text{calib}}$ and let $V_{(1)} \leq \cdots \leq V_{(n)}$ be the order statistic of $V_i$ with the convention $V_{(n+1)} \coloneqq +\infty$. Finally, define the threshold index:
\[
k \coloneqq \lceil(n+1)(1-\alpha\pi)\rceil,
\]
and set a data-dependent threshold $\hat{t}_n \coloneqq V_{(k)}$. The factor $\pi$ inflates the required coverage under the observed treated-selected mixture. Then, for a new always-selected observation with covariate $x$, a level $(1-\alpha)$ counterfactual prediction interval is given by
\[
\hat{C}_1(x) \coloneqq \{y: V(x,y) \leq \hat{t}_n \}.
\]
Propositions (ref) and (ref) identify the population minimax threshold. Furthermore, Theorem (ref) establishes the marginal coverage of this procedure for the unobserved counterfactuals of the principal stratum when the true $\pi = p_0/p_1$ is known. Under the oracle assumption for $\pi$, the theorem shows that the split-conformal order statistic with index $k = \lceil(n+1)(1-\alpha\pi)\rceil$ delivers the desired minimax guarantee in finite samples, distribution-free, conditioning on the training fold.
comment\begin{thm}[Finite-sample counterfactual coverage]
Let $(\mathcal{Z}, \mathcal{B})$ be a standard Borel space. Suppose that Assumptions (ref) and (ref) hold. Fix $\alpha \in (0,1)$ and $\pi \in (0,1]$. Let $P \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$ and $Q \in \mathcal{Q}(P, \pi)$. Conditional on the $\sigma$-field $\mathcal{T}$, let $Z_i = (X_i, Y_i)_{i=1}^m$ be i.i.d. calibration observations from $P = \mathcal{L}((X,Y(1)) \mid S(1)=1)$, and $Z_{m+1} = (X_{m+1}, Y_{m+1})$ be an independent test point drawn from $Q$. Then, the prediction set satisfies
\[
\inf_{Q\in\mathcal{Q}(P,\pi)} \Pr_{P^m \times Q}(Y_{m+1} \in \widehat{C}_1(X_{m+1}) \mid \mathcal{T}) \ge 1-\alpha.
\]
\end{thm}
comment\begin{proof}[Proof of Theorem (ref)]
Fix an arbitrary $Q \in \mathcal{Q}(P,\pi)$. Since $\mathcal{Z}$ is standard Borel, regular conditional laws given $\mathcal{T}$ exist. We work throughout under a version of the conditional probability measure $\Pr(\cdot \mid \mathcal{T})$ on a full-measure set on which $V$ is deterministic and measurable. Thus, conditional on $\mathcal{T}$, $(Z_1,...,Z_m)$ are i.i.d. with law $P$, $Z_{n+1}$ is independent of $(Z_1,...,Z_m)$ with law $Q$, and the threshold $\hat{t} = V_{(k)}$ is measurable with respect to $\mathcal{F}_m = \sigma(Z_1,...,Z_m)$. We want to show that
\[
\Pr_Q \{ Y_{m+1} \notin \hat{C}_1(X_{m+1}) \mid \mathcal{T} \} \le \alpha.
\]
Let $V_{m+1} = V(X_{m+1}, Y_{m+1})$. By the construction of the prediction set, the miscoverage event is equivalent to the test score exceeding the threshold: $$\{Y_{m+1} \notin \widehat{C}_1(X_{m+1})\} = \{V(X_{m+1},Y_{m+1}) > V_{(k)} \} = \{V_{m+1} > \hat{t} \}.$$
Therefore, $$\Pr_{Q}\{ Y_{n+1} \notin \widehat{C}_1(X_{n+1}) \mid \mathcal{T} \} = \Pr_{Q} (V_{n+1} > \hat{t} \mid \mathcal{T}).$$
For $t \in \overline{\mathbb{R}} = \mathbb{R} \cup \{+\infty\}$, define the upper tail event set $A_t = \{ z \in \mathcal{Z} : V(z) > t\} \in \mathcal{B}$. Since $\hat{t}$ is $\mathcal{F}_n-$ measurable and $Z_{m+1}$ is conditionally independent of given $\mathcal{T}$, $$\Pr_{Q}(V_{m+1} > \hat{t} \mid \mathcal{F}_m, \mathcal{T}) = Q(A_{\hat{t}}).$$
Taking conditional expectation with respect to $\mathcal{T}$ yields
\begin{equation}
\Pr_{Q}(V_{m+1} > \hat{t} \mid \mathcal{T}) = E[Q(A_{\hat{t}}) \mid \mathcal{T}].
\end{equation}
First, we establish a tail inequality via likelihood-ratio domination. Because $Q \in \mathcal{Q}(P,\pi)$, we have $Q \ll P$ and $w = dQ/dP \le 1/\pi$ $P-$almost surely. For any $B \in \mathcal{B}$, it follows that
\begin{equation}
Q(B) = \int 1_B dQ = \int_B w dP \le \frac{1}{\pi} P(B)
.
\end{equation}
Applying ((ref)) to the upper tail event $A_{\hat{t}_n}$, we obtain $$Q(A_{\hat{t}}) \le \frac{1}{\pi} P(A_{\hat{t}}) \quad \text{almost surely.}.$$
Substituting into ((ref)) yields
\begin{equation}
\Pr_{Q}(V_{m+1} > \hat{t} \mid \mathcal{T}) \leq \frac{1}{\pi} E[P(A_{\hat{t}}) \mid \mathcal{T}].
\end{equation}
Second, we represent the expected $P-$tail mass using an auxiliary $P-$distributed point. Let $Z_{m+1}^{\circ} \sim P$ be independent of $(Z_1,...,Z_m)$ conditional on $\mathcal{T}$ and define $V_{m+1}^{\circ} = V(Z_{m+1}^{\circ})$. Since $\hat{t}$ is $\mathcal{F}_m-$measurable, $$P(A_{\hat{t}}) = P\{ V(Z) > \hat{t} \} = \Pr( V_{m+1}^{\circ} > \hat{t} \mid \mathcal{F}_m, \mathcal{T} ).$$
Taking conditional expectations yields
\begin{equation}
E[P(A_{\hat{t}}) \mid \mathcal{T}] = \Pr( V_{m+1}^{\circ} > \hat{t} \mid \mathcal{T} ) = \Pr( V_{m+1}^{\circ} > V_{(k)} \mid \mathcal{T}).
\end{equation}
Combining ((ref)) and ((ref)), we have
\begin{equation}
\Pr_{Q}(V_{m+1} > \hat{t} \mid \mathcal{T}) \le \frac{1}{\pi} \Pr( V_{m+1}^{\circ} > V_{(k)} \mid \mathcal{T}).
\end{equation}
Third, we handle ties by introducing auxiliary variables $U_1,...,U_n,U_{m+1} \overset{i.i.d}{\sim} \text{Unif}(0,1),$ independent of everything else conditional on $\mathcal{T}$. Define perturbed scores $$W_i = (V_i,U_i) \quad i=1,...,m, \quad W_{m+1} \coloneqq (V_{m+1}^{\circ}, U_{m+1}),$$
and lexicographic order $$(v,u) <_{\text{lex}}(v',u') \iff [v < v'] \text{ or } [v = v' \text{ and } u < u'].$$
Because the second coordinates are continuously distributed, the random vectors $W_1,...,W_m,W_{m+1}$ are almost surely distinct under $<_{\text{lex}}$. Define the lexicographic rank
\[
\mathcal{R} = 1+ \sum_{i=1}^m \operatorname*{\mathbbm{1}} \{ W_i <_{\text{lex}} W_{m+1}\}.
\]
Then, conditional on $\mathcal{T}$, $W_i$'s are i.i.d. Therefore, by exchangeability, the rank is uniform:
\begin{equation}
\mathcal{R} \sim Unif \{1, ..., m+1\}.
\end{equation}
We now compare the event $\{ V_{m+1}^{\circ} > V_{(k)} \}$ to the rank $\mathcal{R}$. If $V_{m+1}^{\circ} > V_{(k)}$, then at least $k$ calibration scores satisfy $V_i \leq V_{(k)} < V_{m+1}^{\circ}$. For each such $i$, $W_i = (V_i,U_i) <_{\text{lex}} (V_{m+1}^\circ, U_{m+1}) = W_{m+1}$. Therefore, the event $\{ V_{m+1}^{\circ} > V_{(k)} \}$ implies that at least $k$ of the $W_i$'s are lexicographically smaller than $W_{m+1}$, i.e.,
\begin{equation}
\{ V_{m+1}^{\circ} > V_{(k)} \} \subseteq \{ \mathcal{R} \geq k+1\}.
\end{equation}
Using ((ref)) and ((ref)),
\begin{equation}
\Pr( V_{m+1}^{\circ} > V_{(k)} \mid \mathcal{T}) \leq \Pr( \mathcal{R} \geq k+1 \mid \mathcal{T} ) = \frac{m+1-k}{m+1}.
\end{equation}
Finally, we constrain this bounded probability to $\alpha$. By definition of $k$,
\begin{equation}
k = \lceil(m+1)(1-\alpha\pi)\rceil \iff \frac{m+1-k}{m+1} \le \alpha\pi.
\end{equation}
Combining ((ref)), ((ref)), and ((ref)) yields $$\Pr_{Q}(V_{m+1} > \hat{t} \mid \mathcal{T}) \le \frac{1}{\pi} \cdot \frac{m+1-k}{m+1} \le \frac{1}{\pi} \cdot \alpha\pi = \alpha.$$
Therefore, $$\Pr_Q \{ Y_{m+1} \in \widehat{C}_1(X_{m+1}) \mid \mathcal{T} \} = 1 - \Pr_Q (V_{m+1} > \hat{t} \mid \mathcal{T}) \ge 1 - \alpha.$$
Since $Q \in \mathcal{Q}(P,\pi)$ was arbitrary, $$\inf_{Q \in \mathcal{Q}(P,\pi)} \Pr_Q \{ Y_{m+1} \in \widehat{C}_1(X_{m+1}) \mid \mathcal{T} \} \ge 1-\alpha.$$
\end{proof}
thm[Finite sample counterfactual coverage]
Let \((\mathcal Z,\mathcal B)\) be a standard Borel space, $P \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$ and $Q \in \mathcal{Q}(P, \pi)$. Fix \(\alpha\in(0,1)\). Let \(\widehat\pi\in[0,1]\) be the estimated value used by the algorithm. Then, the prediction set satisfies
$$ \Pr_Q \big( Y_{m+1}\in \widehat C_1(X_{m+1}) \mid \mathcal{H} \big)
\ge 1-\alpha-\widehat\Delta,
$$
where the error term is
$$
\widehat\Delta
=
\left[
\frac{m+1-k}{(m+1)\pi}
-\alpha
\right]_+.
$$
Moreover,
$$
\widehat\Delta
\le
\alpha\left(\frac{\widehat\pi}{\pi}-1\right)_+
=
\frac{\alpha}{\pi}(\widehat\pi-\pi)_+.
$$
Theorem (ref) states that if $\widehat\pi\le \pi$, the error term is zero and the usual $1-\alpha$ coverage guarantee remains valid. If $\widehat\pi > \pi$, coverage may fall below \(1-\alpha\) by at most
$$
\alpha\left(\frac{\widehat\pi}{\pi}-1\right).
$$
Thus, the error term $\widehat \Delta$ is one-sided. If $\widehat\pi \le \pi $, then the algorithm behaves as if the always-selected share were smaller than it truly is. Equivalently, it uses
the looser likelihood-ratio bound $1/\widehat\pi \ge 1 / \pi$, so the resulting prediction set is conservative and $\widehat\Delta =0$. Coverage loss can occur only when $\widehat\pi > \pi$, because then the algorithm understates the worst-case amount by which the target law can upweight high-score events. \hyperref[sec:appendix]{Appendix} provides two practical strategies for estimating $\widehat\pi$: Clopper-Pearson construction and H\"{o}effding construction.
The coverage guarantee is monotone in the value of \(\widehat\pi\) used for
calibration. If \(\widehat\pi\le \pi\), then the procedure calibrates against the
larger ambiguity set \(\mathcal Q(P,\widehat\pi)\), since $\mathcal Q(P,\pi)\subseteq \mathcal Q(P,\widehat\pi)$. Thus, the true target law remains covered by the robustness class and the
usual \(1-\alpha\) guarantee is preserved. However, this validity comes at the
cost of wider prediction sets. As \(\widehat\pi\) decreases, the cutoff $V_{(k)} = V_{(\lceil(m+1)(1-\alpha\widehat\pi)\rceil)}$
moves to a higher treated-selected score quantile. In the limit
\(\widehat\pi=0\), the cutoff is \(V_{(m+1)}=+\infty\), yielding the trivial
prediction set with coverage one. Thus, an arbitrarily small \(\widehat\pi\) is
valid but not sharp. For informative inference, \(\widehat\pi\) should be chosen
as tightly as possible while remaining a valid lower bound for the true share
\(\pi\), for example by using a lower confidence bound.
Theorem (ref) proves coverage for a target law that is not the same as the calibration law. In ordinary split conformal prediction, the test point and calibration points are exchangeable because they are all drawn from the same law. In our case, however, the calibration sample comes from $P$, while the target counterfactual test draw comes from $Q \in \mathcal{Q}(P, \pi)$. Thus, the proof of Theorem (ref) introduces an auxiliary independent draw $Z_{m+1}^{\circ} \sim P$ as a proof device. This draw converts the random tail mass $P( V(Z) > t )$ into the exchangeable event $\{ V(Z_{m+1}^{\circ}) > t \}$. Furthermore, the proof augments nonconformity scores $(V_1, \ldots, V_m, V_{m+1})$ with i.i.d. $U_i \sim \text{Unif}(0,1)$ to randomize ties while preserving exchangeability (see, e.g., kuchibhotla2020exchangeability). Ties can occur with positive probability because the score distribution may have atoms. To solve the problem, the augmented scores $W_i = (V_i,U_i)$ are assumed to be ordered lexicographically to implement random tie-breaking. They preserve exchangeability and make the augmented scores almost surely distinct. Thus, the lexicographic rank is uniform.
Let $\Gamma = \mathcal{L}( X, Y(0), Y(1) \mid \mathcal{A})$ denote a joint law of potential outcomes and covariates in the always-selected population, and let $Q_{\Gamma} = \mathcal{L}(X,Y(1) \mid \mathcal{A})$ be its $(X,Y(1))$-marginal distribution. Under Assumptions (ref)--(ref), $Q_\Gamma \in \mathcal Q (P, \pi)$. For the prediction set $\widehat{C}_1$ from Theorem (ref), we define the shifted ITE prediction set.
definition[ITE prediction set]
For $(x,y_0) \in \mathcal X \times \mathcal Y$,
$$
\widehat{C}_{\tau}(x, y_0) = \widehat{C}_1(x) - y_0 = \{y - y_0 : y \in \widehat{C}_1(x)\}.
$$
Corollary (ref) converts marginal counterfactual coverage for $Y(1)$ into valid marginal predictive interval for the realized individual treatment effect of a randomly drawn selected control unit.
coro[Marginal ITE Prediction for Selected Controls]
Let $$(X_{m+1}, Y_{m+1}(0), Y_{m+1}(1)) \sim \Gamma$$ be an independent test unit. Define
$$
\tau_{m+1} = Y_{m+1}(1) - Y_{m+1}(0),
$$
Then,
$$
\Pr_{\Gamma} \big\{ \tau_{m+1} \in \widehat{C}_{\tau}(X_{m+1}, Y_{m+1}(0) ) \mid \mathcal H \big\}\ge 1 - \alpha - \widehat \Delta.
$$
Under Assumptions (ref) and (ref), $\Gamma$ is also the joint law of $(X, Y(0), Y(1))$ for a randomly drawn selected-control unit. Hence, the result gives marginal ITE coverage for an independent future selected controls.
comment\begin{remark}[ITE prediction when both potential outcomes are missing]
Corollary (ref) concerns the one-outcome-missing problem. For a selected-control unit, $Y(0)$ is observed and monotone selection implies membership in the always-selected stratum $\mathcal{A}$. The ITE prediction set can therefore be obtained by shifting the prediction set for the single missing potential outcome $Y(1)$. A distinct prediction problem arises for an independent new unit from the always-selected population for which $X$ is observed but both $Y(0)$ and $Y(1)$ are unobserved. In this case, prediction sets must be constructed for both potential outcomes before they can be combined into a prediction set for ITE.
Let
\[
(X_{m+1},Y_{m+1}(0),Y_{m+1}(1))
\sim
\Gamma
=
\mathcal L(X,Y(0),Y(1)\mid A)
\]
be an independent target unit. For nominal miscoverage levels
\(\alpha_1,\alpha_0\in(0,1)\), construct $\widehat C_{1,\widehat\pi}(x;\alpha_1)$ for $Y(1)$ using the conformalized Lee procedure and the treated-selected calibration sample. Under the conditions of Theorem (ref), this set satisfies
$$
\Pr_\Gamma
\left\{
Y_{m+1}(1)
\in
\widehat C_{1,\widehat\pi}(X_{m+1};\alpha_1)
\mid \mathcal H
\right\}
\ge
1-\alpha_1-\widehat\Delta_{\pi,m}.
$$
For $Y(0)$, use an ordinary split-conformal procedure based on selected-control observations to construct $\widehat C_0(x;\alpha_0)$. No Lee correction is required for this component because, under random assignment and monotone selection,
$$
\mathcal L(X,Y(0)\mid D=0,S=1)
=
\mathcal L(X,Y(0)\mid A).
$$
Hence, selected-control observations and the always-selected target unit have the same marginal law for $(X,Y(0))$, and the ordinary conformal guarantee is
$$
\Pr_\Gamma
\left\{
Y_{m+1}(0)
\in
\widehat C_0(X_{m+1};\alpha_0)
\mid \mathcal H
\right\}
\ge
1-\alpha_0,
$$
provided the corresponding split-conformal sampling conditions hold.
Define the both-outcomes-missing ITE prediction set by the Minkowski difference
\[
\widehat C_\tau^{\,\mathrm{both}}(x)
=
\widehat C_{1,\widehat\pi}(x;\alpha_1)
-
\widehat C_0(x;\alpha_0),
\]
where
\[
A-B=\{a-b:a\in A,\ b\in B\}.
\]
Whenever both component prediction events occur,
\[
Y_{m+1}(1)
\in
\widehat C_{1,\widehat\pi}(X_{m+1};\alpha_1)
\quad\text{and}\quad
Y_{m+1}(0)
\in
\widehat C_0(X_{m+1};\alpha_0),
\]
it follows that
\[
\tau_{m+1}
=
Y_{m+1}(1)-Y_{m+1}(0)
\in
\widehat C_\tau^{\,\mathrm{both}}(X_{m+1}).
\]
Therefore, the union bound gives
\[
\Pr_\Gamma
\left\{
\tau_{m+1}
\in
\widehat C_\tau^{\,\mathrm{both}}(X_{m+1})
\mid \mathcal H
\right\}
\ge
1-\alpha_1-\alpha_0-\widehat\Delta_{\pi,m}.
\]
This argument does not require independence between \(Y(0)\) and \(Y(1)\), nor
does it require identification of their joint dependence within the
always-selected stratum.
When \(\pi\) is known, or when the value used in calibration satisfies
\(\widehat\pi\le\pi\), the error term vanishes. Choosing
\[
\alpha_1+\alpha_0=\alpha,
\]
for example
\[
\alpha_1=\alpha_0=\frac{\alpha}{2},
\]
then yields
\[
\Pr_\Gamma
\left\{
\tau_{m+1}
\in
\widehat C_\tau^{\,\mathrm{both}}(X_{m+1})
\mid \mathcal H
\right\}
\ge
1-\alpha.
\]
If \(\widehat\pi=\pi_L\) is a lower confidence bound satisfying
\[
\Pr(\pi_L\le\pi)\ge 1-\delta_\pi,
\]
the corresponding unconditional guarantee is
\[
\Pr_\Gamma
\left\{
\tau_{m+1}
\in
\widehat C_\tau^{\,\mathrm{both}}(X_{m+1})
\right\}
\ge
1-\alpha_1-\alpha_0-\delta_\pi.
\]
Thus final coverage \(1-\alpha\) can be obtained by allocating the total error
budget so that
\[
\alpha_1+\alpha_0+\delta_\pi\le\alpha.
\]
If the two component sets are intervals,
\[
\widehat C_{1,\widehat\pi}(x;\alpha_1)
=
[L_1(x),U_1(x)]
\]
and
\[
\widehat C_0(x;\alpha_0)
=
[L_0(x),U_0(x)],
\]
then their Minkowski difference is the interval
\[
\widehat C_\tau^{\,\mathrm{both}}(x)
=
[L_1(x)-U_0(x),\,
U_1(x)-L_0(x)].
\]
This construction may be conservative because it combines the Lee-robust correction for \(Y(1)\) with a Bonferroni allocation across the two potential outcomes.
The target qualification is important. The above result concerns an independent new unit drawn from the always-selected population \(A\); it does not provide ITE coverage for an arbitrary external unit whose principal stratum is unknown. For a generic new unit with only \(X\) observed, membership in \(A\) is latent under the maintained assumptions. Extending the guarantee to the overall population, or to units not known to belong to \(A\), would require an additional identification argument or assumptions relating that target population to the always-selected law. As with Corollary 1, the coverage statement is marginal over the target distribution \(\Gamma\), rather than conditional coverage for a fixed individual with covariates \(X=x\).
\end{remark}
Optimality of the Split-Conformal Threshold
Theorem (ref) shows that the split-conformal threshold $\hat{t} = V_{(k)}$ delivers finite-sample coverage for every counterfactual law $Q \in \mathcal{Q}(P, \pi)$. This subsection explains why the same adjustment from the confidence level $1 - \alpha$ to $1 - \alpha \pi$ is not merely a convenient sufficient correction, but the sharp population target implied by the identification region.
To isolate the identification issue from sampling variability, fix a trained nonconformity score function $V$ and consider prediction sets of the form $C_t=\{z: V(z) \le t \}$. For such sets, the only remaining question is how large $t$ must be to guarantee $Q(V \le t) \ge 1 - \alpha$ uniformly over all admissible counterfactual laws $Q \in \mathcal{Q}(P, \pi)$. Since $Q$ is not identified, this is a minimax calibration problem: $$ \inf_{Q \in \mathcal{Q}(P,\pi)} Q(V \le t) \ge 1 - \alpha,$$or equivalently, $$ \sup_{Q \in \mathcal{Q}(P,\pi)} Q(V > t) \le \alpha.$$ The likelihood-ratio bound defining $\mathcal{Q}(P,\pi)$ implies that a counterfactual law can concentrate up to $1 / \pi$ times as much mass as $P$ on any score-tail event. Hence, for every threshold $t$, $$Q(V>t) \le \frac{1}{\pi} P(V>t).$$
Proposition (ref) shows that this upper bound is sharp over the reduced-information class $\mathcal{Q}(P, \pi)$: for every score-tail event $\{V>t \}$, there exists an admissible law $Q \in \mathcal{Q}(P,\pi)$ that attains the worst-case tail probability. Therefore, the minimax coverage constraint is equivalent to $$P(V>t) \le \alpha \pi.$$ This immediately identifies the smallest uniformly valid population threshold as the lower $(1 - \alpha \pi)-$quantile of the score distribution under the observable treated-selected law $P$. Relative to the full observed law, the procedure remains valid but may be conservative.
prop[Sharp tail bound over the Lee ambiguity set]
Let $(\mathcal{Z}, \mathcal{B})$ be a standard Borel space. Let $P \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$, $\pi \in (0,1]$. Then, for any measurable nonconformity score $V : \mathcal Z \to \mathbb{R}$ and every threshold $t \in \mathbb{R}$,
$$
\sup_{Q \in \mathcal{Q}(P,\pi)} Q(V > t) = \min \Big\{ 1, \frac{P(V > t)}{\pi} \Big\}.
$$ Moreover, for every $t$, there exists a maximizing law $Q_t \in \mathcal{Q}(P,\pi)$ such that
$$
Q_t(V>t) = \min \Big\{ 1, \frac{P(V>t)}{\pi} \Big\}.
$$ Therefore, the tail bound is sharp.
Let $F(t) = P(V \le t)$ and $F^{-1}(u) = \inf \{ t \in \overline{\mathbb{R}} : F(t) \ge u\}$, where $\overline{\mathbb{R}} = \mathbb{R} \cup \{+\infty\}$. We define the population minimax threshold over the Lee ambiguity set.
definition[Population minimax threshold]
\[
t^* = \inf \{ t \in \overline{\mathbb{R}} : \inf_{Q \in \mathcal{Q}(P,\pi)} Q(V \leq t) \geq 1-\alpha\}.
\]
Proposition (ref) translates Proposition (ref) into an explicit minimax target threshold. Thus, among all threshold rules that form prediction sets by accepting scores below a cutoff, $t^*$ is the least conservative threshold that is uniformly valid over the sharp identification region. Any smaller threshold fails for some admissible counterfactual law $Q \in \mathcal{Q}(P,\pi)$. In this sense, the split-conformal threshold $\hat{t}$ used in Theorem (ref) is a finite-sample, distribution-free estimate of the sharp population minimax threshold.
prop[Minimax population threshold]
Let $(\mathcal{Z}, \mathcal{B})$ be a standard Borel space. Let $P \in \mathcal{P}(\mathcal{Z}, \mathcal{B})$, $\pi \in (0,1]$. For any measurable nonconformity score $V : \mathcal{Z} \rightarrow \mathbb{R}$ and $\alpha \in (0,1)$, consider score-threshold prediction sets $$A_t = \{z \in \mathcal{Z} : V(z) \le t\}.$$ Then, $t^*$ is the lower $(1 - \alpha\pi)-$quantile of $V$ under $P$, i.e., $ t^* = F^{-1}( 1- \alpha \pi)$. Equivalently, $t^*$ is the smallest threshold satisfying $ P(V > t^*) \leq \alpha \pi$ and for every $t < t^*$, $P(V > t) > \alpha \pi.$
Furthermore, for every $t \in \mathbb{R}$, $$
\inf_{Q \in \mathcal{Q}(P,\pi)} Q(V \le t) = \max \Big\{ 0, 1 - \frac{P(V > t)}{\pi} \Big\},$$ and this lower bound is attained by some $Q_t \in \mathcal{Q}(P,\pi)$. Consequently, no threshold $t < t^*$ can guarantee $Q(V \le t) \ge 1-\alpha$ uniformly over $Q \in \mathcal{Q}(P, \pi)$.
This minimax interpretation connects the present construction to the robust conformal sensitivity-analysis framework of jin2023sensitivity. Their probably approximately correct (PAC) robust conformal procedure chooses a threshold through a conservative lower envelope for the worst-case distribution function of the nonconformity score, and their sharpness analysis studies when this envelope coincides with the true worst-case CDF over the relevant identification set. In the present monotone-selection problem, the ambiguity set is not generated by an analyst-chosen sensitivity parameter. Instead, monotonicity and the observed selection rates identify the likelihood-ratio restriction $0 \le dQ/dP \le 1/\pi,$ where $P$ is the treated-selected law, $Q$ is the always-selected target law, and $\pi$ is the selection rate. This yields the Lee-specific worst-case CDF:
$$
G_{\mathrm Lee}(t) = \inf_{Q \in \mathcal{Q}(P, \pi)} Q(V \le t) = \max \Big\{ 0, 1 - \frac{P(V>t)}{\pi} \Big\}.
$$
Thus, the conformalized Lee cutoff can be viewed as a closed-form robust conformal threshold obtained from the identified monotone-selection ambiguity set. Algorithm 2 of jin2023sensitivity constructs a PAC lower envelope $\hat{G}_n$ and threshold at $\inf\{t: \hat{G}_n(t) \ge 1 - \alpha \}$. However, the conformalized Lee procedure instead exploits the special scalar upper-tail structure to obtain the explicit split-conformal order statistic.
Simulation Studies
table[table omitted — 1,176 chars of source]
table[table omitted — 1,021 chars of source]
comment\begin{table}[t]
\caption{Prediction methods compared in the simulations}
\begin{threeparttable}
\begin{tabular}{p{0.20\linewidth}p{0.34\linewidth}p{0.35\linewidth}}
\toprule
Method & Value used in Lee cutoff & Interpretation \\
\midrule
Naive & $\eta_{\mathrm{used}}=1$ & Ordinary split conformal calibrated on treated-selected units \\
Oracle Lee & $\eta_{\mathrm{used}}=\eta$ & Infeasible benchmark using the true always-selected share \\
Plug-in Lee & $\eta_{\mathrm{used}}=\widehat\eta=(\widehat p_0/\widehat p_1)\wedge 1$ & Feasible direct estimate of the Lee share \\
CP-Lee & $\eta_{\mathrm{used}}=\widehat\eta_{\mathrm{CP},L}$ & Clopper--Pearson lower confidence bound for $\eta$ \\
H\"{o}effding-Lee & $\eta_{\mathrm{used}}=\widehat\eta_{\mathrm{H},L}$ & H\"{o}effding lower confidence bound for $\eta$ \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]
• Notes: All Lee methods use the order statistic $k=\lceil(m+1)(1-\alpha\eta_{\mathrm{used}})\rceil$. Under Option A, CP-Lee and H\"{o}effding-Lee are compared against the formal target $1-\alpha-\delta_\pi$.
\end{tablenotes}
\end{threeparttable}
\end{table}
table[table omitted — 2,020 chars of source]
table[table omitted — 1,781 chars of source]
table[table omitted — 1,137 chars of source]
comment\begin{table}[t]
\caption{Sharpness check under conditional-tail selection with oracle residual scores}
\begin{threeparttable}
\begin{tabular}{rrrrrrrr}
\toprule
$m$ & $\eta$ & Naive theory & Naive & Oracle Lee & Plug-in Lee & CP-Lee & H\"{o}effding-Lee \\
\midrule
100 & 0.75 & 0.867 & 0.871 & 0.908 & 0.908 & 0.925 & 0.934 \\
100 & 0.50 & 0.800 & 0.807 & 0.907 & 0.912 & 0.934 & 0.942 \\
100 & 0.25 & 0.600 & 0.632 & 0.919 & 0.917 & 0.950 & 0.972 \\
200 & 0.75 & 0.867 & 0.867 & 0.904 & 0.904 & 0.915 & 0.919 \\
200 & 0.50 & 0.800 & 0.804 & 0.903 & 0.903 & 0.917 & 0.925 \\
200 & 0.25 & 0.600 & 0.619 & 0.906 & 0.901 & 0.923 & 0.946 \\
\bottomrule
\end{tabular}
\begin{tablenotes}[flushleft]
• Notes: Results use $1-\alpha=0.90$ and the oracle residual score. The naive-theory column is $1-\min\{1,\alpha/\eta\}$. CP-Lee and H\"{o}effding-Lee are evaluated against the formal target $0.85$ under Option A, but their empirical coverage is displayed on the same scale for comparison.
\end{tablenotes}
\end{threeparttable}
\end{table}
table[table omitted — 1,641 chars of source]
We conduct a Monte Carlo study of conformalized Lee inference under the conditional score-tail selection design. The goal is to test the mechanism of Algorithm 1 in the setting where ordinary conformal prediction is expected to fail.
Table (ref) summarizes the Monte Carlo design. The target sample in each replication is drawn independently from the always-selected population and is used only to evaluate coverage and length. The outcome designs for $Y(d), d \in \{0,1\}$ is heteroskedastic and linear. We draw
$$
X\sim \mathrm{Unif}([0,1]^4),
$$
and generate
$$
U_1\mid X\sim N(0,\sigma^2(X)),\qquad
\sigma^2(X)=1+\frac12(2.5X_1)^2.
$$
The treated potential outcome is
$$
Y(1)=\beta^\top X+U_1,\qquad \beta=(-0.531,0.126,-0.312,0.018)^\top.
$$
For selected-control treatment-effect evaluation, we also generate
$$
Y(0) = \beta_0^\top X+0.5U_1+U_0,\qquad U_0\sim N(0,1), \qquad \beta_0=(0.2,-0.1,0.1,0.05)^\top.
$$
The shared component $U_1$ induces dependence between $Y(1)$ and $Y(0)$.
Table (ref) lists the principal-stratum assignment rules. In all designs, we first draw
$$
D \sim \mathrm{Bernoulli}(1/2),\qquad S(1) \sim \mathrm{Bernoulli}(0.8),
$$
then assign the always-selected indicator $A$ only among units with $S(1)=1$. We set
$S(0) = \operatorname*{\mathbbm{1}} \{A=1\}$ Thus, monotonicity $S(1)\ge S(0)$ holds by construction. The four DGPs differ only in how $A$ is assigned. The benign DGP draws $A$ independently of $(X,U_1,Y(1))$, so the always-selected target law coincides with the treated-selected calibration law. The conditional-tail DGP assigns always-selected units to the conditional upper tail of $|U_1|$. The unconditional-tail DGP assigns always-selected units to the unconditional upper tail of $|U_1|$, which also induces covariate shift because $\sigma(X)$ depends on $X_1$. The smooth DGP uses a logistic selection probability depending on both $X_1$ and the standardized absolute residual $|U_1|/\sigma(X)$, with $(b_X,b_Y)=(1,1.5)$.
We consider the five prediction methods for the selection share $\pi$. Naive conformal uses $\pi_{\mathrm{used}}=1$. Oracle Lee uses the true value $\pi_{\mathrm{used}}=\pi$. Plug-in Lee uses $\widehat \pi= \min\{1, \widehat p_0/ \widehat p_1\}$,
where $\widehat p_d$, $d \in \{0,1\}$ is the empirical selection rate. CP-Lee uses a Clopper--Pearson lower confidence bound $\hat \pi_{\mathrm CP}$, and H\"{o}effding-Lee uses a H\"{o}effding lower confidence bound $\widehat \pi_{H}$. For the lower-bound methods, the simulations use the same conformal miscoverage $\alpha$ as the other methods and set the confidence-budget parameter to $\delta_\pi=0.05$. Thus, CP-Lee and H\"{o}effding-Lee are interpreted under the formal target $1-\alpha-\delta_\pi.$ This is the same-$\alpha$ convention used in the reported tables and figures.
Finally, we consider three nonconformity scores. The oracle residual score is
$$
V(x,y)=|y-\beta^\top x|,
$$
where $\beta$ is the true parameter. This score isolates the Lee calibration logic from model-estimation error. The fitted residual score is
$$
V(x,y)=|y-\widehat\mu_1(x)|,
$$
where $\widehat\mu_1$ is fit on the treated-selected training fold. The conformalized quantile regression (CQR) score (romano2019conformalized) is
$$
V(x,y)=
\max\{\widehat q_{\alpha/2}(x)-y,\; y-\widehat q_{1-\alpha/2}(x)\},
$$
where the lower and upper conditional quantile functions are fit on the treated-selected training fold (see, meinshausen2006quantile). These three scores allow us to separate the ideal oracle calibration experiment from more practical implementations using estimated prediction functions.
Table (ref) gives the overall coverage summary across the full focused grid. Since this table averages over the five target coverage levels, the coverage column should be interpreted together with the reported gap column. Under benign selection, naive conformal is valid and its average coverage gap is $0.003$. This is expected because benign selection does not create a difference between the treated-selected calibration law and the always-selected target law. In the same benign design, Lee-adjusted methods over-cover. This overcoverage is the cost of guarding against target laws that are possible under the monotone-selection model but are not realized in the benign DGP.
The pattern is different under selection-induced distribution shift. In the conditional tail DGP, naive conformal has an average gap of $-0.273$. In the unconditional tail DGP, the average gap is $-0.284$. In the smooth DGP, the average gap is $-0.135$. Thus, ordinary conformal prediction fails whenever always-selected units are more likely to have large nonconformity scores than the treated-selected calibration sample suggests. The failure is strongest in the two tail designs and more moderate in the smooth design, which is less adversarial.
Oracle Lee and plug-in Lee largely correct this failure. In the conditional-tail DGP, oracle Lee and plug-in Lee both have average gaps around $0.014$. In the unconditional-tail DGP, their average gaps are around $0.011$ and $0.012$. In the smooth DGP, the average gaps are around $0.066$ and $0.067$. The systematic pattern is that the Lee adjustment restores coverage close to the target, whereas naive conformal remains substantially below target in the shifted designs.
As expected, the lower-bound methods are conservative. CP-Lee has positive average gaps in every DGP, and H\"{o}effding-Lee is even more conservative. This ordering reflects the construction of the lower confidence bounds. Smaller values of $\pi_{\mathrm{used}}$ move the cutoff farther into the upper tail of the treated-selected score distribution, increasing coverage and interval length. The H\"{o}effding bound is looser in these sample sizes, so H\"{o}effding-Lee tends to produce the largest coverage.
Table (ref) reports the same comparison at nominal target coverage $1-\alpha=0.90$. In the benign DGP, naive conformal has coverage $0.903$, while oracle and plug-in Lee cover around $0.953$ and $0.954$. In the shifted designs, naive conformal under-covers: its coverage is $0.771$ in the conditional tail DGP, $0.769$ in the unconditional tail DGP, and $0.821$ in the smooth DGP. Oracle Lee and plug-in Lee recover coverage close to $0.90$, with empirical coverage between $0.909$ and $0.919$ across the shifted designs. CP-Lee and H\"{o}effding-Lee are evaluated against the formal target $0.85$ under the same-$\alpha$ convention and cover well above this target.
Table (ref) summarizes selection-rate diagnostics at nominal coverage $0.90$. The plug-in estimator is nearly unbiased in these designs: for example, when $\pi=0.25$, the average $\widehat\pi$ is about $0.251$ for $m=100$ and $0.251$ for $m=200$. The CP and H\"{o}effding lower bounds are deliberately smaller. When $m=100$ and $\pi=0.25$, the average CP lower bound is $0.181$, while the average H\"{o}effding lower bound is $0.120$. This explains why H\"{o}effding-Lee is the most conservative method. It also explains the only notable infinite-interval issue: for $m=100$ and $\pi=0.25$, H\"{o}effding-Lee has $\Pr(k=m+1)=0.242$ on average. These infinite intervals are mathematically valid but should be interpreted as uninformative. For this reason, the tables report finite average length separately from the probability of $k=m+1$.
commentTable (ref) gives a sharpness check under the conditional-tail DGP with oracle residual scores. This is the cleanest setting because the selection rule is constructed directly from the oracle residual score. For naive conformal, the theoretical benchmark is
\[
1-\min\{1,\alpha/\eta\}.
\]
At nominal coverage $0.90$, the predicted naive coverage is $0.867$ for $\eta=0.75$, $0.800$ for $\eta=0.50$, and $0.600$ for $\eta=0.25$. The simulated naive coverage closely follows these values. For example, with $m=200$, the naive coverages are $0.867$, $0.804$, and $0.619$. By contrast, oracle Lee and plug-in Lee remain close to $0.90$. This confirms the sharp tail-inflation mechanism motivating the $(1-\alpha\eta)$-quantile correction.
Table (ref) summarizes score robustness at nominal coverage $0.90$, averaging over the three shifted DGPs and excluding the benign no-shift design. Naive conformal under-covers for all three scores, with coverage around $0.782$, $0.783$, and $0.794$ for oracle residual, fitted residual, and CQR scores, respectively. Oracle Lee and plug-in Lee cover around $0.91$ across all three score classes. CP-Lee and H\"{o}effding-Lee are conservative relative to their formal target. These results indicate that the Lee correction is not an artifact of oracle residual scoring; it remains effective with fitted residual and CQR scores.
The four figures for coverage-curve consolidate the graphical evidence by DGP. Figure (ref) shows the no-shift benchmark. In the naive column, all rows lie essentially on the diagonal for all values of $\pi$ and for both calibration sizes. This is the expected result because benign selection assigns always-selected status independently of $(X,U_1,Y(1))$ among treated-selected units, so the treated-selected calibration law and always-selected target law coincide. The Lee-adjusted columns lie above the diagonal, with larger overcoverage as $\pi$ decreases. This overcoverage is not a failure of the method; it is the price of robustness when the Lee ambiguity set is larger than the realized target shift.
Figure (ref) gives the sharpest visual test of the Lee correction. In the naive column, coverage is far below the diagonal, and the shortfall grows as $\pi$ decreases. The deterioration is especially severe for $\pi=0.25$, where always-selected units are concentrated in the highest-score part of the treated-selected distribution. In contrast, plug-in Lee is close to the diagonal across oracle residual, fitted residual, and CQR scores. CP-Lee and H\"{o}effding-Lee are conservative, with H\"{o}effding-Lee generally lying highest because the H\"{o}effding lower bound for $\pi$ is the most conservative. This figure provides the clearest evidence that the usual conformal cutoff is calibrated to the wrong law, while the Lee-adjusted cutoff corrects the tail inflation.
Figure (ref) shows the same conclusion under joint covariate and outcome shift. The naive column again exhibits substantial undercoverage, particularly for smaller $\pi$. This is consistent with the construction of the unconditional-tail DGP: selection on the unconditional tail of $|U_1|$ shifts the residual distribution and, because the residual variance depends on $X_1$, also shifts the covariate distribution. Plug-in Lee restores coverage close to the target, while CP-Lee and H\"{o}effding-Lee remain conservative. The similarity between the conditional-tail and unconditional-tail figures indicates that the Lee correction handles shifts in the joint law of $(X,Y(1))$, not only shifts in marginal covariates.
Figure (ref) reports the smooth heterogeneous-selection design. Relative to the two tail DGPs, naive conformal under-covers less severely, but the undercoverage remains systematic and becomes larger as $\pi$ decreases. This pattern is expected because smooth selection is probabilistic rather than a hard tail rule. Plug-in Lee remains close to the target line, and the lower-bound methods over-cover. The results are similar across the oracle residual, fitted residual, and CQR rows, suggesting that the Lee adjustment is not tied to a particular score construction.
Across the four figures, the simulation results are internally coherent. Naive conformal is valid in the benign DGP, where $P=Q$, but fails under conditional-tail, unconditional-tail, and smooth selection, where the always-selected target law differs from the treated-selected calibration law. Plug-in Lee performs close to the oracle-Lee benchmark reported in the tables, while CP-Lee and H\"{o}effding-Lee provide conservative lower-bound implementations. The solid and dashed curves are close in most panels, indicating that increasing the target calibration size from $m=100$ to $m=200$ mainly improves finite-sample stability rather than changing the qualitative conclusions. Overall, the figures and tables support the same message: conformalized Lee inference restores coverage under monotone selection by calibrating to the higher treated-selected quantile implied by the identified always-selected share.
Conclusion
This paper develops conformalized Lee inference for distribution-free predictive inference under monotone sample selection. The starting point is the observation that, when treatment affects selection, the treated-selected sample is generally not drawn from the same population as the selected controls. Under Lee monotonicity, selected controls are always-selected units, while treated-selected observations are a mixture of always-selected and marginal-in units. This mixture structure identifies the share of always-selected units among treated-selected units, $\pi=p_0/p_1$, but not the identity of those units. As a result, the counterfactual treated-outcome law for the always-selected population is only partially identified.
The paper shows that this partial identification problem has a simple distributional form. The always-selected treated law $Q_0$ must be absolutely continuous with respect to the observed treated-selected law $P$, with likelihood ratio bounded above by $1/\pi$. This Lee ambiguity set is shown to be sharp relative to the reduced observables: every law satisfying the likelihood-ratio bound can arise from a data-generating process satisfying random assignment, monotone selection, and the observed selection rates. Thus, the ambiguity set is not an artifact of proof technique, but the exact distributional uncertainty generated by the monotone-selection model.
Building on this identification result, the paper proposes a split-conformal procedure that replaces the usual $(1-\alpha)$ calibration quantile with the higher $(1-\alpha\pi)$ treated-selected score quantile. This adjustment accounts for the fact that the always-selected target law may concentrate up to $1/\pi$ times more probability on high-score events than the treated-selected calibration law. The resulting prediction sets have finite-sample marginal coverage uniformly over all target laws in the Lee ambiguity set. For selected-control units, the observed untreated outcome can be subtracted from the counterfactual treated-outcome interval, yielding distribution-free prediction intervals for realized individual treatment effects in the always-selected population.
The paper also establishes a minimax interpretation of the conformalized Lee cutoff. The worst-case law in the Lee ambiguity set places as much mass as allowed on the upper tail of the nonconformity score. Consequently, the $(1-\alpha\pi)$ population quantile is the smallest score threshold that guarantees uniform coverage over the sharp identification region. Any smaller threshold fails for some target law that is observationally indistinguishable under the maintained monotone-selection assumptions. This result clarifies that the Lee adjustment is not merely conservative; it is the sharp robust calibration rule implied by the identified selection structure.
The simulation evidence supports the theoretical results. Ordinary conformal prediction performs well when selection is benign and the treated-selected and always-selected laws coincide, but it under-covers substantially when always-selected units are more likely to have large nonconformity scores. The Lee-adjusted procedures restore coverage in conditional-tail, unconditional-tail, and smooth selection designs, including designs that induce both outcome-tail shift and covariate shift. The lower-bound implementations based on Clopper--Pearson and Hoeffding confidence bounds are conservative, as expected, and illustrate the practical tradeoff between guaranteed coverage and interval informativeness.
Several directions for future research follow naturally. First, the main procedure uses the reduced information $(P,p_0,p_1)$, but the full observed law contains additional structure through the selected-control covariate distribution and the conditional selection rates $s_0(x)$ and $s_1(x)$. A promising next step is to develop fully covariate-adaptive conformalized Lee methods based on the conditional ambiguity set
\[
Q_x \ll P_x, \qquad 0 \leq \frac{dQ_x}{dP_x}(y) \leq \frac{1}{\pi(x)},
\qquad \pi(x)=\frac{s_0(x)}{s_1(x)}.
\]
Such methods could exploit the identified covariate marginal of the always-selected population and may deliver shorter intervals than the reduced-information procedure while preserving distribution-free validity.
Second, future work should study the statistical problem of estimating local selection shares $\pi(x)$ in high-dimensional settings. Flexible machine-learning estimators, cross-fitting, and one-sided confidence bands for selection probabilities could make the full-law version practically useful. The main challenge is to retain finite-sample or high-probability coverage while avoiding excessive conservativeness from weak lower bounds on $\pi(x)$.
Third, the framework can be extended beyond binary treatment and monotone sample selection. Continuous or multi-valued treatments, dynamic treatment regimes, attrition over time, and competing selection mechanisms all generate related principal-stratum prediction problems. Extending the Lee ambiguity set logic to these settings would broaden the scope of distribution-free causal prediction under post-treatment selection.
Fourth, there is room to connect conformalized Lee inference with sensitivity analysis for violations of monotonicity. The present paper treats Lee monotonicity as the maintained identifying restriction. In applications, however, monotonicity may be credible only approximately. A useful extension would allow a controlled amount of defier mass or bidirectional selection and derive the corresponding robust conformal cutoff. This would place the method between sharp Lee-style identification and analyst-chosen sensitivity models.
Finally, empirical implementation raises important questions about diagnostics and reporting. Because the Lee adjustment becomes more conservative as $\pi$ decreases, applied work should report selection rates, estimated always-selected shares, probabilities of infinite or uninformative intervals, and sensitivity of interval length to the choice of lower confidence bound for $\pi$. Future empirical applications can clarify when conformalized Lee intervals are informative in realistic sample-selection settings and how their conclusions differ from average-effect Lee bounds.
Overall, the paper shows that Lee's monotone-selection logic can be used not only to bound average treatment effects, but also to construct finite-sample, distribution-free prediction intervals for counterfactual outcomes and individual treatment effects. The broader research agenda is to make this form of robust causal prediction more adaptive, more informative, and applicable to richer forms of selection while preserving the central advantage of the present approach: coverage guarantees driven by identified features of the selection problem rather than by unverifiable model specification.
figure[figure omitted — 480 chars of source]
figure[figure omitted — 389 chars of source]
figure[figure omitted — 379 chars of source]
figure[figure omitted — 336 chars of source]