Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
81,731 characters · 25 sections · 131 citation commands
Generalized Lee Bounds
\linespread{1.5}
Randomized controlled trials are often complicated by endogenous sample selection and non-response. This problem occurs when treatment affects the researcher's ability to observe an outcome (a selection effect) in addition to the outcome itself (the causal effect of interest). For example, being randomized into a job training program affects both an individual's wage and employment status. Since wages exist only for employed individuals, treatment-control wage difference is contaminated by selection bias. A common way to proceed is to bound the average causal effect from above and below, focusing on subjects whose outcomes are observed regardless of treatment receipt (the always-observed principal strata, FrangakisRubin or the always-takers, LeeBound).
Seminal work by LeeBound proposes nonparametric bounds assuming the selection effect is non-negative for all subjects (monotonicity). For example, if JobCorps cannot deter employment, basic Lee lower bound is the treatment-control difference in wages, where the top wages in the treated group are trimmed until treated and control employment rates are equal. Furthermore, LeeBound shows that the covariate density-weighted conditional Lee bound is weakly tighter than the basic bound that does not involve any covariates. However, only a handful of discrete covariates can be utilized to tighten the bound, since each covariate cell is required to have a positive number of treated and control outcomes.
This paper generalizes Lee bounds under conditional monotonicity, which allows the sign of selection effect to be determined by pre-treatment covariates. Now, baseline covariates now have two roles: (1) to define the subspaces of positive and negative selection response and (2) to tighten the bound on each subspace. For the first one, using the full covariate vector makes the conditional monotonicity assumption the least restrictive. For the second one, any covariate subvector -- even an empty one -- would suffice to define a valid bound. This paper studies the sharp -- the tightest possible -- generalized bound, where all covariates are used for (1) and (2).
I represent the sharp bound as a ratio of two semiparametric moments whose nuisance functions are the conditional probability of selection, the conditional quantile, and the propensity score (i.e., the conditional probability of treatment). If the nonparametric functions are smooth functions of covariates, they can be estimated by logistic series regression of HIR2003 and quantile series of CherChetQuant, respectively. Alternatively, if these functions have a sparse representation with respect to some transformations of covariates, one could employ their $\ell_1$-penalized analogs proposed in orthogStructural and belloni:11, Program. To make the second stage moments insensitive to the first-stage estimation error, I derive Neyman-orthogonal (Neyman:1959, Neyman:1979, AiChen2003, Newey1994, chernozhukov2016double, LRSP) moment functions for the numerator and the denominator. Combining Neyman-orthogonality and sample splitting, I derive a root-$N$ consistent and asymptotically normal estimator of the sharp bounds in the presence of (1) data-driven covariate selection, (2) possible misclassification into the regions of positive ($\mathcal{X}_{\text{help}}$) and negative ($\mathcal{X}_{\text{hurt}}$) selection response and (3) (any) point mass on the region where the selection response is exactly zero (i.e., the boundary between $\mathcal{X}_{\text{help}}$ and $\mathcal{X}_{\text{hurt}}$). This proposal neither requires the covariates to be discrete nor the propensity score to be known, and considerably expands the scope for using Lee bounds in practice.
As an empirical application, I revisit Lee's JobCorps study as in LeeBound using the data from Schochet. The paper's first major empirical finding is to that the unconditional monotonicity of selection (i.e., employment) is unlikely to hold. After imposing conditional monotonicity (and accounting for the differential JobCorps effect on employment), I find that the average JobCorps effect on the always-takers' week 90 wages ranges between $-5\%$ and $1\%$. Finally, I provide evidence of mean reversion of the expected log wage for the always-takers in the control status. This mean reversion corroborates Ashenfelter pattern and shows that earnings would have recovered even without JobCorps training. Therefore, evaluating JobCorps would have been very difficult without a randomized experiment, as one would need to explicitly model mean reversion in the potential wage in the control status.
Appendix B extends Lee's trimming approach to the case of multiple outcomes. A naive approach to construct an identified set is to take a Cartesian product of scalar bounds for each component of the causal parameter. However, since outcomes may be correlated, such a set may contain points that do not correspond to a data generating process. I characterize the sharp identified set for the causal parameter as well as its support function (BM,BMM). Furthermore, I establish debiased uniform over the boundary inference on the support function based on first-stage regularized estimators. The use of this theory is demonstrated for the parameters involving multi-dimensional outcomes, such as standardized treatment effect and wage growth.
\paragraph{Literature review. } The conditional monotonicity assumption has been previously proposed in Kolesar and sloczynski2020not in a treatment choice context, accommodating differential sign of a binary treatment response to a binary instrument. To generalize the Local Average Treatment Effect parameter, sloczynski2020not combines the estimates from no-defier and no-complier regions with signs $1$ and $-1$, respectively, while the boundary (i.e., the no-defier-and-no-complier region) has zero identification power and, therefore, receives zero weight. In contrast, the sample selection problem focuses on the always takers -- a different principal strata -- whose selection behavior does not change in response to treatment.
This paper combines ideas from various branches of economics and statistics, including bounds on causal effects (Manski89, Manski90, HorowitzManski, HorowitzManski2000, FrangakisRubin, angrist:2002, ZhangRubin, angrist:2006, ChenFlores, feller2016weak, APW2013, APW2018, Honore, kamat2021identifying), convex analysis and support function (CherRigStoker, Stoye, BM, BMM, KaidoSantos, Kaido, StoyeSpread, KaidoMolinariStoye, Gafarov, MolinariStoye, MolinariHandbook), monotonicity and latent index models (Vytlacil, KWLATE, kamat2019identifying, sloczynski2020not, MTW, MTW2, Ura), including the bounds on the same empirical context -- JobCorps job training program -- (LeeBound, FloresLagunes, ChenFlores).
Next, this paper contributes to a large body of work on debiased/orthogonal inference for parameters following regularization or model selection (Neyman:1959, Neyman:1979, HardleStoker1989, NeweyStoker, andrews:1994, Newey1994, Robins, robinson:88, ChenAck, ZhangZhang, JM, chernozhukov2016double, LRSP, Program, sasaki2018estimation, sasaki2020unconditional, Sasaki, chiang2019multiway, ning2020doubly, chernozhukov2021debiased, chernozhukov2021automatic, CherSem, NSS, singh2020debiased, Colangelo, Lieli, ZimLech). In many classic cases, such as Robins or robinson:88, orthogonalization expands the set of first-stage parameters to be estimated. In contrast, the set of first-stage nuisance for the truncated conditional mean functional does not expand after orthogonalization. Finally, the paper contributes to a growing literature on machine learning for bounds and partially identified models (kallus2019assessing, jeong2020robust, Bonvini_2021, ZhouSmith, SemJoE). The causal parameter is not a special case of a set-identified linear model of BM,BMM, and the identification and estimation approaches of CCMS and SemJoE do not apply.
An emerging body of research has validated the usefulness of this paper's results by both expanding theoretical framework and/or employing them in applications. For instance, olma2021nonparametric proposes a nonparametric estimator of truncated conditional expectation functions by plugging an orthogonal moment for the truncated mean into a locally linear regression. While interesting on its own, this parameter also enters the correction term for the unknown propensity score in Section (ref). Finally, Heiler2 proposes an estimator of heterogeneous treatment effects using least squares regression.
The paper is organized as follows. Section (ref) reviews basic Lee bounds and Lee's estimator under the standard unconditional monotonicity assumption. Section (ref) generalizes Lee bounds under conditional monotonicity. Section (ref) defines the debiased moment functions for the numerator and the denominator of each boun. Section (ref) overviews the proposed estimator and provides inference results for this parameter. Section (ref) states the asymptotic theory for generalized Lee bounds. Section (ref) presents empirical application. The Supplementary Appendix contains proofs (Appendix A), an extension to the multiple outcome case (Appendix B), and auxiliary empirical details (Appendix C).
Consider the sample selection model in LeeBound. Let $D=1$ be an indicator for treatment receipt. Let $Y(1)$ and $Y(0)$ denote the potential outcomes if an individual is treated or not, respectively. Likewise, let $S(1)=1$ and $S(0)=1$ be dummies for whether an individual's outcome is observed with and without treatment. The data vector $W=(D,X,S,S \cdot Y)$ consists of the treatment status $D$, a baseline covariate vector $X$, the selection status $S=D \cdot S(1) + (1-D) \cdot S(0)$ and the outcome $S \cdot Y = S \cdot (D \cdot Y(1) + (1-D) \cdot Y(0))$ for selected individuals. LeeBound focuses on the average treatment effect (ATE)
for subjects who are selected into the sample regardless of treatment receipt---the always-takers.
The independence assumption holds by random assignment. In addition, it requires all subjects to have the same probability of being treated. The monotonicity requires all subjects to have the same direction of selection response. In particular, a subject that is selected into the sample when untreated must remain selected if treated: $$ S(0) = 1 \quad \Rightarrow \quad S(1) =1. $$ As a result, $$\mathbb{E} [ Y(0) \mid S(1)=1, S(0)=1] = \mathbb{E} [ Y(0) \mid S(0)=1]. $$ By complete independence, $$ \mathbb{E} [ Y(0) \mid S(0)=1] = \mathbb{E} [ Y \mid S=1, D=0], $$ and $\mathbb{E} [ Y(0) \mid S(1)=1, S(0)=1] $ is point-identified.
In contrast to the control group, a treated outcome can be either an always-taker's outcome or a complier's outcome. The always-takers' share among the treated outcomes is
In the best case, the always-takers comprise the top $p_0$-quantile of the treated outcomes. The largest possible value of $\beta_0$ is basic upper bound
where $Q^1 (u)$ is the $u$-quantile of $Y \mid D=1,S=1$ and $p_0$ in (ref) is the trimming threshold. Likewise, the smallest possible one is $$ \bar{\beta}_L = \mathbb{E} [ Y \mid Y \leq Q^1 (p_0), D = 1, S=1] - \mathbb{E} [ Y \mid D = 0, S=1]. $$
Lee's identification strategy can be implemented conditional on covariates. Denote the conditional trimming threshold $p(x)$ as
and the conditional upper bound $\bar{\beta}_U (x) $ as
where $Q^{1}(u,x)$ is the conditional $u$-quantile of $Y$ in $S=1,D=1,X=x$ group. The covariate-based bound is
which, as Lee has shown, is weakly tighter than the basic one (LeeBound2).
The formula for (ref) involves a covariate density function. When data has a single continuous covariate, accounting for the estimation error of the nonparametric density estimate may be challenging both in theory and practice. An alternative approach, proposed in Algorithm (ref), is to discretize covariates. In absence of a better name, it is referred to as discrete Lee bounds and summarizes the discretization procedure Lee used to report covariate-based bound (Table 5, LeeBound).
This Section describes an alternative -- moment-based -- approach to bounds, which no longer requires estimating the always-takers' covariate density function.
Assumption (ref) (a) requires complete independence. Equivalently, the propensity score
must be constant in $X$. Assumption (ref) relaxes the complete independence to its conditional analog.
I represent (ref) as a ratio of two moments. Let $W=(D, X, S, S\cdot Y)$ be the data vector. Define the numerator moment function
where the true value of the first-stage nuisance function $\xi=\xi(x)$ is
Lemma (ref) represents the sharp bound as a ratio of two moments. To get more insight into the result, note that
Noting that $p(X) s(1,X)=(s(0,X)/s(1,X)) s(1,X)$ simplifies to $s(0,X)$ gives
Invoking Bayes rule gives
Applying Law of Iterated Expectations to (ref) gives the representation (ref). This representation makes it possible to employ continuous covariates under various semiparametric assumptions, discussed in Section (ref).
Remark (ref) points out an interesting connection between the selection problem, studied here, and the treatment choice.
In this Section, I generalize Lee bounds under conditional monotonicity. Abusing notation, I shall denote the generalized upper bound by $\beta_U$. The conditional probability of selection is
The conditional average treatment effect (ATE) on selection is
The sets of positive and negative selection response are
and the boundary is
Let $Q^{d}(u,x)$ is the conditional $u$-quantile of $Y$ in $S=1,D=d, X=x$. Given a quantile level $u \in (0,1)$, define the upper-trimmed mean for the treated group
and the lower-trimmed mean for the control group
Assumption (ref) generalizes the regular unconditional monotonicity assumption to the conditional analog. It allows the direction of the selection effect to vary across covariate values. Defiers $(S(1) = 0, S(0) = 1)$ are ruled out on the covariate set $\mathcal{X}_{\text{help}}$, compliers $(S(1) = 1, S(0) = 0)$ are ruled out on $ \mathcal{X}_{\text{hurt}} $, and both compliers and defiers are ruled out on the set $ \mathcal{X}_0$. In many practical cases, the boundary set $\mathcal{X}_0 $ has zero mass, but the proposed asymptotic theory accommodates the boundary of arbitrary size.
Let us show that the always-takers' share is point-identified under Assumptions (ref) and (ref). For any covariate value $x$ on the positive half-space $\mathcal{X}_{\text{help}}$, we have
Likewise, for any covariate value $x \in \mathcal{X}_{\text{hurt}}$,
On the boundary $\mathcal{X}_0$, $\tau (x) = 0 \Rightarrow s (1,x) = s (0,x)$. Combining the results
and aggregating over the covariate space gives the always-takers' share
\qed
Next, let me generalize the bound itself. For $x \in \mathcal{X}_{\text{help}}$, the conditional always-takers' share is $$ \min (s(0,x), s(1,x)) = s(0,x), \quad x \in \mathcal{X}_{\text{help}}.$$ The conditional upper bound takes the form
where $\beta^{\text{basic}}_U(x)$ is given in (ref). Likewise, for $x \in \mathcal{X}_{\text{hurt}}$, the roles of the treated and control group are reversed:
By Assumption (ref), both defiers and compliers are ruled out on the boundary $\mathcal{X}_0$. Therefore, a selected individual with $S=1$ must be an always-taker, and there is no trimming. The bound reduces to treatment-control difference
Aggregating the conditional sharp bound $\beta_U(x)$ gives
Invoking Bayes rule gives
I conclude this section by representing the generalized bound (ref) as a ratio of two moments. Let $m^{\text{help}}_U (W, \xi)$ be as in (ref). Define the moment function for the negative half-space $\mathcal{X}_{\text{hurt}}$:
If $p(X) = 1$, the moment functions coincide $$ m^{\text{help}}_U (W, \xi) = m^{\text{hurt}}_U (W, \xi) = \dfrac{D}{\mu_1(X)} \cdot S \cdot Y - \dfrac{1-D}{\mu_0(X)} \cdot S \cdot Y $$ Combining the moment equations gives
where the true value $\xi_0(x)$ of the first-stage nuisance parameter is
On the boundary $\mathcal{X}_0$, the moment functions $m^{\text{help}}_U (W, \xi)$ and $m^{\text{hurt}}_U (W, \xi)$ coincide. They reduce to the treatment-control difference
which involves no trimming.
Section (ref) states additional identification results, needed for estimation. Section (ref) introduces additional notation for the lower bound. Section (ref) describes population moment functions for the numerators $N_U$ and $N_L$ (Section (ref)) and the always-takers' share (Section (ref)).
Let me introduce some additional notation for the lower bound. The partition-specific moment functions are
where the true value of the first-stage nuisance parameter is
The combined moment function is
Let $N_L := \mathbb{E} m_L(W, \xi_0)$ and $N_U := \mathbb{E} m_U(W, \xi_0)$. As shown in Lemma (ref), the upper bound is a ratio. The same argument applies to the lower bound
The bounds (and their numerators) are ordered by construction
Consider the upper bound numerator $N_U = \mathbb{E}[m_U(W, \xi_0)]$. The moment function $m_U(W, \xi_0)$ is non-orthogonal to the biased estimation of $\xi_0$. To overcome the transmission of this bias, I replace $m^{\text{help}}_U (W, \xi)$ and $m^{\text{hurt}}_U (W, \xi)$ in (ref) by their orthogonal counterparts $g^{\text{help}}_U (W, \xi)$ and $g^{\text{hurt}}_U (W, \xi)$, defined below. In this Section, the propensity score is assumed known, and the correction term for it is not provided.
The bias correction term (ref) consists of three summands, corresponding to the bias correction of $s(0,x)$, $s(1,x)$, and $Q^1(u,x)$ ( Newey1994). As stated in olma2021nonparametric, adding the correction terms and simplifying gives
Remarkably, the function $Q^1(1-p(X), X)$ is the only nuisance component of both the original and the debiased moment function. This function maps covariate vector $X$ into the always-taker's borderline wage in the best case $Q^1(1-p(X), X)$: the lowest wage earned by an always-taker in the extreme case when all always-takers' treated wages are above compliers' treated wages for each covariate value $x \in \mathcal{X}_{\text{help}}$.
Definition (ref) describes the bias correction term for the moment function on $\mathcal{X}_{\text{hurt}}$. The respective terms are obtained mirroring those in (ref) with the roles of the treated and control group reversed. On the boundary $\mathcal{X}_0$, the moment function $m_U(W, \xi)$ involves no trimming, and the correction is not needed. As a result, we have
In this Section, I state a debiased moment function for the always-takers' share. Let $d=1$ and $d=0$ denote the treated and the contol state, respectively. The function
is Robins debiased moment function for the average potential outcome $\mathbb{E} [ s(d,X )] = \mathbb{E} [ S(d) ]$. Indeed, for each $d$, we have $$ E [g^0(W, s) \mid X ] = s(0,X), \quad E [g^1(W, s) \mid X ] = s(1,X). $$ Combining $g^1(W, s)$ and $g^0(W, s)$ gives
By Law of Iterated Expectations, we have
and $g_D(W, s)$ is a valid moment function for the always-takers' share.
In this Section, I describe the estimator of the bounds as well as the confidence region for the identified set. Section (ref) presents examples of the nonparametric estimators of the first-stage nuisance functions. Section (ref) describes the estimator of generalized Lee bounds. Section (ref) explains the use of asymptotic results.
In this Section, I provide examples of the first-stage estimators for the selection probability and for the conditional quantile.
\paragraph{Conditional selection probabilities. } Suppose the selection probability $s(d,x)$ for $d \in \{1, 0\}$ can be approximated by a logistic function
where $\Lambda(\cdot) = \dfrac{\exp (\cdot)}{1 + \exp (\cdot)}$ is the logistic CDF, $B(x) = (B_1(x), B_2(x), \dots B_{p}(x))'$ is a vector of basis functions (e.g., polynomial series or splines), $\gamma^d_{0} \in \mathrm{R}^{p}$ is the pseudo-true value of the logistic parameter, and $r_d(x)$ is its approximation error. The logistic likelihood function is
Given an estimate $\widehat{\gamma}^d$ of $\gamma^d$, define the estimated selection probabilities as
and the estimated CATE on selection
Suppose there exists a vector $\gamma_0 \in \mathrm{R}^{p}$ with only $s_{\gamma}$ non-zero coordinates such that the approximation error $r_d(x)$ in (ref) decays sufficiently fast relative to the sampling error:
If this condition holds, the $\ell_1$-regularized logistic series estimator of Program applies. It takes the following form.
This penalty term $\lambda \| \gamma \|_1$ prevents overfitting in high dimensions by shrinking the estimate toward zero. Program provides practical choices for the penalty $\lambda$ that provably guard against overfitting. An imminent cost of applying the penalty $\lambda$ is regularization, or shrinkage, bias, that does not vanish faster than root-$N$ rate. To prevent this bias from affecting the second stage, I construct a Neyman-orthogonal moment equation for each bound.
\paragraph{Conditional outcome quantiles. } Let $\rho_N = N^{-1/4} \log^{-1} N$ and $U=U_N = [\rho_N, 1-\rho_N]$ be a compact set in $(0,1)$. For each $u \in U$, suppose the $u$-th conditional quantile can be approximated as
where the conditional quantile is defined as
where $t \rightarrow \rho_u(t)$ is a check function. The quantile loss function takes the form
\setcounter{definition}{0}
Section (ref) outlines the estimator for the generalized Lee bounds. Definition (ref) describes cross-fitting. Once the first-stage cross-fitted values are obtained, I estimate parametric components of the bound $N_L, N_U, \pi_{\text{AT}}$, as described in Algorithm (ref).
The estimator $\widehat \pi_{\text{AT}}$ in Definition (ref) shifts the classification threshold from zero to a close point with zero mass. This shift accommodates positive mass at the boundary. However, if the point mass is assumed to be zero, the sequence $\rho_N$ should be replaced by zero. In this case, the estimator (ref) reduces to the debiased estimator proposed in kallus2020assessing in the context of algorithmic fairness.
Definition (ref) combines the debiased moment functions. The covariate space is divided into three parts, as shown in Equation (ref). If the fitted value $\widehat{\tau}(X)$ falls outside the range $[-\rho_N, \rho_N]$, the covariate value $X$ is assumed to be classified correctly with high probability. In this case, the moment sample estimate $g_U(W_i, \widehat{\xi}_i)$ is calculated using the debiased moment function in Definition (ref) or Definition (ref). Otherwise, if $|\widehat{\tau}(X)|$ is too small, the covariate value $X$ is deemed to be difficult to classify. In this case, the moment sample estimate $g_U(W_i, \widehat{\xi}_i)$ is set to its boundary limit value.
Consider the vector $(N_L, N_U, \pi_{\text{AT}})$ of second-stage parameters. In the large sample, the asymptotic distribution of $(\widehat N_L, \widehat N_U, \widehat \pi_{\text{AT}})$ is
Define the asymptotic variance matrix
Suppose $\pi_{\text{AT}}>0$. Delta method gives the asymptotic approximation for $(\widehat{\beta}_L, \widehat{\beta}_U)$:
where the asymptotic covariance matrix $\Omega $ is
Let $\widehat \Gamma$ be a consistent estimator of $\Gamma$, which I assume exists. Define $ \widehat \Omega:= \widehat Q \widehat \Gamma \widehat Q^{T} $ and let the diagonal elements of $\widehat \Omega$ by $ \widehat \Omega_{LL}$ and $\widehat \Omega_{UU}$.
\paragraph{Confidence Region for the identified set $[\beta_L, \beta_U]$. } Given a significance level $\alpha$, a $(1-\alpha)$-Confidence Region for the identified set $[\beta_L, \beta_U]$ takes the form
where $c_{1-\alpha/2}$ is $(1-\alpha/2)$-quantile of $N(0,1)$. the endpoints of the confidence region $CR^{1-\alpha}(c_{\alpha/2}, c_{1-\alpha/2})$ may not be ordered by construction. As shown in chernozhukovmelly, sorting the endpoints can only improve the coverage
property.
In this Section, I describe the assumptions and state the asymptotic results. Section (ref) describes the regularity conditions on the data generating process. Section (ref) outlines the first-stage rate requirements. Section (ref) presents the asymptotic results. Section (ref) verifies Assumption (ref) in the context of high-dimensional sparse design. Section (ref) introduces basic generalized bound, a non-sharp alternative to the proposed bound. Section (ref) sketches the moment equation for the case of an unknown propensity score.
Assumption (ref) requires the selection probabilities and the propensity score to be bounded away from zero and one, which is a standard condition in the literature.
Assumption (ref) assumes that the distribution of the function $\tau(X)$ is sufficiently smooth. For example, if $\tau(X)$ is continuously distributed with a bounded density, (ref) holds. This assumption is routinely assumed in classification analysis (MammenTsybakov, Tsybakov) and empirical welfare maximization (KitagawaTetenov, MbakopTabord).
Assumption (ref) requires the outcome to have bounded support and to be continuously distributed without point masses. This condition is routinely imposed for the consistency of unpenalized (CherChetQuant) and $\ell_1$-penalized (belloni:11) quantile estimators. Furthermore, the conditional density must be bounded away from zero on its support. For example, Assumption (ref) accommodates truncated normal and uniform distributions, but rules out regular normal distribution.
Assumption (ref) places bounds on selection and quantile rates in various norms. For Examples (ref) and (ref), the rates are defined in terms of the model primitives (i.e., the sparsity indices) and are verified below.
Assumption (ref) states that the functions $s(0,x), s(1,x), Q^1(u,x), Q^0(u,x)$ converge in mean square and sup-rate with sufficiently fast rate. The first condition (ref) controls the higher-order bias; it is a classic assumption in the semiparametric literature (see, e.g., Newey1994, chernozhukov2016double).
Theorem (ref) delivers a root-$N$ consistent, asymptotically normal estimator of $(\beta_L, \beta_U)$ assuming the conditional probability of selection and conditional quantile are estimated at a sufficiently fast rate. In particular, this assumption is satisfied when only a few covariates affect selection and the outcome.
\setcounter{remark}{0}
In this Section, I verify Assumption (ref) in the context of high-dimensional sparse models.
A major challenge of this paper is to verify the mean square quantile rate on the set of quantile levels $U_N = [2 \rho_N/\kappa, 1-2 \rho_N/\kappa]$ which involves extreme quantiles. Here, Assumption (ref) focuses on bounded outcomes whose density is bounded away from zero on the support. For example, $Y \sim U[0, M] \mid X, S=1, D=1$, we have
As a result, extreme quantiles of level $\rho_N$ and $1-\rho_N$ can be consistently estimated at a mean square rate $\sqrt{ s\log p/ (N \rho_N)}$, where $N \rho_N$ acts as an effective sample size. Choosing the trimming threshold $\rho_N = N^{-1/4} \log^{-1} N$ makes the mean square quantile rate $q_N = o(N^{-1/4})$ plausible in Example (ref).
In this Section, I present an alternative generalization of Lee bounds under conditional monotonicity, which does not require trimming (and, therefore, quantiles) to be conditional on covariates. Suppose Assumptions (ref) and (ref)(b) hold. Let $\bar{\beta}^{\text{help}}_U$ and $\bar{\beta}^{\text{hurt}}_U$ be the basic Lee bounds of Section (ref), defined on $\mathcal{X}_{\text{help}}$ and $\mathcal{X}_{\text{hurt}}$, respectively. Focusing on the boundary $ \mathcal{X}_0$, define the treatment-control difference as $$\bar{\beta}^0 := \mathbb{E}[ Y \mid D=1, S=1, X \in \mathcal{X}_0] - \mathbb{E}[ Y \mid D=0, S=1, X \in \mathcal{X}_0].$$ Likewise, let
and
By construction, $S_{\text{help}} +S_{\text{hurt}}+ S_0 = \pi_{\text{AT}}$. Aggregating over the covariate space gives basic generalized bound:
If Assumption (ref) holds, $\bar{\beta}_U = \bar{\beta}^{\text{help}}_U$ reduces to basic Lee bound in (ref). This bound is a direct generalization of basic Lee bound to the case of conditional monotonicity.
In this Section, I consider the case when the propensity score is unknown and needs to be estimated. Let $\beta^{\text{help}}_U(x)$ be the conditional Lee bound defined in (ref), and let $\beta^{1\text{help}}_U(x)$ and $\beta^{0\text{help}}_U(x)$ be its first and second summand, respectively. Likewise, let $\beta^{\text{hurt}}_U(x) $ be the conditional Lee bound defined in (ref), and let $\beta^{1\text{hurt}}_U(x) $ and $ \beta^{0\text{hurt}}_U(x)$ be its first and second summand. Below, I describe the debiased moment function for the upper bound $\beta_U$ in (ref).
For $x \in \mathcal{X}_{\text{help}}$, define the Riesz representer function
For $x \in \mathcal{X}_{\text{hurt}}$, define
On the boundary, there is no trimming, and the two functions coincide $$ \Lambda_U(x) = \Lambda^{\text{help}}_U(x) =\Lambda^{\text{hurt}}_U(x), \quad x \in \mathcal{X}_0. $$ The bias correction term for the propensity score is
The debiased moment function takes the form
In particular, the propensity score correction term depends on the conditional trimmed mean function $\Lambda_U(x)$. In a low-dimensional smooth setting, the function $\Lambda_U(x)$ can be estimated by the local linear regression estimator proposed in olma2021nonparametric. In a high-dimensional sparse setting, one could use the automatic debiasing approach of chernozhukov2021automatic.
\paragraph{Data description. } LeeBound studies the effect of winning a lottery to attend JobCorps, a federal vocational and training program, on applicants' wages. In the mid-1990s, JobCorps used lottery-based admission to assess its effectiveness. The control group of $5, 977$ applicants was essentially embargoed from the program for three years, while the remaining applicants (the treated group) could enroll in JobCorps as usual. The sample consists of $9,145$ JobCorps applicants and has data on lottery outcome, hours worked and wages for 208 consecutive weeks after random assignment. In addition, the data contain educational attainment, employment, recruiting experiences, household composition, income, drug use, arrest records, and applicants' background information. These data were collected as part of a baseline interview, conducted by Mathematica Policy Research (MPR) shortly after randomization (Schochet). Lee has condensed this information to 28 covariates, including demographic characteristics, parental education, and income, wages, and hours of work at baseline (see Table (ref) in Appendix or Table 2, LeeBound). This section considers a richer specification, which includes frequency and type of drug use, arrest experiences, reasons for joining JobCorps, and occupation at baseline. I shall refer to these covariate choices as Lee's covariates (28) and All covariates (>1,000), respectively.
Baseline covariates can detect violations of unconditional monotonicity. If this assumption holds, the conditional average treatment effect on employment $\tau(x)$ in (ref) must be either non-positive or non-negative for all covariate values. Consequently, it cannot be the case that
for any $X$ taken to be the subset of all covariates.
The first exercise is to estimate $s(1,x)$ and $s(0,x)$ by a week-specific cross-sectional logistic regression. Let
where $\Lambda(\cdot) = \dfrac{\exp (\cdot)}{1 + \exp (\cdot)}$ is the logistic CDF, $X$ is a vector of baseline covariates that includes a constant and 28 covariates Lee selected, $D \cdot X$ is a vector of covariates interacted with treatment, and $\alpha$ and $\gamma$ are fixed vectors. Figures (ref) and (ref) show the results: the share of subjects with positive selection effect (Figure (ref), solid black line) and the average employment effect for subjects with $\tau(X)<0$ and $\tau(X)>0$ (Figure (ref)).
The second exercise is to test monotonicity without relying on a particular logistic specification
. For each week, I select a small number of discrete covariates and partition the sample into discrete cells $C_j, \quad j \in \{1,2,\dots, J\}$, determined by covariate values. For example, one binary covariate corresponds to $J=2$ two cells. By monotonicity, the vector of cell-specific treatment-control differences in employment rates, $\mathbf{\mu}= (\mathbb{E}[\tau(X) | X \in C_j])_{j=1}^J$, must be non-negative:
The test statistic for the hypothesis in equation (ref) is
and the critical value is the self-normalized critical value of CCK.
Figure (ref) plots the fraction of subjects with a positive JobCorps effect on employment in each week (that is, the fraction of applicants in black dots in Figure (ref)). In the first weeks after random assignment, there is no evidence of a positive JobCorps effect on employment for any group. By the end of the second year (week 104), JobCorps increases employment for nearly $75\%$ of the individuals, and this fraction rises to $0.9$ by the end of the study period (week 208). This pattern is consistent with the JobCorps program description. While being enrolled in JobCorps, participants cannot hold a job, which is known as the lock-in effect (FloresLagunes). After finishing the program, JobCorps graduates may have gained employment skills that help them outperform the control group. However, the share of subjects with positive employment effect never reaches $100\%$, even four years after RA.
Figure (ref) shows the results of testing the inequality in (ref) for each week. The direction of the employment effect varies with socio-economic factors. For example, the applicants who received AFDC benefits during the 8 months before RA or who belonged to median income and yearly earnings groups experience a significantly positive ($p \leq 0.05$) employment effect at weeks $60$--$89$, although the average effect is significantly negative. As another example, the applicants who answered “1: Very important” to the question “How important was getting away from community on the scale from $1$ (very important) to $3$ (not important)?” ($\text{R\_Home}=1$) and who smoke marijuana or hashish a few times each months experience a significantly negative ($p\leq 0.05$) employment effect at week $117$--$152$ despite the average effect being positive. Finally, at week $153$--$186$, the average JobCorps effect is significantly negative for subjects whose most recent arrest occurred less than 12 months ago ($\text{MARRCAT=1}$), despite the average effect being positive.
Figure (ref) plots the average effects on employment rates across weeks. The effect is shown to be negative for weeks $1-89$ and positive thereafter. Remarkably, week 90 is also the only week whose average wage effect was found significant out of four weeks considered (LeeBound). For this reason, the rest of the section focuses on week 90 as the most interesting one.
\paragraph{Results under unconditional monotonicity.} Table (ref) replicates Lee's results under unconditional monotonicity. The estimated effect on week-90 employment is $0.001$. Among treated individuals, approximately 99.9% of wages are attributed to always-takers. As a result, the basic Lee bounds collapse to a near-point estimate, which coincides with the observed treatment-control difference in log wages. Assuming JobCorps does not reduce employment, the wage effect lies between $4.8\%$ and $4.9\%$ on average. The discrete bounds in Column (2) are wider than the basic ones. The sharpness property of Lee bounds holds in population but may fail in sample, as cell-specific employment effects are positive in some cells and negative in others. Under unconditional monotonicity, such sign reversals can only arise from sampling noise; negative effects are truncated at zero.
\paragraph{Results under conditional monotonicity.} Table (ref) presents results under conditional monotonicity using different covariate sets. Columns (1)–(2) use only the original covariates from Lee’s analysis, relying on standard nonparametric assumptions. Columns (3)–(4) incorporate the full set of available covariates, imposing sparsity assumptions on employment (both columns) and wage (Column (4) only) equations. Columns (5)–(6) restrict attention to covariates selected by Lasso from the full set in Columns (3)–(4). The confidence region in Column (6) does not account for uncertainty due to covariate selection. The resulting bounds are not directly comparable since the assumptions underlying the different specifications are not nested.
Our empirical findings are as follows. First, the estimated share of always-takers is 44%. This estimate remains remarkably stable across different covariate specifications. Second, the upper bound on the wage effect ranges from 0.04% in Column (2) to 3.9% in Column (4). Notably, it does not overlap with the basic lower bound reported in Table (ref), presenting further evidence against unconditional monotonicity. Finally, the generalized upper bound is substantially tighter than its basic counterpart. In terms of width, using covariates to tighten the bound reduces the width by approximately 40% in Columns (3)–(4) and nearly 80% in Columns (5)–(6). Assuming structure on the employment equation—such as sparsity or smoothness—is essential for recovering the direction of the selection effect. Further assuming smoothness or sparsity in the wage equation helps rule out implausibly large wage effects. Because the wage effect is close to zero, its sign remains unknown.
Figure (ref) reports the upper and lower bounds on the average log wage for the always-takers in the control state. The lower (upper) bound grows from $1.63$ ($1.92$) in week $5$ to $1.96$ ($1.96$) in week $208$. The bounds' width decreases from $0.3$ in week 14 to $0.01$ in week 208. The gap between the lower and the upper bound shrinks over time as the share of applicants with a positive employment effect, where the average control log wage is point-identified, increases. The upward trend in the control wages suggests that evaluating JobCorps would have been very difficult without a randomized experiment, as one would need to explicitly model mean reversion in the baseline potential wage.
\paragraph{Conclusion. } Lee bounds are a popular empirical strategy for addressing post-randomization selection bias, assuming treatment's effect on selection has the same sign (unconditional monotonicity). This paper generalizes Lee bounds under conditional monotonicity, which allows the direction of selection effect to differ only with observed covariates. This generalization has proven especially useful for JobCorps job training program, where unconditional monotonicity is unlikely to hold. Relaxing conditional monotonicity is left for the future work.