EconBase
← Back to paper

Generalized Lee Bounds

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,731 characters · 25 sections · 131 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Generalized Lee Bounds

\linespread{1.5}

abstractLee (2009) is a common approach to bound the average causal effect in the presence of selection bias, assuming the treatment effect on selection has the same sign for all subjects. This paper generalizes Lee bounds to allow the sign of this effect to be identified by pretreatment covariates, relaxing the standard (unconditional) monotonicity to its conditional analog. Asymptotic theory for generalized Lee bounds is proposed in low-dimensional smooth and high-dimensional sparse designs. The paper also generalizes Lee bounds to accommodate multiple outcomes. Focusing on JobCorps job training program, I first show that unconditional monotonicity is unlikely to hold, and then demonstrate the use of covariates to tighten the bounds.

Introduction

Randomized controlled trials are often complicated by endogenous sample selection and non-response. This problem occurs when treatment affects the researcher's ability to observe an outcome (a selection effect) in addition to the outcome itself (the causal effect of interest). For example, being randomized into a job training program affects both an individual's wage and employment status. Since wages exist only for employed individuals, treatment-control wage difference is contaminated by selection bias. A common way to proceed is to bound the average causal effect from above and below, focusing on subjects whose outcomes are observed regardless of treatment receipt (the always-observed principal strata, FrangakisRubin or the always-takers, LeeBound).

Seminal work by LeeBound proposes nonparametric bounds assuming the selection effect is non-negative for all subjects (monotonicity). For example, if JobCorps cannot deter employment, basic Lee lower bound is the treatment-control difference in wages, where the top wages in the treated group are trimmed until treated and control employment rates are equal. Furthermore, LeeBound shows that the covariate density-weighted conditional Lee bound is weakly tighter than the basic bound that does not involve any covariates. However, only a handful of discrete covariates can be utilized to tighten the bound, since each covariate cell is required to have a positive number of treated and control outcomes.

This paper generalizes Lee bounds under conditional monotonicity, which allows the sign of selection effect to be determined by pre-treatment covariates. Now, baseline covariates now have two roles: (1) to define the subspaces of positive and negative selection response and (2) to tighten the bound on each subspace. For the first one, using the full covariate vector makes the conditional monotonicity assumption the least restrictive. For the second one, any covariate subvector -- even an empty one -- would suffice to define a valid bound. This paper studies the sharp -- the tightest possible -- generalized bound, where all covariates are used for (1) and (2).

I represent the sharp bound as a ratio of two semiparametric moments whose nuisance functions are the conditional probability of selection, the conditional quantile, and the propensity score (i.e., the conditional probability of treatment). If the nonparametric functions are smooth functions of covariates, they can be estimated by logistic series regression of HIR2003 and quantile series of CherChetQuant, respectively. Alternatively, if these functions have a sparse representation with respect to some transformations of covariates, one could employ their $\ell_1$-penalized analogs proposed in orthogStructural and belloni:11, Program. To make the second stage moments insensitive to the first-stage estimation error, I derive Neyman-orthogonal (Neyman:1959, Neyman:1979, AiChen2003, Newey1994, chernozhukov2016double, LRSP) moment functions for the numerator and the denominator. Combining Neyman-orthogonality and sample splitting, I derive a root-$N$ consistent and asymptotically normal estimator of the sharp bounds in the presence of (1) data-driven covariate selection, (2) possible misclassification into the regions of positive ($\mathcal{X}_{\text{help}}$) and negative ($\mathcal{X}_{\text{hurt}}$) selection response and (3) (any) point mass on the region where the selection response is exactly zero (i.e., the boundary between $\mathcal{X}_{\text{help}}$ and $\mathcal{X}_{\text{hurt}}$). This proposal neither requires the covariates to be discrete nor the propensity score to be known, and considerably expands the scope for using Lee bounds in practice.

As an empirical application, I revisit Lee's JobCorps study as in LeeBound using the data from Schochet. The paper's first major empirical finding is to that the unconditional monotonicity of selection (i.e., employment) is unlikely to hold. After imposing conditional monotonicity (and accounting for the differential JobCorps effect on employment), I find that the average JobCorps effect on the always-takers' week 90 wages ranges between $-5\%$ and $1\%$. Finally, I provide evidence of mean reversion of the expected log wage for the always-takers in the control status. This mean reversion corroborates Ashenfelter pattern and shows that earnings would have recovered even without JobCorps training. Therefore, evaluating JobCorps would have been very difficult without a randomized experiment, as one would need to explicitly model mean reversion in the potential wage in the control status.

Appendix B extends Lee's trimming approach to the case of multiple outcomes. A naive approach to construct an identified set is to take a Cartesian product of scalar bounds for each component of the causal parameter. However, since outcomes may be correlated, such a set may contain points that do not correspond to a data generating process. I characterize the sharp identified set for the causal parameter as well as its support function (BM,BMM). Furthermore, I establish debiased uniform over the boundary inference on the support function based on first-stage regularized estimators. The use of this theory is demonstrated for the parameters involving multi-dimensional outcomes, such as standardized treatment effect and wage growth.

\paragraph{Literature review. } The conditional monotonicity assumption has been previously proposed in Kolesar and sloczynski2020not in a treatment choice context, accommodating differential sign of a binary treatment response to a binary instrument. To generalize the Local Average Treatment Effect parameter, sloczynski2020not combines the estimates from no-defier and no-complier regions with signs $1$ and $-1$, respectively, while the boundary (i.e., the no-defier-and-no-complier region) has zero identification power and, therefore, receives zero weight. In contrast, the sample selection problem focuses on the always takers -- a different principal strata -- whose selection behavior does not change in response to treatment.

This paper combines ideas from various branches of economics and statistics, including bounds on causal effects (Manski89, Manski90, HorowitzManski, HorowitzManski2000, FrangakisRubin, angrist:2002, ZhangRubin, angrist:2006, ChenFlores, feller2016weak, APW2013, APW2018, Honore, kamat2021identifying), convex analysis and support function (CherRigStoker, Stoye, BM, BMM, KaidoSantos, Kaido, StoyeSpread, KaidoMolinariStoye, Gafarov, MolinariStoye, MolinariHandbook), monotonicity and latent index models (Vytlacil, KWLATE, kamat2019identifying, sloczynski2020not, MTW, MTW2, Ura), including the bounds on the same empirical context -- JobCorps job training program -- (LeeBound, FloresLagunes, ChenFlores).

Next, this paper contributes to a large body of work on debiased/orthogonal inference for parameters following regularization or model selection (Neyman:1959, Neyman:1979, HardleStoker1989, NeweyStoker, andrews:1994, Newey1994, Robins, robinson:88, ChenAck, ZhangZhang, JM, chernozhukov2016double, LRSP, Program, sasaki2018estimation, sasaki2020unconditional, Sasaki, chiang2019multiway, ning2020doubly, chernozhukov2021debiased, chernozhukov2021automatic, CherSem, NSS, singh2020debiased, Colangelo, Lieli, ZimLech). In many classic cases, such as Robins or robinson:88, orthogonalization expands the set of first-stage parameters to be estimated. In contrast, the set of first-stage nuisance for the truncated conditional mean functional does not expand after orthogonalization. Finally, the paper contributes to a growing literature on machine learning for bounds and partially identified models (kallus2019assessing, jeong2020robust, Bonvini_2021, ZhouSmith, SemJoE). The causal parameter is not a special case of a set-identified linear model of BM,BMM, and the identification and estimation approaches of CCMS and SemJoE do not apply.

An emerging body of research has validated the usefulness of this paper's results by both expanding theoretical framework and/or employing them in applications. For instance, olma2021nonparametric proposes a nonparametric estimator of truncated conditional expectation functions by plugging an orthogonal moment for the truncated mean into a locally linear regression. While interesting on its own, this parameter also enters the correction term for the unknown propensity score in Section (ref). Finally, Heiler2 proposes an estimator of heterogeneous treatment effects using least squares regression.

The paper is organized as follows. Section (ref) reviews basic Lee bounds and Lee's estimator under the standard unconditional monotonicity assumption. Section (ref) generalizes Lee bounds under conditional monotonicity. Section (ref) defines the debiased moment functions for the numerator and the denominator of each boun. Section (ref) overviews the proposed estimator and provides inference results for this parameter. Section (ref) states the asymptotic theory for generalized Lee bounds. Section (ref) presents empirical application. The Supplementary Appendix contains proofs (Appendix A), an extension to the multiple outcome case (Appendix B), and auxiliary empirical details (Appendix C).

Bounds under unconditional monotonicity

Overview of LeeBound's results

Consider the sample selection model in LeeBound. Let $D=1$ be an indicator for treatment receipt. Let $Y(1)$ and $Y(0)$ denote the potential outcomes if an individual is treated or not, respectively. Likewise, let $S(1)=1$ and $S(0)=1$ be dummies for whether an individual's outcome is observed with and without treatment. The data vector $W=(D,X,S,S \cdot Y)$ consists of the treatment status $D$, a baseline covariate vector $X$, the selection status $S=D \cdot S(1) + (1-D) \cdot S(0)$ and the outcome $S \cdot Y = S \cdot (D \cdot Y(1) + (1-D) \cdot Y(0))$ for selected individuals. LeeBound focuses on the average treatment effect (ATE)

align[align omitted — 90 chars of source]

for subjects who are selected into the sample regardless of treatment receipt---the always-takers.

assumptionp{1}[Assumptions of LeeBound] The following statements hold. \begin{compactenum}[(1)] • (Complete Independence). The potential outcomes vector $(Y(1),Y(0),S(1),S(0),X)$ is independent of $D$. • (Monotonicity). \begin{equation} S(1) \geq S(0) \quad a.s. \end{equation} \end{compactenum}

The independence assumption holds by random assignment. In addition, it requires all subjects to have the same probability of being treated. The monotonicity requires all subjects to have the same direction of selection response. In particular, a subject that is selected into the sample when untreated must remain selected if treated: $$ S(0) = 1 \quad \Rightarrow \quad S(1) =1. $$ As a result, $$\mathbb{E} [ Y(0) \mid S(1)=1, S(0)=1] = \mathbb{E} [ Y(0) \mid S(0)=1]. $$ By complete independence, $$ \mathbb{E} [ Y(0) \mid S(0)=1] = \mathbb{E} [ Y \mid S=1, D=0], $$ and $\mathbb{E} [ Y(0) \mid S(1)=1, S(0)=1] $ is point-identified.

In contrast to the control group, a treated outcome can be either an always-taker's outcome or a complier's outcome. The always-takers' share among the treated outcomes is

align[align omitted — 143 chars of source]

In the best case, the always-takers comprise the top $p_0$-quantile of the treated outcomes. The largest possible value of $\beta_0$ is basic upper bound

align[align omitted — 136 chars of source]

where $Q^1 (u)$ is the $u$-quantile of $Y \mid D=1,S=1$ and $p_0$ in (ref) is the trimming threshold. Likewise, the smallest possible one is $$ \bar{\beta}_L = \mathbb{E} [ Y \mid Y \leq Q^1 (p_0), D = 1, S=1] - \mathbb{E} [ Y \mid D = 0, S=1]. $$

Lee's identification strategy can be implemented conditional on covariates. Denote the conditional trimming threshold $p(x)$ as

align[align omitted — 145 chars of source]

and the conditional upper bound $\bar{\beta}_U (x) $ as

align[align omitted — 156 chars of source]

where $Q^{1}(u,x)$ is the conditional $u$-quantile of $Y$ in $S=1,D=1,X=x$ group. The covariate-based bound is

align[align omitted — 188 chars of source]

which, as Lee has shown, is weakly tighter than the basic one (LeeBound2).

The formula for (ref) involves a covariate density function. When data has a single continuous covariate, accounting for the estimation error of the nonparametric density estimate may be challenging both in theory and practice. An alternative approach, proposed in Algorithm (ref), is to discretize covariates. In absence of a better name, it is referred to as discrete Lee bounds and summarizes the discretization procedure Lee used to report covariate-based bound (Table 5, LeeBound).

algorithm[algorithm omitted — 717 chars of source]

Moment-based approach

This Section describes an alternative -- moment-based -- approach to bounds, which no longer requires estimating the always-takers' covariate density function.

assumptionp{2(a)}[Conditional Independence] The potential outcomes vector is independent of the treatment $D$ conditional on $X$: $$(Y(1),Y(0),S(1),S(0)) \perp\!\!\!\perp D \mid X.$$

Assumption (ref) (a) requires complete independence. Equivalently, the propensity score

align[align omitted — 97 chars of source]

must be constant in $X$. Assumption (ref) relaxes the complete independence to its conditional analog.

I represent (ref) as a ratio of two moments. Let $W=(D, X, S, S\cdot Y)$ be the data vector. Define the numerator moment function

align[align omitted — 160 chars of source]

where the true value of the first-stage nuisance function $\xi=\xi(x)$ is

align[align omitted — 99 chars of source]
lemma[Upper bound under unconditional monotonicity] Suppose Assumptions (ref)(b) and (ref) hold. Then, the upper bound $\beta_U$ is a ratio of two moments: \begin{align} \beta_U = \dfrac{ \mathbb{E} m_U(W, \xi_0) }{ \mathbb{E} s(0,X) }. \end{align}

Lemma (ref) represents the sharp bound as a ratio of two moments. To get more insight into the result, note that

align*[align* omitted — 186 chars of source]

Noting that $p(X) s(1,X)=(s(0,X)/s(1,X)) s(1,X)$ simplifies to $s(0,X)$ gives

align[align omitted — 97 chars of source]

Invoking Bayes rule gives

align[align omitted — 263 chars of source]

Applying Law of Iterated Expectations to (ref) gives the representation (ref). This representation makes it possible to employ continuous covariates under various semiparametric assumptions, discussed in Section (ref).

Remark (ref) points out an interesting connection between the selection problem, studied here, and the treatment choice.

remark[Lee's trimming function and Abadie kappa] The moment function for the numerator $N_U$ is a trimmed version of Abadie's kappa weights, that is, $$ m_U(W, \xi) = \left( \dfrac{D}{\mu_1(X)} \cdot 1{\{ Y \geq Q^1(1-p(X),X) \} } - \dfrac{1-D}{\mu_0(X)} \right) \cdot S \cdot Y, $$ where $D$ is the exogenous variable (i.e., “instrument”) and $S$ is the endogenous variable (i.e., “choice” variable).

Generalized Lee Bounds

In this Section, I generalize Lee bounds under conditional monotonicity. Abusing notation, I shall denote the generalized upper bound by $\beta_U$. The conditional probability of selection is

align[align omitted — 83 chars of source]

The conditional average treatment effect (ATE) on selection is

align[align omitted — 61 chars of source]

The sets of positive and negative selection response are

align[align omitted — 133 chars of source]

and the boundary is

align[align omitted — 69 chars of source]

Let $Q^{d}(u,x)$ is the conditional $u$-quantile of $Y$ in $S=1,D=d, X=x$. Given a quantile level $u \in (0,1)$, define the upper-trimmed mean for the treated group

align[align omitted — 117 chars of source]

and the lower-trimmed mean for the control group

align[align omitted — 117 chars of source]
assumptionp{2(b)} [Conditional monotonicity] The covariate set $\mathcal{X}=\mathcal{X}_{\text{help}} \sqcup \mathcal{X}_{\text{hurt}} \sqcup \mathcal{X}_0$ can be partitioned into the sets $\mathcal{X}_{\text{help}} $, $\mathcal{X}_{\text{hurt}} $ and $\mathcal{X}_0$ so that \begin{align*} X &\in \mathcal{X}_{help} \Rightarrow S(1) \geq S(0) a.s. , \\ X &\in \mathcal{X}_{hurt} \Rightarrow S(1) \leq S(0) a.s. , \\ X &\in \mathcal{X}_0 \Rightarrow S(1) = S(0) a.s. . \end{align*}

Assumption (ref) generalizes the regular unconditional monotonicity assumption to the conditional analog. It allows the direction of the selection effect to vary across covariate values. Defiers $(S(1) = 0, S(0) = 1)$ are ruled out on the covariate set $\mathcal{X}_{\text{help}}$, compliers $(S(1) = 1, S(0) = 0)$ are ruled out on $ \mathcal{X}_{\text{hurt}} $, and both compliers and defiers are ruled out on the set $ \mathcal{X}_0$. In many practical cases, the boundary set $\mathcal{X}_0 $ has zero mass, but the proposed asymptotic theory accommodates the boundary of arbitrary size.

Let us show that the always-takers' share is point-identified under Assumptions (ref) and (ref). For any covariate value $x$ on the positive half-space $\mathcal{X}_{\text{help}}$, we have

align*[align* omitted — 78 chars of source]

Likewise, for any covariate value $x \in \mathcal{X}_{\text{hurt}}$,

align*[align* omitted — 76 chars of source]

On the boundary $\mathcal{X}_0$, $\tau (x) = 0 \Rightarrow s (1,x) = s (0,x)$. Combining the results

align[align omitted — 90 chars of source]

and aggregating over the covariate space gives the always-takers' share

align[align omitted — 113 chars of source]
lemma[Always-takers' Share] Suppose Assumptions (ref) and (ref) hold. Then, the always-takers' share $ \pi_{\text{AT}}$ is point-identified, and \begin{align} \pi_{AT} &:=\Pr (S(1) = S(0) = 1)= \mathbb{E} \min (s(0,X), s(1,X)), \end{align} where the expectation is taken over the covariate distribution.

\qed

Next, let me generalize the bound itself. For $x \in \mathcal{X}_{\text{help}}$, the conditional always-takers' share is $$ \min (s(0,x), s(1,x)) = s(0,x), \quad x \in \mathcal{X}_{\text{help}}.$$ The conditional upper bound takes the form

align[align omitted — 133 chars of source]

where $\beta^{\text{basic}}_U(x)$ is given in (ref). Likewise, for $x \in \mathcal{X}_{\text{hurt}}$, the roles of the treated and control group are reversed:

align[align omitted — 125 chars of source]

By Assumption (ref), both defiers and compliers are ruled out on the boundary $\mathcal{X}_0$. Therefore, a selected individual with $S=1$ must be an always-taker, and there is no trimming. The bound reduces to treatment-control difference

align[align omitted — 119 chars of source]

Aggregating the conditional sharp bound $\beta_U(x)$ gives

align[align omitted — 108 chars of source]

Invoking Bayes rule gives

align*[align* omitted — 240 chars of source]

I conclude this section by representing the generalized bound (ref) as a ratio of two moments. Let $m^{\text{help}}_U (W, \xi)$ be as in (ref). Define the moment function for the negative half-space $\mathcal{X}_{\text{hurt}}$:

align[align omitted — 161 chars of source]

If $p(X) = 1$, the moment functions coincide $$ m^{\text{help}}_U (W, \xi) = m^{\text{hurt}}_U (W, \xi) = \dfrac{D}{\mu_1(X)} \cdot S \cdot Y - \dfrac{1-D}{\mu_0(X)} \cdot S \cdot Y $$ Combining the moment equations gives

align[align omitted — 320 chars of source]

where the true value $\xi_0(x)$ of the first-stage nuisance parameter is

align[align omitted — 173 chars of source]

On the boundary $\mathcal{X}_0$, the moment functions $m^{\text{help}}_U (W, \xi)$ and $m^{\text{hurt}}_U (W, \xi)$ coincide. They reduce to the treatment-control difference

align[align omitted — 135 chars of source]

which involves no trimming.

lemma[Generalized Lee Bound] Suppose Assumptions (ref) and (ref) hold, and $\pi_{\text{AT}}>0$. Then, the upper bound $\beta_U$ in (ref) is a sharp upper bound on the average treatment effect $\beta_0$ in (ref). The bound is \begin{align} \beta_U &= \dfrac{ \mathbb{E} [\beta_U(X) \min (s(0,X), s(1,X))] }{ \mathbb{E} \min (s(0,X),s(1,X))} = \dfrac{ \mathbb{E} [m_U(W, \xi_0)] }{ \mathbb{E} \min (s(0,X),s(1,X)) }. \end{align} where $\beta_U(x)$ is given in (ref)--(ref).

Debiased Moment Functions

Section (ref) states additional identification results, needed for estimation. Section (ref) introduces additional notation for the lower bound. Section (ref) describes population moment functions for the numerators $N_U$ and $N_L$ (Section (ref)) and the always-takers' share (Section (ref)).

Additional Notation for Lower Bound.

Let me introduce some additional notation for the lower bound. The partition-specific moment functions are

align[align omitted — 328 chars of source]

where the true value of the first-stage nuisance parameter is

align[align omitted — 173 chars of source]

The combined moment function is

align*[align* omitted — 117 chars of source]

Let $N_L := \mathbb{E} m_L(W, \xi_0)$ and $N_U := \mathbb{E} m_U(W, \xi_0)$. As shown in Lemma (ref), the upper bound is a ratio. The same argument applies to the lower bound

align[align omitted — 116 chars of source]

The bounds (and their numerators) are ordered by construction

align[align omitted — 127 chars of source]

Second Stage Moment Functions

The Upper Bound Numerator

Consider the upper bound numerator $N_U = \mathbb{E}[m_U(W, \xi_0)]$. The moment function $m_U(W, \xi_0)$ is non-orthogonal to the biased estimation of $\xi_0$. To overcome the transmission of this bias, I replace $m^{\text{help}}_U (W, \xi)$ and $m^{\text{hurt}}_U (W, \xi)$ in (ref) by their orthogonal counterparts $g^{\text{help}}_U (W, \xi)$ and $g^{\text{hurt}}_U (W, \xi)$, defined below. In this Section, the propensity score is assumed known, and the correction term for it is not provided.

definition[Orthogonal Moment Function $g^{\text{help}}_U (W, \xi)$ on $\mathcal{X}_{\text{help}}$] Let $X \in \mathcal{X}_{\text{help}}$. Define the bias correction term \begin{align} cor^{help}_U (W, \xi) &= Q^1(1-p(X), X) \bigg( \dfrac{1-D}{\mu_0(X) }\cdot ( S - s(0,X)) \\ &-\dfrac{D}{\mu_1(X) } \cdot p(X) \cdot ( S - s(1,X)) \nonumber \\ &+ \dfrac{D S}{\mu_1(X) } ( 1 \{ Y \leq Q^1(1-p(X), X) \} - 1+ p(X)) \bigg). \nonumber \end{align} and the debiased moment function \begin{align} g^{help}_U (W, \xi):= m^{help}_U (W, \xi) + cor^{help}_U (W, \xi). \end{align}

The bias correction term (ref) consists of three summands, corresponding to the bias correction of $s(0,x)$, $s(1,x)$, and $Q^1(u,x)$ ( Newey1994). As stated in olma2021nonparametric, adding the correction terms and simplifying gives

align[align omitted — 253 chars of source]

Remarkably, the function $Q^1(1-p(X), X)$ is the only nuisance component of both the original and the debiased moment function. This function maps covariate vector $X$ into the always-taker's borderline wage in the best case $Q^1(1-p(X), X)$: the lowest wage earned by an always-taker in the extreme case when all always-takers' treated wages are above compliers' treated wages for each covariate value $x \in \mathcal{X}_{\text{help}}$.

definition[Orthogonal Moment Function $g^{\text{hurt}}_U(W, \xi)$ on $\mathcal{X}_{\text{hurt}}$] Let $X \in \mathcal{X}_{\text{hurt}}$. Define the bias correction term \begin{align} &cor^{hurt}_U (W, \xi) = -Q^0(1/p(X), X) \left(- \dfrac{1-D}{\mu_0(X)} (1/p(X)) \cdot (S - s(0,X)) \right. \\ &\quad\quad + \left. \dfrac{D}{\mu_1(X)} \cdot (S - s(1,X)) - \dfrac{(1-D)S}{\mu_0(X)} \left( 1\{Y \leq Q^0(1/p(X), X)\} - 1/p(X) \right) \right) \nonumber \end{align} and the debiased moment function \begin{align*} g^{hurt}_U (W, \xi) := m^{hurt}_U (W, \xi) + cor^{hurt}_U (W, \xi). \end{align*}

Definition (ref) describes the bias correction term for the moment function on $\mathcal{X}_{\text{hurt}}$. The respective terms are obtained mirroring those in (ref) with the roles of the treated and control group reversed. On the boundary $\mathcal{X}_0$, the moment function $m_U(W, \xi)$ involves no trimming, and the correction is not needed. As a result, we have

align[align omitted — 308 chars of source]

The Lower Bound Numerator

definition[Moment Functions for Lower Bound] Let $X \in \mathcal{X}_{\text{help}}$. Define the bias correction term \begin{align} cor^{help}_L (W, \xi) &= Q^1(p(X), X) \bigg( \dfrac{1-D}{\mu_0(X)} \cdot (S - s(0,X)) \\ &- \dfrac{D}{\mu_1(X)} \cdot p(X) \cdot (S - s(1,X)) - \dfrac{D S}{\mu_1(X)} \left( 1\{Y \leq Q^1(p(X), X)\} - p(X) \right) \bigg) \end{align} Let $X \in \mathcal{X}_{\text{hurt}}$. Define the bias correction term \begin{align} &cor^{hurt}_L (W, \xi) = -Q^0(1-1/p(X), X) \left( -\dfrac{1-D}{\mu_0(X)} (1/p(X)) \cdot (S - s(0,X)) \right. \\ &+\quad\quad \left. \dfrac{D}{\mu_1(X)} \cdot (S - s(1,X)) + \dfrac{(1-D)S}{\mu_0(X)} \left( 1/p(X) - 1\{Y \geq Q^0(1-1/p(X), X)\} \right) \right) \nonumber \end{align} The debiased moment function is \begin{align} g_L(W, \xi):= \begin{cases} m^{help}_L (W, \xi) + cor^{\text{help}}_L (W, \xi) , \quad X \in \mathcal{X}_{\text{help}} \\ m^{\text{hurt}}_L (W, \xi) + \text{cor}^{\text{hurt}}_L (W, \xi) , \quad X \in \mathcal{X}_{\text{hurt}} \\ \left(\dfrac{DS}{\mu_1(X)} - \dfrac{(1-D)S}{\mu_0(X)} \right)Y, \qquad X \in \mathcal{X}_0. \end{cases} \end{align}

The Always-Takers' Share (Denominator)

In this Section, I state a debiased moment function for the always-takers' share. Let $d=1$ and $d=0$ denote the treated and the contol state, respectively. The function

align[align omitted — 87 chars of source]

is Robins debiased moment function for the average potential outcome $\mathbb{E} [ s(d,X )] = \mathbb{E} [ S(d) ]$. Indeed, for each $d$, we have $$ E [g^0(W, s) \mid X ] = s(0,X), \quad E [g^1(W, s) \mid X ] = s(1,X). $$ Combining $g^1(W, s)$ and $g^0(W, s)$ gives

align[align omitted — 166 chars of source]

By Law of Iterated Expectations, we have

align*[align* omitted — 172 chars of source]

and $g_D(W, s)$ is a valid moment function for the always-takers' share.

Overview of Estimation and Inference

In this Section, I describe the estimator of the bounds as well as the confidence region for the identified set. Section (ref) presents examples of the nonparametric estimators of the first-stage nuisance functions. Section (ref) describes the estimator of generalized Lee bounds. Section (ref) explains the use of asymptotic results.

Examples of First Stage Estimators

In this Section, I provide examples of the first-stage estimators for the selection probability and for the conditional quantile.

\paragraph{Conditional selection probabilities. } Suppose the selection probability $s(d,x)$ for $d \in \{1, 0\}$ can be approximated by a logistic function

align[align omitted — 99 chars of source]

where $\Lambda(\cdot) = \dfrac{\exp (\cdot)}{1 + \exp (\cdot)}$ is the logistic CDF, $B(x) = (B_1(x), B_2(x), \dots B_{p}(x))'$ is a vector of basis functions (e.g., polynomial series or splines), $\gamma^d_{0} \in \mathrm{R}^{p}$ is the pseudo-true value of the logistic parameter, and $r_d(x)$ is its approximation error. The logistic likelihood function is

align[align omitted — 178 chars of source]

Given an estimate $\widehat{\gamma}^d$ of $\gamma^d$, define the estimated selection probabilities as

align[align omitted — 109 chars of source]

and the estimated CATE on selection

align*[align* omitted — 71 chars of source]

Suppose there exists a vector $\gamma_0 \in \mathrm{R}^{p}$ with only $s_{\gamma}$ non-zero coordinates such that the approximation error $r_d(x)$ in (ref) decays sufficiently fast relative to the sampling error:

align*[align* omitted — 149 chars of source]

If this condition holds, the $\ell_1$-regularized logistic series estimator of Program applies. It takes the following form.

examplep{1}[$\ell_1$-penalized LR, orthogStructural] Given the penalty parameter $\lambda_S$, the $\ell_1$-regularized logistic estimator of $\gamma^d$ is \begin{align} \widehat{\gamma}^d_{L}= \arg \max_{\gamma^d \in \mathrm{R}^{p}} \ell_d(\gamma^d) + \lambda_S \| \gamma^d \|_1. \end{align}

This penalty term $\lambda \| \gamma \|_1$ prevents overfitting in high dimensions by shrinking the estimate toward zero. Program provides practical choices for the penalty $\lambda$ that provably guard against overfitting. An imminent cost of applying the penalty $\lambda$ is regularization, or shrinkage, bias, that does not vanish faster than root-$N$ rate. To prevent this bias from affecting the second stage, I construct a Neyman-orthogonal moment equation for each bound.

\paragraph{Conditional outcome quantiles. } Let $\rho_N = N^{-1/4} \log^{-1} N$ and $U=U_N = [\rho_N, 1-\rho_N]$ be a compact set in $(0,1)$. For each $u \in U$, suppose the $u$-th conditional quantile can be approximated as

align[align omitted — 78 chars of source]

where the conditional quantile is defined as

align*[align* omitted — 146 chars of source]

where $t \rightarrow \rho_u(t)$ is a check function. The quantile loss function takes the form

align*[align* omitted — 107 chars of source]
examplep{2}[$\ell_1$-penalized QR, belloni:11] Let $\widehat \sigma^2_j:= N^{-1} \sum_{i=1}^N B^2_j(X_i), \quad j=1,2,\dots, p$. Given the penalty parameter $\lambda_Q$, the $\ell_1$-penalized quantile regression estimator is \begin{align} \widehat{\xi}_L^d(u)=\arg \min_{\xi^d \in \mathrm{R}^{p} }\ell_u(\xi^{d}) + \lambda_{Q}/N \sqrt{ u(1-u) } \sum_{j=1}^{p} \widehat {\sigma}_j | \xi^d | \end{align} and the quantile estimate takes the form \begin{align*} \widehat{Q}^d(u,x):= B(x)'\widehat{\xi}^d_L(u), \quad d \in \{1, 0\}. \end{align*}

\setcounter{definition}{0}

The Estimator of Generalized Lee Bounds

Section (ref) outlines the estimator for the generalized Lee bounds. Definition (ref) describes cross-fitting. Once the first-stage cross-fitted values are obtained, I estimate parametric components of the bound $N_L, N_U, \pi_{\text{AT}}$, as described in Algorithm (ref).

definition[Cross-Fitted Values] \begin{compactenum} • For a random sample of size $N$, denote a $K$-fold random partition of the sample indices $[N]=\{1,2,...,N\}$ by $(J_k)_{k=1}^K$, where $K$ is the number of partitions and the sample size of each fold is $n = N/K$. For each $k \in [K] = \{1,2,...,K\}$ define $J_k^c = \{1,2,...,N\} \setminus J_k$. • For each $k \in [K]$, construct an estimator $\widehat{\xi}_k = \widehat{\xi}( W_{i \in J_k^c})$ of the nuisance parameter $\xi_0$ using only the data $\{ W_{j}: j \in J_k^c \}$. For any observation $i \in J_k$, define the cross-fitted value $\widehat s_i := (\widehat s_k(1, X_i), \widehat s_k(0,X_i)), \widehat \tau_i := \widehat \tau_k (X_i) = \widehat s_k(1,X_i) - \widehat s_k(0,X_i), \widehat \xi_i := \widehat \xi_k (X_i)$. \end{compactenum}
definition[Debiased Estimator of the Always-Takers' Share] Let $\rho_N:= N^{-1/4} \log^{-1} N$. Given the first-stage fitted values $(\widehat s_i)_{i=1}^N$ and $(\widehat \tau_i)_{i=1}^N$, define \begin{align} \widehat \pi_{AT}:&= N^{-1} \sum_{i=1}^N g^0(W_i, \widehat s_i) 1{ \{ \widehat \tau_i \geq \rho_N \} } + g^1(W_i, \widehat s_i) 1{ \{ \widehat \tau_i < \rho_N \} } . \end{align}

The estimator $\widehat \pi_{\text{AT}}$ in Definition (ref) shifts the classification threshold from zero to a close point with zero mass. This shift accommodates positive mass at the boundary. However, if the point mass is assumed to be zero, the sequence $\rho_N$ should be replaced by zero. In this case, the estimator (ref) reduces to the debiased estimator proposed in kallus2020assessing in the context of algorithmic fairness.

definition[Debiased Estimator of the Numerator $N_U$ and $N_L$] Let $\widehat \xi_i = \widehat \xi(X_i)$ be the nuisance parameter cross-fit estimates. Given the sequence $\rho_N:= N^{-1/4} \log^{-1} N$, define the estimated moment \begin{align} g_{\star}(W_i, \widehat{\xi}_i) :&= \begin{cases} g^{help}_{\star}(W_i, \widehat{\xi}_i), \qquad \widehat{\tau} (X_i) \geq \rho_N \\ g^{hurt}_{\star}(W_i, \widehat{\xi}_i), \qquad \widehat{\tau} (X_i) \leq -\rho_N \\ \left(\dfrac{D_i}{\mu_1(X_i)} - \dfrac{1-D_i}{\mu_0(X_i)}\right) S_i Y_i, \qquad |\widehat{\tau} (X_i)| \leq \rho_N, \end{cases}, \qquad \star \in \{L, U\} \end{align} where $g^{\text{help}}_{\star}(W, \xi_0)$ and $g^{\text{hurt}}_{\star}(W, \xi_0)$ are the debiased moment functions defined on the covariate partitions $\mathcal{X}_{\text{help}}$ and $\mathcal{X}_{\text{hurt}}$, respectively.

Definition (ref) combines the debiased moment functions. The covariate space is divided into three parts, as shown in Equation (ref). If the fitted value $\widehat{\tau}(X)$ falls outside the range $[-\rho_N, \rho_N]$, the covariate value $X$ is assumed to be classified correctly with high probability. In this case, the moment sample estimate $g_U(W_i, \widehat{\xi}_i)$ is calculated using the debiased moment function in Definition (ref) or Definition (ref). Otherwise, if $|\widehat{\tau}(X)|$ is too small, the covariate value $X$ is deemed to be difficult to classify. In this case, the moment sample estimate $g_U(W_i, \widehat{\xi}_i)$ is set to its boundary limit value.

algorithm[algorithm omitted — 1,059 chars of source]

Asymptotic Distribution of Second-Stage Parameters

Consider the vector $(N_L, N_U, \pi_{\text{AT}})$ of second-stage parameters. In the large sample, the asymptotic distribution of $(\widehat N_L, \widehat N_U, \widehat \pi_{\text{AT}})$ is

align*[align* omitted — 219 chars of source]

Define the asymptotic variance matrix

align[align omitted — 96 chars of source]

Suppose $\pi_{\text{AT}}>0$. Delta method gives the asymptotic approximation for $(\widehat{\beta}_L, \widehat{\beta}_U)$:

align*[align* omitted — 184 chars of source]

where the asymptotic covariance matrix $\Omega $ is

align[align omitted — 162 chars of source]

Let $\widehat \Gamma$ be a consistent estimator of $\Gamma$, which I assume exists. Define $ \widehat \Omega:= \widehat Q \widehat \Gamma \widehat Q^{T} $ and let the diagonal elements of $\widehat \Omega$ by $ \widehat \Omega_{LL}$ and $\widehat \Omega_{UU}$.

\paragraph{Confidence Region for the identified set $[\beta_L, \beta_U]$. } Given a significance level $\alpha$, a $(1-\alpha)$-Confidence Region for the identified set $[\beta_L, \beta_U]$ takes the form

align[align omitted — 200 chars of source]

where $c_{1-\alpha/2}$ is $(1-\alpha/2)$-quantile of $N(0,1)$. the endpoints of the confidence region $CR^{1-\alpha}(c_{\alpha/2}, c_{1-\alpha/2})$ may not be ordered by construction. As shown in chernozhukovmelly, sorting the endpoints can only improve the coverage

footnote{This paper focuses on the pointwise coverage, where the true values of $N_U, N_L, \pi_{\text{AT}}$ do not change with sample size. }

property.

Asymptotic Theory for Generalized Lee Bounds

In this Section, I describe the assumptions and state the asymptotic results. Section (ref) describes the regularity conditions on the data generating process. Section (ref) outlines the first-stage rate requirements. Section (ref) presents the asymptotic results. Section (ref) verifies Assumption (ref) in the context of high-dimensional sparse design. Section (ref) introduces basic generalized bound, a non-sharp alternative to the proposed bound. Section (ref) sketches the moment equation for the case of an unknown propensity score.

Assumptions

assumptionp{3}[Strict Overlap] (SO). There exists an absolute constant $\kappa \in (0, 1/2)$ so that $s(d,x) \in ( \kappa, 1- \kappa ) \quad$ for all $d \in \{1, 0\}$ and all covariate values $x$. Likewise, the propensity score $\mu_1(x):= \Pr (D=1 \mid X=x) \in ( \kappa, 1- \kappa )$ for any $x$.

Assumption (ref) requires the selection probabilities and the propensity score to be bounded away from zero and one, which is a standard condition in the literature.

assumptionp{4}[Margin Assumption] (MA). There exist absolute finite constants $\bar{B}_f$ and $\eta$ so that \begin{align} {\Pr} ( 0< | \tau(X) | \leq t) \leq \bar{B}_f t, \quad 0 \leq t \leq \eta. \end{align}

Assumption (ref) assumes that the distribution of the function $\tau(X)$ is sufficiently smooth. For example, if $\tau(X)$ is continuously distributed with a bounded density, (ref) holds. This assumption is routinely assumed in classification analysis (MammenTsybakov, Tsybakov) and empirical welfare maximization (KitagawaTetenov, MbakopTabord).

assumptionp{5}[Continuously Distributed Bounded Outcome] (BO) Bounded Outcome: There exists a constant $M < \infty$ such that $|Y| \leq M$ almost surely. (REG): For $d \in \{1, 0\}$, there exist constants $C_f$ and $B_f$ such that for the support $\mathcal{Y}^d_x$ of the conditional distribution $Y \mid D=d, X=x$, we have \begin{enumerate} • The conditional density $f^d(y \mid x) := f_{Y \mid S=1, D=d, X=x}(y \mid x)$ is uniformly bounded from above by $C_f$ for all $y \in \mathcal{Y}^d_x$. • The infimum of $f^d(y \mid x)$ over $x \in \mathcal{X}$ and $y \in \mathcal{Y}_x$ is bounded away from zero by $B_f$. • The derivative of $y \rightarrow f^d(y \mid x)$ is continuous and bounded from above in absolute value by $C_f$ uniformly over $y \in \mathcal{Y}^d_x$. \end{enumerate}

Assumption (ref) requires the outcome to have bounded support and to be continuously distributed without point masses. This condition is routinely imposed for the consistency of unpenalized (CherChetQuant) and $\ell_1$-penalized (belloni:11) quantile estimators. Furthermore, the conditional density must be bounded away from zero on its support. For example, Assumption (ref) accommodates truncated normal and uniform distributions, but rules out regular normal distribution.

First-Stage Rate Requirements

definition[Selection Rate] There exist a sequence of numbers $\phi_N = o(1)$ and a sequence of sets $S^d_N, d \in \{1, 0\}$ such that the first-stage estimates $\widehat{s}(d, x)$ of the true function $s(d, x)$ belong to $S^d_N$ with probability at least $1 - \phi_N$. The sets $S^d_N$ shrink at the following rate: \begin{align*} s_N^p:=\sup_{d \in \{1,0\}} \sup_{\bar{s} \in S^d_N} \left(\mathbb{E}_{X} | \bar{s}(d, X) - s(d,X)|^p\right)^{1/p}, \quad 1 \leq p \leq \infty, \end{align*} and the functions in $S^d_N$ satisfy $\inf_{x \in \mathcal{X}} \inf_{d \in \{1,0\}} s(d,x) > \kappa/2 > 0$. Let $s_N$ and $s^1_N$ and $s_N^{\infty}$ be the mean square, the $L_1$- and sup-norm rates, respectively.
definition[Quantile Rate] There exist a sequence of numbers $\phi_N = o(1)$ and a sequence of sets $Q^d_N$ such that the first-stage estimate $\widehat{Q}^d(u,x)$ of $Q^d(u,x)$ shrinks uniformly over $\mathcal{U}_N = [2/\kappa \rho_N, 1- 2/\kappa \rho_N]$ with $\rho_N = N^{-1/4} \log^{-1} N$ at the following rate: \begin{align*} q_N^p:= \sup_{d \in \{1, 0\}}\sup_{ \bar{Q}^d \in Q_N^d} \sup_{u \in \mathcal{U}_N} \left(\mathbb{E} | \bar{Q}^d(u, X) - Q^d(u,X) |^p\right)^{1/p}, \quad 1 \leq p \leq \infty, \end{align*} where the sets $Q_N^d$ consist of almost surely $M$-bounded functions. Let $q_N$ and $q_N^1$ be the mean square and $L_1$ rates, respectively.

Assumption (ref) places bounds on selection and quantile rates in various norms. For Examples (ref) and (ref), the rates are defined in terms of the model primitives (i.e., the sparsity indices) and are verified below.

assumptionp{6}[First-Stage Rates] The sequences $s_N$, $q_N$, $s_N^{\infty}$, $s_N^1$, and $q_N^1$ obey the following bounds: \begin{enumerate} • Mean square rates are sufficiently fast: \begin{align} s_N + q_N = o(N^{-1/4}). \end{align} • Worst-case selection rate $s_N^{\infty}$ is sufficiently fast \begin{align} s_N^{\infty} = o(N^{-1/4} \log^{-1} N). \end{align} • Estimators are consistent in $L_1$ norm \begin{align*} s_N^1 = o(1), \quad q_N^1 = o(1). \end{align*} \end{enumerate}

Assumption (ref) states that the functions $s(0,x), s(1,x), Q^1(u,x), Q^0(u,x)$ converge in mean square and sup-rate with sufficiently fast rate. The first condition (ref) controls the higher-order bias; it is a classic assumption in the semiparametric literature (see, e.g., Newey1994, chernozhukov2016double).

Main Result

theorem[Generalized Lee bounds: Asymptotic Theory] Suppose Assumptions (ref), (ref), (ref)--(ref) hold. Then, the estimator $(\widehat{N}_L, \widehat{N}_U, \widehat{\pi}_{\text{AT}})$ is consistent and asymptotically normal: \begin{align*} \sqrt{N} \left( \begin{matrix} \widehat{N}_L - N_L \\ \widehat{N}_U - N_U \\ \widehat{\pi}_{AT} - \pi_{AT} \\ \end{matrix} \right) \Rightarrow N \left(0, \Gamma \right), \end{align*} where the asymptotic variance matrix $\Gamma$ is given in (ref). As a result, if $\pi_{\text{AT}} > 0$, the preliminary bounds of Algorithm (ref) are asymptotically Gaussian: \begin{align*} \sqrt{N} \left( \begin{matrix} \widehat{\beta}_L - \beta_L \\ \widehat{\beta}_U - \beta_U \\ \end{matrix} \right) \Rightarrow N \left(0, \Omega \right), \end{align*} where $\Omega = Q \Gamma Q^T$ as in (ref).

Theorem (ref) delivers a root-$N$ consistent, asymptotically normal estimator of $(\beta_L, \beta_U)$ assuming the conditional probability of selection and conditional quantile are estimated at a sufficiently fast rate. In particular, this assumption is satisfied when only a few covariates affect selection and the outcome.

\setcounter{remark}{0}

remark[Strong\begin{footnote}{The 2020 version of the manuscript was based on this assumption. The author thanks to the discussants and referees who pointed out its weaknesses. } \end{footnote} separation from the boundary] Consider Assumption (ref) holding with $\Pr(\mathcal{X}_0) = 0$. Given a fixed $\epsilon > 0$, a separation condition \begin{align} \inf_{x \in \mathcal{X}} |\tau(x)| = \inf_{x \in \mathcal{X}} |s(1,x) - s(0,x)| > \epsilon \end{align} may be plausible in settings with discrete covariates. If $s_N^{\infty} = o(1)$, the subjects are correctly classified into $\mathcal{X}_{\text{help}}$ and $\mathcal{X}_{\text{hurt}}$ with probability approaching one. Then, the statement of Theorem (ref) holds under Assumptions (ref), (ref) (REG), and Assumption (ref) (1)-(2), while (MA) and (BO) are no longer required. Furthermore, the relevant set of estimated quantiles $U$ reduces to $U := [2\epsilon/\kappa, 1-2\epsilon/\kappa]$ and no longer approaches $(0,1)$ as the sample size grows. As a result, unbounded outcome distributions such as Gaussian satisfy Assumption (ref) (REG).

Verification of Assumption (ref)

In this Section, I verify Assumption (ref) in the context of high-dimensional sparse models.

examplep{1'}[Example (ref), cont.] Consider the model (ref) with $p=\dim(B(X)) \gg N$. Suppose there exists a vector $\gamma^d_{0} \in \mathrm{R}^{p}$ with only $s_{\gamma}$ non-zero coordinates such that the approximation error $r_d(x)$ in (ref) decays sufficiently fast relative to the sampling error: \begin{align*} \sup_{d \in {1, 0 }} \left(\frac{1}{N} \sum_{i=1}^N r_d^2(X_i) \right)^{1/2} \lesssim_P \sqrt{\dfrac{s_{\gamma}^2 \log p }{N}}. \end{align*} Then, the $\ell_1$-regularized LR of Example (ref) with the data-driven choice of penalty as in Program attains the mean square rate $s_N: = O\left(\sqrt{s_{\gamma} \log p /N} \right)$ and $s^1_N :=s^{\infty}_N: = O \left(\sqrt{s_{\gamma}^2 \log p /N} \right)$. Thus, \begin{align*} s_{\gamma}^2 \log^2 N \log p = o (N^{1/2}) \end{align*} is sufficient for $s^{\infty}_N = o(N^{-1/4} \log^{-1} N)$.

A major challenge of this paper is to verify the mean square quantile rate on the set of quantile levels $U_N = [2 \rho_N/\kappa, 1-2 \rho_N/\kappa]$ which involves extreme quantiles. Here, Assumption (ref) focuses on bounded outcomes whose density is bounded away from zero on the support. For example, $Y \sim U[0, M] \mid X, S=1, D=1$, we have

align*[align* omitted — 149 chars of source]

As a result, extreme quantiles of level $\rho_N$ and $1-\rho_N$ can be consistently estimated at a mean square rate $\sqrt{ s\log p/ (N \rho_N)}$, where $N \rho_N$ acts as an effective sample size. Choosing the trimming threshold $\rho_N = N^{-1/4} \log^{-1} N$ makes the mean square quantile rate $q_N = o(N^{-1/4})$ plausible in Example (ref).

examplep{2'}[Example (ref), cont.] Consider the model (ref). Suppose Assumption (ref) holds. Then, the $\ell_1$-regularized quantile regression of belloni:11 with data-driven choice of penalty attains the mean square rate and $L_1$ rates: \begin{align} q_N= O\left( \sqrt{ \dfrac{ s_Q \log p }{N \rho_N} }\right), \quad q^1_N=O \left( \sqrt{ \dfrac{ s^2_Q \log p }{N \rho_N} } \right), \end{align} where $\rho_N$ is the trimming threshold in Definition (ref), and $U_N = [2 \rho_N/\kappa, 1-2 \rho_N/\kappa]$. Here, the quantity $N \rho_N$ is the effective sample size used to estimate the conditional $(2 \rho_N/\kappa)$-quantile. The proposed choice $\rho_N:= N^{-1/4} \log^{-1} N$ ensures that \begin{align*} \log N s^2_Q \log p = o(N^{1/4}) \end{align*} holds, which suffices for $q_N = o(N^{-1/4})$ and $q^1_N = o(1)$.

Basic Generalized Bound

In this Section, I present an alternative generalization of Lee bounds under conditional monotonicity, which does not require trimming (and, therefore, quantiles) to be conditional on covariates. Suppose Assumptions (ref) and (ref)(b) hold. Let $\bar{\beta}^{\text{help}}_U$ and $\bar{\beta}^{\text{hurt}}_U$ be the basic Lee bounds of Section (ref), defined on $\mathcal{X}_{\text{help}}$ and $\mathcal{X}_{\text{hurt}}$, respectively. Focusing on the boundary $ \mathcal{X}_0$, define the treatment-control difference as $$\bar{\beta}^0 := \mathbb{E}[ Y \mid D=1, S=1, X \in \mathcal{X}_0] - \mathbb{E}[ Y \mid D=0, S=1, X \in \mathcal{X}_0].$$ Likewise, let

align[align omitted — 171 chars of source]

and

align[align omitted — 142 chars of source]

By construction, $S_{\text{help}} +S_{\text{hurt}}+ S_0 = \pi_{\text{AT}}$. Aggregating over the covariate space gives basic generalized bound:

align[align omitted — 185 chars of source]

If Assumption (ref) holds, $\bar{\beta}_U = \bar{\beta}^{\text{help}}_U$ reduces to basic Lee bound in (ref). This bound is a direct generalization of basic Lee bound to the case of conditional monotonicity.

lemma[Basic Generalized Bound] Suppose Assumptions (ref) and (ref)(b) hold. Then, $\bar{\beta}_U $ is a valid bound on $\beta_0$ obeying $$ \beta_0 \leq \beta_U \leq \bar{\beta}_U. $$

Unknown propensity score

In this Section, I consider the case when the propensity score is unknown and needs to be estimated. Let $\beta^{\text{help}}_U(x)$ be the conditional Lee bound defined in (ref), and let $\beta^{1\text{help}}_U(x)$ and $\beta^{0\text{help}}_U(x)$ be its first and second summand, respectively. Likewise, let $\beta^{\text{hurt}}_U(x) $ be the conditional Lee bound defined in (ref), and let $\beta^{1\text{hurt}}_U(x) $ and $ \beta^{0\text{hurt}}_U(x)$ be its first and second summand. Below, I describe the debiased moment function for the upper bound $\beta_U$ in (ref).

For $x \in \mathcal{X}_{\text{help}}$, define the Riesz representer function

align[align omitted — 168 chars of source]

For $x \in \mathcal{X}_{\text{hurt}}$, define

align[align omitted — 168 chars of source]

On the boundary, there is no trimming, and the two functions coincide $$ \Lambda_U(x) = \Lambda^{\text{help}}_U(x) =\Lambda^{\text{hurt}}_U(x), \quad x \in \mathcal{X}_0. $$ The bias correction term for the propensity score is

align[align omitted — 103 chars of source]

The debiased moment function takes the form

align[align omitted — 143 chars of source]

In particular, the propensity score correction term depends on the conditional trimmed mean function $\Lambda_U(x)$. In a low-dimensional smooth setting, the function $\Lambda_U(x)$ can be estimated by the local linear regression estimator proposed in olma2021nonparametric. In a high-dimensional sparse setting, one could use the automatic debiasing approach of chernozhukov2021automatic.

JobCorps revisited

Overview of JobCorps data

\paragraph{Data description. } LeeBound studies the effect of winning a lottery to attend JobCorps, a federal vocational and training program, on applicants' wages. In the mid-1990s, JobCorps used lottery-based admission to assess its effectiveness. The control group of $5, 977$ applicants was essentially embargoed from the program for three years, while the remaining applicants (the treated group) could enroll in JobCorps as usual. The sample consists of $9,145$ JobCorps applicants and has data on lottery outcome, hours worked and wages for 208 consecutive weeks after random assignment. In addition, the data contain educational attainment, employment, recruiting experiences, household composition, income, drug use, arrest records, and applicants' background information. These data were collected as part of a baseline interview, conducted by Mathematica Policy Research (MPR) shortly after randomization (Schochet). Lee has condensed this information to 28 covariates, including demographic characteristics, parental education, and income, wages, and hours of work at baseline (see Table (ref) in Appendix or Table 2, LeeBound). This section considers a richer specification, which includes frequency and type of drug use, arrest experiences, reasons for joining JobCorps, and occupation at baseline. I shall refer to these covariate choices as Lee's covariates (28) and All covariates (>1,000), respectively.

figure[figure omitted — 811 chars of source]

Testing unconditional monotonicity.

Baseline covariates can detect violations of unconditional monotonicity. If this assumption holds, the conditional average treatment effect on employment $\tau(x)$ in (ref) must be either non-positive or non-negative for all covariate values. Consequently, it cannot be the case that

align[align omitted — 114 chars of source]

for any $X$ taken to be the subset of all covariates.

The first exercise is to estimate $s(1,x)$ and $s(0,x)$ by a week-specific cross-sectional logistic regression. Let

align[align omitted — 83 chars of source]

where $\Lambda(\cdot) = \dfrac{\exp (\cdot)}{1 + \exp (\cdot)}$ is the logistic CDF, $X$ is a vector of baseline covariates that includes a constant and 28 covariates Lee selected, $D \cdot X$ is a vector of covariates interacted with treatment, and $\alpha$ and $\gamma$ are fixed vectors. Figures (ref) and (ref) show the results: the share of subjects with positive selection effect (Figure (ref), solid black line) and the average employment effect for subjects with $\tau(X)<0$ and $\tau(X)>0$ (Figure (ref)).

The second exercise is to test monotonicity without relying on a particular logistic specification

footnote{On p. 1085, Lee says “when $\Pr (S=1 \mid D=1) - \Pr (S=1 \mid D=0)=\mathbb{E}[ \tau(X) ] =0$, there is a limited test of whether monotonicity holds”. Lee considers a logistic regression of treatment $D=1$ (as outcome) on $X$ in the selected sample $S=1$. In contrast to Lee, this paper tests monotonicity using covariates, which does not require assuming $\mathbb{E}[ \tau(X) ] =0$. Furthermore, the covariate-based test (ref) may have higher power since it uses the full sample (and not only observations with $S=1$). }

. For each week, I select a small number of discrete covariates and partition the sample into discrete cells $C_j, \quad j \in \{1,2,\dots, J\}$, determined by covariate values. For example, one binary covariate corresponds to $J=2$ two cells. By monotonicity, the vector of cell-specific treatment-control differences in employment rates, $\mathbf{\mu}= (\mathbb{E}[\tau(X) | X \in C_j])_{j=1}^J$, must be non-negative:

align[align omitted — 74 chars of source]

The test statistic for the hypothesis in equation (ref) is

align[align omitted — 115 chars of source]

and the critical value is the self-normalized critical value of CCK.

figure[figure omitted — 1,059 chars of source]

Figure (ref) plots the fraction of subjects with a positive JobCorps effect on employment in each week (that is, the fraction of applicants in black dots in Figure (ref)). In the first weeks after random assignment, there is no evidence of a positive JobCorps effect on employment for any group. By the end of the second year (week 104), JobCorps increases employment for nearly $75\%$ of the individuals, and this fraction rises to $0.9$ by the end of the study period (week 208). This pattern is consistent with the JobCorps program description. While being enrolled in JobCorps, participants cannot hold a job, which is known as the lock-in effect (FloresLagunes). After finishing the program, JobCorps graduates may have gained employment skills that help them outperform the control group. However, the share of subjects with positive employment effect never reaches $100\%$, even four years after RA.

Figure (ref) shows the results of testing the inequality in (ref) for each week. The direction of the employment effect varies with socio-economic factors. For example, the applicants who received AFDC benefits during the 8 months before RA or who belonged to median income and yearly earnings groups experience a significantly positive ($p \leq 0.05$) employment effect at weeks $60$--$89$, although the average effect is significantly negative. As another example, the applicants who answered “1: Very important” to the question “How important was getting away from community on the scale from $1$ (very important) to $3$ (not important)?” ($\text{R\_Home}=1$) and who smoke marijuana or hashish a few times each months experience a significantly negative ($p\leq 0.05$) employment effect at week $117$--$152$ despite the average effect being positive. Finally, at week $153$--$186$, the average JobCorps effect is significantly negative for subjects whose most recent arrest occurred less than 12 months ago ($\text{MARRCAT=1}$), despite the average effect being positive.

Figure (ref) plots the average effects on employment rates across weeks. The effect is shown to be negative for weeks $1-89$ and positive thereafter. Remarkably, week 90 is also the only week whose average wage effect was found significant out of four weeks considered (LeeBound). For this reason, the rest of the section focuses on week 90 as the most interesting one.

table[table omitted — 1,534 chars of source]
landscape\captionof{table}{Bounds on the JobCorps effect on week 90 log wages under conditional monotonicity} \begin{tabularx}{\linewidth}{p{3cm} *{6}{Y}} \toprule \multicolumn{1}{p{3cm}}{ Covariates } & \multicolumn{2}{c}{28 (Lee)} & \multicolumn{2}{c}{>1000 (All)} & \multicolumn{2}{c}{43 (Lasso)} \\ \cmidrule(lr){2-3} \cmidrule(lr){4-5} \cmidrule(lr){6-7} & (1) & (2) & (3) & (4) & (5) & (6) \\ \midrule \multicolumn{1}{p{3cm}}{ Always-takers' share} & \multicolumn{2}{c}{0.44} & \multicolumn{2}{c}{0.44} & \multicolumn{2}{c}{0.44} \\ \midrule \multicolumn{1}{p{3cm}}{ Bounds} & [-0.060, 0.124] & [-0.072, 0.004] & [-0.021, 0.165] & [-0.088, 0.039] & [-0.180, 0.231] & [-0.055, 0.015] \\ \multicolumn{1}{p{3cm}}{ $95\%$ CR} & & (-0.106, 0.036) & & (-0.116, 0.071) & & (-0.096, 0.046) \\ \multicolumn{1}{p{3cm}}{ Width Reduction} & & 41.3% & & 68.3% & & 17.0% \\ \bottomrule \end{tabularx} \caption*{ Notes. Table shows estimated bounds in square brackets and the $95 \%$ confidence region for the identified set in parentheses. Columns (1), (3) and (5) report basic generalized bound in (ref) under Assumption (ref) with different choice of covariate sets. Columns (2), (4) and (6) report generalized Lee bound. First Stage: The employment equation is estimated using logistic regression (LR) with Lee's covariates (Columns (1)-(2)), post-Lasso-logistic regression (post-Lasso LR) with all covariates (Columns (3)-(4)), and LR with $43$ Lasso-selected covariates (Columns (5)-(6)). The wage quantile regression is estimated using quantile regression (QR) with Lee's covariates (Column (2)), $\ell_1$-QR with all covariates (Column (4)), and QR with $43$ covariates Lasso selected in (Column (6)). The automated penalty choice for $\ell_1$-QR is in equation (2.6) of belloni:11. Second stage. The always-takers' share is estimated as in Definition (ref). The bounds are estimated in Algorithm (ref). Computations use design weights. The sample size $N=9, 145$. The asymptotic variance for the $95 \%$ Confidence Region is based on $B=1,00$ bootstrap repetitions. Width reduction is defined as a ratio of sharp bounds' width in an even-numbered column $2j$ to its basic analog in $2j-1$ for $j \in \{1,2,3\}$. }

\paragraph{Results under unconditional monotonicity.} Table (ref) replicates Lee's results under unconditional monotonicity. The estimated effect on week-90 employment is $0.001$. Among treated individuals, approximately 99.9% of wages are attributed to always-takers. As a result, the basic Lee bounds collapse to a near-point estimate, which coincides with the observed treatment-control difference in log wages. Assuming JobCorps does not reduce employment, the wage effect lies between $4.8\%$ and $4.9\%$ on average. The discrete bounds in Column (2) are wider than the basic ones. The sharpness property of Lee bounds holds in population but may fail in sample, as cell-specific employment effects are positive in some cells and negative in others. Under unconditional monotonicity, such sign reversals can only arise from sampling noise; negative effects are truncated at zero.

\paragraph{Results under conditional monotonicity.} Table (ref) presents results under conditional monotonicity using different covariate sets. Columns (1)–(2) use only the original covariates from Lee’s analysis, relying on standard nonparametric assumptions. Columns (3)–(4) incorporate the full set of available covariates, imposing sparsity assumptions on employment (both columns) and wage (Column (4) only) equations. Columns (5)–(6) restrict attention to covariates selected by Lasso from the full set in Columns (3)–(4). The confidence region in Column (6) does not account for uncertainty due to covariate selection. The resulting bounds are not directly comparable since the assumptions underlying the different specifications are not nested.

Our empirical findings are as follows. First, the estimated share of always-takers is 44%. This estimate remains remarkably stable across different covariate specifications. Second, the upper bound on the wage effect ranges from 0.04% in Column (2) to 3.9% in Column (4). Notably, it does not overlap with the basic lower bound reported in Table (ref), presenting further evidence against unconditional monotonicity. Finally,  the generalized upper bound is substantially tighter than its basic counterpart. In terms of width, using covariates to tighten the bound reduces the width by approximately 40% in Columns (3)–(4) and nearly 80% in Columns (5)–(6). Assuming structure on the employment equation—such as sparsity or smoothness—is essential for recovering the direction of the selection effect. Further assuming smoothness or sparsity in the wage equation helps rule out implausibly large wage effects. Because the wage effect is close to zero, its sign remains unknown.

Figure (ref) reports the upper and lower bounds on the average log wage for the always-takers in the control state. The lower (upper) bound grows from $1.63$ ($1.92$) in week $5$ to $1.96$ ($1.96$) in week $208$. The bounds' width decreases from $0.3$ in week 14 to $0.01$ in week 208. The gap between the lower and the upper bound shrinks over time as the share of applicants with a positive employment effect, where the average control log wage is point-identified, increases.  The upward trend in the control wages suggests that evaluating JobCorps would have been very difficult without a randomized experiment, as one would need to explicitly model mean reversion in the baseline potential wage.

\paragraph{Conclusion. } Lee bounds are a popular empirical strategy for addressing post-randomization selection bias, assuming treatment's effect on selection has the same sign (unconditional monotonicity). This paper generalizes Lee bounds under conditional monotonicity, which allows the direction of selection effect to differ only with observed covariates. This generalization has proven especially useful for JobCorps job training program, where unconditional monotonicity is unlikely to hold. Relaxing conditional monotonicity is left for the future work.

figure[figure omitted — 766 chars of source]