EconBase
← Back to paper

Treatment Evaluation at the Intensive and Extensive Margins

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

113,621 characters · 28 sections · 111 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Treatment Evaluation at the Intensive and Extensive Margins

titlepage\thispagestyle{empty} \begin{abstract} \singlespacing This paper provides a solution to the evaluation of treatment effects in selective samples when neither instruments nor parametric assumptions are available. We provide sharp bounds for average treatment effects under a conditional monotonicity assumption for all principal strata, i.e. units characterizing the complete intensive and extensive margins. Most importantly, we allow for a large share of units whose selection is indifferent to treatment, e.g. due to non-compliance. The existence of such a population is crucially tied to the regularity of sharp population bounds and thus conventional asymptotic inference for methods such as Lee bounds can be misleading. It can be solved using smoothed outer identification regions for inference. We provide semiparametrically efficient debiased machine learning estimators for both regular and smooth bounds that can accommodate high-dimensional covariates and flexible functional forms. Our study of active labor market policy reveals the empirical prevalence of the aforementioned indifference population and supports results from previous impact analysis under much weaker assumptions. \end{abstract} Keywords: Double/debiased machine learning; Lee bounds; Partial identification; Principal strata; Sample selection \\ JEL classification: C13, C14, C21

\setcounter{page}{1}

Introduction

This paper deals with the evaluation of causal effects of a binary treatment when the outcome is only selectively observed and no instruments are available. Such missing outcome data or sample selection is ubiquitous and a threat to internal validity heckman1974shadow,heckman1979sample. Typical examples include outcome data such as wages that are only observed if units are employed lee2009training, missing survey responses bernhardt2024howdoes, or attrition zhang2003estimation. In the context of impact evaluation, a treatment can often change the sample selection status for some units. For example, if the treatment is an active labor market policy such as job training, it likely affects the probability of employment for some participants, e.g. due to accumulation of human capital (positive) or opportunity costs to job search (negative). For some units, however, the policy might be irrelevant with regards to employment or sample selection, but still relevant in terms of their earnings or other outcomes. Hence, there are potentially heterogeneous units at the intensive and extensive margins. These units can be classified within principal strata defined by their potential selection status.

We consider set identification, estimation, and inference for all principal strata, encompassing the complete intensive and extensive margins, under a weak conditional monotonicity assumption. Weak monotonicity allows for the simultaneous presence of always-takers, compliers, defiers, and never-takers.\footnote{In contrast to the instrumental variables literature, principal strata here refer to potential selection statuses when treatment is exogenously set to zero and one respectively. See Section (ref).}

The main challenge is the presence of units whose selection probabilities are unaffected by the treatment. The existence of such units is most apparent when the treatment is subject to additional non-compliance after assignment and intent-to-treat effects are analyzed. In this case, units who have information or preferences that lead them to not participate after treatment assignment will likely also not change their selection behavior, e.g. selection into employment, as a consequence. Further, there are policies that do not boost or hinder employability for some units, e.g. due to a lack of signal for their potential employers in the short run, but increase human capital boosting productivity and earnings. Similarly, in field experiments where responses cannot be enforced, an intervention might only affect response probabilities to follow-up surveys for some units.

From a discrete choice perspective, our setup admits units who are indifferent to the treatment at any conceivable level of their private information. From a model selection perspective, unaffected units are equivalent to a sparsity constraint in the selection equation. Modern model selection such as $L_1$-regularization, subset selection, deep neural networks, boosted trees, and random forests are able to recover such potentially sparse structures in the selection equation.

Figure (ref) contains an example from the job training program analyzed in this paper (Job Corps). It provides the histogram of the estimated relative selection probabilities with and without assignment to training using boosted trees. A value of $1.00$ indicates that assignment does not predict employment status for a given unit. In this case we have that 2616 out of 9415 units stacking up at an exact[!] numerical one. This pattern is striking and persistent across evaluation periods and estimation methods, see Sections (ref) and (ref). In this study, there is significant non-compliance. Thus a large presence of estimated unaffected units seems plausible.

figure[figure omitted — 569 chars of source]

In general, sample selection with unaffected units yield mixture distributions for relative conditional selection probabilities obtained in a first stage with point mass at $1.00$. While without harm from an identification perspective, it directly translates into a fundamental problem for estimation and inference. In particular, sharp population bounds are non-smooth functionals for which no uniformly unbiased estimator exists, i.e. the problem is irregular by construction hirano2012impossibility. More concretely, we show that not having a mixture is in fact necessary and sufficient for the population bound to be pathwise differentiable in the sense of bickel1993efficient. This means that standard errors and statistical inference obtained from methods such as Lee bounds with “naive” asymptotics lee2009training are invalid in the presence of unaffected units.

We propose a solution to the problem involving repeated smoothing steps that effectively act as covers for the sharp identified sets. The resulting smooth outer identification regions can be efficiently estimated at the parametric rate and always yield valid confidence regions for the true parameter of interest as the direction of any bias is known. Thus, the degree of smoothing can be chosen to maximize power in finite samples. Under regularity, the smooth bounds can converge to the sharp identified set. On a high level, our results reveal that there is an identification-precision trade-off: Aiming for an outer identification region with high precision over a sharp region with low precision can significantly improve inference in finite samples.

For both the regular and the smoothed irregular case, we provide efficient influence functions and corresponding estimators using debiased machine learning. Under simple high level conditions, the estimators are asymptotically normal and reach their respective semiparametric efficiency bounds. The influence function for the intensive margin is equal to the one implied by the moment estimator in semenova2023generalized under unknown propensity scores and no moment selection. To our knowledge, this is the first paper to provide efficient influence functions that can be used for semiparametric estimation of compliers, defiers, and the combined extensive margin bounds in the sample selection model without exclusion. In contrast to any regular bounds, the smooth counterparts require weaker convergence assumptions for the involved nuisance parameters and, even in the irregular case, no margin condition that controls the distribution of selection effects.

We also explore the role of the propensity score for efficiency, extending the results in hahn1998role to sample selection without exclusion. The efficiency bound for the principal strata treatment effect bounds do not depend on the knowledge of the propensity score. We quantify the efficiency gap when using the moment functions under knowledge of the propensity score as suggested by semenova2023generalized and heiler2024heterogeneous. These estimators are generally inefficient except in knife-edge cases. However, there is a trade-off from a practical perspective: Analogously to inverse probability weighting (IPW) estimation of average treatment effects, these estimators do not require conditional (truncated) outcome means as additional nuisance inputs. Thus, depending on the difficulty in estimating the latter, using these simpler moment functions might be preferable if propensity scores are known. All aforementioned properties regarding (ir)regularity and smoothing also directly apply to these inefficient moment functions and their smoothed counterparts.

Monte Carlo simulations suggest that, in irregular designs, inference using the smooth methods compares favorably to the alternatives from the literature that either ignore the irregularity of the problem or rely on trimming heiler2024heterogeneous, or moment selection/switching methods andrews2010inference,semenova2023generalized.

We extend the theory to other quantile-trimmed bounds obtained under additional stochastic dominance assumptions zhang2003estimation,huber2015sharp. The components of our influence functions can also be used to construct heterogeneous treatment effect bounds in the sense of heiler2024heterogeneous for any principal strata.

The empirical study is a comprehensive re-evaluation of the National Job Corps Study burghardt1999national,schochet2008does. We demonstrate the empirical relevance of flexible estimation via machine learning methods, in particular with regards to sparsity in the selection equation. Aggregate results of the new semiparametrically efficient smoothing bounds with optimized learners rule out moderate to large negative impact of Job Corps on hourly earnings at the end of the evaluation period. The results over various time periods are surprisingly close to the more restrictive impact analysis by lee2009training that relies on a stronger monotonicity assumption rejected by the data semenova2023generalized.

The paper is organized as follows. Section (ref) discusses the relevant literature. Section (ref) introduces the nonparametric sample selection model and provides an in-depth discussion of conditional monotonicity. Section (ref) presents the sharp and smooth outer identification regions for all principal strata. Section (ref) provides the results on pathwise differentiability. Section (ref) contains the assumptions for debiased machine learning estimation and asymptotic inference as well as the efficiency gap under known propensity scores. Section (ref) extends the methodology to stochastic dominance bounds and heterogeneous effect bounds. Section (ref) presents the empirical study. Section (ref) concludes.

Literature

There is a large literature on sample selection models without exclusion under restrictive parametric assumptions heckman1979sample,staub2014tobit. Bounds for causal effects under varying weaker assumptions also have a rich tradition, see molinari2020microeconometrics for a comprehensive survey. We focus on research that does not rely on the use of additional exclusion restrictions or instruments and is connected to the monotonicity assumption. We follow a principal stratification approach that considers identification of causal effects for all latent subgroups characterized by their potential selection behavior as a function of a binary treatment frangakis2002principal. This circumvents comparing systematically different groups when conditioning on realized selection behavior.

Bounds on causal effects for the intensive margin or always-takers under monotonicity have first been introduced by zhang2003estimation, see also lee2009training for an extension to a form of conditional monotonicity that excludes units whose selection behavior is unaffected by the treatment. Such “Lee” or “Zhang-Rubin-Lee” bounds andersen2023guide are defined by conditional expectations trimmed at conditional quantiles. imai2008sharp shows the sharpness of such trimming bounds. semenova2023generalized introduces orthogonal moments for estimation of intensive margin bounds that can be used in high-dimensional setups, see also olma2021nonparametric for nonparametric estimation of truncated conditional expectations. huber2015sharp derive bounds for compliers under strong monotonicity. We extend their ideas allowing for compliers and defiers simultaneously. honore2020selection,honore2022sample also discuss bounds in sample selection models without exclusion restrictions under additional parametric assumptions. Our bounds do not require any parametric structure. bartalotti2021identifying provide bounds using monotonicity and stochastic dominance within a marginal treatment effect (MTE) framework. They focus on strong monotonicity only and do not provide estimation or inference theory. heiler2024heterogeneous considers heterogeneous intensive margin bounds and misspecification robust inference under a strong margin assumption effectively ruling out subgroups for which selection probabilities do not depend on the treatment. Our new moment functions can also be used in conjunction with the approach by heiler2024heterogeneous to provide heterogeneous effect bounds for any margin. okamoto2023bounds discusses sensitivity analysis of intensive margin and MTE bounds under partial violation of the conventional monotonicity assumption. His stochastic monotonicity assumption is distinct from our approach as it assumes strong monotonicity for a limited share of the population instead of a weaker form of monotonicity. The use of intensive margin bounds has also been advocated by chen2023logs when encountering dependent variables with (many) zeroes. All these papers consider either strong monotonicity or the restrictive version of weak monotonicity lee2009training, with semenova2023generalized being a notable exception.

Monotonicity also plays a crucial role for instrumental variable (IV) based methods imbens1994identification,angrist1995two,angrist2000fish,abadie2003semiparametric,heckman2007econometric2,sloczynski2020should,heiler2022efficient. Similar to the sample selection setup, it is an assumption on latent subtypes in the population to react towards a change in the instrument weakly in a particular direction in terms of their treatment selection. It rationalizes the same choices as a structural model for treatment selection that is additive separable in observables and unobservables both in the population vytlacil2002independence as well as numerically if estimated nonparametrically kline2019heckits. This also applies to our setup. For nonparametric identification of local average treatment effects or other MTE-type parameters, the instrument has to be able to move certain units in terms of their treatment choice (first stage). Combining compliers and defiers for IV using a weak monotonicity assumption has been considered by kolesar2013estimation and sloczynski2020should. The key distinction to the IV setup is that we are interested in the causal effect of the treatment accounting for endogenous selection. Thus, our treatment takes on the role of the instrument while the selection is fully endogenous as the treatment in the IV case. The weak monotonicity assumption posited in this paper allows for the presence of all latent subtypes in the population but limits their presence within ex-ante unknown partitions of the covariate space. This can also by motivated by a structural selection model with mixed indices within these partitions. The IV framework does not allow for nonparametric identification in any population without movers (no first stage). We explicitly want to allow for units whose selection is not affected by the treatment. This can then generates the mixture distribution for the “first stage” relative conditional selection probabilities with point mass at one.

Such mixtures imply that the (generalized) ZR-Lee bounds are no longer smooth in the underlying data distribution and thus there exist no estimator sequence that is locally asymptotically unbiased in the sense of hirano2012impossibility. We construct a particular cover for the identification region that circumvents the non-regularity of the underlying functional of interest. Our identified set has non-empty interior and a zero duality gap as discussed in kaido2014asymptotically. Thus, its semiparametric efficiency bound can be characterized and $\sqrt{n}$-consistent regular estimators of the support function exist in a uniform sense. As a by-product of the strict convexity of the smooth identified set, confidence intervals for the effects (not bounds) can be made more precise using modified critical values imbens2004confidence. Under a strong margin assumption, our bounds converge converge to the standard sharp ZR-Lee bounds at a known rate.

Our estimands are ratios of a weighted conditional expectations normalized by their share. For the intensive margin, estimating the share is related to estimating an expected conditional outcome under unconfoundedness when setting the treatment status to the best conditional mean. For the latter in isolation, luedtke2016statistical show that for the estimand to be pathwise differentiable, it is necessary and sufficient that units which are indifferent between treatments have an exceptional law with zero conditional variance under both regimes. We extend their approach to a class of ratio estimands and show that, for the principal strata bounds, the point mass phenomenon is both necessary and sufficient for irregularity, i.e. there are no exceptional laws as the latter would imply a violation of the overlap assumption.

levis2023covariateassisted also suggest smoothing for obtaining regular estimators of balke1997bounds bounds in the presence of covariates. They propose confidence intervals with a worst-case bias correction arising from the particular approximation. Unlike their approach, our smoothing will always widen the identified set and thus one can navigate the identification-precision trade-off without the need for a worst-case bias component.

semenova2023generalized uses the same weak monotonicity assumption as this paper but focuses exclusively on always-takers and does not address regularity and efficiency. To address the problem of misclassifying units, she uses a weak margin assumption and a shrinkage/moment-switching approach. The latter selects between different moment functions based on the estimated selection probabilities with a vanishing sequence of shrinkage parameters that obey a rate condition. There is no guidance on how to choose this parameter in finite samples. Moreover, while asymptotically valid in handling potential misclassification, this moment selection effectively narrows the identified set leading to potential undercoverage in finite samples. Our approach guarantees at least nominal coverage and can be optimized with respect to power. For always-taker bounds, heiler2024heterogeneous suggests to round trimming threshold towards the closest value on a grid. While this approach similarly yields an outer identification region in the context of relative selection probabilities very close to one, it is unclear how to use in the case of a point mass for units whose selection is unaffected by the treatment, i.e. an exact one, as rounding could be done in two directions leading to different impact of misclassification errors.

In other contexts, lee2021bounding and pakel2023bounds also note that there can be a trade-off between identification and precision. They argue that confidence sets obtained from outer bounds can well be tighter compared to using sharp bounds. A similar intuition applies for our smooth outer identification region. For a data combination problem, dHault2022partially also consider outer bounds whose conservativeness arise from regularization.

We characterize the efficient influence functions and semiparametric efficiency bounds with and without knowledge of the propensity score. Our results generalize hahn1998role to sample selection under weak conditional monotonicity and either (i) regular sharp bounds or (ii) regular smooth bounds that cover the irregular sharp bounds. We also contribute and make use of the expanding literature on the use of debiased machine learning methods for estimation of causal effects in economics chernozhukov2018double,chernozhukov2022locally and its combination with partial identification heiler2024heterogeneous,semenova2023set,semenova2023generalized. More specifically, we show that the particularities of the relative selection probabilities and the rate at which they can be learned from data crucially interact with the (ir)regularity of the parameters of interest.

Model and Monotonicity

Assume we observe iid data $W = (YS,S,D,X)'$ where $S \in \{0,1\}$ is a selection variable indicating whether an outcome is observed or not. $Y \in \mathcal{Y}$ is the partially unobserved outcome of interest. $D \in \{0,1\}$ is a binary treatment of interest. $X \in \mathcal{X}$ is a vector of predetermined covariates. We are interested in the causal effect of the treatment on the outcome defined in terms of potential outcomes and potential selection indicators. The realized but partially unobserved outcome is defined as

align[align omitted — 39 chars of source]

Selection is connected to potential selection indicators as

align[align omitted — 39 chars of source]
figure[figure omitted — 1,564 chars of source]

Figure (ref) depicts a prototypical sample selection model with the relevant exclusion restrictions. Such a model implies the following conditional independence assumption for the treatment which we maintain throughout:

ass[Conditional Independence] $$D \perp\!\!\!\perp Y(d),S(d)~|~X=x \textit{ for all } x \in \mathcal{X} \textit{ and } d=0,1,$$

where $\perp\!\!\!\perp$ denotes statistical independence. Thus, treatment is assumed to be exogenous conditional on $X$. However, selection is left essentially unrestricted, i.e. it can be fully endogenous or exogenous. In particular, there are no exclusion restrictions available for $S$, i.e. we cannot use instruments or similar to account for endogenous selection.

The potential selection indicators define principal strata which represent how the treatment affects selection behavior. Table (ref) contains all strata using the nomenclature from the instrumental variables literature.

table[table omitted — 370 chars of source]

To proceed, we postulate the existence of an unknown but identified partitioning of the covariate space $\mathcal{X}$ on which a weak monotonicity assumption applies to these principal strata\footnote{Note that this definition differs slightly from the conditional monotonicity assumption of semenova2023generalized which does not yield a unique partitioning.}.

ass[Weak/Conditional Monotonicity] Assume there are subsets of the covariate space $\tilde{\mathcal{X}}^+,\tilde{\mathcal{X}}^-\subseteq \mathcal{X}$ such that \begin{align*} S(1) \geq S(0) & if x \in \tilde{\mathcal{X}}^+, \\ S(1) \leq S(0) & if x \in \tilde{\mathcal{X}}^-. \end{align*}

Assumption (ref) yields a distinct partitioning of $\mathcal{X} = {\mathcal{X}}^+ \cup {\mathcal{X}}^- \cup {\mathcal{X}}^0$ where

align*[align* omitted — 226 chars of source]

The presence of a potentially non-empty $\mathcal{X}^0$, i.e. $P(\mathcal{X}^0) > 0$, will be crucial in what follows. Substantially, weak monotonicity allows for the presence of all principal strata including defiers. However, it is a partial restriction of types within partitions. Table (ref) contains the mixture of types in the different partitions as well the population.

table[table omitted — 1,520 chars of source]

For example, in partition $\mathcal{X}^0$, the population consists only of always-takers and never-takers while on $\mathcal{X}^+$ and $\mathcal{X}^-$ only defiers or compliers are ruled out, respectively. This assumption seems most relevant if we have some discrete covariates and/or continuous variable that affect selection behavior discontinuously e.g. by threshold effects. Assumption (ref) is also consistent with a single-index structure in the selection equation that allows for sparsity with respect to the treatment on some parts of the covariate space.

\paragraph{Example (Single Index Model):} Consider a simple semiparametric selection model with treatment and selection that are additive separable in observables and unobservables

align[align omitted — 143 chars of source]

where $U_D \perp\!\!\!\perp U_Y,U_S$ conditional on $X=x$ then implies conditional independence as in Assumption (ref). Weak monotonicity as in Assumption (ref) here is implied by an unrestricted $g(D,X)$. This can generate a mixed index structure

align[align omitted — 195 chars of source]

where $g_+(1,x) > g_+(0,x)$ and $g_-(0,x) > g_-(1,x)$. The fact that $g_0(x)$ does not depend on $d$ can be seen as a sparsity assumption within a partition of the covariate space. For example, consider the case of a simple parametric index in the selection equation with $X = (X_1,X_2)'$ where $X_1$ is continuous and $X_2$ discrete with support points $\{-1,0,1\}$

align[align omitted — 143 chars of source]

Here $X_2$ can separate the monotonicity types. In particular, $\gamma_2 < 0 < \gamma_3$ or $\gamma_2 > 0 > \gamma_3$ imply conditional monotonicity. If the signs of $\gamma_2$ and $\gamma_3$ are identical, then either $\mathcal{X}^+ = \{ \emptyset\} $ or $\mathcal{X}^- = \{ \emptyset\}$, implying the conventional unconditional or strong monotonicity $P(S(1) \geq S(0)) = 1$.

Whenever $P(\mathcal{X}^0) > 0$, conditional monotonicity and independence generate a mixture distribution for the potential or relative conditional selection probabilities

align[align omitted — 100 chars of source]

In particular, for $\mathcal{X}^+$ and $\mathcal{X}^-$ we have that $p_0(x) < 1$ and $p_0(x) > 1$, respectively, while there is potentially a point mass for which $p_0(x) = 1$ corresponding to $P(\mathcal{X}^0)>0$. Thus, conditional monotonicity allows for mixture distributions which are not necessarily smooth in $p_0(x)$. We seek to conduct inference on causal effects for all principal strata under such mixtures.

figure[figure omitted — 1,875 chars of source]

Modern model selection methods such as $L_1$-regularization, subset selection, neural networks, boosting, and random forests are able to recover and produce potentially sparse structures in the selection equation. We present an example from relative employment probabilities and a job training treatment in Figure (ref). It contains the predicted relative selection probabilities from Section (ref) using five different estimation methods: (a) logistic regression, (b) logistic regression with fully interacted treatment, (c) gradient boosted trees (XGBoost), (d) neural network, and (e) random forest, see Section (ref) and Appendix (ref) for more details.

The methods capable of sparsity yield a clear mixture pattern for the estimated distribution of $p_0(x)$. In particular, (c), (d), and (e) produce an exact share of numerical ones with up to $45.16\%$ of the sample estimates. (a) imposes strong monotonicity and cannot produce exact ones in finite samples by construction. We suspect that (a) is most likely to be misspecified. (b) has relevant density around one, but also generates overly extreme values for the relative selection probabilities. The sparse methods agree on around 45% -- 60% of the close or equal to one classifications. This qualitative pattern can be observed along a long sequence of post-treatment periods. Over time there is an average distribution shift suggesting more positive employment effects. The sparsity pattern, however, is pervasive. We take this as a signal that the mixture distributions are likely reflecting of or a good approximation to an underlying structure and are not only spurious by-products of particular estimation algorithms motivating Assumption (ref) with $P(\mathcal{X}^0) > 0$.

\paragraph{Remark:} On a cautionary note, empirical distributions of individual predictions such as in Figure (ref) should not be over-interpreted. On one hand, we can have clear theoretical reasons why true $p_0(x)$ ought to be one for many units, e.g. units that did not intend to join the program are likely to be unaffected by a non-binding assignment to treatment. On the other hand, these are just finite-sample estimates from models that are all likely misspecified to a smaller or larger extend. Extreme values that suggest economically unreasonable unit selection effects are likely driven by a high-variance model. We can take a more agnostic perspective with regards to the evidence provided by the empirical distribution of $\hat{p}_0(x)$. In particular, in high dimensions or with complicated functional forms, convergence to true relative selection probabilities is hard to achieve for any model. In finite samples, producing a sparse representation by removing the impact of the treatment in parts of the covariate space could just be a method's solution to optimize fit. Methods for predicting conditional selection probabilities are not necessarily designed to best predict the implied relative $p_0(x)$. However, if a learner performs best among conventional performance metrics for probabilities and does so by producing sparsity with respect the treatment, we treat it as a potentially better approximation and need to deal with the point mass regardless of whether this is a true or approximate model.

Identification of Sharp and Smooth Bounds

Sharp Bounds

The General Case

We denote $a = O(b)$ and $a=O_p(b)$ as $a\lesssim b$ and $a\lesssim_P b$, respectively. We first present identification of conditional causal effect bounds for any principal stratum, i.e. parameters of the form

align[align omitted — 71 chars of source]

This covers all relevant components for any of the principal strata and margin effects.\footnote{Trivially, for a principal stratum effect to be well defined, we require that the corresponding population share is non-zero. We assume this throughout.} Let $q_d(u,x)$ be the $u$-quantile of $Y$ conditional on $S=1, D=d, X=x$. For $d\in\{0,1\}$, we define

align[align omitted — 127 chars of source]

which yields

align[align omitted — 70 chars of source]

In line with the literature on monotonicity bounds, we assume a continuous outcome.

ass[Continuity] For all $x\in \mathcal{X}$ and $d\in \{0,1\}$, the conditional outcome distribution $P(Y \leq y|S=1,D=d,X=x)$ is continuous.

Moreover, there must be comparable units in terms of covariates in both selected treatment groups.

ass[Multiple Overlap] Let $m(x) = P(D=1|X=x)$ and $s(d,x) = P(S=1|D=d,X=x)$. For all $x\in \mathcal{X}$ and any $d \in \{0,1\}$, we have that $0 < m(x) < 1$ and $0 < s(d,x)< 1$.

Assumptions (ref) and (ref) imply the specific bounds in terms of truncated means of observed conditional distributions. Assumptions (ref) and (ref) ensure that the corresponding conditioning sets are non-empty and the relevant truncation quantiles are unique. In particular, we obtain the following identification result for the sharp upper and lower bounds for all principal strata.

prop[Identification] Let $\mathcal{Y}_d(x)$ be the support of $Y(d)$ conditional on $X=x$ with $\overline{y}_d(x)$ and $\underline{y}_d(x)$ its respective infimum and supremum for both $d\in\{0,1\}$. Under Assumptions (ref), (ref), (ref), and (ref), the sharp lower bounds for the principal strata are given by \begin{align*} \beta_L(x,s_0,s_1) = \begin{cases} \beta_{1,1}(x,\min\{p_0(x),1\}) - \beta_{0,0}(x,1-\min\{1/p_0(x),1\}) & if s_0 = 1, s_1 = 1 \\ \beta_{1,1}(x,1-\min\{p_0(x),1\}) - \beta_{0,0}(x,\min\{1/p_0(x),1\}) & if s_0 \neq s_1 \\ {y}_1(x) - \overline{{y}}_0(x) & if s_0 = 0, s_1 = 0, \end{cases} \end{align*} and the sharp upper bounds by \begin{align*} \beta_U(x,s_0,s_1) = \begin{cases} \beta_{0,1}(x,1-\min\{p_0(x),1\}) - \beta_{1,0}(x,\min\{1/p_0(x),1\}) & if s_0 = 1, s_1 = 1 \\ \beta_{0,1}(x,\min\{p_0(x),1\}) - \beta_{1,0}(x,1-\min\{1/p_0(x),1\}) & if s_0 \neq s_1 \\ \overline{{y}}_1(x) - \underline{{y}}_0(x) &\text{ if } s_0 = 0, s_1 = 0. \end{cases} \end{align*}

For never-takers, all lower and upper bounds are generally uninformative due to never observing them as part of a selected group. For defiers and compliers, the bounds are only informative when at least part of the conditional support is bounded, while for always-takers bounds are also informative without any support condition.

Unconditional bounds for any principal stratum are obtained by integrating the conditional bounds with respect to stratum-specific covariate distributions

align[align omitted — 107 chars of source]

for $B\in\{L,U\}$, where

align*[align* omitted — 502 chars of source]

Example I: Intensive Margin Bounds

The upper and lower conditional intensive margin or always-taker bounds correspond to the ones given by lee2009training and semenova2023generalized. Namely, we have that

equation[equation omitted — 68 chars of source]

where

align*[align* omitted — 480 chars of source]

Similarly, the conditional upper bounds are given by

equation[equation omitted — 69 chars of source]

where

align*[align* omitted — 449 chars of source]

The unconditional bounds for the effect at the intensive margin can then be recovered as

align[align omitted — 188 chars of source]

Example II: Extensive Margin Bounds

The extensive margin is the combination of compliers and defiers where $s_0 \neq s_1$. Thus, for $B\in\{L,U\},$ the extensive margin effect bounds can be written as {

align[align omitted — 68 chars of source]

} where

align*[align* omitted — 331 chars of source]

Combining compliers and defiers from Proposition (ref) then yields the conditional lower bound

align*[align* omitted — 232 chars of source]

and

align*[align* omitted — 233 chars of source]

The conditional upper bound $\beta_U(x,em)$ is found analogously from from Proposition (ref). As a result, the unconditional bounds at the extensive margin are given by

align[align omitted — 98 chars of source]

Smooth Bounds

The bounds in Proposition (ref) are sharp. However, they, as well as the strata densities, contain a variety of non-differentiable components. This poses a challenge to statistical inference. We demonstrate the specific regularity problem and its consequences for inference in Section (ref). We now introduce the smooth bounds that provide an outer identification region to circumvent these problems. In particular, the identification region is constructed to eliminate the impact of misclassifying different monotonicity types while still remaining close to the sharp identified set. We focus on the intensive margin bounds for simplicity but the method works equivalently for all principal strata. More specifically, the form of the bounds in Proposition (ref) reveal two sources where classification of the partition will enter, (i) conditional effects bounds $\beta_B(x,1,1)$, and (ii) conditional always-taker share $\min\{s(0,x),s(1,x)\}$. We suggest to apply repeated, asymmetric smoothing to both components. For (i), consider first the conditional bound $\beta_L(x,1,1)$. Note that, for any $x \in \mathcal{X}$, the following parameter is a valid lower bound for $\beta_L(x,1,1)$

align[align omitted — 276 chars of source]

where $g_{1,h}(z)$ is a smooth monotonic function indexed by a smoothing parameter $h \in \mathcal{H}_n \subset \mathbb{R}^+$ such that $0 \leq g_{1,h}(z) \leq \min\{z,1\}.$ This ensures that

equation[equation omitted — 57 chars of source]

Next, consider that the smoothing of the conditional always-taker share is given by $\min\{s(0,x),s(1,x)\} = s(1,x) \min\{p_0(x),1\}$. We obtain lower and upper bounds using smooth functions $g_{1,h}(z)$ and $g_{3,h}(z)$ such that $g_{1,h}(z) \leq \min\{z,1\} \leq g_{3,h}(z).$ In which direction to smooth, depends on the sign of ${\beta}_{L,h}(x,1,1)$. We decompose this into the difference of two non-negative terms via $z=\max\{z,0\}-\max\{-z,0\}.$ We approximate these two functions with a smoothing function $g_{4,h}(z)\leq \max\{z,0\},$ and the earlier $g_{2,h}(-z) \geq \max\{-z,0\}.$ The unconditional smooth lower effect bound is then given by

align[align omitted — 261 chars of source]

The aforementioned relations together yield

equation[equation omitted — 51 chars of source]

Analogously, the conditional smooth upper bound is given by

align[align omitted — 132 chars of source]

which obeys the inequality $\beta_{U,h}(x,1,1) \geq \beta_U(x,1,1)$, and the unconditional smooth upper bound is defined as

align[align omitted — 267 chars of source]

which analogously satisfies the reversed inequality

equation[equation omitted — 51 chars of source]

Next, we propose a class of functions that satisfy the imposed conditions and show that the distance between smooth and sharp unconditional bounds is negligible as $h \to 0$ with bounded support.

ass[Outcome Regularity I] For $d\in\{0,1\}$, the outcome distribution $P(Y \leq y| S=1, D=d, X=x)$ has uniformly bounded support , i.e. $\sup_{x\in \mathcal{X}} (|\underline{y}_d(x)|+|\overline{y}_d(x)|)<\infty.$

We also impose the following assumption about approximation functions.

ass[Approximation Functions] Approximation functions $g_{i,h}, i=1, \ldots, 4$ satisfy for all $z \in \mathbb{R}$ \begin{align*} g_{1,h}(z) \leq \min\{z,1\} \leq g_{3,h}(z) \quad and \quad g_{4,h}(z)\leq \max\{z,0\} \leq g_{2,h}(z), \end{align*} and for each $i=1, \ldots, 4$ \begin{equation*} \sup_{z \in \mathbb{R}} |g_{i,h}(z)-g_i(z)| \lesssim h, \end{equation*} where $g_1(z)=g_3(z)=\min\{z,1\}$ and $g_2(z)=g_4(z)=\max\{z,0\}.$ In addition, $g_{1,h}(z) \geq 0$ for all $z \in \mathbb{R}$ and $h \in \mathcal{H}_n.$\footnote{All quantile trimming thresholds are well defined as long as $\inf_{x \in \mathcal{X}} g_{1,h}(p_0(x)) \geq 0$ and $\inf_{x \in \mathcal{X}} g_{1,h}(1/p_0(x)) \geq 0.$ Given the strong multiple overlap assumption, one can always find an upper bound for $h$ such that this applies. These large-$h$ values are not relevant from a practical perspective and we consider $h$ to be below such a ceiling in what follows.}

There exist many functions with these properties. Figure (ref) depicts an example using a LogSumExp function for smoothing via $g_{1,h}(z)$.

figure[figure omitted — 689 chars of source]

We obtain the following Theorem.

theoremLet $\mathcal{P}$ be the set of probability measures satisfying Assumptions (ref), (ref), (ref), (ref), (ref), and (ref). Denote $\underline{g}_n = \inf_x g_{1,h}(p_0(x))$. Then, for $B \in \{L, U\},$ and $(s_0, s_1) \neq (0,0),$ we have that \begin{equation*} \sup_{P \in \mathcal{P}} |\beta_{B,h}(s_0,s_1) -\beta_B(s_0, s_1)|\lesssim \frac{h}{g_n^2}. \end{equation*}

Theorem (ref) gives the approximation error of the smooth bounds. Under the strong overlap condition, $\underline{g}_n$ is bounded away from zero and thus the difference decays linearly in $h$. Without strong overlap, there is an implicit restriction on the rate at which the smooth relative probabilities $g_{1,h}(p_0(x))$ can approach zero.

Regularity and Semiparametric Efficiency Bounds

In this section, we present our main results on regularity and semiparametric efficiency. Again, we focus on the lower bound for the intensive margin for simplicity. Table (ref) provides the influence function for the always-taker lower bound and its smoothed counterpart. All other influence functions are in Appendix (ref). We obtain the following theorem.

ass[Outcome Regularity II] For all $x\in \mathcal{X}$ and $d \in \{0,1\},$ the conditional outcome distribution $P(Y \leq y| S=1, D=d, X=x)$ has finite support on $[\underline{y}_d, \overline{y}_d]$ and a continuous density $f(y|X=x, D=d, S=1)$ bounded from below and above.
ass[Strong Multiple Overlap] There exist constants $\underline{m},\underline{s} \in (0,1/2)$ such that \begin{align*} m < \inf_{x\in\mathcal{X}} m(x) \leq \sup_{x\in\mathcal{X}} m(x) < 1- m, \end{align*}\begin{align*} s < \inf_{x\in\mathcal{X},d\in{\{0,1\}}} s(d,x) \leq \sup_{x\in\mathcal{X},d\in{\{0,1\}}} s(d,x) < 1- s. \end{align*}

Assumptions (ref) and (ref) are stronger versions of continuity and overlap in Assumptions (ref) and (ref), respectively. Without strong overlap, there is an additional irregularity problem for the population bounds equivalently to irregular identification of average treatment effects under unconfoundedness with many extreme propensities scores khan2010irregular,HEILER2021valid. We abstract from such issues in this paper to focus on the irregularity obtained from $P(\mathcal{X}^0) > 0$.

theoremSuppose that Assumptions (ref), (ref), (ref), and (ref) hold. For $B \in \{L, U \},$ $\beta_B(1,1)$, $\beta_B(0,1)$, $\beta_B(0,1),$ $\beta_B(\mbox{em})$ are pathwise differentiable if and only if $P(\mathcal{X}^0)=0.$

Theorem (ref) characterizes that the necessary and sufficient condition for the unconditional lower bound introduced in (ref) to be pathwise differentiable is the absence of the point mass on the set $\mathcal{X}^0$, i.e. $P(\mathcal{X}^0)=0$.\footnote{ When $P(\mathcal{X}^0)>0$, the result in Theorem (ref) that $\beta_L(1,1)$ is not pathwise differentiable is related but different to luedtke2016statistical. Their target is essentially the denominator of $\beta_L(1,1)$ in (ref). However, divergence of the derivative of denominator and/or divergence of the numerator separately does not necessarily imply divergence of the ratio. Moreover, the numerator includes the conditional treatment effect bounds $\beta_L(1,1,x)$ and consequently conditional quantiles $q_d(u,x)$ which changes the analysis. } We also present the semiparametric efficiency bound for $\beta_L(1,1)$. It does not depend on the knowledge of the propensity score.

corrLet $1_{\mathcal{X}^{+}} = \mathbbm{1}(X \in \mathcal{X}^+)$ and $1_{\mathcal{X}^{-}} = \mathbbm{1}(X \in \mathcal{X}^-)$. Suppose that Assumptions (ref), (ref), (ref) and (ref) hold, and $P(\mathcal{X}^0)=0.$ The semiparametric efficiency bound $Var_a(\beta_L(1,1))$ for $\beta_L(1,1)$ is defined by {\begin{align*} &E[\min(s(0,X), s(1,X)]^2Var_a(\beta_L(1,1)) = E\left[\frac{s(1,X) \sigma_1^2 (X) }{m(X)}+\frac{s(0,X) \sigma_0^2 (X) }{1-m(X)} \right] \\ &+E\left[1_{\mathcal{X}^{+}}(\beta_L(X,1,1)-\beta_L(1,1))^2 \frac{s(0,X)(1-s(0,X) m(X))}{1-m(X)} \right] \\ &+E\left[1_{\mathcal{X}^{-}}(\beta_L(X,1,1)-\beta_L(1,1))^2 \frac{s(1,X)(1-s(1,X)+ s(1,X) m(X))}{m(X)} \right] \\ &+E\left[ 1_{\mathcal{X}^{+}} \frac{s(1,X) q_1(p_0(X),X)^2 p_0(X) (1-p_0(X))}{m(X)} \right] \\ &+E\left[ 1_{\mathcal{X}^{-}} \frac{s(0,X) q_0(1-1/p_0(X),X)^2 p_0(X)^{-1} (1-p_0(X)^{-1})}{1-m(X)} \right]\\ &+E\left[1_{\mathcal{X}^{+}} (q_1(p_0(X),X)-\beta_{1,1}(X,p_0(X)))^2\left(\frac{ s(0,X)(1-s(0,X))}{1-m(X)}+ \frac{p_0(X)^2 s(1,X)(1-s(1,X))}{m(X)} \right) \right] \\ &+E\bigg[1_{\mathcal{X}^{-}}(q_0(1-1/p_0(X),X)-\beta_{0,0}(X,1-1/p_0(X)))^2 \\ &\quad \times \left(\frac{ p_0(X)^{-2} s(0,X)(1-s(0,X))}{1-m(X)}+ \frac{s(1,X)(1-s(1,X))}{m(X)} \right) \bigg]\\ &- 2E\left[ 1_{\mathcal{X}^{+}} \frac{q_1(p_0(X),X) \beta_{1,1}(X,p_0(X)) s(1, X) p_0(X)(1-p_0(X))}{m(X)} \right]\\ &-2E\left[ 1_{\mathcal{X}^{-}} \frac{q_0(1-1/p_0(X),X) \beta_{0,0}(X,1-1/p_0(X)) s(0, X) p_0(X)^{-1}(1-p_0(X)^{-1})}{1-m(X)} \right]\\ &+2E\left[ 1_{\mathcal{X}^{+}} \frac{(\beta_L(X,1,1)-\beta_L(1,1))(q_1(p_0(X),X)-\beta_{1,1}(X,p_0(X))) s(0, X) (1-s(0,X))}{1-m(X)} \right] \\ &-2E\left[ 1_{\mathcal{X}^{-}} \frac{(\beta_L(X,1,1)-\beta_L(1,1))(q_0(1-1/p_0(X),X)-\beta_{0,0}(X,1-1/p_0(X))) s(1, X) (1-s(1,X))}{m(X)} \right], \end{align*} } where { \begin{align*} \sigma_{1}^2(x) &= Var[Y1_{\{ Y \leq q_1(\min\{p_0(x),1\},x) \}}|S=1,D=1,X=x], \\ \sigma_{0}^2(x) &= Var[Y 1_{\{ Y \geq q_0(\max\{1-1/p_0(x),0\},x) \}}|S=1,D=0,X=x]. \end{align*} }

The setting with full selection, $S=1$ almost surely, renders any truncation redundant and thus $\beta_L(x,1,1) = \beta_U(x,1,1) = \beta(x,1,1)$ are equal to the conditional average treatment effect. Consequently, the efficiency bound in Corollary (ref) becomes

equation[equation omitted — 123 chars of source]

with

align*[align* omitted — 64 chars of source]

which is exactly the efficiency bound for the ATE obtained in Theorem 2 of hahn1998role, i.e. Corollary (ref) nests this as a special case. The smooth bounds (ref) circumvent the irregularity problem by relaxing the width of the identified set using smooth approximators, i.e. the $g$-functions. If the latter are sufficiently smooth, then the resulting outer bounds are pathwise differentiable. In particular, we obtain the following Theorem:

theoremSuppose that Assumptions (ref), (ref), (ref), (ref), (ref) hold. If $\sup_{z,j}g_{i,h}'(z) \lesssim 1$ for some $h>0$, then $\beta_{L,h}(1,1)$ is pathwise differentiable.

The corresponding efficiency bound is in Appendix (ref). The additional smoothness assumption applies to the LogSumExp example shown in Section (ref).

Estimation and Inference

Estimators

We propose to construct estimators by using empirical analogues to the efficient influence functions in Table (ref) with estimated nuisances. Let $E_n[X] = n^{-1}\sum_i^nX_i$. For given nuisance estimates $\hat{\eta}$, any estimator $\hat{\beta}_B = \hat{\beta}_B(s_0,s_1)$ solves

align[align omitted — 68 chars of source]

Note that all principal strata and margin influence functions have the following linear structure

align[align omitted — 141 chars of source]

Thus, the influence function-based estimators have a ratio structure of the form

align[align omitted — 144 chars of source]

For the smooth bounds, parameters are given by a sum

align[align omitted — 61 chars of source]

whose components also have linear influence functions with structure

align[align omitted — 312 chars of source]

which combined with estimated nuisances yield moment estimators

align[align omitted — 305 chars of source]

For the intensive margin, we also consider the moment functions suggested by semenova2023generalized and heiler2024heterogeneous under known propensity scores that have influence function

align[align omitted — 109 chars of source]

with mean-zero term $\delta(D,X,\eta)$ characterized by {

align[align omitted — 431 chars of source]

} Solving for the parameter at the sample average then yields the alternative estimator

align[align omitted — 160 chars of source]
table[table omitted — 2,124 chars of source]
table[table omitted — 3,161 chars of source]

Definitions and Assumptions

In this section, we provide and discuss the assumptions and results for large sample inference for the regular and irregular case. Denote $||\cdot||_p$ as the $L_p$ norm. Let $\eta \in T$ where $T$ is a convex subset of some normed vector space. We assume all nuisance functions are cross-fitted as in chernozhukov2018double, Definition 3.1 and use the LogSumExp function for smoothing, see Section (ref). Denote nuisance realization set $\mathcal{T}_n = (\mathcal{S}_{0,n}\times\mathcal{S}_{1,n}\times \mathcal{Q}_{0,n} \times \mathcal{Q}_{1,n} \times \mathcal{M}_n \times \mathcal{B}_{0,0,n}\times \mathcal{B}_{0,1,n}\times \mathcal{B}_{1,0,n}\times \mathcal{B}_{1,0,n}) \subset T$ as the set that with high probability contains estimators $\hat{\eta} = \{\hat{s}(0,X), \hat{s}(1,X),\hat{q}_0(u,X),\hat{q}_1(u,X),\hat{m}(u,X),\hat{\beta}_{0,0}(u,X),$ $\hat{\beta}_{0,1}(u,X),\hat{\beta}_{1,0}(u,X),\hat{\beta}_{1,1}(u,X)\}$ for nuisance quantities $\eta = \{s(0,X),s(1,X),q_0(u,X), $ $q_1(u,X),m(X),{\beta}_{0,0}(u,X), {\beta}_{0,1}(u,X),$ ${\beta}_{1,0}(u,X),{\beta}_{1,1}(u,X)\}$. Let their corresponding $L_p$ error rates be

align*[align* omitted — 538 chars of source]

where $\tilde{U}$ is a compact subset of $(0,1)$ containing the relevant quantile trimming threshold support unions.\footnote{For the regular case, this will be $([\mbox{supp}(p_0(X)) \cup \mbox{supp}(1-p_0(X))] \cap \mathcal{X}^{+}) \cup ([\mbox{supp}(1/p_0(X)) \cup \mbox{supp}(1-1/p_0(X))] \cap \mathcal{X}^{-})$ while for the smoothed bounds, the subset will depend on $h$ and is defined equivalently with all bounds replaced by their $g_{1,h}$-smoothed counterparts.}

ass[Outcome Regularity III] The outcome has a continuous conditional density $f(y|X=x,D=d,S=s)$ that is uniformly bounded from above and away from zero with bounded first derivative for any $x\in \mathcal{X}$ and $s,d\in \{0,1\}$.

Assumption (ref) here strengthens the smoothness and moment assumptions. This is required to bound (higher order) terms in some of the expansions and a central limit theorem for transformations of the outcome with bounded weights.

Regular Bounds

We first consider the regular case that sets $P(\mathcal{X}^0) = 0$. For this, we require the density around of $s(1,x) - s(0,x)$ in a neighborhood around zero to be well-behaved. This neighborhood is allowed shrink at a rate proportional to the error over the realization set for the selection probabilities. In particular,

ass[Margin] There exist a finite constant $C > 0$ and $\alpha > 0$ such that \begin{align*} P(|s(1,x) - s(0,x)| \leq \delta) \leq C\delta^{\alpha} \quad for any \ 0 \leq \delta \leq c_n \end{align*} and $c_n \lesssim \lambda_{s,n,2}^{1/(2+\alpha)}$.

Note that this implies that $P(\mathcal{X}^0) = 0$ which we know is necessary for the existence of a first-order uniformly unbiased estimator by Theorem (ref). Any continuous bounded density for $s(1,x) - s(0,x)$ is sufficient yielding $\alpha = \infty$. However, a density may also diverge around zero at a controlled rate. We further require the following learning rates.

ass[Machine Learning Bias I] Let $e_n = o(1)$. For all folds, the nuisance parameters obtained via cross-fitting belong to a shrinking neighborhood $\mathcal{T}_n$ around $\eta$ with probability of at least $1-e_n$, such that \begin{align*} \lambda_{s,n,1} + \lambda_{s,n,2}^{\frac{\alpha}{2 + \alpha} } + \lambda_{q,n,1} + \lambda_{q,n,2} + \lambda_{m,n,2} + \lambda_{b,n,2} = o(1) \end{align*} and \begin{align*} \sqrt{n}(\lambda_{s,n,2}^{\frac{2\alpha}{2 + \alpha}} + \lambda_{q,n,2}^2 + \lambda_{m,n,2}^2 + \lambda_{b,n,2}^2) = o(1). \end{align*}

One can see that we require $L_1$ and $L_2$ consistency for quantiles and conditional selection probabilities. The $L_1$ requirement is due to the variance of the trimming indicator for the outcome. Moreover, to control the machine learning bias, the squared $L_2$ rates have to be sufficiently fast. For most nuisances, these are equivalent to the usual $o(n^{-1/4})$ root mean squared error (RMSE) requirement in the debiased machine learning literature chernozhukov2018double. The more demanding rate requirement for the selection probabilities is due to difficulties of classifying positive and negative monotonicity types around the boundary. Given these rates, Assumption (ref) is only really plausible for $\alpha > 2$ as, even in the parametric case where $\lambda_{s,n,2} \sim n^{-1/2}$, $\sqrt{n}\lambda_{s,n,2}^{2\alpha/(2+\alpha)} \nrightarrow 0$ when $\alpha \leq 2$. For more demanding selection probabilities in high dimensions or with complicated functional forms, the density needs to be increasingly well-behaved around the margin. As $\alpha \rightarrow \infty$, the requirement reduces to the typical $o(n^{-1/4})$ rate as in heiler2024heterogeneous. We obtain the following Theorem.

theoremUnder Assumptions (ref), (ref), (ref), (ref), (ref), and (ref), the regular estimator is asymptotically normal and semiparametrically efficient, i.e.\begin{align*} \sqrt{n}(\hat{\beta}_L(1,1) - \beta_L(1,1)) \overset{d}{\rightarrow} \mathcal{N}\big(0,Var_a(\beta_L(1,1))\big), \end{align*} where $Var_a(\beta_L(1,1))$ is the semiparametric efficiency bound in Theorem (ref).

Smooth Bounds

We now consider the large-sample behavior of the smooth bounds. We impose the following learning requirements.

ass[Machine Learning Bias II] Let $e_n = o(1)$. For all folds, the nuisance parameters obtained via cross-fitting belong to a shrinking neighborhood $\mathcal{T}_n$ around $\eta$ with probability of at least $1-e_n$, such that \begin{align*} \lambda_{s,n,1} + \lambda_{s,n,2} + \lambda_{q,n,1} + \lambda_{q,n,2} + \lambda_{m,n,2} + \lambda_{b,n,2} = o(1) \end{align*} and \begin{align*} \sqrt{n}(\lambda_{s,n,2}^2 + \lambda_{q,n,2}^2 + \lambda_{m,n,2}^2 + \lambda_{b,n,2}^2 ) = o(1). \end{align*}

We obtain the following Theorem.

theoremUnder Assumptions (ref), (ref), (ref), (ref), and (ref), and $h > 0$ such that $\sup_{z,j}g_{j,h}'(z) \lesssim 1$, the smooth bound estimator is asymptotically normal and semiparametrically efficient, i.e.\begin{align*} \sqrt{n}(\hat{\beta}_{L,h}(1,1) - \beta_{L,h}(1,1)) \overset{d}{\rightarrow} \mathcal{N}\big(0,Var_a(\beta_{L,h}(1,1))\big), \end{align*} where $Var_a(\beta_{L,h})$ is the semiparametric efficiency bound for $\beta_{L,h}$.

We can see that, without restricting the distribution of $s(1,x) - s(0,x)$, i.e. allowing for point mass $P(\mathcal{X}^0) > 0$ or arbitrary density around zero, the smooth bounds are asymptotically normal under weaker assumptions than the non-smooth bounds with restricted distributions. In particular, they only require $L_1$ and $L_2$ consistency as well as root mean squared error rates for all nuisance quantities of order $o(n^{-1/4})$ as in standard debiased machine learning chernozhukov2018double. This is fundamentally due to the smoothing turning the target parameter into a pathwise differentiable object.

Our smooth estimator differs from semenova2023generalized in the following way: semenova2023generalized uses a moment shifting or shrinkage procedure based on the selection probabilities. These shifting methods are known to be quite sensitive to the choice of tuning/selection parameter andrews2010inference. Moreover, even when $P(\mathcal{X}^0) = 0$, the estimator by semenova2023generalized in finite samples should tend to narrow estimated identified sets as observations close to $p_0(x) = 1$ use a moment for the standard difference in conditional outcome means without truncation. Thus, while asymptotically valid due to correct classification in the limit, it will tend to be over-reject a correct null hypothesis in finite samples. The smoothed estimators, on the other hand, are constructed to be always wider than the estimators using no smoothing. They therefore estimate an outer identification region and the estimated identified will always be at least as wide as the naive generalized Lee bounds. This suggests better finite-sample size control at the expense of some power, see Appendix (ref) for Monte Carlo simulations.

The Efficiency Gap and Known Propensity Scores

If propensity scores are known, the moment functions from semenova2023generalized or heiler2024heterogeneous without correction can be used as well. The corresponding estimator (ref) does not require estimation of $\beta_{j,d}(u,X)$ for any $j,d\in\{0,1\}$ and thus all $\lambda_{b,n,p}$ terms that enter Assumption (ref) and (ref) can be omitted. However, there is an information loss from ignoring the conditional variation in the truncated means. In particular, we obtain the following Theorem.

theoremUnder Assumptions (ref), (ref), (ref), (ref), (ref), and (ref) with $\lambda_{b,n,p} = 0$ known, $\tilde{\beta}_{L}(1,1)$ is asymptotically normal \begin{align*} \sqrt{n}(\tilde{\beta}_L(1,1) - \beta_L(1,1)) \overset{d}{\rightarrow} \mathcal{N}\big(0,Var_{b}(\beta_L(1,1))\big), \end{align*} with efficiency gap \begin{align*} &E[\min\{s(0,X),s(1,X)\}]^2(Var_{b}(\beta_L(1,1)) - Var_a(\beta_L(1,1))) \\ &={E\bigg[\mathbbm{1}_{\{X \in \mathcal{X}^+\}}s(0,X)^2 \bigg( \beta_{1,1}(X,p_0(X))\sqrt{\frac{1-m(X)}{m(X)}} - \beta_{0,0}(X,0)\sqrt{\frac{m(X)}{1-m(X)}}\bigg)^2 \bigg]} \\ &\quad + {E\bigg[\mathbbm{1}_{\{X \in \mathcal{X}^-\}}s(1,X)^2 \bigg( \beta_{1,1}(X,1)\sqrt{\frac{1-m(X)}{m(X)}} - \beta_{0,0}(X,1-1/p_0(X))\sqrt{\frac{m(X)}{1-m(X)}}\bigg)^2 \bigg]}. \end{align*}

Theorem (ref) demonstrates that, under known propensity scores, there is a trade-off between efficiency and learning rates for the truncated conditional means $\beta_{j,d}(u,X)$. A lack of efficiency from not including information about the conditional truncated potential outcome means in the respective influence function arises as the variance of the truncated variables are bounded from below by their conditional analogues. This is equivalent to the comparison between the variance of a dependent variable and its residual variance in standard regression analysis. Theorem (ref) also nests the case without sample selection. In particular, when $S=1$ almost surely, the ATE is point identified $\beta_L(1,1) = \beta_U(1,1)$ and all truncated means simplify to their untruncated counterparts $\beta_{d,1}(x,1) = \beta_{d,0}(x,0) = E[Y|D=d,X=x]$. In this case all correction terms vanish as well and the alternative estimator collapses to

align*[align* omitted — 117 chars of source]

This is the standard Horvitz-Thompson-type IPW estimator for the ATE using true propensities which is known to not reach the semiparametric efficiency bound hahn1998role. More precisely, Theorem (ref) yields efficiency gap

align*[align* omitted — 182 chars of source]

This comparison to the point-identified case suggests that, in the regular regime, one could potentially obtain an efficient estimator for the always-taker ATE bounds without modelling the conditional truncated quantiles analogously to hirano2003efficient using nonparametrically estimated treatment propensities with sufficiently rich basis functions that asymptotically encode all information about the truncated conditional outcome means. We leave the development of such an alternative estimation approach for future research.

Extensions

Mean Dominance

The bounds in Proposition (ref) can further be refined using a mean dominance assumptions. Mean dominance to bound causal effects in sample selection models has been previously suggested by zhang2003estimation for always-takers and by huber2015sharp for compliers under strong monotonicity without covariates.

ass[Mean Dominance] The conditional potential outcomes at the intensive margin are larger than at the extensive margin, i.e. \begin{align} E[Y(d)|S(0)=1,S(1)=1,X=x] \geq E[Y(d)|S(0)\neq S(1), X=x] \end{align} for any $d=0,1$.

This yields the following modified sharp bounds\footnote{ If only either complier or defier bounds are of interest, Assumption (ref) can be relaxed to hold for the respective group only.}.

prop[Identification with Mean Dominance] Let $\mathcal{Y}_d(x)$ be the support of $Y(d)$ conditional on $X=x$ and $\overline{y}_d(x)$ and $\underline{y}_d(x)$ its respective infimum and supremum for $d\in\{0,1\}$. Under Assumptions (ref), (ref), (ref), (ref), and (ref), the lower bounds for the principal strata are given by \begin{align*} \beta_L(x,s_0,s_1) = \begin{cases} \beta_{1,1}(x,1) - \beta_{0,0}(x,1-\min\{1/p_0(x),1\}) & if s_0 = 1, s_1 = 1 \\ \beta_{1,1}(x,1-\min\{p_0(x),1\}) - \beta_{0,0}(x,0) & if s_0 \neq s_1 \\ {y}_1(x) - \overline{{y}}_0(x) & if s_0 = 0, s_1 = 0, \end{cases} \end{align*} and the upper bounds by \begin{align*} \beta_U(x,s_0,s_1) = \begin{cases} \beta_{0,1}(x,1-\min\{p_0(x),1\}) - \beta_{1,0}(x,1) & if s_0 = 1, s_1 = 1 \\ \beta_{0,1}(x,0) - \beta_{1,0}(x,1-\min\{1/p_0(x),1\}) & if s_0 \neq s_1 \\ \overline{{y}}_1(x) - \underline{{y}}_0(x) &\text{ if } s_0 = 0, s_1 = 0. \end{cases} \end{align*}

As the share of the principal strata are already identified via conditional monotonicity, aggregation to unconditional bounds uses the same conditional probability weights as in the case without mean dominance in (ref) and thus the same moment function for the denominator or density part in (ref). For the numerator, all truncated quantities can be estimated as in the pure monotonicity case with the components in the influence function corresponding to the now untruncated means replaced by their standard untruncated augmented IPW counterparts for the conditional outcome means robins1994estimation.

Heterogeneous Bounds

In many evaluation problems, objects of interest are heterogeneous effects, e.g. the effect at a particular margin for an observable subgroup. heiler2024heterogeneous argues that, in the nonparametric sample selection model and more broadly, there are two types of heterogeneity, 1) heterogeneous effects and 2) heterogeneity in the severity of the identification problem, i.e. the width of the identified set. Exploiting these in combination can yield more precise inference and significant effect bounds even when unconditional aggregate bounds do not reject a null-effect. Our influence functions for any principal strata or margin can readily be used for the estimation of such heterogeneous effect bounds following the method of heiler2024heterogeneous. In particular, let $f:\mathcal{X}\rightarrow \mathcal{Z}$ and define $Z = f(X)$ a low-dimensional subgroup or mapping from the covariate space. Recall that all presented influence functions have the following linear structure

align[align omitted — 119 chars of source]

where

align[align omitted — 178 chars of source]

Thus, for any $z \in \mathcal{Z}$,

align[align omitted — 319 chars of source]

by the law of iterated expectations. Hence we can obtain conditional effect bounds by aggregating the components of the presented influence functions within covariate partitions or by (nonparametric) projection onto (basis transformations of) $Z$. Regularity and/or the necessity for smoothing can be obtained under analogous conditions to the ones presented in this paper. For more details consider heiler2024heterogeneous.

Empirical Study: Labor Market Policy

Job Corps and Data

Job Corps (JC) is a US Department of Labor program aimed at empowering young individuals of economically or otherwise disadvantaged background with important labor market skills and qualifications. The intense program combines general education, vocational training, extracurricular activities, placement services, and more. Most participants are in residential slots in centers of varying size all over the US. The National Job Corps Study, executed by Mathematica Policy Research in the mid-90s, analyzed the effects of the Job Corps program on labor market prospects. The experiment implemented stratified randomized assignment of applicants, incorporating over 15400 units between ages 16 to 24. Data on various outcomes such as income, job status, educational achievements, and criminal activities were gathered at different intervals. The outcomes of the original Job Corps study were mixed. Initially, the program boosted the participants' educational achievements and income. However, most aggregate effects waned over time. There is evidence for increased earnings in the older participant population schochet2008does. There is also well-document heterogeneity in terms of vocational and academic training returns as well a gender, see e.g. flores2012estimating or heiler2023effect.

We employ the public use files which contain all survey data up to 208 weeks after the initial assignment burghardt1999national,schochet2008does. Our main outcome measure is log hourly earnings as in lee2009training. If workers are priced around their marginal productivity, this measure can be used as a proxy for increases in the latter caused by assignment to Job Corps. We extract an extended set of covariates containing detailed information regarding demographics, employment, criminal history, education, health, expectations, regional characteristics, and other JC related information.\footnote{We collect all variables used in lee2009training and additional information on individual background and behavior similar to flores2012estimating and heiler2023effect, as well as all variables that were used for stratification in the original experiment. Some of the latter have previously been overlooked but are necessary for correct specification of the propensity score, see Appendix (ref) for more details.} All impact analysis is conditional on having a non-missing earnings entry, including zero, at weeks $t\in\{1,\dots,208\}$ identical to lee2009training leaving a sample of 9.415 units.\footnote{lee2009training suggests that non-response at this stage is a second order issue and initial treatment-control balance is preserved, see his Remark 2. We reassess this claim with the extended data using both formal tests for average differences as well as standardized differences and find little evidence for imbalance, see Appendix (ref). The only exception is significantly larger worries to attend Job Corps in the treatment group. Our prior is that worries are more likely negatively related to any treatment effect. Thus, we expect any potential bias to be negative, i.e. earnings and employment effects are likely to be at least as large.}

In the following, results for nuisance functions are presented for the research sample. For any of the impact analysis as well as general descriptive statistics, we use the nationally representative design weights as in schochet2008does. For the impact evaluation, we analyze the effect per eligible participant, i.e. the causal effect of assignment to Job Corps. We focus on the average treatment effect at the intensive margin as studied by lee2009training,semenova2023generalized. As participation or outside alternatives after assignment are not controlled, all Job Corps effects here are relative to the control state of not being assigned, including partaking in any alternative program and/or working.

Figure (ref) contains the weekly average earnings and employment rate for the 208 weeks after treatment assignment. There are systematic trends and differences between treatment and control groups in both earnings and employment over time. There is a clear upwards trend in earnings and employment over time. The control group tends to be ahead early after assignment but catches up later in time exceeding the trajectory of the control group for both measures. However, any differences between treatment and control earnings effectively consists of a mixture of units (the working) that could be made up of varying proportions of always-takers, compliers, and defiers, effectively leading to selected comparisons. Raw differences in employment rates are also uninformative about heterogeneity and composition of the intensive and extensive margins. We take this into account when constructing the bounds in what follows.

figure[figure omitted — 935 chars of source]

Models and Estimation

We give a brief overview over the estimation methods used. See Appendix (ref) for more details. Treatment propensities are known by the design of the stratified experiment, see Appendix (ref). All other nuisance parameters are estimated using 5-fold cross-fitting with all available observations and pre-treatment predictors at each time period reported. No cross-time information was used.

For the conditional selection probabilities we compare five different main specifications: (a) Logistic regression, (b) logistic regression with interacted treatment, (c) gradient boosted trees (XGBoost), (d) artificial neural network, and (e) random forest. These methods differ with respect to their capabilities to (i) allow for weak monotonicity, (ii) produce sparsity, (iii) be robust to extreme observations, and (iv) fit and predict employment probabilities and status well. Table (ref) contains an aggregate comparison between the different methods with respect to these properties. By construction, all models except for the simple logit allow for weak monotonicity. XGBoost, neural network, and random forest can produce sparsity.\footnote{Both the interacted logit and neural network tend to produce many or some relatively extreme estimates for the relative selection probabilities, see also Figure (ref) in Appendix (ref). It seems questionable whether e.g. extreme relative selection probabilities that suggest e.g. a quadrupling in relative employment chances due to Job Corps are credible or mostly a by-product of an overly high variance model.} XGBoost is the model that has, by far, the best out-of-sample accuracy and loss as measured by the negative likelihood function outperforming all other methods throughout 202/206 out of 208 periods, see Appendix (ref) for the detailed results. This is in line with many studies that demonstrate boosted trees as or among the best performing methods for tabular data shwartz2022tabular.

table[table omitted — 1,323 chars of source]

For conditional means, conditional quantiles, and truncated means required for the bound estimators, we use off-the-shelf honest regression and quantile forests across all specifications athey2019generalized. This was chosen to isolate the contribution of differences in relative selection probabilities which is the main issue addressed in this paper.\footnote{Quantile forests are also computationally attractive in high dimensions and have been used for sample selection bounds in previous research heiler2024heterogeneous,andersen2023guide.}

For the impact analysis, we consider the bounds that (i) remove observations with $p_0(x) = 1$ and use the corresponding influence function (“Trim”), (ii) the moment shifting or selection method by semenova2023generalized (“Shift”), (iii) smooth bounds (“Smooth”), and (iv) conventional Lee bounds using strong monotonicity. For the shifting method, we use the tuning parameter as suggested by semenova2023generalized, $\rho_n = n^{-1/4}/\log n$. For both trimming and shifting, we report results with and without using the propensity score correction that affects efficiency. For the smooth bounds, we report the analysis on a grid of smoothing parameters.\footnote{On any fixed grid, inference is valid for any $h$, so one could pick the $h$ that provides the smallest confidence intervals to optimize power.} For inference on the effect we use imbens2004confidence confidence intervals for all methods.

Relative Selection Probabilities $p_0(x)$

In this section, we consider the predicted relative selection probabilities close or equal to one for the methods that can produce sparsity. Figure (ref) depicts the share of probabilities for XGBoost, neural network, and random forest. The average amount of close-to-one units averaged over all time periods are $51.91\%$ (XGBoost), $18.75\%$ (Neural Net), and $42.36\%$ (Random Forest). Thus, there is a significant amount of units that can affect the various bounds differently. In particular, for the XGBoost, almost all these probabilities are exact numerical ones, i.e. the underlying model is sparse in some parts of the covariate space with respect to the treatment. Figure (ref) depicts the close-to-one and exact-one unit shares over all time periods. One can see that the trajectory of exact-one closely follows the close-to-one shares with an average rate of 44.36%.

figure[figure omitted — 913 chars of source]

Employment Effects

Figure (ref) contains the estimated average assignment effect on employment using the efficient estimator for all probability models. Estimates are almost identical. We replicate the finding that, in the short run, assignment reduces the likelihood of employment which only bounces back after about 1.5 years (the average duration of JC participation was around 8 months at the time). After about 2.25 years, positive employment effects tend to stabilize. In particular, most estimates are significantly negative for weeks 2 to 76 and significantly positive for weeks 116 to 208 ($p < 0.05$).

figure[figure omitted — 544 chars of source]

Causal Effects Bounds

We now present the impact results for assignment to JC on log hourly earnings. We first present the results for the main evaluation in $t=208$ using the new methods and contrast them with results based on the literature and replications with comparable sample and design weights. We then present the results using the best performing nuisance specification for methods that allow for conditional monotonicity for periods $t\in\{45,90,135,180,208\}$ as in lee2009training. Robustness checks and results for additional methods, time periods, and smoothing parameters can be found in Appendix (ref).

Main Effects at $t=208$

Table (ref) contains the impact analysis for various methods and specifications at period $t=208$. All entries except for Lee (2009) are based on own calculations and with identical outcome and design weights. schochet2008does is a simple treatment control difference for the sample of reported hourly log earnings. It is of similar magnitude when compared with the equivalent mean difference in the restricted Lee (2009) sample in line with the balancing analysis.

Switch, regular, and smooth bounds all admit flexible estimation of nuisance parameters. Using the best performing XGBoost model for selection, there are 4252 observations with exact $\hat{p}_0(x) = 1$. Thus, trimming bounds effectively have to use a much reduced sample. Nevertheless, estimates are in a similar range but with larger standard errors. The moment switching approach by semenova2023generalized with known propensity scores provides overly wide bounds. This seems to be due to a strong sensitivity when there are many estimates for $p_0(x)$ around one. With unknown scores, estimates are tighter, however with standard errors much larger compared to alternatives.

\afterpage{

landscape\begin{table}[!h] \caption{Impact Results at $t=208$: Overview} \begin{threeparttable} \scriptsize \begin{tabular}{llccccccc} \hline \hline & Method & Impact Estimate & 95% CI & Sample & Selection & Selection & Covariates & Sample \\ & & & & Selection & Model & Key Assumptions & & \\ \hline \\[-0.5ex] Schochet et al. (2008) & Mean Difference & 0.059 & $(0.031,~0.086)$ & \ding{55} & -- & Independence$^1$ & 0 & 10602 \\ & & & & & & & & \\ & & & & & & & & \\ Lee (2009) & Mean Difference & 0.058 & $(0.027,~0.088)$ & \ding{55} & -- & Independence$^1$ & 0 & 9415 \\ & & & & & & & & \\ & Bounds & $[-0.019,~0.093]$ & $(-0.049,~0.114)$ &\ding{51} & -- & Monotonicity & 0 & 9415 \\ & + Covariates & $[-0.012,~0.089]$ & $(-0.037,~0.112)$ &\ding{51} & Nonparametric & Monotonicity & $28/1/5^2$ & 9415 \\ & & & & & & & & \\ & Heckman (1979) & 0.015 & $(-0.008,~0.038)$ &\ding{51} & Probit & Exclusion$^3$, Normality & 28 & 9415 \\ & Das et al. (2003) & 0.014 & $(-0.010,~0.038)$ &\ding{51} & Linear$^4$ & Exclusion$^3$, Single Index & $28^4$ & 9415 \\ & & & & & & & & \\ & & & & & & & & \\ Switch$^{5}$ & Known Propensity & $[-0.646,~1.514]$ & $(-0.438,~1.650)$ &\ding{51} & Nonparametric & Weak Mononicity & 77 & 9415 \\ & Unknown Propensity & $[0.138,~0.162]$ & $(-0.003,~0.303)$ & \ding{51} & Nonparametric & Weak Mononicity & 77 & 9415 \\ & & & & & & & & \\ & & & & & & & & \\ Trim$^{5,6}$ & Known Propensity & $[-0.008,~0.054]$ & $(-0.063,~0.111)$ &\ding{51} & Nonparametric & Weak Mononicity & 77 & 9415 \\ & Unknown Propensity & $[-0.036,~0.077]$ & $(-0.201,~0.245)$ & \ding{51} & Nonparametric & Weak Mononicity & 77 & 9415 \\ & & & & & & & & \\ Smooth Bounds$^{5,7}$ & $h = 5$ & $[-0.057,~0.275]$ & $(-0.095,~0.312)$ &\ding{51} & Nonparametric & Weak Mononicity & 77 & 9415 \\ & $h = 1$ & $[-0.005,~0.113]$ & $(-0.040,~0.148)$ & \ding{51} & Nonparametric & Weak Mononicity & 77 & 9415 \\ & & & & & & & & \\ \hline \end{tabular} Impact estimates using different methods and samples with key assumptions for validity of the impact estimate. All estimates and intervals except for Lee (2009) are based on own calculations using design weights. Confidence intervals for point-identified methods use standard critical values. Confidence intervals for bounds use the imbens2004confidence refinement. All methods use known propensity scores of the stratified experiment if necessary. $^1$Independence of potential selection from potential outcomes. $^2$28 covariates are used to generate a single predictive score. Bounds are calculated within 5 strata of this score. $^3$Months employed is used as excluded variable in the selection model. $^4$Semiparametric model also includes transformations of the original variables, see Lee (2009). $^5$Best performing nuisance parameter specifications, see Appendix (ref) for details. $^6$Trim are the bounds that remove all units for which $\hat{p}_0(x) = 1$ with unknown/efficient and known/inefficient propensity score and related moment function. $^7$Smoothing parameters $h$ scaled by 100. \end{threeparttable} \end{table}

}

Our preferred specification is smooth bounds with $h=1$. The impact estimate of $[-0.005,0.113]$ with 95% confidence interval of $(-0.040,0.148)$ rules out large negative earnings effects and is surprisingly close and of similar width compared to results obtained under strong monotonicity as in lee2009training. As strong monotonicity is rejected by the data semenova2023generalized, this provides more evidence on the robustness of this final impact evaluation. We also include a larger smoothing specification for comparison; for more values see Appendix (ref). As expected, the identified set becomes larger. In this case, the standard errors are still comparable such that overall less smoothing seems preferable.

Secondary Effects and Robustness

We now analyze the effect for multiple periods using all methods. We present the results for the two best nuisance models (XGBoost and neural network). For the switching and the trimming bounds, we only consider the unknown propensity score estimates, as they are more reliable across specifications. For smooth bounds we present two smoothing parameter results as above. Results for the other methods and parameters are robust with the exception of the interacted logistic model, that failed to converge in some folds.\footnote{See Appendix (ref) for some histograms and descriptive statistics.}

table[table omitted — 3,894 chars of source]

Table (ref) contains the estimates and confidence intervals for XGBoost and neural network selection scores. There are a few general observations. The choice of the selection model matters. In particular, the larger dispersion of predicted $p_0(x)$ for the neural net leads to wide intervals at $t=45$ for most methods. Trimming and switching seem more sensitive in general, see also Appendix (ref). For all other periods, results are relatively consistent across selection models. Trimming and switching bounds, however, tend to have relatively large standard errors compared to the smooth bounds. Results for additional methods reveal an identification-precision trade-off in some, but not all cases, see Appendix (ref).

The intervals for the smooth bounds at $h=1$ are, again, surprisingly close to the original lee2009training specification. Even a strictly positive identified set at $t=90$ is recovered with XGBoost, albeit with larger width which seems more credible given the estimates of the surrounding periods. The Lee unconditional monotonicity can be rejected from the data semenova2023generalized. However, it seems that the general finding for the later periods, with log hourly earnings effects in the range of around $-0.03$ to $0.11$, is robust to conditional monotonicity without parametric assumptions.

Conclusion

This paper demonstrates the importance of heterogeneity in selection behavior with respect to estimation and inference on causal effects at the intensive and extensive margins in selected samples. Allowing for different types of monotonicity crucially affects the (ir)regularity of sharp effect bounds and thus the ability to provide precise inference for the associated causal effects. This paper provides a solution in the form of outer identification regions that can be estimated efficiently from a semiparametric perspective. The approach can handle modern machine learning methods that are able to better approximate such important heterogeneity in sample selection. We discover an empirically relevant trade-off between identification strength versus precision that is likely to apply to a much larger set of models and parameters beyond the ones considered in this paper, as in lee2021bounding and pakel2023bounds.

\addcontentsline{toc}{section}{References} {\setstretch{1} }