EconBase
← Back to paper

Lee Bounds with a Continuous Treatment in Sample Selection

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

102,940 characters · 15 sections · 115 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Lee Bounds with a Continuous Treatment in Sample Selection

center[center omitted — 31 chars of source]

We study causal inference in sample selection models where a continuous or multivalued treat- ment affects both outcomes and their observability (e.g., employment or survey response). We generalize the widely used Lee (2009)’s bounds for binary treatment effects. Our key innovation is a “sufficient treatment value" assumption that imposes weak restrictions on selection heterogeneity and is implicit in separable threshold-crossing models, including monotone effects on selection. Our double debiased machine learning estimator enables nonparametric and high-dimensional methods, using covariates to tighten the bounds and capture heterogeneity. Applications to Job Corps and Civilian Conservation Corps (CCC) program evaluations reinforce prior findings under weaker assumptions. \\ \\[6pt] Keywords: Average dose-response, debiased machine learning, multivalued treatment, nonseparable model, partial identification. \\ JEL Classification: C14, C21

Introduction

Sample selection is a common challenge in studying treatment effects. A classic question in empirical economics is estimating the effect of training programs on wages. There is sample selection because a training program not only affects wages via human capital accumulation, but also affects the chance that a worker is eventually employed and hence is selected into samples. We only observe wages of employed workers. If the estimation is based on a selected sample with observed outcomes, then one must isolate the effect on employment status (extensive margin) to learn about the effect on wages (intensive margin). Nonetheless in survey data, when the treatment affects response behavior, those who respond to the survey (i.e., who are selected into samples) in the treatment group are no longer comparable to respondents from the control group. Such problems of sample selection, attrition, or missing data arise in fields other than economics; for example, in medical studies, quality of life after assignment of a new drug is only observed if a patient does not die (truncation by death). In education, final test score is only observed if a student does not drop out.

Due to the non-random sample selection, the causal effect on the outcome is not point-identified without imposing further assumptions on functional forms or distributions; e.g., RePEc:nbr:nberch:10491,Heckman, ImbensAngrist, ZRM, HonoreHu, ChenRoth. Assuming the treatment variable is randomly assigned (conditional on observables), we build on the seminal work of HM95 and LeeBound, who bound the average causal effect of a binary treatment. The setup is fully non-parametric and hence allows for general heterogeneity. Many treatment or policy variables are continuous or discrete multivalued, e.g., hours in a social program, lottery prize, drug dosage, tuition subsidy, cash transfer, air pollution, etc. As LeeBound's bound has been a common practice in empirical economics, we fill in the important gap to provide a corresponding tool to deal with sample selection in studying the causal effect of a continuous or multivalued treatment. The replication package of codes and data for our empirical applications is available on the authors' websites for practical implementation.

We provide the worst-case sharp bounds for the average treatment effect, or the average dose-response function, for {\it always-takers} who are selected into samples, or whose outcomes are observed, regardless of the treatment values they receive.\footnote{ Always-takers are also known as always-observed, always-responders, always-employed, nonattriters, or survivors. The concept of always-takers is from the literature on imperfect compliance of treatment AIR, where “taking" is the taking of the treatment affected by an instrument, rather than selection into the sample affected by the treatment, considered in this paper.} Note that the selected sample given a certain treatment value $d$ consists of always-takers and {\it compliers} who are not selected given other treatment values $d' \neq d$. The key element of the bounds is the proportion of always-takers in the selected $d$-treated sample, which is used to trim the observed outcomes for the worst-case lower (upper) bound when all always-takes' outcomes are smaller (or larger) than all compliers' outcomes. Recall that for a binary treatment, LeeBound assumes monotone treatment effects on selection, i.e., if a subject is selected in the control group, then it must be selected in the treatment group. Then only always-takers can be in the selected untreated sample and are identified. For a continuous or multivalued treatment, we propose a novel {\it sufficient treatment value} assumption on selection: if a subject is selected into samples when it receives the {\it sufficient treatment value}, then it remains selected when receiving {\it any} other treatment values. So we generalize the monotonicity assumption in LeeBound by assuming the sufficient treatment value to be zero for a binary treatment. Then the selected subjects who are treated with the sufficient treatment value are always-takers. So the probability of always-takers is identified by the minimum conditional selection probability over the treatment values.

For example, in a standard setting of survey attrition, if we find the response rate (conditional selection probability) given cash transfer is lowest at \$1,000, then our sufficient treatment value assumption is that everyone who responded to a survey question when receiving a cash transfer of \$1,000 would also have responded if there received any other values (always-takers). And there are some additional people (compliers) who only responded when they received some other transfer values and did not respond given \$1,000. We can interpret the sufficient treatment value \$1000 as the least-favored treatment (cash transfer) to induce selection into samples (responding to the survey) in the sense that if a subject is selected under the sufficient treatment value, then it is selected under any other treatment values.

In fact, the sufficient treatment value is implicit under a separable structural error in selection, or in a widely used class of latent variable threshold-crossing models Vytlacil02. This important observation suggests that the sufficient treatment value assumption is not restrictive, and the subpopulation of always-takers is a natural target under minimal assumptions. The associated always-takers are the largest subpopulation for whom we can partially identify the average effect of switching treatment over a range of values chosen by researchers for treatment intervention.

Moreover we allow for unconfoundness assumption in observational studies, and subjects with different pretreatment covariates can have different sufficient treatment values. So the information of covariates could potentially tighten the bounds and confidence intervals, or capture heterogenous effects that is not revealed without using the covariates, as supported by our empirical illustrations. Our bounds and asymptotic inference are robust to the extensive margin effect on selection, in the sense that when there is no selection bias or no extensive margin, our bounds contain the point-identified causal object. So we avoid pre-testing the treatment effect on selection.

We note that one might be interested in the partial (or marginal) effect defined as the derivative of the average dose response function. However, we cannot bound such derivative by the same approach of LeeBound. Instead, we bound the average effects of switching treatment values within a subset of the support of the treatment, or an average derivatives over two treatment values, which could be interesting in practice.

We illustrate our methodology by two applications. We find significant effects when incorporating covariates, while the bounds estimates without covariates show positive but insignificant effects. So incorporating covariates is useful to increases precision and captures heterogeneity. First, we revisit the Job Corps program, one of the largest federally funded job training programs in the U.S. The evaluation of these programs has been the focus of a substantive methodological literature, due to the high cost of about \$14,000 on average per participant; see SBM, LeeBound, FFGN12ReStat, among others. The participants are exposed to different numbers of actual hours of academic and vocational training. Their labor market outcomes may differ if they accumulate different amounts of human capital acquired through different lengths of exposure. We try to understand whether hours of training raise wages by helping them find a job (extensive margin) or by increasing human capital that would affect the intensive margin effect on the outcome. We find that increasing the training from 1.5 weeks to 9 months increases log weekly earnings by at least 0.224 at 5% significance level for those always employed ({\it always-takers}).

The second application evaluates the Civilian Conservation Corps (CCC) in Aizer, who conduct the first lifetime evaluation of the largest federal youth employment program in U.S. history created to address high youth unemployment during the Great Depression. The Job Corps is a modern-era job training program that was modeled after the CCC and shares many features. We bound the effect of service duration on the age at death and strengthen the findings in Aizer. As differential attrition could bias the OLS estimates, they find that the effect of duration on longevity is consistently positive and statistically significant under various imputation approaches. Our bounds suggest that increasing duration from about 3 months to 14 months increases the average death age by at least 1.17 years at 5% significance level.

Another theoretical contribution of this paper is a weaker sufficiency assumption of a {\it sufficient set} of $M$ treatment values, which is a useful approximation when the error of unobserved heterogeneity is non-separable in the selection equation. For example of $M=2$, if a program participant is employed under {\it both} one week and fifteen months of training, then this participant is always employed. Such a weaker identification assumption results in a tradeoff with less informative (wider) bounds. Furthermore we utilize the well-known Fr\'{e}chet-Hoeffding bounds for the discrete treatment {\it without} imposing any shape restrictions on selection.

We incorporate covariates, following SGLee on the generalized Lee bounds for the binary treatment. Our bound estimator is doubly debiased using an orthogonal moment function and cross-fitting, which enables nonparametric and machine learning methods to handle high-dimensional data, following the recent double debiased machine learning (DML) literature CCDDHNR. Since the average dose-response function, or the mean potential outcome, and its bounds are functions of the continuous treatment, such non-regular estimand cannot be estimated at the regular root-$n$ rate. We use a kernel function for localizing the continuous treatment as in CL.

The paper is organized as follows. Section (ref) describes the sample selection model under the potential outcome framework, or equivalently a nonparametric non-separable structural model (e.g., IN09ETA). We discuss related literature. Section (ref) presents the basic Lee bounds without covariates for a continuous/multivalued treatment under a sufficient value/set assumption on the treatment effect on selection. We give estimation and inference theory in Section (ref). Section (ref) incorporates the covariates and presents the DML inference. Section (ref) and Section (ref) present the empirical illustration on evaluating the Job Corps and the CCC programs. Appendix contains the main proofs of Theorems and Lemmas. In the online supplementary appendix, we present the proofs of Corollaries, and supplementary material for the empirical applications.

Sample selection model and related literature

The researcher chooses a compact subset of the support of the treatment variable $D$, denoted as $\mathcal{D}$, for treatment intervention. So we aim to learn about the intensive margin effect on the outcome of switching the treatment values over $\mathcal{D}$. For a continuous $D$, let $\mathcal{D} = [\underline{\mathcal{D}}, \overline{\mathcal{D}}]$; for a discrete $D$, let $\mathcal{D} = \big\{\underline{\mathcal{D}} =: d_1, d_2,..., d_J:= \overline{\mathcal{D}}\big\}$ with dimension $J$ smaller or equal to the dimension of $D$.

For any treatment value $d\in \mathcal{D}$ that a subject receives, let $Y_d$ be the continuous potential outcome, or the response function of $d$, and $S_d \in \{0,1\}$ be the potential selection indicator for whether the subject's outcome is observed. If a subject is treated at $d$, i.e., $D = d$, let the selection status $S = S_d$ and the outcome $Y = Y_d$. The observed data vector $W = (D,S,S\cdot Y)$, so the outcome is recorded as zero if missing in the sample.

Following the literature and focusing on the sample selection bias, we begin with the independence Assumption (ref) on treatment assignment. After the key results are established, we consider the standard conditional independence assumption given covariates in Section (ref).

assumption[Independence] $D$ is independent of $\big\{(Y_d, S_d): d\in \mathcal{D}\big\}$.

Under Assumption (ref), we identify the selection probability at treatment $d$, $\mathbb{P}(S_d=1) = \mathbb{E}[S_d] = \mathbb{E}[S|D=d]$, and hence the average treatment effect (ATE) on selection $\mathbb{E}[S_d - S_{d'}] = \mathbb{E}[S|D=d] - \mathbb{E}[S|D=d']$, also known as the extensive margin effect of switching treatment from $d'$ to $d$.

Assumption (ref) also identifies the average outcome of the selected population at $d$, $\mathbb{E}[Y_d|S_d = 1] = \mathbb{E}[Y|S=1, D=d]$. But $\mathbb{E}[Y_d|S_d = 1] - \mathbb{E}[Y_{d'}|S_{d'} = 1]$ is not causal if $\{S_d = 1\}$ and $\{S_{d'} = 1\}$ are different subpopulations. There are two common assumptions for $\mathbb{E}[Y_d|S_d = 1] - \mathbb{E}[Y_{d'}|S_{d'} = 1]$ to capture the intensive margin:\footnote{ We can decompose $\mathbb{E}[Y_d|S_d = 1] - \mathbb{E}[Y_{d'}|S_{d'} = 1] = InM + ExM$, where $InM := \mathbb{E}[Y_d|S_d = 1] - \mathbb{E}[Y_{d'}|S_{d} = 1]$ from the intensive margin and $ExM = \mathbb{E}[Y_{d'}|S_d = 1] - \mathbb{E}[Y_{d'}|S_{d'} = 1]$ from the extensive margin that cause the selection bias. Because $\mathbb{E}[Y_{d'}|S_{d} = 1]$ is not identified, we cannot disentangle $InM$ and $ExM$.} (i) Assume no ATE on selection (no extensive margin), or $\{S_d = 1\} = \{S_{d'}=1\}$ have the same distribution for all $d, d'\in\mathcal{D}$. (ii) Assume missing at random Rubin, so $\mathbb{E}[Y_d|S_d=1] - \mathbb{E}[Y_{d'}|S_{d'}=1]=\mathbb{E}[Y_d - Y_{d'}]$ is the population ATE on the outcome. But these two assumptions can be restrictive and unrealistic.

To understand the source of selection bias, note that the selected population $\{S_d = 1\}$ is composed of always-takers and $d$-compliers. Define always-takers AT$:= \{S_{d'}=1\text{ for all } d'\in \mathcal{D}\}$ to be those selected into samples regardless of the treatment value they receive over $\mathcal{D}$. Define $d$-compliers CP$_d := \{S_d = 1, S_{d'}=0 \text{ for some }d' \in \mathcal{D}\}$ to be those induced to selection due to the treatment value $d$ but are not selected at $d'$. Recall our goal of recovering the intensive margin effect of switching treatment between any values in $\mathcal{D}$. Always-takers are the common subpopulation that are selected into samples for {\it all} treatment values in $\mathcal{D}$, while $d$-compliers are missing in some selected samples with $d'$, $\{D=d', S=1\}$. As we never observe $d$-compliers in some samples with $d'$, it is not possible to learn about their causal effect of switching treatment from $d$ to $d'$. So without further assumptions such as functional form for extrapolation, we could only hope to learn about the causal effect for the always-takers. Therefore our target parameter is the mean potential outcome $Y_d$ for always-takers,

align*[align* omitted — 91 chars of source]

for $d\in\mathcal{D}$. Note that the definition of always-takers depends on the range of treatment values of interest $\mathcal{D}$. So the notation $\beta_d$ should depend on $\mathcal{D}$ that is suppressed for simplicity.

When the treatment is continuous, $\beta_{d{}}$ is known as always-takers' {\it average dose-response function}. When the treatment variable is binary, i.e., $\mathcal{D} = \{0,1\}$, always-takers' ATE is $\beta_{1{}} - \beta_{0{}} = \mathbb{E}[Y_1 - Y_0 \mid S_1=1, S_0=1]$, studied in LeeBound. Nonetheless we cannot determine whether a specific subject is an always-taker or a complier. So we take a bound/partial identification approach following LeeBound and ZR03. Our bounds estimation and inference are robust to the extensive margin effect on selection. GRR propose similar worst-case sharp bounds with manipulation-robust inference in regression discontinuity designs. We also discuss how the sharp bounds might be tightened by our sufficient value/set assumption or the covariates (e.g., FanPark).

There are recent developments in sample selection models using the bound/partial identification approach. HonoreHu and HonoreHuJoE consider parametric and semiparametric structural models. In addition to the concern of misspecification, the parametric selection equation often relies on the monotonicity assumption. Estrada studies spillover effects under sample selection, which can be viewed as Lee bound for the multivalued treatment effect. HEILER24 and Olma study Lee bounds for the conditional average binary treatment effect given a continuous covariate, which is a non-regular estimand as ours. KMV extend Lee bounds for multilayered sample selection to account for training affecting workers sorting to firms. KlineSantos assess the sensitivity of empirical conclusions among a continuum of assumptions ordered from strongest (missing at random) to weakest (worst-case bounds). See the literature reviews on partial identification in HoRosen, MolinariReview, KlineTamer, and references therein.

Alternatively another literature on sample selection models with {\it exclusion restrictions} or partial randomization assumes a variable $Z$ in the selection equation of $S$ that does not enter $Y$. In Heckman's classic sample selection model (“Heckit"), the structural equations are linear in $(D,Z)$ and are separable in the normally distributed errors. AhnPowell, DNV03, and EJL16 consider more nonparametric settings. The standard sample selection model is generally not point-identified without exclusion restrictions. Nevertheless, it has been noted to be difficult to find a credible instrument $Z$ in practice; see, for example, HonoreHu. DiNardo proactively create an instrument by ex-ante randomizing the participants of the Moving to Opportunity experiment to differing intensity of follow-up. BCGL use information on the number of calls made to each individual before responding to the survey to identify the ATE of a binary treatment for a subpopulation of respondents, in the absence of instruments. See also Garlick for evaluation of various sample selection correction methods and references therein. ChenRoth discuss problems of log-like transformations with zeros and propose some solutions. We capture general heterogenous causal effects without exclusion restrictions and free from misspecification.

Lee Bounds

We establish the sharp worst-case bounds for $\beta_d$ with a continuous/multivalued treatment, building on HM95 and LeeBound for a binary treatment. The upper bound is when all always-takers' wages are larger than all $d$-compliers' wages. Denote the fraction of always-takers among the selected subjects with treatment $d$ as $p_d$. Then all always-takers' wages are larger than the $(1-p_d)$-quantile of the observed wage distribution at $d$. So we can construct the worst-case bound by trimming the upper and lower tails of the observed outcome distribution by $p_d$. Next we present the well known worst-case bounds based on a given $p_d$, and then we propose identification strategies of $p_d$ for the continuous and multivalued treatment, which is new to the literature.

Independent treatment Assumption (ref) identifies the selection probability at $d$ by the conditional selection probability given $d$, $\mathbb{P}(S_d = 1) = \mathbb{E}[S|D=d] =: s(d)$. If the proportion of always-takers $\pi_{\text{AT}} := \mathbb{P}(S_{d'} = 1: d'\in\mathcal{D})$ is known, then the fraction of always-takers among the selected subjects with treatment $d$ is $p_d = \mathbb{P}(\text{AT}|S_d = 1) = \pi_{\text{AT}{}}/s(d)$. Let $Q^d(u)$ be the $u$-quantile of $Y|D=d, S=1$. Then the bounds of $\beta_d$ are the trimmed means:

align[align omitted — 120 chars of source]

for the upper bound and $\rho_{dL}(\pi_{\text{AT}{}}) := \mathbb{E}[Y|Y \leq Q^d(\pi_{\text{AT}{}}/s(d)), D=d, S=1]$ for the lower bound. The key element of the bounds is the proportion of always-takers $\pi_{\text{AT}{}}$. Once we identify $\pi_{\text{AT}{}}$ and hence $p_d = \pi_{\text{AT}}/s(d)$, we can consistently estimate the bounds.

Note that when there is no complier, the selected sample is composed of always-takers only, so $s(d) = \mathbb{P}(S_{d'}=1: d'\in\mathcal{D})$ is constant and $p_d = \pi_{\text{AT}{}}/s(d)=1$ for all $d\in\mathcal{D}$. That is, there is no extensive margin, and $\beta_d = \rho_{dU}(\pi_{\text{AT}{}}) = \rho_{dL}(\pi_{\text{AT}{}}) = \mathbb{E}[Y|D=d, S=1]$ is point-identified.

Section (ref) and Section (ref) present novel sufficient assumptions on the treatment effect on selection to identify the proportion of always-takers $\pi_{\text{AT}}$. Section (ref) discusses the connection with the structural selection model. For expositional ease, we focus on the upper bound.

Identification of the proportion of always-takers

The key identification Assumption (ref) requires one {\it sufficient treatment value} $d_{\text{AT}}$ such that if a subject is selected at $d_{\text{AT}}$ then it will be selected at any treatment values.

assumption[Sufficient treatment value] There exists a treatment value $d_{\text{AT}} \in\mathcal{D}$ such that $S_d \geq S_{d_{\text{AT}}}$ almost surely (a.s.) for any $d \in\mathcal{D}$.

Assumption (ref) essentially assumes that always-takers are the selected $d_{\text{AT}}$-receipts, $\{S_{d} =1: d\in\mathcal{D}\} = \{S_{d_{\text{AT}}} = 1\}$. Together with Assumption (ref), the proportion of always-takers $\pi_{\text{AT}} = \mathbb{P}(S_{d_{\text{AT}}} = 1) = \mathbb{E}[S|D=d_{\text{AT}}] =: s(d_{\text{AT}})$ is identified.

Assuming $s(\cdot)$ to be continuous for a continuous $D$, the extreme value theorem implies that $d_{\text{AT}} = \arg\min_{d\in\mathcal{D}} s(d)$ exists and $\pi_{\text{AT}} = s(d_{\text{AT}}) = min_{d\in\mathcal{D}} s(d)$ can be estimated from the data. Notice that $d_{\text{AT}}$ depends on $\mathcal{D}$ that is a range of treatment values chosen by the researcher for policy intervention. A larger $\mathcal{D}$ results in a smaller $\pi_{{\text{AT}}}$, which trims more observations, and wider bounds.

Importantly Assumption (ref) allows any shape of the effect on selection. Consider an example of $d_{\text{AT}} = 1$. A 2-complier can have $\{S_{d_{\text{AT}}} = S_1= 0, S_2= 1, S_3 = 0\}$. In contrast, a stronger monotonicity assumption, which assumes $S_{d'} \geq S_{d}$ a.s.\ for any $d' > d$, rules out this event because it requires $S_3 = 1$ if $S_2 = 1$.

Lemma (ref) formally presents our generalized Lee bounds with a continuous treatment or a discrete multivalued treatment.

lemmaLet Assumption (ref) hold. Assuming $s(d) > 0$ for $d\in\mathcal{D}$, then $\beta_{d{}} \in [\rho_{dL}(\pi_{\text{AT}}), \rho_{dU}(\pi_{\text{AT}})]$ with $\pi_{\text{AT}} = \mathbb{P}(S_{d'} = 1: d' \in\mathcal{D})$ given in equation ((ref)). Further let Assumption (ref) hold. Then we identify $\pi_{\text{AT}} = s(d_{\text{AT}})$, the sharp bounds for $\beta_d \in [\rho_{dL}(s(d_{\text{AT}})), \rho_{dU}(s(d_{\text{AT}}))]$, and $\beta_{d_{\text{AT}}} = \mathbb{E}[Y|D=d_{\text{AT}}, S=1]$.
remark[Binary treatment in LeeBound] Assumption (ref) includes the familiar monotonicity assumption for a binary treatment in LeeBound that assumes $S_1 \geq S_0$ a.s., i.e., if a subject in the control group $\{D=0\}$ is selected, then it remains selected if it was in the treated group $\{D=1\}$, i.e, $d_{\text{AT}} = 0$. So defiers (0-compliers) are excluded ImbensAngrist. Then Lemma (ref) implies Proposition 1a in LeeBound: the upper bound of the always-takers' ATE, $\beta_1 - \beta_0 = \mathbb{E}[Y_1 - Y_0| S_0=S_1=1]$, is $\rho_{1U}(s(0)) - \beta_{d_{\text{AT}}} = \mathbb{E}[Y|Y \geq Q^d(1-s(0)/s(1)), D=1, S=1] - \mathbb{E}[Y|D=0, S=1]$ and the lower bound is $\rho_{1L}(s(0)) - \beta_{d_{\text{AT}}} = \mathbb{E}[Y|Y \leq Q^d(s(0)/s(1)), D=1, S=1] - \mathbb{E}[Y|D=0, S=1]$. We remark that if we are only interested in two values of the continuous treatment, then the identification of the bounds for the binary treatment in LeeBound can be directly applied to the case of two continuous treatment values. But policy makers rarely consider only two values, and the always-takers at the two values could be different from the always-takers at another values; for example, $\{S_{d_1} = 1, S_{d_2}=1\} \neq \{S_{d_3} = 1, S_{d_4} = 1\}$. New challenge in identification arises when we consider many treatment values.

Sufficient set assumption

We introduce a weaker sufficient set Assumption (ref) that does not assume one sufficient treatment value $d_{\text{AT}}$ but assumes a set $\mathcal{D}_M$ of $M$ treatment values, which includes Assumption (ref) as a special case with $M=1$. Assumption (ref) essentially assumes that always-takers $\{S_{d} =1: d\in\mathcal{D}\} = \{S_{d} = 1: d\in\mathcal{D}_M\}$.

assumption[Sufficient set] There exists a set of treatment values $\mathcal{D}_M := \big\{d_1, d_2,..., d_M \big\} \subseteq \mathcal{D}$ such that $S_d \geq \min_{d' \in\mathcal{D}_M} S_{d'}$ a.s.\ for any $d\in\mathcal{D}$.

To see how Assumption (ref) is weaker than Assumption (ref) or a larger $M$ is weaker, consider an example: Under Assumption (ref) with $M=3$ and $\mathcal{D}_3 = \{1,2,3\}$ , $\{S_1=0, S_2 = 1, S_3 = 0\}$ and $\{S_1=1, S_2 = 0, S_3 = 1\}$ are both possible values for a complier. But Assumption (ref) with $d_{\text{AT}} = 1$ does not allow a complier to take $\{S_1=1, S_2 = 0, S_3 = 1\}$. What Assumption (ref) does not allow is this event $\{S_1=1, S_2 = 1, S_{2.5} = 0, S_3 = 1\}$ for example. But assuming a larger $M=4$ and $\mathcal{D}_4 = \{1,2,2.5,3\}$ would allow a complier to take that.

However, the proportion of always-takers $\pi_{\text{AT}} = \mathbb{P}(S_{d} = 1: d\in\mathcal{D}_M)$ for $M \geq 2$ is not point-identified as we cannot observe the $M$ potential outcomes $\{S_{d}: d \in\mathcal{D}_M\}$ at the same time. We can use the lower bound of $\pi_{\text{AT}}$ for trimming, because $\rho_{dU}(\pi_{\text{AT}{}})$ is decreasing in $\pi_{\text{AT}{}}$ and $\rho_{dL}(\pi_{\text{AT}{}})$ is increasing in $\pi_{\text{AT}{}}$. Theorem (ref) below provides the Lee bounds formally.

Specifically, consider a practical example of $M=2$ and $\mathcal{D}_2 = \{\underline{\mathcal{D}}, \overline{\mathcal{D}}\}$. It implies that if a subject is selected at the boundary $\underline{\mathcal{D}}$ and $\overline{\mathcal{D}}$, then this subject must be selected at any treatment value in $\mathcal{D}$. This example of $M=2$ includes the special case when the selection's response is a concave function of treatment $D$. The single-peaked pattern may fit well the law of marginal returns. Therefore $S_{\underline{\mathcal{D}}} = 1$ and $S_{\overline{\mathcal{D}}} = 1$ if and only if $S_{d} = 1$ for all $d\in\mathcal{D}$. This observation gives the insight to use the well-known Fr\'{e}chet-Hoeffding bounds for $\mathbb{P}(S_{\underline{\mathcal{D}}} = 1, S_{\overline{\mathcal{D}}} = 1)$ given in Theorem (ref). The bounds resemble the Fr\'{e}chet-Hoeffding bounds for $\mathbb{P}(S_0 = 1, S_1=1)$ for the binary treatment without shape restrictions, as shown in HSC. We require the sufficient set Assumption (ref) due to the continuous treatment variable. It is important to note that when the treatment is multivalued discrete with support $\mathcal{D} = \mathcal{D}_M$, Assumption (ref) holds by construction and is dropped. So Theorem (ref) provides the sharp bounds without restricting the selection response of the multivalued treatment.

theoremLet $\mathcal{D}_M = \big\{d_1, d_2,..., d_M\big\}\subseteq \mathcal{D}$ with a fixed dimension $M$. Let Assumption (ref) hold. Then \begin{align*} \pi_{L}^M := \max\left( \sum_{d\in \mathcal{D}_M} s(d) - M + 1, 0\right) \leq \mathbb{P}(S_{d}=1: d\in\mathcal{D}_M) \leq \min_{d\in \mathcal{D}_M} s(d) =: \pi_{U}^M. \end{align*} Further let Assumption (ref) hold. Then $\pi_{\text{AT}} = \mathbb{P}(S_{d}=1: d\in\mathcal{D}_M) \in \big[\pi_L^M, \pi_U^M\big]$ and $\beta_{d} \in \big[\rho_{dL}(\pi_{L}^M), \rho_{dU}(\pi_{L}^M)\big]$. The bounds are sharp.

A final goal is to derive the bounds without restrictions on the selection response of a continuous treatment, i.e., dropping Assumption (ref). To make progress, we may assume that the treatment effect is a piecewise constant function $\beta_d = \sum_{m=1}^{M-1} \beta_{d_m} \mathbf{1}\{d \in [d_m, d_{m+1})\}$. Then treatment can be effectively discretized and $\beta_{d_m} \in [\rho_{d_mL}(\pi_{L}^M), \rho_{d_mU}(\pi_{L}^M)]$ for $m=1,..,M-1$. In practice, one might discretize the continuous treatment variable into $M$-multivalued variable for sensitivity analysis.

In general and in theory, we'd like $M$ to be large to allow for a general non-separable nonparametric structural selection model, as discussed in Section (ref). However, the bounds can be wide or less informative for a large $M$. As shown in Theorem (ref), $\pi_L^M$ can be small, unless $s(d)$ is close to one when there is selection bias. We illustrate this tradeoff in Section (ref) by evaluating the Job Corps program and discuss how to choose $M$.

remark[Binary outcome] {\rm KMV provide the Lee bounds with a binary outcome and a binary treatment. We can extend their bounds for a binary outcome to a continuous treatment: $\rho_{dL}(\pi_L^M) = \max\{0, 1- \mathbb{P}(Y = 0|S=1, D=d)/p_d\}$ and $\rho_{dU}(\pi_L^M) = \min\{1, \mathbb{P}(Y=1|S=1, D=d)/p_d\}$, where $p_d = \pi_L^M/s(d)$.\footnote{ $\mathbb{P}(Y=1|S=1, D=d) = \mathbb{P}(Y=1|\text{AT}, D=d) p_d + \mathbb{P}(Y=1|\text{CP}_d, D=d) (1-p_d)$. The bounds on $ \mathbb{P}(Y=1|\text{AT}, D=d) = \mathbb{E}[Y_d|\text{AT}]$ are obtained by the worst-case bounds of $\mathbb{P}(Y=1|\text{CP}_d, D=d)\in [0,1]$. } We focus on the continuous outcome in this paper and develop the inference for the binary outcome in a separate paper. }

Structural selection equation

To understand how the sufficient treatment value Assumption (ref) and the sufficient set Assumption (ref) impose conditions on the heterogeneity in the structural selection equation, we discuss its relationship with the threshold-crossing model in Vytlacil02. Recall that our potential outcome framework is equivalent to the structural equation $S = \mathbf{1}\{q(D,\eta) \geq 0\}$ with unobserved non-separable and multi-dimensional error $\eta$. The structural equation $q$ is nonparametric and model-free. Then we can write the potential variable $S_d =\mathbf{1}\{q(d, \eta) \geq 0\}$. \\[5pt] {\bf Assumption 2$^\prime$ (Latent index selection model)} {\it (i) Let $S = \mathbf{1}\{q(D) \geq \eta\}$, where $q(d)$ is measurable and nontrivial function of $d$. (ii) For a continuous treatment with a compact $\mathcal{D}$, there exists $d_{\text{AT}} = \arg\inf_{d\in\mathcal{D}} q(d) \in \mathcal{D}$.}

Assumption 2$^\prime$ implies Assumption (ref). For a multivalued treatment with a finite countable $\mathcal{D}$, $d_{\text{AT}}$ exists under Assumption 2$^\prime$(i). For a continuous $D$, assuming $q(\cdot)$ in Assumption 2$^\prime$(i) to be continuous, the extreme value theorem implies (ii). This important observation suggests our new sufficient treatment value Assumption (ref) not restrictive and implied by a common threshold-crossing model with a separable error, or the latent index selection model in Vytlacil02.\footnote{ Vytlacil02 shows that the latent index selection model Assumption 2$^\prime$(i) is equivalent to the local average treatment effect (LATE) model with Independence (as our Assumption (ref)) and Monotonicity assumptions in ImbensAngrist. The LATE Monotonicity assumes the orders of $S_d$ to be the same for everyone, i.e., for all $(d, d')\in\mathcal{D}\times\mathcal{D}$, either $S_d \geq S_{d'}$ $a.s.$, or $S_d \leq S_{d'}$ $a.s$. Our Assumption (ref) (or Assumption 2$^\prime$(ii)) is weaker than such Monotonicity assumption and only requires the sufficient treatment value $d_{\text{AT}}$ (or a minimizer of $q(d)$) exists. }

In the selected sample at $d$, $\{S=1, D=d\} = \{S_d=1, D=d\}$, the subpopulation $\{S_d=1\} = \{\eta \leq q(d)\} = \text{AT} \cup \text{CP}_d$, where always-takers $\text{AT} = \{\eta \leq q(d_{\text{AT}})\}$ and $d$-compliers $\text{CP}_d = \{q(d_{\text{AT}}) < \eta \leq q(d)\}$. The sufficient treatment value has meaningful economic interpretation. For example, BCGL interpret $\eta$ as the individual reluctance to respond to surveys and call the compliers as the marginal respondents. Then the sufficient treatment value can be the least-favored treatment value to induce responding to surveys. So if subjects are willing to respond to surveys when receiving $d_{\text{AT}}$, then they continue responding to surveys when receiving any other treatment values.

To further appreciate always-takers as the target population, we discuss the gettable ATE (GATE) $\mathbb{E}[Y_1 - Y_0|\eta \leq k]$ for some constant $k$, defined by DiNardo. As $k$ increases, GATE converges to the population ATE $\mathbb{E}[Y_1 - Y_0]$. As subpopulation-specific ATEs are commonplace, DiNardo show how different GATE parameters may be identified under weaker assumptions than in the traditional parametric framework. We choose $k = q(d_{\text{AT}})$ to characterize always-takers that are the largest subpopulation for whom we can partially identify the ATE of switching treatment values over $\mathcal{D}$, without imposing further assumptions on the functional forms or distributions.

Now we consider a more general non-separable nonparametric model in Assumption 3$^\prime$ that implies our sufficient set Assumption (ref). \\[5pt] {\bf Assumption 3$^\prime$ (Latent index selection model with non-separable errors)} {\it Let $S = \mathbf{1}\{q(D,\eta) \geq 0\}$. There exists $d_{\text{AT}}(\eta) = \arg\inf_{d\in\mathcal{D}} q(d, \eta) \in \mathcal{D}_M$ for each $\eta$.}

The unobserved heterogeneity is captured by $\eta$, so there could be an infinite number of types and the corresponding sufficient value $d_{\text{AT}}(\eta)$ in the most general structural model. Our sufficient set Assumption (ref) restricts there to be $M$ types of unobserved heterogeneity $\eta$, in the sense that $d_{\text{AT}}(\eta)$ belongs to $\mathcal{D}_M$.

Under Assumption (ref)$^\prime$, Assumption (ref) can be implied by further assuming $d_{\text{AT}}(\eta) = d_{\text{AT}}$ to be a constant for all individual with $\eta$, and $M=1$. Therefore we argue that the sufficient value Assumption (ref) is reasonable with a separable structural error in selection as in Assumption 2$^\prime$, and the sufficient set Assumption (ref) is a useful approximation under more general non-separable errors.

Estimation and inference

We estimate bounds over an equally spaced grid $\mathcal{D}_J = \{d_1,..,d_J\} \subset \mathcal{D}$ for a continuous treatment. We can view $\mathcal{D}_J$ as the set of treatment values where the policy maker considers treatment intervention. The estimation procedure is easy to implement, as the bounds estimates are sample analogs to the parameters defined in Theorem (ref). When $D$ is continuous, we use a kernel function $K_h(D-d) = k((D-d)/h)/h$, where the kernel function $k$ includes a sub-population whose treatment is around $d$, and the size of the sub-population is controlled by the bandwidth $h$ shrinking to zero as the sample size grows. We provide inference on the causal effect of switching treatment between any two values in $\mathcal{D}_J$. We present the asymptotic theory for a fixed $J$ and also for $J \rightarrow \infty$ so $\mathcal{D}_J \rightarrow \mathcal{D}$.

For a multivalued discrete treatment, it is straightforward to use the treatment indicator $\mathbf{1}\{D=d\}$ in place of the binary treatment indicator $D$ in the estimator in LeeBound and let $\mathcal{D}_J$ be the (sub)support of $D$. We develop the inference theory for a discrete treatment in a separate paper.

We implement the estimation procedure using leave-out estimators for $s(d)$ and $Q^d$ in Step 1 and Step 2. The leave-out estimation is similar to the cross-fitting in the recent double debiased machine learning literature CCDDHNR. The leave-out preliminary estimation achieves stochastic equicontinuity without strong entropy conditions using empirical process theory. Specifically, for some fixed $L \in \{2,...,n\}$, randomly partition the observation indices into $L$ distinct groups $I_\ell, \ell = 1,...,L$, such that the sample size of each group is the largest integer smaller than $n/L$. The number of folds $L$ is not random and typically small, such as five or ten in practice; see, e.g., CCDDHNR, Velez. When there is no sample splitting ($L = 1$), $\hat s_1(d)$ and $\hat Q^d_1$ use all observations in the full sample.\footnote{ When $L = n$, $\hat s_\ell$ uses all observations except for the $\ell^{th}$ observation, and is well-known as the leave-one-out estimator (i.e., leave the $\ell^{th}$ observation out), e.g., PSS89ETA. A large $L$ is computationally costly.}

The estimation procedure under Assumption (ref) follows four steps:

\paragraph{Step 0.} Estimate the sufficient treatment value $\hat d_{\text{AT}_J} = \arg\min_{d \in \mathcal{D}_J} \hat s(d)$, where a kernel estimator $\hat s(d) = \sum_{i=1}^n S_i K_h(D_i - d)/\sum_{i=1}^n K_h(D_i - d)$. For $d = \hat d_{\text{AT}_{J}}$, $\hat\beta_{d} = \sum_{i=1}^n Y_i S_i K_h(D_i - d)/\sum_{i=1}^n S_i K_h(D_i - d)$. Estimate the proportion of always-takers $\pi_{\text{AT}}$ by $\hat\pi = \min_{d\in\mathcal{D}_J}\hat s(d) = \hat s(\hat d_{\text{AT}_J})$.

For $\ell = 1,..., L$, the estimators in Step 1 and in Step 2 use observations not in $I_\ell$, denoted as $I_\ell^c := \{1,...,n\} \setminus I_\ell$. \paragraph{Step 1.} Compute the leave-out kernel estimator $\hat s_\ell(d) = \sum_{i\in I_\ell^c} S_i K_h(D_i - d)/\sum_{i\in I_\ell^c} K_h(D_i - d)$. Estimate the trimming probability $p_d$ by $\hat p_\ell = \min\{\hat s_\ell(\hat d_{\text{AT}_J})/\hat s_\ell(d), 1\} -\nu$ for some small positive constant $\nu$ used for robust inference that we explain in the following.

\paragraph{Step 2.}

For $d \neq \hat d_{\text{AT}_{J}}$, estimate the $(1-\hat p_\ell)$-quantile of $Y|D=d, S=1$ by $\hat Q_\ell^d(1-\hat p_\ell)$. A nonparametric estimator can be the generalized inverse function of the CDF estimate $\hat F_{Y|D=d,S=1_\ell}(y) = \sum_{i\in I_\ell^c}\mathbf{1}\{Y_i \leq y\} S_i K_h(D_i - d)/\sum_{i\in I_\ell^c} S_i K_h(D_i - d)$.

\paragraph{Step 3.}

Compute the kernel estimator of $\mathbb{E}[Y{\bf 1}\{Y \geq Q^d(1-p)\}| D=d, S=1]$ and obtain

align*[align* omitted — 201 chars of source]

where $\hat p = \hat\pi/\hat s(d) - \nu$ using the full sample from Step 0.

Our inference procedure is robust to extensive margin, allowing the selection probability to be of any unknown functional form with respect to treatment. The challenge for the continuous or multivalued treatment relative to the well studied binary treatment case comes from an infinite or multiple number of treatment values ($J$) and the corresponding potential selections and outcomes. Importantly, we allow the sufficient treatment value to be not unique in the sense that there exists a subset $\mathcal{D}_c \subseteq\mathcal{D}$ such that $d= \arg\min_{d'\in\mathcal{D}} s(d')$ for any $d \in \mathcal{D}_c$ and $\mathbb{P}(D \in \mathcal{D}_c) > 0$. So $s(d)$ is constant over $d\in\mathcal{D}_c$, which is implied by no treatment effect on selection (extensive margin) when changing $d$ within the subset $\mathcal{D}_c$. Then we could use any $d\in\mathcal{D}_c$ as a sufficient treatment value $d_{\text{AT}}$. So for any $d \in \mathcal{D}_c$, $s(d_{\text{AT}})/s(d)=1$ and $\beta_d = \rho_{dU}(\pi_{\text{AT}}) = \rho_{dL}(\pi_{\text{AT}})= \mathbb{E}[Y|D=d, S=1] $ is point-identified. However the asymptotic distributions of the bound-estimators $\big[\widehat{\rho_{dL}(\pi_{\text{AT}})}, \widehat{\rho_{dU}(\pi_{\text{AT}})}\big]$ do not converge to the asymptotic distribution of the point-estimator $\hat\beta_d$ as $s(d_{\text{AT}})/s(d)\rightarrow 1$. We avoid the complication in testing $s(d) = s(d_{\text{AT}})$, e.g., if the hypothesis $s(d_{\text{AT}})/s(d)=1$ is not rejected, then compute the point-estimator $\hat \beta_d$. Instead, we estimate the tight bounds using the trimming probability $\hat p_\ell = \min\{\hat s_\ell(\hat d_{\text{AT}_{J}})/\hat s_\ell(d), 1\} - \nu \stackrel{p}{\rightarrow} p_d = s(d_{\text{AT}})/s(d) - \nu \leq 1-\nu < 1$, which contain the untrimmed point-estimator $\hat\beta_{d} = \sum_{i=1}^n Y_i S_i K_h(D_i - d)/\sum_{i=1}^n S_i K_h(D_i - d)$.\footnote{ In the proof of Theorem (ref), we show that $\hat d_{\text{AT}_{J\ell}} = \arg\min_{d\in\mathcal{D}_J}\hat s_\ell(d) = \hat d_{\text{AT}_J}$ when $n$ large enough. So $\hat p_\ell = \hat s_\ell( \hat d_{\text{AT}_J}) /\hat s_\ell(d) - \nu$. However, in finite samples, it is possible that $\hat s_\ell(\hat d_{\text{AT}_J})/\hat s_\ell(d) > 1$ for some $d\in\mathcal{D}_J$. In such case, if $\hat d_{\text{AT}_{J}}$ and $\hat d_{\text{AT}_{J\ell}}$ are in $\mathcal{D}_c$, then we let $\hat p_\ell \leq 1-\nu$ and estimate bounds of $\beta_{ \hat d_{\text{AT}_{J\ell}}}$ that contain its point-estimate. }

Therefore we estimate bounds that are non-sharp with the trimming probability $p_d = s(d_{\text{AT}})/s(d) - \nu$ but still tight with small $\nu$. We choose this practical and conservative strategy so that our asymptotic theorem and inference are valid regardless of the extensive margin effect on selection, and are easy to implement and interpret.

We show in Theorem (ref) that the estimation errors of $\hat d_{\text{AT}_{J}}$ and grid approximation are asymptotically ignorable. That is, for any given $J$ and for $n$ large enough, $\hat d_{\text{AT}_{J}} = d_{\text{AT}_J} := \arg\min_{d \in \mathcal{D}_J} s(d)$. So for a any fixed set of treatment values $\mathcal{D}_J$, we can find the sufficient treatment value $d_{\text{AT}_J}$ given a large enough sample. As $\mathcal{D}$ contains an infinite number of values for a continuous treatment, we give conditions on the grid size $J$ going to infinity to approximate $\mathcal{D}$. We show that as $J, n\rightarrow \infty$, $d_{\text{AT}_J} \rightarrow d_{\text{AT}} := \arg\min_{d \in\mathcal{D}} s(d)$. Assumption (ref) below gives conditions on $J$ that depends on the accuracy of $\hat s_\ell(d)$ and the shape of $s(d)$ characterized by $\bar M$.

assumptionLet $s^{(m)}(d)$ be the $m^{th}$ derivative of $s(d)$ for $m \in \{1,2,...\}$. Let $\mathcal{D} = \mathcal{D}_s \cup \mathcal{D}_c$, where $\mathcal{D}_c := \{d: s^{(m)}(d) = 0, \forall m \geq 1\}$ and $\mathcal{D}_s := \{d: s^{(m)}(d) \neq 0, \exists m < \infty\}$. If $\mathcal{D}_s \neq \emptyset$, let $\bar M = \min\{m: s^{(m)}(d) \neq 0, m=1,2,..., \forall d \in \mathcal{D}_s\} < \infty$. If $\mathcal{D}_s = \emptyset$, let $\bar M = 0$. Let an equally spaced grid $\mathcal{D}_J = \{d_1,..,d_J\} \subseteq \mathcal{D}$ with $J = O({\mathsf{s}_n}^{-1/\bar M})$ and $\mathsf{s}_n := \sup_{d\in\mathcal{D}}|\hat s(d)- s(d)| = o_\mathbb{P}(1)$.

$\mathcal{D}_c$ is the set of treatment values not affecting the selection. For $d \in \mathcal{D}_c$, $s(d) = c$ for some generic constant $c$, so $s'(d) = 0$. For $\bar M = 0$ ($\mathcal{D} = \mathcal{D}_c$), there is no extensive margin, so $\beta_d = \mathbb{E}[Y|D=d, S=1]$ is point-identified, and there is no restriction on $J \rightarrow \infty$. When $s(d)$ is strictly monotone over $\mathcal{D}$, $s'(d) \neq 0$, so $\bar M = 1$ and $\mathcal{D}_c = \emptyset$. When $s(d)$ is strictly concave, i.e., $s'(d) = 0$ for a $d\in\mathcal{D}$ and $\mathcal{D}_c = \emptyset$, we have $\bar M = 2$. A larger $\bar M$ implies that it is harder to compare $s(d_j)$ and $s(d_{j+1})$, so we need to estimate $s(d)$ more accurately, i.e., $\mathsf{s}_n$ needs to go to zero faster for a larger $\bar M$ on a given grid.

The kernel estimation is well-studied, and we use the results in DHB and HansenBook.

assumption\begin{enumerate} • The kernel function $k$ is non-negative symmetric bounded kernel with a compact support such that $\int k(u) du = 1$, $\int u k(u) du = 0$, and $\kappa := \int u^2 k(u) du < \infty$. Let the roughness of the kernel be $R_k := \int k(u)^2du$. • $h\rightarrow 0$, $nh\rightarrow\infty$, $nh^5 \rightarrow c\in [0,\infty)$. • For $d \in \mathcal{D}$ and $y\in\mathcal{Y}$, $s(d) < 1$; $f_{Y|DS}(y|d,1)$ is continuous and bounded away from $0$; $F_{Y|DS}(y|d,1) s(d) f_D(d)$ is bounded and has bounded continuous second derivative with respect to $d$; $var(Y|D=d, S=1)$, $\mathbb{E}\big[|Y|^3\big|D=d, S=1\big]$, and the second derivative of $\mathbb{E}[Y|D=d, S=1]$ are continuous in $d$. \end{enumerate}
theoremLet Assumptions (ref), (ref), and (ref) hold. \begin{enumerate} • Under Assumption (ref), $\pi = s(d_{\text{AT}_J})$. Then for $d\in\mathcal{D}_J$ and $d \neq \hat d_{\text{AT}_J}$, as $n\rightarrow \infty$, \begin{align} \sqrt{nh}\Big(\widehat{\rho_{dU}(\pi)} - \rho_{dU}(\pi) - h^2{B}_{dU}\Big) &\stackrel{d}{\rightarrow}\mathcal{N}(0,V_{dU}) \notag \\ \sqrt{nh}\Big(\widehat{\rho_{dL}(\pi)} - \rho_{dL}(\pi) - h^2{B}_{dL}\Big) &\stackrel{d}{\rightarrow}\mathcal{N}(0,V_{dL}), \end{align} where $p = p_d = s(d_{\text{AT}_J})/s(d) - \nu$, $V_{dU} := p^{-2} \big(V_1 + V_2 + V_{3} + V_{23}\big) R_k /(s(d)f_D(d))$, $V_{3} := var(Y{\bf 1}\{Y \geq Q^d(1-p)\}|D=d, S=1), V_2 := p(1-p)Q^d(1-p)^2, V_{23} := -2p(1-p) \rho_{dU}(\pi) Q^d(1-p), V_1:= \big(V_\pi+ p^2 V_{s(d)}\big)s(d)^{-1} \big( Q^d(1-p) - \rho_{dU}(\pi) \big)^2,$ with $V_{s(d)} := s(d)(1-s(d))$ and $V_\pi = V_{s(d_{\text{AT}_J})}f_D(d)/f_D(d_{\text{AT}_J})$. For the lower bound, $V_{dL} := p^{-2} \big(V_1 + V_2 + V_{3}^L + V_{23}\big) R_k /(s(d)f_D(d))$, where $V_{3}^L := var(Y{\bf 1}\{Y \leq Q^d(p)\}|D=d, S=1)$, and $V_1$, $V_2$, $V_{23}$ are defined as above with $Q^d(p)$ in place of $Q^d(1-p)$ and with $\rho_{dL}(\pi)$ in place of $\rho_{dU}(\pi)$. ${B}_{dU}$ and ${B}_{dL}$ are given explicitly in the proof in the Appendix. For $d = \hat d_{\text{AT}_J}$, $\sqrt{nh}\big(\hat\beta_{d_{\text{AT}_J}} - \beta_{d_{\text{AT}_J}} - h^2 B_3\big) \stackrel{d}{\rightarrow} \mathcal{N}\big(0, V_{d_{\text{AT}_J}}\big)$ as $n \rightarrow \infty$, where $V_{d_{\text{AT}_J}} = var(Y|D=d_{\text{AT}_J}, S=1) R_k /(s(d_{\text{AT}_J})f_D(d_{\text{AT}_J}))$. As $J \rightarrow\infty$, $d_{\text{AT}_J} \rightarrow d_{\text{AT}}$ and the above statements hold with $d_{\text{AT}}$ in place of $d_{\text{AT}_J}$. • Under Assumption (ref), choose $\mathcal{D}_M$ to contain $\hat d_{\text{AT}_J}$. Let $\hat p_\ell = \hat\pi_{L\ell}^M/\hat s_\ell(d)$ in Step 1, and follow Step 2 and Step 3 to obtain the bounds $[\widehat{\rho_{dL}(\pi)}, \widehat{\rho_{dU}(\pi)}]$, where $\pi = \sum_{d\in\mathcal{D}_M} s(d) - M + 1 > 0$. Let $V_\pi = \sum_{m=1}^M V_{s(d_m)} f_D(d)/f_D(d_m)$. For $d \in \mathcal{D}_M^c \cap\mathcal{D}_J$, the above asymptotic distributions ((ref)) hold. For $d \in \mathcal{D}_M \cap\mathcal{D}_J$, the above asymptotic distributions ((ref)) hold with $V_1:=\left(V_\pi + (p^2-2p) V_{s(d)}\right)s(d)^{-1}\left(Q^d(1-p) - \rho_{dU}(\pi)\right)^2$. \end{enumerate}

Notice that we allow for no extensive margin, e.g., there exists $d \neq d_{\text{AT}_J}$ and $s(d) = s(d_{\text{AT}_J})$. So we use $p_d = 1 - \nu$ to estimate tight bounds of $\beta_d$, rather than a point-estimand.

The asymptotic variance $V_{dU}$ can be decomposed to the three steps of the estimation procedure: $V_1$ is from estimating the effect on selection and the trimming probability in Step 1. $V_2$ is from the quantile regression in Step 2. $V_3$ comes from the Step 3 trimmed regression. $V_{23}$ is from the covariance of the Step 2 and Step 3 estimation errors.

Estimation of the variances is easily carried out by replacing all of the above quantities with their sample analogs or by the sample variances of the influence functions given in equation ((ref)) in the proof of Theorem (ref) in the appendix. No additional preliminary estimators are needed. The confidence interval of at least 95% coverage can be computed by $\Big[\widehat{\rho_{dL}(\pi)}-1.96 \times \widehat{\sigma_{dL}}/\sqrt{nh}, \widehat{\rho_{dU}(\pi)} + 1.96 \times \widehat{\sigma_{dU}}/\sqrt{nh} \Big]$ with $\widehat{\sigma_{dL}} := \sqrt{\widehat{V_{dL}}}$ and $\widehat{\sigma_{dU}} := \sqrt{\widehat{V_{dU}}}$. This interval will asymptotically contain the region $\left[\rho_{dL}(\pi), \rho_{dU}(\pi)\right]$ with at least 95% probability.

We can bound the ATE of increasing the treatment from $d_1$ to $d_2$, $\underline{\Delta}_{d_1d_2} := \rho_{d_2L}(\pi) - \rho_{d_1U}(\pi) \leq \beta_{d_2} - \beta_{d_1} \leq \rho_{d_2U}(\pi) - \rho_{d_1L}(\pi) =: \bar\Delta_{d_1d_2}$. We note that one might be interested in the partial (or marginal) effect defined as the derivative of $\beta_d$ with respect to $d$, i.e., $\theta_d = \frac{\partial}{\partial d}\beta_d$. However, we could not bound such derivative by the same approach of LeeBound. We can view $\Delta_{d_1d_2} := \beta_{d_2} - \beta_{d_1} = \int_{d_1}^{d_2}\theta_s ds$ as an average derivate over $d_1$ to $d_2$. Corollary (ref) provides the asymptotic distribution of the bounds estimators for the ATE $\hat{\underline{\Delta}}_{d_1d_2} := \widehat{\rho_{d_2L}(\pi)} - \widehat{\rho_{d_1U}(\pi)}$ and $\hat{\bar\Delta}_{d_1d_2} := \widehat{\rho_{d_2U}(\pi)} - \widehat{\rho_{d_1L}(\pi)}$. Denote the bandwidth in $ \widehat{\rho_{dL}(\pi)}$ and $\widehat{\rho_{dU}(\pi)}$ be $h_{dL}$ and $h_{dU}$, respectively.

corollary[ATE] Let the conditions in Theorem (ref) hold. Then $\sqrt{n}V_{Un}^{-1/2}\big(\hat{\bar\Delta}_{d_1d_2} - \bar\Delta_{d_1d_2} - (h_{d_2U}^2{B}_{d_2U} -h_{d_1L}^2{B}_{d_1L})\big) \stackrel{d}{\longrightarrow} \mathcal{N}(0, 1)$ and $\sqrt{n} V_{Ln}^{-1/2}\big(\hat{\underline{\Delta}}_{d_1d_2} - \underline{\Delta}_{d_1d_2}- (h_{d_2L}^2{B}_{d_2L} -h_{d_1U}^2{B}_{d_1U})\big) \stackrel{d}{\longrightarrow} \mathcal{N}(0, 1)$, where $V_{Un} := \mathbb{E}[(\phi_{d_2U} - \phi_{d_1L})^2]$ and $V_{Ln}:= \mathbb{E}[(\phi_{d_2L} - \phi_{d_1U})^2]$ with the influence functions $\phi_{dL}, \phi_{dU}$ given explicitly in equation ((ref)) in the Appendix. Assume Assumption (ref). When $h_{d2U} = h_{d_1L} = h$, $\lim_{n\rightarrow \infty, h\rightarrow 0} h V_{Un} = \mathsf{V}_{d_2U} + \mathsf{V}_{d_1L} - 2\mathsf{C}_{d_1d_2U}$. When $h_{d_2L}=h_{d_1U} = h$, $\lim_{n\rightarrow \infty, h\rightarrow 0} h V_{Ln} = \mathsf{V}_{d_2L} + \mathsf{V}_{d_1U} -2 \mathsf{C}_{d_1d_2L}$, where $p_{d} = \pi/s(d)$, $\mathsf{C}_{d_1d_2U}:= R_kV_\pi (Q^{d_2}(1-p_{d_2}) - \rho_{d_2U}(\pi))(Q^{d_1}(p_{d_1}) - \rho_{d_1L}(\pi))/(\pi^2f_D(d_{\text{AT}}))$, and $\mathsf{C}_{d_1d_2L}:= R_kV_\pi (Q^{d_1}(1-p_{d_1}) - \rho_{d_1U}(\pi))(Q^{d_2}(p_{d_2}) - \rho_{d_2L}(\pi))/(\pi^2f_D(d_{\text{AT}}))$.

The variance can be estimated by the plug-in sample analogues $\hat{\mathsf{V}}_{Un} = n^{-1}\sum_{i=1}^n(\hat\phi_{d_2Ui} - \hat\phi_{d_1Li})^2$ and $\hat{\mathsf{V}}_{{Ln}} = n^{-1}\sum_{i=1}^n(\hat\phi_{d_2Li} - \hat\phi_{d_1Ui})^2$. And the 95% confidence interval $\big[\hat{\underline{\Delta}}_{d_1d_2} - 1.96\times \sqrt{\hat{\mathsf{V}}_{Ln}}/\sqrt{n}, \hat{\bar\Delta}_{d_1d_2} + 1.96\times \sqrt{\hat{\mathsf{V}}_{Un}}/\sqrt{n} \big]$.

Next we briefly discuss how to choose an undersmoothing bandwidth $h$ smaller than the optimal bandwidth that minimizes the asymptotic mean squared error (AMSE) such that the bias is first-order asymptotically negligible, i.e., $h^2\sqrt{nh}\rightarrow 0$, so the above confidence interval is valid.

We could estimate the leading bias ${B}_{dU}$ by the method in PS96 and CL. Let the notation $\widehat{\rho_{dU,b}(\pi)}$ be explicit on the bandwidth $b$ and $\widehat{{B}}_{dU}:= \big(\widehat{\rho_{dU,b}(\pi)} - \widehat{\rho_{dU,ab}(\pi)}\big)\big/\big(b^2(1-a^2)\big)$ with a pre-specified fixed scaling parameter $a\in(0,1)$. From the proof of Theorem 3.2 in CL, we can choose $a$ by minimizing the leading term of $var(\hat{B}_{dU})$, i.e., minimizing $(1-a^2)^{-2}a^{-d_T}$ for a $d_T$-dimensional continuous treatment. By deriving the first-order and second-order conditions, we obtain the minimizer $a^*= \sqrt{d_T/(d_T+4)} = \sqrt{1/5}$, for $d_T = 1$ in our case.

Then we propose a data-driven bandwidth $\hat h_{dU}:= (\widehat{V_{dU}}/(4\widehat{{B}}_{dU}^2))^{1/5}n^{-1/5}$ to consistently estimate the AMSE optimal bandwidth $h_{dU}^\ast$ by Theorem 3.2 in CL. For the lower bound, the same estimation applies to the leading bias $\widehat{{B}}_{dL}$ and the AMSE optimal bandwidth $\hat h_{dL}$.

Empirical illustration: Job Corps

We illustrate our method by evaluating the Job Corps program. We use the Job Corps dataset in HHLLdata. The continuous treatment variable ($D$) is the total hours spent in academic and vocational training. The outcome variable ($Y$) is the weekly earnings in the fourth year. Our sample consists of 4,024 subjects who completed at least 40 hours (one week) of training. In the online appendix, Figure (ref) shows the distribution of $D$, and Table (ref) provides brief descriptive statistics. As our analysis builds on FFGN12ReStat, HHLP, HHLL, we refer the readers to the reference therein for further details of Job Corps.

We use $J=100$ grid points over $\mathcal{D}= [40, 2400]$ that ranges from one week to fifteen months. We use the Epanechnikov kernel with a undersmoothing bandwidth and ten-fold cross-fitting. The undersmoothing bandwidth at each $d$ is chosen by $0.8\times \min(\hat h_{dL},\hat h_{dU})$, given in Section (ref), where we use the rule-of-thumb bandwidth $h_1 = 1.05\times\hat\sigma_D\times n_\ell^{-1/5}\times c_1$ with a constant $c_1=1$ and the sample standard deviation of $D$, $\hat\sigma_D$, to estimate the initial variance and bias with $a = \sqrt{0.2}$ and $b = h_1/a$. The resulting undersmoothing bandwidth ranges from 349.17 to 626.91. The estimated trimming probability is capped by $\hat p_\ell \leq 1 - \nu$ with $\nu = 0.01$.

To learn about the effects on selection (or the extensive margin effect), we estimate the average dose-response function of selection $\mathbb{E}[S_d] = s(d)$. The left panel of Figure (ref) presents the estimates of the conditional selection probability given $d$, $\hat s(d)$. More training seems to increase the extensive margin. The minimum selection probability estimate is $\hat\pi_{\text{AT}} = \min_{d\in\mathcal{D}_J} \hat s(d) =\hat s(40) \approx 0.8082$. So we use the least treatment value $40$ as the sufficient treatment value, i.e., if a participant is employed with one-week training, then this participant will remain employed when receiving more training between one week to fifteen months.

figure[figure omitted — 268 chars of source]

In Figure (ref) in Section (ref) in the online appendix, the histograms of $Y$ and $\log(Y)$ in the selected sample $\{Y_i > 0$, or $S_i=1, i=1,...,n\}$ show that weekly earnings has a skewed distribution, and $\log$(weekly earnings) is closer to normal. It is well-known (e.g., ChenRoth) that averages can be heavily influenced by observations in the tail, especially when the outcome has a skewed distribution, as the weekly earnings in the Job Corps data in Figure (ref). So we estimate the ATE in log, i.e., let $\beta_d = \mathbb{E}[\log(Y_d)|\text{AT}]$, a concave transformation of the outcome that is less heavily influenced by outcomes in the tail of the distribution. The right panel of Figure (ref) shows the estimated bounds under Assumption (ref) with $d_{\text{AT}} = 40$ ($M=1$). The thinner lines are the 95% confidence intervals using the sample variance of the influence function. We find the largest ATE $\beta_{2400} - \beta_{40} = \mathbb{E}[\log(Y_{2400}) - \log(Y_{40})|\text{AT}]$ bounded by $[0.054, 0.281]$ with a 90% confidence $[-0.251, 0.543]$. That is, increasing the training hours from one week to fifteen months would increase the weekly earnings by at least 5.4%, but this effect is not significant at 10% level. We also estimate the ATE in level and do not find any effects of switching hours within $[40,2400]$; see Figure (ref) in Section (ref) in the online appendix.

For comparison, the black dash-dotted line is the estimate of $\mathbb{E}[Y|D=d, S=1]$ assuming that there is no effect on selection, i.e., the selected sample over $\mathcal{D}$ contains only the always-takers, $\{S_{d'}=1 : d'\in\mathcal{D}\} = \{S_d =1\}$ for all $d\in\mathcal{D}$. That is, $\beta_d = \mathbb{E}[Y|D=d, S=1]$ is point-identified. Or under the missing at random assumption $\mathbb{E}[Y|D=d, S=1] = \mathbb{E}[Y_d]$. We also plot the estimates of $\mathbb{E}[Y|D=d]$ ignoring selection using all observations including $\{i: S_i=0\}$.

\paragraph{Sufficient set} For sensitivity analysis and illustrating the sufficient set Assumption (ref), Figure (ref) presents the estimated bounds with $M=1,2,3$ given in Theorem (ref). Since $d_{\text{AT}} = \arg\min_{d' \in \mathcal{D}} s(d')$ is estimated as $\hat d_{\text{AT}_J} = 40$, we choose $\mathcal{D}_M$ to contain 40 such that the trimming probability $\pi_L^M/s(d) < 1$. It is natural to include the boundary points $\underline{\mathcal{D}} =40$ and $\overline{\mathcal{D}}=2400$ in $\mathcal{D}_M$. Following the discussion of the non-separable structural selection equation in Assumption 3$^\prime$, for a subject with $d_{\text{AT}}(\eta) = \underline{\mathcal{D}}$, more training hours helps employment. On the other hand, for a subject with $d_{\text{AT}}(\eta) = \underline{\mathcal{D}}$, the largest treatment value hurts selection probability (employment). So for $M=2$, we let $\mathcal{D}_2 = \{40, 2400\}$. Assumption (ref) allows a complier to be employed with $40$ training hours but unemployed with 2400 hours. And if a participant is employed with both the smallest and largest hours ($d=40, 2400$), then this participant must be employed at any hours between one week to fifteen months, as an always-taker. The blue dashed line in Figure (ref) uses the lower bound $\hat\pi_L^2 = \hat s(40) + \hat s(2400) - 1\approx 0.658$.

For $M > 2$, we suggest choosing $\mathcal{D}_M$ as a equally spaced sub-grid over $\mathcal{D}_J$ and including $\hat d_{\text{AT}_J}$, if there is no further information on the structural selection model. So for $M=3$, we let $\mathcal{D}_3 = \{40, 1208, 2400\}$, the orange long-dashed line uses the lower bound $\hat\pi_L^3 = \hat s(40) + \hat s(1208) + \hat s(2400) - 1 \approx 0.5$. As expected, the bounds are less informative when we assume more treatment values, i.e., a larger $M$. So we note that the sufficient set assumption is more of theoretical interest than practical. We focus on sufficient value Assumption (ref) in empirical analysis. Next we incorporate covariates to potentially tighten the bounds and confidence intervals.

figure[figure omitted — 168 chars of source]

Estimation and inference conditional on covariates

We weaken the independent treatment assumption and the sufficient value/set assumption by conditioning on the covariates. So the covariates $X$ can serve two purposes: First, in observational data, it is more plausible and standard to assume conditional independence, also known as unconfoundedness, selection on observables, or ignorability. Second, we allow subjects with different pretreatment covariates to have different sufficient treatment values $d_{\text{AT}x}\in\mathcal{D}$. It might be reasonable that participants with some particular covariates benefit from more training, but more training could hurt the employment of some participants with different covariates. So the bounds may be tightened by incorporating the covariates.

After conditional on the covariates, we follow SGLee to represent the generalized Lee bounds as a moment equation and derive an orthogonal moment for it. The advantage of using the orthogonality moment is that the first-stage estimation has no contribution to the asymptotic variance of the bounds. We utilize cross-fitting to remove overfitting bias without strong entropy conditions, following the recent double debiased machine learning (DML) literature CCDDHNR. Then we show the asymptotic theory that accommodates low-dimensional smooth and high-dimensional sparse designs.

Notice that $\beta_{d{}}$ and its bounds are functions of $d$ over $\mathcal{D}$ and hence are of infinite dimension for a continuous treatment. Such non-regular nonparametric estimands cannot be estimated at a regular root-$n$ rate without further assumptions. The asymptotic theory is more involved with the kernel function and bandwidth $h$, compared with the semiparametric inference in CCDDHNR. CL provide DML inference for $\beta_{d{}}$ that is point-identified when there is no selection bias. In particular, if researchers assume “no effect on selection" by using the selected sample excluding those with zero outcomes, then they estimate $\mathbb{E}[\mathbb{E}[Y|S=1, D=d, X]]$. Or if researchers “ignore selection" by including all observations with zero outcomes, then they estimate $\mathbb{E}[\mathbb{E}[Y|D=d, X]]$. See also Kennedy, SUZ, and CNS21EJ for non-regular estimands and machine learning.

The conditional independence Assumption (ref) means that conditional on observables, the treatment variable is as good as randomly assigned, or conditionally exogenous.

assumption[Conditional independence] $D$ is independent of $\big\{(Y_d, S_d): d\in \mathcal{D}\big\}$ conditional on $X$.

Assumption (ref) relaxes Assumption (ref) to allow subjects with different values of pretreatment covariates $x$ to have different sufficient values $d_{\text{AT}x}\in\mathcal{D}$.

assumption[Conditional sufficiency] For $x \in \mathcal{X}$, there exists $d_{\text{AT}x} \in \mathcal{D}$ such that $\mathbb{P}(S_d \geq S_{d_{\text{AT}x}}, \forall d \in\mathcal{D}|X=x)=1$.

As discussed in Assumption 2$^\prime$, a sufficient condition of Assumption (ref) is to assume a separable structural error $\eta$ in selection $S=\mathbf{1}\{q(D,X) \geq \eta\}$, and there exists $d_{\text{AT}x} = \arg\min_{d\in\mathcal{D}} q(d,x)$.

Let the conditional average dose-response function of always-takers be $\beta_d(x) := \mathbb{E}[Y_d|\{S_{d'}=1: d' \in \mathcal{D}\}, X=x]$, so $\beta_d = \int_{\mathcal{X}} \beta_d(x) f_X(x|S_{d'}= 1: d' \in \mathcal{D}) dx.$ Let the conditional probability of always-takers be $\pi_{\text{AT}}(x) := \mathbb{P}(S_{d'} = 1: d' \in\mathcal{D}|X=x)$. Let the corresponding conditional upper bound given in ((ref)) be $\bar\beta_d(x):= \rho_{dU}(\pi_{\text{AT}}(x),x) := \mathbb{E}[Y|Y \geq Q^d(1-\pi_{\text{AT}}(x)/s(d,x),x), D=d, S=1, X=x]$, where the conditional selection probability $s(d,x) := \mathbb{P}(S=1|D=d, X=x)$ and $Q^d(u,x)$ is the $u$-quantile of $Y|D=d, S=1, X=x$. Define the aggregate upper bound for $\beta_{d}$ as $\bar\beta_{d}:=\int_\mathcal{X}\bar\beta_{d}(x) f_X(x|\text{AT}) dx$.

Lemma (ref) below extends Lemma 1 in SGLee to show that the upper bound $\bar\beta_{d}$ is a ratio of two moments, by replacing the binary treatment indicator $D$ with a kernel function $K_h(D-d)$. The moment function for the lower bound is defined analogously in the proof of Lemma (ref) in the Appendix.

lemma[Moment-based representation] Assuming Assumption (ref), $s(d,x) = \mathbb{E}[S_d|X=x]$. Further assuming Assumption (ref), $\pi_{\text{AT}}(x) = \min_{d\in\mathcal{D}} s(d,x) = s(d_{\text{AT}x},x)$. Assume $\pi_{AT}=\mathbb{E}\big[\pi_{\text{AT}}(X)\big] > 0$ and the generalized propensity score $\mu_d(x) := f_{D|X}(d|x) > 0$ with probability one, for $d\in\mathcal{D}$. Then the sharp upper bound of $\beta_d$ is \begin{align*} \bar\beta_{d} = {\mathbb{E}[\bar\beta_{d}(X) \pi_{AT}(X)]}/{\pi_{AT}} = \lim_{h\rightarrow 0} {\mathbb{E}\big[m_{dU}(W,\xi)}\big]/{\pi_{AT}}, where \\ m_{dU}(W, \xi) := \frac{K_h(D-d)}{\mu_d(X)} S Y\cdot\mathbf{1}\{Y \geq Q^d(1- \pi_{AT}(X)/s(d,X), X)\}, \end{align*} the nuisance parameter $\xi(d, x) = \big\{ s(d, x), \mu_d(x), Q^d(1-\pi_{\text{AT}}(x)/s(d,x), x)\big\}$.
remark[Conditional sufficient set]{\rm We discuss incorporating the covariates under the sufficient set Assumption (ref). We provide a non-sharp bound that can be implemented in practice. The sharp bound is out of the scope of this paper. Under Assumption (ref), Theorem (ref) directly implies the conditional version of the bounds on $\pi_{\text{AT}}(x)$: \begin{align*} \pi^M_{L}(x) := \max\left( \sum_{d\in \mathcal{D}_M} s(d, x) - M + 1, 0\right) &\leq \pi_{AT}(x) = \mathbb{P}(S_{d}=1: d\in\mathcal{D}_M|X=x) \\ &\leq \min_{d\in \mathcal{D}_M} s(d,x) =: \pi^M_{U}(x). \end{align*} Then together with the proof of Lemma (ref), the sharp upper bound $ \mathbb{E}[\rho_{dU}(\pi_{\text{AT}}(X), X) \pi_{\text{AT}}(X)]/\pi_{\text{AT}}$ is not point-identified and is smaller than $\mathbb{E}[\rho_{dU}(\pi_{L}^M(X), X) \pi_{U}^M(X)]/\pi_{L}^M = \lim_{h\rightarrow 0}\mathbb{E}[m_{dU}(W, \xi)\cdot {\pi_{U}^M(X)/ \pi_{L}^M(X)}]/\pi_L^J$ with $\pi_{\text{AT}}(X) = \pi_{L}^M(X)$ in $m_{dU}(W,\xi)$, which likely is a non-sharp bound. Similar, the sharp lower bound $\mathbb{E}[\rho_{dL}(\pi_{\text{AT}}(X), X) \cdot \pi_{\text{AT}}(X)]/\pi_{\text{AT}} \geq \mathbb{E}[\rho_{dL}({\color{black}\pi_{L}^M(X)}, X) \cdot {\color{black}\pi_{L}^M(X)}]/{\color{black}\pi_{U}^M} = \lim_{h\rightarrow 0}\mathbb{E}[m_{dL}(W,\xi)\cdot {{\color{black}\pi_{L}^M(X)}/ {\color{black}\pi_{L}^M(X)}}]/{\color{black}\pi_U^M}$, with $\pi_{\text{AT}}(X) = \pi_{L}^M(X)$ in $m_{dL}(W,\xi)$. }

Estimation and inference

We estimate the bounds over an evenly spaced grid $\mathcal{D}_J := \{d_1,...,d_J\} \subseteq \mathcal{D}$. Let the sets $\mathcal{X}_j := \{x: s(d_j,x)/s(d, x) \leq 1, \forall d \in \mathcal{D}_J\}$ for $j=1,...,J$. So for $x\in\mathcal{X}_j$, $d_{\text{AT}x_J} = d_j = \arg\min_{d\in\mathcal{D}_J} s(d, x)$. We can classify the subjects into $J$ groups $\mathcal{X} = \cup_{j=1,...,J}\mathcal{X}_{j}$. By consistently estimating $s(d,x)$ over the grid $\mathcal{D}_J$, we show that the mis-classification error is asymptotically first-order ignorable.

We allow the subsets $\mathcal{X}_j$ to overlap, i.e., the sufficient treatment value can be not unique, so there could be no treatment effect on selection over some range in $\mathcal{D}$. For example, if there exists $x \in \mathcal{X}_{j} \cap \mathcal{X}_{j+1}$ so that $s(d_j,x)/s(d_{j+1},x)=1$, then the bounds degenerate to points at $d_j$ and $d_{j+1}$, i.e., $\bar\beta_{d_j}(x) = \underline{\beta}_{d_j}(x) = \mathbb{E}[Y|D=d_j, S=1, X=x]$ and $\bar\beta_{d_{j+1}}(x) = \underline{\beta}_{d_{j+1}}(x) = \mathbb{E}[Y|D=d_{j+1}, S=1, X=x]$. We estimate the robust bounds using the trimming probability $\hat p_{d_{j+1}d_j}(x) = \hat s(d_j,x)/\hat s(d_{j+1},x)-\nu$ for some small positive $\nu$, rather than the untrimmed point-estimator, as discussed in Section (ref).

For $X_i \in \mathcal{X}_j$, define the moment function to be $m_{dU}^{j}(W_i,\xi) := m_{dU}(W_i,\xi)$ with $\pi_{\text{AT}}(X_i) = s(d_j, X_i)$. To construct orthogonality as $h\rightarrow 0$, we derive the correction terms for $s(d_j, x)$, $s(d, x)$, $Q^d(u,x)$, $\mu_d(x)$, respectively, collected in

align*[align* omitted — 437 chars of source]

for $p_{dd_j}(X) \leq 1 - \nu < 1$ with $\nu > 0$. For $d=d_j$, let $\nu = 0$, $p_{dd_j}(X) = 1$, and

align[align omitted — 242 chars of source]

Then the orthogonal moment function is defined as $g_{dU}^{j} := m_{dU}^{j} + cor_{dU}^{j}$. Let $g_{dU}(W, \xi) = \sum_{j=1}^J g^j_{dU}(W, \xi)\mathbf{1}\{X \in \mathcal{X}_j\}/\sum_{j=1}^J \mathbf{1}\{X \in \mathcal{X}_j\}$ that allows overlapping $\mathcal{X}_j$s.

These correction terms are derived based on equation (4.7) in SGLee. Specifically let the true nuisance parameter be $\xi_0$ and $\xi_r := \xi_0 + r(\xi - \xi_0)$ for some $\xi$ close to $\xi_0$ and $r \in (0,1)$. Then we verify that the partial derivative of the moment function $g_{dU}^{j}$ with respect to $r$ is zero. When $p_{dd_j}(x) =1$, $\bar\beta_d(x)=\underline{\beta}_d(x)$ is point-identified. So the correction term ((ref)) is only for the nuisance function $\mu_{d}(x)$ as derived in CL.

The estimation procedure follows four steps: \paragraph{Step 1.} ($L$-fold Cross-fitting) As defined in Section (ref), $L$-fold cross-fitting randomly partitions the observation indices into $L$ distinct groups $I_\ell, \ell = 1,...,L$. For $\ell = 1,..., L$, the estimator $\hat \xi_\ell(W)$ for the nuisance function $\xi(W) = \big(s(d,X), Q^d(1- p_{dd_j}(X), X),\mu_d(X), \mathbb{E}[Y|Y \geq Q^d(1- p_{dd_j}(X), X), D=d, S=1, X], d\in\mathcal{D}\big)$ uses observations not in $I_\ell$, satisfying Assumption (ref) below.

\paragraph{Step 2.} (Double robustness) For $i\in I_\ell$ and $X_i \in \hat{\mathcal{X}}_{j\ell} = \{x: \hat s_\ell(d_j,x)/\hat s_\ell(d, x) \leq 1, \forall d \in \mathcal{D}_J\}$, estimate the orthogonal moment function by $g_{dU}(W_i, \hat \xi_\ell) = \sum_{j=1}^J g^j_{dU}(W_i,\hat \xi_\ell) \mathbf{1}\{X_i \in \hat{\mathcal{X}}_{j\ell}\}/\sum_{j=1}^J \mathbf{1}\{X_i \in \hat{\mathcal{X}}_{j\ell}\}$.

\paragraph{Step 3.} The DML estimator in CL for $\pi_{\text{AT}}$: $\hat\pi_{\text{AT}} = n^{-1}\sum_{\ell=1}^L\sum_{i\in I_\ell} \psi(W_i, \hat\xi_\ell)$, where $\psi(W_i, \xi) := K_h(D_i - d_j)(S_i - s(d_j, X_i))/ \mu_{d_j}(X_i) + s(d_j, X_i)$ if $X_i \in \hat{\mathcal{X}}_{j\ell}$.

\paragraph{Step 4.} The DML estimator $\hat{\overline{\beta}}_d = n^{-1}\sum_{\ell = 1}^L \sum_{i \in I_\ell} g_{dU}(W_i, \hat \xi_\ell)/\hat \pi_{\text{AT}}$.

Denote the $L_2$-norm $\big\|\hat \xi- \xi\big\|_2 = \|\Delta\hat\xi(W)\|_2= \left(\int (\hat \xi(w) -\xi(w))^2 f_W(w) dw\right)^{1/2}$.

assumption\begin{enumerate} • (Strict overlap) $s(d, x) \in (c, 1-c)$ for some constant $c \in (0, 1/2)$, for all $(d, x) \in \mathcal{D} \times \mathcal{X}$. $\inf_{d\in\mathcal{D}}{\rm ess}\inf_{x\in\mathcal{X}} \mu_d(x) \geq c$ for some positive constant $c$. • $s(d,x)$, $f_{YSDX}(y, 1, d, x)$ and $\mathbb{E}[Y|D=d, S=1, X=x]$ are two-times differentiable with respect to $d$ with all two derivatives being bounded uniformly over $\mathcal{Y}\times\mathcal{D}\times\mathcal{X}$. The derivative of $Q^d(u, x)$ with respect to $u$ and $var(Y|D=d, S=1, X=x)$ is bounded uniformly over $\mathcal{D}\times\mathcal{X}$. $\mathbb{E}[|Y|^3|S=1, D=d, X=x]$ is continuous in $d$ uniformly over $\mathcal{X}$. • The following terms are $o_\mathbb{P}(1)$ for $d\in\mathcal{D}$: $\|\Delta\hat\mu_d(X)\|_2$, $\sup_{y\in\mathcal{Y}_0}\| \Delta\hat{\mathbb{E}}[Y|Y\geq y, S=1, D=d, X]\|_2$, $\|\Delta\hat s(d, X)\|_2$, $\sup_p\|\Delta\hat Q^d(p,X)\|_2$, where $\mathcal{Y}_0$ be a compact subset of the support of $Y$. • The following terms are $o_\mathbb{P}(1/\sqrt{nh})$ for $d\in\mathcal{D}$: \\ $\sup_{y\in\mathcal{Y}_0}\| \Delta\hat{\mathbb{E}}[Y|Y\geq y, S=1, D=d, X]\|_2 \|\Delta\hat\mu_d(X)\|_2$, \\ $\sup_{p\in(0,1)}\|\Delta\hat Q^d(p,X)\|_2\Big( \sup_{p\in(0,1)}\|\Delta\hat Q^d(p,X)\|_2 + \|\Delta\hat\mu_d(X)\|_2 + \|\Delta\hat s(d, X)\|_2\Big)$,\\ $\|\Delta\hat s(d, X)\|_2 \Big( \|\Delta\hat s(d, X)\|_2 + \sup_{y\in\mathcal{Y}_0}\| \Delta\hat{\mathbb{E}}[Y|Y\geq y, S=1, D=d, X]\|_2 + \|\Delta\hat\mu_d(X)\|_2\Big)$. • $h\rightarrow 0$, $nh\rightarrow \infty$, $\sqrt{nh}h^2 \rightarrow c\in[0,\infty)$. \end{enumerate}

Assumption (ref)(i) requires the generalized propensity score (GPS) $\mu_d(X)$ to be bounded away from zero, which is the standard overlap assumption SGLee. We note that such common support assumption should be made with care in practice and is strong especially with many control variables DAmour. In our empirical application, we find a common support by trimming away observations whose estimated GPSs are smaller than some fixed trimming parameter. Details are in Section (ref) and Section (ref) in the online appendix.

Assumption (ref)(iii)(iv) gives tractable high-level rate conditions on the nuisance function estimators. These are standard conditions similarly assumed in the DML literature CCDDHNR, CL. The rate conditions use a “partial $L_2$" norm in the sense that the regressor $D$ is fixed at $d$ and the expectation is based on the marginal distribution of $X$, e.g., $\|\Delta \hat s(d,X)\|_2 = \big(\int_{\mathcal{X}} (\hat s(d,x) - s(d,x))^2 f_X(x) dx\big)^{1/2}$. See CL for detailed discussion on the low-level conditions of the nuisance function estimators that satisfy such high-level conditions, such as kernel, series, and neural network in low-dimensional settings with fixed dimension of $X$. In high-dimensional settings where the dimension of X grows with the sample size, our inference theory is valid as long as the nuisance function estimators satisfy the high-level rate conditions, e.g., Lasso, as we illustrate in Section (ref).

Similar to Assumption (ref), Assumption (ref) imposes conditions on the grid.

assumptionLet $s^{(m)}(d,x)$ be the $m^{th}$ derivative of $s(d,x)$ w.r.t.\ $d$ for $m \in \{1,2,...\}$. Let $\mathcal{D} = \mathcal{D}_{sx} \cup \mathcal{D}_{cx}$ for any $x \in \mathcal{X}$, where $\mathcal{D}_{cx} := \{d: s^{(m)}(d,x) = 0, \forall m \geq 1\}$ and $\mathcal{D}_{sx} := \{d: s^{(m)}(d, x) \neq 0, \exists m < \infty\}$. If $\mathcal{D}_{sx} \neq \emptyset$, let $\bar M_x = \min\{m: s^{(m)}(d,x) \neq 0, m=1,2,..., \forall d \in \mathcal{D}_{sx}\} < \infty$. If $\mathcal{D}_{sx} = \emptyset$, let $\bar M_x = 0$. Let $\bar M = \max_{x\in\mathcal{X}} \bar M_x$. Let an equally spaced grid $\mathcal{D}_J = \{d_1,..,d_J\} \subseteq \mathcal{D}$ with $J = O({\mathsf{s}_n}^{-1/\bar M})$ and $\mathsf{s}_n := \sup_{d\in\mathcal{D}, x\in\mathcal{X}}|\hat s(d,x) - s(d,x)| = o_\mathbb{P}(1)$.
theoremLet Assumptions (ref), (ref), (ref), (ref), and (ref) hold. Then for $d\in\mathcal{D}$, $\hat{\overline{\beta}}_d - \overline{\beta}_d = n^{-1}\sum_{i=1}^n \phi_{dU}(W_i, \xi) - \overline{\beta}_d + o_\mathbb{P}(1/\sqrt{nh})$, where $\phi_{dU}(W_i, \xi):= g_{dU}(W_i, \xi)/\pi_{\text{AT}} - \psi(W_i, \xi) \overline{\beta}_d/\pi_{\text{AT}}$. And $\sqrt{nh}\big(\hat{\bar\beta}_d - \bar\beta_d - h^2\mathsf{B}_{dU}\big)\stackrel{d}{\rightarrow} \mathcal{N}(0,\mathsf{V}_{dU})$. Similarly for the lower bounds, $\hat{\underline{\beta_d}} - \underline{\beta_d} = n^{-1}\sum_{i=1}^n \phi_{dL}(W_i, \xi) - \underline{\beta_d} + o_\mathbb{P}(1/\sqrt{nh})$, where $\phi_{dL}(W_i, \xi):= g_{dL}(W_i, \xi)/\pi_{\text{AT}} - \psi(W_i, \xi) \underline{\beta_d}/\pi_{\text{AT}}$. And $\sqrt{nh}\big(\hat{\underline{\beta_d}} - \underline{\beta_d} - h^2\mathsf{B}_{dL}\big)\stackrel{d}{\rightarrow} \mathcal{N}(0,\mathsf{V}_{dL})$ , where $\mathsf{B}_{dU}$, $\mathsf{V}_{dU}$, $\mathsf{B}_{dL}$, and $\mathsf{V}_{dL}$ are given explicitly in the proof in the Appendix.

We can estimate the asymptotic variance $\mathsf{V}_{dU}$ by the sample variance of the estimated influence function $\hat{\mathsf{V}}_{dU} := hn^{-1}\sum_{i=1}^n\phi_{dU}(W_i,\hat\xi)^2$. As described in Section (ref), we could estimate the leading bias $\mathsf{B}_{dU}$ by the method in PS96 and CL. Let the notation $\hat{\bar \beta}_{d,b}$ be explicit on the bandwidth $b$ and $\widehat{\mathsf{B}}_{dU}:= \big(\hat{\bar \beta}_{d,b} - \hat{\bar \beta}_{d,ab}\big)\big/\big(b^2(1-a^2)\big)$ with a pre-specified fixed scaling parameter $a\in(0,1)$. Then a data-driven bandwidth $\hat h_{dU}:= (\widehat{V_{dU}}/(4\widehat{\mathsf{B}}_{dU}^2))^{1/5}n^{-1/5}$ consistently estimates the optimal bandwidth that minimizes the AMSE. For the lower bound, the same estimation applies to $\hat{\mathsf{V}}_{dL}$, $\widehat{\mathsf{B}}_{dL}$, and $\hat h_{dL}$. Then we can choose an undersmoothing bandwidth $h$ that is smaller than $\hat h_{dU}$ and $\hat h_{dL}$ to construct the 95% confidence interval $\Big[\hat{\underline{\beta_d}}-1.96 \times \widehat{\sigma_{dL}}/\sqrt{nh}, \hat{\bar{\beta_{d}}} + 1.96 \times \widehat{\sigma_{dU}}/\sqrt{nh} \Big]$ with $\widehat{\sigma_{dL}} := \sqrt{\widehat{V_{dL}}}$ and $\widehat{\sigma_{dU}} := \sqrt{\widehat{V_{dU}}}$. Under additional assumptions, we can straightforwardly show consistency of $\hat{\mathsf{V}}_{dU}$, $\widehat{\mathsf{B}}_{dU}$, and $\hat h_{dU}$ following Theorem 3.2 in CL, so we omit the repetition.

We can bound the ATE of increasing the treatment from $d_1$ to $d_2$, $\underline{\Delta}_{d_1d_2} := \underline{\beta_{d_2}} - \overline{\beta}_{d_1} \leq \beta_{d_2} - \beta_{d_1} \leq \overline{\beta}_{d_2} - \underline{\beta_{d_1}} =: \bar\Delta_{d_1d_2}$. Denote the bandwidth in $\hat{\underline{\beta}}_{d}$ and $\hat{\bar{\beta}}_{d}$ be $h_{dL}$ and $h_{dU}$, respectively.

corollary[ATE] Let the conditions in Theorem (ref) hold. Let $\mathsf{V}_{Un} := \mathbb{E}[(\phi_{d_2U} - \phi_{d_1L})^2]$ and $\mathsf{V}_{Ln}:= \mathbb{E}[(\phi_{d_2L} - \phi_{d_1U})^2]$. Then $\sqrt{n}{\mathsf{V}_{Un}}^{-1/2}\big(\hat{\bar\Delta}_{d_1d_2} - \bar\Delta_{d_1d_2} - (h_{d_2U}^2\mathsf{B}_{d_2U} -h_{d_1L}^2\mathsf{B}_{d_1L})\big) \stackrel{d}{\longrightarrow} \mathcal{N}(0, 1)$ and $\sqrt{n}\mathsf{V}_{Ln}^{-1/2}\big(\hat{\underline{\Delta}}_{d_1d_2} - \underline{\Delta}_{d_1d_2}- (h_{d_2L}^2\mathsf{B}_{d_2L} -h_{d_1U}^2\mathsf{B}_{d_1U})\big)\stackrel{d}{\longrightarrow} \mathcal{N}(0, 1)$.

The variance can be estimated by $\hat{\mathsf{V}}_{Un} = n^{-1}\sum_{i=1}^n(\phi_{d_2U}(W_i,\hat\xi) - \phi_{d_1L}(W_i,\hat\xi))^2$ and $\hat{\mathsf{V}}_{{Ln}} = n^{-1}\sum_{i=1}^n(\phi_{d_2L}(W_i,\hat\xi) - \phi_{d_1U}(W_i,\hat\xi))^2$. The 95% confidence interval $\big[\hat{\underline{\Delta}}_{d_1d_2} - 1.96\times \sqrt{\hat{\mathsf{V}}_{Ln}}/\sqrt{n}, \hat{\bar\Delta}_{d_1d_2} + 1.96\times \sqrt{\hat{\mathsf{V}}_{Un}}/\sqrt{n} \big]$.

First-step nuisance parameter estimation

We illustrate our estimation procedure by applying Lasso methods in Step 1 to estimate the nuisance functions, when $X$ is potentially high-dimensional. We provide sufficient conditions to verify the high-level Assumption (ref). We follow SUZ (SUZ, hereafter) to approximate the outcome, selection, and treatment models by varying coefficient linear regressions and logistic regressions. Particularly, the penalized kernel-smoothing least square and maximum likelihood estimations select covariates for each value of the continuous treatment. The approximation errors satisfy Assumption (ref) below that imposes sparsity structures so that the number of effective covariates that can affect them is small. See Farrell15, CNS21ADML, for example, for in-depth discussions on the specification of high-dimensional sparse models.

To estimate the $Q^d(p, x)$, we first estimate the conditional CDF $F_{Y|SDX}(y|1,d,x)$ and then compute the generalized inverse function. To estimate the conditional density $\mu_d(x) = f_{D|X}(d|x)$, we first estimate the conditional CDF $F_{D|X}$ and then take the numerical derivative. We estimate the $s(d,x)$, $F_{D|X}$, and $F_{Y|SDX}$ by the logistic distributional Lasso regression in BCFH17 and SUZ. We modify the penalized local least squares estimators and use the conditional density estimator in SUZ. For completeness, we present the estimators and asymptotic theory in SUZ in the Appendix and refer readers to SUZ for details.

Let $b(X)$ be a $p\times 1$ vector of basis functions and $\Lambda$ be the logistic CDF. Define the approximation errors $r_{d}(x;F_D) = F_{D|X}(d|x) - \Lambda(b(x)'\beta_{d})$, $r_{dy}(x;F_Y) = F_{Y|SDX}(y|1,d,x) - \Lambda(b(x)'\alpha_{dy})$, $r_d(x;s) = s(d,x) - \Lambda(b(x)'\theta_d)$, and $r_d(x;\rho) = \rho_{dU}(\pi, x)- b(x)'\gamma_d$.

Denote the usual $1$-norm $\|\alpha\|_1 = \sum_{j=1}^p |\alpha_j|$ for a vector $\alpha= (\alpha_1,..., \alpha_p)$. Denote $\|W\|_{\mathbb{P},\infty} = \sup_{w\in\mathcal{W}}|w|$ and $\|W\|_{\mathbb{P}_{N_\ell}, 2} = (N_\ell^{-1}\sum_{i\notin I_\ell}W_i^2)^{1/2}$ for a generic random variable $W$ with support $\mathcal{W}$. Assumption (ref) collects the conditions in Theorems 3.1 and 3.2 in SUZ.

assumption[Lasso] Let $\mathcal{D}_0$ be a compact subset of the support of $D$, $\mathcal{Y}_0$ be a compact subset of the support of $Y$, and $\mathcal{X}$ be the support of $X$. \begin{enumerate} • (a) $\|\max_{j \leq p} |b_j(X)|\|_{{\mathbb{P}}, \infty} \leq \zeta_n$ and $\underline{C} \leq {E}\left[b_j(X)^2\right] \leq 1/\underline{C}$, for some positive constant $\underline{C}$, $j=1,...,p$. (b) $\sup_{d \in \mathcal{D}_0, y\in \mathcal{Y}_0} \max(\|\beta_d\|_0, \|\alpha_{dy}\|_0, \|\theta_d\|_0, \|\gamma_d\|_0) \leq \mathfrak{s}$ for some $\mathfrak{s}$ which possibly depends on $n$, where $\|\theta\|_0$ denotes the number of nonzero coordinates of $\theta$. (c) $\sup_{d\in\mathcal{D}_0} \| r_d(X;F_D)\|_{{\mathbb{P}}_n, 2} =O_\mathbb{P}((\mathfrak{s}\log(p \vee n)/n)^{1/2})$ and $\sup_{d\in \mathcal{D}_0, y \in \mathcal{Y}_0}\Big( \big\| r_{dy}(X;F_Y)k((D-d)/h_1)^{1/2} \big\|_{{\mathbb{P}_n},2} + \big\| r_d(X;s)k((D-d)/h_1)^{1/2} \big\|_{{\mathbb{P}_n},2} + \big\| r_d(X;\rho) k((D-d)/h_1)^{1/2} \big\|_{{\mathbb{P}_n},2}\Big)= O_\mathbb{P}\big(\big(\mathfrak{s}\log(p\vee n)/n\big)^{1/2}\big)$. (d) $\sup_{d\in\mathcal{D}_0} \| r_d(X;F_D)\|_{{\mathbb{P}}, \infty} = O((\mathfrak{s}^2\zeta_n^2\log(p \vee n)/n)^{1/2})$ and $\sup_{d\in \mathcal{D}_0, y \in \mathcal{Y}_0}\big( \big\| r_{dy}(X;F_Y) \big\|_{{\mathbb{P}},\infty} + $\\ $\big\| r_d(X;s) \big\|_{{\mathbb{P}},\infty} + \big\| r_d(X;\rho) \big\|_{{\mathbb{P}},\infty}\big)= O\left(\left(\mathfrak{s}^2\zeta_n^2\log(p\vee n)/(nh_1)\right)^{1/2}\right)$. (e) $f_{D|X}(d,x)$ is second-order differentiable w.r.t.\ $d$ with bounded derivatives uniformly over $(d,x)\in\mathcal{D}_0\times\mathcal{X}$. (f) $\zeta_n^2\mathfrak{s}^2\iota_n^2\log(p\vee n)/(nh_1) \rightarrow 0$, $nh_1^5/(\log(p\vee n))\rightarrow 0$. • Uniformly over $(d,x,y)\in\mathcal{D}_0\times\mathcal{X}\times\mathcal{Y}_0$, (a) there exists some positive constant $\underline{C} < 1$ such that $\underline{C} \leq f_{D|X}(d|x) \leq 1/\underline{C}$, $\underline{C} \leq F_{Y|SDX}(y|1,d,x) \leq 1-\underline{C}$, and $\underline{C}\leq s(d,x) \leq 1-\underline{C}$; (b) $s(d,x)$ and $F_{Y|SDX}(y|1,d,x)$ are three times differentiable w.r.t.\ $d$ with all three derivatives being bounded. • There exists a sequence $\iota_n\rightarrow \infty$ such that $0 < \kappa' \leq \inf_{\delta\neq 0, \|\delta\|_0 \leq \mathfrak{s}\iota_n} \|b(X)'\delta\|_{{\mathbb{P}}_n,2}\big/\|\delta\|_2 \leq \sup_{\delta\neq 0, \|\delta\|_0 \leq \mathfrak{s}\iota_n} \|b(X)'\delta\|_{{\mathbb{P}}_n,2}\big/\|\delta\|_2 \leq \kappa'' < \infty$ $w.p.a.1$. \end{enumerate}

Let Assumption (ref) hold. Then Theorems 3.1 and 3.2 in SUZ imply $\sup_{(d, x)\in\mathcal{D}_0\times\mathcal{X}}|\hat s_\ell(d,x) - s(d,x) |= O_\mathbb{P}(A_n)$, $\sup_{(d, x)\in\mathcal{D}_0\times\mathcal{X}, u\in (0,1)}|\hat Q^d_\ell(u,x) - Q^d(u,x) |= O_\mathbb{P}(A_n)$, $\sup_{(d, x,y)\in\mathcal{D}_0\times\mathcal{X}\times\mathcal{Y}_0}|\Delta \hat\mathbb{E}[Y|Y\geq y, S=1, D=d, X=x]|O_\mathbb{P}(A_n)$, where $A_n = \iota_n(\log(p\vee n) \mathfrak{s}^2\zeta_n^2/(nh_1))^{1/2}$. And $\sup_{(d, x)\in\mathcal{D}_0\times\mathcal{X}}|\hat \mu_{d\ell}(x) - \mu_d(x)| = O_\mathbb{P}(R_n)$, where $R_n = h_1^{-1}(\log(p\vee n) \mathfrak{s}^2\zeta_n^2/n)^{1/2}$. Then we can obtain the $L_2$ rates $\|\hat\xi_\ell - \xi\|_2$ to verify Assumption (ref). In particular a sufficient condition of Assumption (ref)(iii) is $A_n \rightarrow 0$ and $R_n \rightarrow 0$. And a sufficient condition of Assumption (ref)(iv) is $\sqrt{nh}A_n R_n \rightarrow 0$ and $\sqrt{nh}A_n^2 \rightarrow 0$.

Empirical illustration with covariates

We illustrate our DML estimator by evaluating the Job Corps and CCC program. First in Step 0, we prepare a sub-sample that satisfies Assumption (ref)(i) for estimating the bounds. We use the full sample to estimate the GPS by $\tilde\mu_d(X_i)$ and the selection probability by $\tilde s(d, X_i)$. We obtain a sub-sample where $\tilde\mu_d(X_i) \geq trim_{GPS}$ and $\tilde s(d, X_i) \geq 5\%$ for {\it all} $i$ in this sub-sample and for {\it all} $d \in \mathcal{D}_J$, where the trimming parameter $trim_{GPS}$ is based on Imbens04REStat and detailed in Section (ref) in the online appendix. Then we use this sub-sample to estimate the bounds following the procedure described in Section (ref). For JC, we remove 11 observations whose $\tilde s(d, X_i) < 5\%$ for some $d \in\mathcal{D}_J$, and no observation is removed by $trim_{GPS}$. For CCC, we remove 94 observation whose $\tilde \mu_d(X_i) < trim_{GPS}$, and all observations have $\tilde s(d, X_i) \geq 5\%$. So we only trim a small number of observations relative to the sample size.

In Step 1, we use the Lasso estimation given in Section (ref). We present the results with a linear basis function $b(X)$ for the regularized varying coefficient linear regression, e.g., $\rho_{dU}(\pi,X) = X'\gamma_d$. Results with a quadratic basis function for specification robustness and implementation details are in Section (ref) in the online appendix. Same as the setting in Section (ref), we use the Epanechnikov kernel with a undersmoothing bandwidth and ten-fold cross-fitting. Let $h_1 = 1.05\times\hat\sigma_D\times n_\ell^{-1/5}\times c_1$, where we use a range of constants $c_1$ for robustness check.

We allow subjects to have different sufficient treatment values depending on $X$, i.e., the minimizer of the conditional selection probability given $X$ and $D=d$. So we might tighten the bounds and the confidence intervals, or capture heterogenous causal effects that are not revealed without the covariates. Figure (ref) reports histograms of the sufficient treatment values, $\{\hat d_{\text{AT}_xi}, i=1,...n\}$. For Job Corps in the left panel, most sufficient treatment values are 40 or 2400, but we also see values over $\mathcal{D}$. In particular, about 20% of the participants have $d_{\text{AT}x} = 40$, i.e., 40 training hours give them the lowest likelihood of employment. For about 9% of the participants who have $d_{\text{AT}x} = 2400$, any hours smaller than 2400 would help employment. For the CCC data in the right panel, about 59.2% of the participants have $d_{\text{AT}x} < 0.5$ years and about 26.19% of the participants have $d_{\text{AT}x} > 1.2$ years.

figure[figure omitted — 230 chars of source]

Job Corps

We follow the literature to assume the conditional independence Assumption (ref), which is indirectly assessed in FFGN12ReStat. It means that receiving different levels of the treatment is random, conditional on a rich set of observed covariates measured at the baseline survey. We may further use a control function with an instrumental variable to address the concern that the conditional independence assumption might not hold IN09ETA, Lee15.

The top left panel of Figure (ref) shows the estimated selection probability $\hat\mathbb{E}[S_d]$ with covariates and without covariates.\footnote{We estimate the selection probability $\mathbb{E}[S_d]$ with covariates $X$ by the DML estimator in CL, $\hat \mathbb{E}[S_d] = n^{-1}\sum_{\ell = 1}^L\sum_{i\in I_\ell} K_h(D_i-d)(S_i - \hat s_\ell(d, X_i))/\hat\mu_{d\ell}(X_i) + \hat s_\ell(d, X_i)$.} The bottom left panel of Figure (ref) presents the estimated bounds $[\hat{\underline{\beta_d}}, \hat{\bar\beta}_d]$ and the 95% confidence intervals. We find a positive intensive margin effect: increasing the training from 1.5 week to 9 months increases log weekly earnings by at least 0.224, at 5% significance level. This is from the largest lower-bound estimate for the ATE of switching hours over $\mathcal{D}$. That is, we find the largest ATE $\beta_{1446.465} - \beta_{63.838}$ bounded by $[0.224, 0.718]$ with the 95% confidence interval $[0.127, 0.865]$.\footnote{$c_1 = 1$. For robustness check, we find positive ATE at $5\%$ significance level over $c_1\in\{0.75, 1, 1.25, 1.5, 1.75\}$ and using a linear or quadratic basis function $b(X)$ in Section (ref) in the online appendix. The undersmoothing bandwidths range from 159.008 to 259.490.} We also not that we do not find significant positive effects on level of weakly earnings.

Consistent with prior empirical research on the Job Corps, e.g., FFGN12ReStat and HHLL have found inverted-$U$ shaped average dose-response functions on labor outcomes. The concave shape is reasonable to consider the optimal treatment intensity in other settings. Our sufficient treatment values assumption does not assume any shape restrictions and hence includes a concave or monotone selection response.

For comparison, the top right panel presents the bounds without $X$, $[\widehat{\rho_{dL}}, \widehat{\rho_{dU}}]$ (red solid line) given in Figure (ref). We also compute two point-estimators of the average dose-response function “ignoring selection" and “no effect on selection" by the DML estimator in CL incorporating $X$. The “ignoring selection" estimates use the full sample including those with zero outcomes, while the “no effect on selection" estimates use the selected sample with positive outcomes. Our bounds are around the “no effect on selection" estimates, as expected.

To demonstrate the usefulness of the DML method, the bottom right panel presents two alternative estimators from Lemma (ref): the regression-type estimator of ${\mathbb{E}[\bar\beta_{d}(X) \pi_{\text{AT}}(X)]}/{\mathbb{E}\big[\pi_{\text{AT}}(X)\big]}$ (Bounds $\rho(\cdot)$) and the inverse-probability-weighting estimator of $\lim_{h\rightarrow 0} {\mathbb{E}\big[m_{dU}(W,\xi)}\big]/{\pi_{AT}}$ (Bounds $m(\cdot)$). These two estimators are not doubly robust, so we see that they are both biased downward, compared with our bounds (Bounds $g(\cdot)$) using a doubly robust moment function.

figure[figure omitted — 388 chars of source]

Civilian Conservation Corps (CCC)

We study the effect of service duration on the age at death, using the data for panel A of Table III in Aizer. They include the 17,639 men (75% of the original sample size 23,722) who have information on death age and who dies after age 45. Aizer investigate the extent of sample selection and the effects of missing data on their estimates, and conclude modest bias from non-random attrition. They estimate an accelerated failure time model with added controls for the characteristics of the enrollees and the camps to show that the estimates remain stable. They find one more year of training increases the death age by one year.

From the histograms of the death age and $\log$(death age) in Figure (ref) in Section (ref) in the online appendix, we see that the distribution of $\log$(death age) is more skewed. So we use the level of {\it death age} as our outcome variable $Y$ and estimate the bounds for $\beta_d = \mathbb{E}[death\ age|\text{AT}]$.

We consider one hundred equally spaced grid points $\mathcal{D}_{100} \subset \mathcal{D} = [0.238, 1.5]$, which are from the 15% to 85% quantiles of duration in years. Figure (ref) reports the estimates without using covariates. The largest ATE is when increasing duration from 0.238 to 1.488 years with the bounds $[0.747, 1.765]$ and 90% confidence interval $[-1.031, 3.476]$.\footnote{ The undersmoothing bandwidths range from 0.1909 to 0.4409. $\nu=0.01$ and $c_1=1$. For robustness check, the lower bound estimates of the largest ATE with $c_1 \in \{0.75, 1,1.25,1.5, 1.75, 2\}$ are positive but insignificant at 10% level. } The proportion of always-takers, or the minimum selection probability, is $\hat\pi_{\text{AT}} = \hat s(0.2382) = 0.8155$. Our sufficient treatment value Assumption (ref) means that if the participants' death ages are observed when they received 0.2382 years of service, then they would remained observed for any duration over these one hundred grid points.

Next we include covariates as in column (3) Add Indiv Controls in Panel A in Table III in Aizer. Figure (ref) reports the estimates with covariates. The largest ATE is when increasing duration from 0.238 to 1.156 years with the bounds $[1.169, 1.755]$ and 95% confidence interval $[0.366, 2.795]$. That is, increasing duration from 0.238 to 1.156 years would at least increase the average death age by 1.169 years and the effect is significant at 5% level, which is consistent with the findings in Aizer.\footnote{ $\nu=0.01$ and $c_1=1.5$. For robustness check, the lower bounds with $c_1 \in \{1.75, 2\}$ and linear/quadratic basis functions range from $0.288$ to $0.54$ but are not significant at 10% level. The undersmoothing bandwidths range from 0.162 to 0.313. } However, the effects are not significant for other choices of $c_1$. By the histogram in Figure (ref) and summary statistics in Table (ref) in the online appendix, the distribution of duration might have mass points at 0.5, 1, 1.5, 2 years. So the caveat of implementing our method on this dataset is that the smoothness conditions could be violated. This may explain why the estimates in Figure (ref) are not smooth. Nevertheless, we are able to implement our method at a coarser grid with 100 grid points and assume the distributions are smooth locally at these grid points. We also consider thirty equally spaced grid points $\mathcal{D}_{30} \subset \mathcal{D} = [0.238, 1.5]$ in Figure (ref) and obtain similar results as $\mathcal{D}_{100}$.\footnote{The largest ATE is when increasing the duration from 0.238 to 1.109 years with the bounds $[1.255, 1.822]$ and 95% confidence interval $[0.43, 2.918]$. $c_1=1.5$. The undersmoothing bandwidths range from 0.158 to 0.294. }

figure[figure omitted — 267 chars of source]
figure[figure omitted — 271 chars of source]
figure[figure omitted — 277 chars of source]

Conclusion

We study causal effects of a treatment/policy variable that could be either continuous, multivalued discrete or binary, in a sample selection model where the treatment affects the outcome and also researchers' ability to observe the outcome. To account for the non-random selection into samples, we provide sharp bounds for the mean potential outcome of always-takers whose outcomes are observed regardless of their treatment value, generalizing LeeBound's bound. We propose a novel sufficient treatment value (set) assumption to (partially) identify the share of always-takers in the observed selected $d$-receipts for each treatment value $d$.

By incorporating pretreatment covariates $X$, we allow for unconfoundedness and allow subjects with different values of $X$ to have different sufficient treatment values, which might tighten the bounds, increase precision (tighten the confidence intervals), and reveal heterogeneity, as we illustrate in our empirical analysis of the Job Corps and CCC programs. The inference procedure is robust to extensive margin, allowing for any unknown functional form of the selection probability with respect to treatment.

There are many potential applications of our bounds for continuous/multivalued treatments. For example, ACL study the multivalued treatments of the Workforce Investment Act program on participants’ earnings. swedish use Swedish lotteries data to study the effect of wealth on labor supply.