EconBase
← Back to paper

Debiased Machine Learning of Aggregated Intersection Bounds and Other Causal Parameters

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

75,324 characters · 14 sections · 87 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Debiased Machine Learning of Aggregated Intersection Bounds and Other Causal Parameters

\linespread{1.5}

abstractThis paper proposes a novel framework of aggregated intersection of regression functions, where the target parameter is obtained by averaging the minimum (or maximum) of a collection of regression functions over the covariate space. Examples of such quantities include the lower and upper bounds on distributional effects (Fr\'echet-Hoeffding, Makarov) as well as the optimal welfare in statistical treatment choice problem QianMurphy. The proposed estimator -- the envelope score estimator -- is shown to have an oracle property, where the oracle knows the identity of the minimizer for each covariate value. I apply this result to the bounds in Roy model and Horowitz-Manski-Lee bounds with discrete outcome. The proposed approach performs well empirically on the data from Oregon Health Insurance Experiment finkelstein.

Keywords: optimal welfare, cross-fitting, double/debiased machine learning, margin assumption, uniformity, Roy model, selection problem, partial identification

Introduction

Economists are often interested in bounds on parameters when parameters themselves are not point-identified Manski89,Manski90,Manski. Examples include quantiles of heterogeneous treatment effects and other distributional measures beyond the mean FanPark. Baseline or pre-treatment covariates often contain valuable information that can tighten these bounds ManskiPepper. However, in practice, sharp bounds are rarely utilized because their estimators usually have non-standard distributions driven by noisy first-stage estimators of unknown conditional distributions. These challenges are not unique to partial identification and also arise in related areas, such as statistical treatment choice LuedtkeLaan, KitagawaTetenov, AtheyWager2, MbakopTabord.

This paper develops estimation and inference methods for the quantities taking the form

align[align omitted — 105 chars of source]

where $X$ is a covariate vector, $\mathcal{T}$ is a finite index set, and $x \mapsto \nu_0(\cdot) = (\nu_{j0}(\cdot))_{j=1}^d$ is a $d$-dimensional nuisance parameter whose elements $\nu_{j0}(x)$ are functions of covariates, such as conditional expectations. As the simplest case, one can think of the optimal welfare which appears in e.g., LuedtkeLaan or sharp bound on distributional effects FanPark. The paper's contribution is to deliver a debiased inference on $\psi_0$ that is first-order insensitive to the misclassification mistake in the identity of the binding constraint. In particular, its distribution is the same as if the true value of the minimizer (ref) were known. Additionally, the paper establishes the validity of a weighted bootstrap method, which holds the estimated minimizer fixed while bootstrapping the second-stage statistic, providing a valid distributional approximation.

The paper illustrates the usefulness of the proposed approach by considering two applications in applied microeconomics. In particular, we discuss in detail a sharp version of Roy model bounds as studied in MourifieHenry as well as Horowitz-Manski-Lee bounds with discrete-valued outcomes a version of which have been also studied in concurrent, independent work of kroft2024leeboundsmultilayeredsample. Revisiting Oregon Health Insurance Experiment finkelstein, we find our methodology useful in determining the direction of treatment effect in the presence of non-response bias as well as tightening the bounds, echoing earlier work in SemSupp2.

The rest of the paper is organized as follows. Section (ref) gives a literature review. Section (ref) introduces the framework and provides two stylized examples. Section (ref) offers an informal preview of the results. Section (ref) presents the formal asymptotic theory and discusses low-level condition for the margin assumption in the context of a single-index model with continuous covariates. Section (ref) applies the proposed theory to sharp bounds in the Roy model and Horowitz-Manski-Lee bounds in selection problems. Section (ref) provides numerical evidence for the methods developed in the article. All proofs are in Appendix (ref).

Literature Review

This paper is related to two lines of research: partial identification and statistical treatment choice.

\paragraph{Bounds, Convex Optimization, and Directionally Differentiable Functionals.} Set identification is a vast area of research, encompassing a wide variety of approaches: linear and quadratic programming, random set theory, support function, and moment inequalities Manski90, ManskiPepper, Manski:2002, HaileTamer, CHT, BM, Molinari2008, CilibertoTamer, LeeBound, Stoye, AndrewsShiECMA, BMM2, CCMS, CherNeweySantos, Gafarov, kallus2020localized, li2022discordant, henry2023role, acerenza2023marginal, ban2021nonparametric, bartalotti2021identifying, JLS, fava2024, see e.g. Molinari:2018 or MolinariHandbook for a review. In the context of distributional effects Makarov, Manski,HSC, FanPark,FanPark2, Tetenov, FanZhu, FirpoRidder, the first discussion of estimation can be traced to FanPark, where, on p.945 they sketch a plug-in estimation approach without statistical guarantees. Targeting the envelope function $\inf_{t \in T} s(t,x)$, the work by CLR proposes a plug-in approach based on the least squares series estimators, where large sample inference is based on the strong approximation of a sequence of series or kernel-based empirical processes. Switching the focus from the envelope function to its best linear predictor, CCMS proposes a root-$N$ consistent and uniformly asymptotically Gaussian estimator of the target parameter, relying on the first-stage series estimators. Finally, recent work by LeeSungwon focuses on bounds on conditional distributions of treatment effects. That is, most inference work focuses on the envelope function, rather than its mean value, which makes the lack of differentiability of $x \mapsto \min (x,0)$ at the kink point $x=0$ a common concern (e.g., FangSantos). Finally, the paper contributes to a growing literature on machine learning for bounds and partially identified models kallus2019assessing, jeong2020robust, SemJoE and sensitivity analysis DornGuo, DornGuoKallus, Bonvini_2021, Bonvini_2022, see e.g., Kennedy_review for the review.

\paragraph{Statistical treatment choice. } Statistical treatment choice is a vast area of research, focusing on two distinct questions: learning the best policy QianMurphy, KitagawaTetenov, MbakopTabord, AtheyWager2, Sun for a given criterion function and inference on the optimal value of the criterion itself LuedtkeLaan in an unconstrained policy class. In the first stream, recent work focused on various robustness aspects of existing criteria functions ishihara2021evidence, adjaho2023externally or targeting welfare criteria that are partially identified Stoye, Pu_2021, Cui_2021, kitagawa2023treatment, yata2023optimal, dadamo2022orthogonal, christensen2023optimal, benmichael2023policy, olea2023decision, cui2024policy. For example, such criteria may arise in asymmetric loss functions babii2021binary, partial welfare ordering han2021optimal,firpo2023loss or distributional welfare based on quantile treatment effects cui2024policy.

Setup

This paper studies the aggregated intersection of regression functions

align[align omitted — 108 chars of source]

where $X$ is a covariate vector, $\mathcal{T}$ is a finite index set, and $x \mapsto \nu_0(\cdot) = (\nu_{j0}(\cdot))_{j=1}^d$ is a $d$-dimensional nuisance parameter whose elements $\nu_{j0}(x)$ are functions of covariates, such as conditional expectations. For each element $t$ of the set $\mathcal{T}$, $\phi(t, \cdot): \mathrm{R}^d \rightarrow \mathrm{R}$ is a known scalar function of the vector $\nu_0$, which could represent a projection onto the Euclidean axis. This function can be expressed as a conditional expectation

align[align omitted — 102 chars of source]

where $W$ is the data vector and $\rho(W, t, \xi_0)$ is an observed random variable that depends on a nuisance parameter $\xi_0$.

Examples of $\psi_0$ include Fréchet-Hoeffding bounds, Makarov bounds on distributional effects, and sharp versions of \citep*{BalkePearl1994, BalkePearl1997} bounds. In statistical treatment choice, examples include optimal welfare in an unconstrained policy class Manski2004,QianMurphy, LuedtkeLaan,KitagawaTetenov. As a stylized example, this section revisits optimal welfare in statistical treatment choice and Makarov bounds on distributional effects FanPark.

example[Optimal Welfare] Let $D$ be a discrete-valued treatment variable taking values in a finite set $\mathcal{D}$. Let $Y(d)$ be a potential outcome, and let $Y = \sum_{d \in \mathcal{D}} Y(d) 1\{D = d\}$ be the observed outcome. The data vector is $W = (D, X, Y)$. Let $m(d, x) = \mathbb{E} \big[ Y \mid D = d, X = x \big]$ be the conditional expectation function. Under the unconfoundedness assumption \begin{align} (Y(d))_{d \in \mathcal{D}} \perp\!\!\!\perp D \mid X, \end{align} the conditional means of potential outcomes are identified as $$ \mathbb{E} \big[ Y(d) \mid X = x \big] = m(d, x). $$ The negative attained welfare is $$ \psi_0 = -\mathbb{E} \big[ \max_{d \in \mathcal{D}} m(d, X) \big] = \mathbb{E} \big[ \min_{d \in \mathcal{D}} -m(d, X) \big], $$ which is a special case of (ref) with $\mathcal{T} = \mathcal{D}$ and \begin{align*} \nu_0(x) &= (m(d, x))_{d \in \mathcal{D}}, \qquad \phi(d, v) = -v_d, \quad d \in \mathcal{D}. \end{align*} The "unbiased" signal for $-m(d, x)$ is the Robins orthogonal score \begin{equation} \rho(W, d, \xi_0) = -\frac{1\{D = d\}}{\mu_{d0}(X)} \big( Y - m(d, X) \big) - m(d, X), \end{equation} where the propensity score is \begin{align*} \mu_{d0}(x) = \Pr(D = d \mid X = x) \end{align*} and the nuisance parameter is \begin{align*} \xi_0(x) = \big(\nu_0(x), (\mu_{d0}(x))_{d \in \mathcal{D}}\big). \end{align*} When the treatment is binary, that is, $\mathcal{D} = \{0, 1\}$, the parameter $\psi_0$ reduces to $$ \psi_0 = -\mathbb{E} \big[ \max(m(1, X), m(0, X)) \big] $$ and $\mu_{10}(X) + \mu_{00}(X) = 1 \text{ a.s. }$
example[Makarov Bounds on Distributional Effects] Consider the setup of Example (ref) with $\mathcal{D} = \{0, 1\}$. Let $F_1(\cdot \mid x)$ and $F_0(\cdot \mid x)$ be the conditional Cumulative Distribution Functions (CDFs) of the potential outcomes $Y(1)$ and $Y(0)$, respectively, identified under (ref). Let $F_{Y(1) - Y(0)}(d)$ be the CDF of the treatment effect $Y(1) - Y(0)$, and let $F_{Y(1) - Y(0)}(d \mid x)$ be the conditional CDF. As shown in Makarov,FanPark,FirpoRidder, the sharp bounds on the CDF of the treatment effect $F_{Y(1) - Y(0)}(d)$ are \begin{align} \pi_L(d) &:= \mathbb{E} \big[\sup_{t \in \mathrm{R}} \max(F_1(t \mid X) - F_0(t - d \mid X)_{-}, 0)\big] \\ \pi_U(d) &:= \mathbb{E} \big[\inf_{t \in \mathrm{R}} \min(F_1(t \mid X) - F_0(t - d \mid X)_{-}, 0) + 1\big] \end{align} where $F(\cdot)$ and $F(\cdot)_{-}$ is the right-hand limit (i.e., regular CDF) and the left-hand limit of the CDF, respectively. If the distributions $Y \mid D = 1, X = x$ and $Y \mid D = 0, X = x$ have finite support, their respective CDFs are step functions with finitely many jumps whose locations on $x$-axis are denoted by $\mathcal{T}_1$ and $\mathcal{T}_0$, respectively. Consider the share of subjects negatively affected by the treatment. It is upper and lower bounded as $$ \pi_L(0) \leq \Pr (Y(1) - Y(0) \leq 0) \leq \pi_U(0). $$ The upper bound $\pi_U(0)$ is a special case of (ref) with $\mathcal{T} = \{0\} \cup \mathcal{T}_1 \cup \mathcal{T}_0$, $\nu^1_0(x) = (F_1(t \mid X = x))_{t \in \mathcal{T}}$ and $\nu^0_0(x) = ({F_0(t \mid X = x)}_{-})_{t \in \mathcal{T}}$ and $$ \phi(t, v^1, v^0) = v^1_{t} - v^0_{t}, \quad t \in \{0\} \cup \mathcal{T}_1 \cup \mathcal{T}_0. $$ The “unbiased” signal is the Robins-type orthogonal score \begin{align} \rho(W, t, \xi_0) &= \frac{D 1\{Y \leq t\}}{\mu_{10}(X)} \big(1\{Y \leq t\} - F_1(t \mid X)\big) + F_1(t \mid X) \\ &\quad - \frac{(1 - D) 1\{Y < t\}}{\mu_{00}(X)} \big(1\{Y < t\} - F_0(t \mid X)_{-}\big) - F_0(t \mid X)_{-}, \quad t \in \mathcal{T}_1 \cup \mathcal{T}_0 \end{align} and $\rho(W, 0, \xi_0) = 0 \text{ a.s.}$. The propensity score is as in Example (ref), and the nuisance parameter is $$ \xi_0(x) = (\nu^1_0(x),\nu^0_0(x), \mu_{10}(x), \mu_{00}(x)) $$ and $\mu_{10}(X) + \mu_{00}(X) = 1 \text{ a.s. }$.

Example (ref) describes Makarov bounds on the treatment effect CDF, previously studied in FanPark and FirpoRidder, among others. A special case of this example with binary outcomes was studied in kallus2022whats, who proposed debiased inference for $\pi_L(0)$ and $\pi_U(0)$. Other interesting examples of this parameter are described in Section (ref).

Overview of Estimation and Inference

In this section, I introduce the estimator of the parameter of interest and describe two inferential approaches. Let me briefly review the notation. Recall that $W$ is a data vector, $\mathcal{T}$ is an index set, and each function $\phi(t, \nu_0(x))$ is a conditional expectation function of an observed random variable $\rho(W, t, \xi_0)$ $$ \phi(t, \nu_0(x)) = \mathbb{E} \big[ \rho(W, t, \xi_0) \mid X = x \big]. $$ The identity of the minimizer is assumed unique

align[align omitted — 88 chars of source]

almost surely in $P_X$. The envelope regression function is $$ \min_{t \in \mathcal{T}} \phi(t, \nu_0(x)) = \phi(t_0(x), \nu_0(x)). $$ Notice that this function can be written as $$ \min_{t \in \mathcal{T}} \phi(t, \nu_0(x)) = \sum_{t \in \mathcal{T}} \phi(t, \nu_0(x)) 1\{t = \arg \min_{t \in \mathcal{T}} \phi(t, \nu_0(x))\}. $$ Replacing each function $\phi(t, \nu_0(x))$ by its respective "unbiased signal" $\rho(W, t, \xi_0)$ gives the envelope moment function $$ \sum_{t \in \mathcal{T}} \rho(W, t, \xi_0) 1\{t = \arg \min_{t \in \mathcal{T}} \phi(t, \nu_0(X))\} $$ where $\xi_0$ is the true value of the nuisance parameter $\xi$. By the law of iterated expectations, $$ \psi_0 = \mathbb{E}[\rho(W, t_0, \xi_0)] = \mathbb{E}_X \big[\mathbb{E}[\rho(W, t_0, \xi_0) \mid X]\big] = \mathbb{E}_X \big[\min_{t \in \mathcal{T}} \phi(t, \nu_0(X))\big]. $$ The paper relies on standard cross-fitting schick1986asymptotically, as commonly used in debiased machine learning chernozhukov2016double, AtheyWager2, LRSP.

definition[Cross-Fitting] \begin{compactenum} • For a random sample of size $N$, denote a $K$-fold random partition of the sample indices $[N] = \{1, 2, \dots, N\}$ by $(J_k)_{k=1}^K$, where $K$ is the number of partitions and the sample size of each fold is $n = N / K$. For each $k \in [K] = \{1, 2, \dots, K\}$, define $J_k^c = [N] \setminus J_k$. • For each $k \in [K]$, construct an estimator $\widehat{\xi}_k = \widehat{\xi}(W_{i \in J_k^c})$ of the nuisance parameter $\xi_0$ using only the data $\{W_i : i \in J_k^c\}$. Define the first-stage fitted values \begin{align} \widehat{t}_i = \widehat{t}(X_i) = \widehat{t}_k(X_i) := \arg \min_{t \in \mathcal{T}} \phi(t, \widehat{\nu}_k(X_i)), \quad i \in J_k, \\ \widehat{\xi}_i := \widehat{\xi}(X_i) = \widehat{\xi}_k(X_i), \quad i \in J_k. \nonumber \end{align} \end{compactenum}
definition[Estimator] Given the first-stage fitted values, define $$ \widehat{\psi} := \frac{1}{N} \sum_{i=1}^N \rho(W_i, \widehat{t}_i, \widehat{\xi}_i). $$
definition[Multiplier Bootstrap] Let $(e_i)_{i=1}^N$ be a sequence of i.i.d. Exp(1) random variables independent of the data. Define $$ \widetilde{\psi} = \frac{1}{N} \sum_{i=1}^N \frac{e_i}{\bar{e}} \rho(W_i, \widehat{t}_i, \widehat{\xi}_i), \label{eq:psi:boot} $$ where $\bar{e} = N^{-1} \sum_{i=1}^N e_i$.

Under some conditions on the nuisance parameter discussed below, the proposed estimator enjoys the following properties

enumerate• With probability (w.p.) $\rightarrow 1$, the estimator converges at a root-$N$ rate: $$ | \widehat{\psi} - \psi_0 | = O_P(1/\sqrt{N}) = o_P(1). \label{eq:urate} $$ • The estimator $\widehat{\psi}$ is asymptotically linear: $$ \sqrt{N} (\widehat{\psi} - \psi_0) = \sqrt{N} \bigg( N^{-1} \sum_{i=1}^N \rho(W_i, t_0, \xi_0) - \psi_0 \bigg) + o_P(1), $$ and, therefore, asymptotically Gaussian: $$ \sqrt{N} (\widehat{\psi} - \psi_0) \Rightarrow^d N(0, V_0). \label{eq:limit} $$ Its asymptotic variance: \begin{equation} V_0 := \mathbb{E} \big[\rho^2(W, t_0(X), \xi_0(X))\big] - \psi_0^2 \end{equation} can be estimated by the sample analog: \begin{equation} \widehat{V} = N^{-1} \sum_{i=1}^N \rho^2(W_i, \widehat{t}_i, \widehat{\xi}_i) - \widehat{\psi}^2. \end{equation}

The paper establishes the theoretical framework for two inferential approaches. A plug-in $100(1-\alpha)\%$ confidence interval (CI) for $\psi_0$ can be constructed as

equation[equation omitted — 171 chars of source]

where $z_{1-\alpha}$ is the $(1-\alpha)$-quantile of $N(0, 1)$. As shown in Theorem (ref), the estimator $\widehat{V}$ is consistent for $V_0$, which implies $$ \Pr \big(\psi_0 \in CI_{1-\alpha}\big) \rightarrow 1-\alpha, \quad N \rightarrow \infty. $$ The plug-in variance estimator may be sensitive to biased estimation of $\xi$, which could affect the coverage of the plug-in confidence interval in small samples. An alternative to the plug-in procedure is to consider a bootstrap analog of the estimator $\widehat \psi$. A bootstrap confidence interval $CI^b_{1-\alpha}$ can be constructed as

equation[equation omitted — 165 chars of source]

where the critical values $\widehat{C}_{\alpha/2}$ and $\widehat{C}_{1-\alpha/2}$ are the $\alpha/2$ and $1-\alpha/2$ quantiles of the bootstrapped statistic $\sqrt{N} (\widetilde{\psi} - \widehat{\psi})$. Thus

equation[equation omitted — 130 chars of source]
remark[Uniqueness of Minimizer] The paper's results rely on the assumption that the minimizer $t_0(X)$ in (ref) is almost surely unique. In the context of Example (ref), this condition simplifies to \begin{align} P_X(m(1, X) - m(0, X) \neq 0) = 1. \end{align} This assumption is fundamental as it allows us to stay within the standard Gaussian framework. If this condition holds, the parameter $\psi_0$ is a pathwise differentiable parameter with a finite efficiency bound LuedtkeLaan (cf. Lemma (ref) in Appendix). Otherwise, regular estimators of the optimal welfare may not exist HiranoPorter2012.

Requiring the minimizer to be unique is a non-trivial restriction on the data generating process. Remark (ref) sketches a smoothing approach that could be used if condition (ref) is not plausible.

remark[Smoothing Alternative] Following levis2023covariateassisted, consider a log-sum-exp (LSE) function \begin{equation} g_{\widetilde{\kappa}}(\boldsymbol{v}) = \frac{1}{\widetilde{\kappa}}\log \left( \exp^{\widetilde{\kappa} v} + 1 \right), for \boldsymbol{v} \in \mathbb{R}, \end{equation} where $\widetilde{\kappa}$ is a tuning parameter and $\boldsymbol{v} \in \mathbb{R}$ is the argument. Noting that $$ \max \left\{ \boldsymbol{v}, 0 \right\}<g_{\widetilde{\kappa}}(\boldsymbol{v}) \leq \max \left\{\boldsymbol{v}, 0 \right\}+\frac{\log 2}{\widetilde{\kappa}} $$ allows to bound the approximation bias $\mathbb{E} [g_{\widetilde{\kappa}}(\nu_0(X)) ] - \mathbb{E} [ \max (m(1,X)- m(0,X),0)]$. The concurrent and independent work by levis2023covariateassisted develops asymptotic efficiency theory for the smoothened analog.

Theoretical Results

Section (ref) states the assumptions required for asymptotic theory. Section (ref) describes a key condition required for asymptotic theory and verifies it for single-index models. Section (ref) states the theoretical results.

Assumptions

Assumption (ref) ensures that the moment functions $\rho(W, t, \xi)$ are robust to first-order biases in the nuisance parameter $\xi$ uniformly over the index set $\mathcal{T}$.

assumption[Small Bias Condition] There exists a sequence $\epsilon_N = o(1)$, such that with probability at least $1-\epsilon_N$, for every partition index $k \in [K]$, the first stage estimate $\widehat{\xi}_k$, obtained by cross-fitting, belongs to a shrinking neighborhood of $\xi_0$, denoted by $\Xi_N$. Uniformly over $\Xi_N$, the following mean square convergence holds: \begin{align} B_N=\sup_{ \xi \in \Xi_N} \sup_{t \in \mathcal{T}} \sup_{x \in \mathcal{X}} \sqrt{N} |\mathbb{E} [ \rho(W,t,\xi)- \rho(W,t,\xi_0) \mid X=x] | = o(1). \end{align} Furthermore, the second-order terms are bounded as \begin{align} \Lambda_N =\sup_{ \xi \in \Xi_N} \sup_{t \in \mathcal{T}} \sup_{x \in \mathcal{X}} \mathbb{E} [(\rho(W,t,\xi)- \rho(W,t,\xi_0))^2 \mid X=x] = o(1). \end{align}

The small bias property is often attained via orthogonalization, a technique that has been widely studied in, e.g. Newey1994, chernozhukov2016double, and LRSP. For example, if $\rho(W, t, \xi) = \rho(W,t)$ does not involve any nuisance parameters, Assumption (ref) is automatically satisfied.

continuance{(ref)} Let $\mathcal{D} =\{0,1\}$. Consider an IPW-type signal of HIR2003 taking the form $$ \rho(W,1) = \dfrac{D}{\mu_{10}(X)} Y, \quad \rho(W,0) = \dfrac{1-D}{\mu_{00}(X)} Y$$ where the propensity score is assumed known. Thus, Assumption (ref) is automatically satisfied with $B_N = \Lambda_N=0$.

Assumption (ref) is shown to be satisfied if the signal is orthogonal with respect to the nuisance function $\xi_0$ whose estimator converges at a $o(N^{-1/4})$ mean square rate (e.g., SemCher).

continuance{(ref)} Let $\mathcal{D} = \{0,1\}$. Consider the Robins-type signal of Robins with $\rho(W,1,\xi)$ and $\rho(W,0,\xi)$ defined as in (ref). Suppose each function $m(1,X)$ and $m(0,X)$ is estimated at a mean square rate $m_N$, and the function $\mu_{10}(X)$ is estimated at rate $\mu_N$. Under Assumption 4.11 in SemCher, Assumption (ref) holds if $$ B_N = O(\sqrt{N} m_N \cdot \mu_N) =o(1), \qquad \Lambda_N = O (m_N + \mu_N) = o(1). $$

Assumption (ref) is a rate condition on the nuisance parameter $\nu$ entering (ref).

assumption[Rate] There exists a sequence $\epsilon_N = o(1)$, such that with probability at least $1-\epsilon_N$, for all $k \in [K]$, the first stage nuisance estimate $\widehat{\nu}_k (\cdot)$ belongs to a shrinking neighborhood of its true value $\nu_0(\cdot)$, denoted by $\mathcal{T}^{\nu}_N$. Uniformly over $\mathcal{T}^{\nu}_N$, the following worst-case rate bound holds. $$ \sup_{ \nu \in \mathcal{T}^{\nu}_N} \sup_{x \in \mathcal{X}} \| \nu(x) - \nu_0(x) \| \leq \nu^{\infty}_N = o (N^{-1/4}). $$

When the convergence is required in mean square ($\ell_2$) norm, Assumption (ref) is a classic assumption in the semiparametric literature (e.g., Newey1994). The example below demonstrates the plausibility of Assumption (ref) in $\ell_{\infty}$ norm.

continuance{(ref)} Let $\mathcal{D} = \{0,1\}$. Suppose regression functions in Example (ref) obey linear index restrictions \begin{align*} m(1,X)= X^{\prime} \gamma_1, \quad m(0,X) = X^{\prime} \gamma_0. \end{align*} Then, the $\ell_1$-regularized estimator of Program provides a convergence rate bound in $\ell_1$-norm in terms of the sparsity index, which suffices for a uniform rate bound on functions $m(1,x)$ and $m(0,x)$.
assumption[Regularity Conditions] The following technical conditions hold. (1) The moments are uniformly bounded a.s. \begin{align} \sup_{ \xi \in \Xi_N}\sup_{t \in \mathcal{T}} \sup_{x \in \mathcal{X}} | \rho(W,t,\xi) | \leq B_{\rho}. \end{align} (2) The bounded derivative condition holds: \begin{align} \sup_{\nu\in\mathcal{T}_N^\nu} \sup_{x\in\mathcal{X}}\sup_{t \in \mathcal{T}} \| \partial \phi (t, v(x))/ \partial v \| \leq B_{\phi}. \end{align}

Assumption (ref) is a standard regularity condition

footnote{The proof of Theorem (ref) only requires that the conditional second moment is uniformly bounded $ \label{eq:boundedsm} \sup_{ \xi \in \Xi_N}\sup_{t \in \mathcal{T}} \sup_{x \in \mathcal{X}} \mathbb{E} [ \rho^2(W,t,\xi) \mid X=x] \leq B_{\rho}. $ However, the a.s. bound (ref) is easier to verify in our examples which is why it is chosen for Assumption (ref).}

. For example, for the case of Example (ref), the condition (ref) automatically holds with $B_{\phi}=1$ since $\phi (d, v)=-\mathbf{e}_{d} v$ corresponds to taking the $d$'th element of the nuisance parameter $\nu_0(x) = (m(d,x))_{d \in \mathcal{D}}$.

assumption[Margin Assumption] There exist finite positive constants $\bar{B}, \delta \in (0, \infty)$ such that \begin{align} \sup_{(j,k) \in \mathcal{T}, \quad k \neq j} \Pr \left( 0 \leq \phi (j, \nu_0(X))- \phi (t, \nu_0(X)) \leq t \right) \leq \bar{B} t, \quad \forall t \in (0, \delta). \end{align}

Assumption (ref) is a margin condition that ensures separation between minimizers and non-minimizers in order to control the first-order effect of classification mistakes. It is a standard assumption in classification literature MammenTsybakov,Tsybakov,QianMurphy, policy learning KitagawaTetenov,MbakopTabord and debiased inference kallus2022whats,SemJoE,SemSupp2. Section (ref) verifies Assumption (ref) for a special case of single-index models.

Discussion of Assumption (ref)

In this section, I demonstrate the plausibility of Assumption (ref) in the context of Example (ref).

remark[Binary Treatment] Suppose $\mathcal{D} = \{0,1\}$. If the conditional average treatment effect $m(1,X) -m(0,X)$ has a bounded density \begin{align} \exists \delta>0, \ s.t. \sup_{t \in (-\delta, \delta) } f_{ m(1,X) - m(0,X)} (t) \leq \bar{B}_f, \end{align} then Assumption (ref) is satisfied with $\bar{B}= 2 \bar{B}_f$.

Remark (ref) verifies that Assumption (ref) holds with $\bar{B} = 2 \bar{B}_f$ as long as the index set $\mathcal{T}$ has two elements and their difference $m(1,X)-m(0,X)$ has a bounded unconditional density. This is a known result in the literature Tsybakov,KitagawaTetenov,kallus2022whats.

Lemma (ref) describes a class of linear models where the presence of covariate vector $X$ with a smooth distribution suffices for Assumption (ref). Suppose the expectation functions are partially linear

align[align omitted — 118 chars of source]

where $\widetilde{X} = (X, \bar{X})$ and $ \bar{X} $ is independent of $X$.

lemma[Linear Model] Suppose \begin{enumerate} • For any $j,k \in \mathcal{T}$, $\gamma_k \neq \gamma_j$. • For some $M < \infty$, $\max_{t \in \mathcal{T}} |g_t(\bar{X})| \leq M$ almost surely. • The vector $X$ obeys a smoothness condition \begin{align} \sup_{\delta \in \mathrm{R}^{p_X}, \| \delta \|=1} \Pr \left( 0 < X^{\prime} \delta < t \right) \leq \bar{B} t, \quad t \rightarrow 0. \end{align} \end{enumerate} Then, Assumption (ref) is satisfied.

Lemma (ref) verifies Assumption (ref) as long as the model (ref) includes a linear component obeying (ref). Condition (i) ensures that the mapping $x \mapsto \min_{t \in \mathcal{T}} (x^{\prime} \gamma_t)$ has a unique minimum. Condition (ii) accommodates the inclusion of arbitrary covariates (i.e., either continuous or discrete), provided their influence on the conditional mean remains almost surely bounded. Condition (iii) is a smoothness condition on $X$ similar to those in the analysis of least absolute deviation in Powell or the analysis of support function in CCMS. For example, if $X \sim N(h, \Sigma)$ is a Gaussian $p_X$-vector, the scalar $X^{\prime} \delta \sim N(\delta^{\prime} h, \delta^{\prime} \Sigma \delta)$, and the condition (ref) holds with $ \bar{B} = \sqrt{2 \pi \delta^{\prime} \Sigma \delta}^{-1} \leq (\sqrt{ 2 \pi})^{-1} \lambda^{-1/2}_{\min} (\Sigma)$.

Lemma (ref) extends the result of Lemma (ref) to nonlinear models. Suppose the expectation functions are

align[align omitted — 95 chars of source]
lemma[Nonlinear Model] Suppose \begin{enumerate} • Conditions (i)-(iii) of Lemma (ref) hold. • The covariate vector $X$ is $B_X$-bounded, and the support of $\cup_{t \in \mathcal{T}} X^{\prime} \gamma_t$, denoted by $ \mathcal{X}$, is compact. • The link function derivative is bounded from below on $\mathcal{X}$ \begin{align*} \inf_{ t \in \mathcal{X}} \frac{dF(t)}{dt} \geq f >0. \end{align*} \end{enumerate} Then, Assumption (ref) is satisfied.

Lemma (ref) verifies Assumption (ref) for a single-index model with a monotone link function provided the linear index obeys (ref). Condition (iii) is satisfied for a large class of link functions, such as probit, logit or uniform $U[-t, t]$ where $\mathcal{X}$ is included in $[-t,t]$.

In conclusion, let me point out that Assumption (ref) can be viewed as a special case of a general form of the margin assumption

align[align omitted — 269 chars of source]

that is routinely imposed in standard debiased inference LuedtkeLaan,kallus2020assessing,SemSupp2. Relaxing this assumption remains an important open question in the literature. When this assumption is violated, the cross-fit plug-in estimators have a non-standard, heavy-tailed distribution which makes standard Wald-type inference not valid LuedtkeLaan,ponomarev2024.

Asymptotic Results

theorem[Asymptotic Theory] Under Assumptions (ref)--(ref), the proposed estimator obeys the oracle property \begin{align} \sqrt{N} ( N^{-1} \sum_{i=1}^N \rho(W_i, \widehat t_i, \widehat \xi_i) -N^{-1} \sum_{i=1}^N \rho(W_i, t_0, \xi_0)) = o_P(1). \end{align} Therefore, it is asymptotically Gaussian $$ \sqrt{N} (N^{-1} \sum_{i=1}^N \rho(W_i, \widehat t_i, \widehat \xi_i) - \psi_0) \Rightarrow^d N(0, V_0), $$ with the asymptotic variance in (ref).

Theorem (ref) is my first main result

footnote{The oracle property is reminiscent of Neyman orthogonality in the double/debiased machine learning literature Neyman:1959,Neyman:1979,chernozhukov2016double. In the context of Example (ref), the derivative calculation appears in the calculation of the efficiency bound (Lemma (ref), see the proof of Theorem 1 in Supplement). }

. As a special case, it nests recent debiased inference results for sharp Makarov bounds with a binary outcome kallus2022whats and for BalkePearl1994,BalkePearl1997 bounds in the concurrent, independent work of levis2023covariateassisted. The paper's contribution is to introduce a general framework of aggregated intersection of regression functions for which the oracle property also applies. The new applications of the framework include Roy model bounds and Horowitz-Manski-Lee bounds with discrete outcomes, discussed in Section (ref).

theorem[Consistent Estimation of Asymptotic Variance] Under Assumptions (ref)--(ref), the sample analog estimator $\widehat V$ in (ref) is consistent for $V_0$ in (ref), that is $\widehat V - V_0 = O_P (\nu^{\infty}_N + \Lambda_N^{1/2}) = o_P(1)$.

Theorem (ref) establishes consistency of the plug-in estimator of variance which suffices for the validity of the confidence interval (ref). Unlike the envelope score estimator $\widehat \psi$, the variance estimator $\widehat V$ is first-order sensitive to the mistakes in estimated minimizers and converges at rate $\nu^{\infty}_N$ rather than $(\nu^{\infty}_N)^2$.

theorem[Bootstrap Inference] Under Assumptions (ref)--(ref), for $\widehat{c}_N(1-\alpha)$ being the $(1-\alpha)$-quantile of $ \widetilde{S}_N= \sqrt{N} (\widetilde \psi - \widehat \psi)$ under $P^e$, \begin{align*} \Pr ( \sqrt{N} (\widehat \psi - \psi_0) \leq \widehat{c}_N(1-\alpha)) \rightarrow 1-\alpha, \end{align*} which implies (ref).

Theorem (ref) establishes the validity of multiplier bootstrap inference. Note that the bootstrap-based inference is possible because the kink point (the point of non-differentiability) occurs with probability zero. A similar validity argument is used to establish bootstrap inference for support function as in CCMS and SemJoE.

Applications

Roy model with binary outcome

I begin by reviewing Roy model. Let $Y(1)$ and $Y(0)$ be two binary potential utility values, corresponding to the choices of treatment $D=1$ and $D=0$, respectively. An individual chooses the treatment value $D$ according to

align[align omitted — 94 chars of source]

If the potential outcomes $Y(1)$ and $Y(0)$ are equal, the choice is unspecified. The observed data $W=(X,D,Y)$ consist of the covariates $X$, the choice variable $D$, and the observed outcome $Y=DY(1) + (1-D) Y(0)$.

proposition[Proposition 1, MourifieHenry] Suppose there exists a vector $Z$ such that it satisfies the exogeneity restriction \begin{align} (Y(1), Y(0)) \perp\!\!\!\perp Z. \end{align} Then, the sharp bounds on the joint distribution of potential outcomes are \begin{align} \Pr (Y(1)=1, Y(0) =0) &\leq \min_{z \in \mathcal{Z}} \Pr (Y=1, D=1\mid Z=z), \\ \Pr (Y(1)=0, Y(0) =1) &\leq \min_{z \in \mathcal{Z}} \Pr (Y=1, D=0\mid Z=z). \end{align}

Proposition (ref) restates the Proposition 1 of MourifieHenry. It gives sharp bounds in Roy model with an instrument $Z$ obeying exclusion restriction. Such variables are akin to typical instrumental variables. The examples of discrete-valued $Z$ provided in MourifieHenry include parental education, distance to a college, and attendance a Catholic high school.

assumption[Instruments and Covariates] (A) Conditional independence. The discrete-valued instrument $Z$ is independent of the potential outcomes conditional on $X$ \begin{align} (Y(1),Y(0)) \perp\!\!\!\perp Z \mid X. \end{align} (B) Complete independence. The discrete-valued instrument $Z$ is independent of the potential outcomes \begin{align} (Y(1),Y(0),X) \perp\!\!\!\perp Z. \end{align}

Assumption (ref) (A) is a relaxation of (ref) which requires that independence holds only conditional on covariates. Assumption (ref) (B) is equivalent to (ref).

proposition[Sharp bounds in Roy model with covariates] (A) Suppose Assumption (ref)(A) holds. Then, the sharp bounds on the joint distribution of potential outcomes are aggregated intersection bounds \begin{align} \Pr (Y(1)=1, Y(0) =0) &\leq \mathbb{E} [\min_{z \in \mathcal{Z}} \Pr (Y=1, D=1\mid Z=z, X)] \\ \Pr (Y(1)=0, Y(0) =1 ) &\leq \mathbb{E} [ \min_{z \in \mathcal{Z}} \Pr (Y=1, D=0\mid Z=z, X)] \end{align} (B) Jensen's inequality implies \begin{align} \mathbb{E} [\min_{z \in \mathcal{Z}} \Pr (Y=1, D=1\mid Z=z, X)] \leq \min_{z \in \mathcal{Z}} {\Pr}_{\mu} (Y=1, D=1\mid Z=z), \\ \mathbb{E} [ \min_{z \in \mathcal{Z}} \Pr (Y=1, D=0\mid Z=z, X)] \leq \min_{z \in \mathcal{Z}} {\Pr}_{\mu} (Y=1, D=0\mid Z=z), \end{align} where \begin{align} {\Pr}_{\mu} (Y=1, D=d\mid Z=z) := \dfrac{ \mathbb{E} \bigg[ \dfrac{ 1\{ D=d \} \cdot Y \cdot 1\{ Z=z \}}{ \mu_{z0}(X)} \bigg]}{ \Pr (Z=z)}, \quad z \in \mathcal{Z}. \end{align} (C) Furthermore, if the propensity score is constant in $X$, $$ \mu_{z0}(X) = \Pr (Z=z), \quad z \in \mathcal{Z}, \quad \text{ a.s. }, $$ (ref)--(ref) reduce to regular bounds in Proposition (ref) \begin{align} {\Pr}_{\mu} (Y=1, D=1\mid Z=z) = \Pr (Y=1, D=d\mid Z=z), \quad \forall d \in \{0, 1\}. \end{align}

Proposition (ref) refines Proposition (ref) by incorporating covariate information. The sharp bounds are derived by applying the argument in Proposition (ref), conditional on $X$, and then aggregating over the covariate space. Interchanging expectation and minimum gives another pair of bounds of the form (ref)--(ref). Since they do not involve any expectation functions and are simpler to estimate, we refer to them as basic (or no-covariate) bounds. If the propensity score is constant, these basic bounds coincide with the original bounds defined in Proposition (ref).

Let me demonstrate the proposed inferential methodology focusing on the first bound in (ref). The bound

align[align omitted — 107 chars of source]

is a special case of (ref) with $\mathcal{T} = \mathcal{Z}$, the nuisance vector-function $$ \nu_0(x) = (\Pr (D =1, Y=1 \mid Z=z, X=x))_{z \in \mathcal{Z}}, $$ and the projection functions $$ \phi(z, v) = v_z \quad z \in \mathcal{Z}. $$ Furthermore, it can be also mapped to the optimal welfare parameter $\psi_0$ in Example (ref) with

align[align omitted — 92 chars of source]

Following Example (ref), define the orthogonal score for each $\phi(z,\nu_0(x))$ as

align*[align* omitted — 149 chars of source]

where the true value $\xi_0$ of the nuisance parameter is $$ \xi_0(x) = (\nu_0(x), \mu_{z0}(x)), \qquad \mu_{z0}(x) = \Pr (Z=z \mid X=x), \quad z \in \mathcal{Z}. $$ Given the true functions $\nu_{z0}(\cdot), \mu_{z0}(\cdot)$ and sequences of shrinking neighborhoods $\mathcal{T}_N^z$ of $\nu_{z0}(\cdot) $ and $M_N^z$ of $\mu_{z0}(x)$, define the following rates:

align*[align* omitted — 411 chars of source]

Assumption (ref) states the regularity conditions. First, it requires the instrument $Z$ to be discrete-valued with finite support so that the propensity score

align[align omitted — 124 chars of source]

is bounded away from zero and one for each distinct instrument value. Continuously supported instrument (i.e., $\mathcal{T}=\mathcal{Z}$) is outside of the scope of this paper since their propensity score violates (ref). In this case, I conjecture that the propensity score needs to be approximated by a kernel density estimator, e.g. as is standard in the results for continuous treatments Colangelo. Second, the mean square rates of the first-stage estimators must decay sufficiently fast, a condition that is standard in semiparametric estimation literature.

assumption[Regularity Conditions for Roy Model with Covariates] Assume that there exists a sequence of numbers $\epsilon_N = o(1)$ and sequences of neighborhoods $\mathcal{T}^z_N$ of $\nu_{0z}(\cdot) $ and $M_N^z$ of $\mu_{0z}(x)$ such that both the true value $\xi_0(x)$ and the first-stage estimate $\{\widehat{\nu}_z(\cdot), \widehat{\mu}_z(\cdot) \} $ belong to the set $\{\mathcal{T}^z_N \bigtimes M_N^z\}$ w.p. at least $1-\epsilon_N$ for each $ z \in \mathcal{Z}$. The functions in $M_N^z$ are bounded uniformly over their domain from above and below by $\kappa$ and $1-\kappa$. Finally, assume that mean square rates $\nu_N,\mu_N$ decay sufficiently fast: $$N^{1/2} \nu_N\mu_N = o(1), \quad \nu_N \vee\mu_N = o(1), \quad \nu_N^{\infty} = o(N^{-1/4}).$$

Corollary (ref) gives an envelope score estimator for Roy model bounds and delivers uniformly valid debiased inference in the presence of covariates. As described in (ref), the sufficient conditions for the margin assumption discussed in Section (ref) are equally applicable to Roy model bounds.

corollary[Asymptotic Theory for Roy Model with Covariates] Suppose Assumptions (ref) (A) and (ref) hold, and Assumption (ref) holds for $\nu_0(X) = (\Pr (D =1, Y=1 \mid Z=z, X))_{z \in \mathcal{Z}}$. Then, the statements of Theorems (ref)--(ref) hold for the estimator of Definition (ref) with $\psi_0$ in (ref).

Horowitz-Manski-Lee bounds with discrete outcomes.

I begin by introducing the sample selection problem. Let $D=1$ be an indicator for treatment receipt. Let $Y(1)$ and $Y(0)$ denote the potential outcomes if an individual is treated or not, respectively. Likewise, let $S(1)=1$ and $S(0)=1$ be dummies for whether an individual's outcome is observed with and without treatment, respectively. The data vector $W=(D,X,S,S \cdot Y)$ consists of the treatment status $D$, a baseline covariate vector $X$, the observed selection status $S=D \cdot S(1) + (1-D) \cdot S(0)$ and the observed outcome $S \cdot Y = S \cdot (D \cdot Y(1) + (1-D) \cdot Y(0))$ for selected individuals. LeeBound focuses on the average treatment effect (ATE)

align[align omitted — 90 chars of source]

for subjects who are selected into the sample regardless of treatment receipt---the always-takers.

assumption[Assumptions of LeeBound] The following statements hold. \begin{compactenum}[(1)] • (Independence). The vector $(Y(1),Y(0),S(1),S(0))$ is independent of $D$ conditional on $X$. The propensity score \begin{align} \mu_{10}(X) := \Pr (D =1 \mid X), \qquad \mu_{00}(X) =1 - \mu_{10}(X) \end{align} is assumed known. • (Monotonicity). \begin{equation} S(1) \geq S(0) \quad a.s. . \end{equation} \end{compactenum}

Assumption (ref)(1) holds by random assignment. Assumption (ref)(2) states that all subjects must exhibit the same direction of selection response. It is frequently imposed in selection and treatment choice models. If covariates are available, this assumption has testable implications. Furthermore, it can be relaxed to conditional monotonicity Kolesar,SemSupp2. In this paper, we consider an unconditional version of the monotonicity assumption so as to focus on the theory for discrete-valued outcomes. As discussed in LeeBound, the average control outcome is point-identified $$ \mathbb{E} [ Y(0) \mid S(0)=1] = \mathbb{E} [ Y(0) \mid S(1)=1, S(0)=1] =\mathbb{E} [ Y \mid S=1, D=0].$$ I focus on the average treated outcome

align[align omitted — 85 chars of source]

In contrast to the control group, a treated outcome can be either an always-taker's outcome or a complier's outcome. The always-takers' share among the treated outcomes is

align[align omitted — 163 chars of source]

\paragraph{Binary outcomes. } Suppose $Y(1)$ and $Y(0)$ take values in $\{1, 0\}$. In the best case, the always-takers comprise the top $p_0$-quantile of the treated outcomes. Let $p_Y:= \Pr (Y =1 \mid D=1, S=1)$. In case when $p_0 <p_Y$, the best-case always-takers' outcome is equal to one for all always-takers. Otherwise, the best-case always-takers' outcome distribution is a mixture of ones and zeroes, with the mixing proportion of ones and zeroes equal to $p_Y/p_0$ and $1-p_Y/p_0$, respectively. In other words, the basic (i.e., no-covariate) upper bound $\bar{\beta}_U$ on $\mathbb{E}[ Y(1) \mid S(1) =1, S(0)=1]$ can be expressed as

align[align omitted — 132 chars of source]

Since $\bar{\beta}_U$ involves no covariates, we refer to it as the basic bound as opposed to the sharp bound derived further.

Lee's identification strategy can be implemented conditional on covariates. Denote the conditional trimming threshold $p_0(x)$ as

align[align omitted — 151 chars of source]

and the conditional probability of outcome one as $$ p_Y(x) = \Pr (Y = 1\mid D=1, S=1, X=x). $$ Finally, the conditional upper bound $\bar{\beta}_U(x)$ can be expressed as

align[align omitted — 87 chars of source]

Aggregating the conditional bound over the always-takers' covariate distribution gives the sharp upper bound

align[align omitted — 201 chars of source]

and a similar argument applies for the lower bound. Proposition (ref) derives the sharp lower and upper bounds for the average potential outcome.

propositionSuppose Assumption (ref) holds for a binary outcome $Y$ taking values in $\{0, 1\}$. Then, the following statements hold: (A) The sharp lower and upper bounds on $\beta_1$ in (ref) take the form of ratios \begin{align} \beta_L = \dfrac{N_L}{ \mathbb{E} [ s_0(0,X) ] }, \qquad \beta_U = \dfrac{N_U}{ \mathbb{E} [ s_0(0,X) ] }, \end{align} whose numerators are aggregated intersection bounds \begin{align} N_L &= \mathbb{E} [ \max (s_0(0,X) - s_0(1,X)(1-p_Y(X)), 0) ] , \\ N_U &= \mathbb{E} [\min (s_0(1,X) p_Y(X)-s_0(0,X), 0) +s_0(0,X) ] . \end{align} (B) Jensen's inequality implies \begin{align} \mathbb{E} [ \max (s_0(0,X) - s_0(1,X)(1-p_Y(X)), 0) ] &\geq \max (s_0 - s_1 (1-p_Y) , 0), \\ \mathbb{E} [\min (s_0(1,X) p_Y(X)-s_0(0,X), 0) ] &\leq \min ( p_Y s_1 - s_0, 0). \end{align}

Proposition (ref) derives basic and sharp bounds on the average potential outcome in a selection problem. The denominators of basic and sharp bounds are the same and equal the always-takers' share $$s_0 = \Pr [S=1 \mid D=0] = \mathbb{E} [ s_0(0,X) ].$$ Their numerators are regular and aggregated intersection bounds, described in the LHS and RHS of (ref)--(ref), respectively. The basic bounds (ref)--(ref) coincide (up to a constant) the bounds in the Lemma 1 of concurrent, independent work of kroft2024leeboundsmultilayeredsample.

\paragraph{Discrete outcomes. }In this section, I allow the outcome $Y$ to take a finite number of discrete values. Proposition (ref) characterizes the numerators of Horowitz-Manski-Lee bounds as a special case of aggregated intersection bounds.

proposition[Horowitz-Manski-Lee Bounds with Discrete Outcomes] Suppose Assumption (ref) holds with a discrete outcome $Y$ whose support is denoted by $\mathrm{T}$. (A) Then, the sharp lower and upper bound on $\beta_1$ are given in (ref) with $N_L$ and $N_U$ given in \begin{align} N_L &= \mathbb{E} [\max_{\beta \in \mathrm{T} } (\beta s_0(0,X) + s_0(1,X) \mathbb{E} [ \min (Y-\beta,0)\mid D=1, S=1, X])] \end{align} and the upper bound numerator \begin{align} N_U &= \mathbb{E} [ \min_{\beta \in \mathrm{T} } ( \beta s_0(0,X) + s_0(1,X) \mathbb{E}[ \max (Y-\beta,0)\ \mid D=1, S=1, X])]. \end{align} (B) The basic bounds on $\beta_1$ are given in (ref) with $N_L$ and $N_U$ given in \begin{align*} N_L &= \max_{\beta \in \mathrm{T} } (\beta s_0 + s_1 \mathbb{E} [ \min (Y-\beta,0)\mid D=1, S=1]) \end{align*} and the upper bound numerator \begin{align*} N_U &= \min_{\beta \in \mathrm{T} } ( \beta s_0 + s_1 \mathbb{E}[ \max (Y-\beta,0) \mid D=1, S=1]). \end{align*}

Proposition (ref) develops a novel representation of Horowitz-Manski-Lee bounds as regular and aggregated intersection bounds, respectively. The outcome distribution is represented using point mass functions (PMFs) rather than quantiles, which is convenient for working with discrete outcomes. As shown in RockUryasev, the minimum in (ref) is attained by the outcome quantile of level $1-s_0/s_1$ or the “borderline” always-takers' outcome.

\paragraph{Debiased inference. } To describe an inferential approach, I derive moment functions for (ref) and (ref). Define the moment functions for the lower bound

align[align omitted — 172 chars of source]

and for the upper bound

align[align omitted — 172 chars of source]

Next, let $ \mathrm{T}$ be the finite support of the outcome $Y$. Define the nuisance parameter

align*[align* omitted — 94 chars of source]

The resulting moment function for $N_U$ is

align[align omitted — 195 chars of source]

Likewise, the moment function for $N_L$ is

align[align omitted — 195 chars of source]

Given the true functions $\xi_{0}(\cdot)$ and sequences of shrinking neighborhoods $S^d_N$ of $s_0(d,x)$ and $\mathcal{P}^{\beta}_N$ of $\pi_{\beta 0}(x)=\Pr (Y = \beta \mid D=1, S=1, X=x)$, define the following rates:

align*[align* omitted — 305 chars of source]
assumption[Regularity Conditions for Horowitz-Manski-Lee Bounds] Assume that there exists a sequence of numbers $\epsilon_N = o(1)$ and sequences of neighborhoods $\mathcal{P}^{\beta}_N$ of $\pi_{\beta0} (\cdot) $ and $\mathcal{S}^d_N$ of $s_0(d,\cdot)$ such that both the true value $\xi_0(x)$ and the first-stage estimate $\{ \widehat{s}(d,\cdot), \widehat{\pi} (\beta, \cdot) \} $ belongs to the set $\{ \mathcal{S}^d_N \bigtimes \mathcal{P}^{\beta}_N \}$ w.p. at least $1-\epsilon_N$ for each $ \beta \in \mathrm{T}$ and $d \in \{1, 0\}$. (i) The rates $ \pi^{\infty}_N$ and $ s^{\infty}_N$ are sufficiently fast: $$ \pi^{\infty}_N + s^{\infty}_N = o(N^{-1/4}).$$ (ii) The functions in each set, as well as the propensity score $\mu_{10}(x)$ and $\pi_{\beta0}(x)$, are bounded uniformly over their domain from above and below by $\kappa$ and $1-\kappa$. The support set $\mathrm{T}$ is finite. (iii) The vector $(s_0(0,X), s_0(1,X), \cup_{ \mathrm{T} } \pi_{\beta0}(X) )$ is continuously distributed with a bounded joint density such that each of its component is supported on $(\kappa, 1- \kappa)$.

Assumption (ref) summarizes regularity conditions for Horowitz-Manski-Lee bounds. Since the propensity score is assumed known, the individual moment functions $\{ \rho_L (W, \beta), \rho_U (W, \beta)\}_{\{\beta \in \mathcal{T} \}}$ in (ref) and (ref) do not involve any nuisance parameters. As a result, Assumption (ref) is automatically satisfied for the individual functions.

corollary[Asymptotic Theory for Horowitz-Manski-Lee Bounds with Discrete Outcome] Suppose Assumptions (ref) and (ref) hold. Then, the statements of Theorems (ref)--(ref) hold for the estimator described in Algorithm (ref) as well as its bootstrap analog outlined in Definition (ref).

Corollary (ref) delivers a root-$N$ consistent, asymptotically Gaussian estimator of sharp Horowitz-Manski-Lee bounds assuming the conditional probability of selection and the conditional PMF are estimated at a sufficiently fast rate. To the best of my knowledge, this is a first example of debiased inference for the trimming bounds with discrete-valued outcome.

In conclusion, I state the Algorithm (ref) for computing the bounds as well as examples of the first-stage estimators.

example[Estimator of Selection Probabilities] Suppose the selection probability $s_0(d,x)$ for $d \in \{1, 0\}$ can be approximated by a logistic function \begin{align} s_0(d,x) = \Lambda ( x' \gamma^d_{0}) + r_d(x), \quad d \in \{1, 0\}, \end{align} where $\Lambda(\cdot) = \dfrac{\exp (\cdot)}{1 + \exp (\cdot)}$ is the logistic CDF, $\gamma^d_{0} \in \mathrm{R}^{p}$ is the pseudo-true value of the logistic parameter, and $r_d(x)$ is its approximation error. The logistic likelihood function is \begin{align} \ell_d(\gamma^d) =\dfrac{1}{N} \sum_{i=1}^N (D_i = d ) \bigg( \log ( 1+ \exp (X_i '\gamma^d)) - S_i X_i'\gamma^d \bigg), \quad d \in \{1, 0\}. \end{align} Given an estimate $\widehat{\gamma}^d$ of $\gamma^d$, define the estimated selection probabilities as \begin{align} \widehat{s}(d,x) &= \Lambda ( x' \widehat {\gamma}^d), \quad d \in \{1, 0 \} \end{align} and the estimated CATE on selection \begin{align*} \widehat{\tau}(x) = \widehat{s}(1,x) - \widehat{s}(0,x). \end{align*} Given the penalty parameter $\lambda_S$, the $\ell_1$-regularized logistic estimator of $\gamma^d$ orthogStructural,Program is \begin{align} \widehat{\gamma}^d_{L}= \arg \max_{\gamma^d \in \mathrm{R}^{p}} \ell_d(\gamma^d) + \lambda_S \| \gamma^d \|_1. \end{align}
example[Estimator of Outcome Probability Mass Function] Suppose the outcome PMF can be approximated by a multinomial logistic regression \begin{align} \pi_{\beta0}(x) = \Lambda_{\beta}( x' \delta_{\beta 0}) + r_{\beta}(x), \quad \beta \in \mathrm{T}, \end{align} where $\delta_{\beta 0} \in \mathrm{R}^{p}$ is the pseudo-true value of the logistic parameter and $$ \Lambda_{\beta}(x' \delta_{\beta 0}) = \frac{\exp(x' \delta_{\beta 0})}{1+\sum_{\beta \in \mathrm{T} \setminus \{0 \}} \exp(x' \delta_{\beta 0})} , $$ and $r_{\beta}(x)$ is the approximation error. The multiclass classification via sparse multinomial logistic regression is developed in abramovich2020multiclassclassificationsparsemultinomial.
algorithm[algorithm omitted — 1,686 chars of source]

Numerical Results

This section provides numerical evidence for the methods developed in this article. Section (ref) offers a Monte Carlo experiment constructed in the context of Example (ref). Section (ref) offers an empirical illustration of the method for Horowitz-Manski-Lee bounds in Section (ref).

Simulation Study

I build a simulation exercise on JTPA dataset Bloom1997 that consists of three elements: baseline covariates, treatment (access to job training), and outcome. The baseline covariate $X_1$ is taken to be the previous earnings PreEarn measured in $10, 000$ USD. The covariate vector $X=(X_1, \dots, X_1^p)$ includes the first $p$ powers of the PreEarn variable. The treatment $D$ is determined by a coin flip with probability $Pr(D = 1) = \frac{2}{3}$, to match the propensity score in JTPA data. The outcome $Y$ follows a linear model

align[align omitted — 73 chars of source]

where $\epsilon \sim N(0, \sigma^2)$ is a Gaussian shock independent of the data. The true parameter values are $$\kappa_0 = \gamma_0 =(2^{-1}, 2^{-2}, \dots, 2^{-p}), \quad \sigma^2 = 1.$$ The population data set size is $9, 223$. In addition to this primary design, we also consider an artificial (Gaussian) design where $X_1$ is drawn from a Gaussian distribution whose mean and variance matches the respective parameters of actual PreEarn variable, with all other steps being the same. Using this setup, we evaluate coverage of plug-in and bootstrap inferential confidence intervals based on the doubly robust estimator described in Example (ref). The first-stage functions $m(0, X)$ and $m(1, X)$ are estimated via linear least squares. The performance metrics include bias, mean squared error (MSE), and coverage rates of the confidence intervals (CIs). These metrics are analyzed across varying sample sizes $N \in \{100, 200, 300, 500\}$ and polynomial degrees $p \in \{1, 3, 5, 7\}$.

Tables (ref) and (ref) summarize the simulation results, highlighting key differences across polynomial degrees ($p$) and designs. For lower degrees ($p \in \{1, 3\}$), both designs yield estimators with low bias, low MSE, and near-nominal coverage rates for both plug-in and bootstrap CIs. The results are consistent with the theoretical results in Theorems (ref) and (ref). For higher degrees ($p \in \{5, 7\}$), the performance diverges significantly between the two designs. In the Gaussian design (Table (ref)), coverage remains close to the nominal rate even for $p=7$ when $N=500$. Conversely, in the primary design (Table (ref)), coverage drops sharply to $48\%$ (Plug-In CI) and $39\%$ (Bootstrap CI) for $p=7$. This discrepancy likely arises from the heavy-tailed nature of PreEarn, which could either affect the quality of the first-stage estimates of $\kappa_0$ and $\gamma_0$, make Assumption (ref) to be a poor fit for the data, or both.

table[table omitted — 1,681 chars of source]
table[table omitted — 1,677 chars of source]

Empirical application

To illustrate the immediate applicability of the proposed method, this study analyzes the effect of Medicaid exposure on healthcare utilization and health outcomes using data from the Oregon Health Insurance Experiment finkelstein. In 2008, Oregon implemented a limited expansion of its Medicaid program, providing insurance coverage to low-income, uninsured adults selected through a lottery system from a waiting list. One year after randomization, a subset of $N = 58,405$ applicants was mailed a survey to assess changes in healthcare utilization and general well-being, with a response rate of approximately $50\%$. Abstracting from potential non-response bias, the study found that Medicaid significantly improved healthcare access, financial security, and mental health outcomes for low-income adults.\footnote{finkelstein reported that the ability to reject the null hypothesis of no effect of health insurance on healthcare utilization or financial strain is generally robust to Lee bounds, while the ability to reject the null hypothesis of no effect on self-reported health outcomes is not robust (see footnote 19).} To examine the robustness of these findings, this section reports various versions of Horowitz-Manski-Lee bounds under various assumptions about subjects' response behavior.

This section focuses on the average treatment effect (ATE) of Medicaid exposure on self-reported mental health. Let $S(1) = 1$ and $S(0) = 1$ be binary indicators denoting whether a subject completes a survey when treated or not treated, respectively, and let $Y(1)$ and $Y(0)$ represent the corresponding potential outcomes. The observed data, $W = (X, D, S, S \cdot Y)$, include $D$ (lottery outcome), $S$ (an indicator of non-missing response), $Y$ (the response itself), and baseline covariates $X$. The propensity score, $\mu(X)$, is determined by household size and survey wave fixed effects.\footnote{If an applicant wins the lottery, all members of their household become eligible to enroll. Consequently, larger households have a higher probability of winning the lottery compared to smaller ones. Additionally, control applicants were oversampled in earlier survey waves, further influencing the propensity score.} Other components of $X$ include 64 predetermined characteristics, such as demographics, enrollment in the Supplemental Nutrition Assistance Program (SNAP) or Temporary Assistance for Needy Families (TANF) as well as the total amount of benefits received in each program, and pre-existing health conditions.

Table (ref) summarizes the findings. The baseline estimate of Medicaid's effect on mental health is $2.27\%$. Interestingly, the control group's response rate ($49.4\%$) exceeds that of the treated group ($48.2\%$). Unconditional monotonicity (Assumption (ref)), which posits that treatment discourages survey completion, lacks an intuitive explanation. Standard Lee bounds (Column (1)) are misleadingly tight. Without monotonicity, the proportion of “always-takers” (respondents regardless of treatment status) cannot be bounded away from zero, and no-monotonicity bounds (Column (5)) do not provide any meaningful restriction on the treatment effect

footnote{ZhangRubin bounds reported in Column (5) are driven by $Y(1), Y(0) \in \{1,0\}$.}

. This pattern holds for all survey outcomes, including questions about healthcare utilization, financial strain, and self-reported health outcomes.

To make progress, the analysis focuses on a subset of the population for whom the direction of selection response aligns with the majority. Specifically, the parameter of interest is:

align[align omitted — 106 chars of source]

where the outcome of interest is mental health as measured by a positive response to the question “Did you not screen positive for depression in the last two weeks?”). The selection equation (ref) is estimated using logistic regression (see Example (ref)). The estimated share of subjects with negative selection response is $82\%$ (Column (2)), consistent with the negative direction of unconditional response effect. A higher likelihood of survey response is associated with being female, requesting English-language materials, and not receiving TANF benefits. Additionally, Medicaid's effect on response appears negative for individuals who experienced injuries or received SNAP benefits prior to randomization, suggesting that control participants' response could be driven by acute health or financial challenges.

Basic bounds on (ref) (Table (ref), Column (3)), derived under conditional monotonicity, cannot determine the direction of the treatment effect. Sharp bounds (Table (ref), Column (4)), constructed using Algorithm (ref) with first-stage outcome fitted values as described in Examples (ref) and (ref), suggest that the Medicaid exposure effect is positive, though the magnitude is attenuated at the lower bound. The proposed approach relies on a smoothness assumption, specifically that the conditional selection and outcome probabilities are sufficiently continuously distributed, which could be plausible since some components of $X$ are continuously distributed, such as total amount of SNAP or TANF benefits or pre-existing ED charges.

table[table omitted — 1,671 chars of source]