Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,324 characters · 14 sections · 87 citation commands
Debiased Machine Learning of Aggregated Intersection Bounds and Other Causal Parameters
\linespread{1.5}
Keywords: optimal welfare, cross-fitting, double/debiased machine learning, margin assumption, uniformity, Roy model, selection problem, partial identification
Economists are often interested in bounds on parameters when parameters themselves are not point-identified Manski89,Manski90,Manski. Examples include quantiles of heterogeneous treatment effects and other distributional measures beyond the mean FanPark. Baseline or pre-treatment covariates often contain valuable information that can tighten these bounds ManskiPepper. However, in practice, sharp bounds are rarely utilized because their estimators usually have non-standard distributions driven by noisy first-stage estimators of unknown conditional distributions. These challenges are not unique to partial identification and also arise in related areas, such as statistical treatment choice LuedtkeLaan, KitagawaTetenov, AtheyWager2, MbakopTabord.
This paper develops estimation and inference methods for the quantities taking the form
where $X$ is a covariate vector, $\mathcal{T}$ is a finite index set, and $x \mapsto \nu_0(\cdot) = (\nu_{j0}(\cdot))_{j=1}^d$ is a $d$-dimensional nuisance parameter whose elements $\nu_{j0}(x)$ are functions of covariates, such as conditional expectations. As the simplest case, one can think of the optimal welfare which appears in e.g., LuedtkeLaan or sharp bound on distributional effects FanPark. The paper's contribution is to deliver a debiased inference on $\psi_0$ that is first-order insensitive to the misclassification mistake in the identity of the binding constraint. In particular, its distribution is the same as if the true value of the minimizer (ref) were known. Additionally, the paper establishes the validity of a weighted bootstrap method, which holds the estimated minimizer fixed while bootstrapping the second-stage statistic, providing a valid distributional approximation.
The paper illustrates the usefulness of the proposed approach by considering two applications in applied microeconomics. In particular, we discuss in detail a sharp version of Roy model bounds as studied in MourifieHenry as well as Horowitz-Manski-Lee bounds with discrete-valued outcomes a version of which have been also studied in concurrent, independent work of kroft2024leeboundsmultilayeredsample. Revisiting Oregon Health Insurance Experiment finkelstein, we find our methodology useful in determining the direction of treatment effect in the presence of non-response bias as well as tightening the bounds, echoing earlier work in SemSupp2.
The rest of the paper is organized as follows. Section (ref) gives a literature review. Section (ref) introduces the framework and provides two stylized examples. Section (ref) offers an informal preview of the results. Section (ref) presents the formal asymptotic theory and discusses low-level condition for the margin assumption in the context of a single-index model with continuous covariates. Section (ref) applies the proposed theory to sharp bounds in the Roy model and Horowitz-Manski-Lee bounds in selection problems. Section (ref) provides numerical evidence for the methods developed in the article. All proofs are in Appendix (ref).
This paper is related to two lines of research: partial identification and statistical treatment choice.
\paragraph{Bounds, Convex Optimization, and Directionally Differentiable Functionals.} Set identification is a vast area of research, encompassing a wide variety of approaches: linear and quadratic programming, random set theory, support function, and moment inequalities Manski90, ManskiPepper, Manski:2002, HaileTamer, CHT, BM, Molinari2008, CilibertoTamer, LeeBound, Stoye, AndrewsShiECMA, BMM2, CCMS, CherNeweySantos, Gafarov, kallus2020localized, li2022discordant, henry2023role, acerenza2023marginal, ban2021nonparametric, bartalotti2021identifying, JLS, fava2024, see e.g. Molinari:2018 or MolinariHandbook for a review. In the context of distributional effects Makarov, Manski,HSC, FanPark,FanPark2, Tetenov, FanZhu, FirpoRidder, the first discussion of estimation can be traced to FanPark, where, on p.945 they sketch a plug-in estimation approach without statistical guarantees. Targeting the envelope function $\inf_{t \in T} s(t,x)$, the work by CLR proposes a plug-in approach based on the least squares series estimators, where large sample inference is based on the strong approximation of a sequence of series or kernel-based empirical processes. Switching the focus from the envelope function to its best linear predictor, CCMS proposes a root-$N$ consistent and uniformly asymptotically Gaussian estimator of the target parameter, relying on the first-stage series estimators. Finally, recent work by LeeSungwon focuses on bounds on conditional distributions of treatment effects. That is, most inference work focuses on the envelope function, rather than its mean value, which makes the lack of differentiability of $x \mapsto \min (x,0)$ at the kink point $x=0$ a common concern (e.g., FangSantos). Finally, the paper contributes to a growing literature on machine learning for bounds and partially identified models kallus2019assessing, jeong2020robust, SemJoE and sensitivity analysis DornGuo, DornGuoKallus, Bonvini_2021, Bonvini_2022, see e.g., Kennedy_review for the review.
\paragraph{Statistical treatment choice. } Statistical treatment choice is a vast area of research, focusing on two distinct questions: learning the best policy QianMurphy, KitagawaTetenov, MbakopTabord, AtheyWager2, Sun for a given criterion function and inference on the optimal value of the criterion itself LuedtkeLaan in an unconstrained policy class. In the first stream, recent work focused on various robustness aspects of existing criteria functions ishihara2021evidence, adjaho2023externally or targeting welfare criteria that are partially identified Stoye, Pu_2021, Cui_2021, kitagawa2023treatment, yata2023optimal, dadamo2022orthogonal, christensen2023optimal, benmichael2023policy, olea2023decision, cui2024policy. For example, such criteria may arise in asymmetric loss functions babii2021binary, partial welfare ordering han2021optimal,firpo2023loss or distributional welfare based on quantile treatment effects cui2024policy.
This paper studies the aggregated intersection of regression functions
where $X$ is a covariate vector, $\mathcal{T}$ is a finite index set, and $x \mapsto \nu_0(\cdot) = (\nu_{j0}(\cdot))_{j=1}^d$ is a $d$-dimensional nuisance parameter whose elements $\nu_{j0}(x)$ are functions of covariates, such as conditional expectations. For each element $t$ of the set $\mathcal{T}$, $\phi(t, \cdot): \mathrm{R}^d \rightarrow \mathrm{R}$ is a known scalar function of the vector $\nu_0$, which could represent a projection onto the Euclidean axis. This function can be expressed as a conditional expectation
where $W$ is the data vector and $\rho(W, t, \xi_0)$ is an observed random variable that depends on a nuisance parameter $\xi_0$.
Examples of $\psi_0$ include Fréchet-Hoeffding bounds, Makarov bounds on distributional effects, and sharp versions of \citep*{BalkePearl1994, BalkePearl1997} bounds. In statistical treatment choice, examples include optimal welfare in an unconstrained policy class Manski2004,QianMurphy, LuedtkeLaan,KitagawaTetenov. As a stylized example, this section revisits optimal welfare in statistical treatment choice and Makarov bounds on distributional effects FanPark.
Example (ref) describes Makarov bounds on the treatment effect CDF, previously studied in FanPark and FirpoRidder, among others. A special case of this example with binary outcomes was studied in kallus2022whats, who proposed debiased inference for $\pi_L(0)$ and $\pi_U(0)$. Other interesting examples of this parameter are described in Section (ref).
In this section, I introduce the estimator of the parameter of interest and describe two inferential approaches. Let me briefly review the notation. Recall that $W$ is a data vector, $\mathcal{T}$ is an index set, and each function $\phi(t, \nu_0(x))$ is a conditional expectation function of an observed random variable $\rho(W, t, \xi_0)$ $$ \phi(t, \nu_0(x)) = \mathbb{E} \big[ \rho(W, t, \xi_0) \mid X = x \big]. $$ The identity of the minimizer is assumed unique
almost surely in $P_X$. The envelope regression function is $$ \min_{t \in \mathcal{T}} \phi(t, \nu_0(x)) = \phi(t_0(x), \nu_0(x)). $$ Notice that this function can be written as $$ \min_{t \in \mathcal{T}} \phi(t, \nu_0(x)) = \sum_{t \in \mathcal{T}} \phi(t, \nu_0(x)) 1\{t = \arg \min_{t \in \mathcal{T}} \phi(t, \nu_0(x))\}. $$ Replacing each function $\phi(t, \nu_0(x))$ by its respective "unbiased signal" $\rho(W, t, \xi_0)$ gives the envelope moment function $$ \sum_{t \in \mathcal{T}} \rho(W, t, \xi_0) 1\{t = \arg \min_{t \in \mathcal{T}} \phi(t, \nu_0(X))\} $$ where $\xi_0$ is the true value of the nuisance parameter $\xi$. By the law of iterated expectations, $$ \psi_0 = \mathbb{E}[\rho(W, t_0, \xi_0)] = \mathbb{E}_X \big[\mathbb{E}[\rho(W, t_0, \xi_0) \mid X]\big] = \mathbb{E}_X \big[\min_{t \in \mathcal{T}} \phi(t, \nu_0(X))\big]. $$ The paper relies on standard cross-fitting schick1986asymptotically, as commonly used in debiased machine learning chernozhukov2016double, AtheyWager2, LRSP.
Under some conditions on the nuisance parameter discussed below, the proposed estimator enjoys the following properties
The paper establishes the theoretical framework for two inferential approaches. A plug-in $100(1-\alpha)\%$ confidence interval (CI) for $\psi_0$ can be constructed as
where $z_{1-\alpha}$ is the $(1-\alpha)$-quantile of $N(0, 1)$. As shown in Theorem (ref), the estimator $\widehat{V}$ is consistent for $V_0$, which implies $$ \Pr \big(\psi_0 \in CI_{1-\alpha}\big) \rightarrow 1-\alpha, \quad N \rightarrow \infty. $$ The plug-in variance estimator may be sensitive to biased estimation of $\xi$, which could affect the coverage of the plug-in confidence interval in small samples. An alternative to the plug-in procedure is to consider a bootstrap analog of the estimator $\widehat \psi$. A bootstrap confidence interval $CI^b_{1-\alpha}$ can be constructed as
where the critical values $\widehat{C}_{\alpha/2}$ and $\widehat{C}_{1-\alpha/2}$ are the $\alpha/2$ and $1-\alpha/2$ quantiles of the bootstrapped statistic $\sqrt{N} (\widetilde{\psi} - \widehat{\psi})$. Thus
Requiring the minimizer to be unique is a non-trivial restriction on the data generating process. Remark (ref) sketches a smoothing approach that could be used if condition (ref) is not plausible.
Section (ref) states the assumptions required for asymptotic theory. Section (ref) describes a key condition required for asymptotic theory and verifies it for single-index models. Section (ref) states the theoretical results.
Assumption (ref) ensures that the moment functions $\rho(W, t, \xi)$ are robust to first-order biases in the nuisance parameter $\xi$ uniformly over the index set $\mathcal{T}$.
The small bias property is often attained via orthogonalization, a technique that has been widely studied in, e.g. Newey1994, chernozhukov2016double, and LRSP. For example, if $\rho(W, t, \xi) = \rho(W,t)$ does not involve any nuisance parameters, Assumption (ref) is automatically satisfied.
Assumption (ref) is shown to be satisfied if the signal is orthogonal with respect to the nuisance function $\xi_0$ whose estimator converges at a $o(N^{-1/4})$ mean square rate (e.g., SemCher).
Assumption (ref) is a rate condition on the nuisance parameter $\nu$ entering (ref).
When the convergence is required in mean square ($\ell_2$) norm, Assumption (ref) is a classic assumption in the semiparametric literature (e.g., Newey1994). The example below demonstrates the plausibility of Assumption (ref) in $\ell_{\infty}$ norm.
Assumption (ref) is a standard regularity condition
. For example, for the case of Example (ref), the condition (ref) automatically holds with $B_{\phi}=1$ since $\phi (d, v)=-\mathbf{e}_{d} v$ corresponds to taking the $d$'th element of the nuisance parameter $\nu_0(x) = (m(d,x))_{d \in \mathcal{D}}$.
Assumption (ref) is a margin condition that ensures separation between minimizers and non-minimizers in order to control the first-order effect of classification mistakes. It is a standard assumption in classification literature MammenTsybakov,Tsybakov,QianMurphy, policy learning KitagawaTetenov,MbakopTabord and debiased inference kallus2022whats,SemJoE,SemSupp2. Section (ref) verifies Assumption (ref) for a special case of single-index models.
In this section, I demonstrate the plausibility of Assumption (ref) in the context of Example (ref).
Remark (ref) verifies that Assumption (ref) holds with $\bar{B} = 2 \bar{B}_f$ as long as the index set $\mathcal{T}$ has two elements and their difference $m(1,X)-m(0,X)$ has a bounded unconditional density. This is a known result in the literature Tsybakov,KitagawaTetenov,kallus2022whats.
Lemma (ref) describes a class of linear models where the presence of covariate vector $X$ with a smooth distribution suffices for Assumption (ref). Suppose the expectation functions are partially linear
where $\widetilde{X} = (X, \bar{X})$ and $ \bar{X} $ is independent of $X$.
Lemma (ref) verifies Assumption (ref) as long as the model (ref) includes a linear component obeying (ref). Condition (i) ensures that the mapping $x \mapsto \min_{t \in \mathcal{T}} (x^{\prime} \gamma_t)$ has a unique minimum. Condition (ii) accommodates the inclusion of arbitrary covariates (i.e., either continuous or discrete), provided their influence on the conditional mean remains almost surely bounded. Condition (iii) is a smoothness condition on $X$ similar to those in the analysis of least absolute deviation in Powell or the analysis of support function in CCMS. For example, if $X \sim N(h, \Sigma)$ is a Gaussian $p_X$-vector, the scalar $X^{\prime} \delta \sim N(\delta^{\prime} h, \delta^{\prime} \Sigma \delta)$, and the condition (ref) holds with $ \bar{B} = \sqrt{2 \pi \delta^{\prime} \Sigma \delta}^{-1} \leq (\sqrt{ 2 \pi})^{-1} \lambda^{-1/2}_{\min} (\Sigma)$.
Lemma (ref) extends the result of Lemma (ref) to nonlinear models. Suppose the expectation functions are
Lemma (ref) verifies Assumption (ref) for a single-index model with a monotone link function provided the linear index obeys (ref). Condition (iii) is satisfied for a large class of link functions, such as probit, logit or uniform $U[-t, t]$ where $\mathcal{X}$ is included in $[-t,t]$.
In conclusion, let me point out that Assumption (ref) can be viewed as a special case of a general form of the margin assumption
that is routinely imposed in standard debiased inference LuedtkeLaan,kallus2020assessing,SemSupp2. Relaxing this assumption remains an important open question in the literature. When this assumption is violated, the cross-fit plug-in estimators have a non-standard, heavy-tailed distribution which makes standard Wald-type inference not valid LuedtkeLaan,ponomarev2024.
Theorem (ref) is my first main result
. As a special case, it nests recent debiased inference results for sharp Makarov bounds with a binary outcome kallus2022whats and for BalkePearl1994,BalkePearl1997 bounds in the concurrent, independent work of levis2023covariateassisted. The paper's contribution is to introduce a general framework of aggregated intersection of regression functions for which the oracle property also applies. The new applications of the framework include Roy model bounds and Horowitz-Manski-Lee bounds with discrete outcomes, discussed in Section (ref).
Theorem (ref) establishes consistency of the plug-in estimator of variance which suffices for the validity of the confidence interval (ref). Unlike the envelope score estimator $\widehat \psi$, the variance estimator $\widehat V$ is first-order sensitive to the mistakes in estimated minimizers and converges at rate $\nu^{\infty}_N$ rather than $(\nu^{\infty}_N)^2$.
Theorem (ref) establishes the validity of multiplier bootstrap inference. Note that the bootstrap-based inference is possible because the kink point (the point of non-differentiability) occurs with probability zero. A similar validity argument is used to establish bootstrap inference for support function as in CCMS and SemJoE.
I begin by reviewing Roy model. Let $Y(1)$ and $Y(0)$ be two binary potential utility values, corresponding to the choices of treatment $D=1$ and $D=0$, respectively. An individual chooses the treatment value $D$ according to
If the potential outcomes $Y(1)$ and $Y(0)$ are equal, the choice is unspecified. The observed data $W=(X,D,Y)$ consist of the covariates $X$, the choice variable $D$, and the observed outcome $Y=DY(1) + (1-D) Y(0)$.
Proposition (ref) restates the Proposition 1 of MourifieHenry. It gives sharp bounds in Roy model with an instrument $Z$ obeying exclusion restriction. Such variables are akin to typical instrumental variables. The examples of discrete-valued $Z$ provided in MourifieHenry include parental education, distance to a college, and attendance a Catholic high school.
Assumption (ref) (A) is a relaxation of (ref) which requires that independence holds only conditional on covariates. Assumption (ref) (B) is equivalent to (ref).
Proposition (ref) refines Proposition (ref) by incorporating covariate information. The sharp bounds are derived by applying the argument in Proposition (ref), conditional on $X$, and then aggregating over the covariate space. Interchanging expectation and minimum gives another pair of bounds of the form (ref)--(ref). Since they do not involve any expectation functions and are simpler to estimate, we refer to them as basic (or no-covariate) bounds. If the propensity score is constant, these basic bounds coincide with the original bounds defined in Proposition (ref).
Let me demonstrate the proposed inferential methodology focusing on the first bound in (ref). The bound
is a special case of (ref) with $\mathcal{T} = \mathcal{Z}$, the nuisance vector-function $$ \nu_0(x) = (\Pr (D =1, Y=1 \mid Z=z, X=x))_{z \in \mathcal{Z}}, $$ and the projection functions $$ \phi(z, v) = v_z \quad z \in \mathcal{Z}. $$ Furthermore, it can be also mapped to the optimal welfare parameter $\psi_0$ in Example (ref) with
Following Example (ref), define the orthogonal score for each $\phi(z,\nu_0(x))$ as
where the true value $\xi_0$ of the nuisance parameter is $$ \xi_0(x) = (\nu_0(x), \mu_{z0}(x)), \qquad \mu_{z0}(x) = \Pr (Z=z \mid X=x), \quad z \in \mathcal{Z}. $$ Given the true functions $\nu_{z0}(\cdot), \mu_{z0}(\cdot)$ and sequences of shrinking neighborhoods $\mathcal{T}_N^z$ of $\nu_{z0}(\cdot) $ and $M_N^z$ of $\mu_{z0}(x)$, define the following rates:
Assumption (ref) states the regularity conditions. First, it requires the instrument $Z$ to be discrete-valued with finite support so that the propensity score
is bounded away from zero and one for each distinct instrument value. Continuously supported instrument (i.e., $\mathcal{T}=\mathcal{Z}$) is outside of the scope of this paper since their propensity score violates (ref). In this case, I conjecture that the propensity score needs to be approximated by a kernel density estimator, e.g. as is standard in the results for continuous treatments Colangelo. Second, the mean square rates of the first-stage estimators must decay sufficiently fast, a condition that is standard in semiparametric estimation literature.
Corollary (ref) gives an envelope score estimator for Roy model bounds and delivers uniformly valid debiased inference in the presence of covariates. As described in (ref), the sufficient conditions for the margin assumption discussed in Section (ref) are equally applicable to Roy model bounds.
I begin by introducing the sample selection problem. Let $D=1$ be an indicator for treatment receipt. Let $Y(1)$ and $Y(0)$ denote the potential outcomes if an individual is treated or not, respectively. Likewise, let $S(1)=1$ and $S(0)=1$ be dummies for whether an individual's outcome is observed with and without treatment, respectively. The data vector $W=(D,X,S,S \cdot Y)$ consists of the treatment status $D$, a baseline covariate vector $X$, the observed selection status $S=D \cdot S(1) + (1-D) \cdot S(0)$ and the observed outcome $S \cdot Y = S \cdot (D \cdot Y(1) + (1-D) \cdot Y(0))$ for selected individuals. LeeBound focuses on the average treatment effect (ATE)
for subjects who are selected into the sample regardless of treatment receipt---the always-takers.
Assumption (ref)(1) holds by random assignment. Assumption (ref)(2) states that all subjects must exhibit the same direction of selection response. It is frequently imposed in selection and treatment choice models. If covariates are available, this assumption has testable implications. Furthermore, it can be relaxed to conditional monotonicity Kolesar,SemSupp2. In this paper, we consider an unconditional version of the monotonicity assumption so as to focus on the theory for discrete-valued outcomes. As discussed in LeeBound, the average control outcome is point-identified $$ \mathbb{E} [ Y(0) \mid S(0)=1] = \mathbb{E} [ Y(0) \mid S(1)=1, S(0)=1] =\mathbb{E} [ Y \mid S=1, D=0].$$ I focus on the average treated outcome
In contrast to the control group, a treated outcome can be either an always-taker's outcome or a complier's outcome. The always-takers' share among the treated outcomes is
\paragraph{Binary outcomes. } Suppose $Y(1)$ and $Y(0)$ take values in $\{1, 0\}$. In the best case, the always-takers comprise the top $p_0$-quantile of the treated outcomes. Let $p_Y:= \Pr (Y =1 \mid D=1, S=1)$. In case when $p_0 <p_Y$, the best-case always-takers' outcome is equal to one for all always-takers. Otherwise, the best-case always-takers' outcome distribution is a mixture of ones and zeroes, with the mixing proportion of ones and zeroes equal to $p_Y/p_0$ and $1-p_Y/p_0$, respectively. In other words, the basic (i.e., no-covariate) upper bound $\bar{\beta}_U$ on $\mathbb{E}[ Y(1) \mid S(1) =1, S(0)=1]$ can be expressed as
Since $\bar{\beta}_U$ involves no covariates, we refer to it as the basic bound as opposed to the sharp bound derived further.
Lee's identification strategy can be implemented conditional on covariates. Denote the conditional trimming threshold $p_0(x)$ as
and the conditional probability of outcome one as $$ p_Y(x) = \Pr (Y = 1\mid D=1, S=1, X=x). $$ Finally, the conditional upper bound $\bar{\beta}_U(x)$ can be expressed as
Aggregating the conditional bound over the always-takers' covariate distribution gives the sharp upper bound
and a similar argument applies for the lower bound. Proposition (ref) derives the sharp lower and upper bounds for the average potential outcome.
Proposition (ref) derives basic and sharp bounds on the average potential outcome in a selection problem. The denominators of basic and sharp bounds are the same and equal the always-takers' share $$s_0 = \Pr [S=1 \mid D=0] = \mathbb{E} [ s_0(0,X) ].$$ Their numerators are regular and aggregated intersection bounds, described in the LHS and RHS of (ref)--(ref), respectively. The basic bounds (ref)--(ref) coincide (up to a constant) the bounds in the Lemma 1 of concurrent, independent work of kroft2024leeboundsmultilayeredsample.
\paragraph{Discrete outcomes. }In this section, I allow the outcome $Y$ to take a finite number of discrete values. Proposition (ref) characterizes the numerators of Horowitz-Manski-Lee bounds as a special case of aggregated intersection bounds.
Proposition (ref) develops a novel representation of Horowitz-Manski-Lee bounds as regular and aggregated intersection bounds, respectively. The outcome distribution is represented using point mass functions (PMFs) rather than quantiles, which is convenient for working with discrete outcomes. As shown in RockUryasev, the minimum in (ref) is attained by the outcome quantile of level $1-s_0/s_1$ or the “borderline” always-takers' outcome.
\paragraph{Debiased inference. } To describe an inferential approach, I derive moment functions for (ref) and (ref). Define the moment functions for the lower bound
and for the upper bound
Next, let $ \mathrm{T}$ be the finite support of the outcome $Y$. Define the nuisance parameter
The resulting moment function for $N_U$ is
Likewise, the moment function for $N_L$ is
Given the true functions $\xi_{0}(\cdot)$ and sequences of shrinking neighborhoods $S^d_N$ of $s_0(d,x)$ and $\mathcal{P}^{\beta}_N$ of $\pi_{\beta 0}(x)=\Pr (Y = \beta \mid D=1, S=1, X=x)$, define the following rates:
Assumption (ref) summarizes regularity conditions for Horowitz-Manski-Lee bounds. Since the propensity score is assumed known, the individual moment functions $\{ \rho_L (W, \beta), \rho_U (W, \beta)\}_{\{\beta \in \mathcal{T} \}}$ in (ref) and (ref) do not involve any nuisance parameters. As a result, Assumption (ref) is automatically satisfied for the individual functions.
Corollary (ref) delivers a root-$N$ consistent, asymptotically Gaussian estimator of sharp Horowitz-Manski-Lee bounds assuming the conditional probability of selection and the conditional PMF are estimated at a sufficiently fast rate. To the best of my knowledge, this is a first example of debiased inference for the trimming bounds with discrete-valued outcome.
In conclusion, I state the Algorithm (ref) for computing the bounds as well as examples of the first-stage estimators.
This section provides numerical evidence for the methods developed in this article. Section (ref) offers a Monte Carlo experiment constructed in the context of Example (ref). Section (ref) offers an empirical illustration of the method for Horowitz-Manski-Lee bounds in Section (ref).
I build a simulation exercise on JTPA dataset Bloom1997 that consists of three elements: baseline covariates, treatment (access to job training), and outcome. The baseline covariate $X_1$ is taken to be the previous earnings PreEarn measured in $10, 000$ USD. The covariate vector $X=(X_1, \dots, X_1^p)$ includes the first $p$ powers of the PreEarn variable. The treatment $D$ is determined by a coin flip with probability $Pr(D = 1) = \frac{2}{3}$, to match the propensity score in JTPA data. The outcome $Y$ follows a linear model
where $\epsilon \sim N(0, \sigma^2)$ is a Gaussian shock independent of the data. The true parameter values are $$\kappa_0 = \gamma_0 =(2^{-1}, 2^{-2}, \dots, 2^{-p}), \quad \sigma^2 = 1.$$ The population data set size is $9, 223$. In addition to this primary design, we also consider an artificial (Gaussian) design where $X_1$ is drawn from a Gaussian distribution whose mean and variance matches the respective parameters of actual PreEarn variable, with all other steps being the same. Using this setup, we evaluate coverage of plug-in and bootstrap inferential confidence intervals based on the doubly robust estimator described in Example (ref). The first-stage functions $m(0, X)$ and $m(1, X)$ are estimated via linear least squares. The performance metrics include bias, mean squared error (MSE), and coverage rates of the confidence intervals (CIs). These metrics are analyzed across varying sample sizes $N \in \{100, 200, 300, 500\}$ and polynomial degrees $p \in \{1, 3, 5, 7\}$.
Tables (ref) and (ref) summarize the simulation results, highlighting key differences across polynomial degrees ($p$) and designs. For lower degrees ($p \in \{1, 3\}$), both designs yield estimators with low bias, low MSE, and near-nominal coverage rates for both plug-in and bootstrap CIs. The results are consistent with the theoretical results in Theorems (ref) and (ref). For higher degrees ($p \in \{5, 7\}$), the performance diverges significantly between the two designs. In the Gaussian design (Table (ref)), coverage remains close to the nominal rate even for $p=7$ when $N=500$. Conversely, in the primary design (Table (ref)), coverage drops sharply to $48\%$ (Plug-In CI) and $39\%$ (Bootstrap CI) for $p=7$. This discrepancy likely arises from the heavy-tailed nature of PreEarn, which could either affect the quality of the first-stage estimates of $\kappa_0$ and $\gamma_0$, make Assumption (ref) to be a poor fit for the data, or both.
To illustrate the immediate applicability of the proposed method, this study analyzes the effect of Medicaid exposure on healthcare utilization and health outcomes using data from the Oregon Health Insurance Experiment finkelstein. In 2008, Oregon implemented a limited expansion of its Medicaid program, providing insurance coverage to low-income, uninsured adults selected through a lottery system from a waiting list. One year after randomization, a subset of $N = 58,405$ applicants was mailed a survey to assess changes in healthcare utilization and general well-being, with a response rate of approximately $50\%$. Abstracting from potential non-response bias, the study found that Medicaid significantly improved healthcare access, financial security, and mental health outcomes for low-income adults.\footnote{finkelstein reported that the ability to reject the null hypothesis of no effect of health insurance on healthcare utilization or financial strain is generally robust to Lee bounds, while the ability to reject the null hypothesis of no effect on self-reported health outcomes is not robust (see footnote 19).} To examine the robustness of these findings, this section reports various versions of Horowitz-Manski-Lee bounds under various assumptions about subjects' response behavior.
This section focuses on the average treatment effect (ATE) of Medicaid exposure on self-reported mental health. Let $S(1) = 1$ and $S(0) = 1$ be binary indicators denoting whether a subject completes a survey when treated or not treated, respectively, and let $Y(1)$ and $Y(0)$ represent the corresponding potential outcomes. The observed data, $W = (X, D, S, S \cdot Y)$, include $D$ (lottery outcome), $S$ (an indicator of non-missing response), $Y$ (the response itself), and baseline covariates $X$. The propensity score, $\mu(X)$, is determined by household size and survey wave fixed effects.\footnote{If an applicant wins the lottery, all members of their household become eligible to enroll. Consequently, larger households have a higher probability of winning the lottery compared to smaller ones. Additionally, control applicants were oversampled in earlier survey waves, further influencing the propensity score.} Other components of $X$ include 64 predetermined characteristics, such as demographics, enrollment in the Supplemental Nutrition Assistance Program (SNAP) or Temporary Assistance for Needy Families (TANF) as well as the total amount of benefits received in each program, and pre-existing health conditions.
Table (ref) summarizes the findings. The baseline estimate of Medicaid's effect on mental health is $2.27\%$. Interestingly, the control group's response rate ($49.4\%$) exceeds that of the treated group ($48.2\%$). Unconditional monotonicity (Assumption (ref)), which posits that treatment discourages survey completion, lacks an intuitive explanation. Standard Lee bounds (Column (1)) are misleadingly tight. Without monotonicity, the proportion of “always-takers” (respondents regardless of treatment status) cannot be bounded away from zero, and no-monotonicity bounds (Column (5)) do not provide any meaningful restriction on the treatment effect
. This pattern holds for all survey outcomes, including questions about healthcare utilization, financial strain, and self-reported health outcomes.
To make progress, the analysis focuses on a subset of the population for whom the direction of selection response aligns with the majority. Specifically, the parameter of interest is:
where the outcome of interest is mental health as measured by a positive response to the question “Did you not screen positive for depression in the last two weeks?”). The selection equation (ref) is estimated using logistic regression (see Example (ref)). The estimated share of subjects with negative selection response is $82\%$ (Column (2)), consistent with the negative direction of unconditional response effect. A higher likelihood of survey response is associated with being female, requesting English-language materials, and not receiving TANF benefits. Additionally, Medicaid's effect on response appears negative for individuals who experienced injuries or received SNAP benefits prior to randomization, suggesting that control participants' response could be driven by acute health or financial challenges.
Basic bounds on (ref) (Table (ref), Column (3)), derived under conditional monotonicity, cannot determine the direction of the treatment effect. Sharp bounds (Table (ref), Column (4)), constructed using Algorithm (ref) with first-stage outcome fitted values as described in Examples (ref) and (ref), suggest that the Medicaid exposure effect is positive, though the magnitude is attenuated at the lower bound. The proposed approach relies on a smoothness assumption, specifically that the conditional selection and outcome probabilities are sufficiently continuously distributed, which could be plausible since some components of $X$ are continuously distributed, such as total amount of SNAP or TANF benefits or pre-existing ED charges.