Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Beyond the Average: Distributional Causal Inference under Imperfect Compliance
abstractWe study the estimation of distributional treatment effects in randomized experiments with imperfect compliance. When participants do not adhere to their assigned treatments, we leverage treatment assignment as an instrumental variable to identify the local distributional treatment effect—the difference in outcome distributions between treatment and control groups for the subpopulation of compliers. We propose a regression-adjusted estimator based on a distribution regression framework with Neyman-orthogonal moment conditions, enabling robustness and flexibility with high-dimensional covariates. Our approach accommodates continuous, discrete, and mixed discrete-continuous outcomes, and applies under a broad class of covariate-adaptive randomization schemes, including stratified block designs and simple random sampling. We derive the estimator’s asymptotic distribution and show that it achieves the semiparametric efficiency bound. Simulation results demonstrate favorable finite-sample performance, and we demonstrate the method’s practical relevance in an application to the Oregon Health Insurance Experiment.
Introduction
Randomized experiments are a cornerstone of causal inference, widely employed in both academic research duflo2007using and industry settings kohavi2020trustworthy. In practice, however, subjects often deviate from their assigned treatments, leading to imperfect compliance. When compliance is not guaranteed, estimating the causal effect for the entire population is generally not possible, without imposing additional assumptions. However, a standard approach to address this issue is to use the random assignment as an instrumental variable (IV). This strategy allows for identification of the causal effect of treatment for the subset of individuals who comply with their assignment—known as the local average treatment effect (LATE) imbens1994—without requiring assumptions about how individuals self-select into treatment.
To improve covariate balance between treatment and control groups, researchers often use covariate-adaptive randomization (CAR), which stratifies individuals based on key covariates before assigning treatments within each stratum. The CAR framework includes various designs, such as stratified block randomization and Efron’s biased coin design imbens2015causal, with simple random sampling as a special case.
While much of the literature focuses on estimating the average effects, this summary measure can obscure important heterogeneity in treatment responses. In this paper, we study the estimation of distributional treatment effects in randomized experiments with covariate-adaptive randomization and noncompliance, focusing on the local distributional treatment effect (LDTE)—defined as the difference in counterfactual outcome distributions for compliers across treatment arms. By examining the entire distribution of outcomes, rather than just the mean, we aim to provide a more nuanced understanding of how treatments affect different segments of the population.
We propose a regression-adjusted estimator for LDTEs that leverages auxiliary covariates beyond stratum indicators to improve efficiency. Our setup accommodates heterogeneous assignment probabilities and heterogeneous treatment effects. Estimation proceeds via a distribution regression framework combined with Neyman-orthogonal moment conditions chernozhukov2018debiased, chernozhukov2022locally, which provide robustness to first-order estimation errors in high-dimensional or complex nuisance components. These nuisance functions—conditional distribution functions given pre-treatment covariates—are estimated using flexible machine learning methods, including random forests, neural networks, and gradient boosting. Incorporating cross-fitting further strengthens robustness against estimation errors.
Despite the growing body of work on CAR and noncompliance in experimental settings, methods that estimate distributional treatment effects in the presence of both CAR and noncompliance remain scarce. For instance, jiang2023regression address quantile treatment effects under full compliance, and jiang2024improving study average treatment effects under CAR with imperfect compliance. However, to our knowledge, there are no existing methods that integrate regression adjustment and IV techniques for estimating full outcome distributions under CAR and noncompliance. This paper addresses that gap and makes the following contributions:
enumerate• We develop a regression-adjusted estimator for distributional treatment effects under CAR with noncompliance, applicable to continuous, discrete, and mixed discrete-continuous outcomes.
• We derive the asymptotic distribution of the estimator under CAR, generalizing beyond the traditional i.i.d. framework in causal inference.
• We establish the semiparametric efficiency bound for the LDTE under CAR and show that our estimator attains this bound.
• We validate our approach through simulation studies and an empirical application to the Oregon Health Insurance Experiment, where only 58% of subjects complied with their treatment assignment.
The remainder of the paper is structured as follows. Section (ref) reviews related literature. Section (ref) describes the problem setup and identification strategy. Section (ref) introduces the proposed estimation method. Section (ref) presents the asymptotic properties of our estimator. Section (ref) reports simulation and empirical results. Section (ref) concludes. The Appendix includes notation, technical proofs, and additional experimental details.
Related Literature
\paragraph{Distributional treatment effects}
Distributional and quantile treatment effects provide a more comprehensive view of treatment impacts beyond average effects. The concept of QTE was first introduced by doksum1974empirical and lehmann1975nonparametrics, and has since inspired a broad literature developing estimation and inference methods for distributional effects across econometrics, statistics, and machine learning. Notable contributions include heckman1997making, imbens1997estimating, koenker2005quantile, bitler2006mean, athey2006identification, firpo2007efficient, chernozhukov2013inference, koenker2017handbook, belloni2017program,
callaway2018quantile, callaway2019quantile, chernozhukov2019generic, ge2020conditional, park2021conditional, zhou2022estimating, gunsilius2023distributional, kallus2023robust, among others. Most of this work focuses on conditional distributional and quantile treatment effects. In contrast, oka2025regression, byambadalai24a, and hirata2025efficient examine unconditional distributional effects, though their analyses are restricted to settings with simple random sampling and full compliance. byambadalai25a also examine unconditional distributional effects under covariate-adaptive randomization, but their framework likewise assumes full compliance.
\paragraph{Instrumental variables estimation of distributional causal effects}
Instrumental variables have a long-standing role in identifying causal effects in the presence of confounding, either by relying on additional structural assumptions haavelmo1943statistical, angrist1996identification or by enabling partial identification under weaker conditions manski1990nonparametric, balke1997bounds. A key development in the estimation of distributional effects is the instrumental variable quantile regression (IVQR) framework, which estimates quantile functions across the outcome distribution under the rank similarity assumption chernozhukov2004effects, chernozhukov2005iv, chernozhukov2006instrumental, kaido2021decentralization. An alternative approach by abadie2002instrumental focuses on local QTEs for the complier subpopulation, under the monotonicity assumption—a setting also considered in our work. frolich2013unconditional similarly estimate unconditional QTEs under endogeneity, assuming monotonicity. wuthrich2020comparison provide a detailed comparison between IVQR and local QTE models. Additionally, abadie2002bootstrap introduce a Kolmogorov–Smirnov-type test for comparing complier outcome distributions in randomized experiments. Other contributions addressing distributional and quantile causal effects using IV methods under assumptions different from ours include chernozhukov2007instrumental, horowitz2007nonparametric, briseno2020flexible, kook2024instrumental, kallus2024localized, chernozhukov2024estimating, among others.
\paragraph{Regression adjustment under covariate-adaptive randomization}
Regression adjustment using pre-treatment covariates to improve precision in average treatment effect (ATE) estimation has been extensively studied under simple random sampling fisher1932statistical, cochran1977sampling, yang2001efficiency, rosenbaum2002covariance, freedman2008regression, freedman2008regression2, tsiatis2008covariate, rosenblum2010simple, lin2013agnostic, berk2013covariance, ding2019decomposing. Recent work extends this to covariate-adaptive randomization. cytrynbaum2024covariate derive optimal linear adjustments for stratified designs, and rafi2023efficient characterize the semiparametric efficiency bound for ATE estimation. Other contributions include covariate adjustment in matched-pair designs bai2024covariate, general form of adjustment in biostatistics bannick2023general, tu2023unified, and methods for parameters defined by estimating equations wang2023model. While most of these focus on ATEs under full compliance, jiang2023regression study regression adjustment for the QTE, and jiang2024improving extend these ideas to the local ATE with imperfect compliance. Our work builds on this rich literature by targeting distributional causal effects under covariate-adaptive randomization and noncompliance.
\paragraph{Semiparametric estimation}
Our work builds on the semiparametric estimation literature, which focuses on estimating low-dimensional parameters in the presence of possibly infinite-dimensional nuisance components. Foundational contributions include robinson1988root, bickel1993efficient, newey1994asymptotic, robins1995semiparametric, with more recent developments in high-dimensional and machine learning settings by chernozhukov2018debiased, ichimura2022influence, among others. We formulate our estimation problem using Neyman-orthogonal moment conditions neyman1959optimal, chernozhukov2022locally, which provide robustness to errors in the estimation of nuisance components.
Setup and Notation
We consider a randomized experiment with binary treatment employing covariate-adaptive randomization, where imperfect compliance creates a discrepancy between treatment assignment and actual treatment receipt. Let $Y$ denote the observed outcome of interest, $Z\in\{0,1\}$ the random assignment, and $D\in\{0,1\}$ the actual treatment received. Within the potential outcome framework rubin1974estimating, imbens2015causal, we define $Y(1)$ and $Y(0)$ as potential outcomes under treatment status $D=1$ and $D=0$, respectively. Similarly, $D(1)$ and $D(0)$ represent potential treatment statuses under assignment $Z=1$ and $Z=0$. In this setup, random assignment $Z$ serves as an instrumental variable affecting treatment $D$, which subsequently influences outcome $Y$. The exclusion restriction holds, as instrument $Z$ affects outcome $Y$ only through treatment $D$. Hence, we can write the observed outcome and treatment as
align[align omitted — 115 chars of source]
Furthermore, we consider a covariate-adaptive randomization (CAR) setup in which each participant is assigned to a stratum
$S \in \mathcal{S} := \{1, \dots, S\}$, with additional covariates
$X\in\mathcal X \subset \mathbb R^{d_x}$
available.
Strata are typically constructed based on certain baseline covariates, and we allow $S$ and $X$ be dependent. We let $\pi_z(s):=P(Z=z\mid S=s)\in(0,1)$ be the target assignment probability for treatment $z\in\{0,1\}$ in stratum $s$ and let $p(s):=P(S=s)>0$ be the stratum size. Figure (ref) depicts the relationship between the variables.
figure[figure omitted — 1,543 chars of source]
We observe a data $\{(Y_i, D_i, Z_i, S_i, X_i)\}_{i=1}^{n}$ with a sample size of $n$. For each stratum $s\in \mathcal{S}$,
let $n(s):= \sum_{i=1}^{n} \mathbbm{1}_{\{S_i=s\}}$
denote the number of observations in stratum $s$,
and
$n_z(s):=\sum_{i=1}^{n} \mathbbm{1}_{\{Z_{i}=z, S_i=s\}}$
represent the number of observations receiving assignment $z \in \{0,1\}$
in stratum $s$. Here, $\mathbbm{1}_{\{\cdot\}}$ denotes the indicator function, which equals 1 if the condition inside is true and 0 otherwise.
Then, define the following empirical estimates: $\widehat \pi_z(s) := n_z(s)/n(s)$ the estimated target assignment and $\widehat p(s) := n(s)/n$ the proportion of observations falling in stratum $s$. We impose the following assumptions on the data generating process and the treatment assignment mechanism.
assumption[Data generating process and treatment assignment]
We have
(i) $\big \{\big (Y_i(0), Y_i(1), D_i(0), D_i(1), S_i, X_i\big)\big\}_{i=1}^{n}$ are independent and identically distributed
(ii) $\big\{\big(Y_i(0), Y_i(1), D_i(0), D_i(1), X_i\big)\big\}_{i=1}^{n} \rotatebox[origin=c]{90}{$\models$} \{Z_i\}_{i=1}^{n} \mid \{S_i\}_{i=1}^{n}$,
(iii) $\widehat{\pi}_z(s) = \pi_z(s) + o_p(1)$ for every $s\in\mathcal{S}$ and $z\in\{0,1\}$.
(iv) $\mathbb{P}\big (D_i(1) \geq D_i(0) \big )=1$.
Assumption (ref) (i) allows for cross-sectional dependence among treatment statuses $\{Z_i\}_{i=1}^{n}$, thereby accomodating many covariate-adaptive randomization schemes. Assumption (ref) (ii) states that the assignment is independent of potential outcomes, potential treatment choices and pre-treatment covariates conditional on strata. Assumption (ref) (iii) states the assignment probabilities converge to the target assignment probabilities as sample size increases.
Common randomization schemes satisfying Assumption (ref) (i) to (iii) include simple random sampling, stratified block randomization, biased-coin design efron1971forcing, and adaptive biased-coin design wei1978adaptive. Assumption (ref) (iv) says that there are no defiers in the population. This assumption is also called the monotonicity assumption in the literature, and is the key assumption that allows for the identification of the causal effect within a specific subpopulation, known as compliers.
To clarify this, we introduce the four treatment compliance types as defined by angrist1996identification. Never-takers consistently avoid the treatment, with \(D(1) = 0\) and \(D(0) = 0\). Defiers exhibit behavior opposite to the intended assignment, receiving the treatment when not encouraged (\(D(0) = 1\)) and avoiding it when encouraged (\(D(1) = 0\)). Compliers follow the assigned treatment status, such that \(D(1) = 1\) and \(D(0) = 0\). Always-takers are individuals who receive the treatment regardless of the instrument assignment, i.e., \(D(1) = 1\) and \(D(0) = 1\). Note that these types are not directly observable by the researcher.
We are interested in the distributional effects of receiving the treatment. To that end, let the distribution function of potential outcomes be denoted by
align[align omitted — 123 chars of source]
Analogous to the local average treatment effect (LATE) of imbens1994, we define the local distributional treatment effect (LDTE) as the difference in the distribution functions of the potential outcomes among compliers:
align[align omitted — 129 chars of source]
for $y \in \mathcal{Y}$. Here, compliers (i.e., those with $D(1)>D(0)$) refer to individuals who receive the treatment if and only if they are assigned to it. The following lemma demonstrates that, under Assumption (ref), a random assignment can be used to identify the distributional causal effect of receiving the treatment for this subgroup.
lemma[Local distributional treatment effect]
Suppose Assumptions (ref) holds.
Then, the local distributional treatment effect can be expressed as, for $y\in\mathcal{Y}$,
\begin{align}
& \beta(y) = \frac{\sum _{s=1}^{S}p(s)\cdot(\mathbb{E}[\mathbbm{1}_{\{Y \leq y\}} \mid Z=1, S=s]- \mathbb{E}[\mathbbm{1}_{\{Y \leq y\}} \mid Z=0, S=s])}{\sum _{s=1}^{S}p(s) \cdot(\mathbb{E}[ D \mid Z=1, S=s] - \mathbb{E}[ D \mid Z=0, S=s])}.
\end{align}
Our formulation in (ref) builds upon and extends the approach of abadie2002bootstrap to accommodate covariate-adaptive randomization through stratum-specific weights.
Both the numerator and the denominator are written as weighted averages across strata indexed by $s$, with weights given by the distribution $p(s)$.
The numerator in (ref) can be interpreted as the intent-to-treat (ITT) distributional effect—that is, the difference in the distribution functions of the outcome $Y$ between treatment and control groups defined by the random assignment $Z$. Importantly, this reflects the effect of being assigned to treatment, not of actually receiving treatment. The denominator in (ref) represents the first stage of the instrumental variable approach. It captures the effect of the assignment $Z$ on the probability of receiving the treatment $D$, conditional on stratum $S = s$, and then averages this across strata. The first stage quantifies the degree of compliance with the assignment and ensures that the instrument is relevant (i.e., affects treatment uptake). A non-zero first stage is necessary for the IV estimator to be well-defined and to identify the treatment effect for compliers. Thus, the LDTE is obtained by scaling the ITT distributional effect by the strength of the first stage. Notably, the denominator is constant in $y$, so the variation in $\beta(y)$ across values of $y \in \mathcal{Y}$ reflects changes in the distribution of outcomes, not in the compliance rate.
Lastly, we also define the local probability treatment effect (LPTE)
align*[align* omitted — 157 chars of source]
for each $j = 1, \dots, J$, where
$\mathcal{Y}_J:=\{y_1, \cdots, y_J\} \subset \mathcal{Y}$ and $y_0=-\infty$.
The LPTE measures treatment-induced changes in the probability mass of the outcome distribution within each interval $(y_{j-1}, y_j]$, effectively comparing the “histograms” of potential outcomes for compliers. The theoretical results developed for the LDTE extend directly to the LPTE by substituting the indicator functions $\mathbbm{1}_{\{Y(d) \leq y_j\}}$ with $\mathbbm{1}_{\{y_{j-1} < Y(d) \leq y_j\}}$ for $d\in\{0,1\}$ in all relevant expressions.
Estimation
We propose a regression-adjusted LDTE estimator for $\{\beta(y)\}_{y\in\mathcal{Y}}$ incorporating the additional covariates $X_i$. For notational convenience, we define the following terms. The conditional probability of treatment given the instrument, stratum, and covariates:
align[align omitted — 73 chars of source]
The conditional distribution function of \( Y \) given the instrument, stratum, and covariates:
align[align omitted — 126 chars of source]
The estimators for these quantities are denoted by \( \widehat{\mu}_{z}(y, s, x) \) and
\( \widehat{\eta}_{z}(s, x) \), respectively. Since \( X_i \) may be a continuous variable, the estimation of \( \widehat{\mu}_{z}(y, s, x) \) and \( \widehat{\eta}_{z}(s, x) \) relies on nonparametric methods, such as logistic regression, random forests, and other flexible machine learning (ML) approaches. In covariate-adaptive randomized experiments, the target assignment probability for treatment $z\in\{0,1\}$ for a given stratum \(s\), denoted by \(\pi_z(s)\), is typically known in advance or can be consistently estimated using its sample analog, defined as \(\widehat{\pi}_z(s) = n_z(s)/n(s)\). Then, our proposed estimator for the LDTE for $y\in\mathcal{Y}$ is given by
align[align omitted — 180 chars of source]
where
align[align omitted — 408 chars of source]
The estimator presented in (ref) follows the structure of the well-known augmented inverse propensity weighting (AIPW) estimator, which relies on a doubly robust moment condition robins1994estimation, robins1995semiparametric. This moment condition satisfies the Neyman orthogonality property chernozhukov2018debiased, chernozhukov2022locally, ensuring that the estimator is first-order insensitive to the estimation errors of the nuisance functions $(\mu_{z}(\cdot), \eta_{z}(\cdot))$. To further improve robustness, we incorporate cross-fitting with $L$ folds $(L>1)$ as proposed by chernozhukov2018debiased. The complete estimation procedure is detailed in Algorithm (ref). Setting the adjustment terms $\widehat{\mu}_z(\cdot)$ and $\widehat{\eta}_z(\cdot)$ to zero yields the empirical (unadjusted) estimator for the LDTE, obtained by replacing each component in (ref) with its sample analog.
algorithm[algorithm omitted — 1,014 chars of source]
Asymptotic Properties
In this section, we derive the asymptotic distribution of our proposed estimator, which enables statistical inference and the construction of confidence intervals. Additionally, we establish the semiparametric efficiency bound for the LDTE and demonstrate that the regression-adjusted estimator achieves this bound under the specified assumptions. We begin by introducing some additional notation to formalize our results. Let $\ell^{\infty}(\mathcal Y)$ be the space of uniformly bounded functions mapping an arbitrary index set $\mathcal{Y}$ to the real line.
assumptionWe have
(i) For $z\in\{0,1\}$ and $s\in\mathcal{S}$, define $I_z(s) := \{i\in [n]: Z_i = z, S_i=s\}$, $\delta^Y_{z}(y,s,X_i) :=
\widehat{\mu}_{z}(y, s,X_i) - \mu_{z}(y, s,X_i)$, and $\delta^D_{z}(s,X_i) :=
\widehat{\eta}_{z}(s,X_i)- \eta_{z}(s,X_i)$. Then, for $z\in\{0,1\}$, we have
\begin{align}
\sup_{y \in \mathcal{Y},s\in \mathcal{S}}&\biggl|\frac{\sum_{i\in I_1(s)}\delta^Y_z(y,s,X_i)}{n_1(s)} - \frac{\sum_{i \in I_{0}(s)}\delta^Y_{z}(y,s,X_i)}{n_{0}(s)}\biggr| = o_p(n^{-1/2}),
\end{align}
\begin{align}
\max_{s\in \mathcal{S}}&\biggl|\frac{\sum_{i\in I_1(s)}\delta^D_z(s,X_i)}{n_1(s)} - \frac{\sum_{i \in I_{0}(s)}\delta^{D}_z(s,X_i)}{n_{0}(s)}\biggr| = o_p(n^{-1/2}).
\end{align}
(ii) For $z\in\{0,1\}$, let $\mathcal{F}_z = \{\mu_{z}(y, s,x): y\in \mathcal{Y} \}$ with an envelope $F_{z}(s,x)$. Then, $\max_{s \in \mathcal{S}}\mathbb{E}[|F_{z}(S_i,X_i)|^q|S_i=s]<\infty$ for $q > 2$ and there exist fixed constants $(\alpha,v)>0$ such that
\begin{align}
\sup_Q N\left(\varepsilon||F_z||_{Q,2}, \mathcal{F}_z, L_2(Q) \right) \leq \left(\frac{\alpha}{\varepsilon}\right)^{v}, \quad \forall \varepsilon \in (0,1],
\end{align}
where $N(\cdot)$ denotes the covering number and the supremum is taken over all finitely discrete probability measures $Q$.
Assumption (ref)(i) provides a high-level condition on the estimation of $\widehat{\mu}_{z}(y, s,X_i)$ and $\widehat \eta_{z}(s,X_i)$. Assumptions (ref)(ii) impose mild regularity condition on
$\mu_{z}(y, s,X_i)$. Specifically, it holds automatically when $\mathcal{Y}$ is a finite set. We now present the weak convergence of our proposed estimator in the following theorem, which provides the theoretical foundation for conducting statistical inference. This asymptotic result enables the construction of confidence intervals using either sample-based estimates of the asymptotic variance or bootstrap methods. Further details on the inference procedure are provided in Appendix (ref).
We define
$Y(D(z)):= D(z) \cdot Y(1) + \big (1-D(z)\big) \cdot Y(0)$. With this notation, the observed outcome $Y$ can be expressed as $Y =Z \cdot Y\big(D(1) \big) + (1 - Z) \cdot Y\big(D(0)\big)$. For $z\in\{0,1\}$, let $Y^z_{i}(y) := \mathbbm{1}_{\{ Y_{i}(D_{i}(z)) \leq y \}}$ and $\tilde{Y}^z_{i}(y) :=Y^z_{i}(y)-\mathbb{E}[Y^z_{i}(y)|S_{i}]$. Also, let $\tilde{D}_{i}(z) :=D_{i}(z)-\mathbb{E}[D_{i}(z)|S_{i}]$,
$\tilde{\mu}_{z}(y, S_{i}, X_{i}) :=\mu_{z}(y, S_{i}, X_{i})-\mathbb{E}[\mu_{z}(y, S_i, X_i)|S_i]$ and $\tilde{\eta}_{z}(S_{i}, X_{i}) :=\eta_{z}(S_{i}, X_{i})-\mathbb{E}[\eta_{z}(S_i, X_i)|S_i]$ for $z\in\{0,1\}$.
Then, we define
align[align omitted — 388 chars of source]
and
align[align omitted — 191 chars of source]
theorem[Asymptotic Distribution]
Suppose Assumptions (ref) and (ref) hold. Then, in $\ell^{\infty}(\mathcal Y)$, uniformly over $y\in\mathcal Y$, the regression-adjusted estimator defined in Algorithm (ref) satisfies
\begin{align}
\sqrt{n} \big
(\widehat \beta(y) - \beta(y) \big ) \rightsquigarrow \mathcal{G}(y),
\end{align}
where $\mathcal{G}(y)$ is a Gaussian process with covariance kernel
\begin{align}
& \Omega(y, y') := \frac{\Omega_{0}(y, y') +\Omega_{1}(y, y') + \Omega_{2}(y, y')}{\mathbb{E}[D(1)-D(0)]^2},
\end{align}
with
$
\Omega_{z}(y, y') :=
\mathbb{E}[\pi_z(S_i)\phi_i(y, z)\phi_{i}(y', z)]$ for $z\in\{0,1\}$
and
$\Omega_{2}(y, y') :=
\mathbb{E}[\xi_{i}(y)\xi_{i}(y')].$
We next derive the semiparametric efficiency bound of the LDTE and show our estimator achieves this bound in the following theorem. This implies that the asymptotic variance of any regular, root-$n$ consistent, and asymptotically normal estimator of the LDTE cannot be lower than this bound.
theorem[Semiparametric Efficiency Bound]
Under Assumption (ref), for every $y\in\mathcal{Y}$,
\begin{itemize}
• the semiparametric efficiency bound for $\beta(y)$ is $\Omega(y)$, which is defined by
\begin{align}
& \Omega(y) := \frac{\Omega_{0}(y, y) +\Omega_{1}(y, y) + \Omega_{2}(y, y)}{\mathbb{E}[D(1)-D(0)]^2},
\end{align}
where $\Omega_{0}(\cdot)$, $\Omega_{1}(\cdot)$ and $\Omega_{2}(\cdot)$ are defined in Theorem (ref).
• furthermore if Assumption (ref) also holds, then the regression-adjusted estimator $\widehat\beta(y)$
attains the semiparametric efficiency bound.
\end{itemize}
As a corollary to the theorem above, the asymptotic variance of the regression-adjusted estimator with known nuisance functions is lower than that of the empirical (unadjusted) estimator, in which the adjustment terms are set to zero.
Experiments
Simulation Study
figure[figure omitted — 705 chars of source]
We assess the finite-sample performance of our estimator through a simulation study designed to reflect a complex, nonlinear data-generating process with high-dimensional covariates and treatment effect heterogeneity.
The data generating process consists of four strata ($S = 4$) constructed by partitioning the support of a covariate $W_i \sim U(0,1)$ into $S$ equal-length intervals, where $S_i$ indicates the interval containing $W_i$. For each unit $i$, we draw an additional 20-dimensional covariate vector $X_i = (X_{1,i}, \dots, X_{20,i})^\top$ from a multivariate normal distribution $\mathcal N(0, I_{20\times20})$. The treatment indicator $Z_i$ follows a Bernoulli distribution with probability 0.5 within each stratum, maintaining a constant target proportion of treated units ($Z_i = 1$) across strata with $\pi_1(s) = 0.5$ for all $s \in \mathcal{S}$.
The complete specification of the data-generating process is given by:
align[align omitted — 318 chars of source]
where
$(a_1, a_0, b_1, b_0, c_1, c_0) = (2,1,1,-1,3,3)$,
and error term $\epsilon_i \sim \mathcal N(0,1)$
with
align[align omitted — 152 chars of source]
This design incorporates nonlinear dependencies, integrates deliberately irrelevant covariates, and preserves the monotonicity assumption by eliminating the possibility of defiers.
We draw a sample of sizes $\{500, 1000, 5000\}$ from the data-generating
process and estimate the LDTE at quantiles $\{0.1, . . . , 0.9\}$ using three methods with 1000 simulations: an unadjusted estimator, a linear regression-adjusted estimator, and a machine learning-adjusted estimator based on gradient boosting. A reference sample of size $10^6$ is used to approximate ground-truth LDTE values. All adjusted estimators use 2-fold cross-fitting.
Figure (ref) reports RMSE, average length and coverage of 95% confidence interval (CI) based on sample estimates. Both adjusted estimators achieve lower RMSE and shorter CIs than the unadjusted estimator. The unadjusted estimator achieves nominal 95% coverage for most quantiles, while ML adjustment exhibits slight over-coverage (up to 0.98–1.00), suggesting conservative intervals that could be tightened with improved nuisance estimation. Figure (ref) shows RMSE reduction (%) relative to the unadjusted estimator. Linear adjustment yields modest gains (1–10%), while ML adjustment achieves up to 50% reduction for some quantiles, with performance improving as sample size increases. These findings highlight the value of flexible regression adjustment in improving finite-sample efficiency for distributional causal effect estimation.
figure[figure omitted — 444 chars of source]
Real Data Analysis: Oregon Health Insurance Experiment
This subsection analyzes the impact of insurance coverage on emergency department (ED) visits using data from the Oregon Health Insurance Experiment.\footnote{The dataset is publicly available at \href{https://www.nber.org/research/data/oregon-health-insurance-experiment-data}{https://www.nber.org/research/data/oregon-health-insurance-experiment-data}.} We replicate the analysis in finkelstein2016effect and estimate distributional treatment effects. In 2008, the state of Oregon conducted a lottery to allocate health insurance to a group of uninsured low-income adults. Treatment assignment in this experiment was randomized based on household size, making the number of household members a stratification variable. However, due to imperfect compliance, not all individuals offered coverage enrolled, while some who were not selected obtained insurance through other means. Table (ref) displays the sample breakdown by assigned and realized treatments, and only 58% of the subjects comply with their random assignment. For a detailed discussion of the experiment and average treatment effect estimates of insurance coverage on various other outcomes, see finkelstein2012oregon.
table[table omitted — 510 chars of source]
figure[figure omitted — 650 chars of source]
Figure (ref) displays the distributional and probability treatment effect of insurance coverage on ED visits. We compute the LDTE and Local Probability Treatment Effect (LPTE) for $y\in\{0, 1, \dots, 15\}$ accounting for the stratified design and imperfect compliance. For regression adjustment, we use gradient boosting with 5-fold cross-fitting, with 28 pre-treatment covariates ($X_i$) including various variables regarding past emergency department visits. The full list of covariates can be found in the Appendix.
The top-left panel of Figure (ref) displays the empirical LDTE, while the top-right panel presents the regression-adjusted LDTE. Shaded areas represent 95% confidence bands, constructed using 500 bootstrap replications. In this case, regression adjustment reduces standard errors by approximately 0.5–15%. Similarly, the bottom-left panel shows the empirical LPTE, and the bottom-right panel shows the regression-adjusted LPTE. Here, the standard errors decrease by about 3.5–26.5% across most of the distribution, except at $y \in \{0,1,2,3\}$, where a slight increase in standard errors is observed.
The regression-adjusted distributional analysis reveals that the probability of having zero emergency department visits decreases by 9 percentage points (pp), with a standard error of 4.2 pp. Beyond this, the only marginally significant effect at the 5% level is an increase of approximately 1.7 pp in the probability of having five ED visits, with a standard error of 0.8 pp. No other statistically significant changes are observed across the rest of the distribution, even after applying regression adjustment.
Conclusion
We introduced a method for estimating local distributional treatment effects in randomized experiments with covariate-adaptive randomization and imperfect compliance. Our approach combines instrumental variable techniques with regression adjustment in a distribution regression framework, leveraging auxiliary covariates and modern machine learning for improved efficiency. The estimator is asymptotically normal, achieves the semiparametric efficiency bound, and performs well in simulations. We also demonstrated its practical relevance using data from the Oregon Health Insurance Experiment.
This work has several limitations. It relies on standard IV assumptions such as monotonicity and the exclusion restriction, and focuses on binary treatments. Performance may vary depending on the quality of nuisance estimation in finite samples. Future research could extend the framework to multi-valued or continuous treatments, relax identifying assumptions, and explore dynamic or longitudinal settings. Furthermore, extending the non-asymptotic frameworks developed by su2023decorrelation, su2023estimated to a distributional setting represents a promising avenue for future research.
Acknowledgements
We are deeply grateful to the four anonymous reviewers and the program chairs for their thoughtful feedback and constructive discussions, which greatly improved the quality of this paper. Oka also acknowledges the financial support provided by JSPS KAKENHI (Grant Number 24K04821).
NeurIPS Paper Checklist
enumerate• {\bf Claims}
• Question: Do the main claims made in the abstract and introduction accurately reflect the paper's contributions and scope?
• Answer: \answerYes
• Justification: The abstract and introduction clearly state the paper’s objectives, proposed approach, and key findings, which are consistently developed and supported throughout the paper.
• Guidelines:
\begin{itemize}
• The answer NA means that the abstract and introduction do not include the claims made in the paper.
• The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations. A No or NA answer to this question will not be perceived well by the reviewers.
• The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalize to other settings.
• It is fine to include aspirational goals as motivation as long as it is clear that these goals are not attained by the paper.
\end{itemize}
• {\bf Limitations}
• Question: Does the paper discuss the limitations of the work performed by the authors?
• Answer: \answerYes
• Justification: The paper discusses the limitations of the work in the Conclusion section and suggests directions for future research.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper.
• The authors are encouraged to create a separate "Limitations" section in their paper.
• The paper should point out any strong assumptions and how robust the results are to violations of these assumptions (e.g., independence assumptions, noiseless settings, model well-specification, asymptotic approximations only holding locally). The authors should reflect on how these assumptions might be violated in practice and what the implications would be.
• The authors should reflect on the scope of the claims made, e.g., if the approach was only tested on a few datasets or with a few runs. In general, empirical results often depend on implicit assumptions, which should be articulated.
• The authors should reflect on the factors that influence the performance of the approach. For example, a facial recognition algorithm may perform poorly when image resolution is low or images are taken in low lighting. Or a speech-to-text system might not be used reliably to provide closed captions for online lectures because it fails to handle technical jargon.
• The authors should discuss the computational efficiency of the proposed algorithms and how they scale with dataset size.
• If applicable, the authors should discuss possible limitations of their approach to address problems of privacy and fairness.
• While the authors might fear that complete honesty about limitations might be used by reviewers as grounds for rejection, a worse outcome might be that reviewers discover limitations that aren't acknowledged in the paper. The authors should use their best judgment and recognize that individual actions in favor of transparency play an important role in developing norms that preserve the integrity of the community. Reviewers will be specifically instructed to not penalize honesty concerning limitations.
\end{itemize}
• {\bf Theory assumptions and proofs}
• Question: For each theoretical result, does the paper provide the full set of assumptions and a complete (and correct) proof?
• Answer: \answerYes
• Justification: The paper provides a complete and correct set of assumptions for each theoretical result, with all theorems and proofs clearly stated and appropriately referenced. Full proofs are included in the supplemental material, with some explanation in the main text to aid understanding.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not include theoretical results.
• All the theorems, formulas, and proofs in the paper should be numbered and cross-referenced.
• All assumptions should be clearly stated or referenced in the statement of any theorems.
• The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition.
• Inversely, any informal proof provided in the core of the paper should be complemented by formal proofs provided in appendix or supplemental material.
• Theorems and Lemmas that the proof relies upon should be properly referenced.
\end{itemize}
• {\bf Experimental result reproducibility}
• Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper (regardless of whether the code and data are provided or not)?
• Answer: \answerYes
• Justification: The paper proposes a new algorithm and provides all necessary details to reproduce the main experimental results, including a clear description of the algorithm, experimental setup, hyperparameters, and evaluation protocols.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not include experiments.
• If the paper includes experiments, a No answer to this question will not be perceived well by the reviewers: Making the paper reproducible is important, regardless of whether the code and data are provided or not.
• If the contribution is a dataset and/or model, the authors should describe the steps taken to make their results reproducible or verifiable.
• Depending on the contribution, reproducibility can be accomplished in various ways. For example, if the contribution is a novel architecture, describing the architecture fully might suffice, or if the contribution is a specific model and empirical evaluation, it may be necessary to either make it possible for others to replicate the model with the same dataset, or provide access to the model. In general. releasing code and data is often one good way to accomplish this, but reproducibility can also be provided via detailed instructions for how to replicate the results, access to a hosted model (e.g., in the case of a large language model), releasing of a model checkpoint, or other means that are appropriate to the research performed.
• While NeurIPS does not require releasing code, the conference does require all submissions to provide some reasonable avenue for reproducibility, which may depend on the nature of the contribution. For example
\begin{enumerate}
• If the contribution is primarily a new algorithm, the paper should make it clear how to reproduce that algorithm.
• If the contribution is primarily a new model architecture, the paper should describe the architecture clearly and fully.
• If the contribution is a new model (e.g., a large language model), then there should either be a way to access this model for reproducing the results or a way to reproduce the model (e.g., with an open-source dataset or instructions for how to construct the dataset).
• We recognize that reproducibility may be tricky in some cases, in which case authors are welcome to describe the particular way they provide for reproducibility. In the case of closed-source models, it may be that access to the model is limited in some way (e.g., to registered users), but it should be possible for other researchers to have some path to reproducing or verifying the results.
\end{enumerate}
\end{itemize}
• {\bf Open access to data and code}
• Question: Does the paper provide open access to the data and code, with sufficient instructions to faithfully reproduce the main experimental results, as described in supplemental material?
• Answer: \answerYes
• Justification: The paper provides open access to both the code and data (we use publicly available data and we included the link to download), along with detailed instructions in the supplemental material for setting up the environment and reproducing the main experimental results.
• Guidelines:
\begin{itemize}
• The answer NA means that paper does not include experiments requiring code.
• Please see the NeurIPS code and data submission guidelines (\url{https://nips.cc/public/guides/CodeSubmissionPolicy}) for more details.
• While we encourage the release of code and data, we understand that this might not be possible, so “No” is an acceptable answer. Papers cannot be rejected simply for not including code, unless this is central to the contribution (e.g., for a new open-source benchmark).
• The instructions should contain the exact command and environment needed to run to reproduce the results. See the NeurIPS code and data submission guidelines (\url{https://nips.cc/public/guides/CodeSubmissionPolicy}) for more details.
• The authors should provide instructions on data access and preparation, including how to access the raw data, preprocessed data, intermediate data, and generated data, etc.
• The authors should provide scripts to reproduce all experimental results for the new proposed method and baselines. If only a subset of experiments are reproducible, they should state which ones are omitted from the script and why.
• At submission time, to preserve anonymity, the authors should release anonymized versions (if applicable).
• Providing as much information as possible in supplemental material (appended to the paper) is recommended, but including URLs to data and code is permitted.
\end{itemize}
• {\bf Experimental setting/details}
• Question: Does the paper specify all the training and test details (e.g., data splits, hyperparameters, how they were chosen, type of optimizer, etc.) necessary to understand the results?
• Answer: \answerYes
• Justification: The paper thoroughly outlines all relevant experimental details, including the use of a data splitting method known as cross-fitting, the selection and tuning of hyperparameters, and other implementation specifics essential for fully understanding and interpreting the experimental results.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not include experiments.
• The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them.
• The full details can be provided either with the code, in appendix, or as supplemental material.
\end{itemize}
• {\bf Experiment statistical significance}
• Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments?
• Answer: \answerYes
• Justification: The paper reports error bars for the key experimental results, clearly explaining the computation methods. Bootstrap resampling and analytical standard error calculations were employed to estimate variance and assess statistical significance.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not include experiments.
• The authors should answer "Yes" if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper.
• The factors of variability that the error bars are capturing should be clearly stated (for example, train/test split, initialization, random drawing of some parameter, or overall run with given experimental conditions).
• The method for calculating the error bars should be explained (closed form formula, call to a library function, bootstrap, etc.)
• The assumptions made should be given (e.g., Normally distributed errors).
• It should be clear whether the error bar is the standard deviation or the standard error of the mean.
• It is OK to report 1-sigma error bars, but one should state it. The authors should preferably report a 2-sigma error bar than state that they have a 96% CI, if the hypothesis of Normality of errors is not verified.
• For asymmetric distributions, the authors should be careful not to show in tables or figures symmetric error bars that would yield results that are out of range (e.g. negative error rates).
• If error bars are reported in tables or plots, The authors should explain in the text how they were calculated and reference the corresponding figures or tables in the text.
\end{itemize}
• {\bf Experiments compute resources}
• Question: For each experiment, does the paper provide sufficient information on the computer resources (type of compute workers, memory, time of execution) needed to reproduce the experiments?
• Answer: \answerYes
• Justification: The paper provides detailed information on the computational resources used for each experiment, including the type of compute, memory specifications, and execution time.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not include experiments.
• The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage.
• The paper should provide the amount of compute required for each of the individual experimental runs as well as estimate the total compute.
• The paper should disclose whether the full research project required more compute than the experiments reported in the paper (e.g., preliminary or failed experiments that didn't make it into the paper).
\end{itemize}
• {\bf Code of ethics}
• Question: Does the research conducted in the paper conform, in every respect, with the NeurIPS Code of Ethics \url{https://neurips.cc/public/EthicsGuidelines}?
• Answer: \answerYes
• Justification:
The research fully conforms to the NeurIPS Code of Ethics. The study considers potential societal and environmental impacts, avoids known risks such as discrimination or misuse, and follows best practices for reproducibility, transparency, and responsible data handling.
• Guidelines:
\begin{itemize}
• The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.
• If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics.
• The authors should make sure to preserve anonymity (e.g., if there is a special consideration due to laws or regulations in their jurisdiction).
\end{itemize}
• {\bf Broader impacts}
• Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed?
• Answer: \answerYes
• Justification: The paper discusses potential societal impacts, noting that the proposed method can lead to improved decision-making in applied settings as a positive outcome. It also acknowledges a potential negative impact, namely that the underlying assumptions of the method may not hold in all real-world scenarios, which could limit its effectiveness or lead to unintended consequences.
• Guidelines:
\begin{itemize}
• The answer NA means that there is no societal impact of the work performed.
• If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact.
• Examples of negative societal impacts include potential malicious or unintended uses (e.g., disinformation, generating fake profiles, surveillance), fairness considerations (e.g., deployment of technologies that could make decisions that unfairly impact specific groups), privacy considerations, and security considerations.
• The conference expects that many papers will be foundational research and not tied to particular applications, let alone deployments. However, if there is a direct path to any negative applications, the authors should point it out. For example, it is legitimate to point out that an improvement in the quality of generative models could be used to generate deepfakes for disinformation. On the other hand, it is not needed to point out that a generic algorithm for optimizing neural networks could enable people to train models that generate Deepfakes faster.
• The authors should consider possible harms that could arise when the technology is being used as intended and functioning correctly, harms that could arise when the technology is being used as intended but gives incorrect results, and harms following from (intentional or unintentional) misuse of the technology.
• If there are negative societal impacts, the authors could also discuss possible mitigation strategies (e.g., gated release of models, providing defenses in addition to attacks, mechanisms for monitoring misuse, mechanisms to monitor how a system learns from feedback over time, improving the efficiency and accessibility of ML).
\end{itemize}
• {\bf Safeguards}
• Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse (e.g., pretrained language models, image generators, or scraped datasets)?
• Answer: \answerNA
• Justification: The paper does not involve the release of models or datasets that pose a high risk for misuse, such as pretrained language models, generative systems, or scraped data, and therefore no specific safeguards are necessary.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper poses no such risks.
• Released models that have a high risk for misuse or dual-use should be released with necessary safeguards to allow for controlled use of the model, for example by requiring that users adhere to usage guidelines or restrictions to access the model or implementing safety filters.
• Datasets that have been scraped from the Internet could pose safety risks. The authors should describe how they avoided releasing unsafe images.
• We recognize that providing effective safeguards is challenging, and many papers do not require this, but we encourage authors to take this into account and make a best faith effort.
\end{itemize}
• {\bf Licenses for existing assets}
• Question: Are the creators or original owners of assets (e.g., code, data, models), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected?
• Answer: \answerYes
• Justification: The Oregon Health Insurance Experiment dataset is publicly available through the NBER website, and we have appropriately cited the original study in our work.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not use existing assets.
• The authors should cite the original paper that produced the code package or dataset.
• The authors should state which version of the asset is used and, if possible, include a URL.
• The name of the license (e.g., CC-BY 4.0) should be included for each asset.
• For scraped data from a particular source (e.g., website), the copyright and terms of service of that source should be provided.
• If assets are released, the license, copyright information, and terms of use in the package should be provided. For popular datasets, \url{paperswithcode.com/datasets} has curated licenses for some datasets. Their licensing guide can help determine the license of a dataset.
• For existing datasets that are re-packaged, both the original license and the license of the derived asset (if it has changed) should be provided.
• If this information is not available online, the authors are encouraged to reach out to the asset's creators.
\end{itemize}
• {\bf New assets}
• Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets?
• Answer: \answerYes
• Justification: The new assets introduced in the paper are well documented, with detailed descriptions of their structure, usage, and limitations. All relevant materials are included as an anonymized zip file in the supplemental submission.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not release new assets.
• Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates. This includes details about training, license, limitations, etc.
• The paper should discuss whether and how consent was obtained from people whose asset is used.
• At submission time, remember to anonymize your assets (if applicable). You can either create an anonymized URL or include an anonymized zip file.
\end{itemize}
• {\bf Crowdsourcing and research with human subjects}
• Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation (if any)?
• Answer: \answerNA
• Justification: The paper does not involve crowdsourcing or research with human subjects, and therefore no participant instructions, screenshots, or compensation details are applicable.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.
• Including this information in the supplemental material is fine, but if the main contribution of the paper involves human subjects, then as much detail as possible should be included in the main paper.
• According to the NeurIPS Code of Ethics, workers involved in data collection, curation, or other labor should be paid at least the minimum wage in the country of the data collector.
\end{itemize}
• {\bf Institutional review board (IRB) approvals or equivalent for research with human subjects}
• Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals (or an equivalent approval/review based on the requirements of your country or institution) were obtained?
• Answer: \answerNA
• Justification: The paper does not involve research with human subjects or crowdsourcing, so there were no participant risks to assess and no need for IRB or equivalent ethical review.
• Guidelines:
\begin{itemize}
• The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.
• Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research. If you obtained IRB approval, you should clearly state this in the paper.
• We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution.
• For initial submissions, do not include any information that would break anonymity (if applicable), such as the institution conducting the review.
\end{itemize}
• {\bf Declaration of LLM usage}
• Question: Does the paper describe the usage of LLMs if it is an important, original, or non-standard component of the core methods in this research? Note that if the LLM is used only for writing, editing, or formatting purposes and does not impact the core methodology, scientific rigorousness, or originality of the research, declaration is not required.
• Answer: \answerNA
• Justification: The core methodology and contributions of the paper do not involve the use of LLMs in any important, original, or non-standard way.
• Guidelines:
\begin{itemize}
• The answer NA means that the core method development in this research does not involve LLMs as any important, original, or non-standard components.
• Please refer to our LLM policy (\url{https://neurips.cc/Conferences/2025/LLM}) for what should or should not be described.
\end{itemize}