EconBase
← Back to paper

Regression Adjustment for Estimating Distributional Treatment Effects in Randomized Controlled Trials

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

56,837 characters · 13 sections · 65 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Regression Adjustment for Estimating Distributional Treatment Effects in Randomized Controlled Trials

\thispagestyle{empty}

abstractIn this paper, we address the issue of estimating and inferring distributional treatment effects in randomized experiments. The distributional treatment effect provides a more comprehensive understanding of treatment heterogeneity compared to average treatment effects. We propose a regression adjustment method that utilizes distributional regression and pre-treatment information, establishing theoretical efficiency gains without imposing restrictive distributional assumptions. We develop a practical inferential framework and demonstrate its advantages through extensive simulations. Analyzing water conservation policies, our method reveals that behavioral nudges systematically shift consumption from high to moderate levels. Examining health insurance coverage, we show the treatment reduces the probability of zero doctor visits by 6.6 percentage points while increasing the likelihood of 3-6 visits. In both applications, our regression adjustment method substantially improves precision and identifies treatment effects that were statistically insignificant under conventional approaches.

{\it Keywords:} Randomized experiment, A/B testing, Distributional treatment effect, Heterogeneous treatment effect, Regression adjustment, Distributional regression \\ {\it JEL Codes:} C14, C32, C53, E17, E44

\setcounter{page}{1} \setstretch{1.5}

Introduction

Since the foundational works of fisher1925statistical, fisher1935design and Neyman1923, randomized controlled trials (RCTs) have become essential in evaluating the effectiveness of treatments or interventions. Over the last few decades, RCTs, also known as A/B testing, have emerged as key tools for causal inference across a wide range of scientific and engineering disciplines. These fields include medicine rubin1997estimating, economics duflo2007using, athey2017econometrics, political science imai2005get, horiuchi2007designing, and sociology baldassarri2017field, among others. For a more thorough exploration, refer to the modern textbook treatments in imbens2015causal.

In RCTs, researchers often encounter the challenge of low detection power, where the variance of the treatment effect estimator is large relative to the actual treatment effect, leading to ambiguous results lewis2015unfavorable. Contrary to the common belief that large sample sizes alone can resolve this issue, research indicates this is not always the case Deng2013. Increasing sample size or experiment duration is costly and not always feasible. Therefore, using estimators with smaller variance can improve precision and lead to more reliable estimates. Additionally, examining distributional treatment effects, not just average treatment effects, is of practical interest to researchers bitler2006mean.

In this paper, we introduce a regression adjustment method to improve the precision of estimating the distributional treatment effect (DTE) in RCTs. The proposed method uses pre-treatment information through distributional regression (DR). Our approach addresses the challenge of low detection power, employing a semiparametric approach in that the underlying distributions are left unspecified. This method accommodates a wide range of outcome distributions, including discrete, continuous, and mixed variables, and is useful for documenting heterogeneous treatment effects.

We theoretically demonstrate that regression adjustment can reduce the variance of distribution estimators and enhance the precision of the DTE estimator. We establish the uniform asymptotic properties of the DTE estimator. We also introduce a practical resampling method for statistical inference and establish its theoretical validity. Through Monte Carlo simulations, we evaluate our method and show that regression adjustment significantly improves the precision of the DTE estimator and outperforms existing methods. Additionally, we apply our method in two empirical studies: a randomized experiment on the impact of nudges on water usage and the Oregon Health Insurance Experiment. Our results highlight the importance of distributional treatment estimation with regression adjustment.

Our paper provides several important contributions to the existing literature. First, our paper contributes to a growing body of literature on heterogeneous treatment effects in econometrics and statistics Athey2016, Wager2018. The quantile treatment effect (QTE) was first introduced by doksum1974empirical and lehmann1975nonparametrics for estimating treatment effects due to unobserved heterogeneity. Since then, various estimation and inference methods for the distributional and quantile treatment effect have been developed and applied, including heckman1997making, abadie2002instrumental, athey2006identification, bitler2006mean, chernozhukov2005iv, djebbari2008heterogeneous, donald2014estimation, callaway2018quantile, callaway2019quantile, chernozhukov2019generic, and kallus2024localized, among others. jiang2023regression study the estimation of regression-adjusted QTE under covariate adaptive randomization schemes. To the best of our knowledge, however, techniques for estimating DTE has not been explored in the context of regression-adjustment method.\footnote{ To estimate heterogeneous treatment effects, researchers commonly use two mututally not exclusive approaches: conditional average treatment effect (CATE) and distributional treatment effect (DTE). CATE estimates the average treatment effect for subgroups based on observed variables Athey2016, Imai2013, Metalearner2019, Nie2020, Wager2018. DTE, on the other hand, compares outcome distributions between treatment and control groups across the entire spectrum, and thus it can capture effects that CATE might miss, particularly those stemming from unobserved factors. As such, DTE remains valuable even after accounting for observed variables in CATE analyses. }\textsuperscript{,}\footnote{jiang2023regression and our paper complement each other. While jiang2023regression considers our experimental scheme as a special case and employs doubly robust moment conditions for quantile regression, our proposed method applies to non-continuous outcome variables and also employs doubly robust moment condition with distribution regression framework. }

Next, the literature has extensively investigated the use of pretreatment covariates to estimate ATE with the regression adjustment fisher1925statistical, cochran1957analysis, cox1982biometrics, frison1992repeated. freedman2008regression,freedman2008regressionA and lin2013agnostic have noted the potential issue of the adjustment due to finite-sample bias and misspecification. Recent works by Yang2001, Tsiatis2008 and Ansel2018 apply linear regression adjustment, while rosenblum2010simple and bartlett2018covariate explore nonlinear regression adjustment. list2022using use machine learning methods for regression adjustment. imbens2015causal and negi2021revisiting, negi2020robust consider regression adjustment for both linear and nonlinear regression models. These existing works focus on average treatment effect (ATE), whereas our research broadens the scope of regression adjustment to estimate the impact of treatment on the entire distribution in randomized experiments.

Thirdly, for our regression adjustment, we utilize the DR approach, which can be considered a semiparametric approach for estimating conditional distributions, in that both marginal and conditional distributions are left unspecified. The semiparametric property is essential in practice since we rarely know the true underlying distributions in many applications. williams1972analysis introduce DR to analyze ordered categorical outcomes by using multiple binary regressions. foresi1995conditional extend the approach to characterize a continuous, conditional distribution and hall1999methods propose a local version of distributional regression. chernozhukov2013inference establish its uniform validity for estimating the entire conditional distribution when an asymptotical continuum of binary regressions is used. See also chernozhukov2020network, chernozhukov2023distribution, wang2023bivariate, wang2023distributional for later developments.

The rest of the paper is organized as follows. Section (ref) introduces the potential outcomes framework and defines the treatment effect parameters of interest within our study. Section (ref) presents theoretical results. In Section (ref), we provide finite-sample properties through Monte Carlo simulations. Section (ref) offers two empirical case studies. Finally, we conclude in Section (ref). All proofs and additional simulation results can be found in Supplemental Material.

Treatment Effect and Estimation Method

This section first introduces the setup and explain a model for regression adjustment. In what follows, we use $\mathbb{E}[\cdot]$ and $\ensuremath{\mathrm{Var}}(\cdot)$ as the expectation and variance operator, respectively. All small-$o$ notation and asymptotic results, including $o_p(1)$ and $O_p(1)$, in the paper are relative to sample size $n$. Let $\ell^{\infty}(T)$ be the collection of all bounded real-valued functions defined on an arbitrary index set $T$. We use $\rightsquigarrow$ to denote convergence in law. That is, given $X_{n} \in \ell^{\infty}(T)$ or $X_{n}(t)$ for $t \in T$, $X_{n} \rightsquigarrow X$ in $\ell^{\infty}(T)$ if $X_{n}$ converge in law to $X$ in $\ell^{\infty}(T)$.

Setup and Treatment Effects

Consider a randomized controlled trial with $K$ treatments, where each individual is randomly assigned to one of the treatments with pre-specified probabilities $\{\pi_{k}\}_{k=1}^{K}$ with $\sum_{k=1}^{K}\pi_{k} =1 $ and $ \pi_{k}> 0$ for all $k \in \mathcal{K}:= \{1, \dots, K \}$. We consider the potential outcome framework Neyman1923, rubin1974estimating. Let $Y$ denote the (observed) outcome of interest with its support $\mathcal{Y} \subset \mathbb{R}$ and $Y(k) $ the potential outcome under treatment $k \in \mathcal{K}$. Then the following relation holds: $ Y = \sum_{k=1}^{K}W_{k} Y(k) $, where $W_{k}$ is set to 1 if an individual is assigned to treatment $k \in \mathcal{K}$ and 0 otherwise and satisfies $\Pr(W_{k}= 1) = \pi_{k}$ and $\sum_{k=1}^{K} W_{k} = 1$.

Suppose we observe a random sample $\{(\bm{W}_{i}, \bm{X}_{i}, Y_{i})\}_{i=1}^{n}$ drawn independently from the distribution of $(\bm{W}, \bm{X}, Y)$ with a sample size of $n$. Here, $\bm{W} := (W_{1}, \dots, W_{K})^{\top}$ denotes the vector of treatment assignment indicators for $K$ treatments with its support $\mathcal{W}:= \{0,1\}^{K}$ and $\bm{X}$ a $p \times 1$ vector of pre-treatment covariates with its support $\mathcal{X} \subset \mathbb{R}^{p}$. We assume that $\bm{X}$ realizes before the randomized controlled trials and is independently and identically distributed even across treatment statuses. Given $\bm{W}_{i}= (W_{1, i}, \dots, W_{K, i})^{\top}$, the number of observations assigned to treatment $k \in \mathcal{K}$ is defined as $n_{k} := \sum_{i=1}^{n} W_{k,i}$ and the empirical probability of treatment assignment is given by $\hat{\pi}_{k}:=n_{k}/n$ for $k= 1, \dots, K$.

One prevalent approach to measure the treatment effect is the average treatment effect (ATE), defined as $ ATE_{k,k'} := \mathbb{E}[Y(k)] - \mathbb{E}[Y(k')] $ for $k, k' \in \mathcal{K}$. ATE is simple to estimate and interpret, while it is incapable of estimating the treatment effect on the entire distribution. To consider the distributional features, let $F_{Y(k)}(y)$ be the distribution function of $Y(k)$ for treatment status $k \in \mathcal{K}$ and $\gamma:= ( \mathbb{F}_{Y(1)}, \dots, \mathbb{F}_{Y(K)} )^{\top} \in \Gamma \subseteq \ell^{\infty}(\mathcal{Y})^{K} $ a vector of the potential outcome distributions. We define the distributional treatment effect (DTE) between treatments $k, k' \in \mathcal{K}$ as

eqnarray*[eqnarray* omitted — 74 chars of source]

where the functional $\Psi_{k, k'}^{DTE}: \Gamma \to \ell^{\infty}(\mathcal{Y})$ is defined as $ \Psi_{k, k'}^{DTE}(\gamma) = F_{Y(k)} - F_{Y(k')}$. One of its advantages is that DTE is well-defined for any type of outcome, including discrete variables.\footnote{ Our approach measures the distributional effect as the difference between the distributions of potential outcomes under distinct treatments. While this paper focuses on DTE, we note that under the rank invariance assumption, our results are also informative about the distribution of individual treatment effects (ITE). Without additional conditions or rank invariance, researchers may consider adopting a partial-identification approach for estimating the distribution of ITE. See heckman1997making among others. }

Moreover, the distributional information is an essential building block for other types of treatment effect parameters. For example, fixing some constant $h>0$, we can write the probability $\Pr\{y < Y(k) \le y+h \} = F_{Y(k)}(y+h) - F_{Y(k)}(y)$, thereby defining the probability treatment effect (PTE) as

align*[align* omitted — 76 chars of source]

where the functional $\Psi_{k, k', h}^{PTE}: \Gamma \to \ell^{\infty}(\mathcal{Y})$ is defined as

eqnarray*[eqnarray* omitted — 184 chars of source]

PTE measures changes in the probability that the outcome variable falls in interval $(y, y+h]$ and can apply for any type of outcome variables. For an integer outcomes and $h=1$, $\Delta_{k, k',1}^{PTE}(y)$ quantifies the probability change at location $y$, or $\Pr\{Y(k) = y\} - \Pr\{Y(k') = y\}$. Alternatively, for continuous outcome variables, we can consider the quantile treatment effect (QTE), which is defined through a simple inversion of the distribution function. For $u \in (0,1)$, let the quantile function of $Y(k)$ be defined as the map $u \mapsto F_{Y(k)}^{-1}(u)$. The QTE between treatments $k, k' \in \mathcal{K}$ is defined as $ \Delta^{QTE}_{k, k'} := \Psi_{k, k'}^{QTE}(\gamma)(u) $ where $ \Psi_{k, k'}^{QTE}(\gamma)(u) := F_{Y(k)}^{-1}(u) - F_{Y(k')}^{-1}(u). $ Since standard analysis based on quantile functions presumes the existence of the outcome density, QTE is not readily applicable in the presence of discreteness in the outcome distributions.

Distributional Regression

In this paper, we extend the idea of a regression-adjustment approach for estimating the DTE and PTE, using the distributional regression (DR) framework with pre-treatment covariates. The true conditional distribution of $Y(k)$ given $\bm{X}$ can be written as

eqnarray*[eqnarray* omitted — 99 chars of source]

where $\1_{\{\cdot\} }$ represents the indicator function taking the value of 1 if the condition within $\{\cdot\}$ is satisfied, and 0 otherwise.

In many practical applications, the true conditional distribution $F_{Y(k)|\bm{X}}(y|\bm{X})$ remains unknown. It is crucial not to impose overly restrictive global parametric constraints. To address this issue, we employ the distribution regression framework, wherein a parametric linear-index model is used to describe the binary outcomes $\1_{ \{Y(k) \le y\} }$ for each $y \in \mathcal{Y}$. More specifically, let $\Lambda: \mathbb{R} \to [0,1]$ be a known inverse link function. For each $y \in \mathcal{Y}$, we define the following conditional distribution based on a linear index model:

eqnarray[eqnarray omitted — 124 chars of source]

where $T: \mathcal{X} \to \mathbb{R}^{d}$ is some transformation, such as polynomial or pair-wise interaction of regressors, specified by researchers, and $\beta_{k}(y) {\in} \mathbb{R}^{d}$ are unknown parameters. Setting $\Lambda(\cdot)$ as either the normal or logistic distribution function yeilds probit or logit models, respectively, for binary outcome $\1_{ \{Y(k) \le y\} }$.

It is worth noting that the unknown parameters $\beta_{k}(y)$ are specific to point $y \in \mathcal{Y}$, and can be considered as pseudo-parameters to characterize the true conditional distribution at each point $y$. These pseudo-parameters capture local information of the distributions of interest, and can be estimated at the parametric rate under certain regularity conditions, as detailed later. Thus, we view the conditional distribution model $G_{Y(k)| \bm{X}}(\cdot| \bm{X})$ in ((ref)) as an approximation model for the true conditional distribution $F_{Y(k)|\bm{X}}(\cdot|\bm{X})$, and our subsequent analysis accommodates potential model misspecifications.

We estimate the DR model in ((ref)) using the quasi-maximum likelihood framework White1982, separately for observations exposed to different treatments. For each $(k, y) \in \mathcal{K}\times\mathcal{Y}$, the estimator for $\beta_{k}(y)$ is defined as the maximizer of the log-likelihood function:

eqnarray[eqnarray omitted — 131 chars of source]

where the log-likelihood function for treatment $k$ is denoted by $\widehat{\ell}_{k}(\beta; y) := n_{k}^{-1} \sum_{i=1}^{n} W_{k,i} \cdot \ell_{i}(\beta;y)$ with $\ell_{i} (\beta:y) = \1_{ \{Y_{i} \le y \} } \log \Lambda \big( T(\bm{X}_i)^{\top} \beta \big ) + (1 - \1_{ \{Y_{i} \le y \} } ) \log \big \{ 1 - \Lambda\big( T(\bm{X}_i)^{\top} \beta \big) \big \}$ and $\mathcal{B} \subset \mathbb{R}^{d}$ is the parameter space. Given the estimator $\widehat{\beta}_{k}(y)$, we define the conditional distribution function estimator as

eqnarray[eqnarray omitted — 143 chars of source]

We can estimate the conditional distribution above for each value of $y \in \mathcal{Y}$ when the outcome variable $Y$ takes finite discrete values, and we select a sufficiently large set of discrete points $\{y_j \in \mathcal{Y} \}_{j=1}^{J}$ for estimation when the outcome variable is a continuous random variable.\footnote{ One important property of the conditional distribution function is non-decreasing by definition, while the estimator $\widehat{G}_{Y(k)|\bm{X}}(y|\bm{x})$ does not necessarily satisfy monotonicity in finite samples due to estimation errors. We can monotonize the conditional distribution estimators using the rearrangement method, for example, as proposed by chernozhukov2009improving. }

Regression-Adjusted Treatment Effect Estimator

By randomization, the potential outcome distributions $F_{Y(k)}$ for $k\in\mathcal{K}$ are identified as the outcome distributions under treatment $k$. Hence, it can be estimated from the data $\{(\bm{W}_{i}, \bm{X}_{i}, Y_{i})\}_{i=1}^{n}$. The simple DTE and PTE estimators are based on the empirical distribution functions $\widehat{\mathbb{F}}_{Y(k)}^{simple}(y) := n_{k}^{-1} \sum_{i=1}^{n} W_{k,i} \cdot \1_{ \{Y_{i} \le y \} } $ for $k \in \mathcal{K}$. We propose a new estimator based on the distributional regression adjustment with pre-treatment covariates. Define the discrete probability measures $ \widehat{\mathbb{P}}_{\bm{X}} := n^{-1}\sum_{i=1}^{n} \delta_{\bm{X}_i} $ and $ \widehat{\mathbb{P}}_{\bm{X}}^{(k)} := n_{k}^{-1}\sum_{i=1}^{n} W_{k,i} \cdot \delta_{\bm{X}_i} $ for covariates $\bm{X}$ and for $k \in \mathcal{K}$, where $\delta_{\bm{x}}$ is the measure that assigns mass 1 at $\bm{x} \in \mathcal{X}$. For an arbitrary real-valued function $f: \mathcal{X} \to \mathbb{R}$, we denote by

align*[align* omitted — 237 chars of source]

Then, for treatment $k \in \mathcal{K}$, we define the regression-adjusted distribution function as,

align[align omitted — 176 chars of source]

The regression-adjusted estimator is the sample average of conditional distribution functions over pre-treatment covariates utilizing all observations, where these conditional distributions are estimated using only observations pertaining to specific treatments.

Letting $\widehat{\gamma} := ( \widehat{\mathbb{F}}_{Y(1)}, \dots, \widehat{\mathbb{F}}_{Y(K)} )^{\top} $, we define the regression-adjusted DTE and PTE estimators:

eqnarray[eqnarray omitted — 230 chars of source]

Similarly, the QTE estimator is defined as $ \widehat{\Delta}^{QTE}_{k,k'}:= \Psi_{k, k'}^{QTE} (\, \widehat{\gamma} \, ) $.

The linear index model in ((ref)) can be viewed as a special case of generalized linear models (GLMs). Within the GLM framework, using the canonical link function yields unbiased estimation of the unconditional mean in finite samples (see Section (ref) for more details and mccullagh1989binary). Specifically, the canonical link function must satisfy

align[align omitted — 182 chars of source]

At the population level, this relation can be expressed as: $ \mathbb{E} \big[ \1_{ \{Y(k) \le y \} } - G_{Y(k)|\bm{X}}(y|\bm{X}) \big ] = 0. $ This property is important for the regression adjustment to guarantee that the regression adjustment does not induce bias to improve the precision of the treatment effect estimator. negi2021revisiting pointed this out for a nonlinear regression adjustment method and derive the asymptotic property when the true distribution is known. In what follows, we focus the GLMs with the canonical link functions for binary outcome, which includes OLS and logit, while we do not assume any global parametric assumptions and allow for model mis-specification.

Theoretical Results

In this section, we first derive the oracle efficiency of the proposed estimator. We subsequently present asymptotic properties and a resampling method.

Oracle Estimator and Efficiency

In this section, we investigate possible efficiency gain from the regression adjustment, when the population counterpart related to the adjustment is available. This analysis will allow us to establish the efficiency property.

Leveraging the finite-sample unbiasedness property in ((ref)), we can rewrite ((ref)) as

align*[align* omitted — 212 chars of source]

where the second term of the right-hand side involves the DR estimator for the adjustment. In the above equations, we meanwhile consider the situation where the estimator $\widehat{G}_{Y(k)|\bm{X}}$ is replaced by the population model $G_{Y(k)|\bm{X}}(y|\bm{X})$ in ((ref)). Then, define the oracle distribution function with the regression adjustment

align*[align* omitted — 240 chars of source]

Theorem (ref) below presents the efficiency gain resulting from the application of our regression adjustment method, given the oracle estimator above. Although this oracle estimator is infeasible, the subsequent section (specifically, in the proof of Theorem (ref)) establishes an asymptotic equivalence between the feasible and oracle estimators.

For Theorem (ref), we make the following assumptions regarding the experimental setup and sampling scheme:

assumption\ \\ (a) The observations $\{(\bm{W}_{i}, \bm{X}_{i}, Y_{i})\in \{0,1\}^{K}{\times}\mathcal{X}{\times}\mathcal{Y}\}_{i=1}^{n}$ are $n$ independent copies of the random variables $(\bm{W}, \bm{X}, Y)$. \\ (b) Treatment assignment is independent of both potential outcomes and pre-treatment covariates: $ \bm{W} \perp \!\!\! \perp \big(Y(1), \dots, Y(K), \bm{X})$. The pre-specified treatment probabilities $\{\pi_{k}\}_{k=1}^{K}$ satisfy $\pi_{k}>0$ for every $k \in \mathcal{K}$ and $\sum_{k=1}^{K} \pi_{k}= 1$. \\

Assumption (ref) (a) requires random sampling and Assumption (ref) (b) is standard under a randomized controlled trial with multiple treatments.

theorem(a) Suppose that Assumption (ref) holds and that $\hat{\pi}_{k} = \pi_{k} + o(1)$ as $n\to \infty$ for every $k \in \mathcal{K}$. Then, for each $k \in \mathcal{K}$ and for any $y \in \mathcal{Y}$, we have in $\ell^{\infty}(\mathcal{Y})$, \begin{align*} \ensuremath{\mathrm{Var}} \big( \hat{\mathbb{F}}_{Y(k)}^{simple} \big) \ge \ensuremath{\mathrm{Var}} \big( \widetilde{\mathbb{F}}_{Y(k)} \big) + o(n^{-1}), \end{align*} provided that $ \mathbb{E} [ ( \1_{ \{Y(k) \le \cdot\} } - F_{Y(k)} )^2 ] \ge \mathbb{E} [ ( \1_{ \{Y(k) \le \cdot\} } - G_{Y(k)|\bm{X}} )^2 ] $. (b) Define $\widehat{\gamma}^{simple} := ( \widehat{\mathbb{F}}_{Y(1)}^{simple}, \dots, \widehat{\mathbb{F}}_{Y(K)}^{simple} )^{\top} $ and $\widetilde{\gamma} := ( \widetilde{\mathbb{F}}_{Y(1)}, \dots, \widetilde{\mathbb{F}}_{Y(K)} )^{\top} $. Additionally, we assume that the model specification for distributional regression in ((ref)) is correct. Then, \begin{align*} \ensuremath{\mathrm{Var}} \big( \widehat{\gamma}^{simple} \big) \succeq \ensuremath{\mathrm{Var}} \big( \widetilde{\gamma} \big) + o(n^{-1}), \end{align*} where $\succeq$ denotes the positive semi-definiteness. When $ \ensuremath{\mathrm{Var}} \big ( F_{Y(k)}(y|\bm{X}) - r \cdot F_{Y(k')}(y|\bm{X}) \big ) > 0 $ for any distinct $k, k' \in \mathcal{K}$ and $r \in \mathbb{R}$, the positive definite result holds.

Theorem (ref)(a) shows the efficiency gains achieved by applying regression adjustment to estimate unconditional distribution functions. More specifically, the variance of the unconditional distribution estimators decreases if $G_{Y(k)|\bm{X}}(y)$ has some predictive power for $\1_{\{Y(k) \le y\}}$ in the $L^2$ sense. When $\Lambda(\cdot)$ is the identity function or in the case of a linear prediction model, this property holds automatically. In the case of logistic regression, the level of model misspecification must be sufficiently small relative to the predictive signal in the data.\footnote{ A simple algebra shows that $ \mathbb{E} [ ( \1_{ \{Y(k) \le \cdot\} } - F_{Y(k)} )^2 ] - \mathbb{E} [ ( \1_{ \{Y(k) \le \cdot\} } - G_{Y(k)|\bm{X}} )^2 ] = \mathbb{E} [ ( F_{Y(k)|\bm{X}} - F_{Y(k)} )^2 ] - \mathbb{E} [ ( F_{Y(k)|\bm{X}} - G_{Y(k)|\bm{X}} )^2 ] $. Here, the first term measures the signal in $\bm{X}$ for predicting $\1_{ \{Y(k) \le \cdot\} }$. The second term quantifies the model misspecification, that is, the $L^2$ distance between the true conditional distribution $F_{Y(k)|\bm{X}}$ and the specified model $ G_{Y(k)|\bm{X}} $. When the model misspecification (second term) is sufficiently small relative to the signal in the data (quantified by the variance of \( F_{Y(k)|\bm{X}} \)), the predictor \( G_{Y(k)|\bm{X}} \) achieves variance reduction for the unconditional estimator. }

Furthermore, Theorem (ref)(b) elaborates on these gains in a positive semi-definite sense for a vector of regression-adjusted estimators. This result encompasses any linear combination of the unconditional distribution estimators and implies potential efficiency gains for both DTE and PTE estimators. While a similar positive-definiteness result for non-linear regression adjustment is reported in Theorem 6 of negi2020robust, our result provides an additional insight that the efficiency gains occur when the regression models are not linearly dependent across treatments. Theorem (ref)(b) requires correct specification, which is analogous to the requirement in negi2020robust for non-linear regression adjustment. We also show the point-wise semiparametric efficiency of the DTE estimator under correct specification in Section (ref) for the sake of completeness.

Asymptotic Distribution of DTE Estimator

In this subsection, we derive the asymptotic distribution of the DTE and PTE estimators, which is useful for statistical inference. We also introduce a resampling technique and show its asymptotic validity. We obtain the results in this section under the following assumptions:

assumptionFor any $y \in \mathcal{Y}$ and $k \in \mathcal{K}$, the following conditions hold. \\ (a) The log-likelihood functions $\beta \mapsto \ell_{k}(\beta;y)$ is a concave function. The link function is a canonical link function, and its inverse map $\Lambda(\cdot)$ is twice continuously differentiable with its first derivative $\lambda(\cdot)$. (b) The true parameters $\beta_{k}(y)$ uniquely solve the maximization problem in $\max_{\beta}\mathbb{E}[\widehat{\ell}_{k}(\beta;y)]$ and are contained in the interior of the compact parameter space $\mathcal{B}$. (c) The maximum eigenvalues of $H_{k}(y)$ is strictly negative uniformly over $y \in \mathcal{Y}$. (d) If $Y$ is a continuous random variable, then the conditional density function $f_{Y|\bm{X} }(y| \bm{x})$ exists, is uniformly bounded, and is uniformly continuous in $(y, x) \in \mathcal{Y} \times \mathcal{X}$.

Assumption (ref) (a)-(c) are needed for deriving the asymptotic property of the DR estimators and are standard for the likelihood estimation of binary-choice models in the literature. Assumption (ref) (d) is required only when the outcome variable is continuous.

In the below theorem, we derive the limiting distribution of the regression-adjusted DTE and PTE estimators.

theoremSuppose that Assumptions (ref) and (ref) hold. \\ (a) Then, we have, for every $k, k' \in \mathcal{K}$ and some constant $h>0$, in $ \ell^{\infty}(\mathcal{Y})$, \begin{align*} \sqrt{n} ( \widehat{\Delta}^{DTE}_{k,k'} - \Delta^{DTE}_{k,k'} ) \rightsquigarrow \Psi^{DTE}_{k,k'}(\mathbb{B}) \ \ \ \mathrm{and} \ \ \ \sqrt{n} ( \widehat{\Delta}^{PTE}_{k,k',h} - \Delta^{PTE}_{k,k',h} ) \rightsquigarrow \Psi^{PTE}_{k,k',h}(\mathbb{B}), \end{align*} where $\mathbb{B}$ is a mean-zero Gaussian process defined in equation ((ref)) of Supplemental Material. (b) Additionally, if the potential outcome $Y_{k}$ has a continuously differentiable distribution function with strictly positive density $f_{Y_k}$ on its support for every $k \in \mathcal{K}$, then \begin{align*} \sqrt{n} ( \widehat{\Delta}^{QTE}_{k,k'} - \Delta^{QTE}_{k,k'} ) \rightsquigarrow \psi^{DTE}_{k,k', \gamma}(\mathbb{B}), \end{align*} where $\psi^{DTE}_{k,k', \gamma}(h)(u) := h_{k} \circ F_{Y_{k}}^{-1}(u) / f_{Y_{k}} \circ F_{Y_{k}}^{-1}(u) - h_{k'} \circ F_{Y_{k'}}^{-1}(u) / f_{Y_{k'}} \circ F_{Y_{k'}}^{-1}(u)$ is the Hadamard derivative of $\Psi^{DTE}_{k,k'}(\cdot)$ at $\gamma$ in direction $h:=(h_{1}, \dots, h_{K})^{\top}$.

The limiting processes in the above theorem rely solely on the sampling error of the original data, $\{(\bm{W}_{i}, \bm{X}_{i}, Y_{i})\}_{i=1}^{n}$, and are free from estimation errors in DR. This is ensured through the use of the stochastic equicontinuity property in the theorem's proof, which is in line with the argument presented in newey1994asymptotic. This result highlights the significance of our DR-based regression adjustment, as it does not add additional variance to the DTE and PTE estimators for large sample sizes.

Although it is simplified, the limiting processes of the DTE and PTE estimators depend on unknown nuisance parameters and may complicate inference in finite samples. To address the challenge of non-pivotal limiting processes in finite samples, we turn to the exchangeable bootstrap praestgaard1993Exchangeably, van1996weak. This approach consistently estimates limit laws of relevant empirical distributions, allowing us to consistently estimate the limit process of the estimator via the functional delta method.

For the resampling scheme, we introduce a vector of random weights $(S_{1}, \dots, S_{n})$. To establish the validity of the bootstrap, we assume that the random weights satisfy the following assumption.

assumption[Bootstrap] Let $(S_{1}, \dots, S_{n})$ be $n$ scalar, nonnegative exchangeable random variables that are independent of the original sample. The random weights $\{S_{i}\}_{i=1}^{n}$ satisfy the followings: $ \mathbb{E}|S_{i}|^{2+\epsilon} < \infty$ for some $\epsilon>0$, $\bar{S}_{n} := n^{-1} \sum_{i=1}^{n} S_{i} \to^p 0 $ and $ n^{-1} \sum_{i=1}^{n} ( S_{i} - \bar{S}_{n} )^2 \to^p 1 $.

This resampling scheme encompasses a variety of bootstrap methods, such as the empirical bootstrap, subsampling, wild bootstrap and so on van1996weak.

Under treatment $k \in \mathcal{K}$, the empirical distribution of the outcome variable with the random weights and the empirical probability measure of $\bm{X}$ with the random weights are respectively given by

align*[align* omitted — 302 chars of source]

Then, for treatment $k \in \mathcal{K}$, we define regression-adjusted distribution functions in $\ell^{\infty}(\mathcal{Y})$ as,

align[align omitted — 147 chars of source]

In the above equation, we use $\widehat{G}_{Y(k)|\bm{X}}^{\ast}$ the conditional distribution estimator based on the original sample, instead of the bootstrap sample. This improves the computational efficiency of our bootstrap inference by eliminating the need to reestimate the parameters for each bootstrap repetition. According to Theorem (ref), the conditional distribution estimator has no impact on the limit distribution and we keep the estimator throughout the resampling procedure. The validity is proven in the theorem below.

Letting $\widehat{\gamma}^{\ast} := ( \widehat{\mathbb{F}}_{Y(1)}^{\ast}, \dots, \widehat{\mathbb{F}}_{Y(K)}^{\ast})^{\top} $, we define the bootstrapped treatment effect estimator with regression adjustment as

eqnarray*[eqnarray* omitted — 367 chars of source]

In the theorem provided below, we show that the exchangeable bootstrap provides a method to consistently estimate the limit process of the regression-adjusted DTE, PTE and QTE estimators, following van1996weak.

theoremSuppose that Assumptions (ref)-(ref) hold. \\ (a) Then, we have, in $ \ell^{\infty}(\mathcal{Y})$, \begin{align*} \sqrt{n} ( \widehat{\Delta}^{DTE \ast}_{k,k'} - \widehat{\Delta}^{DTE}_{k,k'} ) \overset{p}{\rightsquigarrow} \Psi^{DTE}_{k,k'}(\mathbb{B}) \ \ \mathrm{and} \ \ \sqrt{n} ( \widehat{\Delta}^{PTE \ast}_{k,k',h} - \widehat{\Delta}^{PTE}_{k,k',h} ) \overset{p}{\rightsquigarrow} \Psi^{PTE}_{k,k',h}(\mathbb{B}), \end{align*} for every $k, k' \in \mathcal{K}$ and some constant $h>0$, where $\overset{p}{\rightsquigarrow}$ denotes the conditional weak convergence in probability, which is explained in more detail in Supplemental Material (ref) and $\mathbb{B}$ is a mean-zero Gaussian process defined in Lemma (ref) of Supplemental Material. (b) Additionally, if the potential outcome $Y_{k}$ has a continuously differentiable distribution function with strictly positive density $f_{Y_k}$ on its support for every $k \in \mathcal{K}$, then \begin{align*} \sqrt{n} ( \widehat{\Delta}^{QTE \ast}_{k,k'} - \widehat{\Delta}^{QTE}_{k,k'} ) \rightsquigarrow \psi^{DTE}_{k,k', \gamma}(\mathbb{B}), \end{align*} where $\psi^{DTE}_{k,k', \gamma}(\cdot)$ is defined in Theorem (ref)(b).

Theorems (ref) and (ref) present a methodologically rigorous and empirically useful approach for estimating DTE and PTE. Moreover, these theoretical results establish the asymptotic validity of the exchangeable bootstrap procedure for constructing confidence intervals and conducting hypothesis tests for both DTE and PTE, as well as QTE under additional conditions.

The following algorithm describes the construction of confidence intervals for regression-adjusted DTE using empirical bootstrap. The same procedure can be applied to obtain confidence intervals for regression-adjusted PTE, QTE and simple estimators.

algorithm[algorithm omitted — 1,233 chars of source]

In the above algorithm, the standard errors $\widehat{\Sigma}^{DTE}_{k,k'}(y)$ can be computed by (a) using the bootstrap standard deviation of the bootstrapped DTEs $\{\widehat{\Delta}^{DTE,b}_{k,k'}(y)\}_{b=1}^B$, or (b) using the rescaled bootstrap interquartile range $\big ( q_{0.75}(y) - q_{0.25}(y) \big ) / (z_{0.75} - z_{0.25})$, where $q_p(y)$ is the $p$-th quantile of the bootstrapped DTE at location $y$ and $z_p$ is the $p$-th quantile of the standard normal distribution.

Monte-Carlo Simulation Results

For our Monte Carlo simulation, we consider various setups to cover continuous/discrete and symmetric/skewed outcome distributions. We report our results from four types of data generating processes (DGPs): DGP1-DGP4. In all DGPs, we use the same setup for pre-treatment regressors $\bm{X}= (1, X_{1}, X_{2})^{\top}$ with a uniform random variable $X_{1} \sim U(0.5, 1.5)$ and a normal random variable $X_{2} \sim N(0,1)$. The treatment status $W_{1}$ follows the binomial distribution with $\Pr\{W_{1}=1\} = \pi_{1} \in \{0.3, 0.5\}$.

Under DGP1 and DGP2, we generate continuous outcomes from the following model:

align*[align* omitted — 91 chars of source]

where the unobserved component $U$ is specified as $ U \sim N(0,1)$ under DGP1 and $U \sim \chi_{3}^2 $ under DGP2, respectively. Consequently, the unobserved variable is symmetrically distributed under DGP1, while it exhibits a skewed distribution under DGP2. Under DGP3 and DGP4, we consider discrete outcome distributions satisfying the following conditional mean:

align*[align* omitted — 75 chars of source]

Given the conditional mean model above, we set the outcome distribution conditional on $\bm{X}$ and $\bm{W}$ to follow the Poisson distribution under DGP3 and to follow the negative binomial distribution with the dispersion parameter $r=5$ under DGP4, respectively.

figure[figure omitted — 610 chars of source]
figure[figure omitted — 567 chars of source]

We consider three estimators: the simple DTE (Simple) and the DTE with regression adjustment based on the linear regression (OLS) and logit model (Logit). To evaluate the finite sample performance of these DTE estimators, we compute their bias and root mean squared error (RMSE) across 10,000 simulations. Additionally, we assess the average length and coverage probability of their 95% pointwise confidence intervals, based on analytic standard errors derived from the sample analogs. In the main text, we present results for DGP1 and DGP3 under a treatment assignment probability of $\pi_{1} = 0.5$. Figures displaying results for all other combinations of DGPs and treatment assignment probabilities can be found in Supplemental Material (ref). Additional results using bootstrap confidence intervals and polynomial transformations of covariates are presented in Sections (ref) and (ref), respectively.

Figure (ref) summarizes the performance metrics for DGP1, while Figure (ref) presents results for DGP3 under a treatment assignment probability of $\pi_{1} = 0.5$. For continuous outcomes in DGP1 and DGP2, we consider thresholds $y$ at quantiles $\{0.1, \dots, 0.9\}$ of the true outcome distribution. For discrete outcomes in DGP3 and DGP4, we consider $y \in\{1, \dots, 5\}$. The results are displayed using boxplots over these threshold locations for a concise visual representation. To approximate true distributions, we generate datasets with 1,000,000 observations and compute the DTE at each threshold $y$.

As illustrated in Figure (ref), both the RMSE and the average length of the confidence intervals decrease as the sample size $n$ increases, as expected. The regression-adjusted estimators improve upon the simple estimator by reducing the RMSE and the length of confidence intervals by approximately 13%-25% across all sample sizes. However, for smaller sample sizes ($n \in \{300, 500\}$), the regression-adjusted estimators exhibit a slight undercoverage, with a coverage probability around 0.945. Similar trends are observed with the discrete outcome depicted in Figure (ref). In this scenario, the regression-adjusted estimators shorten the confidence intervals by about 9%-20% compared to the simple estimator.

In summary, these findings suggest that while regression adjustment does not affect the bias in DTE estimation, it significantly enhances precision. Additionally, we confirm that regression adjustment reduces the length of confidence intervals, thereby enhancing precision, while maintaining their 95% coverage probability.

Empirical Applications

Nudge and Water Consumption

In this section, we reanalyze data from a randomized experiment conducted by ferraro2013using in 2007 to examine the impact of norm-based messages, or nudges, on water usage in Cobb County, Atlanta, Georgia.\footnote{The dataset is publicly available as DVN1/22633_2013 at \href{{https://doi.org/10.7910/DVN1/22633}}{https://doi.org/10.7910/DVN1/22633}.} The experiment implemented and compared three different interventions aimed at reducing water usage against a control group (no nudge). For more details on the experiment and the results of the average treatment effect analysis, refer to ferraro2013using. For simplicity, following list2022using, we focus exclusively on the nudge that demonstrated the strongest average effect over the control group, which combined prosocial appeal with social comparison. We estimate the regression-adjusted DTE and PTE of this intervention over the control group, using the monthly water consumption in the year prior as pre-treatment covariates $\bm X$. The outcome variable $Y$ represents the level of water consumption from June to September 2007, measured in thousands of gallons and discretely distributed. It is important to note that, while the measure in gallons appears approximately continuous, the presence of subtle discreteness can complicate both theoretical and practical statistical inference. Consequently, QTE is not applicable in this context.

figure[figure omitted — 828 chars of source]

Figure (ref) presents the DTE and PTE of the intervention in comparison to the control group. We compute the DTE for $y\in\{0,1,2,\dots, 200\}$. The top left panel of Figure (ref) displays the simple estimate of the DTE, whereas the top right panel illustrates the regression-adjusted estimate of the DTE obtained with logit model. The shaded areas denote the 95% pointwise confidence intervals, based on analytic standard errors derived from the sample analogs. The results indicate that the regression-adjusted DTE has around 4% to 32% smaller standard errors and consequently tighter confidence intervals compared to the simple DTE.

The bottom left panel of Figure (ref) depicts the simple PTE, while the bottom right panel shows the regression-adjusted PTE, both with 95% pointwise confidence intervals. The PTE is presented in increments of $h=10$ for $y\in\{0, 10, 20,\dots, 200\}$. The findings suggest that the treatment is effective in reducing the probability of high water consumption and increasing the probability of low water consumption. Specifically, the regression-adjusted PTE results indicate an probability increase in water usage in the range of (20,30] and (30, 40] by 1.67 percentage points (pp) with standard errors (s.e.) of 0.42 pp and 1.31 pp (s.e. = 0.46 pp), respectively. On the other hand, it suggests a probability decrease in water usage in the range of (60,70], (70, 80], (90, 100] and (100, 110] by 0.61 pp (s.e. = 0.24 pp), 0.41 pp (s.e. = 0.20 pp), 0.35 pp (s.e. = 0.15 pp), and 0.26 pp (s.e. = 0.12 pp), respectively. The variance reduction is particularly notable in the range between 70 and 110. The probability change for the range (70,80], (90,100], and (100,110] are not significant under simple estimates, but are significantly negative under regression adjustment.\footnote{Some significance levels are difficult to discern in Figure (ref) due to the scale. The probability of decreased water consumption in the range $(80,90]$ is 0.33 pp with a standard error of 0.17 pp under regression-adjusted PTE, despite appearing significantly negative. Additionally, for the ranges $(90,100]$ and $(100,110]$, the probabilities of decreased water consumption are 0.28 pp (s.e. = 0.15 pp) and 0.22 pp (s.e. = 0.12 pp), respectively, under simple PTE.}

Insurance Coverage and Healthcare Utilization

In this subsection, we investigate the impact of insurance coverage on healthcare utilization using data from the Oregon Health Insurance Experiment.\footnote{The dataset is publicly available at \href{https://www.nber.org/research/data/oregon-health-insurance-experiment-data}{https://www.nber.org/research/data/oregon-health-insurance-experiment-data}.} In 2008, the state of Oregon offered insurance coverage to a group of uninsured low-income adults through a lottery. Twelve months after the lottery, a mail survey was conducted to collect data on various outcomes. The treatment in this experiment was randomly assigned based on the number of people in the household. Therefore, we focus on individuals who reported being in single-person households, excluding those who included one or two additional people in their lottery application. Refer to finkelstein2012oregon for more details about the experiment and the treatment effect results on various outcomes.

We examine healthcare utilization by focusing on the number of primary care (or doctor) visits over a six-month period, denoted as our outcome variable $Y$, collected from survey data. Our set of covariates $\bm X$ includes the applicant's age, gender, race, and income. Not everyone who was selected by the lottery to apply for insurance coverage ended up obtaining it. Previous research by finkelstein2012oregon has documented a positive intention-to-treat (ITT) effect and a local average treatment effect (LATE) of insurance coverage on the number of primary care visits. In this illustrative example, we use the lottery assignment (rather than actual insurance enrollment) as our treatment variable and focus solely on estimating the ITT effect. To obtain the local distributional treatment effect, one can use the representation provided in abadie2002bootstrap, applying the ITT effect at each location $y$.

figure[figure omitted — 889 chars of source]

Figure (ref) shows the distributional effect of insurance coverage on the number of primary care visits. The DTE estimates (top left) indicate a statistically significant effect on visits ranging from 0 to 6. The regression-adjusted DTE estimates obtained with logit model are nearly identical to the simple DTE estimates and reduce the standard errors by 0.1%-1.1% across values of $y{\in}\{0, 1, {\dots}, 30\}$. The gains from regression adjustment are modest in this case.

In the bottom panel of Figure (ref), we present the estimated PTE with 95% confidence intervals for each integer value between 0 and 15 (i.e., $h = 1$). The regression-adjusted PTE estimates suggest that insurance coverage reduces the probability of not seeing a doctor by approximately 6.56 pp (s.e. = 0.82 pp) and increases the probability of seeing a doctor three, four, five, and six times by 2.36 pp (s.e. = 0.61 pp), 0.66 pp (s.e. = 0.31 pp), 0.96 pp (s.e. =0.28 pp), and 0.81 pp (s.e. = 0.31 pp), respectively.

Conclusion

In this paper, we introduce a regression adjustment method to improve the precision of estimating distributional treatment effects through the use of distributional regression. We analyze the theoretical properties of our proposed method and provide a practical inference method that has been proven to be theoretically valid. Our simulation results demonstrate the remarkable performance of our method. Additionally, we apply our method to real-world data in two empirical applications and uncover heterogeneous treatment effects. In both applications, we find clearer conclusions compared to existing methods.

\setstretch{0.1}