EconBase
← Back to paper

Rectified Linear Unit Regression

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

52,990 characters · 15 sections · 64 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Rectified Linear Unit Regression

abstractThis paper develops a regression framework for the direct estimation of integrated functionals of conditional outcome distributions. The proposed method, termed rectified linear unit (ReLU) regression, projects the ReLU-transformed outcome onto covariates and admits a closed-form estimator. Its population regression function coincides with the integrated conditional distribution function of the outcome, and its convex conjugate, obtained via the Legendre-Fenchel transformation, recovers the integrated conditional quantile function. Both the regression and its conjugate require only mild distributional assumptions and accommodate non-continuous outcomes. We establish the uniform asymptotic distribution of the estimator and develop inference for the conjugate functional via the delta method for Hadamard directionally differentiable maps. Building on these results, we establish identification and inference for average quantile treatment effects over arbitrary subintervals of probability levels. This broadens the set of distributional parameters available to empirical work.

KEYWORDS: Rectified linear unit, Distributional treatment effect, Integrated quantile function, Legendre-Fenchel transform, Convex duality.

JEL classification codes: C12, C14 \setcounter{section}{0}

Introduction

Regression is a cornerstone of data analysis. While mean regression has long been the standard, modern approaches such as quantile regression koenker1978regression, expectile regression newey1987asymmetric, and distributional regression foresi1995conditional, chernozhukov2013inference offer a more comprehensive view of the underlying process. These methods capture distributional heterogeneity that the conditional mean alone overlooks. Recovering this heterogeneity is often what matters most for causal inference and policy evaluation.

This paper introduces an approach termed rectified linear unit (ReLU) regression, which regresses the ReLU transform $\max\{0, y - Y\}$ of the outcome $Y$ on covariates at each threshold $y$. The ReLU transformation was first explored in neural computation research hahnloser2000digital and later gained prominence in deep learning through the work of nair2010rectified on restricted Boltzmann machines. A primary advantage of this approach is that ReLU regression captures distributional features under mild conditions while admitting a closed-form estimator. Through the Legendre-Fenchel transformation, the regression yields a convex-duality representation that recovers the integrated quantile function. This convex structure allows for a direct characterization of distributional treatment effects.

The ReLU regression model offers three distinct advantages in analyzing distributional features. First, the model admits a closed-form estimator by formulating the estimation as an $L_2$ minimum distance problem with the ReLU-transformed dependent variable. Quantile regression koenker1978regression, expectile regression newey1987asymmetric, and distributional regression foresi1995conditional, by contrast, all require iterative optimization. The closed-form structure simplifies both computation and the derivation of asymptotic properties.

Next, the ReLU regression model accommodates less stringent conditions on the data-generating process. Specifically, the framework requires only a finite second moment of the outcome variable and standard rank conditions on the covariates. This contrasts with quantile regression, which requires a conditional density that is positive at the quantile of interest and additional smoothness conditions on this density koenker2005quantile, angrist2006quantile. The distinction matters most in the analysis of non-continuous outcomes, including discrete, count, and mixed discrete-continuous variables. To estimate the quantile function in such settings, machado2005quantiles introduce a jittering-based approach and chernozhukov2020generic develop generic distributional inference. ReLU regression instead targets an integrated functional, which smooths over the jumps and flat regions of non-continuous distributions.

Finally, our methodology supports a unified treatment of distributional treatment effects through convex duality. By estimating the integrated conditional distribution function across arbitrary locations, the ReLU regression model provides a natural complement to the conditional value-at-risk paradigm of rockafellar2000optimization, rockafellar2002conditional. Applying the Legendre-Fenchel transformation to the regression outcomes recovers the integrated conditional quantile function, which connects ReLU regression to the stochastic dominance principles of ogryczak2002dual and the coherent risk theory of acerbi2002coherence. The resulting framework quantifies average quantile treatment effects (AQTE) over arbitrary subintervals of quantile levels and includes both the average treatment effect (ATE) and the quantile treatment effect (QTE) as special cases, building on firpo2007efficient and chernozhukov2013inference.

Related Literature

Our framework connects to several strands of literature beyond those already mentioned.

First, our paper contributes to a growing literature on heterogeneous treatment effects in econometrics and statistics. The quantile treatment effect (QTE) was first introduced by doksum1974empirical and lehmann1975nonparametrics as a measure of treatment effects under unobserved heterogeneity. Since then, a large body of work has developed identification, estimation, and inference methods for distributional and quantile treatment effects, including heckman1997making, athey2006identification, bitler2006mean, djebbari2008heterogeneous, donald2014estimation, callaway2018quantile, callaway2019quantile, and firpo2007efficient, firpo2009unconditional, among others. For discrete outcomes, chernozhukov2020generic develop generic distributional inference for treatment effects. Our framework instead routes inference through the convex-duality structure of the integrated conditional distribution function and addresses non-continuous outcomes without smoothing or jittering the underlying distribution. To our knowledge, however, the average quantile treatment effect over an arbitrary subinterval has not been studied as a unified parameter that nests both the average treatment effect and the quantile treatment effect, and remains point-identified for non-continuous outcomes.

The second strand of literature studies the intersection of convex analysis and risk measurement, in which expected shortfall, also known as conditional value-at-risk (CVaR) or superquantiles, links risk measure theory and stochastic optimization. The foundations of coherent risk measures were established by artzner1999coherent, followed by the mathematical framework of rockafellar2000optimization, rockafellar2002conditional. pflug2000some developed convexity properties and computational approaches, while acerbi2002coherence proved the coherence of expected shortfall. follmer2016stochastic provide a comprehensive treatment of risk measures, including convex and monetary risk measures. The duality between conjugate functions and integrated quantile functions for unconditional distributions, established by ogryczak2002dual and rockafellar2000optimization, motivates our extension to the conditional case. The conventional two-step approach to estimating the integrated quantile function first estimates the quantile function pointwise and then integrates chen2025estimation, an approach that requires continuous outcome variables to ensure uniqueness of the quantile function. We instead estimate the integrated conditional distribution function directly via ReLU regression and apply the Legendre-Fenchel transformation, thereby extending the framework to non-continuous outcomes.

The third strand of literature concerns inference for functionals that are not fully Hadamard differentiable. The standard delta method requires full Hadamard differentiability of the functional, a condition that fails for the Legendre-Fenchel transformation. Foundational results on Hadamard directional differentiability are due to shapiro1990 and dumbgen1993. fang2019inference develop the delta method for Hadamard directionally differentiable maps and characterize the inconsistency of the standard nonparametric bootstrap in this setting. Inference for shape-constrained or directionally differentiable functionals also appears in delgado2012testing, beare2015transforming, chernozhukov2010estimation, and chen2021shape, among others. We apply this machinery to the conjugate functional in the ReLU regression framework.

Outline

Section (ref) introduces the ReLU regression model and its population interpretation through convex duality. Section (ref) develops estimation and the asymptotic distribution of the proposed estimator, with inference for the conjugate functional. Section (ref) applies the framework to causal inference. Section (ref) reports an empirical application to the Oregon Health Insurance Experiment, and Section (ref) concludes.

ReLU Regression

In this section, we introduce the ReLU regression model and develop its population interpretation under exogeneity of $X$.

Notation. We use $\|\cdot\|$ to denote the Euclidean norm for vectors and the spectral norm for matrices. The superscript $\top$ denotes the transpose, and $\mathbb{R}_{+} := [0, \infty)$. For an arbitrary index set $T$, $\ell^{\infty}(T)$ denotes the space of uniformly bounded real-valued functions on $T$. We also use $\1\{A\}$ to denote the indicator function, taking the value $1$ if event $A$ occurs and $0$ otherwise. We write $X_{n} \rightsquigarrow X$ for weak convergence of $X_{n}$ to $X$ in $\ell^{\infty}(T)$. For a proper convex function $f:\mathbb{R} \to \mathbb{R} \cup \{+\infty\}$, the subdifferential of $f(\cdot)$ at a point $z_{0}$ is defined as $\partial f(z_{0}) := \{\delta \in \mathbb{R} : f(z) \ge f(z_{0}) + \delta(z - z_{0}), \forall z \in \mathbb{R}\}$.

Data and ReLU Regression Model

Let $Y$ be a scalar outcome with support $\mathcal{Y} \subseteq \mathbb{R}$, and let $X \in \mathcal{X} \subseteq \mathbb{R}^{p}$ be a vector of covariates containing a constant term. The conditional distribution function of $Y$ given $X=x$ is defined as $F_{Y \mid X}(y | x) := \Pr\{Y \le y | X=x\}$ for $y \in \mathcal{Y}$, and the corresponding conditional quantile function is $F_{Y \mid X}^{-1}(u | x) := \inf\{y \in \mathcal{Y} : F_{Y \mid X}(y | x) \ge u\}$ for $u \in (0,1)$.

We introduce the ReLU regression model, whose key building block is the ReLU transformation $(a)_{+} := \max\{0,a\}$ for $a \in \mathbb{R}$. For each fixed $y \in \mathcal{Y}$, we treat the ReLU-transformed outcome $(y-Y)_{+}$ as the dependent variable and define $\beta_{0}(y)$ as the minimizer of the population criterion $Q(\, \cdot \,; y): \mathcal{B} \to \mathbb{R}_{+}$ over the parameter space $\mathcal{B} \subseteq \mathbb{R}^{p}$, given by

equation[equation omitted — 120 chars of source]

The ReLU regression model takes the form

equation[equation omitted — 90 chars of source]

where $\epsilon(y)$ is an error term, and both the coefficient $\beta_{0}(y)$ and the error $\epsilon(y)$ are indexed by $y \in \mathcal{Y}$. Throughout the paper, we adopt a specification linear in the covariates for simplicity. All subsequent analysis carries through when $X$ is replaced by any finite-dimensional vector of transformations, such as polynomials, splines, or interactions.

We impose the following regularity conditions on the joint distribution of $(X, Y)$.

assumptionThe data-generating process satisfies: \begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt] • $\mathbb{E}[Y^2] < \infty$ and $\mathbb{E}[\|X\|^{2}] < \infty$. • The matrix $\mathbb{E}[X X^{\top}]$ is positive definite. \end{enumerate}

Assumption (ref) is the standard regularity condition for the population $L^{2}$ problem. Assumption (ref)(a) ensures that the ReLU-transformed outcome $(y-Y)_{+}$ and any linear combination $X^{\top}\beta$ are square integrable for each fixed $y \in \mathcal{Y}$. Assumption (ref)(b) imposes the full-rank condition on the Gram matrix of $X$.

Under Assumption (ref), the criterion $Q(\beta; y)$ is finite and strictly convex in $\beta$ for each $y \in \mathcal{Y}$. Hence, $\beta_{0}(y)$ is the unique minimizer of $Q(\beta; y)$ and admits the closed-form expression

equation[equation omitted — 129 chars of source]

The linear combination $X^{\top}\beta_{0}(y)$ is the best linear predictor of $(y - Y)_{+}$ given $X$ in the $L^{2}$ sense. This characterization remains valid under misspecification, in the sense that the conditional expectation $\mathbb{E}[(y - Y)_{+} \mid X]$ is not necessarily linear in $X$. The interpretation is analogous to that of linear mean regression white1980using and linear quantile regression angrist2006quantile under misspecification. Proposition (ref) in Appendix A states this result formally.

Population Interpretation of ReLU Regression

We next develop the population interpretation of ReLU regression. Through convex duality, the ReLU transformation simultaneously encodes the conditional distribution function and the conditional quantile function of $Y$ given $X$.

For any $y \in \mathcal{Y}$, the integrated conditional distribution function of $Y$ given $X=x$ is defined as

equation*[equation* omitted — 86 chars of source]

The unconditional analogue has been studied in the risk-measurement literature follmer2016stochastic. The map $y \mapsto G_{Y \mid X}(y | x)$ is convex for each fixed $x \in \mathcal{X}$ since the integrand $F_{Y \mid X}(\cdot | x)$ is nondecreasing. Convexity allows us to apply the Legendre-Fenchel transformation. For each $\tau \in [0,1]$ and $x \in \mathcal{X}$, the conjugate function is defined as

equation[equation omitted — 147 chars of source]

The next proposition characterizes the convexity and the subdifferential structure of $G_{Y \mid X}(\cdot | x)$ and its conjugate $G_{Y \mid X}^{\ast}(\cdot | x)$. We write $F_{Y \mid X}(y\,\text{-} | x) := \lim_{s \nearrow y} F_{Y \mid X}(s | x)$ for the left limit of the conditional distribution function and $F_{Y \mid X}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}} | x) := \lim_{u \searrow \tau} F_{Y \mid X}^{-1}(u | x)$ for the right limit of the conditional quantile function.

propositionSuppose that $\mathbb{E}[|Y|] < \infty$. Then the following statements hold. \begin{enumerate}[label=(\alph*), itemsep=0.2em] • The map $y \mapsto G_{Y \mid X}(y | x)$ is convex, and for any $y \in \mathcal{Y}$, \begin{equation*} G_{Y \mid X}(y | x) = \mathbb{E}[(y-Y)_{+} | X=x], \quad a.s. \end{equation*} Moreover, $\partial G_{Y \mid X}(y | x) = [F_{Y \mid X}(y\,\text{-} | x),\, F_{Y \mid X}(y | x)]$ for any $y \in \mathcal{Y}$ and $x \in \mathcal{X}$. • The conjugate $\tau \mapsto G_{Y \mid X}^{\ast}(\tau | x)$ is convex, and for any $\tau \in (0,1)$ and $x \in \mathcal{X}$, \begin{equation*} G_{Y \mid X}^{\ast}(\tau | x) = \int_{0}^{\tau} F_{Y \mid X}^{-1}(u | x)\, du. \end{equation*} Correspondingly, $\partial G_{Y \mid X}^{\ast}(\tau | x) = [F_{Y \mid X}^{-1}(\tau | x),\, F_{Y \mid X}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}} | x)]$ for any $\tau \in (0,1)$ and $x \in \mathcal{X}$. \end{enumerate}

Proposition (ref) establishes the relationship between the ReLU transformation and the integrated quantile function through convex duality. The unconditional counterpart of this relationship is well known in the risk measurement literature ogryczak2002dual, rockafellar2000optimization, rockafellar2002conditional. The proof of the proposition is given in the appendix. The duality structure is summarized in Figure (ref).

figure[figure omitted — 1,644 chars of source]

Proposition (ref) has three implications. First, the result requires only the finite first-moment condition $\mathbb{E}[|Y|] < \infty$ and is agnostic about the existence or smoothness of conditional densities. Second, the integrated functions $G_{Y \mid X}(y | x)$ and $G_{Y \mid X}^{\ast}(\tau | x)$ are well-defined for arbitrary outcome distributions, including discrete and mixed discrete-continuous cases. For non-continuous outcomes the conditional quantile function lacks point identification, since any choice of generalized inverse is arbitrary on flat regions of $F_{Y \mid X}(\cdot | x)$. The integrated conditional quantile function avoids this difficulty entirely, since the choice of generalized inverse affects only a Lebesgue-null set and does not alter the integral. Third, the integrated functions fully characterize the underlying distribution. The conditional distribution function and the conditional quantile function are recovered as elements of the subdifferentials, $F_{Y \mid X}(y | x) \in \partial G_{Y \mid X}(y | x)$ and $F_{Y \mid X}^{-1}(\tau | x) \in \partial G_{Y \mid X}^{\ast}(\tau | x)$. For distributions with positive density, these subdifferentials reduce to singletons, whereas for non-smooth distributions they remain well-defined as set-valued maps.

Estimation and Asymptotic Properties

In this section, we present the ReLU regression estimator and derive its asymptotic distribution. We then develop inference for the integrated conditional quantile function.

Estimation Method

We observe an independent and identically distributed (i.i.d.) sample $\{(X_{i}, Y_{i})\}_{i=1}^{n}$ drawn from the joint distribution of $(X, Y)$. For each fixed $y \in \mathcal{Y}$, we define the objective function $\widehat{Q}(\, \cdot \,; y): \mathcal{B} \to \mathbb{R}_{+}$ as

equation*[equation* omitted — 130 chars of source]

We define the estimator $\hat{\beta}(y)$ as the minimizer of $\widehat{Q}(\beta; y)$ over the parameter space $\mathcal{B}$. Under Assumption (ref)(b), the sample second-moment matrix $n^{-1} \sum_{i=1}^{n} X_{i} X_{i}^{\top}$ is invertible with probability approaching one. The estimator then admits the closed-form expression

equation[equation omitted — 153 chars of source]

Given $\hat{\beta}(y)$, we estimate the integrated conditional distribution function $G_{Y \mid X}(y | X)$ by

equation*[equation* omitted — 80 chars of source]

Evaluating this estimator across $y \in \mathcal{Y}$ produces the collection $\{\widehat{G}_{Y \mid X}(y | X) : y \in \mathcal{Y}\}$, which characterizes the estimated conditional distribution of $Y$ given $X$.

We estimate the integrated conditional quantile function by applying the Legendre-Fenchel transformation to $\widehat{G}_{Y \mid X}(\cdot | x)$. For each fixed $x \in \mathcal{X}$ and $\tau \in [0, 1]$, the conjugate estimator is

equation*[equation* omitted — 146 chars of source]

The population map $y \mapsto G_{Y \mid X}(y | x)$ is convex, but its sample counterpart $y \mapsto \widehat{G}_{Y \mid X}(y | x)$ need not be. To restore convexity, we apply the Legendre-Fenchel transformation a second time. For each $x \in \mathcal{X}$ and $y \in \mathbb{R}$, the biconjugate is $ \widehat{G}_{Y \mid X}^{\ast\ast}(y | x) := \sup_{\tau \in [0,1]} \big\{ \tau y - \widehat{G}_{Y \mid X}^{\ast}(\tau | x) \big\}. $ The biconjugate coincides with the greatest convex minorant of $\widehat{G}_{Y \mid X}(\cdot | x)$ and is convex by construction.

Asymptotic Properties

We now derive the asymptotic distribution of our proposed estimator. To this end, we impose the following regularity conditions.

assumptionThe following conditions hold: \begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt] • $\{(X_{i}, Y_{i})\}_{i=1}^{n}$ is i.i.d.\ from the distribution of $(X, Y)$. • $\mathcal{Y}_{0}$ is a fixed compact subset of $\mathbb{R}$. • $\mathbb{E}[Y^{4}] < \infty$ and $\mathbb{E}[ \|X\|^{4} ] < \infty$. \end{enumerate}

Assumption (ref)(a) requires independent random sampling. Assumption (ref)(b) imposes only compactness of the index set $\mathcal{Y}_{0}$. We do not require $\mathcal{Y}_{0}$ to lie inside the support of $Y$. If $\mathcal{Y}_{0}$ lies strictly below the support of $Y$, then $(y - Y)_{+} = 0$ almost surely for every $y \in \mathcal{Y}_{0}$, and thus $\hat{\beta}(y) = \beta_{0}(y) = 0$ and the empirical process is identically zero. If $\mathcal{Y}_{0}$ extends beyond the upper end of the support of $Y$, the empirical process converges to a perfectly correlated Gaussian process in that region. In both boundary cases the conclusion of Theorem (ref) continues to hold, with a degenerate limit at points outside the support. Assumption (ref)(c) requires finite fourth moments, which ensures that the asymptotic variance is finite.

The next theorem establishes uniform weak convergence of $\hat{\beta}(\cdot)$ to a Gaussian process in $\ell^{\infty}(\mathcal{Y}_{0})^p$.

theoremSuppose that Assumptions (ref) and (ref) hold. Then \begin{equation*} \sqrt{n}\big( \hat{\beta}(\cdot) - \beta_{0}(\cdot) \big) \rightsquigarrow \mathbb{B}(\cdot) \quad in \ell^{\infty}(\mathcal{Y}_{0})^{p}, \end{equation*} and, for every $x \in \mathcal{X}$, \begin{equation*} \sqrt{n}\big( \widehat{G}_{Y \mid X}(\cdot | x) - G_{Y \mid X}(\cdot | x) \big) \rightsquigarrow x^{\top} \mathbb{B}(\cdot) \quad in \ell^{\infty}(\mathcal{Y}_{0}), \end{equation*} where $\mathbb{B}(\cdot)$ is a zero-mean Gaussian process with uniformly continuous sample paths and covariance function $Q_{X}^{-1}\, \Sigma(y_{1}, y_{2})\, Q_{X}^{-1}$ for $(y_{1}, y_{2}) \in \mathcal{Y}_{0}^{2}$, with $Q_{X} := \mathbb{E}[X X^{\top}]$ and $\Sigma(y_{1}, y_{2}) := \mathbb{E}\big[ X X^{\top}\, \epsilon(y_{1})\, \epsilon(y_{2}) \big]$.

Uniform continuity of the sample paths of $\mathbb{B}(\cdot)$ supports the construction of uniform confidence bands for $G_{Y \mid X}(\cdot | x)$ over $\mathcal{Y}_{0}$. The covariance function is the heteroskedasticity-robust covariance of the ReLU regression score.

The limiting distribution in Theorem (ref) depends on unknown nuisance parameters and is not pivotal. To obtain practical inference for $\widehat{G}_{Y \mid X}(y | x)$, we apply the exchangeable bootstrap praestgaard1993exchangeably, which consistently estimates the limit law of the linear functional $x^{\top} \mathbb{B}(\cdot)$. When the data are organized into independent clusters, the same bootstrap applies with weights drawn once per cluster and assigned to every observation in that cluster. Appendix C establishes consistency for the cluster case.

We next study the asymptotic behavior of the conjugate estimator $\widehat{G}_{Y \mid X}^{\ast}(\tau | x)$. The Legendre-Fenchel transformation is not fully Hadamard differentiable. Thus, the standard functional delta method of van1996weak does not apply. We work instead with Hadamard directional differentiability, which extends the delta method to maps that are differentiable only in selected directions shapiro1990, dumbgen1993, fang2019inference.

definitionLet $\mathbb{D}$ and $\mathbb{E}$ be normed spaces. A map $g: \mathbb{D} \to \mathbb{E}$ is Hadamard directionally differentiable at $\phi \in \mathbb{D}$ tangentially to $\mathbb{D}_{0} \subset \mathbb{D}$ if there exists a continuous map $g_{\phi}^{\prime}: \mathbb{D}_{0} \to \mathbb{E}$ such that \begin{equation*} \lim_{n \to \infty} \bigg\| \frac{g(\phi + t_{n} h_{n}) - g(\phi)}{t_{n}} - g_{\phi}^{\prime}(h) \bigg\|_{\mathbb{E}} = 0 \end{equation*} for all sequences $\{h_{n}\} \subset \mathbb{D}$ and $\{t_{n}\} \subset \mathbb{R}_{+}$ with $t_{n} > 0$, $h_{n} \to h \in \mathbb{D}_{0}$, and $t_{n} \downarrow 0$.

Let $\mathcal{L} : \ell^{\infty}(\mathcal{Y}_{0}) \to \ell^{\infty}([0, 1])$ denote the Legendre-Fenchel transformation defined by

equation*[equation* omitted — 172 chars of source]

For each fixed $\tau$, the supremum operation in the Legendre-Fenchel transformation is Hadamard directionally differentiable at $\phi$ tangentially to $C(\mathcal{Y}_{0})$ by Theorem 2.1 of carcamo2020, and the directional derivative $\mathcal{L}_{\phi}^{\prime} : C(\mathcal{Y}_{0}) \to \ell^{\infty}([0, 1])$ is defined as

equation*[equation* omitted — 134 chars of source]

where $A_{\epsilon}(\phi, \tau) := \{y \in \mathcal{Y}_{0} : \tau y - \phi(y) \ge \sup_{y' \in \mathcal{Y}_{0}}\{\tau y' - \phi(y')\} - \epsilon\}$.

The following proposition gives the asymptotic distribution of the conjugate estimator via the delta method for Hadamard directionally differentiable maps fang2019inference.

propositionSuppose that Assumptions (ref) and (ref) hold. Then, for every $x \in \mathcal{X}$, \begin{equation*} \sqrt{n}\big( \mathcal{L}\widehat{G}_{Y \mid X}(\cdot | x) - \mathcal{L} G_{Y \mid X}(\cdot | x) \big) \rightsquigarrow \mathcal{L}_{G_{Y \mid X}(\cdot | x)}^{\prime}\big( x^{\top} \mathbb{B} \big) \quad in \ell^{\infty}([0, 1]), \end{equation*} where $\mathbb{B}(\cdot)$ is the Gaussian process defined in Theorem (ref).

The limit in Proposition (ref) depends on $\mathcal{L}_{G_{Y \mid X}(\cdot | x)}^{\prime}$ and is therefore non-pivotal. Thus, the standard nonparametric bootstrap is inconsistent for the limit law dumbgen1993, fang2019inference. We adopt the resampling approach of fang2019inference for Hadamard directionally differentiable maps. Alternative approaches for inference on shape-constrained functions include delgado2012testing and beare2015transforming, among others.

The argmax set in the conjugate operation coincides with the subdifferential $\partial G_{Y \mid X}^{\ast}(\tau | x) = [F_{Y \mid X}^{-1}(\tau | x),\, F_{Y \mid X}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}} | x)]$ from Proposition (ref)(b). Estimating this set by the strict empirical argmax fails when the subdifferential is a non-degenerate interval, since small sampling fluctuations can place the empirical argmax at any point in the interval. The fattened-argmax device used in chernozhukov2013intersection and cattaneo2020bootstrap addresses this issue by replacing the strict argmax with the level set

equation*[equation* omitted — 273 chars of source]

for a vanishing tuning sequence $\nu_{n}$. The plug-in directional derivative $-\inf\{h(y) : y \in \widehat{\partial} G_{Y \mid X}^{\ast}(\tau | x; \nu_{n})\}$ recovers the standard nonparametric bootstrap when the subdifferential is a singleton. It captures the worst-case Hadamard directional derivative when the subdifferential is a non-degenerate interval. In both cases the procedure exploits the convex structure of the conjugate operation.

Treatment Effect

Setup and the Treatment Effects

We extend the ReLU regression to evaluate distributional treatment effects in a binary treatment setting under the potential outcomes framework neyman1923applications, rubin1974estimating. Let $W \in \{0, 1\}$ denote the binary treatment indicator, with $W = 1$ for treated units and $W = 0$ for control units, and let $Y(0)$ and $Y(1)$ denote the corresponding potential outcomes. Under the stable unit treatment value assumption rubin1980randomization, the observed outcome is $Y = Y(W)$. For each $w \in \{0, 1\}$, let $F_{Y(w)}(y) := \Pr\{Y(w) \le y\}$ denote the distribution function of $Y(w)$, and let $F_{Y(w)}^{-1}(u) := \inf\{y \in \mathbb{R} : F_{Y(w)}(y) \ge u\}$ denote the corresponding quantile function for $u \in (0, 1)$.

For probability levels $\tau_{\ell}, \tau_{u} \in [0, 1]$ with $\tau_{\ell} < \tau_{u}$, we define the average quantile treatment effect (AQTE) over $[\tau_{\ell}, \tau_{u}]$ as

equation*[equation* omitted — 221 chars of source]

The AQTE measures the difference between the quantile functions of the potential outcomes, averaged over the interval $[\tau_{\ell}, \tau_{u}]$, thereby capturing heterogeneous treatment effects across the outcome distributions. Moreover, the integral formulation renders the AQTE robust to set-valuedness of the quantile functions, since such irregular points form a set of Lebesgue measure zero and do not alter the integral.

The AQTE encompasses well-known treatment effect parameters as special cases, as established in the following proposition.

propositionSuppose $\mathbb{E}[|Y(w)|] < \infty$ for each $w \in \{0, 1\}$. Then \begin{enumerate}[label=(\alph*), topsep=3pt, itemsep=3pt, align=left, leftmargin=1em, labelsep=0.5em] • $\theta(0, 1) = \mathbb{E}[Y(1) - Y(0)]$; • for any $\tau \in (0, 1)$, the limit point of $\theta(\tau_{\ell}, \tau_{u})$ as $(\tau_{\ell}, \tau_{u}) \to (\tau, \tau)$ with $\tau_{\ell} < \tau_{u}$ equals $F_{Y(1)}^{-1}(\tau) - F_{Y(0)}^{-1}(\tau)$ if $F_{Y(w)}^{-1}(\cdot)$ is continuous at $\tau$ for each $w \in \{0, 1\}$, and in general, every limit point of $\theta(\tau_{\ell}, \tau_{u} )$ belongs to the interval $[F_{Y(1)}^{-1}(\tau) - F_{Y(0)}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}}),\, F_{Y(1)}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}}) - F_{Y(0)}^{-1}(\tau)]$. \end{enumerate}

Proposition (ref) shows that the AQTE nests two canonical treatment effect parameters. The AQTE recovers the average treatment effect by setting $(\tau_{\ell}, \tau_{u}) = (0, 1)$. It also recovers the $\tau$-th quantile treatment effect in the limit as $(\tau_{\ell}, \tau_{u}) \to (\tau, \tau)$, provided the quantile functions of both potential outcomes are continuous at $\tau$.

Identification

We consider the identification of the AQTE through the integrated distribution and quantile functions of the potential outcomes, defined for each $w \in \{0,1\}$ as

equation*[equation* omitted — 175 chars of source]

By Proposition (ref)(b), the integrated quantile function $G_{Y(w)}^{\ast} (\cdot)$ is the convex conjugate of $G_{Y(w)} (\cdot)$. It therefore suffices to identify $G_{Y(w)}$ for the identification of the AQTE.

We consider identification under random assignment, as in randomized experiments. Identification under unconfoundedness conditional on pre-treatment covariates is developed in Appendix B.

assumptionThe data-generating process satisfies: \begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt] • $Y = Y(W)$ almost surely. • $\mathbb{E}[|Y(w)|] < \infty$ for each $w \in \{0, 1\}$. • $(Y(0), Y(1)) \protect\mathpalette{\protect\independenT}{\perp} W$ and $0 < \Pr(W = 1) < 1$. \end{enumerate}

Assumption (ref) collects the standard conditions of the potential outcomes framework. Condition (a) is the consistency condition that relates the observed outcome to the realized potential outcome. Condition (b) ensures that the integrated distribution function $G_{Y(w)}$ is finite for each $w \in \{0, 1\}$. Condition (c) combines random assignment of $W$ with the overlap condition $0 < \Pr(W = 1) < 1$, which ensures that both treatment arms are observable.

We now state the identification result for the AQTE under random assignment.

theoremSuppose Assumption (ref) holds. Then, $G_{Y(w)}(y) = \mathbb{E}[(y - Y)_{+} \mid W = w]$ for each $w \in \{0, 1\}$ and $y \in \mathbb{R}$, and for any $\tau_{\ell}, \tau_{u} \in [0, 1]$ with $\tau_{\ell} < \tau_{u}$, \begin{equation*} \theta(\tau_{\ell}, \tau_{u}) = \frac{1}{\tau_{u} - \tau_{\ell}} \big \{ \big ( G_{Y(1)}^{\ast}(\tau_u) - G_{Y(1)}^{\ast}(\tau_\ell)\big) - \big ( G_{Y(0)}^{\ast}(\tau_u) - G_{Y(0)}^{\ast}(\tau_\ell)\big) \big \}. \end{equation*}

The theorem establishes identification of the AQTE through the Legendre--Fenchel transformation of the integrated distribution function $G_{Y(w)}(\cdot)$, which is itself identified from the conditional distribution of $Y$ given $W = w$ under random assignment, for each $w \in \{0, 1\}$. The AQTE is point-identified without imposing continuity of the potential outcome distributions or any further regularity conditions, yet it captures heterogeneous treatment effects across the outcome distributions through the choice of the quantile interval $[\tau_{\ell}, \tau_{u}]$.

Estimation and Asymptotic Properties

Suppose that we observe an i.i.d.\ sample $\{(W_{i}, Y_{i})\}_{i=1}^{n}$ drawn from the joint distribution of $(W, Y)$. For the estimation of the AQTE, we consider the ReLU regression with the regressor $X = (1, W)^{\top}$ and use the estimator $\hat{\beta}(y) \in \mathbb{R}^{2}$ from (ref). For each $w \in \{0, 1\}$, the integrated distribution function $G_{Y(w)}(y)$ is estimated by $\widehat{G}_{Y(w)}(y)$, the fitted value at $W = w$, and its convex conjugate $\widehat{G}_{Y(w)}^{\ast}$ is obtained by the Legendre--Fenchel transformation. For quantile indices $\tau_{\ell}, \tau_{u} \in [0, 1]$ with $\tau_{\ell} < \tau_{u}$, the AQTE estimator is

equation[equation omitted — 337 chars of source]

We derive the asymptotic distribution of $\widehat{\theta}(\tau_{\ell}, \tau_{u})$ under the following regularity condition.

assumptionThe following conditions hold: \begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt] • $\{(W_{i}, Y_{i})\}_{i=1}^{n}$ is i.i.d.\ from the joint distribution of $(W, Y)$. • $\mathbb{E}[Y(w)^{2}] < \infty$ for each $w \in \{0, 1\}$. \end{enumerate}

Assumption (ref)(a) imposes independent random sampling. Although Theorem (ref) is stated under the fourth-moment Assumption (ref)(c), the regressor $X=(1,W)^{\top}$ is bounded. A square-integrable envelope therefore holds under the second-moment condition in Assumption (ref)(b), and the conclusion of Theorem (ref) follows by the same argument. The resulting weak limit is a zero-mean Gaussian process $\mathbb{B}_{w}(\cdot)$ in $\ell^{\infty}(\mathcal{Y}_{0})$, for any compact $\mathcal{Y}_{0} \subset \mathbb{R}$ and for each $w \in \{0,1\}$, with covariance kernel

equation*[equation* omitted — 173 chars of source]

and $\mathbb{B}_{0}, \mathbb{B}_{1}$ are independent.

theoremSuppose Assumptions (ref) and (ref) hold, and let $\tau_{\ell}, \tau_{u} \in (0, 1)$ with $\tau_{\ell} < \tau_{u}$. Then we have \begin{equation*} \sqrt{n}\big(\widehat{\theta}(\tau_{\ell}, \tau_{u}) - \theta(\tau_{\ell}, \tau_{u})\big) \rightsquigarrow \dfrac{1}{\tau_{u} - \tau_{\ell}} \big\{ \big( Z_{1}(\tau_{u}) - Z_{1}(\tau_{\ell}) \big) - \big( Z_{0}(\tau_{u}) - Z_{0}(\tau_{\ell}) \big) \big\}, \end{equation*} where $ Z_{w}(\tau) := \sup_{y \in \partial G_{Y(w)}^{\ast}(\tau)} \big\{ -\mathbb{B}_{w}(y) \big\} $ for each $w \in \{0, 1\}$ and $\tau \in \{\tau_{\ell}, \tau_{u}\}$. The delta-method bootstrap based on the subdifferential estimator of $\partial G_{Y(w)}^{\ast}(\tau)$ for $w \in \{0,1\}$ consistently estimates the limit law above.

Theorem (ref) characterizes the limit law as a linear combination of the Hadamard directional derivatives of the Legendre-Fenchel transformation evaluated at $G_{Y(0)}$ and $G_{Y(1)}$. At points $\tau$ where $F_{Y(w)}^{-1}$ is continuous, the subdifferential $\partial G_{Y(w)}^{\ast}(\tau)$ collapses to a singleton and $Z_{w}(\tau) = -\mathbb{B}_{w}(F_{Y(w)}^{-1}(\tau))$, and the limit law is a centered Gaussian. At points where $F_{Y(w)}^{-1}$ is discontinuous at $\tau$, $\partial G_{Y(w)}^{\ast}(\tau)$ has positive length, and $Z_{w}(\tau)$ is the supremum of a Gaussian process over that subdifferential. To ensure valid bootstrap inference, the delta-method bootstrap is required.

Theorem (ref) states the asymptotic distribution of the AQTE estimator under the regressor $X = (1, W)^{\top}$. The construction extends to a regressor of the form $X = (1, W, V^{\top})^{\top}$ under the conditional random-assignment condition $(Y(0), Y(1)) \protect\mathpalette{\protect\independenT}{\perp} W \mid V$. The marginal integrated distribution function of each potential outcome $G_{Y(w)}(\cdot)$ is recovered from the estimator of $G_{Y(w)|V}(\cdot)$ by integrating $V$ out. The conclusions of Theorem (ref) continue to hold with the corresponding Gaussian processes and covariance kernels.

Application

In this empirical application, we examine the effect of public health insurance on health-care utilization. The 2008 expansion of the Oregon Medicaid program offered enrollment to a randomly selected subset of low-income uninsured adults from a waiting list. The resulting Oregon Health Insurance Experiment is a randomized evaluation of the distributional effects of insurance coverage finkelstein2012oregon. The original analysis of finkelstein2012oregon reports intention-to-treat estimates from a linear regression of the outcome on the lottery offer. chernozhukov2020generic revisit the experiment and report quantile treatment effects for the count outcome through quantile bands.

We consider $W = 1$ for individuals offered the lottery and $W = 0$ otherwise. The outcome $Y$ is the number of outpatient visits in the six-month period preceding the survey. Randomization was at the household level, and the offer probability was constant within strata defined by the survey wave and the listed household size. We collect the design features in the vector $V$, which records six survey-wave dummies, two household-size dummies, and their interactions. Conditional on $V$, the assignment is independent of the potential outcomes, $(Y(0), Y(1)) \protect\mathpalette{\protect\independenT}{\perp} W| V$. The marginal distribution of $Y(w)$ is therefore identified by averaging the conditional distribution of $Y$ given $(W = w, V)$ over the marginal distribution of $V$. The estimand is the average quantile treatment effect $\theta(\tau_{\ell}, \tau_{u})$ of Section (ref). It has the intention-to-treat interpretation as the design-averaged quantile difference between the offered and not-offered populations on the subinterval $[\tau_{\ell}, \tau_{u}]$.

The count nature of $Y$ rules out the linear-conditional-quantile assumption that underlies linear quantile regression. It also induces nonsmooth behavior of the quantile-as-inverse map. This nonsmoothness blocks the standard delta method. chernozhukov2020generic address the second issue by inverting a distribution-regression confidence band and taking the Minkowski difference of the two quantile bands. Our approach estimates the integrated distribution function for each treatment arm by ordinary least squares. We integrate out the design controls through their empirical distribution. We then recover the integrated quantile function as the convex conjugate. The estimator is closed-form. The average quantile treatment effect is read off as a difference of conjugates evaluated at two probability levels. Inference uses the delta-method bootstrap of Section (ref) based on the empirical subdifferential interval and requires no numerical differentiation.

Our data come from the public release of the Oregon Health Insurance Experiment finkelstein2012oregon. The sample consists of survey respondents on the lottery list, with the binary indicator $W$ recording whether the individual was drawn in the lottery. The outcome $Y$ is right-skewed and concentrated on small integers. It has a substantial mass at zero and a sparse upper tail extending to several dozen visits. The empirical distributions of $Y$ for the two treatment arms are reported in Figure (ref).

We use the linear specification with regressor $X_{i} := (1, W_{i}, V_{i}^{\top})^{\top}$ following the baseline of finkelstein2012oregon. We evaluate the estimator on the grid $\mathcal{Y}_{0} := \{0, 1, 2, \dots, y_{\max}\}$, where $y_{\max}$ is the largest observed value of $Y$. Letting $\widehat{G}_{Y(w) \mid V}(y \mid v)$ be the estimator of $G_{Y(w) \mid V}(y \mid v)$, we obtain the integrated distribution function for each treatment through the marginalization: $\widehat{G}_{Y(w)}(y) = n^{-1} \sum_{i=1}^{n} \widehat{G}_{Y(w)|V}(y \mid V_i)$ for each $w \in \{0,1\}$. The corresponding integrated quantile function estimator $ \widehat{G}_{Y(w)}^{\ast}(\tau)$ is obtained by the Legendre-Fenchel transform.

We report the average quantile treatment effect $\widehat{\theta}(\tau, \tau + 0.10)$ on the grid $\tau \in \{0, 0.1, \ldots, 0.9\}$. Inference uses the delta-method bootstrap described in Section (ref). We draw the bootstrap weights at the household level following Appendix C. Sampling weights from the public release enter the ReLU regression as observation weights.

figure[figure omitted — 608 chars of source]

Figure (ref) shows that the outcome distribution is concentrated on a small number of integer values and shares the same support across the two treatment arms. The point mass at zero is visible in both panels and is smaller in the lottery-winner panel than in the lottery-loser panel. The mass at zero and the sparse upper tail motivate working with the integrated distribution function rather than with the quantile function directly.

figure[figure omitted — 569 chars of source]

Figure (ref) reports the estimated integrated quantile functions $\widehat{G}_{Y(w)}^{\ast}$ for the two treatment arms. Both curves are nondecreasing and convex by construction. At points of differentiability, the slope at a probability level $\tau$ identifies the marginal quantile $F_{Y(w)}^{-1}(\tau)$. Both curves are flat over the lower portion of the unit interval, reflecting the point mass at zero in the marginal distribution of $Y(w)$, and start to rise around $\tau \approx 0.35$. The treatment-arm curve sits uniformly above the control-arm curve over the upper portion. The change in vertical separation between probability levels $\tau_{\ell}$ and $\tau_{u}$, divided by $\tau_{u} - \tau_{\ell}$, equals the average quantile treatment effect over $[\tau_{\ell}, \tau_{u}]$.

Figure (ref) reports the estimated average quantile treatment effect across the probability range. The estimates are near zero on the three lowest subintervals, mirroring the flat region of Figure (ref) below $\tau \approx 0.35$. They are positive on every subinterval above $\tau = 0.3$ and are larger at higher quantiles on average. The offer of insurance therefore affects the upper tail of utilization more strongly than the lower tail. The confidence band widens with $\tau$ and is widest on the topmost subinterval, reflecting the sparse upper tail of the outcome distribution.

figure[figure omitted — 663 chars of source]

The pattern of the estimates carries an economic interpretation. The flat lower portion mirrors the large share of respondents with zero outpatient visits over the six-month window, whose utilization does not change with the insurance offer. Above $\tau \approx 0.3$, the offer raises utilization at every probability level. The effect is largest in the upper tail, where heavier users of medical care reside. A standard average treatment effect would aggregate the zero effect on non-utilizers and the larger effect on heavy users into a single scalar and would conceal this distributional pattern. The AQTE locates the gain on the relevant subintervals of the utilization distribution and attaches a quantitative magnitude to each. The resulting estimates are the input required for a distributional welfare analysis of insurance coverage.

Conclusion

This paper develops a regression model with the rectified linear unit transformation applied to the outcome variable. The ReLU regression estimates the integrated conditional distribution function in closed form. Through the Legendre-Fenchel transformation, its convex conjugate equals the integrated conditional quantile function. The framework requires only finite moments and standard rank conditions, and it accommodates outcome distributions that are discrete, mixed, or otherwise non-continuous without modification.

We establish the uniform asymptotic distribution of the estimator as a process indexed by the outcome location, and we develop inference for the conjugate functional through the delta method for Hadamard directionally differentiable maps. As an application, the framework identifies average quantile treatment effects over arbitrary quantile subintervals under random assignment of a binary treatment, with bootstrap inference based on the same delta-method procedure.

\setstretch{0.1}

\setstretch{1.2} \setcounter{section}{0} \setcounter{figure}{0} \setcounter{table}{0} \setcounter{equation}{0} \setcounter{lemma}{0}\setcounter{page}{1} \setcounter{proposition}{0}

center[center omitted — 0 chars of source]