The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
52,990 characters
Rectified Linear Unit Regression
\maketitle
\begin{abstract}
This paper develops a regression framework for the direct estimation of integrated functionals of conditional outcome distributions. The proposed method, termed rectified linear unit (ReLU) regression, projects the ReLU-transformed outcome onto covariates and admits a closed-form estimator. Its population regression function coincides with the integrated conditional distribution function of the outcome, and its convex conjugate, obtained via the Legendre-Fenchel transformation, recovers the integrated conditional quantile function.
Both the regression and its conjugate require only mild distributional assumptions and accommodate non-continuous outcomes.
We establish the uniform asymptotic distribution of the estimator and develop inference for the conjugate functional via the delta method for Hadamard directionally differentiable maps.
Building on these results, we establish identification and inference for average quantile treatment effects over arbitrary subintervals of probability levels.
This broadens the set of distributional parameters available to empirical work.
\end{abstract}
\vspace{0.5cm}
\noindent KEYWORDS: Rectified linear unit, Distributional treatment
effect, Integrated quantile function, Legendre-Fenchel transform,
Convex duality.
\noindent JEL classification codes: C12, C14
\setcounter{section}{0}
\section{Introduction}
Regression is a cornerstone of data analysis. While mean regression has long been the standard, modern approaches such as quantile regression \citep{koenker1978regression}, expectile regression \citep{newey1987asymmetric}, and distributional regression \citep{foresi1995conditional, chernozhukov2013inference} offer a more comprehensive view of the underlying process.
These methods capture distributional heterogeneity that the conditional mean alone overlooks.
Recovering this heterogeneity is often what matters most for causal inference and policy evaluation.
This paper introduces an approach termed rectified linear unit (ReLU) regression, which regresses the ReLU transform $\max\{0, y - Y\}$ of the outcome $Y$ on covariates at each threshold $y$.
The ReLU transformation was first explored in neural computation research \citep{hahnloser2000digital} and later gained prominence in deep learning through the work of \cite{nair2010rectified} on restricted Boltzmann machines.
A primary advantage of this approach is that ReLU regression captures distributional features under mild conditions while admitting a closed-form estimator.
Through the Legendre-Fenchel transformation, the regression yields a convex-duality representation that recovers the integrated quantile function. This convex structure allows for a direct characterization of distributional treatment effects.
The ReLU regression model offers three distinct advantages in analyzing distributional features.
First, the model admits a closed-form estimator by formulating the estimation as an $L_2$ minimum distance problem with the ReLU-transformed dependent variable.
Quantile regression \citep{koenker1978regression}, expectile regression \citep{newey1987asymmetric}, and distributional regression \citep{foresi1995conditional}, by contrast, all require iterative optimization. The closed-form structure simplifies both computation and the derivation of asymptotic properties.
Next, the ReLU regression model accommodates less stringent
conditions on the data-generating process. Specifically, the
framework requires only a finite second moment of the outcome
variable and standard rank conditions on the covariates. This
contrasts with quantile regression, which
requires a conditional density that is positive at the quantile of interest and additional smoothness conditions on this density \citep[see, e.g.,][]{koenker2005quantile, angrist2006quantile}. The distinction
matters most in the analysis of non-continuous outcomes, including
discrete, count, and mixed discrete-continuous variables.
To estimate the quantile function in such settings, \cite{machado2005quantiles} introduce a jittering-based approach and \cite{chernozhukov2020generic} develop generic distributional inference. ReLU regression instead targets an integrated functional, which smooths over the jumps and flat regions of non-continuous distributions.
Finally, our methodology supports a unified treatment of
distributional treatment effects through convex duality. By
estimating the integrated conditional distribution function across
arbitrary locations, the ReLU regression model provides a natural
complement to the conditional value-at-risk paradigm of
\cite{rockafellar2000optimization, rockafellar2002conditional}.
Applying the Legendre-Fenchel transformation to the regression
outcomes recovers the integrated conditional quantile function,
which connects ReLU regression to the stochastic dominance
principles of \cite{ogryczak2002dual} and the coherent risk theory
of \cite{acerbi2002coherence}. The resulting framework quantifies
average quantile treatment effects (AQTE)
over arbitrary subintervals of quantile levels
and includes both the average treatment effect (ATE) and the
quantile treatment effect (QTE) as special cases, building on
\cite{firpo2007efficient} and \cite{chernozhukov2013inference}.
\subsection{Related Literature}
Our framework connects to several strands of literature beyond
those already mentioned.
First, our paper contributes to a growing literature on
heterogeneous treatment effects in econometrics and statistics.
The quantile treatment effect (QTE) was first introduced by
\cite{doksum1974empirical} and \cite{lehmann1975nonparametrics} as
a measure of treatment effects under unobserved heterogeneity.
Since then, a large body of work has developed identification,
estimation, and inference methods for distributional and quantile
treatment effects, including \cite{heckman1997making},
\cite{athey2006identification}, \cite{bitler2006mean},
\cite{djebbari2008heterogeneous}, \cite{donald2014estimation},
\cite{callaway2018quantile, callaway2019quantile}, and
\cite{firpo2007efficient, firpo2009unconditional}, among others.
For discrete outcomes, \cite{chernozhukov2020generic} develop
generic distributional inference for treatment effects. Our
framework instead routes inference through the convex-duality
structure of the integrated conditional distribution function and
addresses non-continuous outcomes without smoothing or jittering
the underlying distribution. To our knowledge, however, the
average quantile treatment effect over an arbitrary subinterval
has not been studied as a unified parameter that nests both the
average treatment effect and the quantile treatment effect,
and remains point-identified for non-continuous outcomes.
The second strand of literature studies the intersection of convex
analysis and risk measurement, in which expected shortfall, also
known as conditional value-at-risk (CVaR) or superquantiles, links
risk measure theory and stochastic optimization. The foundations
of coherent risk measures were established by
\cite{artzner1999coherent}, followed by the mathematical framework
of \cite{rockafellar2000optimization, rockafellar2002conditional}.
\cite{pflug2000some} developed convexity properties and
computational approaches, while \cite{acerbi2002coherence}
proved the coherence of expected shortfall.
\cite{follmer2016stochastic} provide
a comprehensive treatment of risk measures, including convex and monetary risk measures. The
duality between conjugate functions and integrated quantile
functions for unconditional distributions, established by
\cite{ogryczak2002dual} and \cite{rockafellar2000optimization},
motivates our extension to the conditional case. The conventional
two-step approach to estimating the integrated quantile function
first estimates the quantile function pointwise and then integrates
\citep[see, e.g.,][]{chen2025estimation}, an approach that
requires continuous outcome variables to ensure uniqueness of the
quantile function. We instead estimate the integrated conditional
distribution function directly via ReLU regression and apply the
Legendre-Fenchel transformation, thereby extending the framework
to non-continuous outcomes.
The third strand of literature concerns inference for functionals
that are not fully Hadamard differentiable. The standard delta
method requires full Hadamard differentiability of the functional,
a condition that fails for the Legendre-Fenchel transformation.
Foundational results on Hadamard directional differentiability are
due to \cite{shapiro1990} and \cite{dumbgen1993}.
\cite{fang2019inference} develop the delta method for Hadamard
directionally differentiable maps and characterize the
inconsistency of the standard nonparametric bootstrap in this
setting. Inference for shape-constrained or
directionally differentiable functionals also appears in
\cite{delgado2012testing}, \cite{beare2015transforming},
\cite{chernozhukov2010estimation}, and \cite{chen2021shape}, among
others. We apply this machinery to the conjugate functional in the
ReLU regression framework.
\subsection{Outline}
Section \ref{sec:relu} introduces the ReLU regression model and
its population interpretation through convex duality. Section
\ref{sec:estimation} develops estimation and the asymptotic
distribution of the proposed estimator, with inference for the conjugate
functional. Section \ref{sec:treatment} applies the framework to causal inference.
Section \ref{sec:application} reports an empirical application to
the Oregon Health Insurance Experiment, and Section
\ref{sec:conclusion} concludes.
\section{ReLU Regression}
\label{sec:relu}
In this section, we introduce the ReLU regression model and develop
its population interpretation under exogeneity of $X$.
\textit{Notation}.
We use $\|\cdot\|$ to denote the Euclidean norm for vectors and the
spectral norm for matrices. The superscript $\top$ denotes the
transpose, and $\mathbb{R}_{+} := [0, \infty)$. For an arbitrary index set
$T$, $\ell^{\infty}(T)$ denotes the space of uniformly bounded
real-valued functions on $T$.
We also use $\1\{A\}$ to denote the indicator function, taking the value $1$ if event $A$ occurs and $0$ otherwise.
We write $X_{n} \rightsquigarrow X$ for
weak convergence of $X_{n}$ to $X$ in $\ell^{\infty}(T)$.
For a proper convex function $f:\mathbb{R} \to \mathbb{R} \cup \{+\infty\}$, the subdifferential of $f(\cdot)$ at a point $z_{0}$ is defined as
$\partial f(z_{0}) := \{\delta \in \mathbb{R} : f(z) \ge f(z_{0}) + \delta(z - z_{0}), \forall z \in \mathbb{R}\}$.
\subsection{Data and ReLU Regression Model}
Let $Y$ be a scalar outcome with support $\mathcal{Y} \subseteq \mathbb{R}$, and let
$X \in \mathcal{X} \subseteq \mathbb{R}^{p}$ be a vector of covariates
containing a constant term.
The conditional distribution function of $Y$ given $X=x$
is defined as
$F_{Y \mid X}(y | x) := \Pr\{Y \le y | X=x\}$
for
$y \in \mathcal{Y}$,
and
the corresponding conditional quantile function is
$F_{Y \mid X}^{-1}(u | x) := \inf\{y \in \mathcal{Y} : F_{Y \mid X}(y | x) \ge u\}$
for $u \in (0,1)$.
We introduce the ReLU regression model,
whose key building block
is the ReLU transformation
$(a)_{+} := \max\{0,a\}$ for $a \in \mathbb{R}$.
For each fixed $y \in \mathcal{Y}$,
we treat the ReLU-transformed outcome
$(y-Y)_{+}$
as the dependent variable
and
define $\beta_{0}(y)$ as the minimizer of
the population criterion
$Q(\, \cdot \,; y): \mathcal{B} \to \mathbb{R}_{+}$ over the parameter space
$\mathcal{B} \subseteq \mathbb{R}^{p}$,
given by
\begin{equation}
\label{eq:bp-pop}
Q(\beta; y) := \mathbb{E}\big[ \big( (y - Y)_{+} - X^{\top} \beta \big)^{2} \big].
\end{equation}
The ReLU regression model takes the form
\begin{equation}
\label{eq:relu-reg}
(y - Y)_{+} = X^{\top} \beta_{0}(y) + \epsilon(y),
\end{equation}
where
$\epsilon(y)$ is an error term,
and
both the coefficient $\beta_{0}(y)$
and
the error $\epsilon(y)$ are indexed by $y \in \mathcal{Y}$.
Throughout the paper, we adopt a specification linear in the covariates for simplicity. All subsequent analysis carries through when $X$ is replaced by any finite-dimensional vector of transformations, such as polynomials, splines, or interactions.
We impose the following regularity conditions on the joint
distribution of $(X, Y)$.
\vspace{0.2cm}
\begin{assumption}
\label{assump:dgp}
The data-generating process satisfies:
\begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt]
\item[\textnormal{(a)}]
$\mathbb{E}[Y^2] < \infty$ and $\mathbb{E}[\|X\|^{2}] < \infty$.
\item[\textnormal{(b)}]
The matrix $\mathbb{E}[X X^{\top}]$ is positive definite.
\end{enumerate}
\end{assumption}
\vspace{0.2cm}
Assumption \ref{assump:dgp} is the standard regularity condition for
the population $L^{2}$ problem.
Assumption \ref{assump:dgp}(a) ensures that the
ReLU-transformed outcome $(y-Y)_{+}$ and any
linear combination $X^{\top}\beta$ are square integrable for each fixed $y \in \mathcal{Y}$.
Assumption \ref{assump:dgp}(b) imposes the full-rank condition on the Gram matrix of $X$.
Under Assumption \ref{assump:dgp},
the criterion
$Q(\beta; y)$ is finite and strictly convex in $\beta$ for each $y \in \mathcal{Y}$.
Hence, $\beta_{0}(y)$ is the unique minimizer of
$Q(\beta; y)$
and admits the
closed-form expression
\begin{equation}
\label{eq:bp-closed}
\beta_{0}(y)
=
\big( \mathbb{E}[X X^{\top}] \big)^{-1}
\mathbb{E} [ X (y-Y)_{+} ].
\end{equation}
The linear combination $X^{\top}\beta_{0}(y)$ is the best linear
predictor of $(y - Y)_{+}$ given $X$ in the $L^{2}$ sense.
This characterization remains valid under
misspecification, in the sense that the conditional expectation
$\mathbb{E}[(y - Y)_{+} \mid X]$ is not necessarily linear in $X$.
The interpretation is analogous to that of linear mean regression
\citep{white1980using} and linear quantile regression
\citep{angrist2006quantile} under misspecification.
Proposition
\ref{proposition:BLP} in Appendix A states this result formally.
\subsection{Population Interpretation of ReLU Regression}
We next develop the population interpretation of ReLU regression.
Through convex duality, the ReLU transformation simultaneously encodes
the conditional distribution function and the conditional quantile
function of $Y$ given $X$.
For any $y \in \mathcal{Y}$, the integrated conditional distribution function
of $Y$
given $X=x$
is defined as
\begin{equation*}
G_{Y \mid X}(y | x) := \int_{-\infty}^{y} F_{Y \mid X}(s | x)\, ds.
\end{equation*}
The unconditional analogue has been studied in the risk-measurement
literature \citep{follmer2016stochastic}. The map
$y \mapsto G_{Y \mid X}(y | x)$ is convex for each fixed $x \in \mathcal{X}$
since the integrand $F_{Y \mid X}(\cdot | x)$ is nondecreasing.
Convexity allows us to apply the Legendre-Fenchel transformation. For
each $\tau \in [0,1]$ and $x \in \mathcal{X}$, the conjugate function is
defined as
\begin{equation}
\label{eq:conjugate-G}
G_{Y \mid X}^{\ast}(\tau | x)
:=
\sup_{y \in \mathcal{Y}}\big\{ \tau y - G_{Y \mid X}(y | x) \big\}.
\end{equation}
The next proposition characterizes the convexity and the
subdifferential structure of $G_{Y \mid X}(\cdot | x)$ and its conjugate
$G_{Y \mid X}^{\ast}(\cdot | x)$.
We write $F_{Y \mid X}(y\,\text{-} | x) := \lim_{s \nearrow y}
F_{Y \mid X}(s | x)$ for the left limit of the conditional distribution
function and $F_{Y \mid X}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}}
| x) := \lim_{u \searrow \tau} F_{Y \mid X}^{-1}(u | x)$ for the
right limit of the conditional quantile function.
\vspace{0.5cm}
\begin{proposition}
\label{proposition:L2-min-1}
Suppose that $\mathbb{E}[|Y|] < \infty$. Then the
following statements hold.
\begin{enumerate}[label=(\alph*), itemsep=0.2em]
\item The map $y \mapsto G_{Y \mid X}(y | x)$ is convex, and for
any $y \in \mathcal{Y}$,
\begin{equation*}
G_{Y \mid X}(y | x) = \mathbb{E}[(y-Y)_{+} | X=x], \quad
\text{a.s.}
\end{equation*}
Moreover, $\partial G_{Y \mid X}(y | x) =
[F_{Y \mid X}(y\,\text{-} | x),\, F_{Y \mid X}(y | x)]$ for any
$y \in \mathcal{Y}$ and $x \in \mathcal{X}$.
\item The conjugate $\tau \mapsto G_{Y \mid X}^{\ast}(\tau | x)$
is convex, and for any $\tau \in (0,1)$ and $x \in \mathcal{X}$,
\begin{equation*}
G_{Y \mid X}^{\ast}(\tau | x)
=
\int_{0}^{\tau} F_{Y \mid X}^{-1}(u | x)\, du.
\end{equation*}
Correspondingly, $\partial G_{Y \mid X}^{\ast}(\tau | x) =
[F_{Y \mid X}^{-1}(\tau | x),\,
F_{Y \mid X}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}} | x)]$
for any $\tau \in (0,1)$ and $x \in \mathcal{X}$.
\end{enumerate}
\end{proposition}
\vspace{0.5cm}
Proposition \ref{proposition:L2-min-1} establishes the relationship
between the ReLU transformation and the integrated quantile function
through convex duality. The unconditional counterpart of this
relationship is well known in the risk measurement literature
\citep{ogryczak2002dual, rockafellar2000optimization,
rockafellar2002conditional}. The proof of the proposition is given
in the appendix. The duality structure is summarized in Figure
\ref{fig:duality_structure}.
\begin{figure}[htbp]
\centering
\caption{Duality of the Integrated Distribution and Quantile Functions}
\label{fig:duality_structure}
\vspace{0.1cm}
\begin{tikzpicture}[
node distance = 3.6cm,
box/.style = {draw=black!75, fill=gray!1, rounded corners=3pt,
rectangle, minimum width=3.2cm, minimum height=1.2cm,
align=center},
arr/.style = {->, >=stealth, thick, draw=black!80,
shorten >=2pt, shorten <=2pt},
darr/.style = {<->, >=stealth, thick, draw=black!80,
shorten >=2pt, shorten <=2pt}
]
\node[box] (G) at (0, 0) {$G_{Y \mid X}$};
\node[box] (Gs) at (9.5, 0) {$G_{Y \mid X}^{*}$};
\node[box] (F) at (0, -3) {$F_{Y \mid X}$};
\node[box] (Fi) at (9.5, -3) {$F_{Y \mid X}^{-1}$};
\draw[darr] (G) -- (Gs)
node[midway, above, font=\small] {Convex Duality};
\draw[darr] (F) -- (Fi)
node[midway, above, font=\small] {Generalized Inversion};
\draw[arr] ([xshift=-0.4cm]F.north) -- ([xshift=-0.4cm]G.south)
node[midway, left, font=\small] {Integration};
\draw[arr] ([xshift=0.4cm]G.south) -- ([xshift=0.4cm]F.north)
node[midway, right, font=\small] {Subdifferential};
\draw[arr] ([xshift=0.4cm]Fi.north) -- ([xshift=0.4cm]Gs.south)
node[midway, right, font=\small] {Integration};
\draw[arr] ([xshift=-0.4cm]Gs.south) -- ([xshift=-0.4cm]Fi.north)
node[midway, left, font=\small] {Subdifferential};
\end{tikzpicture}
\vspace{0.5cm}
\begin{minipage}{0.93\textwidth}
\small
\textit{Note}: The top row illustrates the convex duality between the
integrated functions, while the bottom row reflects the relationship
between the distribution and quantile functions via generalized
inversion.
\end{minipage}
\end{figure}
Proposition \ref{proposition:L2-min-1} has three implications. First,
the result requires only the finite first-moment condition $\mathbb{E}[|Y|]
< \infty$ and is agnostic about the existence or smoothness of
conditional densities. Second, the integrated functions $G_{Y \mid X}(y |
x)$ and $G_{Y \mid X}^{\ast}(\tau | x)$ are well-defined for arbitrary
outcome distributions, including discrete and mixed
discrete-continuous cases. For non-continuous outcomes the
conditional quantile function lacks point identification, since any
choice of generalized inverse is arbitrary on flat regions of
$F_{Y \mid X}(\cdot | x)$. The integrated conditional quantile function
avoids this difficulty entirely, since the choice of generalized
inverse affects only a Lebesgue-null set and does not alter the
integral. Third, the integrated functions fully characterize the
underlying distribution. The conditional distribution function and
the conditional quantile function are recovered as elements of the
subdifferentials, $F_{Y \mid X}(y | x) \in \partial G_{Y \mid X}(y | x)$ and
$F_{Y \mid X}^{-1}(\tau | x) \in \partial G_{Y \mid X}^{\ast}(\tau | x)$.
For distributions with positive density, these subdifferentials reduce to singletons, whereas for non-smooth distributions they remain well-defined as set-valued maps.
\section{Estimation and Asymptotic Properties}
\label{sec:estimation}
In this section, we present the ReLU
regression estimator and derive its asymptotic distribution. We then develop
inference for the integrated conditional quantile function.
\subsection{Estimation Method}
We observe
an independent and identically distributed (i.i.d.) sample
$\{(X_{i}, Y_{i})\}_{i=1}^{n}$ drawn from the joint distribution of
$(X, Y)$.
For each fixed $y \in \mathcal{Y}$, we define the objective function
$\widehat{Q}(\, \cdot \,; y): \mathcal{B} \to \mathbb{R}_{+}$ as
\begin{equation*}
\widehat{Q}(\beta; y)
:=
\frac{1}{n} \sum_{i=1}^{n}
\big( (y - Y_{i})_{+} - X_{i}^{\top} \beta \big)^{2}.
\end{equation*}
We define the estimator $\hat{\beta}(y)$ as the minimizer of
$\widehat{Q}(\beta; y)$ over the parameter space $\mathcal{B}$. Under
Assumption \ref{assump:dgp}(b), the sample second-moment matrix
$n^{-1} \sum_{i=1}^{n} X_{i} X_{i}^{\top}$ is invertible with
probability approaching one. The estimator then admits the
closed-form expression
\begin{equation}
\label{eq:beta-hat}
\hat{\beta}(y)
=
\bigg( \sum_{i=1}^{n} X_{i} X_{i}^{\top} \bigg)^{-1}
\sum_{i=1}^{n} X_{i} (y - Y_{i})_{+}.
\end{equation}
Given $\hat{\beta}(y)$, we estimate the integrated conditional
distribution function $G_{Y \mid X}(y | X)$ by
\begin{equation*}
\widehat{G}_{Y \mid X}(y | X)
:=
X^{\top} \hat{\beta}(y).
\end{equation*}
Evaluating this estimator across $y \in \mathcal{Y}$ produces the collection
$\{\widehat{G}_{Y \mid X}(y | X) : y \in \mathcal{Y}\}$, which characterizes the
estimated conditional distribution of $Y$ given $X$.
We estimate the integrated conditional quantile function by applying
the Legendre-Fenchel transformation to $\widehat{G}_{Y \mid X}(\cdot | x)$.
For each fixed $x \in \mathcal{X}$ and $\tau \in [0, 1]$, the conjugate
estimator is
\begin{equation*}
\widehat{G}_{Y \mid X}^{\ast}(\tau | x)
:=
\sup_{y \in \mathcal{Y}}
\big\{ \tau y - \widehat{G}_{Y \mid X}(y | x) \big\}.
\end{equation*}
The population map $y \mapsto G_{Y \mid X}(y | x)$ is convex, but its
sample counterpart $y \mapsto \widehat{G}_{Y \mid X}(y | x)$ need not be.
To restore convexity, we apply the Legendre-Fenchel transformation a
second time. For each $x \in \mathcal{X}$ and $y \in \mathbb{R}$, the biconjugate is
$
\widehat{G}_{Y \mid X}^{\ast\ast}(y | x)
:=
\sup_{\tau \in [0,1]}
\big\{ \tau y - \widehat{G}_{Y \mid X}^{\ast}(\tau | x) \big\}.
$
The biconjugate coincides with the greatest convex minorant of
$\widehat{G}_{Y \mid X}(\cdot | x)$ and
is convex by construction.
\subsection{Asymptotic Properties}
We now derive the asymptotic distribution of our proposed estimator.
To this end, we impose the following regularity conditions.
\vspace{0.3cm}
\begin{assumption}
\label{assump:regularity}
The following conditions hold:
\begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt]
\item[\textnormal{(a)}]
$\{(X_{i}, Y_{i})\}_{i=1}^{n}$ is i.i.d.\ from the distribution of
$(X, Y)$.
\item[\textnormal{(b)}]
$\mathcal{Y}_{0}$ is a fixed compact subset of $\mathbb{R}$.
\item[\textnormal{(c)}]
$\mathbb{E}[Y^{4}] < \infty$ and $\mathbb{E}[ \|X\|^{4} ] < \infty$.
\end{enumerate}
\end{assumption}
\vspace{0.3cm}
Assumption \ref{assump:regularity}(a) requires independent random
sampling. Assumption \ref{assump:regularity}(b) imposes only compactness
of the index set $\mathcal{Y}_{0}$. We do not require $\mathcal{Y}_{0}$ to lie inside the
support of $Y$.
If $\mathcal{Y}_{0}$ lies strictly below the support of $Y$, then $(y - Y)_{+} = 0$ almost surely for every $y \in \mathcal{Y}_{0}$, and thus $\hat{\beta}(y) = \beta_{0}(y) = 0$ and the empirical process is identically zero. If $\mathcal{Y}_{0}$ extends beyond the upper end of the support of $Y$, the empirical process converges to a perfectly correlated Gaussian process in that region.
In both boundary cases the conclusion of Theorem \ref{theorem:aym}
continues to hold, with a degenerate limit at points outside the
support.
Assumption \ref{assump:regularity}(c) requires finite fourth moments, which ensures that the asymptotic variance is finite.
The next theorem establishes uniform weak convergence of
$\hat{\beta}(\cdot)$ to a Gaussian process in $\ell^{\infty}(\mathcal{Y}_{0})^p$.
\vspace{0.2cm}
\begin{theorem}
\label{theorem:aym}
Suppose that Assumptions \ref{assump:dgp} and
\ref{assump:regularity} hold. Then
\begin{equation*}
\sqrt{n}\big( \hat{\beta}(\cdot) - \beta_{0}(\cdot) \big)
\rightsquigarrow
\mathbb{B}(\cdot)
\quad \text{in } \ell^{\infty}(\mathcal{Y}_{0})^{p},
\end{equation*}
and, for every $x \in \mathcal{X}$,
\begin{equation*}
\sqrt{n}\big(
\widehat{G}_{Y \mid X}(\cdot | x) - G_{Y \mid X}(\cdot | x)
\big)
\rightsquigarrow
x^{\top} \mathbb{B}(\cdot)
\quad \text{in } \ell^{\infty}(\mathcal{Y}_{0}),
\end{equation*}
where $\mathbb{B}(\cdot)$ is a zero-mean Gaussian process with
uniformly continuous sample paths and covariance function
$Q_{X}^{-1}\, \Sigma(y_{1}, y_{2})\, Q_{X}^{-1}$ for
$(y_{1}, y_{2}) \in \mathcal{Y}_{0}^{2}$, with $Q_{X} := \mathbb{E}[X X^{\top}]$
and
$\Sigma(y_{1}, y_{2}) := \mathbb{E}\big[ X X^{\top}\, \epsilon(y_{1})\, \epsilon(y_{2}) \big]$.
\end{theorem}
\vspace{0.3cm}
Uniform continuity of the sample paths of $\mathbb{B}(\cdot)$ supports
the construction of uniform confidence bands for
$G_{Y \mid X}(\cdot | x)$ over $\mathcal{Y}_{0}$.
The covariance function is the heteroskedasticity-robust covariance
of the ReLU regression score.
The limiting distribution in Theorem \ref{theorem:aym} depends on
unknown nuisance parameters and is not pivotal. To obtain practical
inference for $\widehat{G}_{Y \mid X}(y | x)$, we apply the exchangeable
bootstrap \citep{praestgaard1993exchangeably}, which consistently
estimates the limit law of the linear functional
$x^{\top} \mathbb{B}(\cdot)$. When the data are organized into independent clusters, the same
bootstrap applies with weights drawn once per cluster and assigned
to every observation in that cluster. Appendix C establishes
consistency for the cluster case.
We next study the asymptotic behavior of the conjugate estimator
$\widehat{G}_{Y \mid X}^{\ast}(\tau | x)$. The Legendre-Fenchel
transformation is not fully Hadamard differentiable. Thus, the standard
functional delta method of \cite{van1996weak} does not apply. We
work instead with Hadamard directional differentiability, which
extends the delta method to maps that are differentiable only in
selected directions \citep{shapiro1990, dumbgen1993, fang2019inference}.
\vspace{0.3cm}
\begin{definition}
\label{def:Hdirectional}
Let $\mathbb{D}$ and $\mathbb{E}$ be normed spaces. A map
$g: \mathbb{D} \to \mathbb{E}$ is Hadamard directionally
differentiable at $\phi \in \mathbb{D}$ tangentially to
$\mathbb{D}_{0} \subset \mathbb{D}$ if there exists a continuous map
$g_{\phi}^{\prime}: \mathbb{D}_{0} \to \mathbb{E}$ such that
\begin{equation*}
\lim_{n \to \infty}
\bigg\|
\frac{g(\phi + t_{n} h_{n}) - g(\phi)}{t_{n}}
- g_{\phi}^{\prime}(h)
\bigg\|_{\mathbb{E}} = 0
\end{equation*}
for all sequences $\{h_{n}\} \subset \mathbb{D}$ and
$\{t_{n}\} \subset \mathbb{R}_{+}$ with $t_{n} > 0$,
$h_{n} \to h \in \mathbb{D}_{0}$, and $t_{n} \downarrow 0$.
\end{definition}
\vspace{0.3cm}
Let $\mathcal{L} : \ell^{\infty}(\mathcal{Y}_{0}) \to \ell^{\infty}([0, 1])$
denote the Legendre-Fenchel transformation defined by
\begin{equation*}
\mathcal{L}(\phi)(\tau)
:=
\sup_{y \in \mathcal{Y}_{0}}
\{ \tau y - \phi(y) \},
\qquad \phi \in \ell^{\infty}(\mathcal{Y}_{0}),\ \tau \in [0, 1].
\end{equation*}
For each fixed $\tau$, the supremum operation in the
Legendre-Fenchel transformation is Hadamard directionally
differentiable at $\phi$ tangentially to $C(\mathcal{Y}_{0})$ by Theorem 2.1 of
\cite{carcamo2020}, and the directional derivative
$\mathcal{L}_{\phi}^{\prime} : C(\mathcal{Y}_{0}) \to \ell^{\infty}([0, 1])$
is defined as
\begin{equation*}
\mathcal{L}_{\phi}^{\prime}(h)(\tau) = \lim_{\epsilon \downarrow 0}
\sup_{y \in A_{\epsilon}(\phi, \tau)} \{-h(y)\},
\end{equation*}
where
$A_{\epsilon}(\phi, \tau) := \{y \in \mathcal{Y}_{0} : \tau y - \phi(y) \ge
\sup_{y' \in \mathcal{Y}_{0}}\{\tau y' - \phi(y')\} - \epsilon\}$.
The following proposition gives the asymptotic distribution of the
conjugate estimator via the delta method for Hadamard directionally
differentiable maps \citep{fang2019inference}.
\vspace{0.3cm}
\begin{proposition}
\label{proposition:L-weak}
Suppose that Assumptions \ref{assump:dgp}
and
\ref{assump:regularity} hold. Then, for
every $x \in \mathcal{X}$,
\begin{equation*}
\sqrt{n}\big(
\mathcal{L}\widehat{G}_{Y \mid X}(\cdot | x)
- \mathcal{L} G_{Y \mid X}(\cdot | x)
\big)
\rightsquigarrow
\mathcal{L}_{G_{Y \mid X}(\cdot | x)}^{\prime}\big( x^{\top} \mathbb{B} \big) \quad \text{in } \ell^{\infty}([0, 1]),
\end{equation*}
where $\mathbb{B}(\cdot)$ is the Gaussian process defined in
Theorem \ref{theorem:aym}.
\end{proposition}
\vspace{0.3cm}
The limit in Proposition \ref{proposition:L-weak} depends on
$\mathcal{L}_{G_{Y \mid X}(\cdot | x)}^{\prime}$ and is therefore non-pivotal.
Thus, the
standard nonparametric bootstrap is inconsistent for the limit law
\citep{dumbgen1993, fang2019inference}. We adopt the resampling approach of
\cite{fang2019inference} for Hadamard directionally differentiable
maps.
Alternative approaches for inference on shape-constrained functions
include \cite{delgado2012testing} and \cite{beare2015transforming}, among others.
The argmax set in the conjugate operation coincides with the
subdifferential $\partial G_{Y \mid X}^{\ast}(\tau | x) =
[F_{Y \mid X}^{-1}(\tau | x),\,
F_{Y \mid X}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}} | x)]$ from
Proposition \ref{proposition:L2-min-1}(b). Estimating this set by
the strict empirical argmax fails when the subdifferential is a
non-degenerate interval, since small sampling fluctuations can
place the empirical argmax at any point in the interval. The
fattened-argmax device used in \cite{chernozhukov2013intersection}
and \cite{cattaneo2020bootstrap} addresses this issue by replacing
the strict argmax with the level set
\begin{equation*}
\widehat{\partial} G_{Y \mid X}^{\ast}(\tau | x; \nu_{n})
:=
\big\{ y \in \mathcal{Y}_{0} :
\tau y - \widehat{G}_{Y \mid X}(y | x)
\ge
\sup_{y' \in \mathcal{Y}_{0}}
\{\tau y' - \widehat{G}_{Y \mid X}(y' | x)\}
- \nu_{n}
\big\}
\end{equation*}
for a vanishing tuning sequence $\nu_{n}$. The plug-in directional derivative $-\inf\{h(y) : y \in
\widehat{\partial} G_{Y \mid X}^{\ast}(\tau | x; \nu_{n})\}$ recovers the
standard nonparametric bootstrap when the subdifferential is a
singleton. It captures the worst-case Hadamard directional
derivative when the subdifferential is a non-degenerate interval.
In both cases the procedure exploits the convex structure of the
conjugate operation.
\section{Treatment Effect}
\label{sec:treatment}
\subsection{Setup and the Treatment Effects}
We extend the ReLU regression to evaluate distributional treatment
effects in a binary treatment setting under the potential
outcomes framework \citep{neyman1923applications,
rubin1974estimating}. Let $W \in \{0, 1\}$ denote the binary
treatment indicator, with $W = 1$ for treated units and $W = 0$ for control units, and let $Y(0)$ and $Y(1)$ denote the corresponding
potential outcomes. Under the stable unit treatment value
assumption \citep{rubin1980randomization}, the observed outcome is
$Y = Y(W)$.
For each $w \in \{0, 1\}$,
let
$F_{Y(w)}(y) := \Pr\{Y(w) \le y\}$
denote
the distribution function
of $Y(w)$, and
let
$F_{Y(w)}^{-1}(u) := \inf\{y \in \mathbb{R} :
F_{Y(w)}(y) \ge u\}$
denote
the corresponding
quantile function for $u \in (0, 1)$.
For probability levels $\tau_{\ell}, \tau_{u} \in [0, 1]$ with
$\tau_{\ell} < \tau_{u}$,
we define the average quantile treatment effect
(AQTE) over $[\tau_{\ell}, \tau_{u}]$ as
\begin{equation*}
\theta(\tau_{\ell}, \tau_{u})
:=
\frac{1}{\tau_{u} - \tau_{\ell}}
\bigg\{
\int_{\tau_{\ell}}^{\tau_{u}} F_{Y(1)}^{-1}(u)\, du
-
\int_{\tau_{\ell}}^{\tau_{u}} F_{Y(0)}^{-1}(u)\, du
\bigg\}.
\end{equation*}
The AQTE measures the difference between the quantile functions of the potential outcomes, averaged over the interval $[\tau_{\ell}, \tau_{u}]$, thereby capturing heterogeneous treatment effects across the outcome distributions.
Moreover, the integral formulation renders the AQTE robust to set-valuedness of the quantile functions, since
such irregular points form a set of Lebesgue measure zero
and do not alter the integral.
The AQTE encompasses well-known treatment effect parameters as
special cases, as established in the following proposition.
\vspace{0.3cm}
\begin{proposition}
\label{proposition:AQTE}
Suppose $\mathbb{E}[|Y(w)|] < \infty$ for each $w \in \{0, 1\}$. Then
\begin{enumerate}[label=\textnormal{(\alph*)}, topsep=3pt, itemsep=3pt, align=left, leftmargin=1em, labelsep=0.5em]
\item $\theta(0, 1) = \mathbb{E}[Y(1) - Y(0)]$;
\item for any $\tau \in (0, 1)$,
the limit point of
$\theta(\tau_{\ell}, \tau_{u})$ as $(\tau_{\ell}, \tau_{u})
\to (\tau, \tau)$ with $\tau_{\ell} < \tau_{u}$
equals $F_{Y(1)}^{-1}(\tau) -
F_{Y(0)}^{-1}(\tau)$
if $F_{Y(w)}^{-1}(\cdot)$ is
continuous at $\tau$ for each $w \in \{0, 1\}$, and
in general,
every limit point of
$\theta(\tau_{\ell}, \tau_{u} )$ belongs to
the interval $[F_{Y(1)}^{-1}(\tau) -
F_{Y(0)}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}}),\,
F_{Y(1)}^{-1}(\tau\raisebox{0pt}{\scalebox{0.7}{$+$}}) -
F_{Y(0)}^{-1}(\tau)]$.
\end{enumerate}
\end{proposition}
\vspace{0.3cm}
Proposition \ref{proposition:AQTE} shows that the AQTE nests two canonical treatment effect parameters.
The AQTE recovers the average
treatment effect by setting $(\tau_{\ell}, \tau_{u}) = (0, 1)$.
It also recovers the $\tau$-th quantile treatment effect in the
limit as $(\tau_{\ell}, \tau_{u}) \to (\tau, \tau)$, provided the quantile functions of both potential outcomes are continuous at $\tau$.
\subsection{Identification}
We consider the identification of the AQTE through the integrated
distribution and quantile functions of the potential outcomes,
defined for each $w \in \{0,1\}$ as
\begin{equation*}
G_{Y(w)}(y)
:= \int_{-\infty}^{y} F_{Y(w)}(s)\, ds
\quad
\mathrm{and}
\quad
G_{Y(w)}^{\ast}(\tau)
:=
\int_{0}^{\tau}
F_{Y(w)}^{-1}(u)\, du .
\end{equation*}
By Proposition \ref{proposition:L2-min-1}(b),
the integrated quantile function
$G_{Y(w)}^{\ast} (\cdot)$
is the convex conjugate of
$G_{Y(w)} (\cdot)$.
It therefore suffices to identify
$G_{Y(w)}$ for the identification of the AQTE.
We consider identification under random assignment, as in
randomized experiments. Identification under unconfoundedness
conditional on pre-treatment covariates is developed in Appendix B.
\vspace{0.3cm}
\begin{assumption}
\label{assump:random-assignment}
The data-generating process satisfies:
\begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt]
\item[\textnormal{(a)}] $Y = Y(W)$ almost surely.
\item[\textnormal{(b)}] $\mathbb{E}[|Y(w)|] < \infty$ for each $w \in
\{0, 1\}$.
\item[\textnormal{(c)}]
$(Y(0), Y(1)) \protect\mathpalette{\protect\independenT}{\perp} W$ and
$0 < \Pr(W = 1) < 1$.
\end{enumerate}
\end{assumption}
\vspace{0.3cm}
Assumption \ref{assump:random-assignment} collects the standard
conditions of the potential outcomes framework. Condition (a) is
the consistency condition that relates the observed outcome to the
realized potential outcome. Condition (b) ensures that the
integrated distribution function $G_{Y(w)}$ is finite for each
$w \in \{0, 1\}$. Condition (c) combines random assignment of $W$
with the overlap condition $0 < \Pr(W = 1) < 1$, which ensures that
both treatment arms are observable.
We now state the identification result for the AQTE under random
assignment.
\vspace{0.3cm}
\begin{theorem}
\label{theorem:AQTE}
Suppose Assumption \ref{assump:random-assignment} holds. Then,
$G_{Y(w)}(y) = \mathbb{E}[(y - Y)_{+} \mid W = w]$
for each $w \in \{0, 1\}$ and $y \in \mathbb{R}$,
and
for any $\tau_{\ell}, \tau_{u} \in [0, 1]$ with $\tau_{\ell} < \tau_{u}$,
\begin{equation*}
\theta(\tau_{\ell}, \tau_{u})
=
\frac{1}{\tau_{u} - \tau_{\ell}}
\big \{
\big ( G_{Y(1)}^{\ast}(\tau_u) - G_{Y(1)}^{\ast}(\tau_\ell)\big)
-
\big ( G_{Y(0)}^{\ast}(\tau_u) - G_{Y(0)}^{\ast}(\tau_\ell)\big)
\big \}.
\end{equation*}
\end{theorem}
\vspace{0.3cm}
The theorem establishes identification of the AQTE through the
Legendre--Fenchel transformation of the integrated distribution
function $G_{Y(w)}(\cdot)$, which is itself identified from the
conditional distribution of $Y$ given $W = w$ under random
assignment, for each $w \in \{0, 1\}$. The AQTE is point-identified
without imposing continuity of the potential outcome distributions or any further regularity conditions, yet it captures heterogeneous treatment effects across the outcome distributions through the choice of the quantile interval $[\tau_{\ell}, \tau_{u}]$.
\subsection{Estimation and Asymptotic Properties}
Suppose that we observe an i.i.d.\ sample
$\{(W_{i}, Y_{i})\}_{i=1}^{n}$ drawn from the joint distribution of
$(W, Y)$. For the estimation of the AQTE, we consider the ReLU
regression with the regressor $X = (1, W)^{\top}$ and use the
estimator $\hat{\beta}(y) \in \mathbb{R}^{2}$ from \eqref{eq:beta-hat}.
For each $w \in \{0, 1\}$, the integrated distribution function
$G_{Y(w)}(y)$ is estimated by $\widehat{G}_{Y(w)}(y)$, the fitted
value at $W = w$, and its convex conjugate
$\widehat{G}_{Y(w)}^{\ast}$ is obtained by the Legendre--Fenchel
transformation. For quantile indices $\tau_{\ell}, \tau_{u} \in
[0, 1]$ with $\tau_{\ell} < \tau_{u}$, the AQTE estimator is
\begin{equation}
\label{eq:aqte-hat}
\widehat{\theta}(\tau_{\ell}, \tau_{u})
=
\frac{1}{\tau_{u} - \tau_{\ell}}
\big\{
\big( \widehat{G}_{Y(1)}^{\ast}(\tau_{u})
- \widehat{G}_{Y(1)}^{\ast}(\tau_{\ell}) \big)
- \big( \widehat{G}_{Y(0)}^{\ast}(\tau_{u})
- \widehat{G}_{Y(0)}^{\ast}(\tau_{\ell}) \big)
\big\}.
\end{equation}
We derive the asymptotic distribution of $\widehat{\theta}(\tau_{\ell},
\tau_{u})$ under the following regularity condition.
\vspace{0.3cm}
\begin{assumption}
\label{assump:aqte-regularity}
The following conditions hold:
\begin{enumerate}[label=(\alph*), noitemsep, topsep=0pt]
\item[\textnormal{(a)}]
$\{(W_{i}, Y_{i})\}_{i=1}^{n}$ is i.i.d.\ from the joint
distribution of $(W, Y)$.
\item[\textnormal{(b)}]
$\mathbb{E}[Y(w)^{2}] < \infty$ for each $w \in \{0, 1\}$.
\end{enumerate}
\end{assumption}
\vspace{0.3cm}
Assumption \ref{assump:aqte-regularity}(a) imposes independent
random sampling.
Although Theorem \ref{theorem:aym} is stated under the fourth-moment Assumption \ref{assump:regularity}(c), the regressor
$X=(1,W)^{\top}$
is bounded. A square-integrable envelope therefore holds under the second-moment condition in Assumption \ref{assump:aqte-regularity}(b), and the conclusion of Theorem \ref{theorem:aym} follows by the same argument.
The resulting
weak limit is a zero-mean Gaussian process $\mathbb{B}_{w}(\cdot)$
in $\ell^{\infty}(\mathcal{Y}_{0})$, for any compact $\mathcal{Y}_{0} \subset \mathbb{R}$ and for each $w \in \{0,1\}$,
with covariance kernel
\begin{equation*}
\Sigma_{w}(y_{1}, y_{2})
:=
\frac{1}{\Pr(W = w)}
\big[
\mathbb{E}[(y_{1} - Y(w))_{+}(y_{2} - Y(w))_{+}]
- G_{Y(w)}(y_{1}) G_{Y(w)}(y_{2})
\big],
\end{equation*}
and $\mathbb{B}_{0}, \mathbb{B}_{1}$ are independent.
\vspace{0.3cm}
\begin{theorem}
\label{theorem:AQTE-asymp}
Suppose Assumptions
\ref{assump:random-assignment}
and
\ref{assump:aqte-regularity} hold, and
let $\tau_{\ell}, \tau_{u} \in (0, 1)$ with $\tau_{\ell} < \tau_{u}$.
Then we have
\begin{equation*}
\sqrt{n}\big(\widehat{\theta}(\tau_{\ell}, \tau_{u}) -
\theta(\tau_{\ell}, \tau_{u})\big) \rightsquigarrow
\dfrac{1}{\tau_{u} - \tau_{\ell}}
\big\{
\big( Z_{1}(\tau_{u}) - Z_{1}(\tau_{\ell}) \big)
- \big( Z_{0}(\tau_{u}) - Z_{0}(\tau_{\ell}) \big)
\big\},
\end{equation*}
where
$
Z_{w}(\tau)
:=
\sup_{y \in \partial G_{Y(w)}^{\ast}(\tau)}
\big\{ -\mathbb{B}_{w}(y) \big\}
$
for each $w \in \{0, 1\}$ and $\tau \in \{\tau_{\ell}, \tau_{u}\}$.
The delta-method bootstrap
based on the subdifferential estimator
of $\partial G_{Y(w)}^{\ast}(\tau)$ for $w \in \{0,1\}$
consistently estimates the limit law above.
\end{theorem}
\vspace{0.3cm}
Theorem \ref{theorem:AQTE-asymp} characterizes the
limit law as a linear combination of the Hadamard directional
derivatives of the Legendre-Fenchel transformation evaluated at
$G_{Y(0)}$ and $G_{Y(1)}$.
At points $\tau$ where $F_{Y(w)}^{-1}$
is continuous, the subdifferential
$\partial G_{Y(w)}^{\ast}(\tau)$ collapses to a singleton and
$Z_{w}(\tau) = -\mathbb{B}_{w}(F_{Y(w)}^{-1}(\tau))$, and the limit
law is a centered Gaussian.
At points where $F_{Y(w)}^{-1}$ is discontinuous at $\tau$,
$\partial G_{Y(w)}^{\ast}(\tau)$ has positive length, and
$Z_{w}(\tau)$ is the supremum of a Gaussian process over that
subdifferential.
To ensure valid bootstrap inference, the delta-method bootstrap is required.
Theorem \ref{theorem:AQTE-asymp} states the asymptotic
distribution of the AQTE estimator under the regressor
$X = (1, W)^{\top}$. The construction extends to a regressor of the
form $X = (1, W, V^{\top})^{\top}$ under the conditional
random-assignment condition
$(Y(0), Y(1)) \protect\mathpalette{\protect\independenT}{\perp} W \mid V$.
The marginal integrated distribution function of each potential
outcome $G_{Y(w)}(\cdot)$
is recovered from the estimator of $G_{Y(w)|V}(\cdot)$
by integrating $V$ out.
The conclusions of
Theorem \ref{theorem:AQTE-asymp} continue to hold with the corresponding Gaussian processes and covariance kernels.
\section{Application}
\label{sec:application}
In this empirical application, we examine the effect of public
health insurance on health-care utilization. The 2008 expansion
of the Oregon Medicaid program offered enrollment to a randomly
selected subset of low-income uninsured adults from a waiting
list. The resulting Oregon Health Insurance Experiment is a
randomized evaluation of the distributional effects of insurance
coverage \citep{finkelstein2012oregon}. The original analysis of
\citet{finkelstein2012oregon} reports intention-to-treat estimates
from a linear regression of the outcome on the lottery offer.
\citet{chernozhukov2020generic} revisit the experiment and report
quantile treatment effects for the count outcome through
quantile bands.
We consider $W = 1$ for individuals offered the lottery
and $W = 0$ otherwise.
The outcome $Y$ is the number of outpatient visits in the
six-month period preceding the survey. Randomization was at the
household level, and the offer probability was constant within
strata defined by the survey wave and the listed household size.
We collect the design features in
the vector $V$, which records six survey-wave dummies, two
household-size dummies, and their interactions. Conditional on
$V$, the assignment is independent of the potential outcomes,
$(Y(0), Y(1)) \protect\mathpalette{\protect\independenT}{\perp} W| V$. The marginal distribution of
$Y(w)$ is therefore identified by averaging the conditional
distribution of $Y$ given $(W = w, V)$ over the marginal
distribution of $V$. The estimand is the average quantile
treatment effect $\theta(\tau_{\ell}, \tau_{u})$ of
Section \ref{sec:treatment}. It has the intention-to-treat
interpretation as the design-averaged quantile difference between
the offered and not-offered populations on the subinterval
$[\tau_{\ell}, \tau_{u}]$.
The count nature of $Y$ rules out the linear-conditional-quantile
assumption that underlies linear quantile regression. It also
induces nonsmooth behavior of the quantile-as-inverse map. This
nonsmoothness blocks the standard delta method.
\citet{chernozhukov2020generic} address the second issue by
inverting a distribution-regression confidence band and taking
the Minkowski difference of the two quantile bands. Our approach estimates the integrated distribution function for
each treatment arm by ordinary least squares. We integrate out
the design controls through their empirical distribution. We
then recover the integrated quantile function as the convex
conjugate. The estimator is
closed-form. The average
quantile treatment effect is read off as a difference of
conjugates evaluated at two probability levels. Inference uses
the delta-method bootstrap of Section \ref{sec:estimation} based
on the empirical subdifferential interval and requires no
numerical differentiation.
Our data come from the public release of the Oregon Health
Insurance Experiment \citep{finkelstein2012oregon}. The sample
consists of survey respondents on the lottery list, with the
binary indicator $W$ recording whether the individual was drawn
in the lottery. The outcome $Y$ is right-skewed and concentrated on small
integers. It has a substantial mass at zero and a sparse upper
tail extending to several dozen visits. The empirical
distributions of $Y$ for the two treatment arms are reported in
Figure \ref{fig:application-histogram}.
We use the linear specification with regressor
$X_{i} := (1, W_{i}, V_{i}^{\top})^{\top}$ following the baseline
of \citet{finkelstein2012oregon}. We evaluate the estimator
on the grid $\mathcal{Y}_{0} := \{0, 1, 2, \dots, y_{\max}\}$, where
$y_{\max}$ is the largest observed value of $Y$.
Letting $\widehat{G}_{Y(w) \mid V}(y \mid v)$ be the estimator
of $G_{Y(w) \mid V}(y \mid v)$,
we obtain
the integrated distribution function for each treatment through the marginalization:
$\widehat{G}_{Y(w)}(y)
=
n^{-1}
\sum_{i=1}^{n}
\widehat{G}_{Y(w)|V}(y \mid V_i)$
for each $w \in \{0,1\}$.
The corresponding integrated quantile function
estimator
$ \widehat{G}_{Y(w)}^{\ast}(\tau)$
is obtained
by the Legendre-Fenchel transform.
We report the average quantile treatment effect
$\widehat{\theta}(\tau, \tau + 0.10)$ on the grid
$\tau \in \{0, 0.1, \ldots, 0.9\}$. Inference uses the
delta-method bootstrap described in
Section \ref{sec:estimation}. We draw the bootstrap weights at
the household level following Appendix C. Sampling weights from
the public release enter the ReLU regression as observation weights.
\begin{figure}[htbp]
\centering
\caption{Empirical distribution of outpatient visits by treatment status.}
\includegraphics[width=0.9\textwidth]{Fig01.png}
\label{fig:application-histogram}
\begin{minipage}{0.92\textwidth}
\small
\begin{spacing}{1.0}
\textit{Notes:} The figure reports the empirical histogram of
the number of outpatient visits in the six months preceding
the survey. The left panel pools the lottery losers ($W = 0$)
and the right panel the lottery winners ($W = 1$). The
horizontal axis is truncated at thirty visits.
\end{spacing}
\end{minipage}
\end{figure}
Figure \ref{fig:application-histogram} shows that the outcome
distribution is concentrated on a small number of integer values
and shares the same support across the two treatment arms. The
point mass at zero is visible in both panels and is smaller in
the lottery-winner panel than in the lottery-loser panel. The
mass at zero and the sparse upper tail motivate working with the
integrated distribution function rather than with the quantile
function directly.
\begin{figure}[H]
\centering
\caption{Estimated integrated quantile functions.}
\includegraphics[width=0.45\textwidth]{Fig02.png}
\label{fig:application-conjugate}
\begin{minipage}{0.9\textwidth}
\small
\begin{spacing}{1.0}
\textit{Notes:}
The figure reports the estimated integrated
quantile function for
$w \in \{0, 1\}$ on $\tau \in (0, 1)$. The shaded bands
are pointwise $95\%$ confidence bands based on the
delta-method bootstrap with
bootstrap weights drawn at the household level.
\end{spacing}
\end{minipage}
\end{figure}
Figure \ref{fig:application-conjugate} reports the estimated
integrated quantile functions $\widehat{G}_{Y(w)}^{\ast}$ for the
two treatment arms. Both curves are nondecreasing and convex by
construction. At points of differentiability, the slope at a
probability level $\tau$ identifies the marginal quantile
$F_{Y(w)}^{-1}(\tau)$. Both curves are flat over the lower
portion of the unit interval, reflecting the point mass at zero
in the marginal distribution of $Y(w)$, and start to rise around
$\tau \approx 0.35$. The treatment-arm curve sits uniformly above the control-arm
curve over the upper portion.
The change in vertical separation between probability levels
$\tau_{\ell}$ and $\tau_{u}$, divided by $\tau_{u} - \tau_{\ell}$,
equals the average quantile treatment effect over
$[\tau_{\ell}, \tau_{u}]$.
Figure \ref{fig:application-aqte} reports the estimated average
quantile treatment effect across the probability range. The
estimates are near zero on the three lowest subintervals,
mirroring the flat region of Figure
\ref{fig:application-conjugate} below $\tau \approx 0.35$. They are positive on every subinterval above $\tau = 0.3$ and are
larger at higher quantiles on average. The offer of insurance
therefore affects the upper tail of utilization more strongly than
the lower tail. The confidence band widens with $\tau$ and is widest
on the topmost subinterval, reflecting the sparse upper tail of
the outcome distribution.
\begin{figure}[H]
\centering
\caption{Average quantile treatment effect.}
\includegraphics[width=0.7\textwidth]{Fig03.png}
\label{fig:application-aqte}
\begin{minipage}{0.9\textwidth}
\small
\begin{spacing}{1.0}
\textit{Notes:} The figure reports the estimated average
quantile treatment effect
$\{\widehat{\theta}(\tau, \tau + 0.10) : \tau \in
\{0.0, 0.1, \ldots, 0.9\}\}$ together with the pointwise
$95\%$ confidence band based on the delta-method bootstrap with bootstrap weights drawn at the
household level. Each bar is plotted at the midpoint of the
corresponding subinterval.
\end{spacing}
\end{minipage}
\end{figure}
The pattern of the estimates carries an economic interpretation. The flat lower portion mirrors the large share of
respondents with zero outpatient visits over the six-month
window, whose utilization does not change with the insurance
offer. Above $\tau \approx 0.3$, the offer raises utilization at
every probability level. The effect is largest in the upper tail,
where heavier users of medical care reside. A standard average
treatment effect would aggregate the zero effect on non-utilizers
and the larger effect on heavy users into a single scalar and
would conceal this distributional pattern. The AQTE locates the gain on the relevant subintervals of the
utilization distribution and attaches a quantitative magnitude to
each. The resulting estimates are the input required for a
distributional welfare analysis of insurance coverage.
\section{Conclusion}
\label{sec:conclusion}
This paper develops a regression model with the rectified linear
unit transformation applied to the outcome variable. The ReLU regression estimates the
integrated conditional distribution function in closed form.
Through the Legendre-Fenchel transformation, its convex conjugate
equals the integrated conditional quantile function. The framework
requires only finite moments and standard rank conditions,
and it accommodates outcome distributions that are discrete, mixed,
or otherwise non-continuous without modification.
We establish the uniform asymptotic distribution of the estimator
as a process indexed by the outcome location, and we develop
inference for the conjugate functional through the delta method
for Hadamard directionally differentiable maps. As an application,
the framework identifies average quantile treatment effects over
arbitrary quantile subintervals under random assignment of a
binary treatment, with bootstrap inference based on the same
delta-method procedure.
\clearpage
\setstretch{0.1}
\bibliographystyle{chicago}
\bibliography{ReLU}
\clearpage
\setstretch{1.2}
\setcounter{section}{0}
\setcounter{figure}{0}
\setcounter{table}{0}
\setcounter{equation}{0}
\setcounter{lemma}{0}\setcounter{page}{1}
\setcounter{proposition}{0}
\begin{center}