EconBase
← Back to paper

Policy Learning with $α$-Expected Welfare

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

100,489 characters

Policy Learning with $$-Expected Welfare



\title{Policy Learning with $\alpha$-Expected Welfare\footnote{We thank Gregory M. Duncan, Sukjin Han, Toru Kitagawa, Alex Luedtke, Jing Tao, and participants of seminars at University of California San Diego, University of California Irvine, Institute for Advanced Economic Research at Dongbei University of Finance and Economics, Optimal Transport and Distributional Robustness in Banff, and INFORMS Annual Meeting in Seattle for feedback on an earlier version of this paper titled: ``Targeted Policy Learning.''}}
\author{Yanqin Fan\\
[email removed]}
\and Yuan Qi \\
[email removed]}
\and Gaoqian Xu \\
[email removed]}}
\date{{April 2024}}
\date{\today}

\onehalfspacing
\maketitle
\begin{abstract}
    This paper proposes an optimal policy that targets the average welfare of the worst-off $\alpha$-fraction of the post-treatment outcome distribution. We refer to this policy as the $\alpha$-Expected Welfare Maximization ($\alpha$-EWM) rule, where $\alpha \in (0,1]$ denotes the size of the subpopulation of interest. The $\alpha$-EWM rule interpolates between the expected welfare ($\alpha=1$) and the Rawlsian welfare ($\alpha\rightarrow 0$). For $\alpha\in (0,1)$, an $\alpha$-EWM rule can be interpreted as a distributionally robust EWM rule that allows the target population to have a different distribution than the study population. Using the dual formulation of our $\alpha$-expected welfare function, we propose a debiased estimator for the optimal policy and establish its asymptotic upper regret bounds. In addition, we develop asymptotically valid inference for the optimal welfare based on the proposed debiased estimator. We examine the finite sample performance of the debiased estimator and inference via both real and synthetic data.
    \vspace{12pt}

    \textit{JEL codes}: C10, C14, C31, C54

    \textit{Keywords}: Average Value at Risk, Optimal Welfare Inference, Regret Bounds, Targeted Policy, Treatment Effects
\end{abstract}



\newpage
\section{Introduction}
\label{sec:intro}

\subsection{Motivation}

\textit{Targeted/personalized policy rules assign treatments to individuals based on their observable characteristics}.
Learning treatment assignment policies that benefit the relevant population in a desirable way often require careful consideration. The fact that treatment effects tend to vary with individual observable characteristics prompts policy makers to design policies that determine treatment statuses based on individual characteristics. Examples include deciding which patients should receive medical treatment, assigning unemployed workers to training programs, and selecting which students to offer financial aid. Using experimental or observational data from a sample that represents the relevant population, the optimal utilitarian policy maximizes the sum of individual welfare in the sample. The empirical welfare maximization approach in \cite{kitagawa2018a} provides a solution in this regard.

As noted in \cite{kitagawa2021equality}, maximizing the utilitarian social welfare criterion overlooks distributional impacts. This motivates \cite{kitagawa2021equality} to introduce an equality-minded rank-dependent social welfare function that places greater emphasis on individuals with lower-ranked outcomes.
\textit{When the policy class is restricted due to considerations such as implementability, cost, and interpretability, maximizing the utilitarian social welfare may even hurt those who are disadvantaged in the population}. For example, if welfare is measured as the (negative) mean blood sugar level of individuals at risk of diabetes and the treatment is a new medication, a utilitarian policy may prescribe the medication to most individuals because it can substantially benefit the low-risk individuals, who form the majority of the sample, but high-risk individuals who receive the medication may be hurt and end up in even worse situations. Similarly, if welfare is evaluated by the average post-training income, a utilitarian policy is more inclined to select individuals who are high school graduates and have experienced relatively short periods of unemployment to participate in the training program, while overlooking those with lower educational attainment or longer unemployment durations who might also benefit substantially from the training; see Section 5.1 in \cite{athey2021policy}.

\textit{Taking a group-agnostic and risk-averse point of view, this paper proposes to learn an optimal policy that favors individuals on the lower tail of the outcome distribution}.
Specifically, for any $\alpha\in(0,1)$, we introduce the $\alpha$-expected welfare function as the expected outcome among the worst-affected $(\alpha\times100)\%$ of the population, i.e., a lower-tail conditional average. We study non-randomized binary policies which maximize the  $\alpha$-expected welfare and refer to such policies as \textit{$\alpha$-expected welfare maximization ($\alpha$-EWM)} policies.
The choice of $\alpha$ is problem-specific and should be based on domain knowledge. A smaller $\alpha$ means that the policy is tailored for the more disadvantaged, whereas a larger $\alpha$ generates a policy that considers a broader less-advantaged subpopulation but those who are most disadvantaged receive less attention. From a philosophical standpoint, when $\alpha$ is small, our $\alpha$-EWM objective aligns with John Rawls' \textit{difference principle}, which aims to maximize the welfare of the least-advantaged group to maintain social stability and fairness \citep{41549aa6-42b1-36d1-ab6a-84fdd10b1f93}. Indeed, the $\alpha$-expected welfare converges to the essential infimum of the outcome random variable as $\alpha$ approaches zero. We note that the definition of the $\alpha$-expected welfare function also applies to $\alpha = 1$, in which case it reduces to the utilitarian welfare underlying the empirical welfare maximization studied in \cite{kitagawa2018a} and \cite{athey2021policy}. We refer to such policies as $1$-EWM throughout the rest of this paper.

To further motivate our $\alpha$-EWM for $\alpha \in (0,1)$, we provide a simple numerical comparison with the $1$-EWM criterion from \cite{kitagawa2018a}, the equality-minded welfare criterion from \cite{kitagawa2021equality}, and quantile maximization from \cite{wang2018quantile}. Section \ref{section:relation} discusses the relationship between our $\alpha$-EWM and these criteria in more detail. We use a simple data generating process (DGP) similar to the motivating example in \cite{wang2018quantile}:
\begin{align}
\label{Spec: illustrative}
    Y=20+3 A+X-5 A X+\left(1+A+2 A X\right) \epsilon,
\end{align}
where the covariate $X \sim \operatorname{Unif}[0,1] $, the binary treatment $A\sim\operatorname{Bernoulli}(0.5)$, and $\epsilon \sim N(0,1)$. We assume that the propensity score $e_o(\cdot)=0.5$ is known, and the policy class is defined as $\Pi_{\mathrm{c}} = \mathds{1} \{ X \leq c \}$ for the policy parameter $c \in [0,1]$.

We create a superpopulation of size one million. Since we can generate $Y_i$ for both $A_i=0$ and $A_i=1$, we have full knowledge of the true outcome distribution induced by any $c$. For comparison, we select values of $c$ that maximize the following: the $0.1$-expected welfare, the standard Gini social welfare, the $0.1$-outcome quantile, and the mean outcome. These correspond to the $0.1$-EWM, equality-minded, $0.1$-quantile-optimal, and $1$-EWM policies, respectively.
Figure \ref{Figure: illustrative density} displays the probability densities of the post-treatment outcomes induced by these policies. Under this DGP, there is a gradual tightening of the post-treatment outcome distribution as we move from the $1$-EWM policy to the equality-minded policy, then to the $0.1$-quantile-optimal policy, and finally to the $0.1$-EWM policy. The 0.1-EWM policy produces the most concentrated outcome distribution, with the thinnest tails on both the left and right compared to the other policies. This suggests that the $0.1$-EWM policy not only mitigates the risk of extremely poor outcomes but also avoids disproportionately large gains, resulting in a more equitable distribution centered around the median.

\begin{figure}[t]
\centering
   \includegraphics[width=0.75\textwidth]{plots/Illustrative/density_all.jpeg}
   \caption{Distributions of post-treatment outcomes induced by the optimal policies under different welfare criteria.}
   \label{Figure: illustrative density}
\end{figure}

\subsection{Main Contributions}
\textit{This paper makes several contributions to the literature on policy learning}. First,
under the assumption of unconfoundedness,\footnote{The assumption of unconfoundedness is not essential and can be replaced with any assumption that identifies the conditional marginal distributions of the potential outcomes.} we show that the $\alpha$-expected welfare function is identified and propose a debiased estimator. Our debiased estimator utilizes cross-fitted nuisance estimators and the orthogonal moment function based on the dual form of the $\alpha$-expected welfare function. Optimizing the $\alpha$-expected welfare poses noticeable challenges compared with $1$-EWM. Adopting a group-agnostic perspective, the worst-off subpopulation being targeted changes dynamically with different policies. Consequently, estimating the $\alpha$-expected welfare requires the estimation of the $\alpha$-quantile of the welfare, which serves as a ``cutoff" for computing the tail average (see \cref{section: Average Value-at-Risk Welfare Function} for details).

Second, we establish theoretical guarantees of our $\alpha$-EWM for any $\alpha\in (0,1)$ by deriving asymptotic upper regret bounds with an explicit expression for the constant. This complements similar regret bounds for $1$-EWM in \cite{kitagawa2018a} and \cite{athey2021policy}.

Third, we develop asymptotically valid inference for the optimal $\alpha$-expected welfare. When the optimal policy is unique, Wald-type inference is asymptotically valid. When the optimal policy is not unique, we develop inference by applying the generalized delta method for Hadamard directionally differentiable functionals; see, e.g., \cite{belloni2017program, fang2019inference, hong2018numerical}.

Fourth, we demonstrate that more comprehensive policy evaluations can be performed by consistently estimating the welfare of the worst-off $(\alpha\times100)\%$ of the population for \text{any} $\alpha\in(0,1)$ and policy. Put differently, even if a policy does not specifically target the worst-affected $(\alpha\times100)\%$, we can still assess its performance at $\alpha$ to gain insights into the associated trade-offs. We illustrate our $\alpha$-EWM method using experimental data from the National Job Training Partnership Act (JTPA) Study, as analyzed by \cite{bloom1997benefits}. We find that targeting smaller subpopulations—such as the bottom 25\% or 30\% of the outcome distribution—leads to more robust welfare performance across a range of welfare objectives. In contrast, targeting broader groups (e.g., the bottom 80\%) can result in substantial welfare losses for the bottom 25\%, indicating that policies aimed at broader groups may come at the expense of welfare among the most disadvantaged.

Lastly, we conduct simulation studies based on synthetic JTPA data generated using Wasserstein Generative Adversarial Networks (WGANs) developed by \cite{athey2024using}, to evaluate the performance of our estimator and compare policy outcomes. In the WGAN-JTPA setup, both the $0.25$-EWM and equality-minded policies enhance the welfare of lower-ranked individuals while reducing that of higher-ranked individuals relative to the $1$-EWM policy, with the $0.25$-EWM policy placing much greater emphasis on these adjustments. Additional simulation studies based on stylized DGPs from \cite{athey2021policy} are provided in \cref{Section: AW simulations}. Across all simulation setups, the debiased estimator and Wald inference perform satisfactorily for all $\alpha$ values considered.

The rest of the paper is organized as follows. Section \ref{sec:literature} provides an overview of the related literature. Section \ref{section: Average Value-at-Risk Welfare Function} introduces our model preliminaries, including the $\alpha$-expected welfare measure and its identification under the selection-on-observables assumption. We point out relations and differences between four welfare measures: the 1-expected welfare, equality-minded welfare, quantile welfare, and our $\alpha$-expected welfare. Section \ref{Section: debiased} reviews the dual form of the $\alpha$-expected welfare function and presents its debiased estimator, and \cref{Section: theory} establishes an asymptotic upper regret bound for our debiased optimal policy. \cref{section:inference} constructs asymptotically valid inference for the optimal $\alpha$-expected welfare. \cref{Section: empirical} presents numerical results, including an empirical application based on experimental data from the JTPA Study and a simulation study using WGAN-generated JTPA data. Section \ref{Section: conclusion} concludes. Technical proofs are relegated to a series of appendices.

\subsection{Related Literature}
\label{sec:literature}

Our work builds on existing literature on policy learning from experimental and observational data, as well as statistical inference for the mean outcome under the optimal policy. In the following, we provide a brief discussion of related work.

\paragraph{Mean-optimal Policy Learning}
Existing research on policy learning in economics and statistics has mainly focused on the mean-optimal policy under unconfoundedness \citep{qian2011performance,zhao2012estimating,zhang2012estimating,bhattacharya2012inferring,luedtke2016statistical,kallus2018balanced,luedtke2020performance,athey2021policy}. Most work on policy learning focus on establishing theoretical guarantees by deriving regret bounds. The seminal paper by \cite{kitagawa2018a} explores mean-optimal policy learning from experimental data in a nonparametric framework. When propensity scores are known and the policy class denoted as $\Pi$ has a finite VC dimension, they employ inverse propensity weighting to estimate the welfare function, achieving $n^{-1/2}$-rate regret bounds, where $n$ is the sample size.
\cite{athey2021policy} extend this setup to observational studies where propensity scores are unknown and the policy class $\Pi_n$ may vary with $n$. They estimate the objective function using doubly robust scores, a method that is shown to be efficient in the sense of \cite{newey1994asymptotic}. The resulting policies achieve regret bounds of the order $\sqrt{\mathrm{VC}(\Pi_n)/n}$. Notably, their regret bound depends on the convergence rate of nuisance parameter estimation and the semiparametric efficient variance for evaluating an optimal policy. Finally, under mild conditions, \cite{luedtke2020performance} show that the regret can decay faster than $n^{-1/2}$ for a fixed data distribution.

Several studies have examined statistical inference for the mean-optimal welfare associated with the first-best policies. For instance, \cite{luedtke2016statistical} propose an online one-step estimator that is $\sqrt{n}$-consistent for the optimal value function, where the estimated policy and value function are recursively updated using new observations. Similarly, \cite{shi2020breaking} conduct inference for the optimal welfare via subsample aggregating and cross-validation. In contrast, \cite{rai2018statistical} study inference for the optimal mean welfare under a restricted policy class. The author utilizes bootstrap and numerical delta methods in e.g.,
\cite{fang2019inference} and \cite{hong2018numerical}, to approximate the estimator’s limiting distribution. We apply the same set of tools to develop inference for the optimal $\alpha$-expected welfare associated with a pre-specified policy class when the optimal policy may not be unique.

\paragraph{Fairness and Robustness of Policy Learning.}
In many real-world scenarios, alternative objective functions beyond the mean outcome may be more appropriate. Some studies design objective functions with fairness considerations. Besides \cite{kitagawa2021equality} and \cite{wang2018quantile}, other studies focus on distributional robustness or external validity in decision-making by adopting robust objective functions \citep{cui2023individualized, qi2023robustness, adjaho2022externally, fan2023quantifying, lei2023policy}. The optimal policy under a robust objective function can be interpreted as the policy that maximizes the “worst-case” scenario of individualized outcomes when the underlying distribution is perturbed within an uncertainty set. \cite{fang2023fairness},\cite{viviano2024fair}, and \cite{kim2023fair} propose to maximize the average welfare subject to some fairness constraints.

The paper most closely related to ours is \cite{qi2023robustness}, which adopts the average value-at-risk (AVaR) welfare criterion to develop robust individualized decision rules. The AVaR criterion is the same as our $\alpha$-expected welfare criterion, and \cite{qi2023robustness} is motivated by the distributional robust representation of AVaR, see
\cref{equation: AVaR as DRO}. Apart from differences in motivation, the main results in \cite{qi2023robustness} and our paper also differ. First, \cite{qi2023robustness} focus on experimental data with a known propensity score, allowing direct estimation of the objective function. Instead, we consider observational studies with unknown propensity scores and estimate our objective function using doubly robust scores and cross-fitting. Second, we consider a general policy class $\Pi_n$ with a VC-dimension $\mathrm{VC}(\Pi_n)$ that may be changing with $n$. In contrast, \cite{qi2023robustness} consider a more restrictive policy class within a reproducing kernel Hilbert space, which excludes many machine learning algorithms, such as decision trees and neural networks, from being used to learn the optimal policy. Third, applied to the class of policies in \cite{qi2023robustness}, our regret bound  is sharper than theirs. Fourth,
we develop inference for the optimal welfare in experimental and observational setups. Computationally, \cite{qi2023robustness} propose a non-convex optimization algorithm based on a surrogate function that smooths the binary policy function for the use of difference-of-convex optimization, whereas our optimization is done by derivative-free methods.

We close this section by summarizing the notation used in this paper.
We use $O, o, O_{P}, o_{P}, \asymp,  \gtrsim, \lesssim$ in the following sense: $a_n=O\left(b_n\right)$ if $\left|a_n\right| \leq$ $C b_n$ for $n$ large enough; $a_n = o(b_n)$ if $a_n/b_n \rightarrow 0$; $X_n=O_{P}\left(b_n\right)$, if for any $\delta>0$, there exist $M, N>0$, such that $\mathbb{P}\left|\left|X_n\right| \geq\right.$ $\left.M b_n\right] \leq \delta$ for any $n>N ; X_n=o_{P}\left(b_n\right)$, if $\mathbb{P}\left[\left|X_n\right| \geq \epsilon b_n\right] \rightarrow 0$ for any $\epsilon>0 ; a_n \asymp b_n$ if there exist $k_1, k_2>0$ and $n_0$, such that for all $n>n_0, k_1 a_n \leq b_n \leq k_2 a_n$ if $\lim a_n / b_n=\infty$;  $a_n \gtrsim b_n$ if $b_n=O\left(a_n\right) ; a_n \lesssim b_n$ if $a_n=O\left(b_n\right)$. Furthermore, we write $f(n) = \widetilde{O}(g(n))$ if there is a function $h$ that grows poly-logarithmically such that $f(n) \leq h(g(n))g(n)$. The notation $f(n) = \Omega(g(n))$ means that there is a universal constant $c_o >0$ such that $f(n) \geq  c_og(n)$ uniformly in $n$. We use the shorthand $[n] = \{1, \ldots, n\}$, $a \vee b = \max\{a,b\}$ and $a \wedge b = \min\{a,b\}$.
The abbreviation i.i.d. stands for {\it independent and identically distributed}. In the sequel, let $c_o$ denote a generic positive constant, whose value may vary from line to line.

\section{$\alpha$-Expected Welfare Function and Optimal Policy}
\label{section: Average Value-at-Risk Welfare Function}

Suppose that we have a random sample $\left(X_i, Y_i, A_i \right)_{i=1}^n$, where $X_i\in\mathcal{X} \subseteq \mathbb{R}^p$ denotes the observable characteristics of individual $i$ (continuous or discrete), $Y_i\in\mathcal{Y}\subseteq\mathbb{R}$ represents the outcome of individual $i$ (or utility / welfare), and $A_i\in\{0,1\}$ denotes the treatment status of individual $i$, for $i\in [n]$. Without loss of generality, larger values of $Y_i$ are assumed to be preferable. To simplify notation, we define $Z_i\vcentcolon=(X_i,Y_i,A_i)\in\mathcal{Z}$ and $\mathcal{Z}=\mathcal{X}\times \mathcal{Y}\times \{0,1\}$.  Let $Y_{i}(0)$ and $Y_{i}(1)$ denote the potential outcomes that would have been observed if $A_i=0$ and $A_i=1$, respectively. Then $Y_i=A_iY_{i}(1)+(1-A_i)Y_{i}(0)$ is the realized outcome under the Stable Unit Treatment Value Assumption \citep{rubin1978bayesian, rubin1990comment}.

Throughout the rest of this paper, we
assume that $\mathbb{E}|Y_i(0)| < \infty$ and $ \mathbb{E}|Y_i(1)| < \infty$. We denote by $P$ the
distribution of $Z_i \equiv (X_i, Y_i, A_i)$, and by $\mathbb{E}_P$ and $\mathrm{Var}_P$ the expectation and variance under $P$, respectively.

\subsection{\( \alpha \)-Expected Welfare Function and Identification}

We study non-randomized binary policy/rule $\pi: \mathcal{X} \rightarrow \{0,1\}$. Let $\Pi_o$ denote the policy class that contains all Borel measurable functions from $\mathcal{X}$ to $\{0,1\}$.
For any policy $\pi\in \Pi_o$, let $Y_i(\pi):=Y_i(\pi(X_i) )$,  the outcome of individual $i$ when $\pi$ is implemented. Further, let $F_{\pi}(y)$, $y\in \mathcal {Y}$ denote the distribution function of $Y_i(\pi)$ and
$F^{-1}_\pi(\alpha)=\inf \left\{y \in\mathbb{R}:F_\pi(y)\geq\alpha \right\}$ denote the quantile function of $Y_i(\pi)$.

As discussed by \cite{kitagawa2018a} and \cite{athey2021policy}, practitioners  may adopt a pre-specified policy class $\Pi\subseteq \Pi_o$ that incorporates constraints relevant to the problem context, such as budgetary limitations, specific functional forms, fairness considerations, and other pertinent factors.
\begin{definition}[$\alpha$-Expected Welfare and Optimal Policy]\label{definition: Expected Welfare}
    Given a policy class $\Pi$ chosen by the policymaker, we define   \textit{the $\alpha$-expected welfare of $Y_i(\pi)$} as the expected welfare of the worst-off subpopulation of size $\alpha\in (0,1]$, i.e.,
\begin{equation}\label{equation: Expected Welfare given policy}
\mathbb{W}_\alpha(\pi):=\frac{1}{\alpha}\int_0^{\alpha} F_\pi^{-1}(t)dt \quad  \mbox{ for } \pi \in \Pi.
\end{equation}
An \textit{$\alpha$-expected welfare maximization ($\alpha$-EWM) policy} is defined as
\[
\pi_\alpha^*\in\operatorname{argmax}_{ \pi \in \Pi  }\mathbb{W}_\alpha(\pi).
\]
\end{definition}


As discussed in \cref{sec:intro}, $\lim_{\alpha \rightarrow 0}\mathbb{W}_\alpha(\pi)=\operatorname*{ess\,inf} Y_i(\pi) $ and $\mathbb{W}_1(\pi)=\mathbb{E} \left[Y_i(\pi) \right]$. Our welfare function $\mathbb{W}_\alpha(\pi)$ therefore flexibly interpolates between the expected welfare and infimum welfare of the  target population by varying $\alpha\in (0,1]$, where $\alpha=1$ gives the expected welfare of the target population adopted in \cite{kitagawa2018a} and \cite{athey2021policy}.

\begin{remark}
(i) Our welfare function $\mathbb{W}_\alpha(\pi)$ is closely related to a commonly used coherent risk measure in finance and risk management, Conditional Value at Risk (CVaR) or Expected Shortfall (ES) defined as $\mathrm{CVaR}_\alpha(\pi):=\mathbb{E}\left[Y_i(\pi)\mid Y_i(\pi) \leq F_{\pi}^{-1}(\alpha) \right]$, see \cite{rockafellar2000optimization}. If the distribution function of $Y_i(\pi)$ is continuous at $F_{\pi}^{-1}(\alpha)$, then $\mathbb{W}_\alpha(\pi)= \mathrm{CVaR}_{\alpha}(\pi)$, see  \cite{rockafellar2000optimization,shapiro2021lectures}.

(ii)
$\mathbb{W}_\alpha(\pi)$ is also closely related to the generalized Lorenz function, a popular tool for measuring and comparing inequality, see \cite{greselin2018classical} and \cite{shorrocks1983ranking}. Specifically,
let $L_\alpha^{\mathrm{gen}}(Y_i(\pi))$ denote the \textit{generalized} (unnormalized) Lorenz function at level $\alpha$: $L_\alpha^{\mathrm{gen}}(Y_i(\pi))\vcentcolon=\int_0^{\alpha} F_\pi^{-1}(t)dt$. Then
$\mathbb{W}_{\alpha}(\pi)=\frac{1}{\alpha }L_\alpha^{\mathrm{gen}}(Y_i(\pi))$.
\end{remark}

To identify $\mathbb{W}_\alpha(\pi)$ as defined in \cref{equation: Expected Welfare given policy}, we note that $$Y_i(\pi)=\pi(X_i)Y_i(1)+[1-\pi(X_i)]Y_i(0).$$ The conditional (given $X_i = x$) and unconditional distribution functions of $Y_i(\pi)$ are
\begin{equation} \label{equation: identification of cdf_pi}
F_{\pi}( y | x) = \pi(x) F_1(y | x)+(1-\pi(x)) F_0(y | x) \ \mbox{ and } \ F_{\pi}(y) = \int_{\mathcal{X}}  F_{\pi}( y | x)  \mathrm{d} P_X(x),
\end{equation}
where $F_1(y |x)$ and $F_0(y|x)$ are the conditional distribution functions of $Y_i(1)$ and $Y_i(0)$ given $X_i=x$, respectively.

\cref{equation: Expected Welfare given policy} and \cref{equation: identification of cdf_pi} imply that $\mathbb{W}_\alpha(\pi)$ is a function of the policy $\pi(\cdot)$ and the conditional distribution functions  $F_1( \cdot | \cdot)$ and $F_0( \cdot| \cdot)$. Consequently, for any $\pi \in \Pi_o$, $\mathbb{W}_\alpha(\pi)$ is identified as long as  $F_1(  \cdot| \cdot)$ and $F_0( \cdot| \cdot)$ are identified. Any assumption that ensures the identification of $F_{1}(\cdot|\cdot)$ and $F_{0}(\cdot|\cdot)$ is sufficient to identify $\mathbb{W}_\alpha(\pi)$.  In the rest of this paper, we adopt the selection-on-observables assumption, which includes unconfoundedness and common support, as detailed in \cref{Assumption: Selection-on-observables}.

\begin{assumption} \label{Assumption: Selection-on-observables}
\begin{assumpenum}
\item {\it Unconfoundedness:} \label{assumption: Unconfoundedness}$\left(Y_i(0), Y_i(1) \right)\rotatebox[origin=c]{90}{$\models$} A_i \mid X_i$.
\item \label{assumption: Strict Overlap}{\it Strong overlap}: Let  $e_o(x) := \mathbb{P}\left[A_i =1 \mid X_i=x \right]$ denote the propensity score. There is a constant $\kappa \in \left(0,\frac{1}{2} \right)$ such that $e_o(x)\in [\kappa, 1- \kappa]$ for all $x \in \mathcal{X}$.
\end{assumpenum}
\end{assumption}

\Cref{assumption: Unconfoundedness} states that the potential outcomes are independent of the treatments after conditioning on the observed covariates. Heuristically, it requires that all confounders that affect both treatments and potential outcomes simultaneously be observed. For identification, \cref{assumption: Strict Overlap} can be relaxed to the weaker condition that $ e_o(x) \in (0, 1)$ for all $x \in \mathcal{X}$, but the regret bounds and inference developed in later sections of this paper rely on it.

Under \Cref{Assumption: Selection-on-observables}, the distribution functions $F_{a}(\cdot|x)$ for all  $x \in \mathcal{X}$ are point-identified:
\[
F_{a}(y|x)\vcentcolon=\mathbb{P}\left[Y_i(a)\leq y|X_i=x\right]=\mathbb{P}\left[Y_i\leq y|X_i=x, A_i=a\right].
\]
Consequently, $\mathbb{W}_\alpha(\pi)$ is identified  for any $\pi \in \Pi_o$.

\subsection{Relations with Other Welfare Maximization Criteria}
\label{section:relation}
In this subsection, we compare our $\alpha$-expected welfare $\mathbb{W}_\alpha(\pi)$, defined for $\alpha \in (0,1)$, with three welfare functions commonly used in the literature: the expected welfare, the equality-minded welfare, and the quantile welfare functions.

\subsubsection{$1$-Expected Welfare Maximization}
\label{Section: relation EWM}

$1$-EWM in \cite{kitagawa2018a} and \cite{athey2021policy} take the mean outcome \(\mathbb{E}[Y_i(\pi)]\), which equals $\mathbb{W}_1(\pi) $, as the population welfare function, \textit{assuming that the distribution of \(Y_i(\pi)\) in the target population is the same as that in the study population}.

\textit{For $\alpha\in (0,1)$, our $\alpha$-expected welfare $\mathbb{W}_\alpha(\pi) $  represents a distributionally robust version of the $1$-expected welfare function}. To see this, consider the uncertainty set centered at probability distribution $F_{\pi}$ of the outcome under policy $\pi: \mathcal{X} \rightarrow \{0,1\}$:
\[
\begin{aligned}
\mathcal{U}_{\alpha} (F_\pi) & = \left\{Q: D_\infty(Q \| F_\pi ) \leq  \log \frac{1}{\alpha}  \right\}\\
& = \left\{Q  : \exists \ P \in \mathcal{P}(\mathcal{Y}) , t \in[\alpha, 1] \text { s.t. } F_\pi  = t Q+(1-t) P\right\}.
\end{aligned}
\]
where $D_\infty ( Q\| F_\pi) = \mathrm{ess \ sup} \log \frac{d Q}{d F_\pi}$. From \cite{ rockafellar2002deviation} and \cite{duchi2023distributionally}, it follows that
\begin{equation} \label{equation: AVaR as DRO}
\mathbb{W}_\alpha(\pi) = \inf_{Q \in \mathcal{U}_{\alpha}(F_\pi ) } \mathbb{E}_{Z\sim Q} \left[  Z\right].
\end{equation}
The uncertainty set $\mathcal{U}_{\alpha} (F_\pi) $ is the risk envelope capturing the distributional uncertainty of $Y_i(\pi)$ in the target population, comprising distributions with minority subpopulations of at least size $\alpha$. We can therefore interpret $\pi_\alpha^*$ as the distributionally robust policy that maximizes the average welfare under the worst-case perturbation of the study population in $\mathcal{U}_{\alpha} (F_\pi)$. As $\alpha$ decreases, the uncertainty set expands, making the $\alpha$-expected welfare function more robust to potential distributional shifts in $Y_i(\pi)$ within the target population.

\subsubsection{Equality-Minded Welfare Maximization}
\label{Section: relation EM}

Since the $1$-EWM may worsen inequality, \cite{kitagawa2021equality} propose equality-minded policies by maximizing rank-dependent social welfare functions (SWFs), which assign greater weights to lower-ranked individuals.  Given a decreasing function $\Lambda: [0,1] \rightarrow [0,1]$ with $\Lambda(0) = 1$ and $\Lambda(1) = 0$, the equality-minded welfare under policy $\pi$ is defined as
\begin{equation}
W_\Lambda(F_\pi)\vcentcolon=\int_0^\infty \Lambda \left( F_\pi(y) \right) \mathrm{d}y=\int_0^1  F^{-1}_\pi \left(t\right) \omega(t)\mathrm{d}t,
\end{equation}
where $\omega(t) \coloneqq -\frac{\mathrm{d}}{\mathrm{d}t} \Lambda(t)$ is the associated  weight function.
When $\Lambda$ is strictly convex, the associated SWF, $W_\Lambda$, upholds the Pigou-Dalton Principle of Transfers, as rank-preserving transfers from higher-ranked individuals to lower-ranked individuals are preferred under the welfare $W_\Lambda$.  The function $\Lambda$, chosen by practitioners, captures the degree of inequality aversion through its level of complexity.
An important class of rank-dependent SWFs is the extended Gini SWFs, where $\Lambda (t) =\Lambda_k(t) = (1-t)^{k-1}$ for some $k \geq 2$, and the weight function is $\omega(t)=\omega_k(t) = (k-1)(1-t)^{k-2}$.
The expected welfare and the standard Gini SWF correspond to $k = 2$ and $k = 3$, respectively.

 Equality-minded SWFs can, in fact, be expressed in terms of our $\alpha$-expected welfare $\mathbb{W}_\alpha(\pi)$. For example, when $k > 2$, the extended Gini SWF can be written as a weighted average of $\mathbb{W}_\alpha(\pi)$:
\begin{equation}\label{equation: generalized Gini welfare}
W_\Lambda(F_\pi) =(k-2) \int_0^1 \mathbb{W}_{\alpha}(\pi)
\alpha (1-\alpha)^{k-3}\mathrm{d}\alpha.
\end{equation}
Although our $\alpha$-expected welfare can be written as
\begin{equation}\label{equation: weighted}
\mathbb{W}_\alpha(\pi)= \frac{1}{\alpha} \int_0^\alpha F_\pi^{-1}(t) \mathrm{d}t = \int_0^1  F^{-1}_\pi \left(t\right) \sigma(t)\mathrm{d}t,
\end{equation}
where $\Lambda(t)  =  \left(1 - t/\alpha\right) \mathds{1}\{ 0 \leq t \leq \alpha \}$ and $\sigma(t) = \frac{1}{\alpha} \mathds{1} \{ 0 \leq  t \leq \alpha \}$, it does not satisfy the Pigou-Dalton Principle of Transfers, as $\Lambda(t)$ is not strictly convex.
This principle is satisfied only if the rank-preserving transfer happens across the probability level $\alpha$, i.e., from an individual ranked above $\alpha$ to an individual ranked below $\alpha$. Transfers on the same side do not affect $ \mathrm{W}_\alpha(\pi)$ since all the individuals involved have the same weight.

\subsubsection{Quantile Welfare Maximization}
\label{Section: relation QW}
To prioritize the lower tail of population welfare over the (weighted) expected welfare, \cite{wang2018quantile} propose a quantile-optimal policy, defined as
\[
\operatorname{argmax}_{\pi \in \Pi} \mathrm{VaR}_\alpha(Y_i(\pi)) = F_\pi^{-1}(\alpha) ,
\]
where $\alpha \in (0,1)$ is the quantile level of interest. For the class of linear policies with a fixed number of covariates $\Pi$,  \cite{wang2018quantile} establish the cube root asymptotics for the estimator of the parameter that defines the optimal linear policy.

Compared with quantile welfare $F_\pi^{-1}(\alpha)$ that overlooks the welfare of the population with outcomes below it, our $\alpha$-expected welfare function $ \mathbb{W}_\alpha(\pi)$ integrates $F_\pi^{-1}(t)$ over the range $[0, \alpha]$, thereby accounting for welfare levels below the $\alpha$-quantile and providing a more comprehensive assessment of the lower tail of the welfare distribution.

\section{Debiased Estimation and Practical Implementation}
\label{Section: debiased}

The $\alpha$-expected welfare function $\mathbb{W}_\alpha(\pi)$ has a
convenient dual representation, which we will use to construct a debiased estimator of $\mathbb{W}_\alpha(\pi)$.

Let $(u)_-\vcentcolon=\min{(u, 0)}$ and  $(u)_+ \vcentcolon= \max{(u, 0)}  $. Further, let $\theta = (\pi , \eta)$ and
$$
\mathbb{V}_\alpha(\theta) =  \frac{1}{\alpha} \mathbb{E}\left[\left(Y_i(\pi)-\eta\right)_{-}\right]+\eta.
$$

\begin{lemma}[Dual Representation of $\mathbb{W}_\alpha(\pi)$]
\label{lemma: AVaR dual}
For any $\alpha \in (0,1]$ and $\pi \in \Pi$,
\[
\label{dual}
\mathbb{W}_\alpha(\pi)=\sup_{\eta \in \mathbb{R} }\mathbb{V}_\alpha(\pi, \eta).
\]
Furthermore, for $\alpha \in (0, 1)$, the supremum is attained on the interval $[t^\star, t^{\star\star}]$, where $t^\star = \sup\{ y\in \mathbb{R}: F_\pi (y) \leq \alpha \}$ and $t^{\star \star} =F_\pi^{-1}(\alpha)$. When $\alpha = 1$, if the support of $Y_i(\pi)$ is bounded, then the supremum is attained on $\left[ F_\pi^{-1}(1), \infty \right)$. Otherwise, the supremum is unattainable and $\sup_{\eta \in \mathbb{R} }\mathbb{V}_1(\pi, \eta) = \lim_{\eta \rightarrow \infty} \mathbb{V}_1(\pi, \eta)$.

\end{lemma}


Let
\[\mu_a(x, \eta)\vcentcolon= \mathbb{E}\left[\left(Y_i(a)-\eta  \right)_- |X_i = x\right] \mbox{ for } a\in \{0,1\},\]
and $\tau(x, \eta) \vcentcolon=\mu_1(x, \eta)- \mu_0(x, \eta)$ for any $x \in \mathcal{X}$ and $\eta \in \mathbb{R}$. Under \cref{Assumption: Selection-on-observables}, $\tau(x, \eta)$ is identified for any given $\eta$.


\begin{theorem}\label{theorem: Value function identification}
Under \cref{Assumption: Selection-on-observables}, for any $0 <\alpha \leq 1$ and any $\theta = (\pi, \eta) \in \Pi_o \times \mathbb{R}$, it holds that
\begin{equation}\label{equation: V_identification}
    \begin{aligned}
\mathbb{V}_\alpha(\theta)
& = \frac{1}{\alpha}\left\{  \mathbb{E} \left[  \pi (X_i) \mu_1(X_i, \eta)\right] + \mathbb{E} \left[ \left(1-\pi(X_i ) \right) \mu_0(X_i, \eta) \right]\right\}  +\eta, \\
& = \frac{1}{\alpha}\left\{  \mathbb{E} \left[  \pi (X_i) \tau(X_i, \eta)\right] + \mathbb{E} \left[ \mu_0(X_i, \eta) \right]\right\}  +\eta, \\
& = \frac{1}{\alpha} \mathbb{E} \left[  w(X_i, A_i ,\pi) (Y_i - \eta)_{-} \right] + \eta, \\
\end{aligned}
\end{equation}
where the function $w: \mathcal{X} \times \{0,1\} \times \Pi_o \rightarrow [0, \infty)$ is defined as
\[
 w(x, a ,\pi)\vcentcolon=
\frac{a \pi (x)}{e_o(x)}  + \frac{ (1- a) \left(1- \pi(x) \right)}{1- e_o(x)}.
\]
\end{theorem}

\begin{remark}\label{remark: compact feasible set}
(i) When $0<\alpha<1$, the feasible set in the dual representation of $\mathbb{W}_\alpha(\pi)$ in Lemma \ref{lemma: AVaR dual} can be restricted to a compact set. Since $|Y_i(\pi)| \leq |Y_i(0)| + |Y_i(1)|$ for all $\pi \in \Pi_o$, the $\alpha$-quantile of $|Y_i(\pi)|$ is no greater than the $\alpha$-quantile of $|Y_i(0)| + |Y_i(1)|$, while the $\alpha$-quantile of $-|Y_i(\pi)|$ is no less than the $\alpha$-quantile of $-|Y_i(0)| - |Y_i(1)|$. Therefore, the solution to $\sup_{\eta \in \mathbb{R}} \mathbb{V}(\pi, \eta)$ is $\mathrm{VaR}_\alpha(Y_i(\pi))$, which satisfies the bounds
\[
- \mathrm{VaR}_{1-\alpha} \left( |Y_i(0)| + |Y_i(1)|  \right) \leq   \mathrm{VaR}_\alpha(Y_i(\pi)) \leq   \mathrm{VaR}_\alpha \left(|Y_i(0)| + |Y_i(1)|  \right).
\]
Thus, we can express $\mathbb{W}_\alpha(\pi)$ as $\sup_{\eta \in \mathcal{B}_Y } \mathbb{V}_\alpha(\pi, \eta)$ for some compact set $\mathcal{B}_Y \subset \mathbb{R}$.

(ii) When $Y_i$ has a bounded support, the claim in (i) holds for $\alpha=1$ as well.
\end{remark}

As noted in the previous sections, the 1-expected welfare function  $\mathbb{W}_1(\pi) $ is the same as the expected welfare \(\mathbb{E}[Y_i(\pi)]\) in \cite{kitagawa2018a} and \cite{athey2021policy}. In the rest of this paper, we focus on estimation and asymptotic theory for an $\alpha$-EWM rule when $\alpha\in (0,1)$.

Theorem \ref{theorem: Value function identification} suggests two plug-in methods for estimating $\mathbb{V}_\alpha(\theta)$ or the welfare function $\mathbb{W}_\alpha(\pi)$: IPW and outcome equation estimation. It is known that the IPW estimator is sensitive to the estimator of the propensity score and may suffer from severe bias. The outcome equation estimator may be sensitive to the estimators of $\tau$ ($\mu_1$ and $\mu_0$). This motivates the debiased estimator proposed in this section.

\Cref{theorem: Value function identification} implies that under \Cref{Assumption: Selection-on-observables}, the function $\mathbb{V}_{\alpha}(\theta)$ is identified for any fixed $\theta = (\pi, \eta)$.   Following \cite{robins1994estimation} and \cite{robins1995analysis}, we build our doubly robust score for $\mathbb{V}_\alpha(\theta)$   by introducing the augmentation term.
Given any $\theta = (\eta, \pi)$, and for  any function $\check{e}: \mathcal{X} \rightarrow (0,1)$ and  $\check{\mu}_a :\mathcal{X} \times \mathbb{R} \rightarrow\mathbb{R}$ with $a \in \{0,1\}$, define
\begin{equation} \label{equation: g_theta}
\begin{aligned}
g_\theta (z;  \check{\mu}, \check{e} ) & =\frac{1}{\alpha} \left[ \left(1-\pi(x)  \right)\check{\mu}_0 (x,\eta )+ \pi(x) \check{\mu}_1 (x,\eta ) \right]  +\eta \\
&+ \frac{1}{\alpha} \left[\frac{(1-\pi(x))(1-a)}{1-\check{e}\left(x \right)}( (y-\eta)_{-}  - \check{\mu}_0(x,\eta))\right] \\
& + \frac{1}{\alpha} \left[\frac{\pi(x) a}{\check{e}(x)}\left( (y-\eta)_{-}   - \check{\mu}_1(x,\eta)\right)\right] ,
\end{aligned}
\end{equation}
where $\check{\mu} = (\check{\mu}_0, \check{\mu}_1)$ and the augmentation term  is defined as the sum of the last two components in \eqref{equation: g_theta}. The augmentation term has mean zero and the Neyman orthogonality condition holds:
\[
\partial_{\mu} \mathbb{E}_{P} [ g_\theta(Z_i; \mu_o, e_o) ] [ \check{\mu} - \mu_o] = 0,\quad \text{ and }\quad \partial_{e} \mathbb{E}_{P} [ g_\theta(Z_i; \mu_o, e_o) ] [  \check{e}- e_o] = 0,
\]
where $\mu_o = (\mu_0, \mu_1)$. To simplify notation, we let $g_\theta(\cdot) = g_\theta(\cdot; \mu_o, e_o)$, where the function $g_\theta(\cdot)$ indexed by $\theta$ is referred to as the (doubly robust) score function for estimating $\mathbb{V}_\alpha(\theta)$.
It is clear that for any given $\theta$,  the function $g_{\theta} - \mathbb{E}_P[ g_{\theta} (Z_i)]$ is the efficient influence function for $\mathbb{V}_{\alpha} (\theta)$; see \cite{luedtke2016statistical,kennedy2016semiparametric} for more detailed discussion.

Building upon \cite{chernozhukov2018double} and \cite{chernozhukov2022locally}, we construct our doubly robust score \( \widehat{g}_\theta(Z_i) \) for \( \mathbb{V}_{\alpha}(\theta) \) based on \( K \)-fold cross-fitting, a sample-splitting method used to validate asymptotic properties and leverage high-level conditions concerning the predictive accuracy of nuisance estimation methods.

We describe the estimation steps below, see Algorithm \ref{alg:debiased} in \cref{Section: algorithm} for details.
\begin{enumerate}
\item[(a)]
Randomly partition the sample into $K$ folds $\cup_{k=1}^K \mathcal{I}_k$ such that $|\mathcal{I}_k | = n/K$.
\item[(b)] For each \( k \), define \( \mathcal{I}_k^c = [n] \setminus \mathcal{I}_k \). Fit estimators for the nuisance parameters \( e_o(\cdot) \) and \( \mu_a(\cdot, \cdot) \) for \( a \in \{0,1\} \) using the observations in the remaining \( K-1 \) folds, specifically, \( (Z_i)_{i \in \mathcal{I}_k^c} \). Denote these estimators as \( \widehat{e}^{(-k)}(\cdot) \) and \( \widehat{\mu}^{(-k)}_a(\cdot, \cdot) \).

\item[(c)] The doubly robust score is
\begin{equation} \label{equation: Gamma_theta}
\begin{aligned}
\widehat{g}_\theta(Z_i; \widehat{\mu}^{-k(i)}, \widehat{e}^{-k(i)}) & =\frac{1}{\alpha} \left[ \left(1-\pi(X_i)  \right)\widehat{\mu}^{-k(i)}_0 (X_i,\eta )+ \pi \left(X_i \right) \widehat{\mu}^{-k(i)}_1 (X_i,\eta ) \right] +\eta  \\
&+ \frac{1}{\alpha} \left[\frac{\left(1-\pi(X_i)\right)\left(1-A_i \right)}{1-\widehat{e}^{-k(i)} \left( X_i \right)} \left[ ( Y_i -\eta)_{-}  - \widehat{\mu}_0^{-k(i)}(X_i,\eta)    \right ]  \right] \\
& + \frac{1}{\alpha} \left[\frac{\pi(X_i) A_i}{\widehat{e}^{(-k(i))} \left( X_i \right) }\left[ (Y_i-\eta)_{-}   - \widehat{\mu}_{1}^{-k(i)}(X_i,\eta)\right]\right] ,
\end{aligned}
\end{equation}
where $k(i)$ is the index in $[n]$ such that $i \in \mathcal{I}_k$.
\item[(d)] For each $\theta=(\pi, \eta)$, $\mathbb{V}(\theta)$ and $\mathbb{W}_\alpha(\pi)$  can be estimated by
\[
\widehat{\mathbb{V} }_{n} (\theta) = \frac{1}{n} \sum_{i = 1}^n  \widehat{g}_\theta(Z_i) \quad \text{and} \quad
\widehat{\mathbb{W} }_{n}(\pi) = \sup_{\eta \in \mathcal{B}_Y  } \widehat{\mathbb{V} }_{n}  (\pi, \eta),
\]
where $\mathcal{B}_Y$ is introduced in Remark \ref{remark: compact feasible set}.\footnote{If the support of $Y_i$ is bounded, i.e., Assumption \ref{Assumption: Bounded support} holds, then one can take $\mathcal{B}_Y$ as the support of $Y_i$. In our numerical work, we took $\mathcal{B}_Y$ as the closed interval with lower and upper bounds as the minimum and maximum statistics of $Y_i$ respectively.  }

\item[(e)] The debiased estimator $\widehat{\theta}_{n}  = (\widehat{\pi}_{n}, \widehat{\eta}_{n} )$ is the maximizer of $\widehat{\mathbb{V} }_{n} (\theta)$.
\end{enumerate}

\begin{remark}
Since \cref{lemma: AVaR dual} implies that $\mathbb{W}_\alpha(\pi)=\mathbb{V}_\alpha(\pi, F_\pi^{-1}(\alpha))$, a debiased estimator of $\pi^*$ can also be constructed from the expression  $\mathbb{V}_\alpha(\pi, F_\pi^{-1}(\alpha))$. Noting that $F_\pi^{-1}(\alpha)$ is an optimal solution to $\sup_{\eta\in\mathbb{R}}\mathbb{V}_\alpha(\theta)$, the orthogonal moment function based on  $\mathbb{V}_\alpha(\pi, F_\pi^{-1}(\alpha))$ is equal to
\[
g_{\left(\pi,F_{\pi}^{-1}(\alpha)\right)} (z;  \check{\mu}, \check{e} ) - \mathbb{V}_\alpha(\pi, F_\pi^{-1}(\alpha)) .
\]

A cross-fitting estimator can be constructed from $g_{(\pi,F_{\pi}^{-1}(\alpha))} (z;  \check{\mu}, \check{e} )$, but it requires an estimator of the quantile function $F_\pi^{-1}(\alpha)$. We leave a detailed comparison between these two estimators in future work.
\end{remark}

\section{Asymptotic Upper Regret Bounds}
\label{Section: theory}

In this section, we establish asymptotic regret bounds on the debiased $\alpha$-EWM policy proposed in Section \ref{Section: debiased} for any fixed $\alpha\in(0,1)$. They complement similar regret bounds for the $1$-EWM and equality-minded policies established in \cite{kitagawa2018a}, \cite{athey2021policy} and  \cite{kitagawa2021equality}.

\subsection{Policy Class and Examples}
For each $n$, let $\Pi_n$ denote the class of candidate policies and $\Theta_n = \Pi_n \times \mathcal{B}_Y$, where $\mathcal{B}_Y \subset \mathbb{R}$ is a compact set introduced in \cref{remark: compact feasible set}. For brevity, we write $\mathbb{V}(\theta) \equiv \mathbb{V}_\alpha(\theta)$, omitting the subscript $\alpha$.

The following assumption restricts the complexity of the policy class $\Pi_n$.

\begin{assumption}\label{policy size}
There exists a constant $b_o > 0$ such that the VC-dimension of $\Pi_n$ is bounded as $\mathrm{VC}(\Pi_n) \leq n^{b_o}$ for all $n \in \mathbb{N}^+$.
\end{assumption}

A policy is a classifier that assigns the covariate $X_i$ to a binary treatment status. Any machine learning classification model can serve as a candidate policy class. In the following, we list three examples of policy classes and their VC dimensions.


\begin{example}[Linear Rules]  \label{example: Linear Rules}
The linear policy class can be  characterized by
\begin{equation}\label{equation: Linear policy Class}
\Pi_n = \left\{ \mathds{1}\{  x^\prime \beta > 0 \}: \beta \in B \right\},
\end{equation}
where $B$ is compact subset of $\mathbb{R}^{p_n}$, where the dimension of the covariates $p_n$ is allowed to grow with the sample size $n$. Although the eligibility score is linear in $\beta$, it can include intercepts, interaction terms, higher-order terms, and other transformations of the original covariate $X_i$.  The VC-dimension of $\Pi_n$ is $p_n + 1$.
\end{example}

\begin{example}[Decision Trees] \label{example: Decision Tree}
A decision tree is a predictor $\pi: \mathcal{X}\subset \mathbb{R}^p \rightarrow \{0,1\}$ that recursively partitions the feature space $\mathcal{X}$ into a set of rectangles and assigns a label to each resulting partition. Following \cite{bertsimas2017optimal} and \cite{zhou2023offline}, we define a decision tree recursively. A decision tree of depth $L$ consists of $L$ levels, with the first $L-1$ levels containing branch nodes and the final $L$-th level comprising exclusively of leaf nodes.  For any branch node, we choose the split-point $b$ and the variable $x(j)$ that is a single component of $x$ . If $x(j) < b$, the path taken is towards the left; if not, the decision leads to the right branch. Each path will end with a leaf node that is assigned a unique label. \cite{zhou2023offline} show that the  VC-dimension  of the class $\Pi$  of decision trees of depth-$L$ over $\mathbb{R}$ is $\mathrm{VC}(\Pi) = \widetilde{\mathcal{O}}( 2^L \log p)$.
\end{example}

\begin{example}[ReLU Neural Networks]\label{example: ReLU_DNN}
Deep neural networks have achieved significant success in complex classification tasks, especially in image and speech recognition. Formally, a neural network is defined by an activation function $\sigma: \mathbb{R} \rightarrow \mathbb{R}$, structured as a directed acyclic graph, alongside a set of parameters that include a weight for each edge within the graph and a bias for each node.   Common activation functions include the sigmoid,  $\sigma(x) = 1 /(1+ e^{-x})$, and the Rectified Linear Unit (ReLU), $\sigma(x)=\max (0, x) $. Each edge represents a connection that transmits the output from one neuron to the input of another. This input is calculated as a weighted sum of the outputs from all connected neurons, allowing the network to capture complex relationships and patterns in the data.

Let $W$ denote the total number of parameters (weights and biases),
$U$  the total number of computation units (nodes), and $L$ the length of the longest path in the network graph. Let $\Pi$ denote the policy class of deep ReLU  networks characterized by $W$ weights and  $L$ layers. \cite{bartlett2019nearly} establish that $\mathrm{VC}(\Pi) = O(W L \log(W ))$ and $\mathrm{VC}(\Pi) = \Omega\left( W L \log(W/L)  \right)$.
\end{example}

\subsection{Assumptions on Nuisance Estimators and a Preliminary Lemma}

In this section, we establish a fundamental lemma showing that the estimation error of the nuisance parameters can be ignored when $\mathbb{V}(\cdot)$ is estimated using the doubly robust score with cross-fitting. Before presenting the lemma, we introduce additional assumptions.

Let $ \mathbb{V}_{n}(\theta)= \mathbb{P}_n g_\theta$, where $g_\theta (z):= g_\theta(z; \mu_o, e_o)$ is defined in \eqref{equation: g_theta}. We assume that the nuisance parameter estimators $\widehat{\mu}_a\left(\cdot, \cdot\right)$ and $\widehat e(\cdot)$ converge to their true values at sufficiently fast rates.
 \begin{assumption}\label{assumption: nuisance parameter convergence rate}
 \begin{assumpenum}
 \item   $\sup_{ (x,\eta) \in \mathcal{X} \times \mathcal{B}_Y }   \left | \widehat{\mu}_a(x, \eta) - \mu_a(x, \eta) \right|=o_P(1)$ for $a\in \{0,1\}$, and \\
 $\sup_{ x  \in \mathcal{X}  }   \left| \widehat{e} \left(x\right) - e_o(x) \right|  = o_P(1)$.
 \item Suppose there are $\zeta_\mu > 0$ and $\zeta_e > 0$ such that
 \[
  \begin{aligned}
\sup_{\eta \in \mathcal{B}_Y }  \left[ \mathbb{E} \left|  \widehat{\mu}_a(X_i, \eta) - \mu_a(X_i, \eta)  \right|^2  \right]^{1/2} & = O (n^{-\zeta_\mu}), \\
\left[  \mathbb{E} \left| \widehat{e}\left(X_i \right) - e_o(X_i)  \right|^2  \right]^{1/2} & = O( n^{-\zeta_e}). \\
 \end{aligned}
 \]
 \item  $\mathrm{VC}(\Pi_n) = o \left(n^{ 2 \zeta_\mu \wedge  2\zeta_e } \right)$.
\end{assumpenum}
\end{assumption}

\begin{remark}
(i) The regression function $\mu_a(x, \eta)$ can readily be estimated by regressing $\{ (Y_i - \eta)_{-}: A_i = a \}$ on $\{ X_i: A_i = a \}$. The uniformity in $\eta$ does not severely impact the uniform convergence of the estimator.
For example, given a  bandwidth $b_n = o(1)$, a higher-order kernel can be employed to estimate $\mu_a(x, \eta)$ as follows:
\[
\widehat{\mu}_a(x, \eta) = \frac{ \sum_{i: A_i = a}^n (Y_i -\eta)_{-} K\left( \frac{x-X_i}{b_n} \right)  }{ \sum_{i: A_i = a}^n K\left( \frac{x-X_i}{b_n} \right) }.
\]
Under \cref{Assumption: Selection-on-observables} and certain regularity conditions on the kernel function $K(\cdot)$, if $\mathcal{X}$ is a compact subset of $\mathbb{R}^p$ and the function $\mu(\cdot, \cdot)$ belongs to the Hölder space $\mathcal{C}^s\left(\mathcal{X} \times \mathcal{B}_Y\right)$ with smoothness parameter
$s$, it follows that
\[
\sup_{x\in \mathcal{X}, \eta \in \mathcal{B}_Y  }\left| \widehat{\mu}_a(x, \eta) - \mu(x,\eta) \right| = O\left( b_n^{s} \right) + O_P \left( \sqrt{ \frac{(p+1)\log b_n}{n b_n^{p+1} } } \right).
\]
With a careful choice of bandwidth, the optimal convergence rates—both uniform and in $L^2$, can be achieved and are given by $(\log n/n)^{s/(2s+p)}$; see \cite{gine2002rates,gine2021mathematical}. The sieve-based approach can also be applied in this context; see \cite{chen2015optimal, belloni2015some, ai2003efficient, blundell2007semi}.
Furthermore, machine learning techniques can be employed to estimate nuisance parameters. The $L^2$-convergence rates for nonparametric regression using deep neural networks have been extensively studied; see \cite{farrell2021deep, kohler2021rate, schmidt2020nonparametric}.

(ii) Alternatively, for $a\in \{0,1\}$ and $x\in\mathcal{X}$, it holds that
\begin{equation}\label{equation: mua}
\begin{aligned}
\mu_a(x, \eta) &=\mathbb{E}\left[\left(Y_i(a)-\eta  \right)_- |X_i = x\right]=\int _{-\infty}^\eta y \mathrm{d} F_a(y|x)-\eta F_a(\eta|x).
\end{aligned}
\end{equation}
This suggests a plug-in estimator based on an estimator of $F_a(y|x)$.
\end{remark}

We conclude this subsection by demonstrating that $ \widehat{\mathbb{V}}_{n}(\theta)$ is a good approximation to $\mathbb{V}_{n}(\theta) = \mathbb{P}_n g_\theta$ with convergence rate faster than $n^{-1/2}$. Consequently, we can ignore the nuisance parameter estimation errors in subsequent asymptotic analysis.

\begin{lemma}\label{Lemma: M_error}
Suppose \cref{Assumption: Selection-on-observables}, \cref{policy size} and \cref{assumption: nuisance parameter convergence rate} hold. If $b_o / 2 < \zeta_e \wedge \zeta_\mu$, then
\[
\mathbb{E}_{P} \left[ \sup_{\theta \in \Theta_n} \left|\widehat{\mathbb{V}}_{n}(\theta) - \mathbb{V}_{n}(\theta) \right | \right] = O(n^{- 1/2} ).
\]
\end{lemma}

\subsection{Asymptotic Upper Regret Bound}
In this subsection, we study the regret upper bound of implementing $\widehat{\pi}_{n}$ under the following assumption.

\begin{assumption}\label{assumption: L^2 boundedness}
 $Y_i(a)$ is $L^2(P)$-bounded, i.e.,  $\mathbb{E}_P \left[ |Y_i (a) |^2 \right] < \infty$ for $a \in \{0,1\}$.
\end{assumption}

For any policy class $\Pi_n$, which may depend on $n$, the regret of deploying a policy $\pi \in \Pi_n$  relative to the best policy in $\Pi_n$, is defined as
\[
\mathrm{Reg}(\pi, \Pi_n) = \max_{\pi^\prime \in \Pi_n} \mathbb{W}_\alpha(\pi^\prime) - \mathbb{W}_\alpha(\pi).
\]
When $\Pi_n$ is clearly understood from the context, we write $\mathrm{Reg}(\pi)  =\mathrm{Reg}(\pi, \Pi_n)$ for notational simplicity. Our primary result regarding the asymptotic regret of our $\alpha$-EWM policy incorporates the following two key quantities:
\[
\begin{aligned}
\Xi := \sup_{\eta \in \mathcal{B}_Y } \mathbb{E} \left| \gamma_\eta(Z_i ) \right|^2  \quad \text{and} \quad  \Xi^\dagger := \sup_{\eta \in \mathcal{B}_Y } \mathbb{E} \left| \Gamma^\dagger_\eta(Z_i ) \right|^2 ,
\end{aligned}
\]
where
\[
\begin{aligned}
\gamma_\eta(z) & :=   \tau(x, \eta) +  \frac{a - e_o(x)}{e_o(x) \left( 1 - e_o(x) \right)}  \left\{ (y-\eta)_{-}   - \mu_{a}(x,\eta)\right\} \mbox{ and }\\
\gamma^\dagger_\eta(z ) & :=  \mu_0(x, \eta)  +  \frac{(1-a)}{1-e_o \left(x \right)} \left\{ (y-\eta)_{-}  -\mu_0(x,\eta) \right\}.
\end{aligned}
\]

\begin{theorem}\label{theorem: regret bound with semiparametric efficiency score}
Suppose \cref{Assumption: Selection-on-observables}, \cref{policy size}, \cref{assumption: nuisance parameter convergence rate},  and \cref{assumption: L^2 boundedness} hold. Let $\bar{K} = 3 + 2/\kappa$. If $\mathbb{E} \left|\gamma_\eta(Z_i) \right|^2 > c_o > 0$ for all $\eta \in \mathcal{B}_Y$, then for $\alpha\in (0,1)$, the following inequality holds:
\begin{equation}\label{equation: semiparametric efficiency score regret bound}
\limsup_{n \rightarrow \infty} \frac{
 \mathbb{E}\left[   \mathrm{Reg} \left( \widehat{\pi}_{n} \right )    \right] }{   \sqrt{\mathrm{VC}(\Pi_n)  / n}   }   \leq \frac{30}{\alpha}  \sqrt{ \Xi + \Xi^\dagger } + 72 \sqrt{ (\bar{K}/ \alpha + 1)^2 + \Xi /\alpha^2 }     .
\end{equation}
\end{theorem}

\cref{theorem: regret bound with semiparametric efficiency score} complements Theorem 1 in \cite{athey2021policy} for $1$-EWM policy.\footnote{\cite{athey2021policy} also allow for an approximate optimal policy.}  The constant in
\cref{theorem: regret bound with semiparametric efficiency score} depends on $\alpha$: it increases as $\alpha$ decreases, partly due to estimation error. Specifically, estimating the average welfare of the $\alpha$-worst-affected group makes use of only an $\alpha$-fraction of the total sample, leading to greater instability in welfare estimation.

\begin{remark}
Suppose that $\mathcal{B}_Y \subseteq \left[-\eta_B, \eta_B \right]$ for some $\eta_B > 0$. Under the strict overlap condition in \cref{assumption: nuisance parameter convergence rate}, it follows that $\Xi$ and $\Xi^\dagger$ can be upper bounded as
\[
\begin{aligned}
\Xi  \leq  \left(1 + \frac{2}{\kappa}  \right) \left( \mathbb{E} |Y_i(0)|^2 + \mathbb{E} |Y_i(1)|^2  + \eta_B \right) \quad \text{and} \quad
\Xi^\dagger   \leq  \left(1 + \frac{2}{\kappa}  \right) \left( \mathbb{E} |Y_i(0)|^2  + \eta_B \right).
\end{aligned}
\]
\end{remark}

\begin{remark}\label{remark: near-optimal solution}
Recall that we learn the optimal policy by simultaneously solving out $\widehat{\pi}_{n}$ and $\widehat{\eta}_{n}$ from $\max_{(\pi, \eta)\in \Pi_n \times \mathcal{B}_Y } \widehat{\mathbb{V}}_{n}( \pi, \eta)$.
Let $\widehat{\theta} \equiv (\widehat{\pi}, \widehat{\eta}) \in \Pi_n \times \mathcal{B}_Y$ denote any near-optimal solution satisfying
\[
\widehat{\mathbb{V}}_{n}(\widehat{\theta}) \geq \sup_{\theta \in \Theta_n }\widehat{\mathbb{V}}_{n}(\theta)  - o_P\left( r_n \right),
\]
where $r_n =  \sup_{\theta \in \Theta_n} \left|\widehat{\mathbb{V}}_{n}(\theta) - \mathbb{V}_{n}(\theta) \right | $.
In fact, \cref{theorem: regret bound with semiparametric efficiency score} holds if the exact optimizer  $\widehat{\pi}_{n}$ is replaced by any near-optimal welfare maximizer $\widehat{\pi}$ and $r_n = o_P(n^{-1/2})$.
The term $o_P(r_n)$ enables us to find an approximate solution to $\max_{\theta \in \Theta_n} \widehat{\mathbb{V}}_{n} (\theta) $, which is particularly useful when the optimization is non-concave.
\end{remark}

\subsubsection{Technical Comparisons with \cite{kitagawa2018a, athey2021policy}}

The regret bounds in \cref{theorem: regret bound with semiparametric efficiency score} and those in \cite{kitagawa2018a, athey2021policy}  are all of order $\sqrt{\mathrm{VC}(\Pi_n) / n}$. In addition, \cref{theorem: regret bound with semiparametric efficiency score}  and Theorem 1 in \cite{athey2021policy} provide explicit expressions for the constants which require more delicate technical proofs than  \cite{kitagawa2018a, kitagawa2021equality}.

Following \cite{kitagawa2018a, kitagawa2021equality}, the proof of the order of the regret bounds of $\widehat{\pi}_n$ relies on the lemma below.

\begin{lemma}\label{lemma:REG} Suppose \cref{Assumption: Selection-on-observables}, \cref{policy size} and \cref{assumption: nuisance parameter convergence rate} hold. If $b_o / 2> \zeta_e \wedge \kappa_\mu$, then
    \[\mathrm{Reg}(\widehat{\theta}_{n}) \leq 2 \sup _{\theta \in \Theta_n}\left|\left(\mathbb{P}_n-P\right) g_\theta\right| + r_n,\]
    where $r_n=o_P(n^{-1/2}).$
\end{lemma}

\cref{lemma:REG} implies that it is sufficient to study the concentration of the empirical process: $$\mathbb{V}_{n}(\theta) - \mathbb{V}(\theta) = (\mathbb{P}_n - P) g_\theta \
\mbox{ over } \ \theta \in \Theta_n. $$
In contrast to \cite{kitagawa2018a,kitagawa2021equality} and \cite{athey2021policy}, the score function for the $\alpha$-expected welfare  $g_\theta$ is nonlinear in $\theta$ rendering the VC dimension of the function class $\mathcal{G}_{\Theta_n} :=\{g_\theta : \theta \in \Theta_n\}$ difficult to derive. Instead of exploiting the VC dimension of the corresponding function classes as in \cite{kitagawa2018a,kitagawa2021equality} and \cite{athey2021policy}, we directly upper bound the covering number of $\mathcal{G}_{\Theta_n}$ and then apply the classic empirical process maximal inequality, such as Theorem 2.14.1 in \cite{vaart2023empirical}.
\begin{lemma}\label{lemma: covering number of G_theta}
If \cref{assumption: L^2 boundedness} holds,  then there is an envelope function $G$ for $\mathcal{G}_{\Theta_n}$ and constant $c_o > 0$ not depending on $n$ and $p$ such that
\[
N\left( \epsilon  \|G \|_{Q,2} , \mathcal{G}_{\Theta_n}, L^2(Q)\right) \leq\left(c_o / \epsilon\right)^{24 \mathrm{VC}(\Pi_n)+48}, \quad \forall \epsilon > 0,
\]
for all finite discrete probability measures $Q$ on $\mathcal{Z}$.
\end{lemma}
\Cref{assumption: L^2 boundedness} and \Cref{assumption: Strict Overlap} ensure the existence of an envelope function that is bounded in $L^2(P)$.
Applying Theorem 2.14.1 in \cite{vaart2023empirical} and \cref{lemma: covering number of G_theta}, we conclude that there is a universal constant $c_o > 0$ not depending on $n$ such that
\begin{equation}\label{equation:order}
\mathbb{E}_{P} \left[ \sup_{\theta \in \Theta_n}  \left| (\mathbb{P}_n - P) g_{\theta} \right| \right] \leq  c_o\sqrt{ \mathrm{VC}(\Pi_n)/ n}.
\end{equation}

Compared with \cite{kitagawa2018a,kitagawa2021equality}, one of the technical challenges addressed by \cite{athey2021policy} on $1$-EWM policy lies in handling the doubly robust estimator of the welfare function. They show that as long as $\mathrm{VC}(\Pi_n)$ does not grow too rapidly with $n$, the use of cross-fitting and ML/nonparametric estimation of nuisance parameters results in a regret bound of the order $\sqrt{\mathrm{VC}(\Pi_n)/n}$.
Building on \cite{kitagawa2018a,kitagawa2021equality} and  \cite{athey2021policy} on $1$-EWM policy, we establish an upper bound for $\alpha$-EWM for any $\alpha\in (0,1)$ with an explicit expression for the constant $c_o$ in \cref{equation:order}.
Similar to \cite{athey2021policy}, we employ a classical chaining argument to derive an upper bound for the Rademacher complexity of the score function class. However, due to the nonlinearity of score function $g_\theta$ with respect to $\theta$, the slicing technique used in \cite{athey2021policy} is difficult to implement. Instead, we introduce a new conditional semi-metric  and apply the classical Dudley's chaining argument to directly bound the Rademacher complexity of $\mathcal{G}_{\Theta_n}$. We refer interested reader to \cref{section: Proof of regret bound with semiparametric efficiency score} for details.

\section{Inference for the Optimal Welfare}
\label{section:inference}

In this section, we develop asymptotically valid inference for the optimal $\alpha$-expected welfare. Compared with regret bounds, inference on optimal welfare is lacking even for $1$-EWM except for the first-best policy; see  \cite{luedtke2016statistical,luedtke2018parametric,shi2020breaking}, and Appendix B in the supplemental material to \cite{kitagawa2018a}.

We first impose conditions including the uniqueness of the optimal solution denoted as $\theta_o$ to ensure asymptotic normality of $\sup_{\theta \in \Theta}\widehat{\mathbb{V}}_{n}(\theta )$ based on which we construct Wald-type inference. We then summarize a general inference procedure that relaxes the uniqueness assumption. A detailed treatment of the general inference procedure is postponed to \cref{section: inference for the optimal value}.

For simplicity, we assume that the policy class does not change with the sample size $n$, i.e., $\Pi_n = \Pi$ for all $n$, and write $\Theta = \Pi \times \mathcal{B}_Y$.  We define a metric space $(\Theta, \| \cdot \|)$, where
\[
\left\| \theta_1 - \theta_2 \right\| \equiv  |\eta_1- \eta_2|+ \|\pi_1 - \pi_2\|_{P, 2}  = |\eta_1- \eta_2| + \sqrt{ \mathbb{E} |\pi_1(X_i) - \pi_2(X_i)|^2}.
\]
for any $\theta_1, \theta_2 \in \Theta$. This premise will be upheld throughout the subsequent analysis.

\subsection{Assumptions}
We establish asymptotic normality under two assumptions, the bounded support assumption and the uniqueness assumption.

\begin{assumption}\label{assumption: assumption for inference}
\begin{assumpenum}
\item \label{Assumption: Bounded support}
The outcome $Y_i = Y_i(A_i)$ has bounded support, i.e., $\mathbb{P}\left( | Y_i | \leq c_o \right)  = 1$ for some constant $c_o > 0$.
 \item  \label{Assumption: fixed policy class}  The policy class $\Pi$ has finite VC-dimension, i.e.,  $\mathrm{VC}(\Pi) < \infty$.
\end{assumpenum}
\end{assumption}


\cref{assumption: assumption for inference} is widely adopted in policy learning research, see, e.g., \cite{kitagawa2018a, kitagawa2021equality, rai2018statistical, kallus2018confounding, luedtke2016statistical, luedtke2020performance}.\footnote{Although studies like \cite{athey2021policy} do not adopt this assumption for regret bounds, it substantially simplifies the technical analysis for statistical inference. }
\cref{Assumption: Bounded support} implies that the feasible set $\mathcal{B}_Y$ of the dual reformulation of $\mathbb{W}_\alpha(\pi)$ can be restricted to $[-c_o, c_o]$  and the regression functions  $|\mu_a(x, \eta)| \leq 2c_o$ for all $\eta \in \mathcal{B}_Y$ and $a \in \{0,1\}$. Moreover, the functions $g_\theta(\cdot)$ are also uniformly bounded, i.e., $\sup_{\theta \in \Theta}\| g_{\theta} \|_{\infty} < \infty$.

\begin{assumption}[Uniqueness]
\label{Assumption: V_theta smoothness (1)}
There exists a $\theta_o \equiv (\pi_o, \eta_o) \in \Theta$ such that for all $\epsilon > 0$, $   \mathbb{V}(\theta_o) > \sup \{  \mathbb{V}(\theta): \theta \in \Theta,  \| \theta - \theta_o \| > \epsilon   \}$.
\end{assumption}
\cref{Assumption: V_theta smoothness (1)} is a standard condition in extremum estimation. It ensures that $\theta_o\in \Theta$ is a unique and well-separated point of maximum of $\theta \mapsto \mathbb{V}(\theta)$. Lemma 14.4 in \cite{kosorok2008introduction} gives some sufficient conditions for this assumption. If for all $\epsilon > 0$, $   \mathbb{W}(\pi_o) > \sup_{\pi: \| \pi - \pi_o \| > \epsilon }   \mathbb{W}(\pi)$  and $Y_i(\pi)$ has positive density at $\mathrm{VaR}_\alpha(Y_i(\pi))$ for all $\pi \in \Pi$, then \cref{Assumption: V_theta smoothness (1)} is satisfied. For policy learning, \cref{Assumption: V_theta smoothness (1)} is strong, although it is adopted in \cite{wang2018quantile}, Section 2.3 of \cite{kitagawa2018a}, and Section 2.3 of \cite{luedtke2020performance}.
\begin{remark}
For $1$-EWM, uniqueness of the first-best optimal policy excludes a special class of distributions known as exceptional distributions. For  $\alpha$-EWM with $\alpha\in(0,1)$, we show in \cref{lemma: first best policy} that the first best policy is given by $\pi_o = \mathds{1}\{ \tau(x, \eta_o) \geq 0 \}$ with $\eta_o = \eta_{\mathrm{FB}}^*$ defined in \cref{lemma: first best policy}.  \cref{Assumption: V_theta smoothness (1)} excludes the class of exceptional distributions for which $\mathbb{P}\left[\tau(X_i, \eta_o)=0  \right] > 0$. This is because \cref{Assumption: V_theta smoothness (1)} implies that $\theta_o = (\pi_o, \eta_o)$ is the unique and well separated maximizer. As a result,
\[
 \mathds{1}\{ \tau(X_i, \eta_o) \geq 0 \} =  \mathds{1}\{ \tau(X_i, \eta_o) > 0 \},\quad P\text{-a.s.},
\]
and $\mathbb{P}( \tau(X_i, \eta_o) = 0 ) = 0$.
\end{remark}

\subsection{Asymptotic Normality}
To establish asymptotic normality of $\widehat{\mathbb{V} }_{n} (\widehat{\theta}_{n}  )$,  consider the following decomposition:
\begin{equation}\label{equation: V_DML decomposition}
\begin{aligned}
\widehat{\mathbb{V} }_{n} (\widehat{\theta}_{n}  ) -  \mathbb{V}(\theta_o)  =  \underbrace{ \widehat{\mathbb{V} }_{n} (\widehat{\theta}_{n}  ) - \mathbb{V} _{n} (\widehat{\theta}_{n}  ) }_{= o_P(n^{-1/2})}
 + \underbrace{ \mathbb{V} _{n} (\widehat{\theta}_{n}  )  -   \mathbb{V}  (\widehat{\theta}_{n}  )   }_{\approx \mathbb{V} _{n} (\theta_o  )  -   \mathbb{V}  (\theta_o )    } + \underbrace{ \mathbb{V}  (\widehat{\theta}_{n}  )   -  \mathbb{V}  (\theta_o) }_{= -\mathrm{Reg}\left(\widehat{\pi}_{n}, \Pi\right)}.
\end{aligned}
\end{equation}
Note that the first term on the RHS of \cref{equation: V_DML decomposition} is $o_P(n^{-1/2})$ due to \cref{Lemma: M_error}.

In the rest of this section, we will show that
\begin{enumerate}
    \item[(i)] the second term on the RHS of \cref{equation: V_DML decomposition} is asymptotically equivalent to $\mathbb{V}_{n}(\theta_o) - \mathbb{V}(\theta_o)$;
    \item[(ii)] the third term on the RHS of \cref{equation: V_DML decomposition} is of order $o_P(n^{-1/2})$.
\end{enumerate}
Consequently,
\[
\begin{aligned}
\sqrt{n} \left[\mathbb{V}_{n} (\widehat{\theta}_{n}  )  -   \mathbb{V}  (\theta_o) \right] & = \sqrt{n} \left[ \mathbb{V} _{n} (\theta_o  )  -   \mathbb{V}  (\theta_o ) \right] + o_P(1)  \\
& = \sqrt{n} (\mathbb{P}_n - P) g_{\theta_o}  + o_P(1).
\end{aligned}
\]
and asymptotic normality follows.

To show (i), we first prove $\|\widehat{\theta}_{n}  - \theta_o\| = o_P(1)$ in \cref{lemma: consistency} below. Since $\mathcal{G}_\Theta \equiv \{g_\theta : \theta \in \Theta\}$ is $P$-Donsker by \cref{lemma: covering number of G_theta}, (i) follows.

\begin{lemma}\label{lemma: consistency}
Under \cref{Assumption: Selection-on-observables}, \cref{assumption: nuisance parameter convergence rate},  \cref{assumption: assumption for inference} and \cref{Assumption: V_theta smoothness (1)}, it holds that $\|\widehat{\theta}_{n} -\theta_o \|=o_P(1)$.
\end{lemma}

To show (ii),  we note that
\[
\begin{aligned}
\mathrm{Reg}( \widehat{\pi}_{n} )  & = \mathbb{V}(\theta_o) - \mathbb{V}( \widehat{\theta}_n  ) = \mathbb{V}(\theta_o) - \widehat{\mathbb{V}}_{n} (\theta_o) + \widehat{\mathbb{V}}_{n} (\theta_o) - \widehat{\mathbb{V}}_{n} (\widehat{\theta} _{n} ) + \widehat{\mathbb{V}}_{ n} (\widehat{\theta}_{n} )  - \mathbb{V}(\widehat{\theta}_{ n} )  \\
& \leq \mathbb{V}(\theta_o) - \mathbb{V}_{n} (\theta_o)  + \mathbb{V}_{ n} (\widehat{\theta}_{ n} )  - \mathbb{V}(\widehat{\theta}_{ n} )  + r_n  \\
& = (\mathbb{P}_n - P) ( g_{\widehat{\theta}_{n}  } - g_{\theta_o} ) + r_n,
\end{aligned}
\]
where the inequality follows from $\widehat{\mathbb{V}}_{n} (\theta_o) - \widehat{\mathbb{V}}_{n} (\widehat{\theta} _{ n} ) \leq 0$. Similar to \cite{luedtke2020performance}, one can show that under mild conditions including boundedness and uniqueness, asymptotic equicontinuity arguments ensure that $(\mathbb{P}_n - P) ( g_{\widehat{\theta}_n} - g_{\theta_o} ) = o_P(n^{-1/2})$ for any policy class $\Pi$ satisfying $\text{VC}(\Pi)<\infty$.

Summing up, we obtain asymptotic normality of $\widehat{\mathbb{V}}_{n}(\widehat{\theta}_{n}  )$.

\begin{theorem}\label{Theorem: approximation of DML M_n}
Suppose conditions in \cref{lemma: consistency} hold. Then,
\[
\begin{aligned}
\widehat{\mathbb{V}}_{n}(\widehat{\theta}_{n}  )  -  \mathbb{V}_P (\theta_o) &=  (\mathbb{V}_{n} - \mathbb{V}_P )(\theta_o)  + o_P(n^{-1/2})\\
& = \frac{1}{n} \sum_{i=1}^n \left\{ g_{\theta_o} (Z_i) - \mathbb{E}_{P} [g_{\theta_o}(Z_i)] \right\} + o_P(n^{-1/2})   ,
\end{aligned}
\]
where the function $g_{\theta_o}$, defined in \cref{equation: g_theta}, is evaluated at $\theta = \theta_o$. In particular,
\[
\sqrt{n} \left[\widehat{\mathbb{V}}_{n}(\widehat{\theta}_{n}  )  -  \mathbb{V}(\theta_o)  \right] \rightsquigarrow N\left( 0, \sigma_o^2 \right),
\]
where $\sigma_o^2 = \mathrm{Var}\left[g_{\theta_o}(Z_i) \right]$.
\end{theorem}
\begin{remark}
Drawing on \cite{newey1994asymptotic, luedtke2016statistical}, and under the assumptions stated in \cref{Theorem: approximation of DML M_n} and other mild conditions, our optimal welfare estimator achieves semiparametric efficiency bound.
\end{remark}
The next theorem presents a consistent estimator of the asymptotic variance $\sigma_o^2$.
\begin{theorem} \label{theorem: variance estimation}
Consider the following estimator of $\sigma_o^2$:
\[
\widehat{\sigma}_o^2 =   \frac{1}{n} \sum_{i=1}^n \left[ g_{\widehat{\theta}_{n}  }  \big(Z_i ; \widehat{\mu}^{(-k(i))} , \widehat{e}^{(-k(i))} \big )  \right]^2   -   \left[  \frac{1}{n} \sum_{i=1}^n  g_{\widehat{\theta}_{n}  }  \big(Z_i ; \widehat{\mu}^{(-k(i))} , \widehat{e}^{(-k(i))}  \big)  \right]^2.
\]
Under the conditions of  \cref{lemma: consistency},
it holds that $\widehat{\sigma}_o^2 = \sigma_o^2 + o_P(1)$ and
\[
\sqrt{n} \widehat{\sigma}_o^{-1} \left[\widehat{\mathbb{V}}_{n}(\widehat{\theta}_{n}  )  -  \mathbb{V}(\theta_o)  \right] \rightsquigarrow N\left( 0, 1 \right).
\]
\end{theorem}

\subsection{Uniform Inference}

\cref{section: inference for the optimal value} develops uniform inference for the optimal welfare without \cref{Assumption: V_theta smoothness (1)}. It improves upon the inference proposed in Appendix B in the supplemental material to \cite{kitagawa2018a}. We provide a summary of the procedures here and refer to interested reader to \cref{section: inference for the optimal value} for technical details.

Define the supremum functional  $\psi: \ell^{\infty} (\Theta) \rightarrow \mathbb{R}$ as $\psi: h \mapsto \sup_{\theta \in \Theta}  h (\theta)$.
Consider the multiplier bootstrap $ \widehat{\mathbb{G}}_{n}^*: \Theta \rightarrow \mathbb{R}$ defined as
\begin{equation}\label{equation: multiplier boostrap process}
\widehat{\mathbb{G}}_{n}^*: \theta \mapsto  n^{-1/2} \sum_{i=1}^n  \xi_i \left [ \widehat{g}_{\theta}(Z_i) - \widehat{\mathbb{V}}_n (\theta) \right ],
\end{equation}
where $\{\xi_i\}_{i=1}^n$ are i.i.d. random variables independent of $(Z_i)_{i=1}^n$, with $\mathbb{E}( \xi_i) = 0$, $\mathbb{E}(\xi_i^2) =1$ and $\mathbb{E}\left[ \exp |\xi_i| \right] < \infty$.
For given $\epsilon_n = o(1)$ with $n^{1/2} \epsilon_n \rightarrow \infty$, let
\begin{equation}\label{equation: numerical boostrap}
\begin{aligned}
\widehat{\psi}_n^{\prime} ( \widehat{\mathbb{G}}_n^* )  &= \frac{ \psi \big(\widehat{\mathbb{V}}_n + \epsilon_n  \widehat{\mathbb{G}}_n^* \big )  - \psi(\widehat{\mathbb{V}}_n )    }{  \epsilon_n }.
\end{aligned}
\end{equation}

For any $\gamma \in (0,1)$, let $c_{\gamma}$ denote the $\gamma$-empirical quantile of $\widehat{\psi}_n^{\prime} (\widehat{\mathbb{G}}_n^*)$ which can be obtained from a large number of bootstrap samples. The one-sided confidence interval at the desired level $\gamma$ is
\begin{equation}\label{equaton: one-sided_CI}
\left[ \sup_{\theta \in \Theta} \widehat{\mathbb{V}}_n (\theta) - c_{1-\gamma}/\sqrt{n}, \infty \right),
\end{equation}
with correct asymptotic coverage:
\[
\lim_{n \rightarrow \infty} \inf_{P \in \mathcal{P}_n} \mathbb{P} \left[ \mathbb{V}_P(\theta_o) \geq \sup_{\theta \in \Theta} \widehat{\mathbb{V}}_n (\theta) - c_{1-\gamma} /\sqrt{n}  \right] \geq 1- \gamma,
\]
where $\mathcal{P}_n$ is a collection of distributions satisfying some regularity conditions specified in \cref{assumption: assumption for inference} in \cref{section: inference for the optimal value}.
Define $q_{1-\gamma}$ as the $(1-\gamma)$-empirical quantile of $\left|\widehat{\psi}_n^{\prime} ( \widehat{\mathbb{G}}_n^* ) \right|$ for any $\gamma > 0$. The corresponding two-sided confidence interval is
\begin{equation}\label{equaton: two-sided_CI}
\left[ \sup_{\theta \in \Theta } \widehat{\mathbb{V}}_n (\theta) - q_{1- \gamma} / \sqrt{n},  \sup_{\theta \in \Theta } \widehat{\mathbb{V}}_n (\theta) +  q_{1- \gamma} / \sqrt{n}   \right],
\end{equation}
which attains the correct asymptotic coverage for any fixed distribution $P \in \mathcal{P}_n$:
\[
\liminf_{n \rightarrow \infty}  \mathbb{P} \left[ \left| \sup_{\theta \in \Theta} \widehat{\mathbb{V}}_n (\theta) -  \mathbb{V}(\theta_o) \right|  \leq q_{1-\gamma} /\sqrt{n} \right] \geq 1-\gamma.
\]


\section{Empirical Application and Simulations}
This section presents extensive numerical results on the finite sample performance of our debiased estimator and proposed inference using both real data and synthetic data.\footnote{
Data and codes for this section can be accessed at
 \href{https://github.com/yqi3/alpha-EWM}{\texttt{https://github.com/yqi3/alpha-EWM}}.}

\label{Section: empirical}
\subsection{The JTPA Study}
\label{Section: JTPA}
\cite{kitagawa2018a} apply $1$-EWM method to experimental data from the National Job Training Partnership Act (JTPA) Study. The study randomized whether applicants are eligible to receive training and job-search assistance provided by the JTPA. The pre-treatment covariates included in the data are years of education (\textit{edu}) and pre-program earnings (\textit{prevearn}) and the outcome variable is an applicant's earnings 30 months after the assignment (\textit{earnings}). The sample size is 9,223 and the propensity score is known to be $2/3$. We adopt this data studied by \cite{kitagawa2018a} and, similar to \cite{kitagawa2018a}, we analyze welfare from an intent-to-treat standpoint, considering hypothetically making available the training program to eligible individuals, who may decline it. For detailed data description and evaluation of average program effects, we refer the reader to \cite{bloom1997benefits}.

We consider three policy classes: simple (treat all or none) and linear with and without squared and cubic $edu$. More specifically, the two linear policy classes take the form
\begin{equation}
\label{JTPA LES policy class}
    \Pi_{\mathrm{LES}}\vcentcolon=\left\{\{x:\beta_0+\beta_1 edu+\beta_2 prevearn>0\}, (\beta_0,\beta_1,\beta_2)\in\mathbb{R}^3\right\}\text{ and}
\end{equation}
\begin{equation}
\label{JTPA LES policy class 3}
    \Pi^3_{\mathrm{LES}}\vcentcolon=\left\{\begin{array}{lr}\{x:\beta_0+\beta_1 edu+\beta_2 prevearn+\beta_3edu^2+\beta_4edu^3>0\}, \\
\quad\quad\quad\quad\quad\quad\quad(\beta_0,\beta_1,\beta_2,\beta_3,\beta_4)\in\mathbb{R}^5 \end{array} \right\}.
\end{equation}
We investigate $\alpha\in\mathcal{A}\vcentcolon=\{0.25, 0.3, 0.4, 0.5, 0.8\}$. We recommend that researchers interested in the $\alpha=1$ case consider the $1$-EWM in \cite{kitagawa2018a} directly. For each $\alpha \in \mathcal{A}$ and policy class, we estimate $\mu_a(x, \eta)=\mathbb{E}\left[\left(Y_i(a) - \eta\right)_- \mid X_i = x\right] \text{ for } a \in \{0,1\}$ and a given $\eta$, using random forests (RF) developed by \cite{athey2019generalized}. We then apply simulated annealing (SA), proposed by \cite{kirkpatrick1983optimization}, to select the combination of parameters that (approximately) maximizes the objective function.\footnote{We build RF using \texttt{regression\_forest()} in \texttt{R} package \texttt{grf} and implement SA using \texttt{optim\_sa()} in the \texttt{R} package \texttt{optimization} \citep{athey2019generalized, husmann2017r}. We use default tuning parameters for RF. For SA, the specifications are more problem-specific. A good strategy is to plot the loss function and inspect if there is sufficient evidence of convergence.} SA is a derivative-free probabilistic optimization algorithm aiming at finding approximate solutions by iteratively exploring the solution space and gradually decreasing the probability of accepting worse solutions as the algorithm progresses.\footnote{\cite{geman1984stochastic} prove convergence of \textit{generic} SA to a global optimum, provided that the probability of accepting worse solutions shrinks sufficiently slowly, and that all elements in the solution space are equally probable as the number of training epochs goes to infinity.}

Estimation and inference results for $\mathbb{W}_\alpha(\pi_o)$ are organized in Table \ref{Table: JTPA}. The first two columns consist of the class of simple policies and serve as baselines for $\Pi_{\text{LES}}$ and $\Pi^3_{\text{LES}}$ in the third and fourth columns. Detailed expressions for the optimal policies can be found in \cref{Section: JTPA additional}. The observed increase in $\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$ across panels reflects that, as $\alpha$ grows, the lower-tail subpopulation expands to include relatively better outcomes. This raises the average and thus increases the $\alpha$-expected welfare. The percentage of treated individuals tends to increase with $\alpha$ as well. The \(95\%\) confidence intervals (CIs) constructed using normal inference in Algorithm \ref{alg:debiased} are reported in the third row of each panel in Table \ref{Table: JTPA}, and the 95\% CIs from uniform inference obtained via multiplier bootstrap with \(\epsilon = n^{-1/4}\) and $B=100$ are presented in the last row of each panel. For each combination of $\alpha$ and policy class, the CI from uniform inference is wider than that from normal inference. While we cannot verify uniqueness, a simulation study calibrated to the JTPA sample in Section \ref{Section: WGAN simulations} finds that the Wald-type CIs achieve approximately \(95\%\) coverage, offering supporting evidence for their validity in this application.

\begin{table}[H] \centering
  \scriptsize{
\begin{tabular}
{@{\extracolsep{0pt}}lcccc}
 & \multirow{2}{*}{\textbf{Treat None}} & \multirow{2}{*}{\textbf{Treat All}} & \multirow{2}{*}{\textbf{Linear}} & \textbf{Linear with} \\
 &  &  &  & \textbf{$edu^2$ and $edu^3$} \\
\\[-1.5ex]\hline
\multicolumn{5}{c}{\cellcolor{blue!20}$\textbf{Panel 1: }\boldsymbol{\alpha}\boldsymbol{=0.25}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{$\%$ treated} & 0\% & 100\% & 34.761\% & 32.896\%\\
\multicolumn{1}{l}{$\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$} & 376.968 & 451.027 & 530.630 & 546.300
\vspace{0.075cm}\\
\multicolumn{1}{l}{95\% CI (normal)} & $(298.567, 455.368)$ & $(372.626, 529.427)$ & $(439.331, 621.930)$ & $(446.461, 646.138)$\\
\multicolumn{1}{l}{95\% CI ($\epsilon=n^{-1/4}$)} & $(-48.098, 802.033)$ & $(154.412, 747.641)$ & $(146.400, 914.860)$ & $(155.773, 936.826)$\\
\\[-2ex]\hline
\multicolumn{5}{c}{\cellcolor{blue!20}$\textbf{Panel 2: }\boldsymbol{\alpha}\boldsymbol{=0.3}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{$\%$ treated} & 0\% & 100\% & 50.992\% & 32.820\%\\
\multicolumn{1}{l}{$\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$} & 695.647 & 838.930 & 917.718 & 918.011
\vspace{0.075cm}\\
\multicolumn{1}{l}{95\% CI (normal)} & $(585.617, 805.678)$ & $(728.900, 948.961)$ & $(793.695, 1041.741)$ & $(776.708, 1059.315)$\\
\multicolumn{1}{l}{95\% CI ($\epsilon=n^{-1/4}$)} & $(152.922, 1238.373)$ & $(490.538, 1187.322)$ & $(457.579, 1377.858)$ & $(506.934, 1329.088)$\\
\\[-2ex]\hline
\multicolumn{5}{c}{\cellcolor{blue!20}$\textbf{Panel 3: }\boldsymbol{\alpha}\boldsymbol{=0.4}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{$\%$ treated} & 0\% & 100\% & 82.392\% & 81.969\%\\
\multicolumn{1}{l}{$\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$} & 1647.506 & 1947.011 & 2038.321 & 2039.468
\vspace{0.075cm}\\
\multicolumn{1}{l}{95\% CI (normal)} & $(1468.631, 1826.381)$ & $(1768.137, 2125.886)$ & $(1845.888, 2230.754)$ & $(1840.260, 2238.676)$\\
\multicolumn{1}{l}{95\% CI ($\epsilon=n^{-1/4}$)} & $(995.201, 2299.812)$ & $(1519.072, 2374.951)$ & $(1477.132, 2599.510)$ & $(1516.364, 2562.573)$\\
\\[-2ex]\hline
\multicolumn{5}{c}{\cellcolor{blue!20}$\textbf{Panel 4: }\boldsymbol{\alpha}\boldsymbol{=0.5}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{$\%$ treated} & 0\% & 100\% & 83.400\% & 83.379\%\\
\multicolumn{1}{l}{$\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$} & 2981.034 & 3419.311 & 3524.651 & 3527.108
\vspace{0.075cm}\\
\multicolumn{1}{l}{95\% CI (normal)} & $(2746.431, 3215.638)$ & $(3184.708, 3653.915)$ & $(3274.440, 3774.861)$ & $(3269.096, 3785.121)$\\
\multicolumn{1}{l}{95\% CI ($\epsilon=n^{-1/4}$)} & $(2233.145, 3728.923)$ & $(2910.270, 3928.352)$ & $(2951.684, 4097.617)$ & $(2898.115, 4156.101)$\\
\\[-2ex]\hline
\multicolumn{5}{c}{\cellcolor{blue!20}$\textbf{Panel 5: }\boldsymbol{\alpha}\boldsymbol{=0.8}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{$\%$ treated} & 0\% & 100\% & 86.783\% & 79.204\%\\
\multicolumn{1}{l}{$\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$} & 8671.975 & 9522.451 & 9661.526 & 9690.607
\vspace{0.075cm}\\
\multicolumn{1}{l}{95\% CI (normal)} & $(8326.551, 9017.398)$ & $(9177.028, 9867.874)$ & $(9292.969, 10030.082)$ & $(9309.569, 10071.646)$\\
\multicolumn{1}{l}{95\% CI ($\epsilon=n^{-1/4}$)} & $(7816.114, 9527.835)$ & $(8876.617, 10168.285)$ & $(8940.210, 10382.840)$ & $(8983.668, 10397.546)$\\
\\[-2ex]\hline
\end{tabular}
}\\
\caption{Estimated $\mathbb{W}_\alpha(\pi_o)$ for different $\alpha$'s and policy classes that condition on \textit{edu} and \textit{prevearn}. Baseline results for treating none or all of the individuals are shown in the first two columns. The third and fourth rows of each panel report the $95\%$ CI based on normal and uniform inference, respectively. }
\label{Table: JTPA}
\end{table}

Examining the point estimates of welfare, we see that for all $\alpha\in\mathcal{A}$, a simple policy of treating all outperforms treating none. Moreover, relative to treating all, there is a considerable increase in the targeted welfare generated by the optimal policy of class $\Pi_{\mathrm{LES}}$. Linear policies with $edu^2$ and $edu^3$ only bring tiny welfare improvements. Figures \ref{Figure: JTPA linear} and \ref{Figure: JTPA cubic} highlight the optimal treatment regions. Following \cite{kitagawa2018a}, we bin the individuals by $(edu,prevearn)$, and the number of individuals with each combined characteristic is represented by the size of the corresponding dot.
\vspace{0.5cm}
\begin{figure}[H]
    \centering
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.25linear.jpeg}}
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.3linear.jpeg}}
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.4linear.jpeg}}
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.5linear.jpeg}}
        \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.8linear.jpeg}}
        \caption{Optimal policies from the linear class $\Pi_{\mathrm{LES}}$ conditioning on \textit{edu} and \textit{prevearn}. The number of individuals with characteristics closest to each $(edu,prevearn)$ in the grid is represented by the size of the corresponding dot. $\alpha\in\{0.25,0.3,0.4,0.5,0.8\}$. }
    \label{Figure: JTPA linear}
\end{figure}

\vspace{-0.75cm}
\begin{figure}[H]
    \centering
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.25cubic.jpeg}}
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.3cubic.jpeg}}
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.4cubic.jpeg}}
    \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.5cubic.jpeg}}
        \subfigure{
        \includegraphics[width=0.48\textwidth]{plots/0.8cubic.jpeg}}

    \caption{Optimal policies from the linear class $\Pi_{\mathrm{LES}}^3$ conditioning on \textit{edu}, \textit{prevearn}, $edu^2$, and $edu^3$. The number of individuals with characteristics closest to each $(edu,prevearn)$ in the grid is represented by the size of the corresponding dot. $\alpha\in\{0.25,0.3,0.4,0.5,0.8\}$. }
    \label{Figure: JTPA cubic}
\end{figure}

Tables \ref{Table: JTPA comparisons linear} and \ref{Table: JTPA comparisons cubic} examine welfare gains and losses as we switch between different targeting policies and estimate the resulting welfare of different targeted subpopulations. For example, the first row in Table \ref{Table: JTPA comparisons linear} shows the estimated welfare of the worst-off $25\%$ of the population when the optimal linear policies are targeting the worst-off $25\%$, $30\%$, $40\%$, $50\%$, and $80\%$, respectively. The diagonal entries (i.e., the row maximums) are highlighted as these optimal policies are targeting the actual subpopulations of interest. Tables \ref{Table: JTPA comparisons linear} and \ref{Table: JTPA comparisons cubic} demonstrate a valuable strength of our method, as we are able to conduct rich policy evaluations by estimating the expected welfare at any $\alpha$ for any given policy. In other words, even when a policy is not targeting the worst-affected $(\alpha\times100)\%$, we can still evaluate its performance at $\alpha$ to obtain a clear picture of the trade-offs, which opens up possibilities for learning policies that promote greater equality across subpopulations.

From Table \ref{Table: JTPA comparisons linear} below and Table \ref{Table: JTPA linear welfare loss} in \cref{Section: JTPA additional}, adopting the linear policy that targets $\alpha'=0.8$ leads to an $11.9\%$ decrease in the average welfare of the worst-affected quarter of the population ($\alpha=0.25$), compared to implementing the optimal linear policy \textit{targeting} the worst-affected quarter ($\alpha=\alpha'=0.25$). Conversely, adopting the policy targeting the worst-affected quarter ($\alpha'=0.25$) only leads to a $5.3\%$ decrease in the $0.8$-expected welfare ($\alpha=0.8$) relative to implementing the optimal policy targeting the worst-affected $80\%$ ($\alpha=\alpha'=0.8$). In Table \ref{Table: JTPA comparisons cubic}, similar patterns emerge with the inclusion of $edu^2$ and $edu^3$ in treatment assignment. Based on Tables \ref{Table: JTPA comparisons linear} and \ref{Table: JTPA comparisons cubic}, Tables \ref{Table: JTPA linear welfare loss} and \ref{Table: JTPA cubic welfare loss} in \cref{Section: JTPA additional} report the percentage welfare loss for every combination of actual $\alpha$ and $\alpha'$ for policy selection. A notable observation is that the bottom quarter of the population is particularly vulnerable when the policy targets some $\alpha' \geq 0.4$ instead. Thus, policymakers aspiring for greater equality should prioritize smaller levels of $\alpha$, such as $0.25$ or $0.3$, as evidenced by the small percentage welfare losses in the first two columns of Tables \ref{Table: JTPA linear welfare loss} and \ref{Table: JTPA cubic welfare loss}, all of which are below $5.5\%$.
\begin{table}[H]
\centering
\fontsize{9}{11}\selectfont
\begin{tabular}{|l||*{5}{c|}}\hline
\backslashbox{$\alpha$ of Interest}{$\alpha'$ for Policy Selection}
&\makebox[2em]{$0.25$}&\makebox[2em]{$0.3$}&\makebox[2em]{$0.4$}&\makebox[2em]{$0.5$}&\makebox[2em]{$0.8$}\\\hline\hline
\multirow{2}{*}{$0.25$} & \cellcolor{yellow}530.630 & 525.116 & 500.874 & 495.241 & 467.467 \\
& (46.581) & (48.020) & (42.561) & (45.623) & (41.415)\\\hline
\multirow{2}{*}{$0.3$} & 898.609 & \cellcolor{yellow}917.718 & 908.589 & 896.640 & 862.059\\
& (65.824) & (63.277) & (62.561) & (62.399) & (57.638)\\\hline
\multirow{2}{*}{$0.4$} & 1944.643 & 2020.792 & \cellcolor{yellow}2038.321 & 2035.307 & 1992.898\\
& (108.718) & (105.008) & (98.180) & (99.732) & (96.564)\\\hline
\multirow{2}{*}{$0.5$} & 3331.114 & 3485.067 & 3522.112 & \cellcolor{yellow}3524.651 & 3493.405\\
& (147.675) & (141.436) & (131.105) & (127.658) & (131.633)\\\hline
\multirow{2}{*}{$0.8$} & 9146.288 & 9451.340 & 9552.165 & 9588.269 & \cellcolor{yellow}9661.526\\
& (215.680) & (209.083) & (185.883) & (185.388) & (188.039)\\\hline
\end{tabular}\\
\caption{Estimated $\mathbb{W}_\alpha(\pi _o)$ for different actual $\alpha$'s of interest and $\alpha'$'s for linear policy selection (policy class $\Pi_{\text{LES}}$).  Standard errors are reported in parentheses. }
\label{Table: JTPA comparisons linear}
\end{table}
\vskip 0.2cm
\begin{table}[H]
\centering
\fontsize{9}{11}\selectfont
\begin{tabular}{|l||*{5}{c|}}\hline
\backslashbox{$\alpha$ of Interest}{$\alpha'$ for Policy Selection}
&\makebox[2.25em]{$0.25$}&\makebox[2.25em]{$0.3$}&\makebox[2.25em]{$0.4$}&\makebox[2.25em]{$0.5$}&\makebox[2.25em]{$0.8$}\\\hline\hline
\multirow{2}{*}{$0.25$} & \cellcolor{yellow}546.300 & 543.405 & 504.095 & 496.645 & 476.020\\
& (50.938) & (51.359) & (44.861) & (45.001) & (42.162)\\\hline
\multirow{2}{*}{$0.3$} & 917.043 & \cellcolor{yellow}918.011 & 910.930 & 897.083 & 871.931\\
& (68.319) & (72.094) & (62.953) & (61.785) & (61.71)\\\hline
\multirow{2}{*}{$0.4$} & 1972.299 & 1974.302 & \cellcolor{yellow}2039.468 & 2036.425 & 2004.521\\
& (109.090) & (109.322) & (101.637) & (100.647) & (102.276)\\\hline
\multirow{2}{*}{$0.5$} & 3364.695 & 3369.302 & 3525.834 & \cellcolor{yellow}3527.108 & 3509.814\\
& (144.252) & (146.286) & (130.284) & (131.639) & (134.797)\\\hline
\multirow{2}{*}{$0.8$} & 9197.693 & 9191.845 & 9555.919 & 9598.417 & \cellcolor{yellow}9690.607\\
& (216.352) & (214.680) & (186.193) & (185.961) & (194.407)\\\hline
\end{tabular}\\
\caption{Estimated $\mathbb{W}_\alpha(\pi _o)$ for different actual $\alpha$'s of interest and $\alpha'$'s for linear policy selection with $edu^2$ and $edu^3$ (policy class $\Pi^3_{\text{LES}}$).  Standard errors are reported in parentheses.}
\label{Table: JTPA comparisons cubic}
\end{table}

\subsection{Simulations Based on WGAN-Generated JTPA Data}
\label{Section: WGAN simulations}
We next present simulation results based on a superpopulation generated using Wasserstein Generative Adversarial Networks (WGANs) to evaluate the finite-sample performance of our debiased estimator. We focus on this simulation setup in the main text because the generated data more closely resembles real-world data distributions, making it more illustrative of practical applications. For comparison, we also conduct two additional simulation studies inspired by the DGPs in \cite{athey2021policy}, with adjustments that make the treatment assignment exogenous. Since the results across all three designs are qualitatively similar—our estimator consistently exhibits decreasing mean squared error as the sample size increases, and the coverage rates approach the nominal 95\% level in larger samples—we relegate the latter two studies to \cref{Section: AW simulations}.

In all three simulation setups, the propensity scores are assumed to be known, i.e., $\widehat{e}(\cdot) = e(\cdot)$. Cases with unknown propensity scores can be analyzed analogously using an estimator $\widehat{e}(\cdot)$ that satisfies Assumption \ref{assumption: nuisance parameter convergence rate}. Since uniform inference based on the multiplier bootstrap is computationally intensive, we report only the coverage rates based on confidence intervals constructed via Wald inference. We examine values of $\alpha \in \mathcal{A}$ considered in \cref{Section: JTPA}.

We employ WGANs developed by \cite{athey2024using} to construct a hypothetical superpopulation, referred to as WGAN-JTPA, consisting of one million observations based on the JTPA data in Section \ref{Section: JTPA}. As mentioned by \cite{athey2024using}, a benefit of using WGAN-generated data for simulations is that this practice largely rules out the possibility for researchers to choose particular DGPs that favor their proposed methods. This subsection demonstrates robust performance of our debiased estimator even when the underlying superpopulation is built from real datasets like the JTPA, which has highly skewed outcome and covariate distributions. \cref{Section: superpop} discusses the training process in more detail and presents some summary statistics.

While technical details of WGANs can be found in \cite{athey2024using}, we highlight that to build the superpopulation, since we generate $X|A$ followed by $Y|(X,A)$ and apply the same generator on $(X, 1-A)$ to obtain $Y|(X, 1-A)$, both potential outcomes are available for each individual. As a result, we can directly compute the true expected welfare at any $\alpha$ induced by any policy, which is simply a tail average of post-treatment outcomes. For each $\alpha\in\mathcal{A}$, we run SA to find a linear policy $\pi_o\in\Pi_{\mathrm{LES}}$ (as defined in \eqref{JTPA LES policy class}) that maximizes $\mathbb{W}_\alpha(\pi)$ and treat the resulting optimum $\mathbb{W}_\alpha(\pi_o)$ as the population truth.

As an illustration, we use WGAN-JTPA to compare the $0.25$-EWM policy with the $1$-EWM (mean-optimal) and equality-minded (standard Gini social welfare-optimal) policies. Inspired by Figure~3 in \cite{kitagawa2021equality}, Figure \ref{Figure: WGAN-JTPA quantile comparisons all} plots the between-quantile differences in post-treatment outcomes across these policies. The figure shows that both the $0.25$-EWM and equality-minded policies raise the welfare of lower-ranked individuals while lowering the welfare of higher-ranked individuals relative to the $1$-EWM policy at the population level, with the $0.25$-EWM policy placing much greater emphasis on these adjustments.

\begin{figure}[t]
\centering
   \includegraphics[width=0.9\textwidth]{plots/WGAN/q_diff_all.jpeg}
   \caption{Between-quantile differences in outcomes for the 0.25-EWM, $1$-EWM, and equality-minded policies using the WGAN-JTPA data.}
   \label{Figure: WGAN-JTPA quantile comparisons all}
\end{figure}

In the simulations, for each replicate, we draw a sample of size $n \in \{2000, 5000, 10000\}$ without replacement from WGAN-JTPA. The propensity score is fixed at the population mean of $A$, which is approximately $0.66475$.\footnote{This is very close to the mean of $A$ in the actual JTPA data, $0.66497$. In the JTPA Study, treatment was randomized with probability $2/3$, and we assume randomized treatment in WGAN-JTPA as well.} For each pair $(n, \alpha)$, we apply Algorithm \ref{alg:debiased} to 1,000 sample draws and organize the results in Table \ref{Table: WGAN-JTPA}. As shown by the marginal histogram for \textit{earnings} in Figure \ref{Figure: histograms} in \cref{Section: superpop}, WGAN-JTPA inherits the high skewness present in the original JTPA data. Consequently, larger sample sizes are required to achieve satisfactory coverage. From Table \ref{Table: WGAN-JTPA}, our optimal welfare estimator achieves acceptable coverage when $n = 5{,}000$, which is a realistic sample size for both experimental and observational studies (for reference, the original JTPA sample used by \cite{kitagawa2018a} contains 9,223 observations).

\begin{table}[t] \centering
  \footnotesize{
\begin{tabular}
{@{\extracolsep{25pt}}lccc} \multicolumn{1}{l}{\textbf{Sample size}} & \textbf{2,000} & \textbf{5,000} & \textbf{10,000}\\
\\[-1.5ex]\hline
\multicolumn{4}{c}
{\cellcolor{blue!20}$\textbf{Panel 1: }\boldsymbol{\alpha}\boldsymbol{=0.25}, \textbf{truth}\boldsymbol{=1119.195}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{Avg. \% treated using $\widehat{\pi}_{n}$} & 52.655\% & 54.893\% & 55.310\%\\
\multicolumn{1}{l}{Bias} & 191.791 & 82.014 & 43.564\\
\multicolumn{1}{l}{Variance} & 46486.298 & 18778.912 & 9049.531\\
\multicolumn{1}{l}{MSE} & 83269.985 & 25505.129 & 10947.367\\
\multicolumn{1}{l}{95\% Coverage} & 93.1\% & 93.7\% & 94.9\%\\
\\[-2ex]\hline
\multicolumn{4}{c}
{\cellcolor{blue!20}$\textbf{Panel 2: }\boldsymbol{\alpha}\boldsymbol{=0.3}, \textbf{truth}\boldsymbol{=1908.135}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{Avg. \% treated using $\widehat{\pi}_{n}$} & 53.739\% & 54.245\% & 56.121\%\\
\multicolumn{1}{l}{Bias} & 206.873 & 96.268 & 48.651\\
\multicolumn{1}{l}{Variance} & 55137.705 & 22734.450 & 11799.126\\
\multicolumn{1}{l}{MSE} & 97934.121 & 32001.932 & 14166.024\\
\multicolumn{1}{l}{95\% Coverage} & 92.0\% & 94.3\% & 94.8\%\\
\\[-2ex]\hline
\multicolumn{4}{c}{\cellcolor{blue!20}$\textbf{Panel 3: }\boldsymbol{\alpha}\boldsymbol{=0.4}, \textbf{truth}\boldsymbol{=3460.773}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{Avg. \% treated using $\widehat{\pi}_{n}$} & 55.863\% & 57.153\% & 58.069\%\\
\multicolumn{1}{l}{Bias} & 229.273 & 101.223 & 48.033\\
\multicolumn{1}{l}{Variance} & 59133.582 & 24046.124 & 13456.291\\
\multicolumn{1}{l}{MSE} & 111699.467 & 34292.263 & 15763.427\\
\multicolumn{1}{l}{95\% Coverage} & 91.8\% & 93.9\% & 94.3\%\\
\\[-2ex]\hline
\multicolumn{4}{c}{\cellcolor{blue!20}$\textbf{Panel 4: }\boldsymbol{\alpha}\boldsymbol{=0.5}, \textbf{truth}\boldsymbol{=4867.556}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{Avg. \% treated using $\widehat{\pi}_{n}$} & 58.165\% & 60.027\% & 61.596\%\\
\multicolumn{1}{l}{Bias} & 204.781 & 96.355 & 49.745\\
\multicolumn{1}{l}{Variance} & 58786.097 & 22457.617 & 12335.452\\
\multicolumn{1}{l}{MSE} & 100721.497 & 31741.832 & 14810.019\\
\multicolumn{1}{l}{95\% Coverage} & 92.3\% & 94.4\% & 95.3\%\\
\\[-2ex]\hline
\multicolumn{4}{c}
{\cellcolor{blue!20}$\textbf{Panel 5: }\boldsymbol{\alpha}\boldsymbol{=0.8}, \textbf{truth}\boldsymbol{=9475.336}$}
\vspace{.1cm}\\
\multicolumn{1}{l}{Avg. \% treated using $\widehat{\pi}_{n}$} & 75.955\% & 83.083\% & 88.094\%\\
\multicolumn{1}{l}{Bias} & 210.888 & 92.727 & 52.521\\
\multicolumn{1}{l}{Variance} & 80007.017 & 33274.611 & 16694.598\\
\multicolumn{1}{l}{MSE} & 124480.959 & 41872.844 & 19453.079\\
\multicolumn{1}{l}{95\% Coverage} & 93.5\% & 93.9\% & 95.3\%\\
\\[-2ex]\hline
\end{tabular}}
\caption{Simulation results based on WGAN-JTPA data (1,000 replications). }
\label{Table: WGAN-JTPA}
\end{table}

\section{Concluding Remarks}
\label{Section: conclusion}

The $\alpha$-expected welfare function considered in this paper offers a flexible interpolation between the Rawlsian welfare ($\alpha\rightarrow 0$) and the empirical welfare maximization ($\alpha=1$) approach proposed by \cite{kitagawa2018a}.
Like \cite{athey2021policy} for the empirical welfare maximization, our development of the doubly robust scores facilitates asymptotic inference for the optimal welfare and allows practitioners flexibility in how they estimate the nuisance parameters. Besides learning the optimal policies, our estimation strategy also enables more thorough policy evaluations by computing the average welfare of the worst-affected subpopulation of any size (fraction of the population). In addition to establishing regret bounds for the debiased estimator, we also develop inference for the optimal $\alpha$-expected welfare for any $\alpha \in (0,1)$. Results from extensive numerical studies based on both JTPA data and simulated data demonstrate the efficacy and practical value of policy learning through $\alpha$-EWM.

We are currently working on several extensions of this paper. Methodologically, it is important to develop statistical tests to compare whether one policy is superior to another.
Practically, it would be beneficial to determine who is actually targeted by the optimal policy. For example, what characteristics do the worst-affected individuals have? Information like this could present a more comprehensive picture of the relevant population and promote the design of more equitable policies.
\newpage