EconBase
← Back to paper

Data-Automated Policy Learning for Nonlinear Welfare

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

65,665 characters

\begin{frontmatter}
\hypersetup{linkcolor=black}
\begin{center}
{\LARGE\bfseries Data-Automated Policy Learning for Nonlinear Welfare\par}
\vspace{1.2em}
{\large Chunrong Ai\textsuperscript{a}\footnote{Email: [email removed]}.}\quad Zeqi Wu\textsuperscript{b}\footnote{Email: [email removed]}.}\quad Zheng Zhang\textsuperscript{b}\footnote{Email: [email removed]}.}\par}
\vspace{0.6em}
{\small \textsuperscript{a}\,School of Management and Economics, The Chinese University of Hong Kong, Shenzhen\par}
{\small \textsuperscript{b}\,Institute of Statistics and Big Data, Renmin University of China\par}
\end{center}
\setcounter{footnote}{0}
\hypersetup{linkcolor=blue}
\vspace{1em}

\begin{abstract}
This paper explores policy learning from observational data, focusing on a nonlinear welfare criterion in a binary treatment setting. The nonlinear criterion is inspired by scenarios where policymakers prioritize specific population segments. We model this criterion using a utility function that encompasses potential outcomes and intermediate parameters, with the latter capturing higher moments of the outcome distributions. When formulated in the context of observational data, both the intermediate parameters and the welfare criterion depend on the propensity score, which we estimate using machine-learning techniques. To address bias in machine learning estimates, we introduce a novel reweighting-based debiasing approach that offers a promising alternative to traditional orthogonality-based methods. To tackle the complexities of infinite-dimensional policy spaces, we employ sieve approximations and $K$-fold cross-validation for model selection, thereby fully automating the policy-learning process. Despite these complexities, we demonstrate that both the welfare regret and the average welfare regret of our proposed policy learning method satisfy an oracle inequality, thereby providing theoretical guarantees on the performance of the estimated policy relative to the best possible policy. This finding extends the existing results from linear to nonlinear welfare criteria, from finite-dimensional to infinite-dimensional policy spaces, and from a known propensity score to a machine-learned one.
\end{abstract}

\begin{keyword}
\ifkeywordfirst\keywordfirstfalse\else\unskip; \fi Policy Learning\ignorespaces
\ifkeywordfirst\keywordfirstfalse\else\unskip; \fi Oracle Inequality\ignorespaces
\ifkeywordfirst\keywordfirstfalse\else\unskip; \fi Sieve Approximation\ignorespaces
\ifkeywordfirst\keywordfirstfalse\else\unskip; \fi Machine Learning\ignorespaces
\ifkeywordfirst\keywordfirstfalse\else\unskip; \fi Welfare Criterion\ignorespaces
\end{keyword}
\end{frontmatter}

\section{Introduction}

There is an increasing body of literature on policy learning. Most existing studies focus on a linear welfare criterion and a binary treatment setting. In this context, the binary variable $T$ indicates program participation status: $T=1$ if the individual participates and $T=0$ if not. The potential outcome associated with participation status $t$ is represented as $Y^{*}(t)\in\mathcal{Y}\subset\mathbb{R}$, and $\bs X\in\mathcal{X}\subset\mathbb{R}^{d}$ denotes individual characteristics. A policy $\pi:\mathcal{X}\to\{0,1\}$ decides whether to assign an individual with attributes $\bs{X}$ to the program. The potential outcome under the policy is given by:
\[
Y^{*}(\pi(\bs X))=\pi(\bs X)Y^{*}(1)+(1-\pi(\bs X))Y^{*}(0).
\]
Policymakers specify a welfare criterion $W(\pi)$ and a policy space $\Pi_{\infty}$, aiming to learn the optimal policy defined by
\[
\pi^{*}=\arg \max_{\pi \in \Pi_{\infty}} W(\pi).
\]
A common choice for the welfare is the expectation of the potential outcome:
\[
W(\pi)=\mathbb{E}[Y^{*}(\pi(\bs X))]=\mathbb{E}[\pi(\bs X)Y^{*}(1)+(1-\pi(\bs X))Y^{*}(0)],
\]
which is clearly linear with respect to the policy.

A risk-averse policymaker might prefer to utilize the average utility of potential outcomes, represented as
\[
W(\pi)=\mathbb{E}[U(Y^{*}(\pi(\bs X)))]=\mathbb{E}[\pi(\bs X)U(Y^{*}(1))+(1-\pi(\bs X))U(Y^{*}(0))].
\]
However, this extension is only superficial: redefining the potential outcome as $U(Y^*(t))$ demonstrates that the resulting welfare remains linear in policy. By contrast, many significant applications lead to welfare criteria that are genuinely nonlinear in policy. For instance, in income-inequality literature, policymakers aim to reduce disparities through targeted interventions, and the welfare criterion is often a measure of inequality, such as the Gini coefficient \citep{gastwirth1971General,gastwirth1972Estimation}:
\[
W(\pi)= -\frac{2M^{-1}\sum_{r=1}^{M}\alpha_{r}\beta^{*}(\pi;\alpha_{r})}{\mathbb{E} [Y^{*}(\pi(\bs X))]} + 1,\qquad \alpha_{r}=\frac{r}{M+1},
\]
where $M$ is a finite positive integer and $\beta^{*}(\pi;\alpha)$ denotes the $\alpha$-quantile of $Y^{*}(\pi(\bs X))$. Since quantiles are nonlinear in relation to policy choices, the welfare criterion is likewise nonlinear. Nonlinearity also appears in alternative inequality measures, such as the relative standing of specific population segments:
\[
W(\pi)=\frac{\mathbb{E}\left[Y^{*}(\pi(\bs X))\mid Y^{*}(\pi(\bs X))\leq\beta^{*}(\pi;0.5)\right]}{\mathbb{E}\left[Y^{*}(\pi(\bs X))\right]},
\]
or the relative status of a defined subpopulation:
\[
W(\pi)=\frac{\mathbb{E}\left[Y^{*}(\pi(\bs X))1(\text{gender}=\text{``female''})\right]}{\mathbb{E}\left[Y^{*}(\pi(\bs X))\right]},
\]
and in the analysis of upper-tail (90/50) and lower-tail (50/10) ratios \citep{autor2008trends}:
\[
W(\pi)=-\frac{\beta^{*}(\pi;0.9)}{\beta^{*}(\pi;0.5)}\text{ and }W(\pi) = -\frac{\beta^{*}(\pi;0.5)}{\beta^{*}(\pi;0.1)}.
\]

Nonlinear welfare criteria naturally arise in the fields of risk management and public health. In risk management, firms aim to control the risk of significant losses. Let the binary variable $T\in\{0,1\}$ represent an investment decision, where $T=1$ indicates a one-period investment in the asset and $T=0$ indicates no investment. Let $Y^{*}(1)$ denote the corresponding one-period payoff (or return), while we set $Y^{*}(0)=0$. Given characteristics $\bs X$ (e.g., volatility, momentum, or fundamentals), an investment strategy $\pi:\mathcal X\to\{0,1\}$ induces the realized payoff:
\[
Y^{*}(\pi(\bs X))=\pi(\bs X)Y^{*}(1)+(1-\pi(\bs X))Y^{*}(0)=\pi(\bs X)Y^{*}(1).
\]
We denote $L^{*}(\pi)=-Y^{*}(\pi(\bs X))$ and let $\beta^{*}(\pi;\alpha)$ represent the $\alpha$-quantile of $L^{*}(\pi)$. The welfare criterion can be expressed as a measure of risk, such as Value-at-Risk (VaR), defined as $W(\pi)=-\beta^{*}(\pi;\alpha)$, or Conditional Value-at-Risk (CVaR), given by $W(\pi)=-(1-\alpha)^{-1}\mathbb{E}\!\left[L^{*}(\pi)\cdot 1\!\left(L^{*}(\pi)\ge \beta^{*}(\pi;\alpha)\right)\right]$ \citep{rockafellar2000Optimization}, or the spectral risk measure (e.g., \citealp{acerbi2002Spectral}):
\[
W(\pi)=-\sum_{r=1}^{M}\phi(\alpha_{r})\beta^{*}(\pi;\alpha_{r}),\qquad \alpha_{r}=\frac{r}{M+1},
\]
with a weighting function $\phi(\alpha)$ (see \citealp{dowd2006Var}). All of these risk measures are nonlinear in relation to policy.

In public health, policymakers aim to reduce the incidence of severe post-discharge utilization through targeted programs. For example, in the United States, Medicare's Hospital Readmissions Reduction Program (HRRP) specifically targets 30-day unplanned readmissions by penalizing hospitals with higher-than-average readmission rates \citep{zuckerman2016readmissions,ryan2017valuebased,khera2018hrrp}. Research shows that readmission-related utilization is highly right-skewed: a small fraction of patients accounts for a disproportionate share of readmissions and hospital use \citep{fouayzi2022highfrequency,manning2005generalized}. These distributional features motivate policymakers to prioritize the upper tail of the outcome distribution rather than focusing solely on the average. In this context, let $T$ represent the intensity of post-discharge care for a patient, where $T=1$ indicates enhanced transitional care (e.g., intensive monitoring and follow-up) and $T=0$ signifies usual care. The variable $Y^{*}(t)\geq 0$ reflects the readmission burden under treatment $t$, such as the total number of inpatient days due to unplanned readmissions within 30 days post-discharge, with lower values being more desirable. Given patient covariates $\bs X$ (e.g., age, gender, BMI, and comorbidity indices), a policy $\pi:\mathcal X\to\{0,1\}$ assigns care intensity and determines the burden $Y^{*}(\pi(\bs X))$. The welfare criterion being considered is a tail-sensitive measure of the expected burden among the worst-off $(1-\alpha)$ fraction of patients:
\[
W(\pi)
=
-\frac{1}{1-\alpha}\mathbb{E}\!\left[
Y^{*}(\pi(\bs X))\cdot
1\!\left\{Y^{*}(\pi(\bs X))\ge \beta^{*}(\pi;\alpha)\right\}
\right],
\]
where $\beta^{*}(\pi;\alpha)$ denotes the $\alpha$-quantile of $Y^{*}(\pi(\bs X))$ and is nonlinear in relation to the policy.


All the aforementioned examples emphasize the need for a nonlinear welfare criterion that directly depends on the policy through potential outcomes and indirectly through intermediate parameters $\bs{\beta}^*(\pi)$ (such as the quantiles mentioned earlier). Specifically, we model this with:
\begin{equation}
    W(\pi)=\mathbb{E}\left[U(Y^{*}(\pi(\bs X)),\bs X,\bs{\beta}^{*}(\pi))\right],\label{eq:general welfare}
\end{equation}
where $U(\cdot,\cdot,\cdot)$ is a known utility function.

The policy space can be complex and infinite-dimensional, making the optimal policy $\pi^{*}$ challenging to compute. To overcome this issue, we leverage the sieve literature to approximate the policy space using a sequence of finite-dimensional sieve classes $\{\Pi_{\ell}: \ell=1,2,\ldots\}$. Within each class, we determine the best policy as
\[\pi^{*}_{\ell}=\arg\max_{\pi\in \Pi_{\ell}}W(\pi).
\]
Since the welfare criterion is unknown, we estimate it from a training sample, say $I$, by $\widehat{W}_{I}(\pi)$ and find the best policy using:
\[
\widehat{\pi}_{\ell,I}=\arg\max_{\pi\in \Pi_{\ell}}\widehat{W}_{I}(\pi).
\]
We then apply $K$-fold cross-validation to select the optimal policy approximation space $\Pi_{\widehat{\ell}}$ (see (\ref{eq:def of l_hat-general})) and subsequently estimate the optimal policy as $\widehat{\pi}$ (see (\ref{eq:CV criterion-general})). Despite these extensions, we establish the following oracle inequality for the average welfare regret:
\begin{equation}
\begin{aligned}
    & \mathbb{E}\left[W(\pi^{*})-W(\widehat{\pi})\right]\\  \leq \ &\inf_{\ell=1,2,\ldots}\left\{ \underbrace{W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)}_{\mathrm{approximation \ error}}+\underbrace{\max_{\pi\in\Pi_{\ell}}W(\pi)-\mathbb{E}\left[W(\widehat{\pi}_{\ell,I})\right]}_{\mathrm{estimation \ error}}+\frac{\log\ell}{\sqrt{N}}\right\} +\sqrt{\frac{C}{N}}.
\end{aligned}
\label{eq:oracle inequality}
\end{equation}


\textit{Related Literature.} Our work builds on the policy-learning literature initially developed by \citet{manski2004Statistical}. A significant portion of this literature examines treatment choices under linear welfare criteria, in which an average or utilitarian objective defines the policy value. Early contributions to this field include studies by \citet{hirano2009asymptotics,stoye2009minimax,stoye2012minimax,bhattacharya2012inferring,tetenov2012statistical,qian2011Performance,zhao2012Estimating}, and more recent works by \citet{kitagawa2018Who,athey2021Policy,mbakop2021Model,luedtke2016Statistical,crippa2025Regret,liu2025nonparametric,fang2025semiparametric,fang2025model}. With the exception of \citet{mbakop2021Model}, all these studies assume a finite-dimensional policy space (e.g., $\Pi_{\infty}=\Pi_{\ell}$ with $\ell$ fixed or allowed to grow with sample size), which allows them to avoid the complexities of policy-space approximation and data-driven model selection. They either assume a known propensity score or estimate it using machine-learning methods, then apply double debiasing to remove the machine-learning bias. In both cases, they ultimately establish an explicit upper bound on the welfare regret:
\begin{equation}
\mathbb{E}\left[W(\pi^{*})-W(\widehat{\pi}_{\ell,I})\right]\leq C\sqrt{\frac{\mathrm{VC}(\Pi_{\ell})}{N}} \label{eq:first VC},
\end{equation}
where $\mathrm{VC}(\Pi_{\ell})$ represents the Vapnik-Chervonenkis (VC) dimension of the policy class. In contrast, \citet{mbakop2021Model} addresses an infinite-dimensional policy space with a known propensity score, thus avoiding the need for debiasing. They present an upper bound on the average welfare regret that is similar to ours, though still under a linear welfare criterion. Our contribution extends this literature to encompass a nonlinear welfare criterion, a machine-learned propensity score, and an infinite-dimensional policy space.

A smaller body of literature explores policy learning for nonlinear welfare criteria within a finite-dimensional policy class, where the dimension may increase with sample size (that is, $\Pi_{\infty}=\Pi_{\ell}$, with $\ell$ allowed to grow with sample size). Notable examples include \citet{wang2018QuantileOptimal} on quantile-optimal treatment regimes, \citet{chen2025quantileoptimalpolicylearningunmeasured} addressing a quantile-based welfare criterion with an unmeasured confounder, \citet{fan2025policylearningalphaexpectedwelfare} focusing on a conditional value-at-risk criterion, \citet{kitagawa2021EqualityMinded} discussing an equality-minded social welfare criterion, and \citet{terschuur2025locally} examining nonlinear welfare criteria defined through U-statistics. These studies assume a finite-dimensional policy space, treat the propensity score as either known or derived via machine learning, and employ double-debiasing techniques to mitigate machine-learning bias. Despite the nonlinearity, they manage to derive a similar, explicit upper bound on the average welfare regret, akin to the upper bound mentioned in equation (\ref{eq:first VC}). We aim to extend this existing literature to encompass infinite-dimensional policy spaces.

To learn the optimal policy from observational data, it is essential to estimate both the intermediate parameters and the welfare criterion. Since both estimates rely on the unknown propensity score, we use machine-learning algorithms to estimate it. It is well recognized that machine-learned propensity scores introduce bias in both the intermediate parameter and the welfare-criterion estimates, which in turn affects the (average) welfare regret and slows its convergence rate. To correct for this machine-learning bias, existing literature applies double-debiasing procedures based on a Neyman orthogonality condition (see \citet{athey2021Policy,robins1994Estimation,chernozhukov2018Double} for linear welfare criteria and \citet{fan2025policylearningalphaexpectedwelfare,terschuur2025locally} for nonlinear criteria). Our procedure includes an additional step, estimating the intermediate parameters, so we need to debias both these parameters and the welfare criterion estimates simultaneously. We propose a reweighting method inspired by covariate-balancing approaches \citep{imai2014Covariate,chan2016Globally,ai2021Unified}, adapted here for bias correction. Covariate-balancing methods have shown strong performance in finite samples, and we expect our debiasing procedure to exhibit similar effectiveness. To our knowledge, this reweighting-based debiasing technique is novel in the literature and serves as a valuable alternative to double debiasing based on Neyman orthogonality.

The remainder of the paper is organized as follows. Section~\ref{sec:A-General-Learning} presents a data-automated optimal policy learning procedure that employs a generic welfare estimate alongside $K$-fold cross-validation to identify the best policy subclass, while establishing an oracle inequality for both the average welfare regret and the welfare regret. Section~\ref{sec:general framework} formally defines the model, expressing the unknown parameters and the welfare criterion in terms of the observed data, based on the assumptions of unconfoundedness and overlap. Section~\ref{sec:empirical welfare} introduces a machine-learning propensity-score estimator, along with a novel reweighting debiasing procedure for estimating the intermediate parameters and the welfare criterion. Section~\ref{sec:asymptotic properties} verifies that the proposed estimator for the welfare criterion meets the high-level conditions outlined in Section~\ref{sec:A-General-Learning}. Section~\ref{sec:empirical application} applies the proposed methodology to data from the National Job Training Partnership Act (JTPA) Study. Finally, Section~\ref{sec:conclusion} summarizes the findings. All omitted proofs are included in the Appendix.


\section{A Data-Automated Learning Procedure\label{sec:A-General-Learning}}

We will outline a data-driven policy-learning procedure that relies on a generic welfare estimator and cross-validation. Let $\left\{ \left(Y_{i},\bs X_{i},T_{i}\right)\right\} _{i=1}^{N}$ represent an independent and identically distributed (i.i.d.) sample. Throughout the paper, we use $I\subset\{1,2,\ldots,N\}$ to index a generic training subsample and $|I|$ to denote its sample size. Let $\widehat{W}_{I}(\pi)$ be a generic estimator of $W(\pi)$ calculated from the training subsample $\left\{ \left(Y_{i},\bs X_{i},T_{i}\right)\right\} _{i\in I}$. We assume that $\widehat{W}_{I}(\pi)$ is pointwise $\sqrt{\left|I\right|}$-consistent for $W(\pi)$ and that its estimation error satisfies an exponential probability bound, as formalized in Assumption~\ref{assu:general assu on hatW}.

\begin{assumption}
\label{assu:general assu on hatW}
For a fixed policy $\pi\in\Pi_{\infty}$, there exist finite constants $C_{1},\ldots,C_{4}>0$, independent of $\delta$, $I$, and $\pi$, such that the following inequality holds for all $\delta>0$ and $\left|I\right| >C_{4}$:
\[
P\left(\left|\widehat{W}_{I}(\pi)-W(\pi)\right|\geq\delta+\frac{C_{1}}{\sqrt{\left|I\right|}}\right)\leq C_{2}\exp(-C_{3}\left|I\right|\delta^{2}).
\]

\end{assumption}
In applications, users must verify that their welfare estimators satisfy this high-level condition. We will present a welfare estimator and confirm that it indeed satisfies Assumption~\ref{assu:general assu on hatW}.

With $\widehat{W}_{I}(\pi)$ established, a natural policy learning strategy is to maximize it over $\Pi_{\infty}$. However, for an infinite-dimensional $\Pi_{\infty}$, such a strategy can be computationally intractable and prone to overfitting. Following the work of \citet{mbakop2021Model}, we approximate $\Pi_{\infty}$ by a nested sequence of finite-dimensional policy subclasses $\Pi_{\ell}\subset\Pi_{\ell+1}\subset\cdots\subset\Pi_{\infty}$.\footnote{Throughout the paper, we take this approximating sequence to satisfy $\mathrm{VC}(\Pi_{\ell})<\infty$ for every $\ell\geq1$ and $W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)\to0$ as $\ell\to\infty$.} Within each subclass $\Pi_{\ell}$, we estimate the best policy by
\begin{equation}
\widehat{\pi}_{\ell,I}:=\underset{\pi\in\Pi_{\ell}}{\arg\max}\ \widehat{W}_{I}(\pi).\label{eq:def of each optimal pi-general}
\end{equation}
We determine the best subclass using a $K$-fold cross-validation (CV) procedure, as outlined in various studies (\citet{Hall1983Large}, \citet{Stone1974Cross-Validatory}, \citet{lecue2012Oracle}, and \citet{gyorfi2002DistributionFree}). Specifically, for a fixed integer $K\geq2$, we partition the index set $\{1,\ldots,N\}$ into $K$ disjoint folds of equal size. For each pair $(k,\ell)$, let $I_{k}$ denote the $k$th fold and $I_{-k}=\{1,\ldots,N\}\setminus I_{k}$ its complement. We utilize the training subsample $I_{-k}$ to learn the welfare  $\widehat{W}_{I_{-k}}(\pi)$ and the candidate policy $\widehat{\pi}_{\ell,I_{-k}}$. The holdout sample $I_{k}$ is then used to compute $\widehat{W}_{I_{k}}(\pi)$ and to evaluate the performance of the candidate policy on the holdout sample using $\widehat{W}_{I_{k}}(\widehat{\pi}_{\ell,I_{-k}})$. The $K$-fold CV procedure selects the best subclass according to the formula:
\begin{equation}
\widehat{\ell}=\underset{\ell=1,2,\ldots}{\arg\max}\ \left\{ \frac{1}{K}\sum_{k=1}^{K}\widehat{W}_{I_{k}}(\widehat{\pi}_{\ell,I_{-k}})-\frac{\log\ell}{\sqrt{N}}\right\}, \label{eq:def of l_hat-general}
\end{equation}
where the penalty term $\log\ell/\sqrt{N}$ helps to prevent the selection of very large policy classes. The identified policy is then defined as:
\begin{equation}
\widehat{\pi}:=\widehat{\pi}_{\widehat{\ell},I_{-\widehat{k}}},\ \text{where }\widehat{k}:=\underset{1\leq k\leq K}{\arg\max}\ \widehat{W}_{I_{k}}(\widehat{\pi}_{\widehat{\ell},I_{-k}}).\label{eq:CV criterion-general}
\end{equation}
The following theorem establishes the oracle inequality (\ref{eq:oracle inequality}) with $I=I_{-k}$ for any $k$.

\begin{thm}
\label{thm:oracle inequality-general}
Assuming that the conditions outlined in Assumption \ref{assu:general assu on hatW} hold and that $C>0$ is a finite constant, the learned optimal policy $\widehat{\pi}$ defined in (\ref{eq:CV criterion-general}) satisfies the oracle inequality (\ref{eq:oracle inequality}) for sufficiently large $N$.

\end{thm}

\begin{rem}
The study by \citet{mbakop2021Model} utilized a single holdout sample to identify the best policy class, which introduces randomness due to relying on that single sample. The $K$-fold cross-validation (CV) procedure mitigates this randomness by averaging over multiple holdout samples.
\end{rem}
\begin{rem}
\label{rem:finite L}
The penalty term ``$\log\ell/\sqrt{N}$'' in (\ref{eq:def of l_hat-general}) serves as a regularizer that prevents the selection of very large policy classes (as noted in \cite{mbakop2021Model}). This term can be omitted when the number of candidate subclasses is finite and may even increase with $N$. For example, one might use the following equations:
\[
\widehat{\ell}:=\underset{\ell=1,\ldots,L_{N}}{\arg\max}\sum_{k=1}^{K}\widehat{W}_{I_{k}}(\widehat{\pi}_{\ell,I_{-k}}) \text{ and } \widehat{k}:=\underset{1\leq k\leq K}{\arg\max} \ \widehat{W}_{I_{k}}(\widehat{\pi}_{\widehat{\ell},I_{-k}})
\]
where $L_{N}\geq2$ is a sequence of positive integers. In this scenario, the learned optimal policy $\widehat{\pi}$ fulfills the condition:
\begin{align*}
    &\mathbb{E}\left[W(\pi^{*})-W(\widehat{\pi})\right]\\
      \leq &  \ \inf_{\ell=1,\ldots,L_{N}}\Biggl\{
    \underbrace{W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)}_{\mathrm{approximation\ error}}
    +\underbrace{\max_{\pi\in\Pi_{\ell}}W(\pi)-\mathbb{E}\left[W(\widehat{\pi}_{\ell,I_{-1}})\right]}_{\mathrm{estimation\ error}}
    \Biggr\} +C\sqrt{\frac{\log L_{N}}{N}}
\end{align*}
for sufficiently large $N$; the proof is given in Appendix~\ref{sec:proofs-general-learning}. The bound compares this policy with the best tradeoff among the first $L_{N}$ policy classes, but it may not necessarily be the best overall.  \end{rem}
\begin{rem}
    Once the best subclass $\Pi_{\widehat{\ell}}$ has been selected, it is tempting to re-learn the policy using the full sample by maximizing $\widehat{W}_{\{1,\ldots,N\}}(\pi)$ over $\pi\in\Pi_{\widehat{\ell}}$. However, without additional stability-type conditions, as discussed in \citep{bousquet2002stability}, this learned policy may not satisfy the oracle inequality (\ref{eq:oracle inequality}), because the policy selected from a subsample might not be the best after retraining on the full sample, as noted by \citep[Example~2.8]{lecue2012Oracle}.
    \end{rem}

The theorem generalizes the oracle inequality by providing a bound on the average welfare regret $\mathbb{E}[W(\pi^{*})-W(\widehat{\pi})]$. In practice, policymakers can only access a single sample, and are more concerned with the realized welfare regret $W(\pi^{*})-W(\widehat{\pi})$. We show that a similar oracle inequality holds with high probability.

\begin{cor}
\label{cor:oracle inequality prob bound}
Suppose Assumption \ref{assu:general assu on hatW} holds. Let $C,C^{\prime},C_{1},C_{2}>0$ be finite constants. The learned policy $\widehat{\pi}$, defined in (\ref{eq:CV criterion-general}) satisfies the inequality:
\begin{align*}
    &W(\pi^{*})-W(\widehat{\pi})\\
   \leq & \inf_{\ell=1,2,\ldots}\left\{ \underbrace{W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)}_{\mathrm{approximation\ error}}+\underbrace{\max_{\pi\in\Pi_{\ell}}W(\pi)-\frac{1}{K}\sum_{k=1}^{K}W(\widehat{\pi}_{\ell,I_{-k}})}_{\mathrm{estimation\ error}}+2\frac{\log\ell}{\sqrt{N}}\right\} +\sqrt{\frac{C+\delta}{N}}.
   \end{align*}
This holds with a probability of at least $1-C_{1}\exp(-C_{2}\delta)$ for all $\delta>0$ and sufficiently large $N$.
\end{cor}

Corollary~\ref{cor:oracle inequality prob bound} establishes a probability oracle inequality. The constant ``2'' in front of $\log\ell/\sqrt{N}$ arises from a union bound used to control the concentration of $\frac{1}{K}\sum_{k=1}^{K}\widehat{W}_{I_{k}}(\widehat{\pi}_{\ell,I_{-k}})$ around $\frac{1}{K}\sum_{k=1}^{K}W(\widehat{\pi}_{\ell,I_{-k}})$. This constant can be replaced with any value greater than $1$, with corresponding adjustments to $C,C_1,C_2$.

The oracle inequality provides insights into the quality of the learned policy only when both the approximation and estimation errors are minimal. We can quantify the estimation error under a strengthened Assumption~\ref{assu:general assu on hatW}.

\begin{assumption}
\label{assu:general ass uniform convergence}
For the training sample $I$ and a subclass of policies $\Pi\subset\Pi_{\infty}$ with VC dimension satisfying $\mathrm{VC}(\Pi)/|I|\to0$, we can assume that:
\[
P\left(\sup_{\pi\in\Pi}\left|\widehat{W}_{I}(\pi)-W(\pi)\right|\geq\delta+C_{1}\sqrt{\frac{\mathrm{VC}(\Pi)}{\left|I\right|}}\right)\leq C_{2}\exp(-C_{3}\left|I\right|\delta^{2})
\]
holds for all $\delta>0$ and $\left|I\right|>C_{4}$, where $C_{1},\ldots,C_{4}>0$ are finite constants independent of $\delta$, $I$, and $\Pi$.
\end{assumption}

Assumption~\ref{assu:general ass uniform convergence} imposes a uniform convergence rate on the welfare estimator. It allows the VC dimension of the policy class, $\Pi$, to increase with the sample size at a rate that is slower than the sample size itself: $\mathrm{VC}(\Pi)/|I|\to0$. This rate condition aligns, up to logarithmic factors, with the minimax rate for policy learning under linear welfare criteria \citep{athey2021Policy,kitagawa2018Who}. Under this strengthened assumption, we derive an upper bound on the estimation error.

\begin{cor}
\label{cor:estimation-error-bound-general}
Suppose Assumption~\ref{assu:general ass uniform convergence}
holds. Then there exist finite constants $C,C^{\prime}>0$ such that the learned policy $\widehat{\pi}$ defined in (\ref{eq:CV criterion-general}) satisfies, for sufficiently large $N$,
\begin{align*}
 & \mathbb{E}\left[W(\pi^{*})-W(\widehat{\pi})\right]\\
\leq & \inf_{\ell=1,2,\ldots}\left\{ \underbrace{W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)}_{\mathrm{approximation\ error}}+C^{\prime}\sqrt{\frac{K}{K-1}}\sqrt{\frac{\mathrm{VC}(\Pi_{\ell})}{N}}+\frac{\log\ell}{\sqrt{N}}\right\} +\sqrt{\frac{C}{N}}.
\end{align*}
\end{cor}

To quantify the approximation error, however, we need to gather information about the policy space and its approximation. For illustrative purposes, we will compute the approximation error for three common policy classes and their approximations. Throughout this process, we will maintain the following Lipschitz condition on $W(\pi)$.

\begin{assumption}
\label{ass:high-level}
There exists a finite constant $C_{W}>0$ such
that for any two policies $\pi_{1},\pi_{2}\in\Pi_{\infty}$, \(
\left|W(\pi_{1})-W(\pi_{2})\right|\le C_{W}P(\pi_{1}(\bs X)\neq\pi_{2}(\bs X))\).

\end{assumption}
\begin{example}[Monotone policies]
\label{exa:monotone policies}
In practice, shape restrictions on the policy class $\Pi_{\infty}$ capture constraints implied by economic theory or fairness considerations. A canonical example of this is monotonicity. After normalizing the supports, let $\bs X_i=(X_{i1},X_{i2})^{\top}\in[0,1]^2$. The monotone policy class is written as:
\(
\Pi_{\infty}=\left\{ \pi_f(x_1,x_2)=1\{x_2\le f(x_1)\}:f\text{ is non-increasing}\right\}
\).
We approximate this infinite-dimensional class using the monotone piecewise-linear sieve $\Pi_{\ell}$. Its elements are threshold rules defined by non-increasing, continuous, piecewise-linear boundaries with knots on a $2^{\ell}$-grid. Appendix~\ref{sec:approximation-errors} provides the formal definition, following the construction found in \citet[Example~3.2]{mbakop2021Model}. Under Assumption~\ref{ass:high-level}, if the conditional density of $X_1$ given $X_2$ is uniformly bounded, then $\mathrm{VC}(\Pi_{\ell})=O(2^{\ell})$ and the approximation error is $W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)=O(2^{-\ell})$ (see Proposition~\ref{prop:monotone-approx-vc}).
\end{example}
In the next two examples, the sieve classes need not be subsets of $\Pi_{\infty}$. When these classes are used in the oracle bound, Assumptions~\ref{assu:general assu on hatW}--\ref{ass:high-level} are imposed on these sieve classes directly, so Theorem~\ref{thm:oracle inequality-general} continues to apply through the resulting approximation and estimation errors.
\begin{example}[Decision Trees]
\label{exa:Decision Trees}
Decision trees are another popular class of binary-valued policies \citep[e.g.,][]{athey2021Policy,zhou2023Offline}. One example policy space is a treatment rule represented as a smooth threshold in a single covariate, conditional on the remaining covariates. Specifically, with $\bs X=(\bs X_{-d},X_d)\in[0,1]^d$, the policy class is given by
\(
\Pi_{\infty}=\left\{ \pi_f(\bs x_{-d},x_d)=1\{x_d\le f(\bs x_{-d})\}:f\in C^{s}([0,1]^{d-1}),\ 0\le f\le1\right\}
\).
Here $s\in\mathbb{N}_{+}$, and $C^{s}([0,1]^{d-1})$ denotes the class of functions with continuous partial derivatives up to order $s$. Let $\Pi_{\ell}$ be the class of binary decision trees of depth at most $\ell$. Each internal node selects a coordinate $j\in\{1,\ldots,d\}$ and a threshold $b\in\mathbb{R}$, routing observations according to whether $x_j<b$, with each leaf assigned a label in $\{0,1\}$. Under Assumption~\ref{ass:high-level}, if the conditional density of $X_d$ given $\bs X_{-d}$ is uniformly bounded, then $\mathrm{VC}(\Pi_{\ell})=O(2^{\ell}(\ell+\log d))$ and the approximation error is $W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)=O\!\left(2^{-\left\lfloor(\ell-1)/(d-1)\right\rfloor}\right)$ (see Proposition~\ref{prop:dt-approx-vc}).
\end{example}
\begin{example}[Deep Neural Networks]
\label{exa:Neural Networks}
Deep neural networks (DNNs) provide a flexible framework for approximating complex decision boundaries. As in Example~\ref{exa:Decision Trees}, we also consider the smooth decision-boundary policy class, with $\bs X=(\bs X_{-d},X_d)\in[0,1]^d$, given by
\(
\Pi_{\infty}=\left\{ \pi_f(\bs x_{-d},x_d)=1\{x_d\le f(\bs x_{-d})\}:f\in C^{s}([0,1]^{d-1}),\ 0\le f\le1\right\}
\).
Let $\mathcal{F}_{\mathrm{DNN},\ell}$ denote the class of fully connected feedforward ReLU networks with width $\mathcal{H}_{\ell}$ and depth $\mathcal{D}_{\ell}$, where these architectural parameters grow with $\ell$. The sieve policy class is $\Pi_{\ell}=\{\pi_g(\bs x)=1\{g(\bs x)\ge0\}:g\in\mathcal{F}_{\mathrm{DNN},\ell}\}$. Under Assumption~\ref{ass:high-level}, if the conditional density of $X_d$ given $\bs X_{-d}$ is uniformly bounded, then Proposition~\ref{prop:nn-approx-vc} gives \(\mathrm{VC}(\Pi_{\ell})=O\{\mathcal{D}^{2}_{\ell}\mathcal{H}^{2}_{\ell}\allowbreak\log(\mathcal{D}_{\ell}\mathcal{H}^{2}_{\ell})\}\). The approximation error is \(W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)\allowbreak=O\!\{(\mathcal{H}_{\ell}/\log\mathcal{H}_{\ell})^{-2s/(d-1)}\allowbreak(\mathcal{D}_{\ell}/\log\mathcal{D}_{\ell})^{-2s/(d-1)}\}\).
\end{example}

\section{Model}
\label{sec:general framework}

We now formally establish the model for the welfare criterion. We define the intermediate parameters $\bs{\beta}^{*}(\pi)\in\mathbb{R}^{p}$ as the minimizer of a sum of expected convex losses:
\[
\bs{\beta}^{*}(\pi):=\underset{\bs{\beta}=(\beta_{1},\ldots,\beta_{p})^{\top}\in\mathbb{R}^{p}}{\arg\min}\sum_{j=1}^{p}\mathbb{E}\left[\mathcal{L}_{j}(Y^{*}(\pi(\bs X))-\beta_{j})\right],
\]
where $p\geq1$ is an integer and $\mathcal{L}_{1},\ldots,\mathcal{L}_{p}$ are known convex loss functions. Different selections of loss functions capture different distributional features. For example, when $p=2$, if we take $\mathcal{L}_{1}(v)=v^{2}/2$, and $\mathcal{L}_{2}(v)=v\cdot(0.5-1(v\leq0))$, the intermediate parameters correspond to the mean and the median: $\bs{\beta}^{*}(\pi)=\left(\mathbb{E}\left[Y^{*}(\pi(\bs X))\right],\text{median}[Y^{*}(\pi(\bs X))]\right)^{\top}$. Similarly, with the check loss defined as $\mathcal{L}(v)=v\cdot(\alpha - 1(v\leq 0))$, the intermediate parameter corresponds to the $\alpha$-quantile.

Both the intermediate parameters and the welfare criterion are expressed in terms of potential outcomes. To reframe them using the observed data $(Y,\bs X,T)$, where $Y=Y^{*}(T)$ represents the observed outcome, we impose the Stable Unit Treatment Value Assumption (SUTVA) \citep{imbens2015Causal} and  the following conditions regarding the data-generating process \citep{athey2021Policy,kitagawa2018Who,mbakop2021Model}.

\begin{assumption}[Unconfoundedness and Overlap]\label{assu:unconfounded}
\begin{enumerate}[label=(\roman*)]
    \item \textbf{(Unconfoundedness)} Given the covariate $\bs X$, program participation $T$ is independent of the potential outcomes, meaning $Y^{*}(0),Y^{*}(1)\perp T\mid\bs X$.
    \item \textbf{(Overlap)} There exists a constant $0<\kappa<1/2$ such that $\kappa < e^{*}(\bs X) < 1-\kappa$ almost surely, where $e^{*}(\bs X):=\mathbb{E}\left[T\mid\bs X\right]$ is the propensity score.
\end{enumerate}
\end{assumption}

Under Assumption \ref{assu:unconfounded}, we can express the following optimization problem:
\begin{equation}
\bs{\beta}^{*}(\pi)=\underset{\bs{\beta}=(\beta_{1},\ldots,\beta_{p})^{\top}\in\mathbb{R}^{p}}{\arg\min}\sum_{j=1}^{p}\mathbb{E}\left[\left\{ \frac{\pi(\bs X)T}{e^{*}(\bs X)}+\frac{\left(1-\pi(\bs X)\right)\left(1-T\right)}{1-e^{*}(\bs X)}\right\} \mathcal{L}_{j}(Y-\beta_{j})\right].\label{eq:def-of-beta-pi}
\end{equation}
Additionally, we define:
\begin{align}
W(\pi)=\mathbb{E}\left[\left\{ \frac{\pi(\bs X)T}{e^{*}(\bs X)}+\frac{\left(1-\pi(\bs X)\right)\left(1-T\right)}{1-e^{*}(\bs X)}\right\} U(Y,\bs X,\bs{\beta}^{*}(\pi))\right].\label{eq:identification-of-W-pi}
\end{align}
The propensity score $e^{*}(\bs X)$ is unknown and is defined by $e^{*}(\bs X)=\mathbb{E} [T\mid\bs X]$, which can be found by solving the following optimization problem:
\(
e^{*}(\bs X)=\underset{e(\cdot)}{\arg\min}\left\{ \mathbb{E}\left[(T-e(\bs{X}))^2\right]\right\}
\).
However, a sample least-squares estimator based on this characterization may produce fitted values close to 0 or 1. We therefore use the following equivalent population formulation, understood over functions satisfying $0<e(\bs X)<1$:
\begin{align}
e^{*}(\bs X) & =\underset{0<e(\cdot)<1}{\arg\min}\Biggl\{ \mathbb{E}\left[T\left(\frac{1}{e(\bs X)}-\frac{1}{e^{*}(\bs X)}\right)^{2}+(1-T)\left(\frac{1}{1-e(\bs X)}-\frac{1}{1-e^{*}(\bs X)}\right)^{2}\right]\Biggr\} \nonumber \\
  & =\underset{0<e(\cdot)<1}{\arg\min}\left\{ \mathbb{E}\left[\frac{T}{e(\bs X)^{2}}-\frac{2}{e(\bs X)}\right]+\mathbb{E}\left[\frac{1-T}{(1-e(\bs X))^{2}}-\frac{2}{1-e(\bs X)}\right]\right\}. \label{eq:propensity-objective}
\end{align}
Compared with the least-squares characterization, this criterion measures errors on the inverse-propensity scale that enters the inverse-probability-weighted (IPW) terms. Its population excess risk is a weighted sum of squared errors of $1/e(\bs X)$ and $1/\{1-e(\bs X)\}$, with the same minimizer $e^{*}$; in estimation, we minimize its sample analog over a bounded logistic DNN class to keep fitted propensity scores away from $0$ and $1$.

\section{Estimation of Welfare}\label{sec:empirical welfare}

We present an estimator for the welfare criterion based on a generic training sample $I$. Equations (\ref{eq:def-of-beta-pi})--(\ref{eq:propensity-objective}) suggest a three-step sequential estimation procedure. In the first step, we estimate the propensity score using a sample analog of (\ref{eq:propensity-objective}). We then substitute this estimate into a sample analog of (\ref{eq:def-of-beta-pi}) to estimate the intermediate parameters. Finally, we use both estimates to compute the welfare criterion using a sample analog of (\ref{eq:identification-of-W-pi}).

To estimate the propensity score from the training sample $I$, we utilize a deep neural network, denoting the estimate as $\widehat{e}_{I}(\cdot)$ (see (\ref{eq:logistic reg for propensity})). From the existing literature, it follows that $\left\Vert \widehat{e}_{I}-e^{*}\right\Vert_{P,2}= O_{P}\left(\left|I\right|^{-s_{e}/(2s_{e}+d)}\log^{3}\left|I\right|\right)$ (see the proof of Lemma~\ref{lem:converge of propensity}), where $s_{e}$ indicates the smoothness of $e^{*}(\bs X)$ (see Assumption~\ref{assu:Overlap assumption and smoothness}).

It is well-documented that machine learning can induce bias, which then propagates to the welfare criterion through both direct bias in the IPW terms and indirect bias in estimates of intermediate parameters. Simply substituting $e^{*}$ with $\widehat{e}_{I}$ in the sample analog of equations (\ref{eq:def-of-beta-pi})--(\ref{eq:propensity-objective}) may introduce bias in the intermediate parameters and welfare estimates. This, in turn, leads to a violation of Assumption~\ref{assu:general assu on hatW}. To mitigate machine learning bias, we propose a weighted analog of equations (\ref{eq:def-of-beta-pi})--(\ref{eq:propensity-objective}):
\begin{equation}
\widehat{\bs{\beta}}_{I}(\pi)=\underset{\bs{\beta}=(\beta_{1},\ldots,\beta_{p})^{\top}\in\mathbb{R}^{p}}{\arg\min}\sum_{j=1}^{p}\frac{1}{\left|I\right|}\sum_{i\in I}\widehat{w}_{I,i}(\pi)\Bigl\{\frac{\pi(\bs X_{i})T_{i}}{\widehat{e}_{I}(\bs X_{i})}+\frac{(1-\pi(\bs X_{i}))(1-T_{i})}{1-\widehat{e}_{I}(\bs X_{i})}\Bigr\}\mathcal{L}_{j}(Y_{i}-\beta_{j}).\label{eq:def-of-beta-hat-pi-I-unknown-propensity}
\end{equation}
The welfare estimator is defined as:
\begin{equation}
\widehat{W}_{I}(\pi)=\frac{1}{\left|I\right|}\sum_{i\in I}\widehat{w}_{I,i}(\pi)\Bigl\{\frac{\pi(\bs X_{i})T_{i}}{\widehat{e}_{I}(\bs X_{i})}+\frac{(1-\pi(\bs X_{i}))(1-T_{i})}{1-\widehat{e}_{I}(\bs X_{i})}\Bigr\} U(Y_{i},\bs X_{i},\widehat{\bs{\beta}}_{I}(\pi)).\label{eq:def-of-W-hat-pi-I-unknown-propensity}
\end{equation}
In this context, $\left\{ \widehat{w}_{I,i}(\pi):i\in I\right\}$ represents the calibrated weights. We show in Appendix~\ref{sec:app-nuisance-estimation} that these weights must satisfy the following conditions:
\begin{equation}
\begin{aligned}
 & \frac{1}{\left|I\right|}\sum_{i\in I}(1-\pi(\bs X_{i}))\left\{ \frac{w_{i}(1-T_{i})}{1-\widehat{e}_{I}(\bs X_{i})}-1\right\} \mu_{j0}^{*}(\bs X_{i};\bs{\beta}^{*}(\pi))\\
 & +\frac{1}{\left|I\right|}\sum_{i\in I}\pi(\bs X_{i})\left\{ \frac{w_{i}T_{i}}{\widehat{e}_{I}(\bs X_{i})}-1\right\} \mu_{j1}^{*}(\bs X_{i};\bs{\beta}^{*}(\pi))=0,\quad j=0,1,\ldots,p.
\end{aligned}
\label{eq:calibrated-weight-constraints}
\end{equation}
For $t\in\{0,1\}$, we define:
$\mu_{0t}^{*}(\bs x;\bs{\beta}):=\allowbreak\mathbb{E}\bigl[U(Y,\bs X,\bs{\beta})\allowbreak\mid\allowbreak\bs X=\bs x,T=t\bigr]$
and $\mu_{jt}^{*}(\bs x;\bs{\beta}):=\allowbreak\mathbb{E}\bigl[\mathcal{L}_{j}^{\prime}(Y-\beta_{j})\allowbreak\mid\allowbreak\bs X=\bs x,T=t\bigr]$
for $j=1,\ldots,p$, where $\mathcal{L}_{j}^{\prime}$ is understood as specified in Assumption~\ref{assu:regularity-conditions-on-L}.
In practice, $\bs{\beta}^{*}(\pi)$ and $\mu_{jt}^{*}(\cdot;\bs{\beta})$ are replaced with the initial estimator $\widehat{\bs{\beta}}_{I}^{\mathrm{init}}(\pi)$
and the conditional mean estimators $\widehat{\mu}_{I,jt}(\cdot;\widehat{\bs{\beta}}_{I}^{\mathrm{init}}(\pi))$, respectively,
as detailed in Appendix~\ref{sec:app-nuisance-estimation}.


Notice that the weights satisfying the equations (\ref{eq:calibrated-weight-constraints}) are generally not unique. We apply the entropy method to calibrate the weights as the solution to
\begin{equation}
\left(\widehat{w}_{I,i}(\pi):i\in I\right)=\underset{w_{i}>0:i\in I}{\arg\min}\sum_{i\in I}\left(w_{i}\log w_{i}-w_{i}\right)\text{ subject to }(w_{i}:i\in I)\text{ satisfying }(\ref{eq:calibrated-weight-constraints}).\label{eq:balancing for weights}
\end{equation}
In this context, the objective function $D(w)=w\log w-w$ measures the distance of $w$ from $1$, ensuring that the calibrated weights $\widehat{w}_{I,i}(\pi)$, $i\in I$, are unique and always non-negative.

\begin{rem}
    From a computational perspective, the equation (\ref{eq:balancing for weights}) is a convex program with $p+1$ linear constraints, making it straightforward to compute. In particular, the problem can be solved efficiently through its dual formulation using standard convex-optimization solvers, as referenced in \citep{boyd2004Convex}.
    \end{rem}

\section{Properties of the Empirical Welfare Criterion}\label{sec:asymptotic properties}

Having constructed the welfare criterion estimator, we will now verify that it meets the high-level condition outlined in Section~\ref{sec:A-General-Learning}. We require the following conditions.

\begin{assumption}
\label{assu:Overlap assumption and smoothness} Assume that $\mathcal{X}=[0,1]^{d}$. Let $M>1$ be a finite constant. The true propensity score $e^{*}(\bs x)$ satisfies the following condition:
\[
\log\frac{e^{*}(\bs x)}{1-e^{*}(\bs x)}\in C^{s_{e}}\left([0,1]^{d}\right):=\left\{ f:\max_{\bs{\alpha}\in\mathbb{N}^{d},\left\Vert \bs{\alpha}\right\Vert _{1}\leq s_{e}}\underset{\bs x\in[0,1]^{d}}{\sup}\left|\partial^{\bs{\alpha}}f(\bs x)\right|\leq M\right\} ,
\]
where $s_{e}>d/2$ is an integer, $\left\Vert \bs{\alpha}\right\Vert _{1}:=\alpha_{1}+\cdots+\alpha_{d}$ and $\partial^{\bs{\alpha}}f$ represents the partial derivative of $f$.
\end{assumption}
\begin{assumption}
[Regularity assumptions on $\mathcal{L}_{j}$]\label{assu:regularity-conditions-on-L} For
each $j=1,\ldots,p$, let $\beta_{j}^{*}(\pi)$ be the $j$th component of $\bs{\beta}^{*}(\pi)$. Let $c_{0}>0$ be a finite constant. The following conditions hold for any $j=1,\ldots,p$.
\begin{enumerate}[label=(\roman*)]
\item $\mathcal{L}_{j}(v)$ is convex on $\mathbb{R}$. There exists a non-decreasing function $\mathcal{L}_{j}^{\prime}(v):\mathbb{R}\to\mathbb{R}$ such that $\int_{a}^{b}\mathcal{L}_{j}^{\prime}(v)dv=\mathcal{L}_{j}(b)-\mathcal{L}_{j}(a)$ for any $a,b\in\mathbb{R}$. We refer to $\mathcal{L}_{j}^{\prime}(v)$ as the ``derivative'' of $\mathcal{L}_{j}(v)$.
\item Define
\(
Q_{j}(\beta;\pi):=\mathbb{E}\Bigl[\left\{ \frac{\pi(\bs X)T}{e^{*}(\bs X)}+\frac{\left(1-\pi(\bs X)\right)\left(1-T\right)}{1-e^{*}(\bs X)}\right\} \mathcal{L}_{j}(Y-\beta)\Bigr]
\).
Let $\underline{Q}^{\prime\prime}$ and $Q_{lip}^{\prime\prime}$ be two positive and finite constants. $Q_{j}(\beta;\pi)$ is twice differentiable with respect to $\beta$, and we denote its second-order derivative by $Q_{j}^{\prime\prime}(\beta;\pi)$. It holds that
$Q_{j}^{\prime\prime}(\beta_{j}^{*}(\pi);\pi)\geq\underline{Q}^{\prime\prime}$ uniformly over $\pi\in\Pi_{\infty}$ and
$\left|Q_{j}^{\prime\prime}(\beta;\pi)-Q_{j}^{\prime\prime}(\beta_{j}^{*}(\pi);\pi)\right|\leq Q_{lip}^{\prime\prime}\cdot\left|\beta-\beta_{j}^{*}(\pi)\right|$ for all $\beta$ satisfying $\left|\beta-\beta_{j}^{*}(\pi)\right|\leq c_{0}$
and all $\pi\in\Pi_{\infty}$.
\item $\sup_{\pi\in\Pi_{\infty}}\sup_{\beta:\left|\beta-\beta_{j}^{*}(\pi)\right|\leq c_{0}}\left|\mathcal{L}_{j}^{\prime}(Y-\beta)\right|\leq M/4$ almost surely, where $0<M<\infty$ is a constant that may depend on $\underline{Q}^{\prime\prime}$ and $Q_{lip}^{\prime\prime}$.
\end{enumerate}
\end{assumption}
\begin{assumption}
[Regularity assumptions on $U$]\label{assu:regularity-assumptions-on-U} Denote $\Psi(\bs{\beta};\pi)$ by
\(
\mathbb{E}\bigl[\bigl\{ \pi(\bs X)T/e^{*}(\bs X)+\allowbreak(1-\pi(\bs X))(1-T)/(1-e^{*}(\bs X))\bigr\} \allowbreak U(Y,\bs X,\bs{\beta})\bigr]
\).
The following conditions hold.
\begin{enumerate}[label=(\roman*)]
\item There exists a finite constant $M>0$ such that the bound
$\sup_{\pi\in\Pi_{\infty}}\sup_{\bs{\beta}:\left\Vert \bs{\beta}-\bs{\beta}^{*}(\pi)\right\Vert \leq c_{0}}\allowbreak\left|U(Y,\bs X,\bs{\beta})\right|\leq M/4$
holds almost surely.
\item The function class $\mathcal{U}:=\{(Y,\bs X)\mapsto U(Y,\bs X,\bs{\beta}):\allowbreak\left\Vert \bs{\beta}-\bs{\beta}^{*}(\pi)\right\Vert \leq c_{0},\allowbreak\pi\in\Pi_{\infty}\}$
satisfies $\sup_{Q}\log N\left(\frac{M}{4}\epsilon,\mathcal{U},\left\Vert \cdot\right\Vert _{Q,2}\right)\leq\nu\log\left(a/\epsilon\right)\text{ for all }0<\epsilon<1,$
where $a,\nu>0$ are finite constants, and $\sup_{Q}$ is taken over
all finitely discrete measures.
\item Given any policy $\pi\in\Pi_{\infty}$, $\Psi(\bs{\beta};\pi)$ is differentiable
with respect to $\bs{\beta}$, and we denote its gradient by $\nabla\Psi(\bs{\beta};\pi)$.
There exists a finite constant $\overline{\Psi}^{\prime}\geq0$ such that $\left\Vert \nabla\Psi(\bs{\beta};\pi)\right\Vert \leq\overline{\Psi}^{\prime}$
for all $\bs{\beta}$ satisfying $\left\Vert \bs{\beta}-\bs{\beta}^{*}(\pi)\right\Vert \leq c_{0}$ and all $\pi\in\Pi_{\infty}$.
\end{enumerate}
\end{assumption}
\begin{assumption}
\label{assu:smoothness for mu} Let $s_{\mu}>d/2$ be an integer. It
holds that
\begin{align*}
 & \left\{ \bs X\mapsto\mu^{*}_{jt}(\bs X;\bs{\beta}):\left\Vert \bs{\beta}-\bs{\beta}^{*}(\pi)\right\Vert \leq c_{0},\pi\in\Pi_{\infty},t\in\{0,1\},j=0,\ldots,p\right\} \\
\subset & \  C^{s_{\mu}}\left([0,1]^{d}\right):=\left\{ f:\max_{\bs{\alpha}\in\mathbb{N}^{d},\left\Vert \bs{\alpha}\right\Vert _{1}\leq s_{\mu}}\underset{\bs x\in[0,1]^{d}}{\sup}\left|\partial^{\bs{\alpha}}f(\bs x)\right|\leq M\right\}
\end{align*}
for some constant $M>0$.
\end{assumption}
\begin{assumption}
\label{assu:regularity on mu} There exist finite constants $L_{\mu}>0$ and
$c_{\xi}>0$ such that the following conditions hold.
\begin{enumerate}[label=(\roman*)]
\item     For any $\bs{\beta}_{1},\bs{\beta}_{2}\in\mathbb{R}^{p}$, $j=0,\ldots,p$, and $t\in\{0,1\}$: $\left\Vert \mu_{jt}^{*}(\bs X;\bs{\beta}_{1})-\mu_{jt}^{*}(\bs X;\bs{\beta}_{2})\right\Vert _{P,2}\leq L_{\mu}\left\Vert \bs{\beta}_{1}-\bs{\beta}_{2}\right\Vert$.
\item For $t\in\{0,1\}$, define $\bs{\xi}_{t}^{*}(\bs X;\pi):=\left(\mu_{jt}^{*}(\bs X;\bs{\beta}^{*}(\pi)):j=0,\ldots,p\right)^{\top}$. Then
\[
\inf_{\pi\in\Pi_{\infty}}\lambda_{\min}\Bigl\{\mathbb{E}\bigl[(1-\pi(\bs X))\bs{\xi}_{0}^{*}(\bs X;\pi)\bs{\xi}_{0}^{*}(\bs X;\pi)^{\top}+\pi(\bs X)\bs{\xi}_{1}^{*}(\bs X;\pi)\bs{\xi}_{1}^{*}(\bs X;\pi)^{\top}\bigr]\Bigr\}\geq c_{\xi}.
\]
\end{enumerate}
\end{assumption}
\begin{assumption}
\label{assu:regularity-assumptions-on-nuisance} The training sample size satisfies $|I|\to\infty$ as $N\to\infty$.
Throughout this assumption, $\widehat{e}_{I}$, $\widehat{\bs{\beta}}_{I}^{\mathrm{init}}(\pi)$, and $\widehat{\mu}_{I,jt}$, $j=0,\ldots,p$ and $t\in\{0,1\}$, refer to the estimators defined in the respective equations (\ref{eq:logistic reg for propensity}), (\ref{eq:beta-init-pi-I}), and (\ref{eq:DNN reg for L})--(\ref{eq:DNN reg for U}).
\begin{enumerate}[label=(\roman*)]
\item Almost surely, $\left|\log\{\widehat{e}_{I}(\bs X)/(1-\widehat{e}_{I}(\bs X))\}\right|\leq M$, and $\left|\widehat{\mu}_{I,jt}(\bs X;\bs{\beta})\right|\leq M$ uniformly over $\pi\in\Pi_{\infty}$,
$\left\Vert \bs{\beta}-\bs{\beta}^{*}(\pi)\right\Vert \leq c_{0}$, $j=0,\ldots,p$, and $t=0,1$, where $c_{0}$ is from Assumption~\ref{assu:regularity-conditions-on-L}.
\item The DNN classes $\mathcal{F}_{\mathrm{DNN}}(\mathcal{H}_{e},\mathcal{D}_{e})$ and $\mathcal{F}_{\mathrm{DNN}}(\mathcal{H}_{\mu},\mathcal{D}_{\mu})$
defined in (\ref{eq:DNN class}) and used in (\ref{eq:DNN reg for L})--(\ref{eq:DNN reg for U}) are constructed as follows. The first uses $\mathcal{H}_{e}\mathcal{D}_{e}\asymp\left|I\right|^{d/(4s_{e}+2d)}(\log\left|I\right|)^{2}$, and the second uses $\mathcal{H}_{\mu}\mathcal{D}_{\mu}\asymp\left|I\right|^{d/(4s_{\mu}+2d)}(\log\left|I\right|)^{2}$. In addition, their minimum diverges to infinity, and their logarithms are $O(\log |I|)$.
\end{enumerate}
\end{assumption}
Assumption~\ref{assu:Overlap assumption and smoothness} is a smoothness condition that is commonly recognized in the literature on deep neural network estimation \citep{farrell2021Deepa,jiao2023Deep,schmidt-hieber2020Nonparametric}, as well as in the broader nonparametric-estimation literature \citep{chen2007Large}. The conditions outlined in Assumption~\ref{assu:regularity-conditions-on-L} are for estimating $\bs{\beta}^{*}(\pi)$, allowing for non-smooth objectives such as $\mathcal{L}_{j}(v)=v(0.5-1(v\leq0))$. Similarly, the conditions in Assumption~\ref{assu:regularity-assumptions-on-U} are for estimating $W(\pi)$ and are satisfied by many utility functions $U$. Both assumptions are well-established in the literature \citep{vaart1998Asymptotic,vaart1996Weak}.

It is important to note that Assumptions~\ref{assu:regularity-conditions-on-L} and~\ref{assu:regularity-assumptions-on-U} hold uniformly over $\pi\in\Pi_{\infty}$. The uniform condition is critical for Assumption~\ref{assu:general ass uniform convergence}, which requires uniform convergence of $\widehat{W}_{I}(\pi)$. If only Assumption~\ref{assu:general assu on hatW} is necessary, a pointwise-in-$\pi$ version of these conditions is sufficient. Assumption~\ref{assu:smoothness for mu} is a smoothness condition on the conditional mean function $\mu_{jt}^{*}(\bs X_{i};\bs{\beta})$, which is analogous to Assumption~\ref{assu:Overlap assumption and smoothness}. Assumption~\ref{assu:regularity on mu} imposes Lipschitz continuity on $\mu_{jt}^{*}(\bs X_{i};\bs{\beta})$ in relation to $\bs{\beta}$ and includes a population non-singularity condition involving $\mu_{jt}^{*}(\bs X;\bs{\beta}^{*}(\pi))$, $j=0,\ldots,p$ and $t\in\{0,1\}$. The Lipschitz condition is satisfied by both $\mathcal{L}_{j}(v)=v(0.5-1(v\leq0))$ and $\mathcal{L}_{j}(v)=v^{2}/2$. The non-singularity condition eliminates linear redundancy among the functions $\mu_{jt}^{*}(\bs X;\bs{\beta}^{*}(\pi))$; if this condition fails, redundant functions can be removed before applying the calibration step. Assumption~\ref{assu:regularity-assumptions-on-nuisance}(i) imposes boundedness on the nuisance estimates, while Assumption~\ref{assu:regularity-assumptions-on-nuisance}(ii) restricts the width and depth of the deep neural networks. This latter requirement is familiar in the deep neural network estimation literature \citep[see, e.g.,][]{farrell2021Deepa,schmidt-hieber2020Nonparametric,jiao2023Deep}.

Under these sufficient conditions, we show that the proposed welfare criterion estimator meets the high-level assumptions outlined in Section~\ref{sec:A-General-Learning}.

\begin{thm}
\label{thm:oracle holdout unknown propensity} Suppose that Assumptions
\ref{assu:unconfounded}, \ref{assu:Overlap assumption and smoothness},
\ref{assu:regularity-conditions-on-L},
\ref{assu:regularity-assumptions-on-U},
\ref{assu:smoothness for mu}, \ref{assu:regularity on mu}, and
\ref{assu:regularity-assumptions-on-nuisance}
hold. Additionally, suppose that $\left|\widehat{W}_{I}(\pi)\right|\leq C$ almost surely for
any $\pi\in\Pi_{\infty}$ and $I\subset\{1,\ldots,N\}$, where $C>0$
is a finite constant. Under these conditions, the debiased welfare criterion estimator
$\widehat{W}_{I}(\pi)$, defined in (\ref{eq:def-of-W-hat-pi-I-unknown-propensity}),
satisfies Assumptions~\ref{assu:general assu on hatW} and \ref{assu:general ass uniform convergence}.
Furthermore, the welfare function $W(\pi)$, defined in (\ref{eq:identification-of-W-pi}),
satisfies Assumption~\ref{ass:high-level}.
\end{thm}
Combined with Theorem~\ref{thm:oracle inequality-general}, Theorem~\ref{thm:oracle holdout unknown propensity} implies that the average welfare regret of the proposed policy-learning procedure adheres to the oracle inequality stated in~(\ref{eq:oracle inequality}); combined with Corollary~\ref{cor:oracle inequality prob bound}, it also yields the corresponding high-probability welfare regret bound.


\section{Empirical Application} \label{sec:empirical application}

To illustrate the practical value of the proposed policy learning procedure, we apply it to data from the National Job Training Partnership Act (JTPA) Study. This large-scale randomized controlled trial was commissioned by the U.S. Department of Labor to evaluate the effectiveness of publicly funded job-training programs. This dataset has become a benchmark in the policy evaluation and policy learning literature \citep{crippa2025Regret,ai2026data,abadie2002Instrumental,liu2025nonparametric, kitagawa2018Who, mbakop2021Model}.\footnote{The sample we use is taken from the supplementary materials of \cite{mbakop2021Model}, available at \url{https://onlinelibrary.wiley.com/doi/10.3982/ECTA16437}.}

Our analysis uses a sample of $N=11{,}008$ individuals. For each individual, we observe two baseline covariates: years of education ($X_{1}$) and pre-program earnings ($X_{2}$). The outcome of interest, $Y_i$, is the total earnings over the 30-month period following random assignment. Let $T\in\{0,1\}$ denote the randomized treatment assignment (the training offer), which means that the propensity score $e^{*}(\bs X)=P(T=1\mid \bs X)$ is constant and equal to $2/3$. While the true propensity score is known in this sample, we deliberately treat it as unknown to demonstrate the applicability of the method in situations where assignment probabilities are unavailable, partially observed, or require estimation.


We focus on a specific class of monotone, interpretable allocation rules, guided by the principle that, \emph{all else being equal}, individuals with lower socioeconomic status (such as less education or lower earnings) should be (weakly) prioritized for training. Let $\mathcal X_1$ and $\mathcal X_2$ represent the supports of education and pre-program earnings, respectively. We define the policy space as:
\[
\Pi_{\infty}
=
\left\{
\pi: \mathcal{X}_1 \times \mathcal{X}_2 \to \{0,1\}
:
\pi(x_1, x_2) = 1\!\left(f(x_1) \geq x_2\right)
\text{ for some non-increasing } f
\right\}.
\]
The policy $\pi(x_1,x_2)=1\{x_2\le f(x_1)\}$ assigns an individual to training whenever their pre-program earnings fall below an education-specific cutoff $f(x_1)$. The restriction that $f$ is non-increasing ensures that the cutoff is (weakly) higher for individuals with less education, making the earnings criterion more lenient for them. This rule is transparent: it can be represented as a treatment region in the $(x_1,x_2)$-plane or, equivalently, as an estimated cutoff curve $\widehat f(x_1)$.

Previous studies (e.g., \citealp{mbakop2021Model}) have optimized this class of rules using a linear welfare criterion that maximizes average outcomes $\mathbb{E}[Y^{*}(\pi(\bs{X}))]$. However, such an objective neglects distributional concerns: a policy designed to maximize average income may inadvertently increase income disparities. To address this trade-off between \textit{efficiency} (aggregate income) and \textit{equity} (income dispersion), we adopt a nonlinear welfare criterion that penalizes outcome dispersion. Specifically, we aim to maximize the ratio of the mean outcome to its standard deviation, known as the inverse coefficient of variation:
\begin{equation}
    W_{\mathrm{ICV}}(\pi) = \frac{\mathbb{E}\left[Y^{*}(\pi(\bs{X}))\right]}{\sqrt{\mathrm{Var}\left(Y^{*}(\pi(\bs{X}))\right)}},
\end{equation}
where $\mathrm{Var}\left(Y^{*}(\pi(\bs{X}))\right) = \mathbb{E}\left[Y^{*}(\pi(\bs{X}))^2\right] - \left(\mathbb{E}\left[Y^{*}(\pi(\bs{X}))\right]\right)^2$. This objective is rooted in the axiomatic literature on inequality measurement (e.g., \citealp{atkinson1970Measurement}), which emphasizes that social welfare evaluations should balance efficiency (mean outcomes) against equity (distributional fairness). Maximizing $W_{\mathrm{ICV}}(\pi)$ is equivalent to minimizing the coefficient of variation, a scale-invariant measure of inequality that penalizes dispersion relative to the mean.

To reformulate this objective within our framework, we express it using the auxiliary parameters $\bs{\beta}^{*}(\pi)$ and the utility function $U(\cdot)$ introduced in Section~\ref{sec:general framework}. Since earnings are non-negative in our sample and the policy mean is positive for the policies considered here, maximizing $W_{\mathrm{ICV}}(\pi)$ is equivalent to maximizing its square. Simple algebra shows that maximizing $W_{\mathrm{ICV}}(\pi)^2$ is equivalent to maximizing the negative ratio of the second moment to the squared first moment:
\begin{equation} \label{eq:obj_transformed}
    \pi^* = \underset{\pi \in \Pi_{\infty}}{\arg\max}
    \left\{
    - \frac{\mathbb{E}\left[Y^*(\pi(\bs X))^2\right]}{\left(\mathbb{E}\left[Y^*(\pi(\bs X))\right]\right)^2}
    \right\}.
\end{equation}
This problem fits directly into our general framework, with a single auxiliary parameter ($p=1$). We define $\beta^{*}_1(\pi)$ to be the population mean of the potential outcome, corresponding to the quadratic loss $\mathcal{L}_1(v) = v^2/2$:
\[
    \beta_1^*(\pi) := \underset{\beta \in \mathbb{R}}{\arg\min}\ \mathbb{E}\left[ \frac{1}{2}\left( Y^*(\pi(\bs X)) - \beta \right)^2 \right]
    = \mathbb{E}\left[ Y^*(\pi(\bs X)) \right].
\]
The utility function is given by \(U(y, \bs x, \beta_1) := -y^2/\beta_1^2\).
With a slight abuse of notation, we denote the welfare criterion as:
\[
    W(\pi)=\mathbb{E}\left[ U(Y^*(\pi(\bs X)), \bs X, \beta_1^*(\pi)) \right]
    = - \frac{\mathbb{E}[Y^*(\pi(\bs X))^2]}{(\mathbb{E}[Y^*(\pi(\bs X))])^2}.
\]
We learn the optimal policy $\pi^*$ from observational data by maximizing $\widehat{W}_{I}(\pi)$ as defined in equation (\ref{eq:def-of-W-hat-pi-I-unknown-propensity}) over the sieve approximating sequence described in Example~\ref{exa:monotone policies}. We use a $5$-fold cross-validation procedure (\ref{eq:CV criterion-general}) to select the best subclass.\footnote{Following \cite{mbakop2021Model}, this application contains only five candidate subclasses, $\Pi_1,\ldots,\Pi_{5}$.
}
To determine the best policy within each policy subclass, we employ the Strategic Monte Carlo Optimization (SMCO) algorithm as outlined in \citet{chen2026Optimization}. This algorithm demonstrates that, under suitable conditions, it converges to a local optimum from a single starting point and to a global optimum as the number of starting points increases. Consequently, we run SMCO from multiple starting points that are generated quasi-uniformly over a unit hypercube. This approach provides space-filling exploration of the parameter domain and enhances the robustness of the nonconvex search.\footnote{Following the subclass construction in \citet{mbakop2021Model}, the $\ell$th subclass can be indexed, after an appropriate reparametrization, by a vector $\bs{\theta}=(\theta_1,\ldots,\theta_{2^{\ell-1}+1})$ lying in a simplex-type set: $\theta_j\ge 0$ and $\sum_{j=1}^{2^{\ell-1}+1}\theta_j \le 1$. We construct an explicit mapping from the $(2^{\ell-1}+1)$-dimensional unit hypercube onto this set and run SMCO over the hypercube, which lets us generate quasi-uniform starting points while enforcing the simplex constraint by construction.}

 The welfare estimate $\widehat W_I(\pi)$ as described in equation (\ref{eq:def-of-W-hat-pi-I-unknown-propensity}) relies on several nuisance components, including the estimated propensity score $\widehat e_I(\cdot)$ and the estimated conditional mean functions $\widehat \mu_{I,jt}(\cdot;\bs\beta)$. We compute these estimates using deep neural networks, as detailed in Appendix~\ref{sec:app-nuisance-estimation}. The architecture of the network, including its depth and width, is determined through cross-validation on the same training sample used to train $\widehat W_I(\pi)$.

Figure~\ref{fig:best policy in Pi1 and Pi5} displays the best policies found in the simplest ($\Pi_{1}$) and the most complex ($\Pi_{5}$) subclasses of the approximating sequence.

\begin{figure}[h]
    \includegraphics[width=0.5\textwidth]{figures/welfare1_fig_1.pdf}\includegraphics[width=0.5\textwidth]{figures/welfare1_fig_5.pdf}

\caption{\label{fig:best policy in Pi1 and Pi5}The best policy found in the
    simplest ($\Pi_{1}$) and most complicated ($\Pi_{5}$) classes. The $x$-axis represents years of education, while the $y$-axis indicates pre-program earnings. The red and green shaded areas represent individuals assigned to the treatment and control groups, respectively, under the estimated optimal policy. Left
    panel: best policy learned within $\Pi_{1}$; right panel: best policy
    learned within $\Pi_{5}$.}

    \end{figure}

    Our $5$-fold cross-validation procedure identifies $\Pi_{1}$ as the best subclass, and we denote the learned optimal policy as $\widehat{\pi}_{\mathrm{Nonlin}}$. We compare $\widehat{\pi}_{\mathrm{Nonlin}}$ with the benchmark policy $\widehat{\pi}_{\mathrm{Lin}}$, which was derived from penalized welfare maximization under the linear criterion $\mathbb{E}[Y^{*}(\pi(\bs{X}))]$ as discussed in \cite{mbakop2021Model}. To evaluate each policy, we compute the mean and standard deviation of its associated potential outcomes based on the full sample, using a de-biased estimator as described in Section~\ref{sec:empirical welfare}. For a given policy $\widehat{\pi}$, these estimators are defined as follows:
    \[
    \widehat{\mathrm{Mean}}(\widehat{\pi}) = \frac{1}{N}\sum_{i=1}^{N}\widehat{w}_{i}(\widehat{\pi})\left\{\frac{\widehat{\pi}(\bs X_{i})T_{i}}{\widehat{e}(\bs X_{i})}+\frac{(1-\widehat{\pi}(\bs X_{i}))(1-T_{i})}{1-\widehat{e}(\bs X_{i})}\right\} Y_{i},
    \]
    and
    \[
    \widehat{\mathrm{SD}}(\widehat{\pi}) = \sqrt{\frac{1}{N}\sum_{i=1}^{N}\widehat{w}_{i}(\widehat{\pi})\left\{\frac{\widehat{\pi}(\bs X_{i})T_{i}}{\widehat{e}(\bs X_{i})}+\frac{(1-\widehat{\pi}(\bs X_{i}))(1-T_{i})}{1-\widehat{e}(\bs X_{i})}\right\} Y_{i}^{2} - \left(\widehat{\mathrm{Mean}}(\widehat{\pi})\right)^{2}},
    \]
    where $\widehat{e}(\bs X)$ is the DNN estimate of the propensity score using the full sample. The weights $\widehat{w}_{i}(\widehat{\pi})$ are determined by the following optimization problem:
    \[
    \begin{cases}
        (\widehat{w}_{1}(\widehat{\pi}),\ldots,\widehat{w}_{N}(\widehat{\pi}))=\underset{w_{i}>0:i=1,\ldots,N}{\arg\min}\sum_{i=1}^{N}\left(w_{i}\log w_{i}-w_{i}\right)\text{ subject to}\\
    0=\frac{1}{N}\sum_{i=1}^{N}(1-\widehat{\pi}(\bs X_{i}))\left\{ \frac{w_{i}(1-T_{i})}{1-\widehat{e}(\bs X_{i})}-1\right\} \widehat{\mathbb{E}}\left[Y\mid\bs X=\bs X_i,T=0\right]\\
    \quad+\frac{1}{N}\sum_{i=1}^{N}\widehat{\pi}(\bs X_{i})\left\{ \frac{w_{i}T_{i}}{\widehat{e}(\bs X_{i})}-1\right\} \widehat{\mathbb{E}}\left[Y\mid\bs X=\bs X_i,T=1\right],\\
    0=\frac{1}{N}\sum_{i=1}^{N}(1-\widehat{\pi}(\bs X_{i}))\left\{ \frac{w_{i}(1-T_{i})}{1-\widehat{e}(\bs X_{i})}-1\right\} \widehat{\mathbb{E}}\left[Y^{2}\mid\bs X=\bs X_i,T=0\right]\\
    \quad+\frac{1}{N}\sum_{i=1}^{N}\widehat{\pi}(\bs X_{i})\left\{ \frac{w_{i}T_{i}}{\widehat{e}(\bs X_{i})}-1\right\} \widehat{\mathbb{E}}\left[Y^{2}\mid\bs X=\bs X_i,T=1\right],
    \end{cases}
    \]
    where $\widehat{\mathbb{E}}[Y\mid\bs X=\bs x,T=t]$ and $\widehat{\mathbb{E}}[Y^{2}\mid\bs X=\bs x,T=t]$ are the DNN estimates of $\mathbb{E}[Y\mid\bs X=\bs x,T=t]$ and $\mathbb{E}[Y^{2}\mid\bs X=\bs x,T=t]$, respectively.

    The empirical results highlight the trade-off inherent in our method. The benchmark policy $\widehat{\pi}_{\mathrm{Lin}}$ yields a mean outcome of $16{,}201.57$, with a standard deviation of $16{,}763.98$. In contrast, the policy $\widehat{\pi}_{\mathrm{Nonlin}}$, estimated under our nonlinear welfare criterion, yields a mean outcome of $16{,}132.78$ and a standard deviation of $16{,}617.18$. Relative to the benchmark, $\widehat{\pi}_{\mathrm{Nonlin}}$ reduces the mean outcome by approximately $0.42\%$ and the standard deviation by about $0.88\%$, reflecting our goal of achieving lower outcome dispersion under the nonlinear welfare criterion.

\section{Conclusion}\label{sec:conclusion}

This paper presents a data-driven policy-learning procedure that utilizes observational data to address a nonlinear welfare criterion within an infinite-dimensional policy space. The proposed learning procedure expands the existing literature on policy learning by moving from a linear (utilitarian) welfare criterion to a nonlinear one, transitioning from finite-dimensional to infinite-dimensional policy spaces, and shifting focus from a known propensity score to an unknown one. Additionally, we introduce a novel reweighting-based debiasing method, providing a valuable alternative to the current double debiasing approach. We applied this procedure to the JTPA study, where we found a balance between efficiency and equity.

However, a significant challenge remains in the computational aspect: determining the best policy $\widehat{\pi}_{\ell,I}$ is fundamentally a nonconvex optimization problem. Due to the nonlinearity of the welfare criterion and the structure of the policy space, multiple local optima may arise. Future research should aim to develop more efficient optimization techniques, such as tighter convex relaxations or advanced heuristic search algorithms, to tackle these computational challenges.