Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
65,812 characters · 7 sections · 47 citation commands
There is an increasing body of literature on policy learning. Most existing studies focus on a linear welfare criterion and a binary treatment setting. In this context, the binary variable $T$ indicates program participation status: $T=1$ if the individual participates and $T=0$ if not. The potential outcome associated with participation status $t$ is represented as $Y^{*}(t)\in\mathcal{Y}\subset\mathbb{R}$, and $\bs X\in\mathcal{X}\subset\mathbb{R}^{d}$ denotes individual characteristics. A policy $\pi:\mathcal{X}\to\{0,1\}$ decides whether to assign an individual with attributes $\bs{X}$ to the program. The potential outcome under the policy is given by: \[ Y^{*}(\pi(\bs X))=\pi(\bs X)Y^{*}(1)+(1-\pi(\bs X))Y^{*}(0). \] Policymakers specify a welfare criterion $W(\pi)$ and a policy space $\Pi_{\infty}$, aiming to learn the optimal policy defined by \[ \pi^{*}=\arg \max_{\pi \in \Pi_{\infty}} W(\pi). \] A common choice for the welfare is the expectation of the potential outcome: \[ W(\pi)=\mathbb{E}[Y^{*}(\pi(\bs X))]=\mathbb{E}[\pi(\bs X)Y^{*}(1)+(1-\pi(\bs X))Y^{*}(0)], \] which is clearly linear with respect to the policy.
A risk-averse policymaker might prefer to utilize the average utility of potential outcomes, represented as \[ W(\pi)=\mathbb{E}[U(Y^{*}(\pi(\bs X)))]=\mathbb{E}[\pi(\bs X)U(Y^{*}(1))+(1-\pi(\bs X))U(Y^{*}(0))]. \] However, this extension is only superficial: redefining the potential outcome as $U(Y^*(t))$ demonstrates that the resulting welfare remains linear in policy. By contrast, many significant applications lead to welfare criteria that are genuinely nonlinear in policy. For instance, in income-inequality literature, policymakers aim to reduce disparities through targeted interventions, and the welfare criterion is often a measure of inequality, such as the Gini coefficient gastwirth1971General,gastwirth1972Estimation: \[ W(\pi)= -\frac{2M^{-1}\sum_{r=1}^{M}\alpha_{r}\beta^{*}(\pi;\alpha_{r})}{\mathbb{E} [Y^{*}(\pi(\bs X))]} + 1,\qquad \alpha_{r}=\frac{r}{M+1}, \] where $M$ is a finite positive integer and $\beta^{*}(\pi;\alpha)$ denotes the $\alpha$-quantile of $Y^{*}(\pi(\bs X))$. Since quantiles are nonlinear in relation to policy choices, the welfare criterion is likewise nonlinear. Nonlinearity also appears in alternative inequality measures, such as the relative standing of specific population segments: \[ W(\pi)=\frac{\mathbb{E}\left[Y^{*}(\pi(\bs X))\mid Y^{*}(\pi(\bs X))\leq\beta^{*}(\pi;0.5)\right]}{\mathbb{E}\left[Y^{*}(\pi(\bs X))\right]}, \] or the relative status of a defined subpopulation: \[ W(\pi)=\frac{\mathbb{E}\left[Y^{*}(\pi(\bs X))1(\text{gender}=\text{``female''})\right]}{\mathbb{E}\left[Y^{*}(\pi(\bs X))\right]}, \] and in the analysis of upper-tail (90/50) and lower-tail (50/10) ratios autor2008trends: \[ W(\pi)=-\frac{\beta^{*}(\pi;0.9)}{\beta^{*}(\pi;0.5)}\text{ and }W(\pi) = -\frac{\beta^{*}(\pi;0.5)}{\beta^{*}(\pi;0.1)}. \]
Nonlinear welfare criteria naturally arise in the fields of risk management and public health. In risk management, firms aim to control the risk of significant losses. Let the binary variable $T\in\{0,1\}$ represent an investment decision, where $T=1$ indicates a one-period investment in the asset and $T=0$ indicates no investment. Let $Y^{*}(1)$ denote the corresponding one-period payoff (or return), while we set $Y^{*}(0)=0$. Given characteristics $\bs X$ (e.g., volatility, momentum, or fundamentals), an investment strategy $\pi:\mathcal X\to\{0,1\}$ induces the realized payoff: \[ Y^{*}(\pi(\bs X))=\pi(\bs X)Y^{*}(1)+(1-\pi(\bs X))Y^{*}(0)=\pi(\bs X)Y^{*}(1). \] We denote $L^{*}(\pi)=-Y^{*}(\pi(\bs X))$ and let $\beta^{*}(\pi;\alpha)$ represent the $\alpha$-quantile of $L^{*}(\pi)$. The welfare criterion can be expressed as a measure of risk, such as Value-at-Risk (VaR), defined as $W(\pi)=-\beta^{*}(\pi;\alpha)$, or Conditional Value-at-Risk (CVaR), given by $W(\pi)=-(1-\alpha)^{-1}\mathbb{E}\!\left[L^{*}(\pi)\cdot 1\!\left(L^{*}(\pi)\ge \beta^{*}(\pi;\alpha)\right)\right]$ rockafellar2000Optimization, or the spectral risk measure (e.g., acerbi2002Spectral): \[ W(\pi)=-\sum_{r=1}^{M}\phi(\alpha_{r})\beta^{*}(\pi;\alpha_{r}),\qquad \alpha_{r}=\frac{r}{M+1}, \] with a weighting function $\phi(\alpha)$ (see dowd2006Var). All of these risk measures are nonlinear in relation to policy.
In public health, policymakers aim to reduce the incidence of severe post-discharge utilization through targeted programs. For example, in the United States, Medicare's Hospital Readmissions Reduction Program (HRRP) specifically targets 30-day unplanned readmissions by penalizing hospitals with higher-than-average readmission rates zuckerman2016readmissions,ryan2017valuebased,khera2018hrrp. Research shows that readmission-related utilization is highly right-skewed: a small fraction of patients accounts for a disproportionate share of readmissions and hospital use fouayzi2022highfrequency,manning2005generalized. These distributional features motivate policymakers to prioritize the upper tail of the outcome distribution rather than focusing solely on the average. In this context, let $T$ represent the intensity of post-discharge care for a patient, where $T=1$ indicates enhanced transitional care (e.g., intensive monitoring and follow-up) and $T=0$ signifies usual care. The variable $Y^{*}(t)\geq 0$ reflects the readmission burden under treatment $t$, such as the total number of inpatient days due to unplanned readmissions within 30 days post-discharge, with lower values being more desirable. Given patient covariates $\bs X$ (e.g., age, gender, BMI, and comorbidity indices), a policy $\pi:\mathcal X\to\{0,1\}$ assigns care intensity and determines the burden $Y^{*}(\pi(\bs X))$. The welfare criterion being considered is a tail-sensitive measure of the expected burden among the worst-off $(1-\alpha)$ fraction of patients: \[ W(\pi) = -\frac{1}{1-\alpha}\mathbb{E}\!\left[ Y^{*}(\pi(\bs X))\cdot 1\!\left\{Y^{*}(\pi(\bs X))\ge \beta^{*}(\pi;\alpha)\right\} \right], \] where $\beta^{*}(\pi;\alpha)$ denotes the $\alpha$-quantile of $Y^{*}(\pi(\bs X))$ and is nonlinear in relation to the policy.
All the aforementioned examples emphasize the need for a nonlinear welfare criterion that directly depends on the policy through potential outcomes and indirectly through intermediate parameters $\bs{\beta}^*(\pi)$ (such as the quantiles mentioned earlier). Specifically, we model this with:
where $U(\cdot,\cdot,\cdot)$ is a known utility function.
The policy space can be complex and infinite-dimensional, making the optimal policy $\pi^{*}$ challenging to compute. To overcome this issue, we leverage the sieve literature to approximate the policy space using a sequence of finite-dimensional sieve classes $\{\Pi_{\ell}: \ell=1,2,\ldots\}$. Within each class, we determine the best policy as \[\pi^{*}_{\ell}=\arg\max_{\pi\in \Pi_{\ell}}W(\pi). \] Since the welfare criterion is unknown, we estimate it from a training sample, say $I$, by $\widehat{W}_{I}(\pi)$ and find the best policy using: \[ \widehat{\pi}_{\ell,I}=\arg\max_{\pi\in \Pi_{\ell}}\widehat{W}_{I}(\pi). \] We then apply $K$-fold cross-validation to select the optimal policy approximation space $\Pi_{\widehat{\ell}}$ (see ((ref))) and subsequently estimate the optimal policy as $\widehat{\pi}$ (see ((ref))). Despite these extensions, we establish the following oracle inequality for the average welfare regret:
Related Literature. Our work builds on the policy-learning literature initially developed by manski2004Statistical. A significant portion of this literature examines treatment choices under linear welfare criteria, in which an average or utilitarian objective defines the policy value. Early contributions to this field include studies by hirano2009asymptotics,stoye2009minimax,stoye2012minimax,bhattacharya2012inferring,tetenov2012statistical,qian2011Performance,zhao2012Estimating, and more recent works by kitagawa2018Who,athey2021Policy,mbakop2021Model,luedtke2016Statistical,crippa2025Regret,liu2025nonparametric,fang2025semiparametric,fang2025model. With the exception of mbakop2021Model, all these studies assume a finite-dimensional policy space (e.g., $\Pi_{\infty}=\Pi_{\ell}$ with $\ell$ fixed or allowed to grow with sample size), which allows them to avoid the complexities of policy-space approximation and data-driven model selection. They either assume a known propensity score or estimate it using machine-learning methods, then apply double debiasing to remove the machine-learning bias. In both cases, they ultimately establish an explicit upper bound on the welfare regret:
where $\mathrm{VC}(\Pi_{\ell})$ represents the Vapnik-Chervonenkis (VC) dimension of the policy class. In contrast, mbakop2021Model addresses an infinite-dimensional policy space with a known propensity score, thus avoiding the need for debiasing. They present an upper bound on the average welfare regret that is similar to ours, though still under a linear welfare criterion. Our contribution extends this literature to encompass a nonlinear welfare criterion, a machine-learned propensity score, and an infinite-dimensional policy space.
A smaller body of literature explores policy learning for nonlinear welfare criteria within a finite-dimensional policy class, where the dimension may increase with sample size (that is, $\Pi_{\infty}=\Pi_{\ell}$, with $\ell$ allowed to grow with sample size). Notable examples include wang2018QuantileOptimal on quantile-optimal treatment regimes, chen2025quantileoptimalpolicylearningunmeasured addressing a quantile-based welfare criterion with an unmeasured confounder, fan2025policylearningalphaexpectedwelfare focusing on a conditional value-at-risk criterion, kitagawa2021EqualityMinded discussing an equality-minded social welfare criterion, and terschuur2025locally examining nonlinear welfare criteria defined through U-statistics. These studies assume a finite-dimensional policy space, treat the propensity score as either known or derived via machine learning, and employ double-debiasing techniques to mitigate machine-learning bias. Despite the nonlinearity, they manage to derive a similar, explicit upper bound on the average welfare regret, akin to the upper bound mentioned in equation ((ref)). We aim to extend this existing literature to encompass infinite-dimensional policy spaces.
To learn the optimal policy from observational data, it is essential to estimate both the intermediate parameters and the welfare criterion. Since both estimates rely on the unknown propensity score, we use machine-learning algorithms to estimate it. It is well recognized that machine-learned propensity scores introduce bias in both the intermediate parameter and the welfare-criterion estimates, which in turn affects the (average) welfare regret and slows its convergence rate. To correct for this machine-learning bias, existing literature applies double-debiasing procedures based on a Neyman orthogonality condition (see athey2021Policy,robins1994Estimation,chernozhukov2018Double for linear welfare criteria and fan2025policylearningalphaexpectedwelfare,terschuur2025locally for nonlinear criteria). Our procedure includes an additional step, estimating the intermediate parameters, so we need to debias both these parameters and the welfare criterion estimates simultaneously. We propose a reweighting method inspired by covariate-balancing approaches imai2014Covariate,chan2016Globally,ai2021Unified, adapted here for bias correction. Covariate-balancing methods have shown strong performance in finite samples, and we expect our debiasing procedure to exhibit similar effectiveness. To our knowledge, this reweighting-based debiasing technique is novel in the literature and serves as a valuable alternative to double debiasing based on Neyman orthogonality.
The remainder of the paper is organized as follows. Section (ref) presents a data-automated optimal policy learning procedure that employs a generic welfare estimate alongside $K$-fold cross-validation to identify the best policy subclass, while establishing an oracle inequality for both the average welfare regret and the welfare regret. Section (ref) formally defines the model, expressing the unknown parameters and the welfare criterion in terms of the observed data, based on the assumptions of unconfoundedness and overlap. Section (ref) introduces a machine-learning propensity-score estimator, along with a novel reweighting debiasing procedure for estimating the intermediate parameters and the welfare criterion. Section (ref) verifies that the proposed estimator for the welfare criterion meets the high-level conditions outlined in Section (ref). Section (ref) applies the proposed methodology to data from the National Job Training Partnership Act (JTPA) Study. Finally, Section (ref) summarizes the findings. All omitted proofs are included in the Appendix.
We will outline a data-driven policy-learning procedure that relies on a generic welfare estimator and cross-validation. Let $\left\{ \left(Y_{i},\bs X_{i},T_{i}\right)\right\} _{i=1}^{N}$ represent an independent and identically distributed (i.i.d.) sample. Throughout the paper, we use $I\subset\{1,2,\ldots,N\}$ to index a generic training subsample and $|I|$ to denote its sample size. Let $\widehat{W}_{I}(\pi)$ be a generic estimator of $W(\pi)$ calculated from the training subsample $\left\{ \left(Y_{i},\bs X_{i},T_{i}\right)\right\} _{i\in I}$. We assume that $\widehat{W}_{I}(\pi)$ is pointwise $\sqrt{\left|I\right|}$-consistent for $W(\pi)$ and that its estimation error satisfies an exponential probability bound, as formalized in Assumption (ref).
In applications, users must verify that their welfare estimators satisfy this high-level condition. We will present a welfare estimator and confirm that it indeed satisfies Assumption (ref).
With $\widehat{W}_{I}(\pi)$ established, a natural policy learning strategy is to maximize it over $\Pi_{\infty}$. However, for an infinite-dimensional $\Pi_{\infty}$, such a strategy can be computationally intractable and prone to overfitting. Following the work of mbakop2021Model, we approximate $\Pi_{\infty}$ by a nested sequence of finite-dimensional policy subclasses $\Pi_{\ell}\subset\Pi_{\ell+1}\subset\cdots\subset\Pi_{\infty}$.\footnote{Throughout the paper, we take this approximating sequence to satisfy $\mathrm{VC}(\Pi_{\ell})<\infty$ for every $\ell\geq1$ and $W(\pi^{*})-\max_{\pi\in\Pi_{\ell}}W(\pi)\to0$ as $\ell\to\infty$.} Within each subclass $\Pi_{\ell}$, we estimate the best policy by
We determine the best subclass using a $K$-fold cross-validation (CV) procedure, as outlined in various studies (Hall1983Large, Stone1974Cross-Validatory, lecue2012Oracle, and gyorfi2002DistributionFree). Specifically, for a fixed integer $K\geq2$, we partition the index set $\{1,\ldots,N\}$ into $K$ disjoint folds of equal size. For each pair $(k,\ell)$, let $I_{k}$ denote the $k$th fold and $I_{-k}=\{1,\ldots,N\}\setminus I_{k}$ its complement. We utilize the training subsample $I_{-k}$ to learn the welfare $\widehat{W}_{I_{-k}}(\pi)$ and the candidate policy $\widehat{\pi}_{\ell,I_{-k}}$. The holdout sample $I_{k}$ is then used to compute $\widehat{W}_{I_{k}}(\pi)$ and to evaluate the performance of the candidate policy on the holdout sample using $\widehat{W}_{I_{k}}(\widehat{\pi}_{\ell,I_{-k}})$. The $K$-fold CV procedure selects the best subclass according to the formula:
where the penalty term $\log\ell/\sqrt{N}$ helps to prevent the selection of very large policy classes. The identified policy is then defined as:
The following theorem establishes the oracle inequality ((ref)) with $I=I_{-k}$ for any $k$.
The theorem generalizes the oracle inequality by providing a bound on the average welfare regret $\mathbb{E}[W(\pi^{*})-W(\widehat{\pi})]$. In practice, policymakers can only access a single sample, and are more concerned with the realized welfare regret $W(\pi^{*})-W(\widehat{\pi})$. We show that a similar oracle inequality holds with high probability.
Corollary (ref) establishes a probability oracle inequality. The constant “2” in front of $\log\ell/\sqrt{N}$ arises from a union bound used to control the concentration of $\frac{1}{K}\sum_{k=1}^{K}\widehat{W}_{I_{k}}(\widehat{\pi}_{\ell,I_{-k}})$ around $\frac{1}{K}\sum_{k=1}^{K}W(\widehat{\pi}_{\ell,I_{-k}})$. This constant can be replaced with any value greater than $1$, with corresponding adjustments to $C,C_1,C_2$.
The oracle inequality provides insights into the quality of the learned policy only when both the approximation and estimation errors are minimal. We can quantify the estimation error under a strengthened Assumption (ref).
Assumption (ref) imposes a uniform convergence rate on the welfare estimator. It allows the VC dimension of the policy class, $\Pi$, to increase with the sample size at a rate that is slower than the sample size itself: $\mathrm{VC}(\Pi)/|I|\to0$. This rate condition aligns, up to logarithmic factors, with the minimax rate for policy learning under linear welfare criteria athey2021Policy,kitagawa2018Who. Under this strengthened assumption, we derive an upper bound on the estimation error.
To quantify the approximation error, however, we need to gather information about the policy space and its approximation. For illustrative purposes, we will compute the approximation error for three common policy classes and their approximations. Throughout this process, we will maintain the following Lipschitz condition on $W(\pi)$.
In the next two examples, the sieve classes need not be subsets of $\Pi_{\infty}$. When these classes are used in the oracle bound, Assumptions (ref)--(ref) are imposed on these sieve classes directly, so Theorem (ref) continues to apply through the resulting approximation and estimation errors.
We now formally establish the model for the welfare criterion. We define the intermediate parameters $\bs{\beta}^{*}(\pi)\in\mathbb{R}^{p}$ as the minimizer of a sum of expected convex losses: \[ \bs{\beta}^{*}(\pi):=\underset{\bs{\beta}=(\beta_{1},\ldots,\beta_{p})^{\top}\in\mathbb{R}^{p}}{\arg\min}\sum_{j=1}^{p}\mathbb{E}\left[\mathcal{L}_{j}(Y^{*}(\pi(\bs X))-\beta_{j})\right], \] where $p\geq1$ is an integer and $\mathcal{L}_{1},\ldots,\mathcal{L}_{p}$ are known convex loss functions. Different selections of loss functions capture different distributional features. For example, when $p=2$, if we take $\mathcal{L}_{1}(v)=v^{2}/2$, and $\mathcal{L}_{2}(v)=v\cdot(0.5-1(v\leq0))$, the intermediate parameters correspond to the mean and the median: $\bs{\beta}^{*}(\pi)=\left(\mathbb{E}\left[Y^{*}(\pi(\bs X))\right],\text{median}[Y^{*}(\pi(\bs X))]\right)^{\top}$. Similarly, with the check loss defined as $\mathcal{L}(v)=v\cdot(\alpha - 1(v\leq 0))$, the intermediate parameter corresponds to the $\alpha$-quantile.
Both the intermediate parameters and the welfare criterion are expressed in terms of potential outcomes. To reframe them using the observed data $(Y,\bs X,T)$, where $Y=Y^{*}(T)$ represents the observed outcome, we impose the Stable Unit Treatment Value Assumption (SUTVA) imbens2015Causal and the following conditions regarding the data-generating process athey2021Policy,kitagawa2018Who,mbakop2021Model.
Under Assumption (ref), we can express the following optimization problem:
Additionally, we define:
The propensity score $e^{*}(\bs X)$ is unknown and is defined by $e^{*}(\bs X)=\mathbb{E} [T\mid\bs X]$, which can be found by solving the following optimization problem: \( e^{*}(\bs X)=\underset{e(\cdot)}{\arg\min}\left\{ \mathbb{E}\left[(T-e(\bs{X}))^2\right]\right\} \). However, a sample least-squares estimator based on this characterization may produce fitted values close to 0 or 1. We therefore use the following equivalent population formulation, understood over functions satisfying $0<e(\bs X)<1$:
Compared with the least-squares characterization, this criterion measures errors on the inverse-propensity scale that enters the inverse-probability-weighted (IPW) terms. Its population excess risk is a weighted sum of squared errors of $1/e(\bs X)$ and $1/\{1-e(\bs X)\}$, with the same minimizer $e^{*}$; in estimation, we minimize its sample analog over a bounded logistic DNN class to keep fitted propensity scores away from $0$ and $1$.
We present an estimator for the welfare criterion based on a generic training sample $I$. Equations ((ref))--((ref)) suggest a three-step sequential estimation procedure. In the first step, we estimate the propensity score using a sample analog of ((ref)). We then substitute this estimate into a sample analog of ((ref)) to estimate the intermediate parameters. Finally, we use both estimates to compute the welfare criterion using a sample analog of ((ref)).
To estimate the propensity score from the training sample $I$, we utilize a deep neural network, denoting the estimate as $\widehat{e}_{I}(\cdot)$ (see ((ref))). From the existing literature, it follows that $\left\Vert \widehat{e}_{I}-e^{*}\right\Vert_{P,2}= O_{P}\left(\left|I\right|^{-s_{e}/(2s_{e}+d)}\log^{3}\left|I\right|\right)$ (see the proof of Lemma (ref)), where $s_{e}$ indicates the smoothness of $e^{*}(\bs X)$ (see Assumption (ref)).
It is well-documented that machine learning can induce bias, which then propagates to the welfare criterion through both direct bias in the IPW terms and indirect bias in estimates of intermediate parameters. Simply substituting $e^{*}$ with $\widehat{e}_{I}$ in the sample analog of equations ((ref))--((ref)) may introduce bias in the intermediate parameters and welfare estimates. This, in turn, leads to a violation of Assumption (ref). To mitigate machine learning bias, we propose a weighted analog of equations ((ref))--((ref)):
The welfare estimator is defined as:
In this context, $\left\{ \widehat{w}_{I,i}(\pi):i\in I\right\}$ represents the calibrated weights. We show in Appendix (ref) that these weights must satisfy the following conditions:
For $t\in\{0,1\}$, we define: $\mu_{0t}^{*}(\bs x;\bs{\beta}):=\allowbreak\mathbb{E}\bigl[U(Y,\bs X,\bs{\beta})\allowbreak\mid\allowbreak\bs X=\bs x,T=t\bigr]$ and $\mu_{jt}^{*}(\bs x;\bs{\beta}):=\allowbreak\mathbb{E}\bigl[\mathcal{L}_{j}^{\prime}(Y-\beta_{j})\allowbreak\mid\allowbreak\bs X=\bs x,T=t\bigr]$ for $j=1,\ldots,p$, where $\mathcal{L}_{j}^{\prime}$ is understood as specified in Assumption (ref). In practice, $\bs{\beta}^{*}(\pi)$ and $\mu_{jt}^{*}(\cdot;\bs{\beta})$ are replaced with the initial estimator $\widehat{\bs{\beta}}_{I}^{\mathrm{init}}(\pi)$ and the conditional mean estimators $\widehat{\mu}_{I,jt}(\cdot;\widehat{\bs{\beta}}_{I}^{\mathrm{init}}(\pi))$, respectively, as detailed in Appendix (ref).
Notice that the weights satisfying the equations ((ref)) are generally not unique. We apply the entropy method to calibrate the weights as the solution to
In this context, the objective function $D(w)=w\log w-w$ measures the distance of $w$ from $1$, ensuring that the calibrated weights $\widehat{w}_{I,i}(\pi)$, $i\in I$, are unique and always non-negative.
Having constructed the welfare criterion estimator, we will now verify that it meets the high-level condition outlined in Section (ref). We require the following conditions.
Assumption (ref) is a smoothness condition that is commonly recognized in the literature on deep neural network estimation farrell2021Deepa,jiao2023Deep,schmidt-hieber2020Nonparametric, as well as in the broader nonparametric-estimation literature chen2007Large. The conditions outlined in Assumption (ref) are for estimating $\bs{\beta}^{*}(\pi)$, allowing for non-smooth objectives such as $\mathcal{L}_{j}(v)=v(0.5-1(v\leq0))$. Similarly, the conditions in Assumption (ref) are for estimating $W(\pi)$ and are satisfied by many utility functions $U$. Both assumptions are well-established in the literature vaart1998Asymptotic,vaart1996Weak.
It is important to note that Assumptions (ref) and (ref) hold uniformly over $\pi\in\Pi_{\infty}$. The uniform condition is critical for Assumption (ref), which requires uniform convergence of $\widehat{W}_{I}(\pi)$. If only Assumption (ref) is necessary, a pointwise-in-$\pi$ version of these conditions is sufficient. Assumption (ref) is a smoothness condition on the conditional mean function $\mu_{jt}^{*}(\bs X_{i};\bs{\beta})$, which is analogous to Assumption (ref). Assumption (ref) imposes Lipschitz continuity on $\mu_{jt}^{*}(\bs X_{i};\bs{\beta})$ in relation to $\bs{\beta}$ and includes a population non-singularity condition involving $\mu_{jt}^{*}(\bs X;\bs{\beta}^{*}(\pi))$, $j=0,\ldots,p$ and $t\in\{0,1\}$. The Lipschitz condition is satisfied by both $\mathcal{L}_{j}(v)=v(0.5-1(v\leq0))$ and $\mathcal{L}_{j}(v)=v^{2}/2$. The non-singularity condition eliminates linear redundancy among the functions $\mu_{jt}^{*}(\bs X;\bs{\beta}^{*}(\pi))$; if this condition fails, redundant functions can be removed before applying the calibration step. Assumption (ref)(i) imposes boundedness on the nuisance estimates, while Assumption (ref)(ii) restricts the width and depth of the deep neural networks. This latter requirement is familiar in the deep neural network estimation literature farrell2021Deepa,schmidt-hieber2020Nonparametric,jiao2023Deep.
Under these sufficient conditions, we show that the proposed welfare criterion estimator meets the high-level assumptions outlined in Section (ref).
Combined with Theorem (ref), Theorem (ref) implies that the average welfare regret of the proposed policy-learning procedure adheres to the oracle inequality stated in ((ref)); combined with Corollary (ref), it also yields the corresponding high-probability welfare regret bound.
To illustrate the practical value of the proposed policy learning procedure, we apply it to data from the National Job Training Partnership Act (JTPA) Study. This large-scale randomized controlled trial was commissioned by the U.S. Department of Labor to evaluate the effectiveness of publicly funded job-training programs. This dataset has become a benchmark in the policy evaluation and policy learning literature crippa2025Regret,ai2026data,abadie2002Instrumental,liu2025nonparametric, kitagawa2018Who, mbakop2021Model.\footnote{The sample we use is taken from the supplementary materials of mbakop2021Model, available at \url{https://onlinelibrary.wiley.com/doi/10.3982/ECTA16437}.}
Our analysis uses a sample of $N=11{,}008$ individuals. For each individual, we observe two baseline covariates: years of education ($X_{1}$) and pre-program earnings ($X_{2}$). The outcome of interest, $Y_i$, is the total earnings over the 30-month period following random assignment. Let $T\in\{0,1\}$ denote the randomized treatment assignment (the training offer), which means that the propensity score $e^{*}(\bs X)=P(T=1\mid \bs X)$ is constant and equal to $2/3$. While the true propensity score is known in this sample, we deliberately treat it as unknown to demonstrate the applicability of the method in situations where assignment probabilities are unavailable, partially observed, or require estimation.
We focus on a specific class of monotone, interpretable allocation rules, guided by the principle that, all else being equal, individuals with lower socioeconomic status (such as less education or lower earnings) should be (weakly) prioritized for training. Let $\mathcal X_1$ and $\mathcal X_2$ represent the supports of education and pre-program earnings, respectively. We define the policy space as: \[ \Pi_{\infty} = \left\{ \pi: \mathcal{X}_1 \times \mathcal{X}_2 \to \{0,1\} : \pi(x_1, x_2) = 1\!\left(f(x_1) \geq x_2\right) \text{ for some non-increasing } f \right\}. \] The policy $\pi(x_1,x_2)=1\{x_2\le f(x_1)\}$ assigns an individual to training whenever their pre-program earnings fall below an education-specific cutoff $f(x_1)$. The restriction that $f$ is non-increasing ensures that the cutoff is (weakly) higher for individuals with less education, making the earnings criterion more lenient for them. This rule is transparent: it can be represented as a treatment region in the $(x_1,x_2)$-plane or, equivalently, as an estimated cutoff curve $\widehat f(x_1)$.
Previous studies (e.g., mbakop2021Model) have optimized this class of rules using a linear welfare criterion that maximizes average outcomes $\mathbb{E}[Y^{*}(\pi(\bs{X}))]$. However, such an objective neglects distributional concerns: a policy designed to maximize average income may inadvertently increase income disparities. To address this trade-off between efficiency (aggregate income) and equity (income dispersion), we adopt a nonlinear welfare criterion that penalizes outcome dispersion. Specifically, we aim to maximize the ratio of the mean outcome to its standard deviation, known as the inverse coefficient of variation:
where $\mathrm{Var}\left(Y^{*}(\pi(\bs{X}))\right) = \mathbb{E}\left[Y^{*}(\pi(\bs{X}))^2\right] - \left(\mathbb{E}\left[Y^{*}(\pi(\bs{X}))\right]\right)^2$. This objective is rooted in the axiomatic literature on inequality measurement (e.g., atkinson1970Measurement), which emphasizes that social welfare evaluations should balance efficiency (mean outcomes) against equity (distributional fairness). Maximizing $W_{\mathrm{ICV}}(\pi)$ is equivalent to minimizing the coefficient of variation, a scale-invariant measure of inequality that penalizes dispersion relative to the mean.
To reformulate this objective within our framework, we express it using the auxiliary parameters $\bs{\beta}^{*}(\pi)$ and the utility function $U(\cdot)$ introduced in Section (ref). Since earnings are non-negative in our sample and the policy mean is positive for the policies considered here, maximizing $W_{\mathrm{ICV}}(\pi)$ is equivalent to maximizing its square. Simple algebra shows that maximizing $W_{\mathrm{ICV}}(\pi)^2$ is equivalent to maximizing the negative ratio of the second moment to the squared first moment:
This problem fits directly into our general framework, with a single auxiliary parameter ($p=1$). We define $\beta^{*}_1(\pi)$ to be the population mean of the potential outcome, corresponding to the quadratic loss $\mathcal{L}_1(v) = v^2/2$: \[ \beta_1^*(\pi) := \underset{\beta \in \mathbb{R}}{\arg\min}\ \mathbb{E}\left[ \frac{1}{2}\left( Y^*(\pi(\bs X)) - \beta \right)^2 \right] = \mathbb{E}\left[ Y^*(\pi(\bs X)) \right]. \] The utility function is given by \(U(y, \bs x, \beta_1) := -y^2/\beta_1^2\). With a slight abuse of notation, we denote the welfare criterion as: \[ W(\pi)=\mathbb{E}\left[ U(Y^*(\pi(\bs X)), \bs X, \beta_1^*(\pi)) \right] = - \frac{\mathbb{E}[Y^*(\pi(\bs X))^2]}{(\mathbb{E}[Y^*(\pi(\bs X))])^2}. \] We learn the optimal policy $\pi^*$ from observational data by maximizing $\widehat{W}_{I}(\pi)$ as defined in equation ((ref)) over the sieve approximating sequence described in Example (ref). We use a $5$-fold cross-validation procedure ((ref)) to select the best subclass.\footnote{Following mbakop2021Model, this application contains only five candidate subclasses, $\Pi_1,\ldots,\Pi_{5}$. } To determine the best policy within each policy subclass, we employ the Strategic Monte Carlo Optimization (SMCO) algorithm as outlined in chen2026Optimization. This algorithm demonstrates that, under suitable conditions, it converges to a local optimum from a single starting point and to a global optimum as the number of starting points increases. Consequently, we run SMCO from multiple starting points that are generated quasi-uniformly over a unit hypercube. This approach provides space-filling exploration of the parameter domain and enhances the robustness of the nonconvex search.\footnote{Following the subclass construction in mbakop2021Model, the $\ell$th subclass can be indexed, after an appropriate reparametrization, by a vector $\bs{\theta}=(\theta_1,\ldots,\theta_{2^{\ell-1}+1})$ lying in a simplex-type set: $\theta_j\ge 0$ and $\sum_{j=1}^{2^{\ell-1}+1}\theta_j \le 1$. We construct an explicit mapping from the $(2^{\ell-1}+1)$-dimensional unit hypercube onto this set and run SMCO over the hypercube, which lets us generate quasi-uniform starting points while enforcing the simplex constraint by construction.}
The welfare estimate $\widehat W_I(\pi)$ as described in equation ((ref)) relies on several nuisance components, including the estimated propensity score $\widehat e_I(\cdot)$ and the estimated conditional mean functions $\widehat \mu_{I,jt}(\cdot;\bs\beta)$. We compute these estimates using deep neural networks, as detailed in Appendix (ref). The architecture of the network, including its depth and width, is determined through cross-validation on the same training sample used to train $\widehat W_I(\pi)$.
Figure (ref) displays the best policies found in the simplest ($\Pi_{1}$) and the most complex ($\Pi_{5}$) subclasses of the approximating sequence.
Our $5$-fold cross-validation procedure identifies $\Pi_{1}$ as the best subclass, and we denote the learned optimal policy as $\widehat{\pi}_{\mathrm{Nonlin}}$. We compare $\widehat{\pi}_{\mathrm{Nonlin}}$ with the benchmark policy $\widehat{\pi}_{\mathrm{Lin}}$, which was derived from penalized welfare maximization under the linear criterion $\mathbb{E}[Y^{*}(\pi(\bs{X}))]$ as discussed in mbakop2021Model. To evaluate each policy, we compute the mean and standard deviation of its associated potential outcomes based on the full sample, using a de-biased estimator as described in Section (ref). For a given policy $\widehat{\pi}$, these estimators are defined as follows: \[ \widehat{\mathrm{Mean}}(\widehat{\pi}) = \frac{1}{N}\sum_{i=1}^{N}\widehat{w}_{i}(\widehat{\pi})\left\{\frac{\widehat{\pi}(\bs X_{i})T_{i}}{\widehat{e}(\bs X_{i})}+\frac{(1-\widehat{\pi}(\bs X_{i}))(1-T_{i})}{1-\widehat{e}(\bs X_{i})}\right\} Y_{i}, \] and \[ \widehat{\mathrm{SD}}(\widehat{\pi}) = \sqrt{\frac{1}{N}\sum_{i=1}^{N}\widehat{w}_{i}(\widehat{\pi})\left\{\frac{\widehat{\pi}(\bs X_{i})T_{i}}{\widehat{e}(\bs X_{i})}+\frac{(1-\widehat{\pi}(\bs X_{i}))(1-T_{i})}{1-\widehat{e}(\bs X_{i})}\right\} Y_{i}^{2} - \left(\widehat{\mathrm{Mean}}(\widehat{\pi})\right)^{2}}, \] where $\widehat{e}(\bs X)$ is the DNN estimate of the propensity score using the full sample. The weights $\widehat{w}_{i}(\widehat{\pi})$ are determined by the following optimization problem: \[
\] where $\widehat{\mathbb{E}}[Y\mid\bs X=\bs x,T=t]$ and $\widehat{\mathbb{E}}[Y^{2}\mid\bs X=\bs x,T=t]$ are the DNN estimates of $\mathbb{E}[Y\mid\bs X=\bs x,T=t]$ and $\mathbb{E}[Y^{2}\mid\bs X=\bs x,T=t]$, respectively.
The empirical results highlight the trade-off inherent in our method. The benchmark policy $\widehat{\pi}_{\mathrm{Lin}}$ yields a mean outcome of $16{,}201.57$, with a standard deviation of $16{,}763.98$. In contrast, the policy $\widehat{\pi}_{\mathrm{Nonlin}}$, estimated under our nonlinear welfare criterion, yields a mean outcome of $16{,}132.78$ and a standard deviation of $16{,}617.18$. Relative to the benchmark, $\widehat{\pi}_{\mathrm{Nonlin}}$ reduces the mean outcome by approximately $0.42\%$ and the standard deviation by about $0.88\%$, reflecting our goal of achieving lower outcome dispersion under the nonlinear welfare criterion.
This paper presents a data-driven policy-learning procedure that utilizes observational data to address a nonlinear welfare criterion within an infinite-dimensional policy space. The proposed learning procedure expands the existing literature on policy learning by moving from a linear (utilitarian) welfare criterion to a nonlinear one, transitioning from finite-dimensional to infinite-dimensional policy spaces, and shifting focus from a known propensity score to an unknown one. Additionally, we introduce a novel reweighting-based debiasing method, providing a valuable alternative to the current double debiasing approach. We applied this procedure to the JTPA study, where we found a balance between efficiency and equity.
However, a significant challenge remains in the computational aspect: determining the best policy $\widehat{\pi}_{\ell,I}$ is fundamentally a nonconvex optimization problem. Due to the nonlinearity of the welfare criterion and the structure of the policy space, multiple local optima may arise. Future research should aim to develop more efficient optimization techniques, such as tighter convex relaxations or advanced heuristic search algorithms, to tackle these computational challenges.