Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
100,584 characters · 24 sections · 121 citation commands
Policy Learning with $$-Expected Welfare
\and Yuan Qi \\ [email removed]} \and Gaoqian Xu \\ [email removed]}}
\onehalfspacing
Targeted/personalized policy rules assign treatments to individuals based on their observable characteristics. Learning treatment assignment policies that benefit the relevant population in a desirable way often require careful consideration. The fact that treatment effects tend to vary with individual observable characteristics prompts policy makers to design policies that determine treatment statuses based on individual characteristics. Examples include deciding which patients should receive medical treatment, assigning unemployed workers to training programs, and selecting which students to offer financial aid. Using experimental or observational data from a sample that represents the relevant population, the optimal utilitarian policy maximizes the sum of individual welfare in the sample. The empirical welfare maximization approach in kitagawa2018a provides a solution in this regard.
As noted in kitagawa2021equality, maximizing the utilitarian social welfare criterion overlooks distributional impacts. This motivates kitagawa2021equality to introduce an equality-minded rank-dependent social welfare function that places greater emphasis on individuals with lower-ranked outcomes. When the policy class is restricted due to considerations such as implementability, cost, and interpretability, maximizing the utilitarian social welfare may even hurt those who are disadvantaged in the population. For example, if welfare is measured as the (negative) mean blood sugar level of individuals at risk of diabetes and the treatment is a new medication, a utilitarian policy may prescribe the medication to most individuals because it can substantially benefit the low-risk individuals, who form the majority of the sample, but high-risk individuals who receive the medication may be hurt and end up in even worse situations. Similarly, if welfare is evaluated by the average post-training income, a utilitarian policy is more inclined to select individuals who are high school graduates and have experienced relatively short periods of unemployment to participate in the training program, while overlooking those with lower educational attainment or longer unemployment durations who might also benefit substantially from the training; see Section 5.1 in athey2021policy.
Taking a group-agnostic and risk-averse point of view, this paper proposes to learn an optimal policy that favors individuals on the lower tail of the outcome distribution. Specifically, for any $\alpha\in(0,1)$, we introduce the $\alpha$-expected welfare function as the expected outcome among the worst-affected $(\alpha\times100)\%$ of the population, i.e., a lower-tail conditional average. We study non-randomized binary policies which maximize the $\alpha$-expected welfare and refer to such policies as $\alpha$-expected welfare maximization ($\alpha$-EWM) policies. The choice of $\alpha$ is problem-specific and should be based on domain knowledge. A smaller $\alpha$ means that the policy is tailored for the more disadvantaged, whereas a larger $\alpha$ generates a policy that considers a broader less-advantaged subpopulation but those who are most disadvantaged receive less attention. From a philosophical standpoint, when $\alpha$ is small, our $\alpha$-EWM objective aligns with John Rawls' difference principle, which aims to maximize the welfare of the least-advantaged group to maintain social stability and fairness 41549aa6-42b1-36d1-ab6a-84fdd10b1f93. Indeed, the $\alpha$-expected welfare converges to the essential infimum of the outcome random variable as $\alpha$ approaches zero. We note that the definition of the $\alpha$-expected welfare function also applies to $\alpha = 1$, in which case it reduces to the utilitarian welfare underlying the empirical welfare maximization studied in kitagawa2018a and athey2021policy. We refer to such policies as $1$-EWM throughout the rest of this paper.
To further motivate our $\alpha$-EWM for $\alpha \in (0,1)$, we provide a simple numerical comparison with the $1$-EWM criterion from kitagawa2018a, the equality-minded welfare criterion from kitagawa2021equality, and quantile maximization from wang2018quantile. Section (ref) discusses the relationship between our $\alpha$-EWM and these criteria in more detail. We use a simple data generating process (DGP) similar to the motivating example in wang2018quantile:
where the covariate $X \sim \operatorname{Unif}[0,1] $, the binary treatment $A\sim\operatorname{Bernoulli}(0.5)$, and $\epsilon \sim N(0,1)$. We assume that the propensity score $e_o(\cdot)=0.5$ is known, and the policy class is defined as $\Pi_{\mathrm{c}} = \mathds{1} \{ X \leq c \}$ for the policy parameter $c \in [0,1]$.
We create a superpopulation of size one million. Since we can generate $Y_i$ for both $A_i=0$ and $A_i=1$, we have full knowledge of the true outcome distribution induced by any $c$. For comparison, we select values of $c$ that maximize the following: the $0.1$-expected welfare, the standard Gini social welfare, the $0.1$-outcome quantile, and the mean outcome. These correspond to the $0.1$-EWM, equality-minded, $0.1$-quantile-optimal, and $1$-EWM policies, respectively. Figure (ref) displays the probability densities of the post-treatment outcomes induced by these policies. Under this DGP, there is a gradual tightening of the post-treatment outcome distribution as we move from the $1$-EWM policy to the equality-minded policy, then to the $0.1$-quantile-optimal policy, and finally to the $0.1$-EWM policy. The 0.1-EWM policy produces the most concentrated outcome distribution, with the thinnest tails on both the left and right compared to the other policies. This suggests that the $0.1$-EWM policy not only mitigates the risk of extremely poor outcomes but also avoids disproportionately large gains, resulting in a more equitable distribution centered around the median.
This paper makes several contributions to the literature on policy learning. First, under the assumption of unconfoundedness,\footnote{The assumption of unconfoundedness is not essential and can be replaced with any assumption that identifies the conditional marginal distributions of the potential outcomes.} we show that the $\alpha$-expected welfare function is identified and propose a debiased estimator. Our debiased estimator utilizes cross-fitted nuisance estimators and the orthogonal moment function based on the dual form of the $\alpha$-expected welfare function. Optimizing the $\alpha$-expected welfare poses noticeable challenges compared with $1$-EWM. Adopting a group-agnostic perspective, the worst-off subpopulation being targeted changes dynamically with different policies. Consequently, estimating the $\alpha$-expected welfare requires the estimation of the $\alpha$-quantile of the welfare, which serves as a “cutoff" for computing the tail average (see (ref) for details).
Second, we establish theoretical guarantees of our $\alpha$-EWM for any $\alpha\in (0,1)$ by deriving asymptotic upper regret bounds with an explicit expression for the constant. This complements similar regret bounds for $1$-EWM in kitagawa2018a and athey2021policy.
Third, we develop asymptotically valid inference for the optimal $\alpha$-expected welfare. When the optimal policy is unique, Wald-type inference is asymptotically valid. When the optimal policy is not unique, we develop inference by applying the generalized delta method for Hadamard directionally differentiable functionals; see, e.g., belloni2017program, fang2019inference, hong2018numerical.
Fourth, we demonstrate that more comprehensive policy evaluations can be performed by consistently estimating the welfare of the worst-off $(\alpha\times100)\%$ of the population for any $\alpha\in(0,1)$ and policy. Put differently, even if a policy does not specifically target the worst-affected $(\alpha\times100)\%$, we can still assess its performance at $\alpha$ to gain insights into the associated trade-offs. We illustrate our $\alpha$-EWM method using experimental data from the National Job Training Partnership Act (JTPA) Study, as analyzed by bloom1997benefits. We find that targeting smaller subpopulations—such as the bottom 25% or 30% of the outcome distribution—leads to more robust welfare performance across a range of welfare objectives. In contrast, targeting broader groups (e.g., the bottom 80%) can result in substantial welfare losses for the bottom 25%, indicating that policies aimed at broader groups may come at the expense of welfare among the most disadvantaged.
Lastly, we conduct simulation studies based on synthetic JTPA data generated using Wasserstein Generative Adversarial Networks (WGANs) developed by athey2024using, to evaluate the performance of our estimator and compare policy outcomes. In the WGAN-JTPA setup, both the $0.25$-EWM and equality-minded policies enhance the welfare of lower-ranked individuals while reducing that of higher-ranked individuals relative to the $1$-EWM policy, with the $0.25$-EWM policy placing much greater emphasis on these adjustments. Additional simulation studies based on stylized DGPs from athey2021policy are provided in (ref). Across all simulation setups, the debiased estimator and Wald inference perform satisfactorily for all $\alpha$ values considered.
The rest of the paper is organized as follows. Section (ref) provides an overview of the related literature. Section (ref) introduces our model preliminaries, including the $\alpha$-expected welfare measure and its identification under the selection-on-observables assumption. We point out relations and differences between four welfare measures: the 1-expected welfare, equality-minded welfare, quantile welfare, and our $\alpha$-expected welfare. Section (ref) reviews the dual form of the $\alpha$-expected welfare function and presents its debiased estimator, and (ref) establishes an asymptotic upper regret bound for our debiased optimal policy. (ref) constructs asymptotically valid inference for the optimal $\alpha$-expected welfare. (ref) presents numerical results, including an empirical application based on experimental data from the JTPA Study and a simulation study using WGAN-generated JTPA data. Section (ref) concludes. Technical proofs are relegated to a series of appendices.
Our work builds on existing literature on policy learning from experimental and observational data, as well as statistical inference for the mean outcome under the optimal policy. In the following, we provide a brief discussion of related work.
\paragraph{Mean-optimal Policy Learning} Existing research on policy learning in economics and statistics has mainly focused on the mean-optimal policy under unconfoundedness qian2011performance,zhao2012estimating,zhang2012estimating,bhattacharya2012inferring,luedtke2016statistical,kallus2018balanced,luedtke2020performance,athey2021policy. Most work on policy learning focus on establishing theoretical guarantees by deriving regret bounds. The seminal paper by kitagawa2018a explores mean-optimal policy learning from experimental data in a nonparametric framework. When propensity scores are known and the policy class denoted as $\Pi$ has a finite VC dimension, they employ inverse propensity weighting to estimate the welfare function, achieving $n^{-1/2}$-rate regret bounds, where $n$ is the sample size. athey2021policy extend this setup to observational studies where propensity scores are unknown and the policy class $\Pi_n$ may vary with $n$. They estimate the objective function using doubly robust scores, a method that is shown to be efficient in the sense of newey1994asymptotic. The resulting policies achieve regret bounds of the order $\sqrt{\mathrm{VC}(\Pi_n)/n}$. Notably, their regret bound depends on the convergence rate of nuisance parameter estimation and the semiparametric efficient variance for evaluating an optimal policy. Finally, under mild conditions, luedtke2020performance show that the regret can decay faster than $n^{-1/2}$ for a fixed data distribution.
Several studies have examined statistical inference for the mean-optimal welfare associated with the first-best policies. For instance, luedtke2016statistical propose an online one-step estimator that is $\sqrt{n}$-consistent for the optimal value function, where the estimated policy and value function are recursively updated using new observations. Similarly, shi2020breaking conduct inference for the optimal welfare via subsample aggregating and cross-validation. In contrast, rai2018statistical study inference for the optimal mean welfare under a restricted policy class. The author utilizes bootstrap and numerical delta methods in e.g., fang2019inference and hong2018numerical, to approximate the estimator’s limiting distribution. We apply the same set of tools to develop inference for the optimal $\alpha$-expected welfare associated with a pre-specified policy class when the optimal policy may not be unique.
\paragraph{Fairness and Robustness of Policy Learning.} In many real-world scenarios, alternative objective functions beyond the mean outcome may be more appropriate. Some studies design objective functions with fairness considerations. Besides kitagawa2021equality and wang2018quantile, other studies focus on distributional robustness or external validity in decision-making by adopting robust objective functions cui2023individualized, qi2023robustness, adjaho2022externally, fan2023quantifying, lei2023policy. The optimal policy under a robust objective function can be interpreted as the policy that maximizes the “worst-case” scenario of individualized outcomes when the underlying distribution is perturbed within an uncertainty set. fang2023fairness,viviano2024fair, and kim2023fair propose to maximize the average welfare subject to some fairness constraints.
The paper most closely related to ours is qi2023robustness, which adopts the average value-at-risk (AVaR) welfare criterion to develop robust individualized decision rules. The AVaR criterion is the same as our $\alpha$-expected welfare criterion, and qi2023robustness is motivated by the distributional robust representation of AVaR, see (ref). Apart from differences in motivation, the main results in qi2023robustness and our paper also differ. First, qi2023robustness focus on experimental data with a known propensity score, allowing direct estimation of the objective function. Instead, we consider observational studies with unknown propensity scores and estimate our objective function using doubly robust scores and cross-fitting. Second, we consider a general policy class $\Pi_n$ with a VC-dimension $\mathrm{VC}(\Pi_n)$ that may be changing with $n$. In contrast, qi2023robustness consider a more restrictive policy class within a reproducing kernel Hilbert space, which excludes many machine learning algorithms, such as decision trees and neural networks, from being used to learn the optimal policy. Third, applied to the class of policies in qi2023robustness, our regret bound is sharper than theirs. Fourth, we develop inference for the optimal welfare in experimental and observational setups. Computationally, qi2023robustness propose a non-convex optimization algorithm based on a surrogate function that smooths the binary policy function for the use of difference-of-convex optimization, whereas our optimization is done by derivative-free methods.
We close this section by summarizing the notation used in this paper. We use $O, o, O_{P}, o_{P}, \asymp, \gtrsim, \lesssim$ in the following sense: $a_n=O\left(b_n\right)$ if $\left|a_n\right| \leq$ $C b_n$ for $n$ large enough; $a_n = o(b_n)$ if $a_n/b_n \rightarrow 0$; $X_n=O_{P}\left(b_n\right)$, if for any $\delta>0$, there exist $M, N>0$, such that $\mathbb{P}\left|\left|X_n\right| \geq\right.$ $\left.M b_n\right] \leq \delta$ for any $n>N ; X_n=o_{P}\left(b_n\right)$, if $\mathbb{P}\left[\left|X_n\right| \geq \epsilon b_n\right] \rightarrow 0$ for any $\epsilon>0 ; a_n \asymp b_n$ if there exist $k_1, k_2>0$ and $n_0$, such that for all $n>n_0, k_1 a_n \leq b_n \leq k_2 a_n$ if $\lim a_n / b_n=\infty$; $a_n \gtrsim b_n$ if $b_n=O\left(a_n\right) ; a_n \lesssim b_n$ if $a_n=O\left(b_n\right)$. Furthermore, we write $f(n) = \widetilde{O}(g(n))$ if there is a function $h$ that grows poly-logarithmically such that $f(n) \leq h(g(n))g(n)$. The notation $f(n) = \Omega(g(n))$ means that there is a universal constant $c_o >0$ such that $f(n) \geq c_og(n)$ uniformly in $n$. We use the shorthand $[n] = \{1, \ldots, n\}$, $a \vee b = \max\{a,b\}$ and $a \wedge b = \min\{a,b\}$. The abbreviation i.i.d. stands for {\it independent and identically distributed}. In the sequel, let $c_o$ denote a generic positive constant, whose value may vary from line to line.
Suppose that we have a random sample $\left(X_i, Y_i, A_i \right)_{i=1}^n$, where $X_i\in\mathcal{X} \subseteq \mathbb{R}^p$ denotes the observable characteristics of individual $i$ (continuous or discrete), $Y_i\in\mathcal{Y}\subseteq\mathbb{R}$ represents the outcome of individual $i$ (or utility / welfare), and $A_i\in\{0,1\}$ denotes the treatment status of individual $i$, for $i\in [n]$. Without loss of generality, larger values of $Y_i$ are assumed to be preferable. To simplify notation, we define $Z_i\vcentcolon=(X_i,Y_i,A_i)\in\mathcal{Z}$ and $\mathcal{Z}=\mathcal{X}\times \mathcal{Y}\times \{0,1\}$. Let $Y_{i}(0)$ and $Y_{i}(1)$ denote the potential outcomes that would have been observed if $A_i=0$ and $A_i=1$, respectively. Then $Y_i=A_iY_{i}(1)+(1-A_i)Y_{i}(0)$ is the realized outcome under the Stable Unit Treatment Value Assumption rubin1978bayesian, rubin1990comment.
Throughout the rest of this paper, we assume that $\mathbb{E}|Y_i(0)| < \infty$ and $ \mathbb{E}|Y_i(1)| < \infty$. We denote by $P$ the distribution of $Z_i \equiv (X_i, Y_i, A_i)$, and by $\mathbb{E}_P$ and $\mathrm{Var}_P$ the expectation and variance under $P$, respectively.
We study non-randomized binary policy/rule $\pi: \mathcal{X} \rightarrow \{0,1\}$. Let $\Pi_o$ denote the policy class that contains all Borel measurable functions from $\mathcal{X}$ to $\{0,1\}$. For any policy $\pi\in \Pi_o$, let $Y_i(\pi):=Y_i(\pi(X_i) )$, the outcome of individual $i$ when $\pi$ is implemented. Further, let $F_{\pi}(y)$, $y\in \mathcal {Y}$ denote the distribution function of $Y_i(\pi)$ and $F^{-1}_\pi(\alpha)=\inf \left\{y \in\mathbb{R}:F_\pi(y)\geq\alpha \right\}$ denote the quantile function of $Y_i(\pi)$.
As discussed by kitagawa2018a and athey2021policy, practitioners may adopt a pre-specified policy class $\Pi\subseteq \Pi_o$ that incorporates constraints relevant to the problem context, such as budgetary limitations, specific functional forms, fairness considerations, and other pertinent factors.
As discussed in (ref), $\lim_{\alpha \rightarrow 0}\mathbb{W}_\alpha(\pi)=\operatorname*{ess\,inf} Y_i(\pi) $ and $\mathbb{W}_1(\pi)=\mathbb{E} \left[Y_i(\pi) \right]$. Our welfare function $\mathbb{W}_\alpha(\pi)$ therefore flexibly interpolates between the expected welfare and infimum welfare of the target population by varying $\alpha\in (0,1]$, where $\alpha=1$ gives the expected welfare of the target population adopted in kitagawa2018a and athey2021policy.
To identify $\mathbb{W}_\alpha(\pi)$ as defined in (ref), we note that $$Y_i(\pi)=\pi(X_i)Y_i(1)+[1-\pi(X_i)]Y_i(0).$$ The conditional (given $X_i = x$) and unconditional distribution functions of $Y_i(\pi)$ are
where $F_1(y |x)$ and $F_0(y|x)$ are the conditional distribution functions of $Y_i(1)$ and $Y_i(0)$ given $X_i=x$, respectively.
(ref) and (ref) imply that $\mathbb{W}_\alpha(\pi)$ is a function of the policy $\pi(\cdot)$ and the conditional distribution functions $F_1( \cdot | \cdot)$ and $F_0( \cdot| \cdot)$. Consequently, for any $\pi \in \Pi_o$, $\mathbb{W}_\alpha(\pi)$ is identified as long as $F_1( \cdot| \cdot)$ and $F_0( \cdot| \cdot)$ are identified. Any assumption that ensures the identification of $F_{1}(\cdot|\cdot)$ and $F_{0}(\cdot|\cdot)$ is sufficient to identify $\mathbb{W}_\alpha(\pi)$. In the rest of this paper, we adopt the selection-on-observables assumption, which includes unconfoundedness and common support, as detailed in (ref).
(ref) states that the potential outcomes are independent of the treatments after conditioning on the observed covariates. Heuristically, it requires that all confounders that affect both treatments and potential outcomes simultaneously be observed. For identification, (ref) can be relaxed to the weaker condition that $ e_o(x) \in (0, 1)$ for all $x \in \mathcal{X}$, but the regret bounds and inference developed in later sections of this paper rely on it.
Under (ref), the distribution functions $F_{a}(\cdot|x)$ for all $x \in \mathcal{X}$ are point-identified: \[ F_{a}(y|x)\vcentcolon=\mathbb{P}\left[Y_i(a)\leq y|X_i=x\right]=\mathbb{P}\left[Y_i\leq y|X_i=x, A_i=a\right]. \] Consequently, $\mathbb{W}_\alpha(\pi)$ is identified for any $\pi \in \Pi_o$.
In this subsection, we compare our $\alpha$-expected welfare $\mathbb{W}_\alpha(\pi)$, defined for $\alpha \in (0,1)$, with three welfare functions commonly used in the literature: the expected welfare, the equality-minded welfare, and the quantile welfare functions.
$1$-EWM in kitagawa2018a and athey2021policy take the mean outcome \(\mathbb{E}[Y_i(\pi)]\), which equals $\mathbb{W}_1(\pi) $, as the population welfare function, assuming that the distribution of \(Y_i(\pi)\) in the target population is the same as that in the study population.
For $\alpha\in (0,1)$, our $\alpha$-expected welfare $\mathbb{W}_\alpha(\pi) $ represents a distributionally robust version of the $1$-expected welfare function. To see this, consider the uncertainty set centered at probability distribution $F_{\pi}$ of the outcome under policy $\pi: \mathcal{X} \rightarrow \{0,1\}$: \[
\] where $D_\infty ( Q\| F_\pi) = \mathrm{ess \ sup} \log \frac{d Q}{d F_\pi}$. From rockafellar2002deviation and duchi2023distributionally, it follows that
The uncertainty set $\mathcal{U}_{\alpha} (F_\pi) $ is the risk envelope capturing the distributional uncertainty of $Y_i(\pi)$ in the target population, comprising distributions with minority subpopulations of at least size $\alpha$. We can therefore interpret $\pi_\alpha^*$ as the distributionally robust policy that maximizes the average welfare under the worst-case perturbation of the study population in $\mathcal{U}_{\alpha} (F_\pi)$. As $\alpha$ decreases, the uncertainty set expands, making the $\alpha$-expected welfare function more robust to potential distributional shifts in $Y_i(\pi)$ within the target population.
Since the $1$-EWM may worsen inequality, kitagawa2021equality propose equality-minded policies by maximizing rank-dependent social welfare functions (SWFs), which assign greater weights to lower-ranked individuals. Given a decreasing function $\Lambda: [0,1] \rightarrow [0,1]$ with $\Lambda(0) = 1$ and $\Lambda(1) = 0$, the equality-minded welfare under policy $\pi$ is defined as
where $\omega(t) \coloneqq -\frac{\mathrm{d}}{\mathrm{d}t} \Lambda(t)$ is the associated weight function. When $\Lambda$ is strictly convex, the associated SWF, $W_\Lambda$, upholds the Pigou-Dalton Principle of Transfers, as rank-preserving transfers from higher-ranked individuals to lower-ranked individuals are preferred under the welfare $W_\Lambda$. The function $\Lambda$, chosen by practitioners, captures the degree of inequality aversion through its level of complexity. An important class of rank-dependent SWFs is the extended Gini SWFs, where $\Lambda (t) =\Lambda_k(t) = (1-t)^{k-1}$ for some $k \geq 2$, and the weight function is $\omega(t)=\omega_k(t) = (k-1)(1-t)^{k-2}$. The expected welfare and the standard Gini SWF correspond to $k = 2$ and $k = 3$, respectively.
Equality-minded SWFs can, in fact, be expressed in terms of our $\alpha$-expected welfare $\mathbb{W}_\alpha(\pi)$. For example, when $k > 2$, the extended Gini SWF can be written as a weighted average of $\mathbb{W}_\alpha(\pi)$:
Although our $\alpha$-expected welfare can be written as
where $\Lambda(t) = \left(1 - t/\alpha\right) \mathds{1}\{ 0 \leq t \leq \alpha \}$ and $\sigma(t) = \frac{1}{\alpha} \mathds{1} \{ 0 \leq t \leq \alpha \}$, it does not satisfy the Pigou-Dalton Principle of Transfers, as $\Lambda(t)$ is not strictly convex. This principle is satisfied only if the rank-preserving transfer happens across the probability level $\alpha$, i.e., from an individual ranked above $\alpha$ to an individual ranked below $\alpha$. Transfers on the same side do not affect $ \mathrm{W}_\alpha(\pi)$ since all the individuals involved have the same weight.
To prioritize the lower tail of population welfare over the (weighted) expected welfare, wang2018quantile propose a quantile-optimal policy, defined as \[ \operatorname{argmax}_{\pi \in \Pi} \mathrm{VaR}_\alpha(Y_i(\pi)) = F_\pi^{-1}(\alpha) , \] where $\alpha \in (0,1)$ is the quantile level of interest. For the class of linear policies with a fixed number of covariates $\Pi$, wang2018quantile establish the cube root asymptotics for the estimator of the parameter that defines the optimal linear policy.
Compared with quantile welfare $F_\pi^{-1}(\alpha)$ that overlooks the welfare of the population with outcomes below it, our $\alpha$-expected welfare function $ \mathbb{W}_\alpha(\pi)$ integrates $F_\pi^{-1}(t)$ over the range $[0, \alpha]$, thereby accounting for welfare levels below the $\alpha$-quantile and providing a more comprehensive assessment of the lower tail of the welfare distribution.
The $\alpha$-expected welfare function $\mathbb{W}_\alpha(\pi)$ has a convenient dual representation, which we will use to construct a debiased estimator of $\mathbb{W}_\alpha(\pi)$.
Let $(u)_-\vcentcolon=\min{(u, 0)}$ and $(u)_+ \vcentcolon= \max{(u, 0)} $. Further, let $\theta = (\pi , \eta)$ and $$ \mathbb{V}_\alpha(\theta) = \frac{1}{\alpha} \mathbb{E}\left[\left(Y_i(\pi)-\eta\right)_{-}\right]+\eta. $$
Let \[\mu_a(x, \eta)\vcentcolon= \mathbb{E}\left[\left(Y_i(a)-\eta \right)_- |X_i = x\right] \mbox{ for } a\in \{0,1\},\] and $\tau(x, \eta) \vcentcolon=\mu_1(x, \eta)- \mu_0(x, \eta)$ for any $x \in \mathcal{X}$ and $\eta \in \mathbb{R}$. Under (ref), $\tau(x, \eta)$ is identified for any given $\eta$.
As noted in the previous sections, the 1-expected welfare function $\mathbb{W}_1(\pi) $ is the same as the expected welfare \(\mathbb{E}[Y_i(\pi)]\) in kitagawa2018a and athey2021policy. In the rest of this paper, we focus on estimation and asymptotic theory for an $\alpha$-EWM rule when $\alpha\in (0,1)$.
Theorem (ref) suggests two plug-in methods for estimating $\mathbb{V}_\alpha(\theta)$ or the welfare function $\mathbb{W}_\alpha(\pi)$: IPW and outcome equation estimation. It is known that the IPW estimator is sensitive to the estimator of the propensity score and may suffer from severe bias. The outcome equation estimator may be sensitive to the estimators of $\tau$ ($\mu_1$ and $\mu_0$). This motivates the debiased estimator proposed in this section.
(ref) implies that under (ref), the function $\mathbb{V}_{\alpha}(\theta)$ is identified for any fixed $\theta = (\pi, \eta)$. Following robins1994estimation and robins1995analysis, we build our doubly robust score for $\mathbb{V}_\alpha(\theta)$ by introducing the augmentation term. Given any $\theta = (\eta, \pi)$, and for any function $\check{e}: \mathcal{X} \rightarrow (0,1)$ and $\check{\mu}_a :\mathcal{X} \times \mathbb{R} \rightarrow\mathbb{R}$ with $a \in \{0,1\}$, define
where $\check{\mu} = (\check{\mu}_0, \check{\mu}_1)$ and the augmentation term is defined as the sum of the last two components in (ref). The augmentation term has mean zero and the Neyman orthogonality condition holds: \[ \partial_{\mu} \mathbb{E}_{P} [ g_\theta(Z_i; \mu_o, e_o) ] [ \check{\mu} - \mu_o] = 0,\quad \text{ and }\quad \partial_{e} \mathbb{E}_{P} [ g_\theta(Z_i; \mu_o, e_o) ] [ \check{e}- e_o] = 0, \] where $\mu_o = (\mu_0, \mu_1)$. To simplify notation, we let $g_\theta(\cdot) = g_\theta(\cdot; \mu_o, e_o)$, where the function $g_\theta(\cdot)$ indexed by $\theta$ is referred to as the (doubly robust) score function for estimating $\mathbb{V}_\alpha(\theta)$. It is clear that for any given $\theta$, the function $g_{\theta} - \mathbb{E}_P[ g_{\theta} (Z_i)]$ is the efficient influence function for $\mathbb{V}_{\alpha} (\theta)$; see luedtke2016statistical,kennedy2016semiparametric for more detailed discussion.
Building upon chernozhukov2018double and chernozhukov2022locally, we construct our doubly robust score \( \widehat{g}_\theta(Z_i) \) for \( \mathbb{V}_{\alpha}(\theta) \) based on \( K \)-fold cross-fitting, a sample-splitting method used to validate asymptotic properties and leverage high-level conditions concerning the predictive accuracy of nuisance estimation methods.
We describe the estimation steps below, see Algorithm (ref) in (ref) for details.
In this section, we establish asymptotic regret bounds on the debiased $\alpha$-EWM policy proposed in Section (ref) for any fixed $\alpha\in(0,1)$. They complement similar regret bounds for the $1$-EWM and equality-minded policies established in kitagawa2018a, athey2021policy and kitagawa2021equality.
For each $n$, let $\Pi_n$ denote the class of candidate policies and $\Theta_n = \Pi_n \times \mathcal{B}_Y$, where $\mathcal{B}_Y \subset \mathbb{R}$ is a compact set introduced in (ref). For brevity, we write $\mathbb{V}(\theta) \equiv \mathbb{V}_\alpha(\theta)$, omitting the subscript $\alpha$.
The following assumption restricts the complexity of the policy class $\Pi_n$.
A policy is a classifier that assigns the covariate $X_i$ to a binary treatment status. Any machine learning classification model can serve as a candidate policy class. In the following, we list three examples of policy classes and their VC dimensions.
In this section, we establish a fundamental lemma showing that the estimation error of the nuisance parameters can be ignored when $\mathbb{V}(\cdot)$ is estimated using the doubly robust score with cross-fitting. Before presenting the lemma, we introduce additional assumptions.
Let $ \mathbb{V}_{n}(\theta)= \mathbb{P}_n g_\theta$, where $g_\theta (z):= g_\theta(z; \mu_o, e_o)$ is defined in (ref). We assume that the nuisance parameter estimators $\widehat{\mu}_a\left(\cdot, \cdot\right)$ and $\widehat e(\cdot)$ converge to their true values at sufficiently fast rates.
We conclude this subsection by demonstrating that $ \widehat{\mathbb{V}}_{n}(\theta)$ is a good approximation to $\mathbb{V}_{n}(\theta) = \mathbb{P}_n g_\theta$ with convergence rate faster than $n^{-1/2}$. Consequently, we can ignore the nuisance parameter estimation errors in subsequent asymptotic analysis.
In this subsection, we study the regret upper bound of implementing $\widehat{\pi}_{n}$ under the following assumption.
For any policy class $\Pi_n$, which may depend on $n$, the regret of deploying a policy $\pi \in \Pi_n$ relative to the best policy in $\Pi_n$, is defined as \[ \mathrm{Reg}(\pi, \Pi_n) = \max_{\pi^\prime \in \Pi_n} \mathbb{W}_\alpha(\pi^\prime) - \mathbb{W}_\alpha(\pi). \] When $\Pi_n$ is clearly understood from the context, we write $\mathrm{Reg}(\pi) =\mathrm{Reg}(\pi, \Pi_n)$ for notational simplicity. Our primary result regarding the asymptotic regret of our $\alpha$-EWM policy incorporates the following two key quantities: \[
\] where \[
\]
(ref) complements Theorem 1 in athey2021policy for $1$-EWM policy.\footnote{athey2021policy also allow for an approximate optimal policy.} The constant in (ref) depends on $\alpha$: it increases as $\alpha$ decreases, partly due to estimation error. Specifically, estimating the average welfare of the $\alpha$-worst-affected group makes use of only an $\alpha$-fraction of the total sample, leading to greater instability in welfare estimation.
The regret bounds in (ref) and those in kitagawa2018a, athey2021policy are all of order $\sqrt{\mathrm{VC}(\Pi_n) / n}$. In addition, (ref) and Theorem 1 in athey2021policy provide explicit expressions for the constants which require more delicate technical proofs than kitagawa2018a, kitagawa2021equality.
Following kitagawa2018a, kitagawa2021equality, the proof of the order of the regret bounds of $\widehat{\pi}_n$ relies on the lemma below.
(ref) implies that it is sufficient to study the concentration of the empirical process: $$\mathbb{V}_{n}(\theta) - \mathbb{V}(\theta) = (\mathbb{P}_n - P) g_\theta \ \mbox{ over } \ \theta \in \Theta_n. $$ In contrast to kitagawa2018a,kitagawa2021equality and athey2021policy, the score function for the $\alpha$-expected welfare $g_\theta$ is nonlinear in $\theta$ rendering the VC dimension of the function class $\mathcal{G}_{\Theta_n} :=\{g_\theta : \theta \in \Theta_n\}$ difficult to derive. Instead of exploiting the VC dimension of the corresponding function classes as in kitagawa2018a,kitagawa2021equality and athey2021policy, we directly upper bound the covering number of $\mathcal{G}_{\Theta_n}$ and then apply the classic empirical process maximal inequality, such as Theorem 2.14.1 in vaart2023empirical.
(ref) and (ref) ensure the existence of an envelope function that is bounded in $L^2(P)$. Applying Theorem 2.14.1 in vaart2023empirical and (ref), we conclude that there is a universal constant $c_o > 0$ not depending on $n$ such that
Compared with kitagawa2018a,kitagawa2021equality, one of the technical challenges addressed by athey2021policy on $1$-EWM policy lies in handling the doubly robust estimator of the welfare function. They show that as long as $\mathrm{VC}(\Pi_n)$ does not grow too rapidly with $n$, the use of cross-fitting and ML/nonparametric estimation of nuisance parameters results in a regret bound of the order $\sqrt{\mathrm{VC}(\Pi_n)/n}$. Building on kitagawa2018a,kitagawa2021equality and athey2021policy on $1$-EWM policy, we establish an upper bound for $\alpha$-EWM for any $\alpha\in (0,1)$ with an explicit expression for the constant $c_o$ in (ref). Similar to athey2021policy, we employ a classical chaining argument to derive an upper bound for the Rademacher complexity of the score function class. However, due to the nonlinearity of score function $g_\theta$ with respect to $\theta$, the slicing technique used in athey2021policy is difficult to implement. Instead, we introduce a new conditional semi-metric and apply the classical Dudley's chaining argument to directly bound the Rademacher complexity of $\mathcal{G}_{\Theta_n}$. We refer interested reader to (ref) for details.
In this section, we develop asymptotically valid inference for the optimal $\alpha$-expected welfare. Compared with regret bounds, inference on optimal welfare is lacking even for $1$-EWM except for the first-best policy; see luedtke2016statistical,luedtke2018parametric,shi2020breaking, and Appendix B in the supplemental material to kitagawa2018a.
We first impose conditions including the uniqueness of the optimal solution denoted as $\theta_o$ to ensure asymptotic normality of $\sup_{\theta \in \Theta}\widehat{\mathbb{V}}_{n}(\theta )$ based on which we construct Wald-type inference. We then summarize a general inference procedure that relaxes the uniqueness assumption. A detailed treatment of the general inference procedure is postponed to (ref).
For simplicity, we assume that the policy class does not change with the sample size $n$, i.e., $\Pi_n = \Pi$ for all $n$, and write $\Theta = \Pi \times \mathcal{B}_Y$. We define a metric space $(\Theta, \| \cdot \|)$, where \[ \left\| \theta_1 - \theta_2 \right\| \equiv |\eta_1- \eta_2|+ \|\pi_1 - \pi_2\|_{P, 2} = |\eta_1- \eta_2| + \sqrt{ \mathbb{E} |\pi_1(X_i) - \pi_2(X_i)|^2}. \] for any $\theta_1, \theta_2 \in \Theta$. This premise will be upheld throughout the subsequent analysis.
We establish asymptotic normality under two assumptions, the bounded support assumption and the uniqueness assumption.
(ref) is widely adopted in policy learning research, see, e.g., kitagawa2018a, kitagawa2021equality, rai2018statistical, kallus2018confounding, luedtke2016statistical, luedtke2020performance.\footnote{Although studies like athey2021policy do not adopt this assumption for regret bounds, it substantially simplifies the technical analysis for statistical inference. } (ref) implies that the feasible set $\mathcal{B}_Y$ of the dual reformulation of $\mathbb{W}_\alpha(\pi)$ can be restricted to $[-c_o, c_o]$ and the regression functions $|\mu_a(x, \eta)| \leq 2c_o$ for all $\eta \in \mathcal{B}_Y$ and $a \in \{0,1\}$. Moreover, the functions $g_\theta(\cdot)$ are also uniformly bounded, i.e., $\sup_{\theta \in \Theta}\| g_{\theta} \|_{\infty} < \infty$.
(ref) is a standard condition in extremum estimation. It ensures that $\theta_o\in \Theta$ is a unique and well-separated point of maximum of $\theta \mapsto \mathbb{V}(\theta)$. Lemma 14.4 in kosorok2008introduction gives some sufficient conditions for this assumption. If for all $\epsilon > 0$, $ \mathbb{W}(\pi_o) > \sup_{\pi: \| \pi - \pi_o \| > \epsilon } \mathbb{W}(\pi)$ and $Y_i(\pi)$ has positive density at $\mathrm{VaR}_\alpha(Y_i(\pi))$ for all $\pi \in \Pi$, then (ref) is satisfied. For policy learning, (ref) is strong, although it is adopted in wang2018quantile, Section 2.3 of kitagawa2018a, and Section 2.3 of luedtke2020performance.
To establish asymptotic normality of $\widehat{\mathbb{V} }_{n} (\widehat{\theta}_{n} )$, consider the following decomposition:
Note that the first term on the RHS of (ref) is $o_P(n^{-1/2})$ due to (ref).
In the rest of this section, we will show that
Consequently, \[
\] and asymptotic normality follows.
To show (i), we first prove $\|\widehat{\theta}_{n} - \theta_o\| = o_P(1)$ in (ref) below. Since $\mathcal{G}_\Theta \equiv \{g_\theta : \theta \in \Theta\}$ is $P$-Donsker by (ref), (i) follows.
To show (ii), we note that \[
\] where the inequality follows from $\widehat{\mathbb{V}}_{n} (\theta_o) - \widehat{\mathbb{V}}_{n} (\widehat{\theta} _{ n} ) \leq 0$. Similar to luedtke2020performance, one can show that under mild conditions including boundedness and uniqueness, asymptotic equicontinuity arguments ensure that $(\mathbb{P}_n - P) ( g_{\widehat{\theta}_n} - g_{\theta_o} ) = o_P(n^{-1/2})$ for any policy class $\Pi$ satisfying $\text{VC}(\Pi)<\infty$.
Summing up, we obtain asymptotic normality of $\widehat{\mathbb{V}}_{n}(\widehat{\theta}_{n} )$.
The next theorem presents a consistent estimator of the asymptotic variance $\sigma_o^2$.
(ref) develops uniform inference for the optimal welfare without (ref). It improves upon the inference proposed in Appendix B in the supplemental material to kitagawa2018a. We provide a summary of the procedures here and refer to interested reader to (ref) for technical details.
Define the supremum functional $\psi: \ell^{\infty} (\Theta) \rightarrow \mathbb{R}$ as $\psi: h \mapsto \sup_{\theta \in \Theta} h (\theta)$. Consider the multiplier bootstrap $ \widehat{\mathbb{G}}_{n}^*: \Theta \rightarrow \mathbb{R}$ defined as
where $\{\xi_i\}_{i=1}^n$ are i.i.d. random variables independent of $(Z_i)_{i=1}^n$, with $\mathbb{E}( \xi_i) = 0$, $\mathbb{E}(\xi_i^2) =1$ and $\mathbb{E}\left[ \exp |\xi_i| \right] < \infty$. For given $\epsilon_n = o(1)$ with $n^{1/2} \epsilon_n \rightarrow \infty$, let
For any $\gamma \in (0,1)$, let $c_{\gamma}$ denote the $\gamma$-empirical quantile of $\widehat{\psi}_n^{\prime} (\widehat{\mathbb{G}}_n^*)$ which can be obtained from a large number of bootstrap samples. The one-sided confidence interval at the desired level $\gamma$ is
with correct asymptotic coverage: \[ \lim_{n \rightarrow \infty} \inf_{P \in \mathcal{P}_n} \mathbb{P} \left[ \mathbb{V}_P(\theta_o) \geq \sup_{\theta \in \Theta} \widehat{\mathbb{V}}_n (\theta) - c_{1-\gamma} /\sqrt{n} \right] \geq 1- \gamma, \] where $\mathcal{P}_n$ is a collection of distributions satisfying some regularity conditions specified in (ref) in (ref). Define $q_{1-\gamma}$ as the $(1-\gamma)$-empirical quantile of $\left|\widehat{\psi}_n^{\prime} ( \widehat{\mathbb{G}}_n^* ) \right|$ for any $\gamma > 0$. The corresponding two-sided confidence interval is
which attains the correct asymptotic coverage for any fixed distribution $P \in \mathcal{P}_n$: \[ \liminf_{n \rightarrow \infty} \mathbb{P} \left[ \left| \sup_{\theta \in \Theta} \widehat{\mathbb{V}}_n (\theta) - \mathbb{V}(\theta_o) \right| \leq q_{1-\gamma} /\sqrt{n} \right] \geq 1-\gamma. \]
This section presents extensive numerical results on the finite sample performance of our debiased estimator and proposed inference using both real data and synthetic data.\footnote{ Data and codes for this section can be accessed at \href{https://github.com/yqi3/alpha-EWM}{https://github.com/yqi3/alpha-EWM}.}
kitagawa2018a apply $1$-EWM method to experimental data from the National Job Training Partnership Act (JTPA) Study. The study randomized whether applicants are eligible to receive training and job-search assistance provided by the JTPA. The pre-treatment covariates included in the data are years of education (edu) and pre-program earnings (prevearn) and the outcome variable is an applicant's earnings 30 months after the assignment (earnings). The sample size is 9,223 and the propensity score is known to be $2/3$. We adopt this data studied by kitagawa2018a and, similar to kitagawa2018a, we analyze welfare from an intent-to-treat standpoint, considering hypothetically making available the training program to eligible individuals, who may decline it. For detailed data description and evaluation of average program effects, we refer the reader to bloom1997benefits.
We consider three policy classes: simple (treat all or none) and linear with and without squared and cubic $edu$. More specifically, the two linear policy classes take the form
We investigate $\alpha\in\mathcal{A}\vcentcolon=\{0.25, 0.3, 0.4, 0.5, 0.8\}$. We recommend that researchers interested in the $\alpha=1$ case consider the $1$-EWM in kitagawa2018a directly. For each $\alpha \in \mathcal{A}$ and policy class, we estimate $\mu_a(x, \eta)=\mathbb{E}\left[\left(Y_i(a) - \eta\right)_- \mid X_i = x\right] \text{ for } a \in \{0,1\}$ and a given $\eta$, using random forests (RF) developed by athey2019generalized. We then apply simulated annealing (SA), proposed by kirkpatrick1983optimization, to select the combination of parameters that (approximately) maximizes the objective function.\footnote{We build RF using regression_forest() in R package grf and implement SA using optim_sa() in the R package optimization athey2019generalized, husmann2017r. We use default tuning parameters for RF. For SA, the specifications are more problem-specific. A good strategy is to plot the loss function and inspect if there is sufficient evidence of convergence.} SA is a derivative-free probabilistic optimization algorithm aiming at finding approximate solutions by iteratively exploring the solution space and gradually decreasing the probability of accepting worse solutions as the algorithm progresses.\footnote{geman1984stochastic prove convergence of \textit{generic} SA to a global optimum, provided that the probability of accepting worse solutions shrinks sufficiently slowly, and that all elements in the solution space are equally probable as the number of training epochs goes to infinity.}
Estimation and inference results for $\mathbb{W}_\alpha(\pi_o)$ are organized in Table (ref). The first two columns consist of the class of simple policies and serve as baselines for $\Pi_{\text{LES}}$ and $\Pi^3_{\text{LES}}$ in the third and fourth columns. Detailed expressions for the optimal policies can be found in (ref). The observed increase in $\widehat{\mathbb{W}}_\alpha(\widehat\pi_{n})$ across panels reflects that, as $\alpha$ grows, the lower-tail subpopulation expands to include relatively better outcomes. This raises the average and thus increases the $\alpha$-expected welfare. The percentage of treated individuals tends to increase with $\alpha$ as well. The \(95\%\) confidence intervals (CIs) constructed using normal inference in Algorithm (ref) are reported in the third row of each panel in Table (ref), and the 95% CIs from uniform inference obtained via multiplier bootstrap with \(\epsilon = n^{-1/4}\) and $B=100$ are presented in the last row of each panel. For each combination of $\alpha$ and policy class, the CI from uniform inference is wider than that from normal inference. While we cannot verify uniqueness, a simulation study calibrated to the JTPA sample in Section (ref) finds that the Wald-type CIs achieve approximately \(95\%\) coverage, offering supporting evidence for their validity in this application.
Examining the point estimates of welfare, we see that for all $\alpha\in\mathcal{A}$, a simple policy of treating all outperforms treating none. Moreover, relative to treating all, there is a considerable increase in the targeted welfare generated by the optimal policy of class $\Pi_{\mathrm{LES}}$. Linear policies with $edu^2$ and $edu^3$ only bring tiny welfare improvements. Figures (ref) and (ref) highlight the optimal treatment regions. Following kitagawa2018a, we bin the individuals by $(edu,prevearn)$, and the number of individuals with each combined characteristic is represented by the size of the corresponding dot.
Tables (ref) and (ref) examine welfare gains and losses as we switch between different targeting policies and estimate the resulting welfare of different targeted subpopulations. For example, the first row in Table (ref) shows the estimated welfare of the worst-off $25\%$ of the population when the optimal linear policies are targeting the worst-off $25\%$, $30\%$, $40\%$, $50\%$, and $80\%$, respectively. The diagonal entries (i.e., the row maximums) are highlighted as these optimal policies are targeting the actual subpopulations of interest. Tables (ref) and (ref) demonstrate a valuable strength of our method, as we are able to conduct rich policy evaluations by estimating the expected welfare at any $\alpha$ for any given policy. In other words, even when a policy is not targeting the worst-affected $(\alpha\times100)\%$, we can still evaluate its performance at $\alpha$ to obtain a clear picture of the trade-offs, which opens up possibilities for learning policies that promote greater equality across subpopulations.
From Table (ref) below and Table (ref) in (ref), adopting the linear policy that targets $\alpha'=0.8$ leads to an $11.9\%$ decrease in the average welfare of the worst-affected quarter of the population ($\alpha=0.25$), compared to implementing the optimal linear policy targeting the worst-affected quarter ($\alpha=\alpha'=0.25$). Conversely, adopting the policy targeting the worst-affected quarter ($\alpha'=0.25$) only leads to a $5.3\%$ decrease in the $0.8$-expected welfare ($\alpha=0.8$) relative to implementing the optimal policy targeting the worst-affected $80\%$ ($\alpha=\alpha'=0.8$). In Table (ref), similar patterns emerge with the inclusion of $edu^2$ and $edu^3$ in treatment assignment. Based on Tables (ref) and (ref), Tables (ref) and (ref) in (ref) report the percentage welfare loss for every combination of actual $\alpha$ and $\alpha'$ for policy selection. A notable observation is that the bottom quarter of the population is particularly vulnerable when the policy targets some $\alpha' \geq 0.4$ instead. Thus, policymakers aspiring for greater equality should prioritize smaller levels of $\alpha$, such as $0.25$ or $0.3$, as evidenced by the small percentage welfare losses in the first two columns of Tables (ref) and (ref), all of which are below $5.5\%$.
\vskip 0.2cm
We next present simulation results based on a superpopulation generated using Wasserstein Generative Adversarial Networks (WGANs) to evaluate the finite-sample performance of our debiased estimator. We focus on this simulation setup in the main text because the generated data more closely resembles real-world data distributions, making it more illustrative of practical applications. For comparison, we also conduct two additional simulation studies inspired by the DGPs in athey2021policy, with adjustments that make the treatment assignment exogenous. Since the results across all three designs are qualitatively similar—our estimator consistently exhibits decreasing mean squared error as the sample size increases, and the coverage rates approach the nominal 95% level in larger samples—we relegate the latter two studies to (ref).
In all three simulation setups, the propensity scores are assumed to be known, i.e., $\widehat{e}(\cdot) = e(\cdot)$. Cases with unknown propensity scores can be analyzed analogously using an estimator $\widehat{e}(\cdot)$ that satisfies Assumption (ref). Since uniform inference based on the multiplier bootstrap is computationally intensive, we report only the coverage rates based on confidence intervals constructed via Wald inference. We examine values of $\alpha \in \mathcal{A}$ considered in (ref).
We employ WGANs developed by athey2024using to construct a hypothetical superpopulation, referred to as WGAN-JTPA, consisting of one million observations based on the JTPA data in Section (ref). As mentioned by athey2024using, a benefit of using WGAN-generated data for simulations is that this practice largely rules out the possibility for researchers to choose particular DGPs that favor their proposed methods. This subsection demonstrates robust performance of our debiased estimator even when the underlying superpopulation is built from real datasets like the JTPA, which has highly skewed outcome and covariate distributions. (ref) discusses the training process in more detail and presents some summary statistics.
While technical details of WGANs can be found in athey2024using, we highlight that to build the superpopulation, since we generate $X|A$ followed by $Y|(X,A)$ and apply the same generator on $(X, 1-A)$ to obtain $Y|(X, 1-A)$, both potential outcomes are available for each individual. As a result, we can directly compute the true expected welfare at any $\alpha$ induced by any policy, which is simply a tail average of post-treatment outcomes. For each $\alpha\in\mathcal{A}$, we run SA to find a linear policy $\pi_o\in\Pi_{\mathrm{LES}}$ (as defined in (ref)) that maximizes $\mathbb{W}_\alpha(\pi)$ and treat the resulting optimum $\mathbb{W}_\alpha(\pi_o)$ as the population truth.
As an illustration, we use WGAN-JTPA to compare the $0.25$-EWM policy with the $1$-EWM (mean-optimal) and equality-minded (standard Gini social welfare-optimal) policies. Inspired by Figure 3 in kitagawa2021equality, Figure (ref) plots the between-quantile differences in post-treatment outcomes across these policies. The figure shows that both the $0.25$-EWM and equality-minded policies raise the welfare of lower-ranked individuals while lowering the welfare of higher-ranked individuals relative to the $1$-EWM policy at the population level, with the $0.25$-EWM policy placing much greater emphasis on these adjustments.
In the simulations, for each replicate, we draw a sample of size $n \in \{2000, 5000, 10000\}$ without replacement from WGAN-JTPA. The propensity score is fixed at the population mean of $A$, which is approximately $0.66475$.\footnote{This is very close to the mean of $A$ in the actual JTPA data, $0.66497$. In the JTPA Study, treatment was randomized with probability $2/3$, and we assume randomized treatment in WGAN-JTPA as well.} For each pair $(n, \alpha)$, we apply Algorithm (ref) to 1,000 sample draws and organize the results in Table (ref). As shown by the marginal histogram for earnings in Figure (ref) in (ref), WGAN-JTPA inherits the high skewness present in the original JTPA data. Consequently, larger sample sizes are required to achieve satisfactory coverage. From Table (ref), our optimal welfare estimator achieves acceptable coverage when $n = 5{,}000$, which is a realistic sample size for both experimental and observational studies (for reference, the original JTPA sample used by kitagawa2018a contains 9,223 observations).
The $\alpha$-expected welfare function considered in this paper offers a flexible interpolation between the Rawlsian welfare ($\alpha\rightarrow 0$) and the empirical welfare maximization ($\alpha=1$) approach proposed by kitagawa2018a. Like athey2021policy for the empirical welfare maximization, our development of the doubly robust scores facilitates asymptotic inference for the optimal welfare and allows practitioners flexibility in how they estimate the nuisance parameters. Besides learning the optimal policies, our estimation strategy also enables more thorough policy evaluations by computing the average welfare of the worst-affected subpopulation of any size (fraction of the population). In addition to establishing regret bounds for the debiased estimator, we also develop inference for the optimal $\alpha$-expected welfare for any $\alpha \in (0,1)$. Results from extensive numerical studies based on both JTPA data and simulated data demonstrate the efficacy and practical value of policy learning through $\alpha$-EWM.
We are currently working on several extensions of this paper. Methodologically, it is important to develop statistical tests to compare whether one policy is superior to another. Practically, it would be beneficial to determine who is actually targeted by the optimal policy. For example, what characteristics do the worst-affected individuals have? Information like this could present a more comprehensive picture of the relevant population and promote the design of more equitable policies.