EconBase
← Back to paper

Leave No One Undermined: Policy Targeting with Regret Aversion

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

79,320 characters

Leave No One Undermined: Policy Targeting with Regret Aversion



\sloppy

\newgeometry{verbose,tmargin=1in,bmargin=1in,lmargin=1.25in,rmargin=1.25in,footskip=1cm}

\title{Leave No One Undermined:  Policy Targeting with Regret Aversion{\thanks{We would like to thank Don Andrews, Isaiah Andrews, Tim Armstrong, Yong Cai, Kevin Chen,  Xiaohong Chen, Harold Chiang, Tim Christensen, Bruce Hansen, Marc Henry, Lihua Lei, Yuan Liao, Doug Miller, Francesca Molinari, Jos\'e Luis Montiel Olea, Mikkel Plagborg-M{\o}ller, Jack Porter, Sophie Sun, Chris Taber, Max Tabord-Meehan,   Aleksey Tetenov, Davide Viviano, Ed Vytlacil, Kohei Yata, and the participants at the Bravo/JEA/SNSF Workshop on
``Using Data to Make Decisions'' (Brown), Madison, Yale, Greater NY Metropolitan Area Econometrics Colloquium (Rochester), SEA 2025 (Tampa), and SETA 2025 (Macau) for helpful comments.
Chen Qiu gratefully acknowledges financial support from NSF grant (number SES-2315600).
}}}
\author{Toru Kitagawa\thanks{ Department of Economics, Brown University. Email: toru\[email removed]} \and Sokbae Lee\thanks{Department of Economics, Columbia University. Email: [email removed]} \and Chen Qiu\thanks{Department of Economics, Cornell University. Email: [email removed]}}
\date{April 2026}
\maketitle


\begin{abstract}
While the importance of personalized policymaking is widely recognized, fully personalized implementation remains rare in practice, often due to legal, fairness or cost concerns. We study the problem of policy targeting for a regret-averse planner when training data gives a rich set of observables while the assignment rules can only depend on its subset. Our regret-averse criterion reflects a planner's concern about regret inequality across the population. This, in general, leads to a fractional optimal rule due to treatment effect heterogeneity beyond the average treatment effects conditional on the subset of  observables. We propose a debiased empirical risk minimization approach to learn the optimal rule from data and establish favorable, new upper and lower bounds for the excess risk, indicating a convergence rate of $1/n$ and asymptotic efficiency in certain cases. We apply our approach to the National JTPA Study and the International Stroke Trial.
\end{abstract}

\newpage
\section{Introduction}

Personalization and policy targeting gained wide recognition in social and medical sciences. Yet, practice of complete personalization is rare. For example, as noted by \cite{manski2022patient}, President's Council of Advisors on Science and Technology (\citeyear{PCAST}) defines  ``personalized medicine'' as:

\emph{``...the tailoring of medical treatment to the specific characteristics of each patient. In an operational sense, however, personalized medicine does not literally mean the creation of drugs or medical devices that are unique to a patient...''}


Indeed, even though a rich set of observables $X$ are present in the data, policy makers (PM) often only form allocation rules based on a few covariates $W$, a subset of and not as rich as $X$. If treatment responses are significantly different in $X$ even after conditioning on $W$, how should the PM design its optimal treatment policy?


For instance, consider the mentoring program studied by \cite*{resnjanskij2024can} that aims to improve the labor market prospects for disadvantaged  adolescents in Germany. \cite{resnjanskij2024can} collected a rich set of covariate information on the participants of their randomized control trial. Their study highlights drastically different treatment responses depending on the social economic status (SES, classified based on answers to questions like how many books are at home, whether the adolescent is a first-generation migrant or has a single parent, etc.) of the adolescents: Low SES adolescents respond positively to the mentoring program while  higher SES adolescents respond negatively (see left panel of Figure \ref{fig:mentor}).

\begin{figure}[h!]
    \centering
    \includegraphics[width=0.46\linewidth]{figures/mentor.1.png}
    \includegraphics[width=0.46\linewidth]{figures/mentor.2.png}
    \caption{Left: Treatment effects taken from \citet[][Column 4, Table 2]{resnjanskij2024can}; Right: Our calculation of welfare regrets for three hypothetical decision rules (note for the rule that treats all, regret of the low SES group is zero).}
    \label{fig:mentor}
\end{figure}



Motivated by their heterogeneity analyses,  imagine a PM who decides what fraction of the population should be treated. However, suppose the PM cannot condition their decision on SES, because including it may be perceived as  illegal or unfair, or measuring it precisely at the decision making stage is too costly. This corresponds to a scenario in which $X$ refers to SES and there is no $W$. Since the full population average  treatment effect is positive, the common approach that aims to maximize the average welfare (e.g., \citealt{manski2004statistical,kitagawa2018should,AW20,MT17}) would inform the PM to treat all adolescents, even though higher SES adolescents lose from the program. Is the common approach still reasonable?  And if not, what alternative decision criterion can reflect the heterogeneous treatment responses along SES?


In this paper, we answer these questions by studying the problem of policy targeting for a regret averse PM. The PM aims to find an optimal allocation rule $\delta$ that maps subset characteristics $W$ to  $[0,1]$ indicating treatment fractions, with a loss function $L(\delta):=\mathbb{E}\left[Reg^{\alpha}(X,\delta)\right]$,
where $Reg(x,\delta)$ refers to the welfare regret of group $X=x$ when applied with $\delta$, i.e., the efficiency loss compared to its best achievable welfare,     $\alpha\geq1$ and $\mathbb{E}[\cdot]$ denotes expectation with respect to the marginal distribution of $X$. The case of $\alpha>1$ is what we call regret aversion, while $\alpha=1$ corresponds to the  common approach, i.e., minimizing mean welfare regret  (equivalent to  maximizing mean welfare).



We link the regret-averse loss to the PM's aversion to unequal distributions of regret among groups defined by covariates that are not allowed to use as input to the treatment rule.\footnote{\cite*{liang2021algorithm} provide microfoundation for preference types that yield a preference for fairness, which in turn can be interpreted as inequality aversion. \cite*{auerbach2024testing,liu2024inference} conduct statistical inference based on the theoretical characterization of the fairness-accuracy frontier in \cite{liang2021algorithm}.}  For example, viewing the index of labor market prospects as a proxy for welfare, the right panel of Figure \ref{fig:mentor} plots distributions of welfare regrets  for three hypothetical rules (treat all, treat 88\%, and treat 82\%).  Treating all benefits and induces no regret for low SES adolescents, but hurts and generates a high regret for higher SES adolescents.\footnote{As explained by \cite{resnjanskij2024can}, mentoring may actaully crowd out more useful inputs offered by higher SES families, such as parental attachment or participation in other useful activities.} The other two fractional rules enjoy a more equal distribution of regret between the two groups, although they also have higher mean regrets. If $\alpha = 1$, PM displays no aversion to inequality of regrets and evaluates rules solely based on the mean, indicating treating all is indeed optimal. However, if $\alpha>1$, the policymaker dislikes and penalizes rules with higher inequality. We formally quantify such aversion  via an Atkinson inequality index applied to regret (calculated in Table \ref{tab:mentor} according to \eqref{eq:inequality}).\footnote{\cite{atkinson1970measurement}'s measure of inequality is at an individual level rather than group level. We
may interpret $X$ in our setup as individuals in his framework.} The higher the value of  $\alpha$, the more the PM dislikes inequality. In fact, treating  88\% (82\%) of the population is optimal if $\alpha=2$ (3).  As one of the fundamental goals of sustainable development policy is to  ``reduce the inequalities (\ldots) that (\ldots) undermine the potential of individuals''\footnote{United Nations Sustainable Development Group, \url{https://unsdg.un.org/2030-agenda/universal-values/leave-no-one-behind}.}, and a wide literature in  program evaluation \citep*[e.g.,][]{angrist2012benefits,angrist2022marginal,gray2023long} documents significant subgroup treatment heterogeneity along gender, race and other costly or sensitive attributes, our approach develops a new perspective in the current policy learning literature.

\begin{table}[htbp]
\centering
\caption{Regret and inequality of various allocation rules in \cite{resnjanskij2024can}}
\begin{tabular}{lcccccc}
\toprule
& \multicolumn{3}{c}{Regret} & \multicolumn{3}{c}{Atkinson inequality index} \\
\cmidrule(lr){2-4} \cmidrule(lr){5-7}
Allocation rule & Low SES & Higher SES & Mean & $\alpha=1$ & $\alpha=2$ & $\alpha=3$ \\
\midrule
Treat all  & 0 & 0.22 & 0.12 & 0 & 0.36 & 0.51 \\
Treat 88\% & 0.08 & 0.19 & 0.14 & 0 & 0.08 & 0.14 \\
Treat 82\% & 0.12 & 0.18 & 0.15 & 0 & 0.02 & 0.05 \\
\bottomrule
\end{tabular}\label{tab:mentor}
\end{table}









In general, both $W$ and $X$ can be continuous and multidimensional. We show the key insights persist:  If $\alpha=1$, the optimal rule is to treat everyone or no one in the same $W$ group,  depending on the  sign of the corresponding CATE($W$), even if significantly  heterogeneous treatment responses remains  beyond $W$. Whenever $\alpha>1$,  the optimal rule is fractional, if the treatment effect heterogeneity is severe enough to alter the sign of  CATE($X$) within the same $W$.\footnote{We stress that the mechanism for fractional rules is different from that in \cite{kitagawa2022treatment}, in which fractional rules arise due to sampling uncertainty. Fractional rules also show up with  partially identified welfare with \citep{Manski2000,manski2005social,manski2007identification, Manski2007} or without \citep{ stoye2012minimax,yata2021,manski2022identification,montielolea2023decision, kitagawa2023treatment} true knowledge of the identified set, with nonlinear welfare \citep{manski2007admissible,manski20092009}, when the decision maker targets a functional of the outcome distribution that is not quasi-convex \citep*{kock2022functional,kock2023treatment}, or when agents respond with strategic behavior \citep{munro2020learning}.}  This insight carries over analogously even with a capacity constraint.

The preceding example  ignores the statistical uncertainty in the estimation of heterogeneous treatment effects. Focusing on $\alpha=2$ for tractability, we  propose an empirical squared regret minimization approach with debiasing and cross fitting to properly account for the estimation uncertainty from training data. Our procedure accommodates a wide range of black-box machine learning methods for estimating the heterogeneous treatment effects. We develop new asymptotic upper and lower bounds for the excess risk of our proposal.
In the case of a correctly specified linear sieve policy class, our procedure achieves a fast convergence rate of  $O(1/n)$  and is in fact asymptotically efficient.
In terms of computation, our squared regret approach is attractive due to the weighted least squares structure of the objective function. It will be straightforward to accommodate policy classes with convex constraints,  compared to the mean regret approach (which often relies on maximum score type optimization).


We further illustrate the value of our approach using two real datasets. For the National JTPA Study that measured the benefit and cost of employment and training programs,
we consider a PM  who designs treatment policies based on  pre-program years of education and earnings only, even though a wider range of characteristics like gender, race and marital status are available in the data. For the International Stroke Trial data that assessed
the effect of aspirin treatment for patients with acute ischemic stroke, we considered a hypothetical scenario  in which a doctor determines whether a patient should be treated with aspirin based on their age only, even though the training data contains many covariates of the patients, including their  demographic and medical history.
In both datasets, the estimated optimal policy fractions with our squared regret approach reveal considerable treatment effect heterogeneity  for some population subgroups, which can lead to significant regret inequality  should a singleton policy be applied instead.
These exercises suggest that our squared regret approach reveals  additional important information that may not be assessed with the mean regret paradigm alone.


The treatment choice literature has become an area of active research since the pioneering works of \cite{Manski2000, manski2002treatment,manski2004statistical} and \cite{Dehejia2005}. When the policy maker cannot differentiate individuals based on observable characteristics, finite and asymptotic results are developed by  \cite{schlag2006eleven}, \cite{stoye2009minimax},  \cite{HiranoPorter2009}, \cite{tetenov2012statistical}, \cite{masten2023minimax}, and \cite{chen2024note}.  \cite{manski2014quantile} and \cite*{guggenberger2024minimax}  in point-identified situations, and by  \cite{Manski2000,manski2005social,manski2007identification, Manski2007, manski20092009},  \cite{stoye2012minimax}, \cite{christensen2020}, \cite{ishihara2021}, \cite{yata2021}, \cite{manski2022identification}, \cite{ishihara2023bandwidth} and \cite{montielolea2023decision} in partially-identified settings. When the policy maker is able to condition on individual characteristics, studies on personalization and policy targeting include \cite{manski2004statistical}, \cite{BhattacharyaDupas2012}, \cite{kitagawa2018should, KT21}, \cite{MT17}, \cite{AW20}, \cite{sun2021empirical}, \cite{adjaho2022externally}, \cite{han2023optimal}, \cite{cui2023individualized},
\cite{terschuur2024locally}, \cite{viviano2024policy} and \cite{viviano2024fair}, among others, for analyses in different settings.

Our squared regret criterion coincides with a quadratic surrogate criterion for a cost-sensitive classification problem to predict the sign of CATE($X$) with cost being squared CATE($X$). See \cite{Zhang_2004}, \cite*{Bartlett_et_al_2006}, and references therein. Note, however, that our approach fundamentally differs from classification with the quadratic surrogate since our approach takes the squared regret as the ultimate objective function to minimize rather than a surrogate for the binary classification loss. As a result, our analysis can allow constrained $W$-individualized rules without raising the inconsistency issue of the constrained classification studied in \cite*{KST21}.

The rest of the paper is organized as follows: Section \ref{sec:setup} sets the stage, motivates our regret-averse loss function via inequality aversion and  studies the population optimal allocation rule. Section \ref{sec:proposal} presents our main proposal. Section \ref{sec:stat} develops the statistical performance guarantee. Section \ref{sec:capacity} discusses the case with capacity constraint. Empirical applications are in Section \ref{sec:emp.app}. Additional proofs, lemmas and technical results  are reserved in the Appendix and Online Supplement.








\section{Setup}\label{sec:setup}


Consider a policymaker who has access to a random sample of size $n$:
$\ensuremath{Z^{n}:=\left\{ Z_i\right\} {}_{i=1}^{n}\in\mathcal{Z}^{n}}$,
where $Z_{i}:=\{X_{i},D_{i},Y_{i}\}$, $X_{i}\in\mathcal{X}\subseteq\mathbb{R}^{d_X}$ is
the observed pre-treatment characteristics (covariates) of unit $i$,
e.g., their gender, race, pre-treatment education level, etc., $D_{i}\in\left\{ 0,1\right\} $
is the binary treatment indicator ($D_{i}=1$ means unit $i$ is under
treatment and $D_{i}=0$ means under control), and $Y_{i}\in\mathbb{R}$
is the observed outcome of interest of unit $i$, generated as
\begin{equation}\label{eq:observed.outcome}
Y_{i}=D_{i}Y_{i}(1)+(1-D_{i})Y_{i}(0),
\end{equation}
in which $Y_{i}(1),Y_i(0)\in\mathcal{Y}\subseteq\mathbb{R}$ are the
potential outcomes under treatment and control, respectively. Denote
by $P\in\mathcal{P}$ the joint distribution of $\left\{ X_{i},D_{i},Y_{i}(1),Y_{i}(0)\right\} $.
Then, the random sample $Z^{n}$ follows a joint distribution written as $P^{n}\in\mathcal{P}^n$, determined jointly by $P$, $n$ and \eqref{eq:observed.outcome}. In this paper, we maintain the following unconfoundeness and overlap assumptions:
\begin{assumption}\label{asm:unconfounded}
For each $i=1,...,n$, we have:
\begin{itemize}
\item[(i)] $Y_i(1),Y_i(0)\perp D_i\mid X_i$, i.e., $Y_i(1)$ and $Y_i(0)$ are independent
of $D_i$ conditional on $X_i$;

\item[(ii)] $\pi(x):=Pr\{D_i=1|X_i=x\}$ is bounded away from zero and one, i.e.,
$0<\underline{\pi}<\pi(x)<\bar{\pi}<1$ for some $\underline{\pi}$ and
$\bar{\pi}$, for all $x\in\mathcal{X}$.
\end{itemize}
\end{assumption}


The policymaker wishes to allocate a binary treatment $D\in\left\{ 0,1\right\} $
to a future population that shares the same marginal distribution
of $\left\{ X_{i},Y_{i}(1),Y_{i}(0)\right\} $ induced by $P$. For
each subpopulation group $X=x$ in which the covariate takes a specific
value $x\in\mathcal{X}$, write
\begin{align*}
\gamma_{1}(x)  :=\mathbb{E}[Y(1)\mid X=x],\quad
\gamma_{0}(x)  :=\mathbb{E}[Y(0)\mid X=x],
\end{align*}
where $\mathbb{E}[\cdotp|\cdotp]$ denotes the conditional expectation
under $P$. Write
$\tau(x):=\gamma_{1}(x)-\gamma_{0}(x)
$
as the conditional average treatment effect (CATE) for subgroup $X=x$. Under Assumption \ref{asm:unconfounded}, the CATE $\tau(x)$ is point identified for
all $x\in\mathcal{X}$.
We focus on the situation when --- at the decision-making
stage --- the set of covariates that the policymaker can actually
condition on, denoted by $W\in\mathcal{W}\subseteq\mathbb{R}^{d_W}$, is only a subset of and
not as rich as $X$, i.e., we may partition $X=\{W,X_1\}$ for some $X_1\in \mathbb{R}^{d_{X_1}}$.
\begin{example}[Legal or fairness concern] A randomized control trial may collect sensitive characteristics (e.g., gender and race), while  the policymaker cannot differentiate treatment decisions  based on them due to legal or fairness concerns. For example, many countries have anti-discrimination laws  that prohibit  treating an individual differently because of their membership to a protected class. Calls for the removal of race in many clinical diagnoses are also growing,  see, e.g., \cite*{briggs2022healing,manski2022patient,manski2023using} and the debates therein.
\end{example}




\begin{example}[Costly or manipulated variables at decision-making stage] In some scenarios, certain covariates are known to be important and are diligently recorded at the data-collecting stage. However, these variables could be costly to collect in practice and as a result, the decision maker does not observe these variables at the actual decision-making stage. For example, for patients in severe life-threatening
conditions such as sepsis, a physician must make a timely bedside intervention before lab measurements regarding key conditions of the patients can be returned \citep*{tan2022rise}. Moreover, some covariates may also be manipulated easily (e.g., they are not reported precisely) at the decision making stage, which makes them unsuitable to be included for treatment allocation.
\end{example}


\begin{example}[Single-index rules and subgroup analyses] Even in the absence of legal, fairness, or cost concerns, policy makers may prefer simple and interpretable rules. For example, the decision maker may determine treatment eligibility based on a single scalar variable $W:=\varphi(X)$, a function of  the whole observed covariate $X$.
See, e.g., \cite{kitagawa2018should,crippa2024regret} and references therein. Policy makers may also have particular pre-defined subgroups of interest for policy making, e.g., subgroups based on income or education brackets that are much coarser than observed income and education levels. These subgroups of interest may  be determined ex-ante in the pre-analysis plan prior to the data-collecting stage, or  determined ex-post by certain machine learning algorithms with data collected from earlier studies \citep{chernozhukov2018generic}.
\end{example}




\subsection{A regret averse planner's problem}

We start from the planner's problem without sample data $Z^{n}$.
Since the policymaker can only allocate policy decisions conditional
on $W$, we call their action plan a \emph{$W$-individualized decision
rule}, i.e., a mapping $
\delta:\mathcal{W}\rightarrow[0,1]$ from the support of $W$ to the unit interval. Here, $\delta(w)$
is the fraction of the subpopulation $W=w$ to be treated. For example,
$\delta(w)=0.5$ means half of the subpopulation
with $W=w$ will be treated, leaving the rest untreated. Although the policy
rule of the planner can only condition on $W$, treatment effect heterogeneity
may still vary at the more refined level $X$. For each group with $X=x$, let its corresponding  covariate $W$
take a value at $W=w$. Suppose that applying a generic $W$-individualized rule
$\delta$ to $X=x$ yields a linearly additive welfare for
the planner:
\begin{equation}
W(x,\delta,\gamma_{1},\gamma_{0}):=\gamma_{1}(x)\delta(w)+\gamma_{0}(x)(1-\delta(w)).\label{eq:welfare}
\end{equation}
Note the form of the welfare in (\ref{eq:welfare}) implies that the
optimal level of the welfare for $X=x$ is achieved by the
infeasible rule $\mathbf{1}\left\{ \tau(x)\geq0\right\} $. Then, for  $X=x$, define the regret of rule $\delta$ as the welfare
gap between $\delta$ and $\mathbf{1}\left\{ \tau(x)\geq0\right\}$:
\begin{align*}
Reg(x,\delta,\tau) & :=\max\left\{ \gamma_{1}(x),\gamma_{0}(x)\right\} -W(x,\delta,\gamma_{1},\gamma_{0})\\
 & =\tau(x)\left[\mathbf{1}\left\{ \tau(x)\geq0\right\} -\delta(w)\right].
\end{align*}
We consider a regret-averse policy maker who chooses an optimal $W$-individualized
policy rule by minimizing a nonlinear
transformation of regret, a notion advocated by \cite{kitagawa2022treatment}
and axiomatized by \cite{hayashi2008regret,stoye2011axioms}. Specifically, the policy maker
aggregates regrets among different subpopulation groups via the average
nonlinear regret loss:
\begin{align*}
L(\delta,\tau)  :=\int Reg^{\alpha}(x,\delta,\tau)dF_{X}(x) :=\mathbb{E}\left[Reg^{\alpha}(X,\delta,\tau)\right],
\end{align*}
where $\alpha\geq1$ is the degree of regret aversion, $F_{X}(\cdotp)$
denotes the marginal distribution of $X$ induced by the population
distribution $P$, and $\mathbb{E}[\cdotp]$ is the expectation operator
under $P$. A nonlinear regret optimal decision rule $\delta^{*}$
then solves
\begin{equation}\label{eq:population.optimal}
\min_{\delta:\mathcal{W}\rightarrow[0,1]}L(\delta,\tau).
\end{equation}
The rule $\delta^{*}$ characterizes an optimal action plan (conditional
on $W$) for the planner  if $P$ were known. Since $P$ in fact is unknown and needs to be learned from
data $\ensuremath{Z^{n}\in\mathcal{Z}^{n}}$,
the decision of the planner becomes statistical, i.e., selecting a
$W$-individualized \emph{statistical} decision rule:
\[
\hat{\delta}:\mathcal{Z}^{n}\times\mathcal{W}\rightarrow[0,1]
\]
that instructs an action for each subgroup $W=w$ given each possible
realization of data $Z^{n}=z^{n}\in\mathcal{Z}^{n}$. Denote by $\mathbb{E}_{P^{n}}[\cdotp]$
the expectation with respect to the randomness of
$Z^{n}\sim P^{n}$. The planner's ultimate goal is to find a ``good''
rule $\hat{\delta}$ from the training data so that
\begin{equation}\label{eq:goal.risk}
\sup_{P^{n}\in\mathcal{P}^{n}}\mathbb{E}_{P^{n}}\left[L(\hat{\delta},\tau)-L(\delta^{*},\tau)\right]
\end{equation}
is small and converges to zero at a fast rate (hopefully the
fastest) uniformly across a set of possible distributions $P^{n}\in\mathcal{P}^{n}$.

\subsection{Regret aversion as inequality aversion}\label{sec:inequalty.aversion}

We argue that our regret aversion loss
$L(\delta,\tau)$ has baked in an aversion to regret  inequality
in the population.\footnote{Regret measures the welfare loss of a group compared to what \emph{could have achieved} in terms of its best potential. Thus, our $L(\delta,\tau)$ reflects the preference of a planner who cares about to what extent personalized policies are equally fulfilling each sub-population's potential.}
More concretely, write $Reg_{\delta}(x):=Reg(x,\delta,\tau)$. Inspired
by the seminal work of \cite{atkinson1970measurement}, let
\begin{equation}
I_{\alpha}(Reg_{\delta}):=\frac{\left\{ \mathbb{E}[Reg_{\delta}(X)^{\alpha}]\right\} ^{1/\alpha}}{\mathbb{E}[Reg_{\delta}(X)]}-1\label{eq:inequality}
\end{equation}
be the Atkinson inequality measure of the regret
distribution $Reg_{\delta}(\cdotp)$ in the population induced by
rule $\delta$.\footnote{We may take $I_\alpha(Reg_\delta)=0$ if $\mathbb{E}[Reg_{\delta}(X)]=0$ as a convention.} Call $\left\{ \mathbb{E}[Reg_{\delta}(X)^{\alpha}]\right\} ^{1/\alpha}$
as the \textit{equally distributed equivalent} level of regret ---
the level of regret, if equally distributed to each subgroup in $X$, would yield
the same level of the loss as the actual distribution $Reg_{\delta}(\cdotp)$.
As $\alpha\geq1$, $I_{\alpha}(Reg_{\delta})\geq0$. The index $I_{\alpha}(Reg_{\delta})$
then measures how much larger the \textit{equally
distributed equivalent} is compared to the actual mean of regret $\mathbb{E}[Reg_{\delta}(X)]$.
A larger value of $I_{\alpha}$ indicates a higher degree of regret inequality. In this context, the regret aversion coefficient
$\alpha$ can be alternatively interpreted as a degree of inequality
aversion of the policymaker. A larger value of $\alpha$ means the
planner is more averse to regret inequality  among the population.
From (\ref{eq:inequality}), we may rewrite our nonlinear regret loss as
\begin{align*}
\mathbb{E}[Reg_{\delta}(X)^{\alpha}] & =\left[\underset{\text{mean regret}}{\underbrace{\mathbb{E}[Reg_{\delta}(X)]}}\left(1+\underset{\text{penalty for regret inequality}}{\underbrace{I_{\alpha}(Reg_{\delta})}}\right)\right]^{\alpha}.
\end{align*}
The mean regret paradigm corresponds
$\alpha=1$: $I_{1}(Reg_{\delta})=0$ for all distributions
of regret, meaning the policy maker displays no aversion to regret inequality
and ranks each distribution of regret only by their mean. If $\alpha>1$,
the planner is averse to regret inequality and penalizes rules that
lead to large inequality among the population.
\begin{figure}
\caption{Equally distributed equivalent of regret in an illustrative example}
\begin{center}
\includegraphics[scale=0.2]{figures/Figure1.png}\label{fig:atkinson.inequality}
\end{center}
{\raggedright\footnotesize \textit{Notes}: The distribution of the regret for rule $\delta=1$ is represented at point $A$. Its equally distributed equivalent $\zeta$ can be found as the $x$-coordinate of point $B$, where the dotted $45^\circ$   line intersects the curved isoquant that shows the same level of the loss as point $A$. The actual mean of
the regret corresponds to the $x$-coordinate of point $C$, where the dotted line intersects the solid black line perpendicular to it through point $A$. Then, the inequality
index of $\delta=1$ is  $\frac{\zeta}{\mu}-1$.}

\end{figure}

\begin{example}\label{ex:simple}
Suppose $X=\left\{ b,r\right\} $ is a binary group identity (blue
or red) with equal shares in the population, with $CATE(b)=\tau_{b}>0$ and $CATE(r)=\tau_{r}<0$ ($0<\left|\tau_{r}\right|<\tau_{b}$).
The policymaker cannot differentiate the two groups and can
only make a single treatment decision $\delta\in[0,1]$ to be applied to
the whole population. See Figure \ref{fig:atkinson.inequality} for an illustration of
the equally distributed equivalent of the regret for rule $\delta=1$ (which benefits  group $b$ but hurts group $r$).

\end{example}

\subsection{Population analysis}\label{sec:population}
We say a $W$-individualized
rule $\delta$ is a singleton rule if for almost all $w\in\mathcal{W}$,
it holds $\delta(w)\in\left\{ 0,1\right\} $; otherwise, we say $\delta$
is fractional.
With a slight modification of notation, we write
$\tau(w):=\mathbb{E}[Y(1)-Y(0)\mid W=w]$.
\begin{prop}
\label{prop:population}Consider the population optimal rule $\delta^{*}$
that solves (\ref{eq:population.optimal}).
\begin{itemize}
    \item[(i)] If $\alpha>1$, $\delta^{*}$ satisfies
\[
\mathbb{E}\left[\tau(X)\left\{ \tau(X)\left[\mathbf{1}\left\{ \tau(X)\geq0\right\} -\delta^{*}(W)\right]\right\} ^{\alpha-1}\mid W=w\right]=0,
\]
for almost all $w\in\mathcal{W}$, and is fractional unless
\[\min\left\{ Pr\{\tau(X)>0|W=w\},Pr\{\tau(X)<0|W=w\}\right\} =0\]
for almost all $w\in\mathcal{W}$;
 \item[(ii)]
If $\alpha=1$,  then $\delta^{*}(w)=1$ if $\tau(w)>0$, $\delta^{*}(w)=0$ if $\tau(w)<0$, and $\delta^{*}(w)\in[0,1]$ if $\tau(w)=0$,
for almost all \textup{$w\in\mathcal{W}$.}
\end{itemize}
\end{prop}


Proposition \ref{prop:population} shows that a regret-averse planner
concerned about regret inequality will
often prefer a fractional $W$-individualized rule.  They would prefer a singleton
rule if, for almost all $w\in\mathcal{W}$, CATE($x$) shares the same sign for all $x$ with  the same value of $w$, which also nests the case when $W=X$.
Our results offer a novel justification of implementing fractional
rules at the population level: Treatment effect heterogeneity at the $X$ level plus a concern for regret inequality induces a planner
to ``diversify'' their treatment allocation. We illustrate the optimal rule of Example \ref{ex:simple} in Figure \ref{fig:ex.2}.
\begin{figure}[h!]
\caption{Optimal treatment rule in Example \ref{ex:simple}}
\begin{center}
\includegraphics[scale=0.27]{figures/Figure2_updated.pdf}\label{fig:ex.2}
\end{center}

{\raggedright\footnotesize
\textit{Notes}: Line AE, viewed as a budget line, collects all the feasible allocations
of regrets between groups $b$ and $r$ with a decision $\delta\in[0,1]$.
Point $A$ represents $\delta=1$, under which the regret of
group $b$ is zero while the regret for group $r$ is $-\tau_{r}$.
Point $E$ is the regret distribution for rule $\delta=0$, in which the regret for group $r$ is zero but now group $b$ incurs
a regret of $\tau_{b}$. Every interior point of Line AE represents
a regret allocation of a fractional rule between the two groups. Since
$-\tau_{r}<\tau_{b}$, line AE tilts more heavily toward the horizontal
line. The green curves are the isoquants, each of them showing all
the allocations giving the same level of the loss function. The planner's
goal is to search for a point on Line AE that yields the smallest
loss. As long as $\alpha>1$, the isoquants will be strictly concave,
yielding an interior solution, i.e., point F in Figure \ref{fig:ex.2} (left), and
its equally distributed equivalent can be found as via point $G$. However, if $\alpha=1$,
the isoquants become linear, yielding a corner solution, i.e., point
A in Figure \ref{fig:ex.2} (right).}
\end{figure}


When $\alpha=2$, $L(\delta,\tau)$ becomes the average squared
regret, and the planner's problem becomes a weighted least squares
problem:
\[
\min_{\delta:\mathcal{W}\rightarrow[0,1]}\mathbb{E}\left\{ \tau^{2}(X)\left[\mathbf{1}\left\{ \tau(X)\geq0\right\} -\delta(W)\right]^{2}\right\} .
\]
Moreover, the associated
optimal rule also has an explicit form
\begin{equation}\label{eq:squared.regret.optimal}
\delta^{*}(w)=\frac{E\left[\tau^{2}(X)\mathbf{1}\{\tau(X)\geq0\}\mid W=w\right]}{E\left[\tau^{2}(X)\mid W=w\right]},
\end{equation}
whenever $\mathbb{E}\left[\tau^{2}(X)\mid W=w\right]\neq0$.\footnote{If $\mathbb{E}\left[\tau^{2}(X)\mid W=w\right]=0$ for some $w\in\mathbb{R}^{d_W}$, then $\delta^*(w)\in[0,1]$, i.e., any action in $[0,1]$ is optimal for those $w$ values. Our theory accommodates this ``non-uniqueness'' situation to some extent. See Assumptions \ref{asm:reg} and \ref{asm:margin} and discussions therein. }
\eqref{eq:squared.regret.optimal} shows that the fractional nature of the optimal rule $\delta^*$ depends not only on the sign and magnitude of the average treatment effect $\tau(x)$, but also the conditional distribution of $X$ given $W$. The optimal treatment assignment would be more fractional (i.e., closer to 0.5) if the values of $\int_{\tau(x)>0}\tau^{2}(x)dF_{X|W}(x\mid w)$ and $\int_{\tau(x)<0}\tau^{2}(x)dF_{X|W}(x\mid w)$ are closer to each other.

\begin{rem}
\cite{atkinson1970measurement}'s original proposal concerns inequality of \emph{welfare levels}. In our context, that corresponds to picking a concave transformation $U(\cdotp)$
and solving:
\begin{equation}\label{eq:concave.welfare}
\max_{\delta:\mathcal{W}\rightarrow[0,1]}\int U\left[W(x,\delta,\gamma_{1},\gamma_{0})\right]dF_{X}(x).
\end{equation}
We adapt
\cite{atkinson1970measurement}'s framework to focus on  inequality of \emph{regret}. It would be easy to construct examples in which low inequality of welfare levels imply high inequality of regret, and vice versa.  Given regret measures how far away each group is compared to their optimal welfare level, we think our approach could be  suitable when the policy maker mainly cares about   supporting each group to their full potential and not considerably hurting any group.
\end{rem}

\begin{rem}
One possibility of incorporating regret aversion is through an alternative loss function $\tilde{L}(\delta, \tau):=\{\mathbb{E}\left[Reg(X,\delta,\tau)\right]\}^{\alpha}$ for the same $\alpha\geq 1$. That is, the planner first aggregates subgroup level regret $Reg(X,\delta,\tau)$ and then takes a nonlinear transformation of the aggregate average regret ${\mathbb{E}\left[Reg(X,\delta,\tau)\right]}$.\footnote{C.f. \cite{manski2016sufficient} for discussions on achieving an optimality criterion for each observed covariate group or within the overall population only.} However, in this case, minimizing $\tilde{L}$ for any $\alpha\geq1$  is the same as minimizing  $\mathbb{E}\left[Reg(X,\delta,\tau)\right]$, i.e., the case of $\alpha=1$ for our loss $L$. Such a planner also does not display aversion to regret inequality and only ranks rules according to their group-wise average regret.  One may also twist our loss function by redefining the regret for each subgroup $X=x$ as:
\begin{align*}
\widetilde{Reg}(x,\delta,\tau):= \tau(w)\left[\mathbf{1}\left\{ \tau(w)\geq0\right\} -\delta(w)\right],
\end{align*}
leading to an alternative loss
$\widetilde{L}(\delta,\tau):=\mathbb{E}\left[\widetilde{Reg}^{\alpha}(X,\delta,\tau)\right]$.
That is, the regret of each subgroup $X=x$ is evaluated according to the welfare gap of its corresponding coarser $W=w$ group compared with the best welfare for the same coarser $W=w$ group. However, a planner with a loss $\widetilde{L}(\delta,\tau)$ is not concerned about regret inequality within the $W$ group.\footnote{For instance, in Example \ref{ex:simple}, as  $0<-\tau_r<\tau_b$, the overall average treatment effect is positive. Hence, according to $\widetilde{L}$, rule $\delta=1$ would yield a regret of $0$ for both $r$ and $b$ groups --- meaning there is no inequality between the two groups. However, we know $\delta=1$ actually hurts group $r$ dramatically as $CATE(r)<0$.}
\end{rem}

\begin{rem}
If $\alpha>1$ but the action space of the planner is restricted, i.e, $\delta(w)\in\left\{ 1,0\right\}$  for all $w\in\mathcal{W}$, the optimal rule $\delta^{*}$ would still be different from that of $\alpha=1$.\footnote{Moreover, to what extent one shall view the action space  as ``restricted'' is subject to the interpretation of the researcher. Even when a planner cannot take a fractional treatment allocation \emph{per se}, considering an extended action space $[0,1]$ is still valuable, as $\delta(w)\in[0,1]$ may be interpreted as a probabilistic recommendation, instead of an actual allocation of treatment. } For example, when $\alpha=2$, the average squared regret for action $1$, conditioning on  $W=w$,  is $
\mathbb{E}\left[\tau^{2}(X)\left(\mathbf{1}\left\{ \tau(X)\geq0\right\} -1\right)^{2}\mid W=w\right]$,
while that of  action $0$ is
$\mathbb{E}\left[\tau^{2}(X)\left(\mathbf{1}\left\{ \tau(X)\geq0\right\} \right)^{2}\mid W=w\right]$.
Thus, the optimal restricted $W$-individualized rule is
$
\delta^{*}_{\text{restricted}}(w)=\mathbf{1}\left\{ \mathbb{E}\left[\tau^{2}(X)\text{sgn}(\tau(X))\mid W=w\right]\geq0\right\}$.
\end{rem}





\section{Main proposal}\label{sec:proposal}


From now on and for the rest of
the paper, we focus on $\alpha=2$ for tractability. Other
values of $\alpha>1$ may  be analyzed analogously  with
more technicalities. Our
 proposal for learning a good rule $\hat{\delta}$
from training data involves two steps. The first step is
the efficient estimation of the loss function for each fixed $\delta$, i.e.,
\begin{equation}
L(\delta,\tau)=\mathbb{E}\left\{ \tau^{2}(X)\left[\mathbf{1}\left\{ \tau(X)\geq0\right\} -\delta(W)\right]^{2}\right\}.\label{eq:loss.proposal}
\end{equation}
Once
$L(\delta,\tau)$ is efficiently estimated from data,
the second step is to minimize the estimated loss
among a class of $W-$individualized rules that must be specified
by the researcher.
\begin{figure}[http]
\centering
\begin{minipage}{0.43\textwidth}
\centering
\includegraphics[width=0.9\textwidth]{figures/plot-1.pdf}
\text{$f(t)=t^{2}\left(\mathbf{1}\left\{ t\geq0\right\} -\delta\right)^{2}$}
\end{minipage}
\begin{minipage}{0.43\textwidth}
\centering
\includegraphics[width=0.9\textwidth]{figures/plot-2.pdf}
        \text{$f^{(1)}(t)=2t\left(\mathbf{1}\left\{ t\geq0\right\} -\delta\right)^{2}$}
    \end{minipage}
    \caption{Average squared regret functional is continuously differentiable}\label{fig:diff}
    \label{fig:enter-label}
\end{figure}

To allow for a wide range of ML algorithms in the first step and to potentially improve the statistical qualities of the second step, we consider
``debiasing'' (in a sense we make precise below) the loss function together with cross-fitting.
\footnote{See Remark \ref{rem:plug.in} for discussions on the connections with more direct ``plug-in'' approaches that do not involve debiasing.} Note although the nuisance function $\tau$ appears inside the indicator function, the loss $L(\delta, \tau)$ is still continuously differentiable in $\tau$ (though not twice continuously differentiable). This is due to the squaring of the term $[\mathbf{1}\{\tau(X) \geq 0\} - \delta(W)]$, which smooths out the discontinuity (See Figure \ref{fig:diff}).  Furthermore, we may view $L(\delta,\tau)$ as a finite dimensional parameter $\theta_0\in\mathbb{R}$
in a moment condition
\begin{equation}\label{eq:moment.loss}
\mathbb{E}\left[m(Z,\theta_{0},\tau)\right]=0,\quad \text{where } m(Z,\theta,\tau):=\tau(X)^{2}\left(\mathbf{1}\{\tau(X)\geq0\}-\delta(W)\right)^{2}-\theta.
\end{equation}
Therefore, following \citet[][Proposition 4]{newey1994asymptotic}, we can still derive the efficient influence function for any regular and asymptotically linear estimator of $L(\delta,\tau)$.  Let $\eta_{0}:=(\gamma_{1},\gamma_{0},\omega_{1},\omega_{0})$, where
 $\omega_{1}(x):=\frac{2\left(\gamma_{1}(x)-\gamma_{0}(x)\right)}{\pi(x)}$,
$\omega_{0}(x):=\frac{2\left(\gamma_{1}(x)-\gamma_{0}(x)\right)}{1-\pi(x)}$. Write
\[\xi(Z,\eta_{0}):=\left[\gamma_{1}(X)-\gamma_{0}(X)\right]^{2}+\left[D\omega_{1}(X)(Y-\gamma_{1}(X))-\left(1-D\right)\omega_{0}(X)(Y-\gamma_{0}(X))\right].\]


\begin{prop}\label{prop:debias}

Suppose Assumption \ref{asm:unconfounded} holds. For each $\delta$, the efficient influence function for any regular and asymptotically linear estimator of $\theta_0:=L(\delta,\tau)$ is
\[
\psi(Z)=\xi(Z,\eta_{0})\left(\mathbf{1}\{\gamma_{1}(X)-\gamma_{0}(X)\geq0\}-\delta(W)\right)^{2}-\theta_0
\]
\end{prop}
As $Var(\psi(Z))$ defines the semiparametric efficiency bound for estimating $L(\tau,\delta)$, we think it makes sense to exploit the structure of $\psi(Z)$ and define our modified loss function as:\footnote{Moreover, with the margin condition in Assumption \ref{asm:margin} and other regularity conditions, $L^o(\delta,\eta_0)$ can also be verified to  satisfy the Neyman orthogonal condition (c.f., \citealt{chernozhukov2018double} and references therein).}
\[
L^{o}(\delta,\eta_{0}):=\mathbb{E}\left[\xi(Z,\eta_{0})\left(\mathbf{1}\{\gamma_{1}(X)-\gamma_{0}(X)\geq0\}-\delta(W)\right)^{2}\right].
\]

Our theory does not restrict how the additional  nuisance  functions, $\omega_{1}$ and $\omega_{0}$, should be estimated. However, we note
they feature the following \emph{balancing} property (see, e.g., \citealt*{hainmueller2012entropy,zubizarreta2015stable,athey2018approximate}): for  all $g(\cdotp)$ such that $\mathbb{E}[g^{2}(X)]<\infty$,
\begin{align}
\mathbb{E}\left[D\omega_{1}(X)g(X)\right] & =\mathbb{E}\left[2\left(\gamma_{1}(X)-\gamma_{0}(X)\right)g(X)\right]=\mathbb{E}\left[(1-D)\omega_{0}(X)g(X)\right],\label{eq:balancing}
\end{align}
which may facilitate their estimation without the need to calculate propensity score. For example, to construct an estimator for $\omega_{1}$, denote
by $b(x):=(b_{1}(x),\ldots,b_{\text{dim}(b)}(x))^{\prime}$ a vector
of $\text{dim}(b)$ basis functions. Let  $\left\Vert \cdot\right\Vert$ be the vector $l_{2}$ norm. Note \eqref{eq:balancing} implies
\begin{equation}\label{eq:practical.omega}
\mathbb{E}\left[D\omega_{1}(X)b(X)\right]=\mathbb{E}\left[2\left(\gamma_{1}(X)-\gamma_{0}(X)\right)b(X)\right],
\end{equation}
suggesting we may estimate $\omega_1$  by solving the following minimum distance estimator with a Tikhonov penalty (c.f., \citealt{chen2012estimation,qiu2022approximate}):
\begin{equation}\label{eq:minimum.distance.penalty}
\min_{\omega_{1}\in\Theta_{n}}\left\Vert \frac{1}{n}\sum_{i=1}^{n}\left[2\left(\hat{\gamma}_{1}(X_i)-\hat{\gamma}_{0}(X_i)\right)b(X_i)-D_i\omega_{1}(X_i)b(X_i)\right]\right\Vert^2 +\lambda_{1,n}\frac{1}{n}\sum_{i=1}^{n}\left[D_i\omega_{1}^{2}(X_i)\right],
\end{equation}
where $\hat{\gamma}_{1},\hat{\gamma}_{0}$ are estimated versions
of $\gamma_{1}$ and $\gamma_{0}$,
\[
\Theta_{n}=\left\{ f:\mathcal{X}\rightarrow\mathbb{R}\mid f(x)=a^{\prime}b(x),a\in\mathbb{R}^{\text{dim}(b)}\right\} ,
\]
and $\lambda_{1,n}\geq0$ is a tuning parameter specified by the researcher.\footnote{In the empirical application, we use cross validation to select the tuning parameter $\lambda_{1,n}$. See Appendix \ref{sec:compute} for additional computational details.}


We now describe our cross-fitting procedure to estimate $L^{o}(\delta,\eta_0)$ from
data $Z^{n}=\left\{ X_{i},D_{i},Y_{i}\right\} {}_{i=1}^{n}$. Let
$[n]:=\{1,\dots,n\}$ be the observation index set.  Randomly partition
$[n]$ into approximately equal-sized $K\geq2$ folds $\left(I_{k}\right)_{k=1}^{K}$. Without loss of generality, we assume each fold is of sample size $m:=n/K$.
For each $k\in[K]:=\{1,\dots,K\}$, let $I_{k}^{c}:=[n]\backslash I_{k}$
only include observations $\textit{not}$ from fold $I_{k}$. For
each $I_{k}$, $k\in[K]$, we estimate $\eta_{0}$ by $\hat{\eta}^{k}:=(\hat{\gamma}_{1}^{k},\hat{\gamma}_{0}^{k},\hat{\omega}_{1}^{k},\hat{\omega}_{0}^{k})$,
where $\hat{\gamma}_{1}^{k}:=\hat{\gamma}_{1}\left(\left(Z_{j}\right)_{j\in I_{k}^{c}}\right)$,
$\hat{\gamma}_{0}^{k}:=\hat{\gamma}_{0}\left(\left(Z_{j}\right)_{j\in I_{k}^{c}}\right)$,
$\hat{\omega}_{1}^{k}:=\hat{\omega}_{1}\left(\left(Z_{j}\right)_{j\in I_{k}^{c}}\right)$
and $\hat{\omega}_{0}^{k}:=\hat{\omega}_{0}\left(\left(Z_{j}\right)_{j\in I_{k}^{c}}\right)$,
i.e., $\hat{\eta}^{k}$ is constructed only using data in $I_{k}^{c}$.
Then, for each $\delta$, an estimator of $L^{o}(\delta,\eta_{0})$
is
\[
\hat{L}_{n}^{o}(\delta):=\frac{1}{n}\sum_{i=1}^{n}\hat{\xi}(Z_{i})\left(\mathbf{1}\{\hat{\tau}(X_{i})\geq0\}-\delta(W_{i})\right)^{2},
\]
where
\begin{align*}
\hat{\xi}(Z_{i}) & :=\sum_{k=1}^{K}\hat{\xi}^{k}(Z_{i})\mathbf{1}\left\{ i\in I_{k}\right\} ,\hat{\xi}^{k}(Z_{i}):=\xi(Z_{i},\hat{\eta}^{k}),\\
\hat{\tau}(X_{i}) & :=\sum_{k=1}^{K}\left(\hat{\gamma}_{1}^{k}(X_{i})-\hat{\gamma}_{0}^{k}(X_{i})\right)\mathbf{1}\left\{ i\in I_{k}\right\}
\end{align*}
are estimated versions of the weight $\xi(Z_{i},\eta_{0})$ and CATE
$\tau(X_{i})$ for each $i\in[n]$. Next, let $p(w):=(p_{1}(w),p_{2}(w),\ldots p_{d_p}(w))^{\prime}$
be a vector of basis functions with dimension
$d_p:=d_p(n)$ that may grow as $n\rightarrow\infty$. Write
\begin{equation}
\hat{A}_{n}:=\frac{1}{n}\sum_{i=1}^{n}\hat{\xi}(Z_{i})p(W_{i})p(W_{i})^{\prime},\quad \hat{B}_{n}:=\frac{1}{n}\sum_{i=1}^{n}\hat{\xi}(Z_{i})p_{i}(W_{i})\mathbf{1}\left\{ \hat{\tau}(X_{i})\geq0\right\}.
\end{equation}
Our final estimated policy is defined as:
\begin{align}\label{eq:our.solution.trimmed}
\hat{\delta}^{\mathcal{T}}(w)	:=\begin{cases}
1, & \hat{\delta}(w)>1,\\
\hat{\delta}(w), & \hat{\delta}(w)\in[0,1],\\
0, & \hat{\delta}(w)<0,
\end{cases}
\end{align}
where
\begin{equation}
\hat{\delta}(w)=\hat{\beta}^{\prime}p(w),\quad \hat{\beta}:={\hat{A}_{n}}^{-}\hat{B}_{n},\label{eq:our.solution}
\end{equation}
and ${(\cdot)}^{-}$ denotes the Moore-Penrose inverse.\footnote{Our cross-fitted procedure is considered as DML2 \citep{chernozhukov2018double}. One may also consider a different cross-fitting approach, in which we solve for the optimal rule in  each fold before taking the averages over all folds. It would be interesting to compare these two approaches in policy learning problems, in light of the recent progress of \cite{velez2024asymptotic} for estimation problems.}

The rationale behind \eqref{eq:our.solution.trimmed} is as follows. Given $\hat{L}_{n}^{o}(\delta)$, one may consider finding an optimal rule by solving
\begin{equation}
\inf_{\delta\in\mathcal{D}}\hat{L}_{n}^{o}(\delta),\label{eq:our.proposal}
\end{equation}
where
\begin{equation}
\mathcal{D}:=\mathcal{D}_{n}:=\left\{ f(w)=\sum_{j=1}^{d_p}\beta_{j}p_{j}(w):\beta_{j}\in\mathbb{R},\forall j=1\ldotsd_p\right\} .\label{eq:policy.class.linear}
\end{equation}
is a  linear sieve policy class.\footnote{It is a common practice to use a class of linear functions to approximate
a probability function. See, e.g., \citet*{chen2008semiparametric}. Our theory in fact can also be extended to other policy classes, e.g.,
a class of logit functions, with more technicalities.} \eqref{eq:our.proposal} may be viewed as a \emph{weighted least squares (empirical projection)}  problem, in which we predict an estimated outcome  $\mathbf{1}\{\hat{\tau}(X_i)\geq0\}$ in space $\mathcal{D}$ with an estimated weight $\hat{\xi}(Z_i)$.  Due to the presence of the adjustment term to ``debias'', the weight $\hat{\xi}(Z_i)$ may be negative, and the Hessian matrix $\hat{A}_n$
may also not be positive semidefinite. As a result, the problem in
\eqref{eq:our.proposal} is not necessarily convex in finite sample. However, our theory shows that, whenever $\hat{\eta}^{k}$ is of
high quality and $n$ is sufficiently large (in a sense we make precise),
the probability of $\hat{A}_{n}$ not being positive definite is exponentially
small. In addition, on the event that $\hat{A}_{n}$ is positive definite, \eqref{eq:our.proposal} has a unique solution  \eqref{eq:our.solution}, in which the Moore-Penrose inverse reduces to a standard inverse.\footnote{Therefore, \eqref{eq:our.solution} also has an interesting IV interpretation with an outcome of interest $\mathbf{1}\{\hat{\tau}(X_i)\geq0\}$,
a vector of endogenous variables $p(W_i)$,
and a vector of instrument $\hat{\xi}(Z_i)p(W_i)$.} Finally, to guarantee the estimated policy is indeed a valid decision rule and also for technical tractability, \eqref{eq:our.solution.trimmed} takes a trimmed form.\footnote{See, e.g., \cite*{newey1994series,newey1999nonparametric} for  examples, in other contexts, of using trimming to improve the  theoretic performances of certain statistics.} Note \eqref{eq:our.solution.trimmed} is well-defined irrespective of whether $\hat{A}_n$ is positive definite or not.


\section{Performance guarantee}\label{sec:stat}

Let $e_{1}:=Y(1)-\gamma_{1}(X)$,
$e_{0}:=Y(0)-\gamma_{0}(X)$ and $A:=\mathbb{E}[\tau^{2}(X)p(W)p^{\prime}(W)]$. We first impose the following regularity
conditions.
\begin{assumption}
\label{asm:reg}
\begin{itemize}
\item[(i)]  There exist some constants $C_{e}$ and $C_{\gamma}$ such that $\left|e_{1}\right|\leq C_{e}$, $\left|e_{0}\right|\leq C_{e}$, $\sup_{x\in\mathcal{X}}\left|\gamma_{1}(x)\right|\leq C_{\gamma}$, $\sup_{x\in\mathcal{X}}\left|\gamma_{0}(x)\right|\leq C_{\gamma}$.
\item[(ii)] All the eigenvalues of $A$
are bounded from above and away from zero.
\end{itemize}
\end{assumption}

Under Assumptions \ref{asm:unconfounded} and \ref{asm:reg}, there exists some $C_{\xi}$ such that $\sup_{z\in\mathcal{Z}}\left|\xi(z,\eta_0)\right|\leq C_{\xi}$, and denote by $\bar{\lambda}<\infty$ and $\underline{\lambda}>0$
the maximum and minimum eigenvalues of $A$. Notably, even if $\mathbb{E}\left[\tau^{2}(X)\mid W=w\right]=0$ for some $w\in\mathbb{R}^{d_W}$, Assumption \ref{asm:reg}(ii) may still hold, thus allowing $\delta^*(w)$ to be non-unique for some $w$ values.
Next, we impose the following
statistical quality requirements on the learners of $\eta_0$. Let
$
\mathbb{E}_{k}\left[\cdotp\right]:=\mathbb{E}_{P^{n}}\left[\cdotp\mid\left\{ Z_{j}\right\} _{j\in[n]\setminus I_{k}}\right]$.
\begin{assumption}
\label{asm:quality} For each $k\in[K]$,
the following holds:
\begin{itemize}
\item[(i)] for some constant $C_{M}$ and  some constants $r_{\gamma_{1}},r_{\gamma_{0}},r_{\omega_{1}}$
and $r_{\omega_{0}}$ in $(0,1]$,
\[
\begin{aligned}\mathbb{E}_{k}\left[\int\left(\hat{\gamma}_{1}^{k}(x)-\gamma_{1}(x)\right)^{2}dF_{X}(x)\right]\leq C_{M}n^{-r_{\gamma_{1}}}, & \mathbb{E}_{k}\left[\int\left(\hat{\gamma}_{0}^{k}(x)-\gamma_{0}(x)\right)^{2}dF_{X}(x)\right]\leq C_{M}n^{-r_{\gamma_{0}}},\\
\mathbb{E}_{k}\left[\int\left(\hat{\omega}_{1}^{k}(x)-\omega_{1}(x)\right)^{2}dF_{X}(x)\right]\leq C_{M}n^{-r_{\omega_{1}}}, & \mathbb{E}_{k}\left[\int\left(\hat{\omega}_{0}^{k}(x)-\omega_{0}(x)\right)^{2}dF_{X}(x)\right]\leq C_{M}n^{-r_{\omega_{0}}};
\end{aligned}
\]
\item[(ii)] conditional on $\left\{ Z_{j}\right\} _{j\in[n]\setminus I_{k}}$ and for some constant $\tilde{C}_{M}$,
\begin{align*}
\sup_{x\in\mathcal{X}}\left|\hat{\gamma}_{1}^{k}(x)-\gamma_{1}(x)\right| & \leq\tilde{C}_{M},\sup_{x\in\mathcal{X}}\left|\hat{\gamma}_{0}^{k}(x)-\gamma_{0}(x)\right|\leq\tilde{C}_{M},\\
\sup_{x\in\mathcal{X}}\left|\hat{\omega}_{1}^{k}(x)-\omega_{1}(x)\right| & \leq\tilde{C}_{M},\sup_{x\in\mathcal{X}}\left|\hat{\omega}_{0}^{k}(x)-\omega_{0}(x)\right|\leq\tilde{C}_{M}.
\end{align*}

\end{itemize}



\end{assumption}

By Assumption \ref{asm:quality}(i), our cross-fitted learners
of $\eta_0$ are mean square consistent with certain
convergence rates. Moreover, Assumptions \ref{asm:unconfounded}, \ref{asm:reg} and \ref{asm:quality}(ii) together imply that  there exists some $\tilde{C}_{\xi}$ such that for all $k\in[K]$, $\sup_{z\in\mathcal{Z}}\left|\hat{\xi}^k(z)\right|\leq\tilde{C}_{\xi}$ conditional on $\left\{ Z_{j}\right\} _{j\in[n]\setminus I_{k}}$.

We now present a high-level stability condition of $\hat{\delta}$ useful for deriving fast convergence rates of our proposal.
\begin{assumption}\label{asm:stability}
For some positive constant $\underline{\lambda}_\varepsilon$, we have $\sup_{w\in\mathcal{W}}|\hat{\delta}(w)|\cdot\mathbf{1}\{\lambda_{\text{min}}(\hat{A}_n)\geq\underline{\lambda}_\varepsilon\}\leq C_{L}$ for some finite constant $C_L$ (which may depend on $\underline{\lambda}_\varepsilon$).
\end{assumption}
Assumption~\ref{asm:stability} essentially requires that the solution to the weighted least squares problem~\eqref{eq:our.proposal} satisfies a stability property with respect to the sup norm.
\footnote{See, for example, the sup-norm stability property of empirical $L_2$ projections using certain basis functions (e.g., splines and wavelets), which has been exploited by \cite*{huang2003local,belloni2015some,chen2015optimal} for sharp bias control in least squares series estimation. Our weighted least squares problem~\eqref{eq:our.proposal}, however, differs from those studied in the preceding literature due to the presence of estimated weights and outcomes.}  With Assumptions \ref{asm:unconfounded}-\ref{asm:quality}, we verify in Appendix \ref{sec:Bspline} that Assumption \ref{asm:stability} holds if $\mathcal{D}$ is constructed with B-spline basis functions.
Finally, we consider the following  margin condition that concerns  the distribution of $\tau(X)$ in the neighborhood of $\tau(X)=0$ (see also \citealt{tsybakov2004optimal,kitagawa2018should}):
\begin{assumption}\label{asm:margin}
There exist positive constants $C_{\tau}$, $\alpha$, and $t^{*}$ such that
\[
P\left( \left| \tau(X) \right| \leq t \right) \leq C_{\tau} t^{\alpha}, \quad \text{for all } 0 \leq t \leq t^{*}.
\]
\end{assumption}
Note Assumption \ref{asm:margin} rules out  $P\{\tau(X)=0\}>0$, implying that $\delta^*(w)$ will be unique a.e. with respect to the marginal distribution of $W$. We are now ready to state our main result. Denote by $d_p^*$ the dimension of the basis functions for  $\{f^2:f\in\mathcal{D}\}$, where $\mathcal{D}$ is defined in \eqref{eq:policy.class.linear}, and write $\zeta_{p}:=\sup_{w\in\mathcal{W}}\left\Vert p(w)\right\Vert$.\footnote{The magnitudes of $d_p^*$ and $\zeta_p$ depend the choice of the basis functions. An upper bound of $d^*_p$ is $d_p^2$, but  $d_p^*$ may be as small as  $O(d_p)$, e.g., for B-splines. The quantity of $\zeta_p$ is a key object of interest in the series estimation literature. It is well known that $\zeta_p=O(\sqrt{d_p})$ for general spline basis functions (see, e.g., \citealt{newey1997convergence}). For B-splines, its structural properties imply that in fact, $\zeta_p\leq1$. See Appendix \ref{sec:Bspline} for additional treatments.}
\begin{thm}
\label{thm:main}Suppose Assumptions \ref{asm:unconfounded}-\ref{asm:stability}
hold. Fix $0<\varepsilon<\min\left\{ \underline{\lambda},6\tilde{C}_{\xi}\zeta_{p}^{2},3C_{\xi}\zeta_{p}^{2}/2\right\}$, and  let
\begin{align*}
R_{n,O} & :=\frac{d_p^*}{n}+\sup_{P^{n}}\inf_{\delta\in\mathcal{D}}\left[L(\delta,\tau)-L(\delta^{*},\tau)\right],\\
R_{n,B} & :=\text{max}\{\left(\log2d_{p}\right)^{2}\zeta_{p}^{6},d_p\zeta_{p}^{3}\}\left(n^{-2r_{\gamma_{1}}}+n^{-2r_{\gamma_{0}}}+n^{-\left(r_{\omega_{1}}+r_{\gamma_{1}}\right)}+n^{-\left(r_{\omega_{0}}+r_{\gamma_{0}}\right)}\right),\\
R_{n,V} &:=\zeta_{p}^{3}n^{-1},\\
R_{n,F}&:=4C_{\xi}d_{p}\left[\exp\left(\frac{-n\varepsilon^{2}}{4C_{\xi}^{2}\zeta_{p}^{4}}\right)+K\exp\left(\frac{-n\varepsilon^{2}}{16K\tilde{C}_{\xi}^{2}\zeta_{p}^{4}}\right)\right].
\end{align*}
Then, for each $n$ such that
\[
4C_{M}\zeta_{p}^{2}\left(n^{-r_{\gamma_{1}}}+n^{-r_{\gamma_{0}}}+n^{-\frac{r_{\omega_{1}}+r_{\gamma_{1}}}{2}}+n^{-\frac{r_{\omega_{0}}+r_{\gamma_{0}}}{2}}\right)<\varepsilon,
\]
the following statements hold.
\begin{itemize}
\item[(i)] For some constant $\mathcal{C}$ that is independent of $n$, $d_p$, $d^*_p$ and $\zeta_p$,
\begin{align}\label{eq:rate.main}
\sup_{P_{n}}\mathbb{E}_{P_{n}}\left[L(\hat{\delta}^{\mathcal{T}},\tau)-L(\delta^{*},\tau)\right]& \leq\mathcal{C}\left(R_{n,O}+R_{n,B}+R_{n,V}\right)+R_{n,F}.
\end{align}
\item[(ii)] The right-hand side of \eqref{eq:rate.main} improves to
\begin{align}\label{eq:rate.w.margin}
\mathcal{C}\left(R_{n,O}+R_{n,B}+R_{n,V}\left(n^{-r_{\gamma_{1}}}+n^{-r_{\gamma_{0}}}\right)^{\frac{\alpha}{\alpha+2}}\right)+R_{n,F},
\end{align}
with a constant $\mathcal{C}$ suitably redefined (but also independent of $n$, $d_p$, $d^*_p$ and $\zeta_p$), if in addition, Assumption
\ref{asm:margin} holds and $n$ is also such that
$\left(4C_{M}C_{\tau}^{-1}\left(n^{-r_{\gamma_{1}}}+n^{-r_{\gamma_{0}}}\right)\right)^{\frac{1}{\alpha+2}}<t^{*}$.
\end{itemize}
\end{thm}



Theorem \ref{thm:main} provides an upper bound for  the uniform excess risk whenever $n$ is sufficiently large. As long as  $R_{n,O}$, $R_{n,B}$ and $R_{n,V}$
all go to zero at a polynomial rate as a function of $n$, the exponential term $R_{n,F}$ will be asymptotically negligible, implying immediately that  if Assumptions \ref{asm:unconfounded}-\ref{asm:stability} hold,
\[
\sup_{P_{n}}\mathbb{E}_{P_{n}}\left[L(\hat{\delta}^{\mathcal{T}},\tau)-L(\delta^{*},\tau)\right]=O(R_{n,O}+R_{n,B}+R_{n,V}),
\]
and with the additional Assumption \ref{asm:margin},
\[
\sup_{P_{n}}\mathbb{E}_{P_{n}}\left[L(\hat{\delta}^{\mathcal{T}},\tau)-L(\delta^{*},\tau)\right]=O\left(R_{n,O}+R_{n,B}+R_{n,V}\left(n^{-r_{\gamma_{1}}}+n^{-r_{\gamma_{0}}}\right)^{\frac{\alpha}{\alpha+2}}\right).
\]
Each part of  \eqref{eq:rate.main} and \eqref{eq:rate.w.margin} is interpretable. Term $R_{n,F}$ controls the excess risk even when $\hat{A}_n$ and its oracle version $A_n:=\frac{1}{n}\sum_{i=1}^{n}\xi(Z_{i},\eta_0)p(W_{i})p(W_{i})^{\prime}$
do not ``behave nicely'' (i.e., when they are not positive definite). When they do ``behave nicely'', consider the following oracle ``empirical risk minimization''
(ERM) problem with known $\eta_0$:
\begin{equation}
\min_{\delta\in\mathcal{D}}L_{n}^{o}(\delta,\eta_0),\label{eq:oracle}
\end{equation}
where
\begin{align}
L_{n}^{o}(\delta,\eta_0):=\frac{1}{n}\sum_{i=1}^{n}\left[\xi(Z_i,\eta_0)\left(\mathbf{1}\{\gamma_{1}(X_i)-\gamma_{0}(X_i)\geq0\}-\delta(W_i)\right)^{2}\right]\nonumber\label{eq:empirical.average.debiased}.
\end{align}
The oracle excess risk is of  $O\left(R_{n,O}\right)$,
containing an approximation error term $\sup_{P^{n}}\inf_{\delta\in\mathcal{D}}\left[L(\delta,\tau)-L(\delta^{*},\tau)\right]$ due to  using $\mathcal{D}$ to approximate $\delta^*$.\footnote{This approximation quality term may depend on whether Assumption \ref{asm:margin} is imposed or not, and may be further analyzed provided with suitable smoothness conditions on $\delta^*$, which we leave for future research.
}
Since $\eta_0$ is in fact unknown and needs to be estimated, we have to pay an additional price from the ``remainder estimation error''. Interestingly, the asymptotic order of this remainder error depends on whether the margin condition is imposed or not. Without margin condition, the ``remainder estimation error'' is $O\left(R_{n,B}+R_{n,V}\right)$, where $R_{n,B}$ is a bias term
while $R_{n,V}$ is a variance term.
If  the margin condition is imposed
with some $\alpha>0$, the rate for the variance term improves to $O(R_{n,V}\left(n^{-r_{\gamma_{1}}}+n^{-r_{\gamma_{0}}}\right)^{\frac{\alpha}{\alpha+2}})$.


The optimality of our proposal depends on the complexity of $\delta^*$. If there exists some $\delta\in\mathcal{D}$ that solves \eqref{eq:loss.proposal} with $d_p$ fixed, we say $\delta^{*}$ is parametric. If no rule in $\mathcal{D}$ solves \eqref{eq:loss.proposal}, we say $\delta^{*}$ is nonparametric.
The following proposition suggests that,  when  $\delta^*$ is parametric,  our procedure
is asymptotically optimal in terms of the rate with Assumptions \ref{asm:unconfounded}-\ref{asm:stability}. Moreover, it is also asymptotically semiparametrically efficient with the additional Assumption \ref{asm:margin}.

\begin{prop}\label{prop:parametric}
Consider the case when $\delta^*$ is parametric.
\begin{itemize}
\item[(i)] Suppose Assumptions \ref{asm:unconfounded}, \ref{asm:reg}, \ref{asm:quality} and \ref{asm:stability} hold true, $r_{\gamma_{1}}>1/2$, $r_{\gamma_{0}}>1/2$, $r_{\omega_{1}}+r_{\gamma_{1}}>1$ and $r_{\omega_{0}}+r_{\gamma_{0}}>1$.
Then,  $R_{n,O}=R_{n,V}=O(\frac{1}{n})$, $R_{n,B}=o(\frac{1}{n})$, and
\begin{equation}\label{eq:rate.parametric}
\sup_{P_{n}}\mathbb{E}_{P_{n}}\left[L(\hat{\delta}^{\mathcal{T}},\tau)-L(\delta^{*},\tau)\right]=O\left(\frac{1}{n}\right).
\end{equation}

\item[(ii)] If in addition, Assumption \ref{asm:margin} also holds, then $R_{n,O}=O(\frac{1}{n})$, $R_{n,B}=R_{n,V}=o(\frac{1}{n})$ and \eqref{eq:rate.parametric} is still true. Moreover, suppose $\Omega :=A^{-1}VA^{-1}$ is  positive definite, where $V:=\mathbb{E}\left[SS^{\prime} \right]$,
\begin{align*}
S&:=\xi(Z)p(W)\mathbf{1}\{\tau(X)\geq 0\}-\mathbb{E}[\xi(Z)p(W)\mathbf{1}\{\tau(X)\geq 0\}].
\end{align*}
Then, as $n\rightarrow\infty$,
\[
n\left(L(\hat{\delta}^{\mathcal{T}},\tau)-L(\delta^{*},\tau)\right)\overset{d}{\rightarrow}N(0,\Omega)^{\prime}AN(0,\Omega),
\]
where $N(0,\Omega)$ denotes a multivariate normal distribution with mean $0$ and covariance matrix $\Omega$.
\end{itemize}
\end{prop}

Note if $\delta^*$ is parametric, Assumption \ref{asm:reg}(ii) implies a unique $\beta^*\in\mathbb{R}^{d_p}$ such that $\left(\beta^{*}\right)^{\prime}p(w)$ solves \eqref{eq:loss.proposal}. The study of $\beta^*$ has a known semiparametric efficiency bound $\Omega$. See e.g., \cite*{newey1994asymptotic,ackerberg2014asymptotic}.  By Proposition \ref{prop:parametric}(ii),  our procedure is asymptotically equivalent to the oracle that solves \eqref{eq:oracle}. In particular, $\sqrt{n}\left(\hat{\beta}-\beta^{*}\right)\overset{d}{\rightarrow}N(0,\Omega)$, achieving the semiparametric efficiency bound asymptotically. Moreover, with a parametric $\delta^*$,
\[n\left(L(\hat{\delta},\tau)-L(\delta^{*},\tau)\right)	=n\left(\hat{\beta}-\beta^{*}\right)^{\prime}A\left(\hat{\beta}-\beta^{*}\right).\]
The asymptotic efficiency of $\beta^*$ translates to that of $L(\hat{\delta},\tau)$, implying that our  procedure is asymptotically efficient as well.\footnote{A parametric $\delta^*$ is not necessarily restrictive. Even if $\delta^*$ is nonparametric, one may wish to target the ``second best'' rule, i.e., $\delta^{SB}\in\arg\inf_{\delta\in\mathcal{D}}L(\delta,\tau)$, for which Proposition \ref{prop:parametric} can be shown to still apply.}

When $\delta^*$ is nonparametric, the discussion of the optimality of our procedure is more involved.   In Appendix \ref{sec:lower.bound}, we derive a minimax lower bound for $\delta^*$ in the style of \cite{stone1982optimal}, which we suspect is attainable by our procedure for certain high smoothness class of $\delta^*$ when $d_p$ grows sufficiently slowly. We leave the verification of this conjecture, as well as the asymptotic distribution of $(L(\hat{\delta},\tau)-L(\delta^{*},\tau))$ for future research.

\begin{rem}\label{rem:proof.strategy}
The proof strategy of Theorem \ref{thm:main} is significantly different from the existing approaches in the policy learning literature (c.f.  \citealt{kitagawa2018should,AW20}). For the oracle problem, one may follow the classic theory developed by \cite{vapnik1971uniform,vapnik1974theory} to bound
\[
\sup_{\delta\in\mathcal{D}}\mathbb{E}_{P^{n}}\left|L_{n}^{o}(\delta,\eta_0)-L^o(\delta,\eta_0)\right|,
\]
resulting in an order of
$O(1/\sqrt{n})$ in general even in the case of a parametric $\delta^*$. Instead, we exploit the weighted least squares structure embedded in $L^o_{n}$ and adapt (in Lemma \ref{lem:main.1}) a refined maximal inequality developed by \citet[][Theorem 2]{KOHLER20001}, leading to a  convergence rate for the oracle that can be as fast as $O(1/n)$. For the remainder estimation error part, one possibility is to follow
\citet[][Section 3.2]{AW20} and control the estimation error uniformly over all rules in $\mathcal{D}$. This approach, however, would only lead to a rate of $o(1/\sqrt{n})$ even with a parametric $\delta^*$,  much slower than our result of  $O(R_{B,n}+R_{V,n})$ even without margin condition. We, instead, utilize the fact that both \eqref{eq:our.proposal} and \eqref{eq:oracle} have explicit solutions in large sample that  satisfy certain first order optimality conditions, which allows us to derive a faster rate (Lemma \ref{lem:main.2}). These nice structures for the remainder estimation errors are only valid in large samples. Therefore, our results are asymptotic in nature, as opposed to the finite sample performance guarantee in \cite{kitagawa2018should}.
\end{rem}


\begin{rem} \label{rem:plug.in}Currently, it is not  entirely clear  to what extent our debiased approach is strictly needed for some of the results in Theorem \ref{thm:main} to hold. Indeed, debiasing is costly: $\hat{L}_n^o(\delta)$ is a low-bias, but more noisy estimator of $L(\delta, \tau)$ due to the indefiniteness of $\hat{A}_n$, which may compromise the finite-sample performance of debiasing. A natural alternative is to solve
\begin{equation}
\inf_{\delta\in\mathcal{D}}\hat{L}_{n}(\delta),\label{eq:plug-in.1}
\end{equation}
where
\begin{equation}
\hat{L}_{n}(\delta):=\frac{1}{n}\sum_{i=1}^{n}\hat{\tau}^2(X_{i})\left(\mathbf{1}\{\hat{\tau}(X_{i})\geq0\}-\delta(W_{i})\right)^{2},
\end{equation}
and $\hat{\tau}$ is any estimator of $\tau$ that may or may not be cross-fitted. This plug-in approach maintains the positive semi-definiteness of the associated Hessian matrix. It is straightforward to extend our theory and establish the oracle rate for this plug-in approach, which would be the same as $R_{n,O}$.  Analogous analyses (c.f. proof of Lemma \ref{lem:main.2}) imply that the remainder estimation error is determined asymptotically by
\begin{equation}\label{eq:plug.in.bias}
\left(\mathbb{E}_{P^{n}}\left\Vert \hat{A}_{n}^{\text{plug-in}}-A_{n}^{\text{plug-in}}\right\Vert ^{2}+\mathbb{E}_{P^{n}}\left\Vert \hat{B}_{n}^{\text{plug-in}}-B_{n}^{\text{plug-in}}\right\Vert ^{2}\right),
\end{equation}
where
\begin{align*}
A_{n}^{\text{{plug-in}}}&:=\frac{1}{n}\sum_{i=1}^{n}\tau^{2}(X_{i})p(W_{i})p(W_{i})^{\prime},\\
\hat{A}_{n}^{\text{{plug-in}}}&:=\frac{1}{n}\sum_{i=1}^{n}\hat{\tau}^{2}(X_{i})p(W_{i})p(W_{i})^{\prime},\\
B_{n}^{\text{{plug-in}}}&:=\frac{1}{n}\sum_{i=1}^{n}\tau^{2}(X_{i})p(W_{i})\mathbf{1}\left\{ \tau(X_{i})\geq0\right\},\\
\hat{B}_{n}^{\text{{plug-in}}}&:=\frac{1}{n}\sum_{i=1}^{n}\hat{\tau}^{2}(X_{i})p_(W_{i})\mathbf{1}\left\{ \hat{\tau}(X_{i})\geq0\right\}.
\end{align*}
If cross-fitting is used, \eqref{eq:plug.in.bias} in general presents an asymptotic bias larger  than $R_{B,n}$ (c.f. Lemmas \ref{lem:remainder.5} and \ref{lem:remainder.6}). However, if cross-fitting is not used and conditional on the specific structure of the estimator $\hat{\tau}$, we suspect the asymptotic bias in \eqref{eq:plug.in.bias}  may be as fast as $R_{B,n}$, in light of the classic ``low bias'' results of certain plug-in approaches in the semiparametric estimation literature, e.g., \cite*{ai2003efficient,chen2003estimation,hirano2003efficient}.  Whether the plug-in approach preserves the same asymptotic remainder estimation error is an intricate but fascinating question that we leave for future research. In the empirical applications below, we present results with our main debiased approach as well as the plug-in alternative.
\end{rem}

\section{Capacity constraint}\label{sec:capacity}

In this section, we consider a decision maker facing convex constraints for the allocation rules. As a leading case, suppose $W$ is discrete and a capacity constraint exists on how many people in the population can get treatment. With such capacity constraint, the problem is convex with
differentiable objective and constraint functions,  and the Slater's condition can be verified to hold. Therefore, the optimal solution is characterized by the well-known KKT condition (e.g., \citealt{boyd2004convex}, Chapter
5, p.244), as we show below.
\begin{prop}\label{prop:capacity}
Suppose $W$ is discrete and takes values $\left\{ w_{j}\right\} _{j=1}^{J}$
with corresponding probabilities $\left\{ p_{j}\right\} _{j=1}^{J}$, where
$p_{j}>0$ for all $j=1,\ldots, J$. Consider solving (\ref{eq:population.optimal}) with $\alpha=2$
and a capacity constraint $\mathbb{E}[\delta(W)]\leq t$ for some
$0<t<1$. Let
\begin{align*}
b_{j} & :=\mathbb{E}\left[\tau^{2}(X)\mathbf{1}\left\{ \tau(X)\geq0\right\} \mid W=w_{j}\right],\\
a_{j} & :=\mathbb{E}\left[\tau^{2}(X)\mid W=w_{j}\right].
\end{align*}
Wlog, suppose $a_{j}>0$ for all $j=1\ldots J$\footnote{The case of $a_{j}=0$ can be excluded as an action of 0 would be optimal and  not add to the capacity.}, and index groups
so that $b_{1}\geq b_{2}\ldots\geq b_{J}$. If the capacity constraint
is not binding (i.e., $\sum\limits _{j=1}^{J}(p_{j}b_{j}/a_{j})\leq t$),
then the unconstrained solution $\left\{ {b_{j}}/{a_{j}}\right\} _{j=1}^{J}$
is optimal. Otherwise, the optimal decision is
\begin{align*}
\delta_{j}^{*}  =\frac{b_{j}}{a_{j}}-\frac{\lambda^{*}}{2a_{j}},\text{ for all }j\leq J^{*},\quad
\delta_{j}^{*} & =0,\text{ for all }j>J^{*},
\end{align*}
where $J^{*}\in\left\{ 1,\ldots,J\right\} $ and $\lambda^{*}\geq0$
are jointly determined such that
\[
\lambda^{*}=\frac{\sum_{j=1}^{J^{*}}\frac{p_{j}b_{j}}{a_{j}}-t}{\sum_{j=1}^{J^{*}}\frac{p_{j}}{2a_{j}}}.
\]
\end{prop}
Proposition \ref{prop:capacity} highlights an interesting insight:   with a capacity constraint, a regret-averse decision maker would reduce the fractional treatment for \emph{all} groups, possibly with some groups with smallest $b_j$ not treated at all if the capacity constraint is too severe. In contrast, when $\alpha=1$, the decision maker always prioritizes treating the $W$ groups with the largest  positive average treatment effect until the capacity constraint is filled, possibly with fractional allocation for the marginal group. In the hypothetical policy question  from \cite{resnjanskij2024can} considered in the introduction, suppose we have a capacity constraint that at most a $t\leq1$ fraction of the population can be offered with the mentoring program. Since there is only one $W$ group whose average treatment effect is positive, the optimal constrained rule is easy to calculate (see Table \ref{tab:constraint}). For example, if $\alpha=1$, the optimal rule is to treat $t$ fraction of the population; if $\alpha=2$, the optimal rule is to treat $t$ fraction of the population if $t<0.88$ and to treat 0.88 of the population if $t\geq0.88$ (as 0.88 is the unconstrained optimal which does not violate the capacity constraint). In this simple case with one $W$ group, $\alpha=1$ and $2$  would share the same optimal rule if $t<0.88$.
\begin{table}[htbp]
\centering
\caption{Optimal allocation rule for the hypothetical policy question in \cite{resnjanskij2024can} with a capacity constraint}\label{tab:constraint}
\begin{tabular}{lccc}
\toprule
& \multicolumn{3}{c}{Atkinson inequality index} \\
 \cmidrule(lr){2-4}
Capacity constraint &  $\alpha=1$ & $\alpha=2$ & $\alpha=3$ \\
\midrule
$t\in[0.88,1]$  & $t$ & 0.88 & 0.82 \\
$t\in[0.82,0.88)$ &  $t$ &  $t$ & 0.82 \\
$t\in[0,0.82)$ &  $t$ &  $t$ &  $t$ \\
\bottomrule
\end{tabular}
\end{table}



In the setup of Proposition \ref{prop:capacity}, we can still learn the optimal constrained rule from data by solving \eqref{eq:our.proposal} and incorporating  additional constraints:\footnote{We do not need to impose  the constraints that  $\delta(w_{j})\leq1$ for $j=1,...,J$, as they  will be non-binding  at the true population constrained optimal rule.}
\begin{equation}\label{eq:add.constraint}
\frac{1}{n}\sum_{i=1}^{n}\delta(W_{i})\leq t,\delta(w_{j})\geq0,j=1,...,J,
\end{equation}
which is still a convex program with  differentiable objective and constraints and can be efficiently computed. However, establishing the statistical performance guarantee is more involved due to the known technical difficulty associated with not knowing  whether the constraints in \eqref{eq:add.constraint} are binding or not in general.

\section{Empirical applications}\label{sec:emp.app}

\subsection{Job Training Partnership Act (JTPA) Study}\label{sec:JTPA}

We revisit the experimental dataset of the National JTPA Study that aimed to
measure the benefit and cost of employment and training programs. Our sample consists of 9223 observations, in which the treatment $D$ was randomized to generate the applicants'  eligibility for receiving a mix of training, job-search assistance, and other
services provided by the JTPA. The outcome of interest $Y$ is the total individual earnings in the 30 months after program
assignment.\footnote{We take the intention-to-treat perspective. One may also consider an net-of-cost outcome, which would further deduct 774 dollars for each of  those assigned to treatment.} The study also collected a variety of the applicants' background information ($X$), some of which might be perceived  as sensitive, e.g.,  gender, race and marital
status. Following \cite{kitagawa2018should}, we consider a scenario in which a policymaker can only design treatment policies based on pre-program
years of education (``education'') and the pre-program earnings (``income'') --- these two variables become the $W$ in our setup.  As an illustration of our debiased approach, we choose $K=5$ and  estimate $\gamma_1$ and  $\gamma_0$ via lasso with 10-fold cross-validation, with all
interactions and squared terms of $X$. We estimate $\omega_1$ and $\omega_0$ with the minimum distance estimator with a Tikhonov
penalty \eqref{eq:minimum.distance.penalty}, where the tuning parameter is selected via cross validation. See Section \ref {sec:compute} for computational details and our algorithm to calculate $\hat{\xi}(Z_i)$.

\begin{figure}[!t]]
\includegraphics[width=0.49\linewidth]{figures/sq.debiased.pdf}
\includegraphics[width=0.49\linewidth]{figures/linear.debiased.pdf}
\medskip
\includegraphics[width=0.49\linewidth]{figures/sq.plugin.ols.pdf}
\includegraphics[width=0.49\linewidth]{figures/linear.ipw.pdf}
{\footnotesize \textit{Notes}: Income brackets are defined according to pre-program  earnings as follows: 1 ($\leq\$220$),  2 ($>\$220$ and  $\leq\$3800$),  3 ($>\$3800$). Education brackets are defined according to pre-program years of schooling: 1 ($\leq11$, high school dropouts); 2 ($=12$,  high school graduates); 3 ($>12$, with higher education. Top left: our squared-regret approach with debiasing; Top right: linear regret approach, with $CATE(W)$ estimated with debiasing  and $\eta_0$ fitted with lasso. Down left: squared-regret approach with $\gamma_1$ and  $\gamma_0$ estimated by OLS; Down right: linear regret approach, with $CATE(W)$ estimated with inverse propensity score weighting with the known propensity score of $2/3$. The numbers in each of the brackets in the left two graphs refer to the corresponding estimated treatment assignment fractions, while the numbers in each of the brackets in the right two graphs refer to the estimated $CATE(W)$.
}
\caption{JTPA: estimated simple bracket rules }\label{fig:JTPA.bracket}
\end{figure}

To start with, suppose the policymaker is interested in implementing a simple rule based on nine pre-determined income and education brackets (defined in the note of Figure \ref{fig:JTPA.bracket}). In this case, $W$ is discrete, and the optimal rule can be solved bracket-by-bracket. Figure \ref{fig:JTPA.bracket} reports the results for our squared regret debiased approach,   a squared regret plug-in approach, as well as two  linear regret approaches.
Although the majority of the estimated CATEs conditional on $W$ are positive, the fractional nature of our estimated policies reveals plausible and considerable treatment effect heterogeneity at the $X$ level for some brackets, demonstrating the value of our approach compared to the standard mean regret paradigm. For example, for those units in education bracket 3 and income bracket 3, the debiased CATE estimate is slightly positive (41), implying all units shall be treated. However, an IPW estimate of the same CATE (-3361) would imply that no-one should be treated. For this group of workers, our squared regret debiased optimal policy is 0.63, indicating that workers in the high-education and high-income bracket may display drastically different treatment effects
from each other, which can lead to a high regret-inequality should a non-fractional policy be applied. The pattern of the squared-regret policy estimates between the plug-in and debiased approaches are similar for many brackets, although some disparities do exist.

Next, we consider a policymaker designing a class of linear sieve policies based on education and income. As an illustration, for each of the education and income variables, we create cubic B-splines with a total of 5 degrees of freedom. The multivariate B-splines are then generated as tensor products of the two. We present estimated policies of the debiased and plug-in approaches for selected values of the income and education variables in Figure \ref{fig:JTPA.heatmap}. Both approaches again indicate considerable effect heterogeneity in the population, although  disparities remain in the exact fitted values for some $W$ groups.

\begin{figure}[http]
\centering
\includegraphics[width=0.49\linewidth]{figures/sq.debiased.spline.pdf}
\includegraphics[width=0.49\linewidth]{figures/sq.plug.in.ols.spline.pdf}
\caption{JTPA: estimated linear sieve policy rules with multivariate B-splines}
\label{fig:JTPA.heatmap}
\end{figure}




\subsection{International Stroke Trial}
As a second application and to demonstrate the value of our approach in medical studies, we analyze the International Stroke Trial (IST, \citealt{international1997international}) that assessed the effect of aspirin and other treatments  for patients with presumed acute ischemic stroke. Following \cite*{yadlowsky2025evaluating}, we focus on the treatment of aspirin only on the outcome of whether there is death
or dependency at 6 months. This leaves us with a sample of 18304 patients  from over 30 countries. For each patient, we also observe a vector of 39 covariates ($X$), including  their gender, age as well as some of their medical history and geographical information. In this exercise, we consider a hypothetical scenario in which a doctor determines whether a patient should be treated with aspirin only based on their age ($W$). The aim is to assess whether our approach would generate significantly different treatment fractions compared to the mean regret approach.

\begin{figure}[http]
\centering
\includegraphics[width=0.7\linewidth]{figures/RATE.pdf}
    \caption{IST: estimated optimal treatment fractions based on age with B-splines}
    \label{fig:rate}
\end{figure}


We estimate the nuisance parameters with the same methodology described in Section \ref{sec:JTPA}. For the age variable, we create a cubic B-spline with a total degree freedom of 6. Figure \ref{fig:rate} reports our estimated optimal treatment fractions for patients with age between 39 and 92.  As the CATEs are all positive for all considered age groups, a linear regret approach will recommend to treat everyone for all age groups. In sharp contrast, the estimated treatment fractions with our debiased approach is between  25\% and 75\% for most age values, revealing considerable treatment heterogeneity among those sharing the same ages. The treatment proportion is especially close to 0.5 for patients with age between 75 to 85, suggesting that  a singleton ``treat everyone''  rule would potentially harm significantly some of those patients, leaving some of them with large regrets. The fitted curve with the plug-in approach shares the same downward sloping pattern as the debiased approach, although the estimated treatment proportions is slightly higher for all age groups. In light of these findings, we think that our squared regret approach to policy learning reveals additional important information that cannot be assessed with the mean regret approach alone.


\bibliographystyle{ecta}
\bibliography{bib}