EconBase
← Back to paper

Policy Learning with Distributional Welfare

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

68,656 characters

Policy Learning with Distributional Welfare


\title{Policy Learning with Distributional Welfare\thanks{For helpful discussions, the authors are grateful to Debopam Bhattacharya,
Kei Hirano, Nathan Kallus, Toru Kitagawa, Ashesh Rambachan, Vira Semenova, and participants
at the 2024 Econometric Society Interdisciplinary Frontiers (ESIF)
Conference, the 2024 International Society for Clinical Biostatistics
Conference, the Advances in Econometrics Conference 2023 and the seminar
participants at Brown and Essex. We also thank Xuanman Li for her
excellent research assistance. Yifan Cui acknowledges the financial
support from the National Key R\&D Program of China (2024YFA1015600), and the National Natural Science Foundation of China (12471266 and U23A2064).}}
\author{Yifan Cui\\
 Center for Data Science\\
 Zhejiang University\\
 \UrlFont[email removed]\and
Sukjin Han\\
 School of Economics\\
 University of Bristol\\
 \UrlFont[email removed]}
\date{\today}

\maketitle
\vspace{-0.2cm}

\begin{abstract}
In this paper, we explore optimal treatment allocation policies that
target distributional welfare. Most literature on treatment choice
has considered utilitarian welfare based on the conditional average
treatment effect (ATE). While average welfare is intuitive, it may
yield undesirable allocations especially when individuals are heterogeneous
(e.g., with outliers)---the very reason individualized treatments
were introduced in the first place. This observation motivates us
to propose an optimal policy that allocates the treatment based on
the conditional \emph{quantile of individual treatment effects} (QoTE).
Depending on the choice of the quantile probability, this criterion
can accommodate a policymaker who is either prudent or negligent.
The challenge of identifying the QoTE lies in its requirement for
knowledge of the joint distribution of the counterfactual outcomes,
which is not generally point-identified. We introduce minimax policies
that are robust to this model uncertainty. A range of identifying
assumptions can be used to yield more informative policies. For both
stochastic and deterministic policies, we establish the asymptotic
bound on the regret of implementing the proposed policies. The framework
can be generalized to any setting where welfare is defined as a functional
of the joint distribution of the potential outcomes.

\vspace{0.1in}

\noindent \textit{JEL Numbers:} C14, C31, C54.

\noindent \textit{Keywords:} Treatment regime, treatment rule, individualized
treatment, distributional treatment effects, quantile treatment effects,
partial identification, sensitivity analysis.
\end{abstract}

\section{Introduction\label{sec:Introduction}}

Individuals are heterogeneous, so are their responses to treatments
or programs. When designing policies (e.g., rules of allocating treatments
or programs), it is important to reflect the heterogeneity of individual
treatment effects. A policymaker (PM), or equivalently an analyst,
would devise a policy to achieve a specific objective (e.g., welfare).
Depending on how the PM aggregates individual gains, her objective
can be viewed as either \emph{utilitarian} or \emph{non-utilitarian}.
A utilitarian PM would consider welfare that takes the sum or average
of individual gains to ensure the greatest benefits for the greatest
number, whereas a non-utilitarian (e.g., prioritarian, maximin) PM
would prioritize specific groups of individuals. The utilitarian objective
has been the most widely-used criterion in the literature of treatment
allocations and policy learning (e.g., \citet{manski2004statistical};
see below for a further review). However, there may be settings where
the utilitarian goal is less sensible. For example, the target population
may exhibit skewed heterogeneity (e.g., outliers). As another example,
the PM may want to target a vulnerable population or privileged individuals,
or a certain share of benefited individuals.\footnote{The possibility of non-utilitarian welfare is also briefly mentioned
in \citet{manski2004statistical}.} The purpose of this paper is to explore objectives of a (non-utilitarian)
PM who is concerned with certain aspects of the distribution (e.g.,
tails) of treatment effects or who has political incentives and thus
makes decisions influenced by vote shares.

In this paper, we develop a policy learning framework that concerns
distributional welfare. A policy is defined as a mapping from individuals'
observed characteristics to either a deterministic or stochastic decision
of treatment allocation. Intuitively, the knowledge of individual
treatment effects conditional on characteristics plays a crucial role
in learning such a policy. We propose an objective function that is
formulated based on the conditional quantile of individual treatment
effects (QoTE). This objective function is robust to outliers of treatment
effects and, more importantly, can reflect the PM's level of prudence
toward the target population. As quantifying the uncertainty of allocation
decisions is intrinsically difficult (e.g., \citet{chen2023inference}),
the ability to adjust the level of prudence can be practically valuable
to the PM.

Suppose the PM employs the utilitarian welfare, which can be written
as a function of the conditional average treatment effect (ATE). If
the policy class is unconstrained, it is optimal for the utilitarian
PM to treat each subgroup (defined by observed characteristics) whenever
their ATE is positive. Suppose that this PM faces a target subgroup,
say black females, whose distribution of treatment effects exhibits
that a small share of individuals enjoys positive treatment effects
that dominate the negative effects of the remaining majority. If the
resulting ATE is positive, then the PM would treat \emph{all} black
females, harming the majority. The objective function based on the
QoTE with the quantile probability $\tau=0.5$ (i.e., the median of
treatment effects) would not suffer from this sensitivity to outliers.
Moreover, the PM can choose the quantile probability $\tau$ (i.e.,
the rank in individual treatment effects) to set a reference group.
A large $\tau$ corresponds to a PM who is willing to focus on privileged
individuals in each subgroup, ignoring the majority of less advantaged,
thus being a \emph{negligent} PM. A small $\tau$ corresponds to a
PM who is concerned with the disadvantaged, treating each subgroup
only if most benefit from the treatment, thus being a \emph{prudent}
PM. Relatedly, we show that the PM equipped with the QoTE can be interpreted
as being concerned with vote shares when each individual casts a vote
whenever he or she experiences a positive gain from the treatment.

An alternative objective function that can be robust to certain outliers
is the one based on the conditional quantile treatment effect (QTE)
which contrasts the quantiles of treated and untreated outcomes. We
argue that this quantity may not be a desirable basis for individualized
treatment decisions, because an individual represented by the quantile
of treated outcomes is not necessarily the same individual represented
by the same quantile of untreated outcomes. For example, as shown
below, it is difficult for the PM to aim the level of prudence when
the criterion is based on the QTE. On the other hand, the QoTE by
definition captures an individual with a specific rank in gains and
thus naturally accommodates the notion of prudence of PM.

Despite the desirable properties of the PM's objective function constructed
from the QoTE, the challenge is that the QoTE is not generally point-identified
even when the PM has access to experimental data. This is due to the
fact that the joint distribution of counterfactual outcomes is involved
in the definition of the QoTE. We therefore propose a minimax criterion
that is robust to model ambiguity. In particular, we propose to minimize
the worst-case regret calculated over the class of joint distributions
of counterfactual outcomes that are compatible with the data and identifying
assumptions. We then show that a range of identifying assumptions
that can be imposed to tighten the identified set of the QoTE, sometimes
to a singleton, leading to more informative policies. These assumptions
can be imposed by practitioners depending on their specific settings.
For some assumptions, bounds on the QoTE may not have a closed-form
expression. In this case, an optimization algorithm can be used to
compute the bounds. By using a Bernstein approximation, we show how
the optimization problem becomes a simple linear programming.

We establish theoretical properties of the proposed minimax policy
by providing asymptotic bounds on the regret of implementing the estimated
policy. First, when the policy class is unconstrained, we show that
the estimated policy is consistent if either the bounds on the QoTE
are sign-determining or the QoTE is point-identified. Otherwise, the
leading term of the regret bound has a magnitude that depends on the
relative location of zero in the QoTE bounds. We provide the theory
for both stochastic and deterministic policies. The leading term with
the stochastic policy is smaller than that with the deterministic
policy, consistent with the findings in the literature \citep{manski2007minimax,stoye2007minimax,cui2021individualized}.
It is important to allow the policy class to be constrained as the
PM may prefer a parsimonious policy or face institutional or budget
constraints. In this case of constrained policy classes, we propose
to use the machine learning (ML) technique of the outcome-weighting
framework with a surrogate loss \citep{zhao2012estimating}. We
then show that the ML-estimated policy is consistent and characterize
the rate in terms of approximation and estimation errors.

In this paper, we consider empirical applications in two well-known
randomized control trials in medicine and economics. The first application
concerns the allocation of a diagnostic procedure for critically ill
patients using data from the Study to Understand Prognoses and Preferences
for Outcomes and Risks of Treatments \citep{hirano2001estimation}.
The second application examines the allocation of job training using
data from the US National Job Training Partnership Act \citep{bloom1997benefits}.
In both applications, a common finding is that there exists substantial
heterogeneity in the distributional treatment effects and thus in
the corresponding allocation decisions based on the QoTE. To deliver
the main messages of this paper, we show in the space of covariates
how the allocation decisions take place (see Figures \ref{fig:application1-2}--\ref{fig:application1-3}
below). As expected, the allocation becomes more aggressive as the
quantile probability $\tau$ increases. We compare this result with
the decisions based on the QTE and ATE. The QTE decisions do not exhibit
the change in the degree of prudence in $\tau$. Comparing the ATE
decisions with the QoTE decisions with $\tau=0.5$, we can inspect
whether outliers are problematic in calculating the ATE decision in
these data sets. In this sense, we view the QoTE decisions as a means
of a robustness check for the ATE decisions prevalent in the literature.

The policy learning framework of this paper can be generalized to
any setting where welfare is defined as a functional of the joint
distribution of potential outcomes. Towards the end of the paper,
we introduce a general framework and propose other examples of welfare
criteria that may be interest a non-utilitarian PM. These include
criteria targeting individuals who are either worst off in the counterfactual
baseline or worst-affected by the treatment.

\subsection{Related Literature}

Learning optimal treatment regimes has received considerable interest
in the past few years across multiple disciplines including computer
science \citep{dudik2011doubly}, econometrics \citep{manski2004statistical,hirano2009asymptotics,stoye2009minimax,kitagawa2018should,athey2021policy,mbakop2021model,ida2022choosing},
and statistics \citep{murphy2003optimal,kosorok2015adaptive,kosorok2019precision,tsiatis2019dynamic,jiang2019entropy}.
In statistics, existing methods for learning optimal treatment regimes
are mostly through either Q-learning \citep{watkins1992q,qian2011performance}
or A-learning \citep{murphy2003optimal,robins2004optimal,shi2018high}.
Alternative approaches have emerged from a classification perspective
\citep{zhao2012estimating,zhang2012robust,rubin2012statistical},
which has proven more robust to model misspecification in some settings.

Recently, there is a growing literature on learning optimal treatment
allocations that aims to relax the unconfoundedness assumption. Within
this literature, a strand of work considers cases where the welfare
and optimal treatment regime is point-identified, that is, the treatment
decision is free from ambiguity given the observed data. \citet{cui2021semiparametric,qiu2021optimal}
consider instrumental variable (IV) approaches under a point identification
and \citet{han2021comment,cui2021necessary} consider IV methods under
a sign identification. \citet{kallus2021causal,qi2023proximal,shen2023optimal}
consider optimal policy learning under the proximal causal inference
framework. Another strand of work considers robust policy learning
under ambiguity. \citet{kallus2021minimax} propose to learn an optimal
policy in the presence of partially identified treatment effects under
a sensitivity model. \citet{pu2021estimating} consider a minimax
regret policy for IV models under partial identification. \citet{cui2021individualized}
and \citet{d2021orthogonal} consider a variety of decision rules
in general settings where treatment effects are partially identified.
\citet{stoye2012minimax,yata2021optimal} develop finite-sample minimax
regret rules under partial identification of welfare. Moreover, \citet{han2023optimal}
proposes optimal dynamic treatment regimes through a partial welfare
ordering when the sequential randomization assumption is violated.
Policy learning under ambiguity is not limited to confounded settings.
There are other settings of robust decisions under ambiguity, for
example, when the treatment positivity assumption is violated \citep{ben2021safe},
when data sets are aggregated in meta analyses \citep{ishihara2021evidence}
and when the target population is shifted from the experiment population
\citep{adjaho2022externally}. The present paper contributes to this
literature of model ambiguity by considering a distributional welfare
that is partially identified.

There is also work focused on policy learning based on distributional
properties under point identification. \citet{leqi2021median} consider
the QTE as a criterion and \citet{qi2023robustness} consider maximizing
the average outcomes that are below a certain quantile. \citet{wang2018quantile,linn2017interactive}
consider maximizing the quantile of global welfare, which can be viewed
as a special case of \citet{kitagawa2021equality}. The latter study
considers estimating the optimal treatment allocation based on individual
characteristics when the objective is to maximize an equality-minded
rank-dependent welfare function, which essentially puts higher weights
on individuals with lower-ranked outcomes. Our work complements this
line of literature by introducing a different type of distributional
welfare using the distribution of treatment effects and proposing
decision-making under ambiguity. Further comparisons to this line
of work are made in Section \ref{sec:Treatment-Rules-and}. \citet{kock2022functional,kock2023treatment,kock2024regularizing} consider choosing an optimal treatment among a discrete set of treatments with distributional targets. Finally, \citet{manski2023statistical,kitagawa2023treatment} consider a distribution
or nonlinear function of regret and establish admissible treatment
rules within that framework. Although our welfare has distributional
aspects, when showing the theoretical guarantee of the estimated rules,
we use the standard the notion of the (mean) regret.

\subsection{Organization of the Paper}

The paper is organized as follows. The next section formally introduces
our welfare criterion and compare it with criteria previously considered
in the literature. Then the minimax framework is proposed. Section
\ref{sec:Possible-Identifying-Assumptions} provides identifying assumptions
that can be used to narrow the bounds on the QoTE. Section \ref{sec:Theoretical-Properties-of}
presents the theoretical properties of the estimated policies for
constrained and unconstrained policy classes. Section \ref{sec:Calculating-Bounds}
discusses how to systematically calculate the bounds on the QoTE using
linear programming. Section \ref{sec:Empirical-Applications} presents
the two empirical applications. Finally, Section \ref{sec:Extensions}
concludes the paper by generalizing the paper's framework to other
related non-utilitarian welfare criteria. In the Supplemental Appendix,
Section A lists further identifying assumptions for tightening bounds
on the QoTE. Section B presents an additional empirical application,
and Section C contains numerical exercises. Section D further discusses
stochastic rules and Section E contains all proofs.

\section{Treatment Rules and Distributional Welfare\label{sec:Treatment-Rules-and}}

Let $Y\in\mathcal{Y}$ be the outcome, $X\in\mathcal{X}$ be covariates,
and $D\in\{0,1\}$ be binary treatment in respective supports. Let
$Y_{d}$ be the potential outcome that is consistent with the observed
outcome, that is, $Y=DY_{1}+(1-D)Y_{0}$. We define a treatment allocation
rule, or equivalently a \emph{policy}, as $\delta:\mathcal{X}\rightarrow\mathcal{A}\subseteq[0,1]$
where $\mathcal{A}$ is the action space. A deterministic rule corresponds
to $\mathcal{A}=\{0,1\}$ and a stochastic rule corresponds to $\mathcal{A}=[0,1]$.
Unless noted otherwise, we allow both in our general framework. Let
$\delta\in\mathcal{D}$ where $\mathcal{D}$ is the (potentially constrained)
space of $\delta$. For the allocation problem, a policymaker (PM)
would set an objective function that she maximizes to find the optimal
allocation rule.

\subsection{Introducing Distributional Welfare\label{subsec:Introducing-Distributional-Welfa}}

To motivate the objective function we propose, we first review the
most common objective function considered in the literature: the average
welfare.\footnote{Welfare is sometimes called a value function in the literature.}
The optimal policy under this welfare criterion can be defined as
$\delta_{ATE}^{*}\in\arg\max_{\delta\in\mathcal{D}}E[\delta(X)Y_{1}+(1-\delta(X))Y_{0}]$.
With deterministic rules in particular, the welfare can be written
as $E[\delta(X)Y_{1}+(1-\delta(X))Y_{0}]=E[Y_{\delta(X)}]$. See Section
E in the Appendix that shows how $E[\delta(X)Y_{1}+(1-\delta(X))Y_{0}]$
(and other welfare criteria appearing below) is compatible with stochastic
rules. Because $E[\delta(X)Y_{1}+(1-\delta(X))Y_{0}]=E[Y_{0}+\delta(X)(Y_{1}-Y_{0})]=E[Y_{0}]+E[\delta(X)E[Y_{1}-Y_{0}|X]],$
$\delta_{ATE}^{*}$ also satisfies
\begin{align}
\delta_{ATE}^{*} & \in\arg\max_{\delta\in\mathcal{D}}E[\delta(X)E[Y_{1}-Y_{0}|X]],\label{eq:mean_welfare}
\end{align}
where the objective function corresponds to the \emph{welfare gain}.
Therefore, subject to the constraints, $\delta_{ATE}^{*}$ maximizes
the average of conditional average treatment effects (ATEs) either
chosen (in the case of deterministic policies) or weighted (in the
case of stochastic policies) by $\delta$, thus the notation ``$\delta_{ATE}^{*}$.''
For example, when $\mathcal{D}$ is not constrained, $\delta_{ATE}^{*}(x)=1\{E[Y_{1}-Y_{0}|X=x]\ge0\}$
for both deterministic and stochastic policies. In general, the formulation
\eqref{eq:mean_welfare} reveals an important fact: the conditional
treatment effect is the important basis for the policy choice. This
makes sense because the treatment should be allocated to those who
would benefit the most from it. This idea becomes important in introducing
our distributional welfare later.

Although it is the most common form of welfare, the average welfare
is obviously sensitive to outliers. For example, a small share of
individuals with $X=x$ and substantially large $Y_{1}-Y_{0}$ can
make $E[Y_{1}-Y_{0}|X=x]$ positive, suggesting to treat \emph{all}
individuals with $X=x$ even though the majority suffers from receiving
the treatment. This can be especially problematic when the distribution
of $Y_{1}-Y_{0}|X=x$ is skewed and heavy-tailed. This motivates us
to alternatively consider the quantile of individual treatment effects
$Y_{1}-Y_{0}$ (QoTE) as the basis for a welfare criterion (analogous
to \eqref{eq:mean_welfare}) and a corresponding optimal policy. Let
$Q_{\tau}(Y)\equiv\inf\{y:F_{Y}(y)\ge\tau\}$ be the $\tau$-quantile
of $Y$ and $Q_{\tau}(Y|X)\equiv\inf\{y:F_{Y|X}(y)\ge\tau\}$ be the
$\tau$-quantile of $Y$ conditional on $X$. We consider an optimal
policy that satisfies
\begin{align}
\delta^{*}\equiv\delta_{\tau}^{*} & \in\arg\max_{\delta\in\mathcal{D}}E[\delta(X)Q_{\tau}(Y_{1}-Y_{0}|X)],\label{eq:quantile_welfare}
\end{align}
where $Q_{\tau}(Y_{1}-Y_{0}|X)$ is the $\tau$-quantile of $Y_{1}-Y_{0}$
given $X$. That is, $\delta^{*}$ maximizes the average of conditional
QoTEs chosen (in the case of deterministic policies) or weighted (in
the case of stochastic policies) by $\delta$. With no constraint
in $\mathcal{D}$, $\delta^{*}(x)=1\{Q_{\tau}(Y_{1}-Y_{0}|X=x)\ge0\}$
for both deterministic and stochastic policies. The QoTE is less sensitive
to outliers than the ATE, so for example \eqref{eq:quantile_welfare}
with $\tau=0.5$ may be preferred to \eqref{eq:mean_welfare}. This
aspect makes the allocation decision within the $X=x$ group not driven
by treatment effects of a small share of individuals. In this sense,
this aspect of robustness can be viewed as the ``within-group robustness''
\citep{leqi2021median}. In general, $\tau$ (i.e., the rank in
individual treatment effects) represents individuals in that specific
quantile as a \emph{reference group} chosen by the PM. For example,
by choosing low $\tau$, the PM allocates the treatment only if most
individuals benefit from it because $Q_{\tau'}(Y_{1}-Y_{0}|X)\ge Q_{\tau}(Y_{1}-Y_{0}|X)$
for any $\tau'>\tau$. In other words, she ensures that disadvantaged
individuals with poor treatment effects are not harmed from receiving
the allocation. In this sense, low $\tau$ corresponds to a \emph{prudent
PM}. On the other hand, by choosing high $\tau$, the PM focuses on
benefiting solely the top-ranked individuals even though the majority
would suffer from it. In this sense, high $\tau$ corresponds to a
\emph{negligent PM}. Therefore, the choice of $\tau$ reflects the
level of prudence of the policy that the PM commits to.

The proposed optimal policy has another interesting interpretation
that relates to the PM's incentive. Let $\delta_{\tau}^{\dagger}\equiv1\{Q_{\tau}(Y_{1}-Y_{0}|X)\ge0\}\in\arg\max_{\delta:\mathcal{X\rightarrow\mathcal{A}}}E[\delta(X)Q_{\tau}(Y_{1}-Y_{0}|X)]$
be the first-best rule for $\mathcal{A}$ being \emph{either} $[0,1]$
or $\{0,1\}$. As mentioned above, $\delta_{\tau}^{\dagger}$ is an
optimal rule when no restriction is imposed on the class of $\delta$.
Suppose individuals who benefit from the treatment would vote for
it. Also suppose $\tau=0.5$. Then $\delta_{0.5}^{\dagger}(X)=1\{Q_{0.5}(Y_{1}-Y_{0}|X)\ge0\}$
can be viewed as a policy that obeys \emph{majority vote}. To see
this, note the following is true for continuously distributed $Y_{d}$:
$Q_{0.5}(Y_{1}-Y_{0}|X)\ge0$ if and only if $P[Y_{1}\ge Y_{0}|X]\ge P[Y_{1}<Y_{0}|X]$.
Therefore, the distributional welfare criterion \eqref{eq:quantile_welfare}
is consistent with a PM who has political incentive and whose decision
is influenced by vote shares. This interpretation can be generalized
by considering $Q_{0.5-\alpha/2}(Y_{1}-Y_{0}|X)\ge0$ for $0\le\alpha\le1$,
which is equivalent to $P[Y_{1}\ge Y_{0}|X]\ge P[Y_{1}<Y_{0}|X]+\alpha$
where $\alpha$ can be viewed as the vote share margin.

Exploring this interpretation further, we can show that the first-best
policy for the median can be viewed as the one that maximizes the
share of positively affected individuals or the correct classification
rate over a class of deterministic policies:

\begin{theorem}\label{thm:interpretation}Suppose $Y_{d}$ is continuously
distributed and $\mathcal{A}$ is either $[0,1]$ or $\{0,1\}$. Then,
the first best rule $\delta_{\tau}^{\dagger}(x)\equiv1\{Q_{\tau}(x)\ge0\}$
for $\tau=0.5$ satisfies
\begin{align}
\delta_{0.5}^{\dagger}\in\arg\max_{\delta:\mathcal{X}\rightarrow\mathcal{A}}E[\delta(X)Q_{0.5}(X)] & =\arg\max_{\delta:\mathcal{X}\rightarrow\{0,1\}}P\left[Y_{\delta(X)}-Y_{1-\delta(X)}>0\right]\label{eq:interpretation1}\\
 & =\arg\max_{\delta:\mathcal{X}\rightarrow\{0,1\}}P\left[\delta(X)\in\arg\max_{d}Y_{d}\right].\label{eq:interpretation2}
\end{align}
\end{theorem}

In the theorem, \eqref{eq:interpretation1} holds by the equivalence
result in the previous paragraph and \eqref{eq:interpretation2} is
immediate. Note that $P\left[\delta(X)\in\arg\max_{d}Y_{d}\right]$
is the correct classification rate. We can equivalently say that $\delta_{0.5}^{\dagger}$
minimizes the \emph{fraction negatively affected} by switching from
$1-\delta$ to $\delta$, namely, $P[Y_{\delta(X)}-Y_{1-\delta(X)}<0]$,
or the misclassification rate, $P\left[\delta(X)\notin\arg\max_{d}Y_{d}\right]$.
The latter extends \citet{kallus2022s}'s definition which focuses
on binary $Y_{d}$.

Related to the proposed welfare criterion, one can consider alternative
criteria that are robust to outliers. Focusing on a deterministic
policy (i.e., $\mathcal{A}=\{0,1\}$), \citet{wang2018quantile} consider
the marginal quantile of $Y_{\delta(X)}$ as their criterion, while
\citet{leqi2021median} focus on the average of conditional quantile
$Y_{\delta(X)}$. First, \citet{wang2018quantile} explore the optimal
policy under $Q_{\tau}(Y_{\delta(X)})$, which can be viewed as a
sensible quantity robust to outliers. Note that the randomness in
$Y_{\delta(X)}$ arises from both $Y_{d}$ and $X$. Because of that,
the optimal policy under $Q_{\tau}(Y_{\delta(X)})$ does not have
a closed form solution, which make the interpretation of the optimal
policy somewhat elusive. Moreover, \citet{leqi2021median} demonstrate
that the policy under this welfare criterion lacks ``across-group
fairness,'' in that the allocation decision for one group (defined
by $X=x$) can be influenced by the treatment effects of other groups
(defined by other $X=x'$). This issue stems from the difficulty in
associating the objective function $Q_{\tau}(Y_{\delta(X)})$ with
a clear notion of treatment effects or gains, unlike the other criteria
discussed in this section. To overcome this issue, \citet{leqi2021median}
consider the optimal policy under $E[Q_{\tau}(Y_{\delta(X)}|X)]$,
which achieves across-group fairness as $X$ is fixed in the calculation
of quantile. They show the optimal policy also satisfies $\delta_{QTE}^{*}\in E[\delta(X)\{Q_{\tau}(Y_{1}|X)-Q_{\tau}(Y_{0}|X)\}].$
That is, $\delta_{QTE}^{*}$ maximizes the average of conditional
QTEs chosen by $\delta$. However, a PM may find the allocation decision
based on the QTE undesirable because the individual at the $\tau$-quantile
of $Y_{1}$ may not be the same individual as the one at the $\tau$-quantile
of $Y_{0}$. Since introduced in \citet{doksum1974empirical} and
\citet{lehmann1975statistical}, the QTE has been a popular causal
parameter. However, its limitation is also acknowledged in the literature,
which seems more pronounced in the context of treatment allocation.
This aspect implies that it is difficult for the PM to aim the level
of prudence (e.g., to be conservative) as there is no clear notion
of a negligence or prudence associated with the level of $\tau$;
see Figure \ref{fig:application1-3} in the application (Section \ref{sec:Empirical-Applications})
for related discussions.

\subsection{Policies Robust to Model Ambiguity\label{subsec:Policies-Robust-to}}

Despite the desirable properties of our proposed objective function,
the main challenge of using \eqref{eq:quantile_welfare} as the welfare
criterion is that the QoTE is generally not point-identified even
under unconfoundedness. This is because, in general, the QoTE is not
equal to the QTE and, while the latter can be identified from
the marginal distributions of $Y_{1}$ and $Y_{0}$, the former can
only be identified from the joint distribution of $(Y_{1},Y_{0})$.
Therefore, we propose optimal policies that are robust to this ambiguity.
One may consider maximizing the worst-case gain: $\delta_{mmw}^{*}\in\arg\max_{\delta\in\mathcal{D}}\min_{F_{Y_{1},Y_{0}|X}\in\mathcal{F}}E[\delta(X)Q_{\tau}(Y_{1}-Y_{0}|X)],$
where $F_{Y_{1},Y_{0}|X}$ is the joint distribution of $(Y_{1},Y_{0})$
conditional on $X$ and $\mathcal{F}\equiv\mathcal{F}(P)$ is the
identified set of $F_{Y_{1},Y_{0}|X}$ given the data $P$. However,
this criterion is known to be overly pessimistic \citep{savage1951theory}.
Therefore, one may instead consider minimizing the worst-case regret:
\begin{align}
\delta_{mmr}^{*} & \in\arg\min_{\delta\in\mathcal{D}}\max_{F_{Y_{1},Y_{0}|X}\in\mathcal{F}}E[\{\delta^{\dagger}(X)-\delta(X)\}Q_{\tau}(Y_{1}-Y_{0}|X)],\label{eq:mmr}
\end{align}
where $\delta^{\dagger}\equiv\delta_{\tau}^{\dagger}\equiv1\{Q_{\tau}(Y_{1}-Y_{0}|X)\ge0\}$
is the first-best rule. The minimax regret criterion is free from
priors and thus avoids the feature of maximin mentioned above. The
two criteria becomes identical under point identification (i.e., when
$\mathcal{F}(P)$ is a singleton). Therefore, our primary focus is
the minimax policy.

For each $x$, define the identified interval for $Q_{\tau}(Y_{1}-Y_{0}|X=x)$
as
\begin{align*}
[Q_{\tau}^{L}(x),Q_{\tau}^{U}(x)] & =\{Q_{\tau}(Y_{1}-Y_{0}|X=x):F_{Y_{1},Y_{0}|X}\in\mathcal{F}\}.
\end{align*}
Using these lower and upper bounds, we can derive closed-form expressions
for the inner optimization in \eqref{eq:mmr} (and similarly in the
objective function for $\delta_{mmw}^{*}$). To this end, we impose
a very weak assumption on the identified interval.

\begin{myas}{RC}\label{as:RC}The identified set $\mathcal{Q}(P)$
of $Q_{\tau}(Y_{1}-Y_{0}|X=\cdot)$ is rectangular, that is, $\mathcal{Q}(P) =\{Q_{\tau}(Y_{1}-Y_{0}|X=\cdot):Q_{\tau}(Y_{1}-Y_{0}|X=x)\in[Q_{\tau}^{L}(x),Q_{\tau}^{U}(x)]\}$.
\end{myas}

This assumption holds for the identified sets we derive in this paper.
It will be violated if one imposes certain shape restrictions on $Q_{\tau}(Y_{1}-Y_{0}|X=\cdot)$
such as monotonicity. We do not consider shape restrictions in this
paper as allowing for unrestricted heterogeneity across $X$ is important
in the context of optimal allocations. Essentially, this assumption
allows us to interchange the maximum or minimum over $\mathcal{F}$
with the expectation over $X$ \citep{kasy2016partial,d2021orthogonal}.\footnote{To illustrate this, consider a simple case of binary $X\in\{0,1\}$
and let $Q_{\tau}(x)\equiv Q_{\tau}(Y_{1}-Y_{0}|X=x)$ and $p_{x}\equiv P[X=x]$.
Then Assumption \ref{as:RC} imposes that $\{(Q_{\tau}(0),Q_{\tau}(1)):Q_{\tau}(x)\in[Q_{\tau}^{L}(x),Q_{\tau}^{U}(x)],x\in\{0,1\}\}$
is rectangular, which implies that, for example,
\begin{align*} \min_{F_{Y_{1},Y_{0}|X}}E[\delta(X)Q_{\tau}(X)] & =p_{1}\delta(1)\min_{F_{Y_{1},Y_{0}|X}}Q_{\tau}(1)+p_{0}\delta(0)\min_{F_{Y_{1},Y_{0}|X}}Q_{\tau}(0)=E[\delta(X)\min_{F_{Y_{1},Y_{0}|X}}Q_{\tau}(X)].
\end{align*}
} Under Assumption \ref{as:RC}, we can easily show that $\delta_{mmr}^{*}$
equivalently satisfies
\begin{align}
\delta_{mmr}^{*} & \in\arg\max_{\delta\in\mathcal{D}}E[\delta(X)\bar{Q}_{\tau}(X)],\label{eq:mmr2}
\end{align}
where $\bar{Q}_{\tau}(x)=Q_{\tau}^{U}(x)1\{Q_{\tau}^{L}(x)\ge0\}+Q_{\tau}^{L}(x)1\{Q_{\tau}^{U}(x)\le0\}+\left(Q_{\tau}^{U}(x)+Q_{\tau}^{L}(x)\right)1\{Q_{\tau}^{L}(x)<0<Q_{\tau}^{U}(x)\}.$
Also, we can show $\delta_{mmw}^{*}\in\arg\max_{\delta\in\mathcal{D}}E[\delta(X)Q_{\tau}^{L}(X)].$
In general, finding the optimal $\delta$ for \eqref{eq:mmr2} does
not yield a closed-form expression when the policy class $\mathcal{D}$
is constrained. Additionally, solving $\max_{\delta\in\mathcal{D}}E[\delta(X)\bar{Q}_{\tau}(X)]$
proves to be a challenging task as $\bar{Q}_{\tau}(\cdot)$ incorporates
an indicator function. Nonetheless, allowing the policy class to be
constrained is important because the PM may prefer a more parsimonious
rule (e.g., a linear rule) or be limited by certain institutional
constraints. Following \citet{zhao2012estimating}, we consider a
convex and continuous relaxation of \eqref{eq:mmr2} by utilizing
the hinge loss function $\phi(t)=\max(1-t,0)$ and introducing a regularization
term. This is done in Section \ref{subsec:Regret-Bounds-with-1} below.
The consistency of hinge loss is shown even when the class of $\delta$
is restricted \citep{kitagawa2021constrained}.

\section{Possible Identifying Assumptions\label{sec:Possible-Identifying-Assumptions}}

We now provide identifying assumptions that researchers may want to
consider imposing to shrink $\mathcal{F}$ (i.e., the identified set
for the joint distribution of $(Y_{1},Y_{0})$ conditional on $X$)
and thus $[Q_{\tau}^{L}(x),Q_{\tau}^{U}(x)]$. First, there are ways
to identify the marginal distribution of $Y_{d}$. The most obvious
approach is to impose conditional independence.

\begin{myas}{CI}[Conditional Independence]\label{as:CI}For $d\in\{0,1\}$,
$Y_{d}\perp D|X$.\end{myas}

Alternative to Assumption \ref{as:CI}, local copula modeling \citep{chernozhukov2024estimating}
or panel quantile regression models \citep{chernozhukov2013average}
can be used to identify $Q_{\tau}(Y_{d}|X)$. Given the identification
of the marginal distribution of $Y_{d}$, the \emph{Makarov bounds}
\citep{fan2010sharp} can be derived (see Section A), although they
tend to be uninformative. We now consider identifying assumptions
that can be used to yield tighter bounds, leading to more informative
decisions.

\begin{myas}{PD}[Positive Dependence]\label{as:PD}For $x\in\mathcal{X}$,
either (i) $P[Y_{1}\le y_{1},Y_{0}\le y_{0}|X=x]\le P[Y_{1}\le y_{1}|X=x]P[Y_{0}\le y_{0}|X=x]$,
(ii) $P[Y_{1}>y_{1}|Y_{0}>\cdot,X=x]$ and $P[Y_{0}>y_{0}|Y_{1}>\cdot,X=x]$
are non-decreasing and $P[Y_{1}\le y_{1}|Y_{0}\le\cdot,X=x]$ and
$P[Y_{0}\le y_{0}|Y_{1}\le\cdot,X=x]$ are non-increasing, or (iii)
$P[Y_{1}>y_{1}|Y_{0}=\cdot,X=x]$ and $P[Y_{0}>y_{0}|Y_{1}=\cdot,X=x]$
are non-decreasing, for all $y_{1},y_{0}\in\mathcal{Y}$.\end{myas}

Assumption \ref{as:PD} imposes various versions of positive dependence
between $Y_{1}$ and $Y_{0}$. This assumption makes sense when individuals
with high $Y_{1}$ (e.g., potential health with the treatment) tend
to have high $Y_{0}$ (e.g., potential health without the treatment)
and vice versa. \ref{as:PD}(iii), called \emph{stochastic increasingness} (SI), implies (ii), and (ii)
implies (i) \citep{joe2014dependence}. The following lemma provides a model for $Y_d$ that implies \ref{as:PD}:
\begin{lemma}
    Assumption \ref{as:PD} holds if $Y_{d}=g(d,X,U)$
where $g(d,x,\cdot)$ is non-decreasing.
\end{lemma} In the previous example, $U$ may capture
underlying health conditions. The model assumption in the lemma trivially holds
with additively separable $U$ common in regression specification,
although it is substantially weaker than that. Due to its plausibility,
we consider this assumption as our leading one in later analyses. Maintaining Assumption \ref{as:CI},
Assumption \ref{as:PD} is helpful to obtain more informative bounds
on the conditional QoTE. For example, \citet{frandsen2021partial}
derive bounds on the distribution of treatment effects under an unconditional
version of \ref{as:PD}(iii). Instead of assuming positive dependence
between $Y_{1}$ and $Y_{0}$, one may want to impose stochastic dominance
of $Y_{d}$ between treatment and control groups or stochastic dominance
between $Y_{1}$ and $Y_{0}$ for each subgroup:

\begin{myas}{SD}[Stochastic Dominance]\label{as:SD}For $x\in\mathcal{X}$,
either (i) $P[Y_{d}\le y|D=1,X=x]\le P[Y_{d}\le y|D=0,X=x]$, or (ii)
$P[Y_{1}\le y|D=d,X=x]\le P[Y_{0}\le y|D=d,X=x]$.\end{myas}

Either under Assumption \ref{as:CI} or the existence of instrumental
variables (IVs), Assumption \ref{as:SD}(i) or \ref{as:SD}(ii) can
be used to narrow the bounds on the distribution of treatment effects
\citep{blundell2007changes,lee2021partial} and thus on
the QoTE.

Next, we present assumptions that help point-identify the conditional
QoTE. The following assumption is a special instance of Assumption
\ref{as:PD}.

\begin{myas}{RI}[Rank Invariance]\label{as:RI}For $d\in\{0,1\}$,
$Y_{d}=m_{d}(X,U_{d})$ where $m_{d}(x,\cdot)$ is strictly increasing
and $U_{d}|X=x$ is absolutely continuous and satisfies $U_{1}|_{X=x}=U_{0}|_{X=x}$
for given $x\in\mathcal{X}$. \end{myas}

\citet{heckman1997making} and \citet{chernozhukov2005iv} show the
identifying power of Assumption \ref{as:RI}. This assumption essentially
restricts heterogeneity by holding the ranks in $Y_{1}$ and $Y_{0}$
the same. This implies that, under this assumption, the QTE can be
interpreted as the difference between $Y_{1}$ and $Y_{0}$ for the
same individual. Yet, the QTE is \emph{not} identical to the OoTE
even under this assumption. Moreover, Assumption \ref{as:RI} implies
Assumption \ref{as:PD} because, suppressing $X$, $P[Y_{1}\le y_{1}|Y_{0}=y_{0}]=P[m_{1}(U)\le y_{1}|m_{0}(U)=y_{0}]=P[U\le m_{1}^{-1}(y_{1})|U=m_{0}^{-1}(y_{0})]$
and thus the probability is 1 when $y_{0}\le m_{0}(m_{1}^{-1}(y_{1}))$
and 0 otherwise. Under Assumptions \ref{as:CI} and \ref{as:RI},
$F_{\Delta|X}(\delta)$ is point identified. More generally, \citet{heckman1997making}
consider Markov kernels $M$ and $\tilde{M}$ so that $F_{Y_{1}|X}(y_{1})=\int M(y_{1},y_{0}|X)dF_{Y_{0}|X}(y_{0})$
and $F_{Y_{0}|X}(y_{0})=\int\tilde{M}(y_{1},y_{0}|X)dF_{Y_{1}|X}(y_{1})$.
Also see \citet{vuong2017counterfactual} for the case of endogenous
treatment with IVs. \citet[Section 2.5.3]{abbring2007econometric}
also consider perfect \emph{negative} dependence.\footnote{This corresponds to $U_{1}|_{X=x}=-U_{0}|_{X=x}$.}
Section A of the Supplemental Appendix contains an extended list of
assumptions for point identification, which includes deconvolution,
symmetry, and Roy models.

\section{Theoretical Properties of Estimated Policy\label{sec:Theoretical-Properties-of}}

Henceforth, let $Q_{\tau}(X)\equiv Q_{\tau}(Y_{1}-Y_{0}|X)$ for notational
simplicity. Focusing on the optimal policy $\delta_{mmr}^{*}$ based
on the minimax criterion, we provide theoretical guarantees for the
estimated policy. The policy can be readily estimated once the bounds
$[Q_{\tau}^{L}(X),Q_{\tau}^{U}(X)]$ on $Q_{\tau}(X)$ are estimated
using parametric or nonparametric methods with the sample of $(Y,D,X)$.
The theory includes the case of point identification as a special
case in which $Q_{\tau}(X)=Q_{\tau}^{L}(X)=Q_{\tau}^{U}(X)$.

Recall that our objective function is $V(\delta)\equiv E[\delta(X)Q_{\tau}(X)].$
To define the regret, we introduce a r.v. $A(x)$ that is distributed
as $Bernoulli(\delta(x))$. For a stochastic policy $\delta(x)\in[0,1]$,
$\delta(x)$ is the probability that $A(x)=1$. For a deterministic
policy $\delta(x)\in\{0,1\}$, the distribution of $A(x)$ is degenerate
and thus $A(x)=\delta(x)$. Define the regret of the ``classification''
as $R(\delta)\equiv V(\delta^{\dagger})-V(\delta)=E[|Q_{\tau}(X)|1\{A(X)\neq sign(Q_{\tau}(X))\}],$
where $\delta^{\dagger}(X)=1\{Q_{\tau}(X)\ge0\}$ and $sign(q)=1$
when $q\geq0$ and $sign(q)=0$ when $q<0$. Note that $R(\delta)$
is generally not point-identified and thus we define maximum regret
as $\bar{R}(\delta) \equiv\max_{Q_{\tau}(\cdot)\in[Q_{\tau}^{L}(\cdot),Q_{\tau}^{U}(\cdot)]}E[|Q_{\tau}(X)|1\{A(X)\neq sign(Q_{\tau}(X))\}]$. The maximum regret can be expressed in different ways, which are useful
in the analysis below.

\begin{lemma}\label{lem:max_regret}Suppose Assumption \ref{as:RC}
hold. For a stochastic or deterministic rule $\delta$, the maximum
regret can be expressed as
\begin{align}
\bar{R}(\delta) & =E\left[\max\left\{ [1-\delta(X)]\max(Q_{\tau}^{U}(X),0),\delta(X)\max(-Q_{\tau}^{L}(X),0)\right\} \right]\label{eq:max_regret1}\\
 & =-E[\delta(X)\bar{Q}_{\tau}(X)]+E\left[Q_{\tau}^{U}(X)1\{Q_{\tau}^{U}(X)\ge0\}\right]\label{eq:max_regret2}\\
 & =E[|\bar{Q}_{\tau}(X)|1\{A(X)\neq sign(\bar{Q}_{\tau}(X))\}]\nonumber \\
 & \qquad+E\bigg[\min(Q_{\tau}^{U}(X),-Q_{\tau}^{L}(X))1\{Q_{\tau}^{L}(X)<0<Q_{\tau}^{U}(X)\}\bigg].\label{eq:max_regret3}
\end{align}
\end{lemma}Note
that \eqref{eq:max_regret2} is used in expressing \eqref{eq:mmr2}.
Below, \eqref{eq:max_regret1} is used in Section \ref{subsec:Regret-Bounds-with}
and \eqref{eq:max_regret3} in Section \ref{subsec:Regret-Bounds-with-1}.
Now, we provide asymptotic bounds on these regrets evaluated at the
estimated stochastic and deterministic policies when $\mathcal{D}$
is unconstrained and constrained.

\subsection{Regret Bounds with Unconstrained Policy Class\label{subsec:Regret-Bounds-with}}

We assume that we are equipped with the consistent estimators for
$Q_{\tau}^{L}(X)$ and $Q_{\tau}^{U}(X)$.

\begin{myas}{EST}\label{as:EST}$Q_{\tau}(X)$ is bounded almost
surely, and $\hat{Q}_{\tau}^{L}(X)-Q_{\tau}^{L}(X)=o_{p}(1)$ and
$\hat{Q}_{\tau}^{U}(X)-Q_{\tau}^{U}(X)=o_{p}(1)$.\end{myas}

When $Q_{\tau}^{L}(X)$ and $Q_{\tau}^{U}(X)$ are known functions
of $F_{Y_{1}|X}$ and $F_{Y_{0}|X}$, Assumption \ref{as:EST} is
implied from the consistency of $\hat{F}_{Y_{1}|X}$ and $\hat{F}_{Y_{0}|X}$
by the continuous mapping theorem; see Section \ref{sec:Calculating-Bounds}
for the case of bounds that are computationally derived. Let $\delta^{*,stoch}\equiv\delta_{mmr}^{*,stoch}$
and $\delta^{*,determ}\equiv\delta_{mmr}^{*,determ}$ are the optimal
policies that minimize $\bar{R}(\delta)$ when $\delta$ is stochastic
and deterministic policies, respectively. Given the expression \eqref{eq:max_regret1},
a simple calculation yields
\begin{align*}
\delta^{*,stoch}(x) & =\begin{cases}
1 & \text{if }Q_{\tau}^{L}(x)\ge0,\\
0 & \text{if }Q_{\tau}^{U}(x)\le0,\\
\frac{Q_{\tau}^{U}(x)}{Q_{\tau}^{U}(x)-Q_{\tau}^{L}(x)} & \text{if }Q_{\tau}^{L}(x)<0<Q_{\tau}^{U}(x),
\end{cases}
\end{align*}
and
\begin{align*}
\delta^{*,determ}(x) & =\begin{cases}
1 & \text{if }Q_{\tau}^{L}(x)\ge0, \text{ or } \{Q_{\tau}^{L}(x)<0<Q_{\tau}^{U}(x)\text{ and }|Q_{\tau}^{L}(x)|<|Q_{\tau}^{U}(x)|\},\\
0 & \text{if }Q_{\tau}^{U}(x)\le0, \text{ or } \{Q_{\tau}^{L}(x)<0<Q_{\tau}^{U}(x)\text{ and }|Q_{\tau}^{L}(x)|>|Q_{\tau}^{U}(x)|\}.
\end{cases}
\end{align*}
Let $\hat{\delta}^{stoch}$ and $\hat{\delta}^{determ}$ are the estimates
of $\delta^{*,stoch}$ and $\delta^{*,determ}$, respectively.

\begin{theorem}\label{thm:regret_bound} Suppose Assumptions \ref{as:RC}
and \ref{as:EST} hold and $|Y|\le M$ for some constant $M$. The
regret of $\hat{\delta}^{stoch}$ is bounded by
\begin{align*}
R(\hat{\delta}^{stoch}) & \leq E\left[\frac{Q_{\tau}^{L}(X)Q_{\tau}^{U}(X)}{Q_{\tau}^{L}(X)-Q_{\tau}^{U}(X)}1\{Q_{\tau}^{L}(X)<0<Q_{\tau}^{U}(X)\}\right]+o_{p}(1),
\end{align*}
where the ratio is defined to be $0$ whenever its denominator is
$0$. The regret of $\hat{\delta}^{determ}$ is bounded by
\begin{align*}
R(\hat{\delta}^{determ}) & \leq E\left[\min(\max(Q_{\tau}^{U}(X),0),\max(-Q_{\tau}^{L}(X),0))\right]+o_{p}(1).
\end{align*}
\end{theorem}

The proof of this theorem and all other proofs are collected in the
appendix. The leading term in each asymptotic regret bound collapses
to zero when either (i) the bounds on $Q_{\tau}(X)$ exclude zero
almost surely or (ii) $Q_{\tau}(X)$ is point-identified. These are
the situations in which we can identify the sign of $Q_{\tau}(X)$.
Recalling $\delta^{\dagger}(X)=1\{Q_{\tau}(X)\ge0\}$, this is enough
to achieve consistency $R\rightarrow0$ as the second term in each
regret bound is the sampling error. In general, the leading term becomes
larger as the endpoints $[Q_{\tau}^{L}(X),Q_{\tau}^{U}(X)]$ are farther
away from zero, which is intuitive. Finally, the leading term with
the stochastic rule ($\frac{Q_{\tau}^{L}(X)Q_{\tau}^{U}(X)}{Q_{\tau}^{L}(X)-Q_{\tau}^{U}(X)}$)
is weakly smaller than that with the deterministic rule ($\min(\max(Q_{\tau}^{U}(X),0),\max(-Q_{\tau}^{L}(X),0))$),
suggesting that the stochastic rule is more preferred when there is
model ambiguity. This is consistent with the findings in the literature
\citep{manski2007minimax,stoye2007minimax,cui2021individualized}.

An immediate corollary of Theorem \ref{thm:regret_bound} establishes
the bound for the regret averaged over the sample of estimated policies.
Let $\mathbb{E}_{n}$ denote the expectation over the sample of $(Y,D,X)$.

\begin{corollary}Suppose Assumptions \ref{as:RC} and \ref{as:EST}
hold. Then,
\begin{align*}
\mathbb{E}_{n}\left[R(\hat{\delta}^{stoch})\right] & \leq E\left[\frac{Q_{\tau}^{L}(X)Q_{\tau}^{U}(X)}{Q_{\tau}^{L}(X)-Q_{\tau}^{U}(X)}1\{Q_{\tau}^{L}(X)<0<Q_{\tau}^{U}(X)\}\right]+o_{p}(1),
\end{align*}
where the ratio is defined to be $0$ whenever its denominator is
$0$, and
\begin{align*}
\mathbb{E}_{n}\left[R(\hat{\delta}^{determ})\right] & \leq E\left[\min(\max(Q_{\tau}^{U}(X),0),\max(-Q_{\tau}^{L}(X),0))\right]+o_{p}(1).
\end{align*}
\end{corollary}

\subsection{Regret Bounds with Constrained Policy Class\label{subsec:Regret-Bounds-with-1}}

As mentioned, allowing for a constrained policy class is crucial for
practical and institutional reasons. Our proposed method readily extends
to a scenario in which the policy class $\mathcal{D}$ is constrained.
Define the estimator of $\bar{Q}_{\tau}(\cdot)$ as $\widehat{\bar{Q}}_{\tau}(X)\equiv\hat{Q}_{\tau}^{U}(X)1\{\hat{Q}_{\tau}^{U}(X)\ge0\}+\hat{Q}_{\tau}^{L}(X)1\{\hat{Q}_{\tau}^{L}(X)\le0\}$
by noting that $\bar{Q}_{\tau}(x)$ also satisfies $\bar{Q}_{\tau}(x)=Q_{\tau}^{U}(x)1\{Q_{\tau}^{U}(x)\ge0\}+Q_{\tau}^{L}(x)1\{Q_{\tau}^{L}(x)\le0\}$.
We assume that the consistent estimators $\hat{Q}_{\tau}^{L}(X)$
and $\hat{Q}_{\tau}^{U}(X)$ are consistent with the specified rate
of convergence.

\begin{myas}{EST2}\label{as:EST2}$Q_{\tau}(X)$ is bounded almost
surely and $\hat{Q}_{\tau}^{L}(X)-Q_{\tau}^{L}(X)=O_{p}(n^{-\alpha})$
and $\hat{Q}_{\tau}^{U}(X)-Q_{\tau}^{U}(X)=O_{p}(n^{-\alpha})$ for
some $\alpha>0$.\end{myas}To overcome the computational problem
of obtaining $\delta_{mmr}^{*}$, we adopt the outcome weighted learning
framework \citep{zhao2012estimating}. We are interested in finding
a decision function $f:\mathcal{X}\rightarrow\mathbb{R}$ such that
$\delta(x)=1\{f(x)\ge0\}$. Note that by \eqref{eq:max_regret3},
we have $\bar{R}(f)\equiv\bar{R}(1\{f(\cdot)\ge0\})=E[|\bar{Q}_{\tau}(X)|1\{sign(f(X))\neq sign(\bar{Q}_{\tau}(X))\}]+E\bigg[\min(Q_{\tau}^{U}(X),-Q_{\tau}^{L}(X))1\{Q_{\tau}^{L}(X)<0<Q_{\tau}^{U}(X)\}\bigg].$
Accordingly, we define the surrogate regret as $\bar{R}^{S}(f)\equiv E[|\bar{Q}_{\tau}(X)|\phi\{sign(\bar{Q}_{\tau}(X))f(X)\}]+E\bigg[\min(Q_{\tau}^{U}(X),-Q_{\tau}^{L}(X))1\{Q_{\tau}^{L}(X)<0<Q_{\tau}^{U}(X)\}\bigg].$
Motivated by this expression, let $\hat{f}$ be the ML estimator of
$f$ from the following problem:
\begin{align}
\hat{f} & =\arg\min_{f\in\mathcal{H}}\left\{ \frac{1}{n}\sum_{i=1}^{n}\left|\widehat{\bar{Q}}_{\tau}(X_{i})\right|\phi\{sign(\widehat{\bar{Q}}_{\tau}(X_{i}))f(X_{i})\}+\lambda_{n}||f||^{2}\right\} ,\label{eq:surrogate}
\end{align}
where $\phi(t)=\max\{1-t,0\}$ is the hinge loss, $\lambda_{n}$ is
the regularization parameter, and $||\cdot||$ is the norm in a function
space. We focus on the reproducing kernel Hilbert space (RKHS) $\mathcal{H}_{k}$
associated with Gaussian radial basis function kernels $k(x,z)=\exp(-\sigma_{n}^{2}||x-z||^{2})$.
By Theorem 2.1 of \citet{steinwart2007fast}, the complexity of $\mathcal{H}_{k}$
in terms of the covering number satisfies $\sup_{P_{n}}\log N\{B_{\mathcal{H}_{k}},\epsilon,L_{2}(P_{n})\}\leq c_{n}\epsilon^{-p},$
where $P_{n}$ is the distribution of $(Y,D,X)$, $c_{n}=c_{p,\delta,d}\sigma_{n}^{(1-p/2)(1+\delta)d}$,
$B_{\mathcal{H}_{k}}$ is the closed unit ball of $\mathcal{H}_{k}$,
$p\in(0,2]$, $\delta>0$, and $c_{p,\delta,d}$ is a constant. Define
the approximation error function as $a(\lambda_{n})=\inf_{f\in\mathcal{H}_{k}}E[|\bar{Q}_{\tau}(X)|\phi\{sign(\bar{Q}_{\tau}(X))f(X)\}+\lambda_{n}||f||^{2}]-\inf_{f}E[|\bar{Q}_{\tau}(X)|\phi\{sign(\bar{Q}_{\tau}(X))f(X)\},$
where the second infimum is over the unrestricted space of $f$. Note
that $a(\lambda_{n})$ goes to zero if the RKHS is rich enough. The
following theorem establishes the asymptotic bound on $\bar{R}(f)$.
The asymptotic bound on the true regret can be similarly obtained.

\begin{theorem}\label{thm:regret_bound_ML} Suppose Assumptions \ref{as:RC}
and \ref{as:EST2} hold, and suppose that $\lambda_{n}=o(1)$ and
$\lambda_{n}n^{\min(2\alpha,1)}\rightarrow\infty$. Then, with probability
larger than $1-\exp(-2\eta)$, we have
\begin{align*}
\bar{R}(\hat{f})\leq & \inf_{f}\bar{R}(f)+a(\lambda_{n})+O_{p}(n^{-\alpha}\lambda_{n}^{-1/2})+M_{p}c_{n}^{\frac{2}{p+2}}n^{-\frac{2}{p+2}}\left(\lambda_{n}^{-\frac{2}{p+2}}+\lambda_{n}^{-1/2}\right)+\frac{K\eta}{n\lambda_{n}}(1+2\lambda_{n}^{1/2}),
\end{align*}
where $M_{p}$ and $K$ are constants. \end{theorem}

The leading term satisfies $\inf_{f}\bar{R}(f)=\bar{R}(f^{*})=E\left[\min(\max(Q_{\tau}^{U}(X),0),\max(-Q_{\tau}^{L}(X),0))\right]$,
because $f$ is not restricted and $f^{*}(x)=1$ if $Q_{\tau}^{L}(x)\ge0$,
$f^{*}(x)=0$ if $Q_{\tau}^{U}(x)\le0$ and $f^{*}(x)=sign(Q_{\tau}^{L}(x)+Q_{\tau}^{U}(x))$
if $Q_{\tau}^{L}(x)<0<Q_{\tau}^{U}(x)$. Note that this term coincides
with the leading term derived in Theorem \ref{thm:regret_bound} for
the deterministic rule. The second term is the approximation error
due to using the RKHS. The third term is the estimation error in estimating
the bounds. The rest of the terms are statistical errors in estimating
the policy.

\section{Calculating Bounds\label{sec:Calculating-Bounds}}

When $Q_{\tau}(X)$ is partially identified, we need a practical way
of calculating its bounds $[Q_{\tau}^{L}(x),Q_{\tau}^{U}(x)]=\{Q_{\tau}(x):F_{Y_{1},Y_{0}|X}\in\mathcal{F}\}.$
Unlike the Makarov bounds, the closed-form expression of the bounds
is not always available especially under Assumption \ref{as:PD}.
Therefore, it is fruitful to have a systematic procedure of calculating
the bounds. To this end, let $C(u_{1},u_{2}|X)$ be the copula for
$(U_{1},U_{2})\equiv(F_{Y_{1}}(Y_{1}),F_{Y_{0}}(Y_{0}))$ conditional
on $X$. By Sklar's Theorem, $C(u_{1},u_{2}|X)=F_{Y_{1},Y_{0}|X}(Q_{u_{1}}(Y_{1}|X),Q_{u_{2}}(Y_{0}|X))$.
Then, it satisfies $P[Y_{1}-Y_{0}\le t|X]=\int1\{Q_{u_{1}}(Y_{1}|X)-Q_{u_{2}}(Y_{0}|X)\le t\}dC(u_{1},u_{2}|X).$
Therefore, we can calculate the lower and upper bounds on the distribution
of $\Delta|X$ (recalling $\Delta\equiv Y_{1}-Y_{0}$) by
\begin{align}
F_{\Delta|X}^{L}(t) & =\inf_{C(\cdot,\cdot|X)\in\mathcal{C}}\int1\{Q_{u_{1}}(Y_{1}|X)-Q_{u_{2}}(Y_{0}|X)\le t\}dC(u,v|X),\label{eq:LB}
\end{align}
and similarly for $F_{\Delta|X}^{U}(t)$ by taking supremum over $\mathcal{C}$,
where $\mathcal{C}$ is the class of copulas $C(\cdot,\cdot|X=x)$
restricted by identifying assumptions. Note that \eqref{eq:LB} can
be viewed as the (constrained version of the) Monge-Kantorovich problem
of finding the optimal coupling of marginal distributions in the optimal
transport theory \citep{villani2009optimal}. Then, for $\tau$-quantile
$Q_{\tau}$ of $\Delta$, we can obtain its lower and upper bounds
as $Q_{\tau}^{L}(X)=F_{\Delta|X}^{U,-1}(\tau)$ and $Q_{\tau}^{U}(X)=F_{\Delta|X}^{L,-1}(\tau)$,
where the inverse denotes the generalized inverse. In practice, \eqref{eq:LB}
is an infinite dimensional program, and thus infeasible. To transform
them into a linear program, we propose to approximate $C(u,v|x)$
using the Bernstein copula $C_{B}(u,v|x)$ \citep{sancetta2004bernstein}.

\begin{definition}[Bernstein Copula]For $j\in\{1,2\}$, let $P_{v_{j}}^{m_{j}}(u_{j})\equiv\left(\begin{array}{c}
m_{j}\\
v_{j}
\end{array}\right)u_{j}^{v_{j}}(1-u_{j})^{m_{j}-v_{j}}$. Then, $C_{B}:[0,1]^{2}\rightarrow[0,1]$ is a conditional Bernstein
copula for any $m_{j}\ge1$ and $x\in\mathcal{X}$ if $C_{B}(u_{1},u_{2}|x)=\sum_{v_{1}=0}^{m_{1}}\sum_{v_{2}=0}^{m_{2}}\beta\left(\frac{v_{1}}{m_{1}},\frac{v_{2}}{m_{2}},x\right)P_{v_{1}}^{m_{1}}(u_{1})P_{v_{2}}^{m_{2}}(u_{2})$
satisfies the usual properties of the copula function.\end{definition}

Then we can compute a feasible version of \eqref{eq:LB} as
\begin{align}
 & \min_{\beta\in\mathcal{B}}\sum_{v_{1}=0}^{m_{1}}\sum_{v_{2}=0}^{m_{2}}\beta\left(\frac{v_{1}}{m_{1}},\frac{v_{2}}{m_{2}},X\right)\int 1\{Q_{u_{1}}(Y_{1}|X)-Q_{u_{2}}(Y_{0}|X)\le t\}dP_{v_{1}}^{m_{1}}(u_{1})dP_{v_{2}}^{m_{2}}(u_{2}),\label{eq:LB1-1}
\end{align}
and similarly for the upper bound by taking maximum over $\mathcal{B}$,
where $\mathcal{B}$ is the restricted set of $\beta(\cdot)$ to impose
identifying assumptions and guarantee that $C_{B}$ is a proper copula.
We omitted the latter restrictions for succinctness; see Theorem 2
in \citet{sancetta2004bernstein} for details. For example, to impose
Assumption \ref{as:PD}(iii) it is necessary to ensure that $C_{B}(u_{1}|u_{2},x)=\partial C_{B}(u_{1},u_{2},x)/\partial u_{2}$
and $C_{B}(u_{2}|u_{1},x)=\partial C_{B}(u_{1},u_{2},x)/\partial u_{1}$
are non-increasing. Then, by the desirable property of Bernstein,
this corresponds to $\beta\left(\frac{v_{1}}{m_{1}},\frac{v_{2}}{m_{2}},X\right)$
being weakly increasing in $v_{1}$ and $v_{2}$. The use of Bernstein
approximation for the systematic calculation of bounds on treatment
effects also appears in \citet{han2023optimal} and \citet{han2020sharp}
in different contexts. As an alternative to the Bernstein approximation,
one can discretize the space of $(U_{1},U_{2})\in[0,1]^{2}$ \citep{blundell2007changes,frandsen2021partial}.
This approach can be viewed as a simple local approximation involving
a uniform kernel. Finally, in practice, the inputs $Q_{u_{1}}(Y_{1}|X)$
and $Q_{u_{2}}(Y_{0}|X)$ of the linear program can be estimated using
standard nonparametric or parametric methods. When they are estimated
consistently, we can show that Assumption \ref{as:EST} holds for
the estimated outputs, $\hat{Q}_{\tau}^{L}(X)$ and $\hat{Q}_{\tau}^{U}(X)$,
of the linear program:

\begin{lemma}\label{lem:consistency}Suppose that, for $d\in\{0,1\}$,
$F_{Y_{d}|X}(y|X)$ and $Q_{\tau}(Y_{d}|X)$ are absolutely continuous
in $y\in\mathcal{Y}$ and $\tau\in(0,1)$, respectively, and $\hat{Q}_{\tau}(Y_{d}|X)$
is a consistent estimator of $Q_{\tau}(Y_{d}|X)$ for any $\tau\in(0,1)$,
almost surely. Then, $|\hat{Q}_{\tau}^{L}(X)-Q_{\tau}^{L}(X)|=o_{p}(1)$
and $|\hat{Q}_{\tau}^{U}(X)-Q_{\tau}^{U}(X)|=o_{p}(1)$.\end{lemma}

\section{Numerical Illustrations}

We numerically show the performance of treatment allocations across
welfare criteria, especially when the QoTE is partially identified.
For succincness, we only present a subset of results here; the full
results are contained in the Supplemental Appendix. As data-generating
processes (DGPs), we conside normal and lognormal distributions for
$(Y_{1},Y_{0})$ and Bernoulli for $D$. In simulation, the bounds
$Q_{\tau}^{L}$ and $Q_{\tau}^{U}$ are calculated under either no
assumption (i.e., Makarov bounds) or SI. For the policies $\delta_{mmr}^{*}$,
$\delta_{QTE}^{*}$ and $\delta_{ATE}^{*}$, we estimate their sample
counterparts $\hat{\delta}^{*}$, $\hat{\delta}_{QTE}^{*}$ and $\hat{\delta}_{ATE}^{*}$
by estimating $Q_{\tau}^{U}$, $Q_{\tau}^{L}$, $Q_{\tau}(Y_{d})$,
and $E[Y_{d}]$ ($d=0,1$).

Table \ref{tab:classification_rate1_short} presents the simulated
correct classification rates of the estimated policies. Under DGP
1 (with a normal distribution), both intervals under stochastic
increasingness (SI) (i.e., Assumption \ref{as:PD}(iii)) and no assumption
exclude $0$ and lie relatively far from it; therefore, both $\hat{\delta}$
and $\hat{\delta}^{SI}$ perform well. Under DGP 2 (with a log-normal
distribution),\emph{ }the interval under SI excludes $0$ while the
interval under no assumption covers $0$; therefore, $\hat{\delta}^{SI}$
performs better than $\hat{\delta}$; under this log-normal setting
and SI, $Q_{\tau}(Y_{1}-Y_{0})<E(Y_{1})-E(Y_{0})$ may be violated,
which occurs in the current subgroup and thus $\hat{\delta}^{SI}$
performs better than $\hat{\delta}_{ATE}$.

\begin{table}[ht]
\begin{centering}
\begin{tabular}{|c|cccccc|}
\hline
{\scriptsize{}\diagbox{Optimal Policy}{Estimated Policy}} & {\scriptsize{}$\hat{\delta}^{stoch,SI}$} & {\scriptsize{}$\hat{\delta}^{stoch}$} & {\scriptsize{}$\hat{\delta}^{determ,SI}$} & {\scriptsize{}$\hat{\delta}^{determ}$} & {\scriptsize{}$\hat{\delta}_{QTE}$} & {\scriptsize{}$\hat{\delta}_{ATE}$}\tabularnewline
\hline
\hline
 & \multicolumn{6}{c|}{{\scriptsize{}DGP 1 (Normal)}}\tabularnewline
\hline
{\scriptsize{}$\delta^{*}$} & {\scriptsize{}$100\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$1.5\%$} & {\scriptsize{}$100\%$}\tabularnewline
{\scriptsize{}$\delta_{QTE}^{*}$} & {\scriptsize{}$0\%$} & {\scriptsize{}$0\%$} & {\scriptsize{}$0\%$} & {\scriptsize{}$0\%$} & {\scriptsize{}$98.5\%$} & {\scriptsize{}$0\%$}\tabularnewline
{\scriptsize{}$\delta_{ATE}^{*}$} & {\scriptsize{}$100\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$1.5\%$} & {\scriptsize{}$100\%$}\tabularnewline
\hline
 & \multicolumn{6}{c|}{{\scriptsize{}DGP 2 (Log-Normal)}}\tabularnewline
\hline
{\scriptsize{}$\delta^{*}$} & {\scriptsize{}$99.5\%$} & {\scriptsize{}$48\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$32.5\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$4.5\%$}\tabularnewline
{\scriptsize{}$\delta_{QTE}^{*}$} & {\scriptsize{}$99.5\%$} & {\scriptsize{}$48\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$32.5\%$} & {\scriptsize{}$100\%$} & {\scriptsize{}$4.5\%$}\tabularnewline
{\scriptsize{}$\delta_{ATE}^{*}$} & {\scriptsize{}$0\%$} & {\scriptsize{}$52\%$} & {\scriptsize{}$0\%$} & {\scriptsize{}$67.5\%$} & {\scriptsize{}$0\%$} & {\scriptsize{}$95.5\%$}\tabularnewline
\hline
\end{tabular}
\par\end{centering}
\caption{Correct Classification Rate ($\tau=0.25$, $n=1000$)}
\label{tab:classification_rate1_short}
\end{table}


\section{Empirical Applications\label{sec:Empirical-Applications}}

We consider two empirical applications to illustrate our method: (i)
the allocation of right heart catheterization to patients and (ii)
the allocation of job training to workers. We only present (i) here;
(ii) is contained in the Supplemental Appendix.

We consider the right heart catheterization (RHC) dataset from the
Study to Understand Prognoses and Preferences for Outcomes and Risks
of Treatments (SUPPORT) \citep{hirano2001estimation}. The treatment
$D$ in question is the RHC ($1$ if received and $0$ otherwise),
a diagnostic procedure for critically ill patients. The outcome $Y$
is the number of days from admission to death within 30 days (t3d30),
whose value ranges from 2 to 30. In contrast to the belief of practitioners
that the RHC is beneficial, studies like \citet{connors1996effectiveness}
found that patient survival is lower with the RHC than without. Therefore,
a relevant policy question in this critical situation is to find patients
for whom allocating (or avoiding) the RHC is life-saving. In the dataset,
5735 patients are divided into a treatment group (2184 patients) and
a control group (3551 patients). We consider the following covariates
as $X$: age, sex, coma in primary disease 9-level category (cat1\_coma),
coma in secondary disease 6-level category, (cat2\_coma), do not resuscitate
(DNR) status on day 1 (i.e., DNR when heart stops) (dnr1), estimated
probability of surviving 2 months (surv2md1), and APACHE III score
ignoring coma (i.e., ICU mortality score) (aps1).

\begin{figure}[H]
\begin{centering}
\centering \includegraphics[scale=0.3]{rhc}
\par\end{centering}
\begin{centering}
{\small \begin{tablenotes}
    \item \emph{Note}: In the figure, the vertical line indicates zero and the horizontal line indicates $\tau=0.25$.
\end{tablenotes}}
\par\end{centering}
\caption{Bounds on the QoTE of Six Representative Patients}
\label{fig:application1-1}
\end{figure}

To estimate the counterfactual distributions $F_{Y_{1}|X}$ and $F_{Y_{0}|X}$
of the outcome (t3d30) for different groups defined by the covariates,
we conduct a kernel regression in the treatment and control groups
separately with bandwidth under Scott's rule of thumb.\footnote{To simplify this process, we run the regression $P[Y<y_{j}|X=x]=E[1\{Y<y_{j}\}|X=x]$
on a series of $y_{j}=F_{Y}^{-1}(\frac{2j-1}{2k})$, where $k=1000$
and $j=1,...,k$.} Then we calculate the upper and lower bounds of the QoTE under SI (i.e., \ref{as:PD}(iii)) and no assumption
and make the decisions by using the proposed criterion based on the
QoTE. As shown in the simulation results in the Supplemental Appendix,
the SI and no-assumption bounds will not always give the same decisions,
and the information provided by the bounds differs from person to
person.

In Figure \ref{fig:application1-1}, we present six cases to show
the SI and no-assumption bounds of the QoTE. We only focus on deterministic
policies and $\tau=0.25$. These results illustrate how the actual
implementation of our proposed policies would look like for each individual.
It is shown that there is much heterogeneity in terms of the QoTEs
and thus the corresponding optimal decisions.

Next, in Figure \ref{fig:application1-2}, we present the decisions
of allocating the RHC in terms of age and survival rate, which are
two important covariates for the allocation decision. We focus on
the male group whose primary and secondary disease categories are
not coma and APACHE score at day 1 is 54 and with resuscitate status.
We use the 0.25-quantile, median, and 0.75-quantile QoTE bounds to
represent prudent, majority-minded, and negligent PMs, respectively.
As expected, the 0.75-quantile bounds suggest the treatment option
more often than the bounds with the other quantile probabilities.
Given that the 0.25-quantile bounds suggest the most prudent decisions,
the suggested treatment option can be viewed as a compelling recommendation.
\begin{figure}[htbp]
\centering \subfigure[QoTE with $\tau=0.25$]{ \label{Fig.sub.1-1}
\includegraphics[width=3.9cm,height=3.9cm]{pi1_1st_quantile_rhc}}
\subfigure[QoTE with $\tau=0.5$]{ \label{Fig.sub.2-1} \includegraphics[width=3.9cm,height=3.9cm]{pi1_median_rhc}}
\subfigure[QoTE with $\tau=0.75$]{ \label{Fig.sub.3-1} \includegraphics[width=3.9cm,height=3.9cm]{pi1_3rd_quantile_rhc}}
\caption{Treatment Decisions for Male Patients with Specific Health Conditions}
\label{fig:application1-2}
\end{figure}
\begin{figure}[htbp]
\centering \subfigure[QTE with $\tau=0.25$]{ \label{Fig2.sub.1-1}
\includegraphics[width=3.9cm,height=3.9cm]{pi2_1st_quantile_rhc}}
\subfigure[QTE with $\tau=0.5$]{ \label{Fig2.sub.2-1} \includegraphics[width=3.9cm,height=3.9cm]{pi2_median_rhc}}
\subfigure[QTE with $\tau=0.75$]{ \label{Fig2.sub.3-1} \includegraphics[width=3.9cm,height=3.9cm]{pi2_3rd_rhc}}
\subfigure[ATE]{ \label{Fig2.sub.4-1} \includegraphics[width=3.9cm,height=3.9cm]{pi3_decision_rhc}}
\caption{Treatment Decisions Based on the QTEs and the ATE}
\label{fig:application1-3}
\end{figure}

For comparison, in Figure \ref{fig:application1-3}, we present the
allocation decisions based on the 0.25-quantile, median, and 0.75-quantile
QTE and the ATE. Interestingly, there is no obvious tendency in decisions
when the quantile probability increases from 0.25 to 0.75, which reflects
the possible limitation of using the QTE as the basis for decisions
(e.g., the quantile probability does not capture the level of prudence).
The decisions based on the ATE show how they can be viewed as the
most common approach in the literature. They look very similar to
the decisions based on the median QoTE bounds, although there are
a few points that differ from the latter. Note that the policy based
on the median QoTE bounds can be viewed as a robustness check for
the policy based on the ATE.

\section{Generalization\label{sec:Extensions}}

The joint distribution of the potential outcomes may contain other
useful information about treatment effect heterogeneity for policy
learning. The theoretical results of this paper can be generalized
to any setting where welfare is defined as a functional of the joint
distribution of potential outcomes. Consider an optimal rule that
satisfies
\begin{align}
\delta^{**} & \in\arg\max_{\delta\in\mathcal{D}}E\left[\delta(X)\Lambda(F_{Y_{1},Y_{0}|X})\right],\label{eq:extension0}
\end{align}
where $\Lambda:\tilde{\mathcal{F}}\rightarrow\mathbb{R}$ is some
functional of the joint distribution of $(Y_{1},Y_{0})$ given $X$.
Our original criterion \eqref{eq:quantile_welfare} is a special case
of \eqref{eq:extension0} with $\Lambda(F_{Y_{1},Y_{0}|X})=Q_{\tau}(Y_{1}-Y_{0}|X)$.
We propose other examples of the criterion $\Lambda(F_{Y_{1},Y_{0}|X})$
that may interest a non-utilitarian PM.

\begin{example}\label{ex1}Consider
\begin{align}
\delta_{\bar{y}}^{**} & \in\arg\max_{\delta\in\mathcal{D}}E\left[\delta(X)E[Y_{1}-Y_{0}|Y_{0}<\bar{y},X]\right]\label{eq:extension1}
\end{align}
for some predetermined $\bar{y}$. This can be motivated by maximizing
the average of $\delta(X)Y_{1}+(1-\delta(X))Y_{0}$ (or simply $Y_{\delta(X)}$
with deterministic $\delta$), conditional on $Y_{0}<\bar{y}$: $E\left[\delta(X)Y_{1}+(1-\delta(X))Y_{0}|Y_{0}<\bar{y}\right]$,
because $E[\delta(X)Y_{1}+(1-\delta(X))Y_{0}|Y_{0}<\bar{y}]=E[Y_{0}|Y_{0}<\bar{y}]+E[\delta(X)E[Y_{1}-Y_{0}|Y_{0}<\bar{y},X]]$.
The PM with the latter criterion focuses on the welfare
of a disadvantaged population, defined by the baseline outcome $Y_{0}$
being less than $\bar{y}$. A similar intuition applies to the criterion
\eqref{eq:extension1}, which can be interpreted as addressing the
average gain for the disadvantaged. The ATE for the disadvantaged,
$E[Y_{1}-Y_{0}|Y_{0}<\bar{y}]$, appears as a parameter of interest
in the context of policy evaluation \citep{heckman1995assessing},
which is adapted here for the context of optimal allocation; also
see \citet{kaji2023assessing} for recent related work. \end{example}

\begin{example}\label{ex2}Alternative to Example \ref{ex1}, one
can consider $\Lambda(F_{Y_{1},Y_{0}|X})=Q_{\tau}(Y_{1}-Y_{0}|Y_{0}<\bar{y},X)$.
This would make the criterion robust to outliers and add an additional
dimension, $\tau$, to target a specific subgroup. Analogous to Theorem
\ref{thm:interpretation}, for continuously distributed $Y_{d}$ and
deterministic policy $\delta$, the first-best policy under $Q_{0.5}(Y_{1}-Y_{0}|Y_{0}<\bar{y},X)$
is the one that maximizes $P\left[Y_{\delta(X)}-Y_{1-\delta(X)}>0|Y_{0}<\bar{y}\right]=P\left[\delta\in\arg\max_{d}Y_{d}|Y_{0}<\bar{y}\right]$.\end{example}

\begin{example}\label{ex3}One may wish to target individuals worst-affected
by the treatment rather than those who are worst off at baseline.
The conditional value at risk (CVaR) of the distribution of individual
treatment effects can serve just that: $\Lambda(F_{Y_{1},Y_{0}|X})=E[Y_{1}-Y_{0}|Y_{1}-Y_{0}<\bar{\Delta},X]$.
\citet{kallus2023treatment} considers the CVaR, $E[Y_{1}-Y_{0}|Y_{1}-Y_{0}<\bar{\Delta}]$
with $\bar{\Delta}=Q_{\tau}(Y_{1}-Y_{0})$, as a parameter related
to distributional treatment effects and provides a sharp upper bound
and, under restricted heterogeneity, a sharp lower bound on the CVaR.\end{example}

In all these examples, $\Lambda(F_{Y_{1},Y_{0}|X})$ is not generally
point-identified, so one can follow the approach in Section \ref{subsec:Policies-Robust-to}
by considering $\delta_{mmr}^{**} \in\arg\min_{\delta\in\mathcal{D}}\max_{F_{Y_{1},Y_{0}|X}\in\mathcal{F}}E\left[\delta(X)\Lambda(F_{Y_{1},Y_{0}|X})\right]$. Then, one can apply the identifying assumptions listed in Section
\ref{sec:Possible-Identifying-Assumptions} and the computational
method proposed in Section \ref{sec:Calculating-Bounds} to systematically
calculate the bounds on $\Lambda(F_{Y_{1},Y_{0}|X})$ and to eventually
estimate $\delta_{mmr}^{**}$. Let $\Lambda^{L}(X)$ and $\Lambda^{U}(X)$
be the lower and upper bounds on $\Lambda(F_{Y_{1},Y_{0}|X})$. Then
the theoretical properties of the estimated $\delta_{mmr}^{**}$ with
constrained and unconstrained policy classes can be established based
on Section \ref{sec:Theoretical-Properties-of} by simply replacing
$Q_{\tau}^{L}(X)$ and $Q_{\tau}^{U}(X)$ with $\Lambda^{L}(X)$ and
$\Lambda^{U}(X)$ throughout the section.

\newpage
\begin{center}
    \LARGE Supplemental Appendix\\[1ex]
\end{center}