EconBase
← Back to paper

Policy Learning under Biased Sample Selection

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

62,320 characters · 0 sections · 47 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Policy Learning under Biased Sample Selection

abstractPractitioners often use data from a randomized controlled trial to learn a treatment assignment policy that can be deployed on a target population. A recurring concern in doing so is that, even if the randomized trial was well-executed (i.e., internal validity holds), the study participants may not represent a random sample of the target population (i.e., external validity fails) – and this may lead to policies that perform suboptimally on the target population. We consider a model where observable attributes can impact sample selection probabilities arbitrarily but the effect of unobservable attributes is bounded by a constant, and we aim to learn policies with the best possible performance guarantees that hold under any sampling bias of this type. In particular, we derive the partial identification result for the worst-case welfare in the presence of sampling bias and show that the optimal max-min, max-min gain, and minimax regret policies depend on both the conditional average treatment effect (CATE) and the conditional value-at-risk (CVaR) of potential outcomes given covariates. To avoid finite-sample inefficiencies of plug-in estimates, we further provide an end-to-end procedure for learning the optimal max-min and max-min gain policies that does not require the separate estimation of nuisance parameters.
section{Introduction} Practitioners often use data from a randomized controlled trial (RCT) to a learn treatment assignment policy that can be deployed on a target population. Formally, in the study population, each unit $i$ has covariates $X_{i} \in \mathcal{X}$ and potential outcomes $Y_{i}(0), Y_{i}(1) \in \mathcal{Y}$ under control and treatment, respectively, that are distributed according to a study potential outcome distribution $P$. After administering the trial, the practitioner has access to samples $(X_{i}, Y_{i}, W_{i}) \sim P_{\text{obs}}$, where $P_{\text{obs}}$ is the observed data distribution that is consistent with $P$ under an internally-valid RCT, meaning that the treatments $W_{i} \in \{0, 1\}$ are randomly assigned and the observed outcome $Y_{i}$ equals $Y_{i}(W_{i}),$ the potential outcome under the prescribed treatment. A common criticism of RCTs is that they lack external validity. External validity captures the extent to which conclusions drawn from the RCT study population can generalize to a broader population or other target populations. An RCT may lack external validity for a variety of reasons. For example, bell2016estimates demonstrates that non-random site selection in the evaluation of the Reading First educational program would have led to lower impact estimates than representative site selection. In addition, wang2018efficacy discuss that randomized trials for measuring the effects of anti-depressants rely on volunteers to opt in to participate, so the findings of these studies may not apply to non-volunteers. In addition, the time lag between data collection and deployment may cause the target population to differ from the study population due to temporal shift. Standard policy learning approaches ignore the concern that the randomized control trial may lack external validity. These methods use data from $P_{\text{obs}}$ to learn the policy that maximizes the mean outcome over the study population athey2021policy,bhattacharya2012inferring, kitagawa2018should, manski2004statistical: \begin{equation} \pi^{*} = \argmax_{\pi \in \Pi} \EE[P]{Y(\pi(X))},\end{equation} where $\pi: \mathcal{X}\mapsto \{0, 1\}$ denotes a treatment assignment policy and $\Pi$ denote a policy class. The implicit assumption of these methods is that the target population is the same as the study population. The learned policy in (ref) does not have performance guarantees for target potential outcome distributions $Q$ that differ from $P$. In this work, we aim to identify and learn policies that are robust to certain failures of external validity, namely sampling bias in the selection of study participants. As before, we assume that the practitioner has access to data from $P_{\text{obs}}$, but we instead aim to learn a policy $\pi$ that has performance guarantees on an unknown target distribution $Q$, while the study potential outcome distribution $P$ may be biased relative to $Q$ due to sampling bias. Following manski2003partial, we model sample selection using a binary selection indicator. We quantify the strength of bias in sample selection via the $\Gamma$-biased sampling model aronow2013interval, nie2021covariate, sahoo2022learning. This model allows the probability of sample selection for each unit to depend arbitrarily on covariates but only a bounded amount on unobservables. The parameter $\Gamma \geq 1$ captures the strength of the sampling bias due to unobservables, where larger values of $\Gamma$ permit larger amounts of bias. Note that, when $\Gamma=1$, this model corresponds to “unconfounded sample selection,” which is studied in the literature on generalizability stuart2011use, tipton2013improving, tipton2014generalizable. \begin{defi} Let $\Gamma \geq 1$. For any pair of distributions $P$ and $Q$ over $(X, Y(0), Y(1))$, we say that $Q$ can generate $P$ under $\Gamma$-biased sampling if there exists a distribution $\tilde{Q}$ over $(X, Y(0), Y(1), S)$, where $S \in \{0, 1\}$ is a “selection indicator” that satisfies the following properties: The $(X,Y(0), Y(1))$-marginal of $\tilde{Q}$ is equal to $Q$, the $(X,Y(0), Y(1))$-marginal of $\tilde{Q}$ conditionally on $S = 1$ is equal to $P$, and \begin{equation} \frac{\PP[\tilde{Q}]{S=1 \mid X=x, Y(0)=y_{0}, Y(1)=y_{1}}}{\PP[\tilde{Q}]{S=1 \mid X=x}} \in [\Gamma^{-1}, \Gamma] \quad \forall x \in \mathcal{X}, y_{0}, y_{1} \in \mathcal{Y}. \end{equation} \end{defi} \begin{figure} \caption{We visualize the set of plausible target distributions $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$. The set $\mathcal{T}(P_{\text{obs}})$ is the set of study potential outcome distributions $P$ that are consistent with $P_{\text{obs}}$ under an internally-valid RCT. The set $\mathcal{R}^{\Gamma}(P, Q_{X})$ is the set of target potential outcome distributions that can generate $P$ under $\Gamma$-biased sampling and have covariate distribution $Q_{X}$.} \end{figure} There are two key challenges when learning policies that are robust to biased sample selection. The first challenge is the “missing data problem” that arises in the standard policy learning setting. In a randomized controlled trial, units are assigned to either control or treatment, so $Y_{i}(0), Y_{i}(1)$ are not simultaneously observed for any unit. Although we can identify the marginal distributions of the study potential outcome distribution $P_{X, Y(0)}, P_{X, Y(1)}$ from $P_{\text{obs}}$, the joint distribution $P$ cannot be identified from $P_{\text{obs}}$. So, the true study potential outcome distribution is unknown, and there are many potential outcome distributions $P$ that yield $P_{\text{obs}}$ under an internally-valid RCT. Let $\mathcal{T}(P_{\text{obs}})$ be the set of study potential outcome distributions that yield $P_{\text{obs}}$ under an internally-valid RCT, i.e., \[\mathcal{T}(P_{\text{obs}}) = \{P: P_{X, Y(0)} = P_{\text{obs}, X, Y \mid W=0}, P_{X, Y(1)} = P_{\text{obs}, X, Y \mid W=1} \}.\] The second challenge arises from biased sampling. We note that the true target potential outcome distribution is also unknown, and given a particular study potential outcome distribution $P$, there are many possible target distributions $Q$ that can generate $P$ under $\Gamma$-biased sampling for $\Gamma > 1$. Let $\mathcal{R}^{\Gamma}(P, Q_{X})$ be the set of potential outcome distributions $Q$ that have covariate distribution $Q_{X}$ and can generate $P$ via $\Gamma$-biased sampling. Thus, we can define $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$ to be the robustness set, the set of all plausible target potential outcome distributions consistent with $P_{\text{obs}}$ under $\Gamma$-biased sampling and an internally-valid RCT: \[ \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X}) = \{ Q \mid Q \in \mathcal{R}^{\Gamma}(P, Q_{X}) \text{ where } P \in \mathcal{T}(P_{\text{obs}}) \}.\] This set is visualized in Figure (ref). To obtain performance guarantees on the unknown target population, we apply the distributionally robust optimization (DRO) framework ben2013robust to learn a policy that is robust to all plausible target distributions in the robustness set. As in manski2011choosing, we consider a number of different ways to quantify good performance of a policy that account for ambiguity under partial identification. The max-min objective considers worst-case welfare over the partially identified set, and the max-min policy is \begin{equation} \pi_{\Gamma, maxmin}^{*} \in \argmax_{\pi \in \Pi} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{obs}, Q_{X})} \EE[Q]{Y(\pi(X))}. \end{equation} Optimizing max-min guarantees is conceptually simple; however, in many cases, the max-min objective is viewed as too pessimistic manski2011choosing,savage1951theory. To avoid this pessimism we also consider two other objectives. When a natural baseline policy $\pi_{0}$ is available (e.g., there is a clear status quo), we can optimize max-min improvements over the baseline; the max-min gain policy is \begin{equation} \pi_{\Gamma, gain}^{*} \in \argmax_{\pi \in \Pi} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{obs}, Q_{X})} \EE[Q]{Y(\pi(X))} - \EE[Q]{Y(\pi_{0}(X))}. \end{equation} As discussed in kallus2021minimax, when the status quo policy is already considered to be a reasonably good decision rule, then the max-min gain approach may be considered desirable in that it provides “safe” improvements over the status quo, i.e., we will remain with the status quo unless the data lets us unequivocally prefer another action. Finally, we also consider a minimax regret criterion; the minimax regret policy is \begin{equation} \pi_{\Gamma, regret}^{*} \in \argmin_{\pi \in \Pi} \sup_{Q \in \mathcal{S}^{\Gamma}(P_{obs}, Q_{X})} R_{Q}(\pi(X)), \end{equation} where regret under distribution $Q$ is the gap between the mean outcome of a policy and that of the best-performing policy under distribution $Q$: \begin{equation} R_{Q}(\pi(X)) = \sup_{\pi' \in \Pi} \EE[Q]{Y(\pi'(X))} - \EE[Q]{Y(\pi(X))}. \end{equation} Several authors have argued in favor of minimax regret is a relevant criterion for guiding action under ambiguity manski2011choosing,savage1951theory; this criterion does not suffer from pessimistic failure modes of the max-min criterion, and also can be used in the absence of a strong baseline policy. Our work has two main contributions. In our first contribution, we demonstrate that the optimal policies defined in (ref), (ref), and (ref) are identifiable under $P_{\text{obs}}$ and give closed-form expressions for these policies when the policy class $\Pi$ consists of deterministic, unconstrained, binary-valued functions. In our second contribution, we give end-to-end procedures for learning the optimal max-min and max-min gain policies using data from $P_{\text{obs}}$ that do not require separate estimation nuisance parameters. In Section (ref), we demonstrate that when $\Pi$ consists of deterministic, unconstrained, binary-valued functions, the optimal policy under the max-min, max-min gain, and minimax regret objectives are identifiable using only data from $P_{\text{obs}}$, and we give closed-form expressions for these policies. In Section (ref), we give an equivalence result that is essential for demonstrating that our learning procedures yield the optimal policies. The result shows that robust optimization of a potential outcome value function $v^{*}(\pi(X); X, Y(0), Y(1))$ over $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$ is equivalent to robust optimization of a observed data value function $v(\pi(X); X, Y, W)$ over a robustness set of observed data distributions when the mean of the observed data value function under $P_{\text{obs}}$ is equal to the mean of the potential outcome value function under any potential outcome distribution $P \in \mathcal{T}(P_{\text{obs}})$. In Section (ref), we demonstrate that the robust optimization of an observed data value function over a robustness set of observed data distributions can be solved with an extension of Rockafellar-Uryasev (RU) Regression sahoo2022learning. We combine this result with the equivalence result from the previous section to give a procedure for learning the optimal max-min and max-min gain policies. In Section (ref), we evaluate our learning procedure empirically in simulations and a semi-synthetic experiment with the voting dataset of gerber2008social. \begin{subsection}{Related Work} The study of optimal treatment allocation has received attention in economics athey2021policy, bhattacharya2012inferring, kitagawa2018should, manski2004statistical, statistics qian2011performance, zhao2012estimating, and computer science swaminathan2015batch. The standard setting in this literature does not address the concern that the study population may differ from the target population of interest. The distinctive feature of our setting is that we permit the study potential outcome distribution $P$ to differ from the target potential outcome distribution $Q$ in the sense of $\Gamma$-biased sampling (Definition (ref)). sahoo2022learning introduces the $\Gamma$-biased sampling model for the supervised learning setting and focuses on statistical and algorithmic considerations for learning robust regression models. In this work, we consider optimal treatment choice using the potential outcomes framework and identify optimal policies under various robust objectives, including minimax regret. Our contribution of policy learning under biased sample selection is part of the literature on policy learning under partial identification adjaho2022externally,christensen2022optimal,ben2021safe, hansen2001robust,hatt2022generalizing,higbee2022policy, kallus2022doubly, kallus2021minimax, mufactored, manski2000identification, manski2007minimax, si2020distributional, watson2016approximate. The main challenge in this area is that the optimal policy is not identifiable because it depends quantities that are unknown, such as missing outcome data ben2021safe,higbee2022policy,manski2007minimax, unknown propensity weights kallus2021minimax, or an unknown target distribution adjaho2022externally, hatt2022generalizing, si2020distributional,mufactored. These works take a robust optimization approach to solve this problem; they define a plausible set of values for the unknown quantities and find a policy that performs well when the unknown quantities take on their worst-case values. Our work is in line with these previous works because we define a set of plausible values for target potential outcome distributions that is consistent with our observed data and the assumed model of sampling bias and identify policies that are robust to the worst-case target distribution. Of this literature, our work is most related to adjaho2022externally and si2020distributional. Similar to our work, they use a robust optimization approach to develop policies that have welfare guarantees for target populations that may differ from the study population. Nonetheless, the robustness set we consider in this paper has a different form than the other works (e.g., Wasserstein balls adjaho2022externally and KL-divergence balls si2020distributional about the study distribution). Unlike adjaho2022externally and si2020distributional which only discuss the max-min policy, we are able to derive the closed-form expressions for the max-min gain policy and the minimax regret policy that are arguably more relevant in practice manski2000identification. Another difference between our work and that of adjaho2022externally and si2020distributional is that they both define their robustness set by placing constraints on the joint distribution over $X, Y(0), Y(1)$, while the robustness sets that we consider place constraints on the conditional distribution $Y(0), Y(1) \mid X$. The approach of constraining the joint distribution over outcomes and covariates is overly conservative for a number of reasons. First, we typically observe the shift in covariate distribution at test-time ($Q_{X}$ is observed at test-time), so it is often not necessary to protect against arbitary covariate shifts. Second, when the policy class $\Pi$ consists of unconstrained binary-valued functions, the optimal policy solves the optimal treatment assignment problem for every $x \in \mathcal{X}$, so the covariate shift does not impact the optimal policy. Third, empirical evaluations have revealed that constraining the joint distribution over outcomes and covariates can yield conservative results compared to handling covariate shift and conditional shift separately mufactored. \end{subsection}
section{Identification} First, we define the notation and setup of our policy learning problem and define the set of target distributions that we aim to be robust to. Second, we define various robustness criteria that we will consider in this work. Finally, we identify the optimal policies under each robustness criterion and discuss connections between them. \begin{subsection}{Robustness Set} We consider the following policy learning problem. In our study population, each unit $i$ is associated with covariates $X_{i} \in \mathcal{X}$ and potential outcomes $Y_{i}(0), Y_{i}(1) \in \mathcal{Y}$ which are drawn independently from an (unknown) study distribution $P$. Throughout the paper we assume that \[P_{(Y(0), Y(1))\mid X = x}\text{ is absolutely continuous with respect to the Lebesgue measure for every }x\in \mathcal{X}.\] We will not mention this assumption again for simplicity. We have access to data $(X_{i}, Y_{i}, W_{i}) \sim P_{\text{obs}}$ from a randomized controlled trial, where $X_{i} \in \mathcal{X}$ are covariates, $W_{i} \in \{0, 1\}$ is the treatment assignment, and $Y_{i} \in \mathcal{Y}$ is the observed outcome. We assume that $P_{\text{obs}}$ is consistent with $P$ under an randomized control trial with treatment probability $e$. This means that treatments are randomly assigned according to $\text{Bernoulli}(e)$ and the observed outcomes satisfy SUTVA (stable unit treatment value assignment), meaning that $Y_{i} = Y_{i}(W_{i}).$ In other words, if $P_{\text{obs}}$ is generated via an internally-valid RCT, then \begin{align} P_{obs, X, Y \mid W=w} &= P_{X, Y(w)} \quad \forall w \in \{0, 1\}, \\ P_{obs, W \mid X=x} &= Bernoulli(e) \quad \forall x \in \mathcal{X}. \end{align} Let $Q$ be the target potential outcome distribution. We seek to learn a policy $\pi: \mathcal{X} \rightarrow \{0, 1\}$ such that the mean outcome $\EE[Q]{Y(\pi(X))}$ is large when $(X, Y(0), Y(1))$ are sampled from $Q$. In this work, we assume that $\pi \in \Pi$ where \[\Pi = \{\pi: \mathcal{X}\mapsto \{0, 1\}: \pi \text{ is measurable and deterministic}\}.\] The two key challenges of this setting include that the true study potential outcome distribution $P$ cannot be identified from the observed data, and furthermore, the target potential outcome distribution $Q$ is not the same as $P$. In particular, we assume $Q$ is unknown and generates $P$ under $\Gamma$-biased sampling, in the sense of Definition (ref). For any marginal covariate distribution $Q_{X},$ we define $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$ to be the set of all target potential outcome distributions that can generate $P_{\text{obs}}$ under $\Gamma$-biased sampling and an internally-valid RCT with treatment probability $e$, have covariate distribution $Q_{X}$, and are absolutely continuous with respect to Lebesgue measure. Let $\mathcal{T}(P_{\text{obs}})$ be the set of potential outcome distributions that can generate $P_{\text{obs}}$ under an internally-valid RCT with treatment probability $e$ and are absolutely continuous with respect to Lebesgue measure. Essentially, this is the set of couplings of $P_{X, Y(0)}$ and $P_{X, Y(1)}$ that satisfy (ref). In addition, we can define $\mathcal{R}^{\Gamma}(P, Q_{X})$ as the set of potential outcome distributions that can generate $P$ under $\Gamma$-biased sampling and have covariate distribution $Q_{X}$. While it is possible to place additional restrictions on the possible values of $P$ and $Q$, we focus on the most general setting where \begin{equation} \mathcal{S}^{\Gamma}(P_{obs}, Q_{X}) = \{ Q \mid Q \in \mathcal{R}^{\Gamma}(P, Q_{X}) where P \in \mathcal{T}(P_{obs})\}.\end{equation} We aim to learn policies that are robust to all target potential outcome distributions in $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X}).$ In the next subsection, we define objectives that yield different performance guarantees when optimized over $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X}).$ \end{subsection} \begin{subsection}{Objectives} We examine the robust policies given by three different objectives: max-min, max-min gain, and minimax regret. We refer to manski2011choosing for a detailed discussion of the max-min and minimax regret objectives. When learning robust policies, a natural first step is to consider the max-min policy: \begin{equation} \pi_{\Gamma, \text{maxmin}}^{*} \in \argmax_{\pi \in \Pi} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})} \EE[Q]{Y(\pi(X))}. \end{equation} The max-min policy yields the greatest lower bound on the mean outcome across all states of nature (under all distributions $Q \in \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$). The max-min objective is the objective that is studied most often in the literature on policy learning under partial identification adjaho2022externally, mufactored, si2020distributional. However, it is known to often be ultra-pessimistic savage1951theory because there may be distributions of $Q$ for which all policies perform poorly--and the max-min objective will focus on these instances. In some cases, a decision maker may want to protect against these worst-case scenarios and for that reason, the max-min objective may be reasonable choice manski2011choosing. As an alternative to the max-min objective, we also consider the policy that maximizes the worst-case gain over a baseline $\pi_{0}$, which is given by \begin{equation} \pi_{\Gamma, \text{gain}}^{*} \in \argmax_{\pi \in \Pi} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})} \EE[Q]{Y(\pi(X))} - \EE[Q]{Y(\pi_{0}(X))}. \end{equation} The max-min gain policy is a natural choice if we aim to improve robustness relative to a status quo policy. This type of objective has been recently considered by ben2021safe,kallus2021minimax. The minimax regret objective, introduced by savage1951theory, is closely related to the max-min gain objective. In minimax regret, we essentially aim to maximize the worst-case gain relative to the best possible baseline policy, a baseline policy that obtains the maximum mean outcome for every choice of $Q$. The minimax regret policy is defined to be \begin{equation} \pi^{*}_{\Gamma, \text{regret}} \in \argmin_{\pi \in \Pi} \sup_{Q \in \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})} R_{Q}(\pi(X)), \end{equation} where the regret of a policy $\pi(X)$ under a distribution $Q$ is given by \[ R_{Q}(\pi(X)) = \sup_{\pi' \in \Pi} \EE[Q]{Y(\pi'(X))} - \EE[Q]{Y(\pi(X))}.\] When considering policy classes that include non-deterministic policies, we often find that the minimax regret policy is often non-deterministic manski2011choosing. We emphasize that in this work, we define the minimax regret objective with respect to $\Pi$, the class of deterministic, unconstrained, binary-valued functions, so we aim to find the minimax regret policy among the class of deterministic functions. The minimax regret rule yields the least upper bound on the loss in mean outcome that results from not knowing $Q$ and is viewed as less pessimistic than the max-min rule manski2011choosing, savage1951theory. Nevertheless, due to the complex structure of the minimax regret policy, it can often be difficult to identify and learn the minimax regret policy. A key contribution of our work is that we identify a closed-form expression for the minimax regret policy under the class of deterministic, unconstrained, binary-valued policies. \end{subsection} \begin{subsection}{Results} Our identification results rely on the conditional average treatment effect (CATE) function and the conditional value-at-risk rockafellar2000optimization of the potential outcomes. We define the CATE below \begin{equation} \tau(x) = \EE[P]{Y(1)- Y(0) \mid X=x}. \end{equation} We consider the conditional value-at-risk at a particular quantile \begin{equation} \zeta(\Gamma) = \frac{1}{\Gamma + 1}.\end{equation} Recall that for a continuous random variable $Z$ with quantile function (inverse cdf) $q_{Z}$ and $\zeta \in (0, 1)$ \[ \text{CVaR}_{\zeta}(Z) = \EE[]{Z \mid Z \geq q_{\zeta}(Z)}.\] In addition, we recall that in the absence of sampling bias, the optimal policy to deploy is \begin{equation} \pi_{\text{non-robust}}(x) = \mathbb{I}(\tau(x) \geq 0),\end{equation} which treats units that have nonnegative conditional average treatment effect under $P$ kitagawa2018should. Note that this policy is identified under $P_{\text{obs}}$ because the CATE is identified under $P_{\text{obs}}.$ \begin{theo} Define \begin{equation} H_{\Gamma}(x) = (1 - \Gamma^{-1}) \cdot \left(\mathrm{CVaR}_{\zeta(\Gamma)}(Y(1) \mid x) - \mathrm{CVaR}_{\zeta(\Gamma)}(Y(0) \mid x)\right). \end{equation} The policy \begin{equation} \pi^{*}_{\Gamma, \mathrm{maxmin}}(x)=\mathbb{I}(\tau(x) \geq H_{\Gamma}(x)) \end{equation} solves (ref) for any $Q_{X}$ such that $Q_{X} \ll P_{\mathrm{obs}, X}$ and $\sup_{x \in \mathcal{X}} \frac{dP_{\mathrm{obs}, X}(x)}{dQ_{X}(x)} < \infty.$ \hyperref[subsec:opt_maxmin]{Proof in Appendix (ref).} \end{theo} We recall that $\Gamma=1$ corresponds to the case where there is no variation in the probability of sample selection due to unobservables. We find that (ref) reduces to $\pi_{\text{non-robust}}$ when $\Gamma=1.$ Interestingly, we note that when $\Gamma > 1$, $H_{\Gamma}(x)$ is not necessarily positive, meaning that (ref) may not necessarily yield a higher threshold for treatment than $\pi_{\text{non-robust}}.$ So, depending on the tail behavior of $Y(1), Y(0)$, the max-min rule may treat more units than $\pi_{\text{non-robust}}$ which assumes no sampling bias due to unobservables. We note that another interpretation of the max-min policy is that it compares the lower bounds on $\EE[Q]{Y(1) \mid X=x}$ and $\EE[Q]{Y(0) \mid X=x}$ over the robustness set and selects that treatment that yields the higher lower bound. \begin{lemm} The max-min policy $\pi_{\Gamma, \mathrm{maxmin}}^{*}$ can be written as \begin{equation} \mathbb{I}\left(\inf_{Q \in \mathcal{S}^{\Gamma}(P_{\mathrm{obs}}, Q_{X})} \EE[Q]{Y(1) \mid X=x} \geq \inf_{Q \in \mathcal{S}^{\Gamma}(P_{\mathrm{obs}}, Q_{X})} \EE[Q]{Y(0) \mid X=x})\right). \end{equation} \hyperref[subsec:maxmin_lower_bounds]{Proof in Appendix (ref).} \end{lemm} We compare and contrast the max-min policy to the policy that maximizes the worst-case gain over a baseline policy and the minimax regret policy. \begin{theo} Define \begin{align} H_{\Gamma}^{+}(x) &= (1 - \Gamma^{-1}) \cdot \left(\mathrm{CVaR}_{\zeta(\Gamma)}(Y(1) \mid x) + \mathrm{CVaR}_{\zeta(\Gamma)}(-Y(0)\mid x)\right) , \\ H_{\Gamma}^{-}(x) &= (1 - \Gamma^{-1}) \cdot \left(-\mathrm{CVaR}_{\zeta(\Gamma)}(-Y(1) \mid x) - \mathrm{CVaR}_{\zeta(\Gamma)}(Y(0) \mid x)\right). \end{align} The policy \begin{equation} \pi^{*}_{\Gamma, \mathrm{gain}}(x) = \mathbb{I}(\pi_{0}(x) = 0) \cdot \mathbb{I}(\tau(x) \geq H^{+}_{\Gamma}(x)) + \mathbb{I}(\pi_{0}(x) = 1) \cdot \mathbb{I}(\tau(x) \geq H^{-}_{\Gamma}(x))\end{equation} solves (ref) for any $Q_{X}$ such that $Q_{X} \ll P_{\mathrm{obs}, X}$ and $\sup_{x \in \mathcal{X}} \frac{dP_{\mathrm{obs}, X}(x)}{dQ_{X}(x)} < \infty.$ \hyperref[subsec:opt_gain]{Proof in Appendix (ref).} \end{theo} Note that when $\Gamma=1$, the optimal max-min gain policy defaults to $\pi_{\text{non-robust}}.$ When $\Gamma > 1$, the optimal max-min gain policy in (ref) has different thresholds for treatment depending on whether the baseline policy recommends treatment or control. Intuitively, we would expect (ref) to have a higher threshold for treatment when the baseline policy recommends control $(\pi_{0}(x) = 0)$ than when the the baseline policy recommends treatment $(\pi_{0}(x) = 1)$. The following lemma confirms this intuition by demonstrating that $H_{\Gamma}^{+}(x) \geq H_{\Gamma}^{-}(x).$ We also note that threshold for treatment of the max-min policy (ref) falls between $H^{-}_{\Gamma}(x)$ and $H^{+}_{\Gamma}(x).$ \begin{lemm} For $x \in \mathrm{supp}(P_{\mathrm{obs}, X}),$ \begin{equation} H_{\Gamma}^{-}(x) \leq H_{\Gamma}(x) \leq H_{\Gamma}^{+}(x).\end{equation} \hyperref[subsec:threshold_comparison]{Proof in Appendix (ref).} \end{lemm} An additional consequence of the above lemma is that if the baseline policy is given by $\pi_{0}(x) = \mathbb{I}(\tau(x) \geq b(x))$ where $H^{-}_{\Gamma}(x) \leq b(x) \leq H^{+}_{\Gamma}(x)$ for all $x \in \mathcal{X}$, then the baseline policy $\pi_{0}$ maximizes the max-min gain objective, i.e., \[ \pi_{\Gamma, \text{gain}}^{*} =\pi_{0}.\] Lastly, we consider the minimax regret rule. \begin{theo} Define $H_{\Gamma}^{+}(\cdot), H_{\Gamma}^{-}(\cdot)$ as in (ref), (ref). The policy \begin{equation} \pi^{*}_{\Gamma, \mathrm{regret}}(x) = \mathbb{I}\Big(\tau(x) \geq \frac{H^{+}_{\Gamma}(x) + H^{-}_{\Gamma}(x)}{2}\Big) \end{equation} solves (ref) for any $Q_{X} \ll P_{\mathrm{obs}, X}$ and $\sup_{x \in \mathcal{X}} \frac{dP_{\mathrm{obs}, X}(x)}{dQ_{X}(x)} < \infty.$ \hyperref[subsec:opt_regret]{Proof in Appendix (ref).} \end{theo} We give another characterization of the minimax regret policy that is related to the standard policy learning problem. In the absence of sampling bias, solving the following optimization problem \begin{equation} \sup_{\pi \in \Pi} \EE[P]{(2\pi(X) - 1) \cdot (Y(1) - Y(0))} \end{equation} yields $\pi_{\text{non-robust}}$, which is the optimal policy when there is no sampling bias zhao2012estimating. In the following theorem, we see that the maximizer of this objective over the robustness set yields the minimax regret policy. We find that robust optimization of the objective $\EE[Q]{v^{*}(\pi(X); X, Y(0), Y(1))} = \EE[Q]{(2\pi(X) - 1) \cdot (Y(1) - Y(0))}$ over the robustness set $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$ yields the minimax regret policy. \begin{theo} Let \[\tilde{v}_{\mathrm{regret}}(z; x, y_{0}, y_{1}) = (2z - 1) \cdot (y_{1}- y_{0}).\] The policy $\pi$ that solves \begin{equation} \sup_{\pi \in \Pi} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{\mathrm{obs}}, Q_{X})} \EE[Q]{\tilde{v}_{\mathrm{regret}}(\pi(X); X, Y(0), Y(1))} \end{equation} is equal to the optimal policy under the minimax regret objective (ref). \hyperref[subsec:po_regret]{Proof in Appendix (ref).} \end{theo} Note that the policies in (ref), (ref), (ref) can be estimated by first estimating the nuisance parameters $\tau(\cdot), H_{\Gamma}(\cdot), H_{\Gamma}^{+}(\cdot), H_{\Gamma}^{-}(\cdot)$ and then forming the defined policies. While this is one potential approach for learning policies, implementing this method may be onerous and give poor empirical performance. In the following sections, we consider an alternative approach for learning max-min and max-min gain policies. \end{subsection}
section{Equivalence} In this section, we prove an equivalence result that will link robust optimization over potential outcome distributions with robust optimization over observed data distribution. Such an equivalence is useful when we aim to learn robust policies using only data from $P_{\text{obs}}$ without knowledge of the true study population $P$. Robust optimization over potential outcome distributions can generally be formulated as the following problem \begin{equation} \sup_{\pi \in \Pi} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{obs}, Q_{X})} \EE[Q]{v^{*}(\pi(X); X, Y(0), Y(1))}.\end{equation} Note that $v^{*}$ is a potential outcome value function, which is defined in terms of $X, Y(0), Y(1).$ Now, we consider robust optimization over observed data distributions. Let $\mathcal{S}_{\text{obs}}^{\Gamma}(P_{\text{obs}}, Q_{X})$ be a robustness set that consists of observed data distributions. A distribution $Q_{\text{obs}} \in \mathcal{S}_{\text{obs}}^{\Gamma}(P_{\text{obs}}, Q_{X})$ if $Q_{\text{obs}, X} = Q_{X}$ and \begin{equation} \Gamma^{-1} \leq \frac{dQ_{obs, Y \mid W=w, X=x}(y)}{dP_{obs, Y \mid W=w, X=x}(y)} \leq \Gamma \quad \forall w \in \{0, 1\}, x \in \mathcal{X}, y \in \mathcal{Y}. \end{equation} We consider \begin{equation} \sup_{\pi \in \Pi} \inf_{Q_{obs} \in \mathcal{S}^{\Gamma}_{obs}(P_{obs}, Q_{X})} \EE[Q_{\text{obs}}]{v(\pi(X); X, Y, W)}.\end{equation} Note that $v$ is an observed data value function, which is defined in terms of $X, Y, W.$ We require the following assumption to establish the link between robust optimization over potential outcome distributions and robust optimization over observed data distributions. \begin{assumption} We assume that $v$ is defined so that \[\EE[P_{\text{obs}}]{v(\pi(X); X, Y, W)} = \EE[P]{v^{*}(\pi(X); X, Y(0), Y(1))}, \] for any $P \in \mathcal{T}(P_{\text{obs}})$ and arbitrary $\pi$. \end{assumption} \begin{theo} Suppose $v$ is a value function that satisfies Assumption (ref) for $v^{*}$. Then a policy $\pi$ that solves (ref) also solves (ref) for any $Q_{X} \ll P_{\text{obs}, X}$ and $\sup_{x \in \mathcal{X}} \frac{dP_{\mathrm{obs}, X}(x)}{dQ_{X}(x)}< \infty.$ \hyperref[subsec:equivalence]{Proof in Appendix (ref).} \end{theo} Theorem (ref) is a nontrivial and somewhat surprising result. While $Q\in \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$ implies $Q_{\text{obs}}\in \mathcal{S}_{\text{obs}}^{\Gamma}(P_{\text{obs}}, Q_{X})$, the converse is not true generally because $Q_{\text{obs}}$ does not carry any information on the copula between potential outcomes (conditional on $X$). Nonetheless, we show that the worse-case copula for this optimization problem yields a joint distribution in $\mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})$ using an argument that is analogous to Neyman-Pearson Lemma. We note that Theorem (ref) also holds when we optimize over continuous-valued functions $h$. Let $\mathcal{H} = L^{2}(P_{\text{obs}, X}, \mathcal{X}).$ We can extend Assumption (ref): we assume that $v, v^{*}$ are defined so that \[ \EE[P_{\text{obs}}]{v(h(X); X, Y, W)} = \EE[P]{v^{*}(h(X); X, Y(0), Y(1))}, \] for $P \in \mathcal{T}(P_{\text{obs}})$ and $h$ arbitrary. Under this assumption, a function $h$ that solves \begin{equation} \sup_{h \in \mathcal{H}} \inf_{Q_{\text{obs}} \in \mathcal{S}_{\text{obs}}^{\Gamma}(P_{\text{obs}}, Q_{X})} \EE[Q_{\text{obs}}]{v(h(X); X, Y, W)}. \end{equation} also solves \begin{equation} \sup_{h \in \mathcal{H}} \inf_{Q \in \mathcal{S}^{\Gamma}(P_{\text{obs}}, Q_{X})} \EE[Q]{v^{*}(h(X); X, Y(0), Y(1))}. \end{equation}
section{End-to-End Policy Learning via RU Regression} In this section, we first demonstrate that (ref) can be solved via RU Regression sahoo2022learning. Next, we define smooth observed data value functions can be used in combination with the RU Regression procedure to learn the optimal max-min and max-min gain policies. We also define an observed data value function to learn the optimal minimax regret policy, however it is nonconcave, which results in a nonconvex RU Regression problem. First, we demonstrate that RU Regression can be used to solve (ref). \begin{theo} Let $v(z;x, y, w)$ be a value function. For any $\Gamma > 1$, we define the following augmented loss function \begin{equation} L_{\mathrm{RU}}^{\Gamma}(z, a; x, y, w) = -\Gamma^{-1} v(z; x, y, w) + (1 - \Gamma^{-1}) \cdot a + (\Gamma - \Gamma^{-1})(-v(z;x, y, w) - a)_{+}. \end{equation} Then, any solution \begin{equation} \{ h_{\Gamma}^{*}, \alpha_{\Gamma}^{*} \} \in \argmin_{h, \alpha} \EE[P_{\mathrm{obs}}]{L_{RU}^{\Gamma}(h(X), \alpha(X, W); X, Y, W)} \end{equation} is also a solution to (ref) for any $Q_{\mathrm{obs}} \ll P_{\mathrm{obs}}$. \hyperref[subsec:ru_reg]{Proof in Appendix (ref).} \end{theo} \begin{rema} If $v(z; x, y, w)$ is concave in $z$ for any $(x, y, w) \in \mathcal{X} \times \mathcal{Y} \times \{0, 1\}$, then the augmented loss $L_{\mathrm{RU}}^{\Gamma}(z, a; x, y, w)$ is jointly convex in $(z, a)$ for any $(x, y, w) \in \mathcal{X} \times \mathcal{Y} \times \{0, 1\}.$ \end{rema} Now, it remains to define the appropriate observed data value functions $v$ that when plugged into the RU Regression procedure yield the optimal max-min and optimal max-min gain policies. \begin{subsection}{Observed Data Value Functions for Policy Learning} First, we define potential outcome value functions that when maximized yield the max-min policy, the max-min gain policy, and the minimax regret policy. \begin{theo} Define $\mathcal{H} = \{h \in L^{2}(P_{\mathrm{obs}, X}, \mathcal{X}) \mid 0 \leq h(x) \leq 1 \}.$ Let \[v^{*}_{\text{maxmin}}(z; x, y_{0}, y_{1}) = \log(1 + \exp(2z -1)) \cdot y_{1} + \log (1 + \exp(-2z + 1)) \cdot y_{0}.\] Then the policy $\pi(x) = \mathbb{I}\left(h^{*}(x) \geq \frac{1}{2}\right)$, where \begin{equation} h^{*} \in \argsup_{h \in \mathcal{H}} \inf_{Q \in \mathcal{S}_{\Gamma}(P_{\mathrm{obs}}, Q_{X})} \EE[Q]{v^{*}(h(X); X, Y(0), Y(1))}, \end{equation} is the optimal policy under the max-min objective (ref). \hyperref[subsec:po_maxmin]{Proof in Appendix (ref).} \end{theo} \begin{theo} Define $\mathcal{H} = \{h \in L^{2}(P_{\mathrm{obs}, X}, \mathcal{X}) \mid 0 \leq h(x) \leq 1 \}.$ Let \[v^{*}_{\text{gain}}(z; x, y_{0}, y_{1}) = (1 - \pi_{0}(x)) \log(1 + \exp(2z-1)) \cdot (y_{1} - y_{0}) + \pi_{0}(x) \log (1 + \exp(-2z + 1)) \cdot (y_{0} - y_{1}).\] Then $\pi(x) = \mathbb{I}\left(h^{*}(x) \geq \frac{1}{2}\right)$, where \begin{equation} h^{*} \in \argsup_{h \in \mathcal{H}} \inf_{Q \in \mathcal{S}_{\Gamma}(P_{\mathrm{obs}}, Q_{X})} \EE[Q]{v^{*}_{gain}(\pi(X); X, Y(0), Y(1))} \end{equation} is the optimal policy under the max-min gain objective (ref). \hyperref[subsec:po_gain]{Proof in Appendix (ref).} \end{theo} \begin{theo} Define $\mathcal{H} = \{h \in L^{2}(P_{\mathrm{obs}, X}, \mathcal{X}) \mid 0 \leq h(x) \leq 1 \}.$ Let \[v^{*}_{\text{regret}}(z; x, y_{0}, y_{1}) = (2z-1)(y_{1} - y_{0}) + \log\left(\left|2z- 1\right|\right).\] Then $\pi(x) = \mathbb{I}(h^{*}(x) \geq \frac{1}{2})$, where \begin{equation} h^{*} \in \argsup_{h \in \mathcal{H}} \inf_{Q \in \mathcal{S}_{\Gamma}(P_{\mathrm{obs}}, Q_{X})} \EE[Q]{v^{*}_{regret}(\pi(X); X, Y(0), Y(1))} \end{equation} is the optimal policy under the minimax regret objective (ref). \hyperref[subsec:po_regret_loss]{Proof in Appendix (ref).} \end{theo} In the following lemma, we define observed data value functions that satisfy Assumption (ref) for $v^{*}_{\text{maxmin}}, v^{*}_{\text{gain}}, v^{*}_{\text{regret}}.$ \begin{lemm} Define \begin{equation} v_{maxmin}(z; x, y, w) = \log(1 + \exp((2z - 1) \cdot (2w-1))) \cdot \Big( \frac{y}{w e + (1-w)(1-e)} \Big). \end{equation} \begin{equation} \begin{aligned} v_{gain}(z; x, y, w) &= (1 - \pi_{0}(x)) \log(1 + \exp(2z-1)) \cdot \Big(\frac{y \cdot w}{e} - \frac{y \cdot (1 - w)}{1 -e} \Big) \\ &+ \pi_{0}(x) \log(1 + \exp(-2z+1)) \cdot \Big(\frac{y \cdot (1-w)}{1 -e} - \frac{y \cdot w}{e}\Big). \end{aligned} \end{equation} \begin{equation} v_{regret}(z; x, y, w) = (2z-1) \cdot \Big(\frac{y \cdot w}{e} - \frac{y \cdot (1 - w)}{1 -e}\Big) + \log\left(\left|2z - 1\right|\right) . \end{equation} The value functions $v_{\text{maxmin}}, v_{\text{gain}}, v_{\text{regret}}$ satisfy Assumption (ref) with $v^{*}_{\text{maxmin}}, v^{*}_{\text{gain}}, v^{*}_{\text{regret}}$, respectively. \hyperref[subsec:equal_value]{Proof in Appendix (ref).} \end{lemm} Since the the observed data value functions $v_{\text{maxmin}}, v_{\text{gain}}$ are smooth and concave in $z$, we can combine the results from Section (ref) and Theorem (ref), to show that finding the optimal max-min policy and optimal max-min gain policy amounts to solving convex RU Regression problems. This allows us to learn these policies directly without needing to separately estimating multiple nuisance parameters. We propose to parametrize the functions $h, \alpha$ using neural networks and train them with the RU loss (ref) jointly. In contrast, we realize that $v_{\text{regret}}$ is nonconcave in $z$ and has a singularity at $z=\frac{1}{2}$. So, RU Regression with $v_{\text{regret}}$ as the observed data value function is a nonconvex problem and may be difficult to solve via continuous optimization methods. We leave the development of an algorithm to solve the minimax regret problem for future work. \end{subsection}
section{Experiments} In this section, we evaluate the learning procedures proposed in Section (ref). First, we describe how to implement RU Regression for policy learning. Second, we visualize the max-min policy and the max-min gain policy learned via RU Regression on a one-dimensional synthetic dataset and compare these policies to the true policies given by (ref) and (ref). We also examine the true minimax regret policy given by (ref). Third, in a semi-synthetic experiment, we compare the behavior of the max-min and max-min gain policies to the non-robust policy (ref). \begin{subsection}{Implementation of RU Regression for Policy Learning} \begin{figure} \caption{RU Regression architecture for policy learning. We note that the function $h$ takes $X$ as input and the auxiliary function $\alpha$ takes $X, W$ as input.} \end{figure} There are a few subtle differences between the RU Regression implementation in this work and that of sahoo2022learning. From Theorem (ref), the auxiliary function $\alpha$ in the policy learning setting depends on covariates $X$ and the treatment assignment $W$, while in the regression setting of sahoo2022learning, $\alpha$ only depends on $X$. This is reflected in our architecture for learning $h, \alpha$ (Figure (ref)). We note that the value functions $v_{\text{maxmin}}$ in (ref) and $v_{\text{gain}}$ in (ref) are unbounded above, so they do not have maximizers over unbounded function classes. To ensure that the RU loss (ref) is not unbounded below when $v_{\text{maxmin}}$ or $v_{\text{gain}}$ are plugged into (ref), we perform the optimization of $h$ over a class of bounded functions. In particular, we aim to learn a regression function $h_{\Gamma}(x): \mathcal{X} \rightarrow [0, 1].$ Then we determine treatment assignment by computing \[ \pi_{\Gamma}(x) = \mathbb{I}\Big(h_{\Gamma}(x) \geq \frac{1}{2}\Big).\] Since the policy only depends on whether $h$ is greater or less than $\frac{1}{2}$, the restriction of $h$ to the class of bounded functions does not impact the optimal policy value. To enforce that $h$ must have outputs in $[0, 1]$, we add a sigmoid activation as the final activation function of the neural network that represents $h$, which ensures that the network outputs values in $[0, 1].$ \end{subsection} \begin{subsection}{Toy Experiment} First, we consider a simple one-dimensional experiment, so that we can visualize the data distributions and compare the learned policies to the optimal policies computed under the data model. \begin{subsubsection}{Data} \begin{figure} \caption{We visualize the potential outcome distribution as $p$ varies. When $p$ is low, our potential outcome distribution mostly consists of units for which $Y_{i}(1) > Y_{i}(0)$. As $p$ increases, a larger fraction of units have $Y_{i}(0) > Y_{i}(1).$ } \end{figure} We consider a toy example where the study and target potential outcome distribution are generated as follows. Let $U_{i}$ be an unobserved variable that impacts $Y_{i}(1).$ \begin{equation} \begin{split} X_{i} &\sim Uniform[-3, 3] \\ U_{i} &\sim Bernoulli(p) \\ Y_{i}(0) \mid X_{i} &\sim N(\sin(X_{i}) \,, \sigma^{2}) \\ Y_{i}(1) \mid X_{i} &\sim N(\sin(X_{i}) + 1.5 - U_{i} \cdot (5X_{i})_{+} \,, \sigma^{2}), \end{split} \end{equation} where $\sigma=0.2.$ We visualize the potential outcome distributions in Figure (ref). We generate study and target potential outcome distributions by varying the Bernoulli parameter $p$ of the distribution of $U_{i}$. Let the study potential outcome distribution $P$ correspond to the potential outcome distribution where $p_{\text{study}}=0.2$. Given $P$, we generate the observed data distribution $P_{\text{obs}}$ over $(X_{i}, Y_{i}, W_{i})$ as follows \[ W_{i} \sim \text{Bernoulli}(0.5), \quad Y_{i} = Y_{i}(W_{i}).\] The target potential outcome distributions $Q$ are generated by varying $p_{\text{target}} \in [0, 1].$ \end{subsubsection} \begin{subsubsection}{Policies} We use RU Regression, with the implementation described in Section (ref), to learn the optimal max-min and max-min gain policies with access only to data from the observed data distribution. Since this is a one-dimensional synthetic example, we also use the data model (ref) and the closed-form formula (ref) to compute the minimax regret policy. We evaluate the following 4 following policies. \begin{enumerate} • Max-Min Policy (learned via RU Regression). • Max-Min Gain Policy with Always Control Baseline (learned via RU Regression). • Max-Min Gain Policy with Always Treat Baseline (learned via RU Regression). • Minimax Regret (computed using (ref) and (ref)). \end{enumerate} \end{subsubsection} \begin{subsubsection}{Results} \begin{figure} \caption{Max-min policy results for one-dimensional toy example. Left: We visualize the true max-min policies, computed using the data model (ref). Middle Left: For one random seed, we visualize the learned maxmin policy. Middle Right: We visualize the mean outcome of the true policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation. Right: We visualize the mean outcome of the learned policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation.} \end{figure} \begin{figure} \caption{Max-min gain (over always control) policy results for one-dimensional toy example. \textbf{Left}: We visualize the true max-min gain policies where the baseline is $\pi_{0}(X)= 0$ (always control), computed using the data model (ref). \textbf{Middle Left}: For one random seed, we visualize the learned maxmin gain over always control policy. \textbf{Middle Right}: We visualize the mean outcome of the true policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation. \textbf{Right}: We visualize the mean outcome of the learned policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation.} \end{figure} \begin{figure} \caption{Max-min gain (over always treat) policy results for one-dimensional toy example. \textbf{Left}: We visualize the true max-min gain policies where the baseline is $\pi_{0}(X)= 1$ (always treat), computed using the data model (ref). \textbf{Middle Left}: For one random seed, we visualize the learned maxmin gain over always control policy. \textbf{Middle Right}: We visualize the mean outcome of the true policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation. \textbf{Right}: We visualize the mean outcome of the learned policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation.} \end{figure} \begin{figure} \caption{Minimax regret policy results for one-dimensional toy example. \textbf{Left}: We visualize the true minimax regret policies, computed using the data model (ref). \textbf{Right}: We visualize the mean outcome of the learned policy for target distributions with different $p_{\text{target}}$, aggregated over 6 random trials, where the randomness is over dataset generation.} \end{figure} We use the data model in (ref) to compute the nuisance parameters $\tau(\cdot), H_{\Gamma}(\cdot), H_{\Gamma}^{+}(\cdot), H_{\Gamma}^{-}(\cdot)$ and plug them into (ref) and (ref) to compute the true optimal policies. A comparison of the left and middle left plots of Figures (ref), (ref), (ref) reveals that our learned policies match the true policies. A comparison of the middle right and right plots of Figures (ref), (ref), (ref) demonstrates that over 6 random trials, the learned policies and the true policies achieve similar mean outcome. Note that $\pi_{\text{non-robust}}$ from (ref) is equal to the optimal max-min, max-min gain, or minimax regret policies when $\Gamma=1$. The non-robust policy recommends to treat units with nonnegative CATE. This strategy performs well when $p_{\text{target}} < 0.3$ but it is outperformed by the max-min policy and the max-min gain over always control policy when $p_{\text{target}} > 0.3.$ We note that the max-min gain over always treat policy does not deviate from the baseline for $\Gamma> 1.$ Somewhat surprisingly, in this particular example, the minimax regret policy is less conservative that $\pi_{\text{non-robust}}$ (Figure (ref)) because it recommends treatment to a higher proportion of the population than $\pi_{\text{non-robust}}$ as $\Gamma$ increases. \end{subsubsection} \end{subsection} \begin{subsection}{High-Dimensional Experiment} Second, we evaluate our methods in a high-dimensional $(d=10)$ simulation. \begin{subsubsection}{Data} We consider a high-dimensional $(d=10)$ example where the study and target potential outcome distribution are generated as follows. Let $U_{i}$ be an unobserved variable that impacts $Y_{i}(1).$ \begin{equation} \begin{split} X_{i} &\sim \text{Uniform}[-3, 3]^{d} \\ U_{i} &\sim \text{Bernoulli}(p) \\ Y_{i}(0) \mid X_{i} &\sim N(\sin(a^{T} X_{i}) \,, \sigma^{2}) \\ Y_{i}(1) \mid X_{i} &\sim N(\sin(a^{T}X_{i}) + 1.5 - U_{i} \cdot (2 (X_{i, 1} + X_{i, 2} + X_{i, 3}))_{+} \,, \sigma^{2}), \end{split} \end{equation} where $\sigma=0.2$ and $a$ is a constant vector defined in Section (ref). \end{subsubsection} \begin{subsubsection}{Policies} We evaluate the following policies. \begin{enumerate} • Max-Min (learned via RU Regression) • Max-Min Gain over Baseline Policy (learned via RU Regression) \[ \pi_{0}(X) = \mathbb{I}(X_{1} \leq 0).\] \end{enumerate} \end{subsubsection} \begin{subsubsection}{Results} \begin{figure} \caption{Max-min policy results from high-dimensional simulation.} \end{figure} \begin{figure} \caption{Max-min gain policy results from high-dimensional simulation. We note that when $\Gamma=1, 2$, the max-min gain policy does not deviate from $\pi_{\text{non-robust}}$. For higher values of $\Gamma$, the max-min gain policy outperforms $\pi_{\text{non-robust}}$ for high values of $p_{\text{target}}.$} \end{figure} As in the toy experiments, we observe that when $\Gamma=1$ the optimal robust policies corresponds to $\pi_{\text{non-robust}}.$ As before, we observe that $\pi_{\text{non-robust}}$ performs well when $p_{\text{target}}$ is small but deteriorates in performance for larger values of $p_{\text{target}}.$ In contrast, when $\Gamma> 1$, the max-min (Figure (ref)) and max-min gain (Figure (ref)) policies are more robust to changes in $p_{\text{target}}$. \end{subsubsection} \end{subsection} \begin{subsection}{Semi-Synthetic Experiment with Voting Dataset} We compare the behavior of the max-min and max-min gain policies in a semi-synthetic experiment, where we use the voting dataset of gerber2008social. \begin{subsubsection}{Data} The voting dataset of gerber2008social was collected from a randomized controlled trial-style study that aimed to effect of various actions on voter turnout. The researchers designed one control and four treatment actions that involved mailing the selected units a letter ahead of the 2006 Michigan primary election. In the original field experiment, the probability of the control action was $\frac{5}{9}$ and the probability of each treatment action was $\frac{1}{9}.$ In our experiment we focus on only on the “Control” action $(W_{i}=0)$ and the “Neighbors” action ($W_{i}=1$). Note that this implies that the probability of treatment is $e=\frac{1}{9} / (\frac{1}{9} + \frac{5}{9}) = \frac{1}{6}$. The “Neighbors” action entailed mailing a letter to the individual that contained voting participation records of the individual, the other members of the individual’s household, as well as the neighbors. The letter also mentioned a follow-up letter will be sent after the election with everyone’s updated participation, so the individual’s participation will be made known among the neighbors. The “Control” action sends no letter to the individual. In our experiment, the covariates $X_{i}$ include the household size of the individual, age, sex, and whether the individual $i$ voted in the previous primary elections from 2000-2002 and the previous general elections from 2000-2004. The outcome $Y_{i}$ is whether individual $i$ voted in the 2006 primary election. Although the dataset also contains information on whether individual $i$ voted in the 2004 primary election, we treat this as an unobserved variable $U_{i}$ and use it generate synthetic target and study populations. Through a dataset generation procedure described in Appendix (ref), we create a combined training and validation dataset where 67$\%$ of the units have $U_{i}=1$. In contrast, units with $U_{i}=1$ make up the only 18$\%$ of the test dataset. The training, validation, and test sets consist of 62044, 41364, and 126036 samples. By design, units who voted in the 2004 election are over-represented in the study population compared to the target population. We fit policies using data from the synthetic study population and evaluate the mean outcome under the synthetic target population. \end{subsubsection} \begin{subsubsection}{Policies} We evaluate the following policies. \begin{enumerate} • Max-Min (learned via RU Regression) • Max-Min Gain with Always Control Baseline (learned via RU Regression) \[ \pi_{0}(X) = 0.\] • Max-Min Gain with Always Treat Baseline (learned via RU Regression) \[ \pi_{0}(X) = 1.\] \end{enumerate} \end{subsubsection} \begin{subsubsection}{Results} Recall that as before, $\Gamma=1$ corresponds to $\pi_{\text{non-robust}}$, so we report results for $\pi_{\text{non-robust}}$ (learned via RU Regression with $v_{\text{maxmin}}$ and $\Gamma=1$). We evaluate the robust policies learned for $\Gamma=1.1, 1.2, 1.3, 1.5, 2, 3, 4$. We simply report results for $\Gamma=1.1, 1.2, 1.3, 1.5$ because when $\Gamma \geq 1.5$ we observe the same results as when $\Gamma=1.5$. This can be explained by the fact that the outcomes $Y_{i}(0), Y_{i}(1)$ are binary-valued, so $q_{\zeta(\Gamma)}(Y(w) \mid X=x)$ may take on the same value for many values of $\Gamma$. Note that our theoretical results only apply to the case where the conditional study potential outcome distribution is absolutely continuous with respect to Lebesgue measure, but in this experiment, the study potential outcome distribution is a discrete distribution. Nevertheless, we observe results that are in line with our results for continuous-valued outcomes. We report the target policy value and the proportion of target population treated under the different learned policies in Table (ref). We find that the non-robust policy recommends recommends to treat about 66% of the population. We find that the max-min and max-min gain over always treat baseline policy recommends to treat the entire population. In contrast, the max-min gain over the always control baseline recommends to treat a decreasing fraction of the population as $\Gamma$ increases. In this particular example, we observe that different robust objective functions can yield very different policies. \begin{table} \scriptsize \begin{tabular}{|c|c|c|} \hline Method & \multicolumn{1}{c}{Target Policy Value} & Proportion of Target Population Treated\\ \hline Non-Robust & 0.3106 $\pm$ 0.003 & 0.66\\ \hline Max-Min $(\Gamma=1.1)$ & 0.3272 $\pm$ 0.003 & 0.87 \\ Max-Min $(\Gamma=1.2)$ & 0.3375 $\pm$ 0.004 & 1.0 \\ Max-Min $(\Gamma=1.3)$ & 0.3375 $\pm$ 0.004 & 1.0 \\ Max-Min $(\Gamma=1.5)$ & 0.3375 $\pm$ 0.004 & 1.0 \\ \hline Max-Min Gain over Always Treat $(\Gamma=1.1)$ & 0.3375 $\pm$ 0.004 & 1.0 \\ Max-Min Gain over Always Treat $(\Gamma=1.2)$ & 0.3375 $\pm$ 0.004 & 1.0\\ Max-Min Gain over Always Treat $(\Gamma=1.3)$ & 0.3375 $\pm$ 0.004 & 1.0\\ Max-Min Gain over Always Treat $(\Gamma=1.5)$ & 0.3375 $\pm$ 0.004 & 1.0\\ \hline Max-Min Gain over Always Control $(\Gamma=1.1)$ & 0.2934 $\pm$ 0.003 & 0.44 \\ Max-Min Gain over Always Control $(\Gamma=1.2)$ & 0.2780 $\pm$ 0.002 & 0.25\\ Max-Min Gain over Always Control $(\Gamma=1.3)$ & 0.2705 $\pm$ 0.002 & 0.15\\ Max-Min Gain over Always Control $(\Gamma=1.5)$ & 0.2653 $\pm$ 0.001 & 0.0\\ \hline \end{tabular} \caption{The non-robust policy recommends to treat 66% of the population. The max-min and max-min gain over always treat policies recommend to treat the entire population. In contrast, the max-min gain over always control policies recommends treatment to a much smaller proportion of the population.} \end{table} \end{subsubsection} \end{subsection}