Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,350 characters · 23 sections · 62 citation commands
Policy Learning under Unobserved Confounding: A Robust and Efficient Approach
\and Gaoqian Xu} \and Xi Zheng} \and Yahong Zhou} }
Policy targeting, allocating interventions based on individual characteristics, has become increasingly influential in applied economics and econometrics. For example, a local government may need to decide which workers receive job training, a school district may identify students most in need of educational support, and a regulatory agency has to determine which workplaces warrant safety inspections.
Policymakers often rely on observational data to guide targeting decisions. However, unobserved confounding, unobserved factors that simultaneously affect both treatment and potential outcomes, can severely bias causal inference from such data. Ignoring these confounders can result in misleading estimates and suboptimal policy decisions. For example, johnson2023improving note that safety inspections tend to target workplaces with a history of accidents or informal complaints, events that may reflect unobserved traits, such as poor safety culture or weak management, which also affect injury risk. Given the potentially large benefits of well-targeted inspections,\footnote{Workplace injuries and illnesses place a significant economic burden on the United States, with annual social costs estimated at \$250 billion leigh2011economic.}designing targeting without accounting for unobserved confounding may misidentify high-risk workplaces, leading to the inefficient use of limited regulatory resources. Similar challenges arise in education, where observational data guide targeting decisions for specific academic programs. However, actual enrollment often hinges on parental choice, which is based on factors inaccessible to policymakers, such as a child’s latent abilities, behavioral traits, or home environment. This self-selection process based on unobserved characteristics complicates the design of effective targeting policies.
Many existing policy learning (or targeting) methods using observational data assume no unobserved confounding, that is, treatment assignment is random conditional on the observed covariates; see kitagawa2018a,athey2021policy. This assumption, commonly referred to as unconfoundedness, is often unreasonable, as historical policies or decisions may have been based on additional, unobserved information. When this assumption breaks down, researchers often turn to instrumental variable (IV) methods, which enable causal inference and policy targeting; see cui2021semiparametric,qiu2021optimal,sasaki2024welfare. Yet in practice, identifying a valid and credible instrument is often challenging. Moreover, even with a valid binary instrument, heterogeneous treatment effects are typically only partially identified, thereby limiting the ability to effectively target interventions.
As a supplement to IV methods, sensitivity analysis is used to assess the impact of unobserved confounding on policy effects, as originally proposed by rosenbaum1983assessing. Building on this idea, kallus2021minimax incorporate sensitivity analysis into policy learning. Leveraging the marginal sensitivity model (MSM) of tan2006distributional, they relax the unconfoundedness assumption and formulate a robust optimization framework to learn policies that remain effective under unobserved confounding. Specifically, building on the framework proposed by zhao2019sensitivity, they construct the objective function as the largest IPW estimator of expected welfare, where the putative propensity scores are restricted by the MSM. While the method offers robustness to unobserved confounding, it is not without limitations. First, as shown by dorn2023sharp, it may yield overly conservative estimates of the upper bounds on expected welfare. Second, the method requires propensity scores to be estimated at the parametric rate, which precludes the use of many popular machine learning methods, such as random forests and LASSO. Third, the policy optimization algorithm is tailored to differentiable parametric policy classes and relies on a computationally intensive iterative routine.
This work proposes a method for policy learning that is robust to unobserved confounding, serving as a complement to existing IV-based approaches. Building on the MSM framework, we construct a distributionally robust welfare criterion that captures the worst-case policy value over a plausible set of counterfactual outcome distributions. Based on this criterion, we develop a two-stage policy learning procedure: the first stage estimates nuisance components using flexible nonparametric or modern ML methods; the second stage optimizes the estimated criterion over a constrained policy class. The resulting procedure is both computationally efficient and straightforward to implement.
In (ref), building on dorn2023sharp and dorn2024doubly, we derive closed-form expressions for the worst-case (i.e., sharp lower bounds of) average welfare and welfare improvement\footnote{These are commonly used objective functions in policy learning, which are formally defined in (ref).} compatible with the MSM. Given the identification of these two robust criteria/objectives using moment conditions, we can construct the doubly robust estimation of them, even when the nuisance components are estimated at moderately slow rate, i.e., $o_P(n^{-1/4})$.
(ref) presents algorithms for learning optimal confounding-robust policies, where the policy class may be constrained by practical considerations such as implementability, cost, and interpretability. Our method follows the mainstream two-stage paradigm of athey2021policy, zhou2023offline. Unlike kallus2021minimax, our second-stage policy optimization supports standard machine learning methods and readily available software packages, such as logistic regression, policy trees, and neural networks, to learn the optimal policy by maximizing the estimated criterion. For example, given the doubly robust scores, our method can be implemented using the widely adopted R-package policytree sverdrup2020policytree. Moreover, (ref) provides asymptotic upper regret bounds as performance guarantees for our estimated confounding-robust policies.
We validate our method via simulations in (ref), comparing it to kallus2021minimax and the empirical welfare maximization (EWM) of athey2021policy. We then demonstrate its practical relevance in (ref) by applying it to two classic empirical settings where unobserved confounding is a primary concern, focusing on the contrast with the EWM.
Our applications suggest that ignoring such confounding may lead to misguided policy recommendations. First, in the National Job Training Partnership Act (JTPA) study, our method refines the naive EWM policy of treating nearly all participants, instead identifying more selective policies that target subgroups with higher education but lower prior earnings as concerns for unobserved confounding grow. In our Head Start application, the EWM policy assigns a mere 1.23% of slots to healthier, higher-income children. In contrast, our robust method prioritizes disadvantaged children with lower family income and maternal education, aligning with the program's mission. These findings demonstrate that accounting for unobserved confounding can fundamentally alter policy targeting, highlighting the importance of robust methods in practice.
Our work contributes to the growing literature on policy learning. Many policy learning works focus on mean-optimal policies under unconfoundedness kitagawa2018a,athey2021policy,zhou2023offline. Alternative objectives under the same assumption have also been explored, including quantile-optimal policies wang2018quantile,leqi2021median as well as those motivated by fairness kitagawa2021equality,fang2023fairness,viviano2023fair,fan2025policy.
To address unobserved confounding, several studies have proposed IV-based approaches for policy targeting. To avoid partial identification of the welfare function, sasaki2024welfare assume the availability of a continuous instrument with sufficiently large support and identify the average welfare via the marginal treatment effect (MTE). A similar approach has been applied to the design of optimal encouragement rules under endogeneity; see chen2022personalized,liu2022policy. In the statistics literature, other works investigate optimal treatment rules or encouragement interventions using binary instruments; see cui2021semiparametric,qiu2021optimal. However, these methods rely on quite stringent assumptions for a valid IV to point identify the conditional average treatment effect (CATE).
Binary instruments, though widely used, typically allow identification only of the local average treatment effect (LATE) for the subpopulation of compliers; see imbens1994identification,angrist1996identification,athey2017state. This presents major challenges for policy learning: the resulting policy intervention is effective only for compliers, a subpopulation that cannot be identified in advance in a new sample, thus limiting external validity. Moreover, compliance status is instrument-dependent, leading to instability and lack of generalizability across different IV designs. To address this issue, pu2021estimating and d2021orthogonal, propose methods for binary IV settings that partially identify heterogeneous treatment effects, learning robust policies by optimizing policy criteria constructed from IV-identified bounds on the CATE.
A related line of work studies distributionally robust policy learning, where uncertainty sets, often defined via Wasserstein distance or Kullback–Leibler divergence, are used to estimate policies that maximize worst-case expected outcomes adjaho2022externally, si2023distributionally, qi2023robustness. However, these approaches are often overly conservative and lack interpretability, as they do not explicitly model the source of distributional shift. lei2023policy address external validity under sample selection bias but do not explicitly account for latent confounding in treatment assignment. Finally, our work also connects to the growing literature on sensitivity analysis, which investigates how violations of unconfoundedness affect causal identification and inference; see tan2006distributional, masten2018identification, zhao2019sensitivity,rambachan2022counterfactual, dorn2023sharp, dorn2024doubly, masten2025general.
We consider observational data consisting of random samples $\{Z_i\}_{i=1}^n=\{X_i, Y_i, A_i\}_{i=1}^n$, where $X_i\in \mathcal{X} \subseteq \mathbb{R}^d$ represents the observed covariates, $A_i \in \{0,1\}$ denotes a binary intervention/treatment, and $Y_i \in \mathbb{R}$ is the real-valued observed outcome. Within the Neyman-Rubin potential outcomes framework, let $Y_i(0)$ and $Y_i(1)$ denote the potential outcomes under control ($A_i = 0$) and treatment ($A_i = 1)$, respectively. The observed outcome satisfies $Y_i = Y_i(A_i)$. Throughout this work, we interpret $Y_i$ as a measure of welfare or utility, where higher values correspond to more desirable outcomes.
Our objective is to use the observational data $\{ Z_i \}_{i=1}^n$ to guide personalized policy interventions in settings where unobserved confounding may be present.
Our analysis builds on the marginal sensitivity model (MSM) introduced by tan2006distributional to control the selection bias caused by unobserved confounders. This model relaxes the unconfoundedness assumption by allowing for unobserved confounders $U \in \mathbb{R}^k$, where $k\in \mathbb{N}^+$ is unknown. Here, we assume that $U$ is a vector but have no prior knowledge of its dimension or other characteristics. We denote $P_{o}$ as the true distribution of $O \equiv \left(X,Y(1), Y(0), A,U \right) $. The true propensity score is given by $e_o(x, u) =\mathbb{P}_{P_{o}}[A = 1 | X = x, U = u]$, which accounts for both observed covariates and unobserved confounders. In contrast, the nominal propensity score, defined as $e(x) = \mathbb{P}_{P_{o}}[A = 1 | X = x]$. In the MSM, unobserved confounders are assumed to have a bounded influence on the odds of treatment assignment, restricting the extent of selection bias introduced by unobserved confounders.
We now present the formal specification of the MSM, a widely used framework for sensitivity analysis in the causal inference literature (see, e.g., zhao2019sensitivity, dorn2023sharp, oprescu2023b, dorn2024doubly).
In practice, the sensitivity parameter $\Lambda$ in (ref) is selected by policymakers based on their prior beliefs regarding the extent of unobserved confounding and the resulting selection bias. Implicitly, it is assumed that policymakers choose $\Lambda$ such that (ref) is satisfied. The special case of $\Lambda=1$ corresponds to the unconfoundedness, also known as the selection-on-observables assumption. Higher values of $\Lambda$ impose fewer restrictions on the degree of unobserved confounding, allowing for greater potential selection bias. Calibration methods can be used to select the sensitivity parameter; see hsu2013calibrating.
In this subsection, we provide a brief review of policy learning under the assumption of unconfoundedness ($\Lambda = 1$). We then explore how to learn a robust policy when this assumption is violated, within the framework of the marginal sensitivity model. Throughout the rest of this work, we fix the sensitivity parameter $\Lambda \geq 1$.
A possibly randomized policy $\pi$ maps covariates $x \in \mathcal{X}$ to treatment assignment probabilities, with $\pi(x)$ denoting the probability of receiving treatment $A = 1$. The policymaker can pre-specify a policy class $\Pi$, defined as a collection of Borel measurable functions mapping from $\mathcal{X}$ to $[0,1]$. This class typically incorporates application-specific constraints, including budgetary limitations, structural assumptions, and fairness requirements; see (ref) for concrete examples.
We begin by recalling the framework of policy learning under the unconfoundedness assumption. Given a predefined policy class $\Pi$, the mean-optimal policy learning aims to identify a policy that maximizes the expected outcome:
where $Y\left(\pi \left(X \right)\right) = \pi(X)Y(1) + (1-\pi(X))Y(0)$. For simplicity, let $W_1: \pi \mapsto \mathbb{E} \left[ Y\left(\pi(X) \right) \right]$ denote the expected welfare function. Moreover, the regret of deploying a policy $\pi \in \Pi$ is defined as
It is evident that a policy maximizing $W_1(\pi)$ simultaneously minimizes the regret given in (ref).
A central objective in policy learning is policy evaluation, that is, identifying and estimating the criterion function. Once an estimate of $W_1(\cdot)$ is available, the optimal policy can be estimated by maximizing the estimated welfare over the policy class. Unconfoundedness (i.e., $\Lambda = 1$) is a widely adopted assumption for identifying the expected welfare function; see kitagawa2018a,athey2021policy, zhou2023offline. Under this assumption, various methods such as inverse probability weighting (IPW) can be employed to identify and estimate $W_1(\pi)$. In particular, the expected welfare function can be expressed as
where $e(x)$ is the nominal propensity score, and $\tau(x) = \mathbb{E}[Y(1) - Y(0) \mid X = x]$ is the conditional average treatment effect (CATE). When $\Lambda = 1$, the outcome regression, nominal propensity score, and CATE are all identifiable.
When unconfoundedness fails (i.e., $\Lambda > 1$), the expected welfare $W_1(\pi)$ becomes unidentifiable, and policies based on (ref) or (ref) lack reliable guarantees. As a result, $W_1(\pi)$ is not a suitable criterion for policy learning in the presence of unobserved confounding.
We adopt a distributionally robust optimization (DRO) framework for policy learning, replacing $W_1(\pi)$ with its worst-case counterpart over distributions consistent with MSM specified in (ref). Within this framework, we construct two welfare criteria that are both robust to unobserved confounding and identifiable, and derive the policy by optimizing one of them.
To implement the DRO, we first characterize the distributional uncertainty set over the counterfactual distribution of $\left(Y(1), Y(0), X, U, A\right)$ subject to the constraints imposed by (ref). Formally, let $\mathcal{P}(\Lambda)$ denote the set of all probability distributions $Q$ on $\mathbb{R}^2 \times \mathcal{X}\times \{0,1\} \times \mathbb{R}^k$ satisfying the following conditions:
We now present two complementary policy learning methods that are robust to unobserved confounding under the MSM framework. Our first approach is the max-min expected welfare method. Specifically, the {\it worst-case welfare function} $W_{\Lambda}:\Pi\rightarrow\mathbb{R}$ is defined as \[ W_{\Lambda}(\pi)=\inf_{ Q \in \mathcal{P}(\Lambda) } \mathbb{E}_Q \left[ Y(\pi (X) )\right]. \] The corresponding max-min welfare (MMW) policy is given by
Under the worst-case welfare function $W_\Lambda$, we define the confounding-robust welfare regret (CRW-regret) of a policy $\pi \in \Pi$, relative to the best possible policy in $\Pi$, as
The second approach builds on the conditional average treatment effect (CATE). When $\Lambda = 1$, unconfoundedness holds and, the policy learning problem reduces to maximizing the policy improvement function $\mathbb{E}[\tau(X)\pi(X)]$. To handle the more general case with $\Lambda \geq 1$, we extend this objective to the worst-case policy improvement function $\Delta_{\Lambda}: \Pi \rightarrow \mathbb{R}$, defined as \[ \Delta_{\Lambda}(\pi)=\inf_{ Q \in \mathcal{P}(\Lambda) } \mathbb{E}_Q \left[ \pi(X) ( Y(1) - Y(0) ) \right]. \] The corresponding max-min improvement (MMI) policy is given by
Under the criterion $\Delta_\Lambda$, we define the confounding-robust policy improvement regret (CRI-regret) of a policy $\pi \in \Pi$, relative to the best possible policy in $\Pi$, as
The quantity $\Delta_{\Lambda}(\pi) = \inf_{Q \in \mathcal{P}(\Lambda)} \mathbb{E}_Q \left[Y(\pi(X)) - Y(0)\right]$ represents the worst-case welfare gain of policy $\pi$ relative to the baseline policy $\pi_0(x) = 0$. As a result, this approach can be readily generalized to the policy improvement of $\pi$ against any given baseline policy $\pi_0$, with the optimal policy obtained by:
Within our framework, learning a confounding-robust policy requires identifying and estimating either the worst-case welfare function $W_{\Lambda}(\pi)$ or the worst-case policy improvement function $\Delta_{\Lambda}(\pi)$. In this section, we formally characterize these criterion functions by leveraging partial identification results for conditional means and CATE under the MSM. We then develop doubly robust/orthogonal moment functions for identifying $W_{\Lambda}(\pi)$ and $\Delta_{\Lambda}(\pi)$, which enable the construction of efficient estimators.
For notational simplicity, we define two quantile functions $q^{\pm}_{\Lambda}(x,a)$ as \[
\] where $F(\cdot|x,a)$ denotes the CDF of $Y$ given $X=x$ and $A=a$. Moreover, let \( e_a(x) = \mathbb{P}(A = a | X = x) \) for $a \in \{0,1\}$, so that \( e_1(x) = e(x) \) and \( e_0(x) = 1 - e(x) \).
In this subsection, we identify and characterize the robust criterion functions $W_\Lambda(\pi)$ and $\Delta_\Lambda(\pi)$. Additionally, we derive the first-best policies optimized under these robust criteria. Let $\Pi_o$ denote the set of all measurable policies: \[ \Pi_o = \left\{ \pi: \mathcal{X} \rightarrow [0,1] : \pi \text{ is Borel measurable} \right\}. \] A first-best policy refers to a policy that maximizes the criterion function over the unrestricted policy class $\Pi_o$.
We begin by identifying the worst-case welfare $W_\Lambda(\pi)$. When $\Lambda >1$, the true conditional mean function $\mu_{o}(x,a) = \mathds{E}_{P_o}[ Y(a) | X =x ]$ is no longer point identified due to unobserved confounding. Instead, we characterize the sharp bounds for $\mu_{o}(x,a)$ using (ref) in (ref), which also implies a sharp lower bound for the true welfare function $\mathbb{E}_{P_o}[Y(\pi)]$. Specifically, the proposition provides closed-form expressions for the upper and lower bounds, denoted $\mu^{\pm}_{\Lambda}(x,a)$, which are given by \[ \mu^{\pm}_{\Lambda}(x,a) = \mathbb{E}\left[ Y \mathds{1}\{ A = a \} \left[1+\frac{1- e_a(X)}{e_a(X)}\Lambda^{\pm\text{sgn}\left(Y-q^{\pm}_{\Lambda}(X,a)\right)}\right] \Big|X=x \right], \\ \] where $\text{sgn}(t)=1$ if $t\geq0$ and $-1$ otherwise.
The identification of $\Delta_\Lambda(\pi)$ is established in (ref), which follows as a corollary of (ref). To formalize the result, define $\tau_\Lambda^-(x) = \mu^-_\Lambda(x,1) - \mu^+_\Lambda(x,0)$.
(ref) characterizes the first-best MMW policy $\pi_{W,\Lambda}^\star(x)$, which assigns treatment by comparing the lower bounds $\mu_{\Lambda}^-(x,1)$ and $\mu_{\Lambda}^-(x,0)$. In contrast, (ref) shows that the first-best MMI policy $\pi_{\Delta,\Lambda}^\star(x)$ assigns treatment by comparing $\mu_{\Lambda}^-(x,1)$ with the upper bound $\mu_{\Lambda}^+(x,0)$. As a result, the MMI policy is more conservative: it treats only when the worst-case treated outcome exceeds the best-case control outcome. In the context of workplace inspections, the MMI policy $\pi_{\Delta,\Lambda}^\star(x)$ is particularly suitable when regulatory resources are severely limited. Conversely, when resources are abundant, the more expansive MMW policy may be preferable.
We conclude this subsection by comparing our first-best MMW and MMI policies with the Bayes decision rule (referred to as a policy in our framework) of pu2021estimating, hereafter referred to as PZ-policy.
To effectively learn the confounding-robust policy, it is essential to estimate the entire worst-case welfare function with minimal estimation error. To facilitate the use of modern ML methods, we derive doubly robust scores for the worst-case criteria, $W_\Lambda (\pi)$ and $\Delta_\Lambda (\pi)$, such that the estimation of nuisance parameters has no first-order influence on the resulting policy evaluation. For notational convenience, we may write $\pi(a|x) = a \pi(x) + (1-a) \left(1-\pi(x)\right)$, which compactly represents the probability that treatment $a \in \{0,1\}$ is assigned under policy $\pi$ given covariate $x$.
We begin by presenting the doubly robust score for estimating $W_\Lambda (\pi)$. As shown in (ref), this naturally leads to the following moment condition: \[ \mathbb{E}\left[ \mu_\Lambda^-(X, 1) \pi (1|X) + \mu_\Lambda^-(X, 0) \pi (0|X) \right] - W_\Lambda(\pi) = 0. \] For $t \in \{0,1\}$, let
Using explicit expressions for $\mu_\Lambda^{\pm}$ given in (ref), the moment condition above can be rewritten as
Here, $W_\Lambda(\pi)$ is the parameter of interest, while $e(\cdot)$ and $q_{\Lambda}^-(\cdot, \cdot)$ are two unknown nuisance parameters that can be estimated using ML/nonparametric methods.
A plug-in estimator for $W_\Lambda(\pi)$ can be constructed by substituting estimates into the score function $g_t(Z; \cdot)$ and averaging over the sample, yielding $n^{-1} \sum_{i=1}^n\sum_{t\in\{0,1\}} g_t\left(Z_i; \widehat{e}, \widehat{q}^{-}_{\Lambda}\right)\pi(t|X_{i})$. However, this plug-in estimator is sensitive to estimators of the nominal propensity score and the quantile functions, and may suffer from severe bias. This motivates the doubly robust estimator proposed in this section.
Following chernozhukov2018double,chernozhukov2022locally, we construct the doubly robust score for $W_\Lambda(\pi)$ as follows:
where
Let $\eta_{W, \Lambda} = \left(e,q_\Lambda^- ,\rho^-_{1,\Lambda}, \rho^-_{0,\Lambda} \right) $ denote the tuple of nauisance functions, the doubly robust score for $W_\Lambda(\pi)$ is given by \[ \psi_{W}\left(z,\pi;\eta_{W, \Lambda} \right) = \sum_{t \in \{0,1\} } \phi_{t}^{-} \left(z; \eta_{W, \Lambda} \right) \pi(t|x). \]
The following (ref) establishes the Neyman orthogonality for the moment condition. Moreover, the function $\psi_W(z, \pi; \eta_{W, \Lambda}) - W_\Lambda(\pi)$ serves as the efficient influence function for estimating $W_\Lambda(\pi)$.
We then construct the doubly robust score of $\Delta_{\Lambda}(\pi)$. (ref) implies the moment condition as follows: \[ \mathbb{E}\left[ \tau_{\Lambda}^{-} (X) \pi(X) \right] - \Delta_{\Lambda}(\pi) = 0. \] Let $\eta_{\Delta, \Lambda} = \left(e,q_\Lambda^{\pm}, \rho^{\pm}_{1,\Lambda}, \rho^{\pm}_{0,\Lambda} \right)$ denote the tuple of nuisance parameters. The doubly robust score for $\Delta_\Lambda(\pi)$ can be constructed as
where
We summarize several crucial properties of $\psi_{\Delta}$ in the following proposition.
In this section, we focus on the algorithms for estimating the MMW and MMI policies, as defined in (ref). Let $\widehat{\pi}_{W,\Lambda}$ and $\widehat{\pi}_{\Delta,\Lambda}$ denote the estimated MMW and MMI policies, respectively. We also provide theoretical guarantees for these estimated policies in the form of upper bounds on their associated regrets. Specifically, we derive asymptotic upper bounds on the CRW-regret $\mathrm{Reg}_W\left(\widehat{\pi}_{W, \Lambda} \right)$ and the CRI-regret $\mathrm{Reg}_{\Delta}\left(\widehat{\pi}_{\Delta, \Lambda} \right)$, as defined in (ref) and (ref).
We estimate the MMW and MMI policies using a two-stage procedure: (1) estimation of the robust criterion, either $W_\Lambda(\pi)$ or $\Delta_\Lambda(\pi)$, using $K$-fold cross-fitting; and (2) policy optimization based on the estimated objective.
Given the doubly robust scores in (ref), we first estimate $e(\cdot)$, $q^{\pm}_{\Lambda}(\cdot, \cdot)$ and $\rho_{t,\Lambda}^{\pm}(\cdot, \cdot)$ for $t\in\{0,1\}$. To this end, we divide the sample into $K$ evenly-sized folds $\cup_{k=1}^{K}\mathcal{I}_{k}$. For each fold $k \in [K]$, we estimate $e, q^{\pm}_{\Lambda}$ and $\rho_{t,\Lambda}^{\pm}$ using the rest $K-1$ folds, and denote the resulting estimates by \[ \widehat{\eta}_{W,\Lambda}^{-k} =\left( \widehat{e}^{-k}, q^{-, -k}_{\Lambda}, \rho_{1,\Lambda}^{-, -k }, \rho_{0,\Lambda}^{-, -k } \right) \quad\text{and}\quad \widehat{\eta}_{\Delta,\Lambda}^{-k} =\left( \widehat{e}^{-k}, q^{\pm, -k}_{\Lambda}, \rho_{1,\Lambda}^{\pm, -k }, \rho_{0,\Lambda}^{\pm, -k } \right). \] Using these estimates, we construct the cross-fitted estimator for $W_\Lambda(\pi)$ and $\Delta_{\Lambda}(\pi)$ as
The final policy is selected from a pre-specified policy class $\Pi_n$ by maximizing the estimated objective, either $\widehat{W}_{\Lambda,n}(\pi)$ or $\widehat{\Delta}_{\Lambda,n}(\pi)$. The complete procedures for estimating the MMW and MMI policies are summarized in (ref) and (ref), respectively.
In policy design, it is crucial to account for multiple constraints, including budget, simplicity, interpretability, and functional form. Statistically speaking, to achieve regret bounds that decay at the rate of $n^{-1/2}$, one must control the complexity of the policy class. In the remainder of the paper, we allow the policy class $\Pi\equiv \Pi_n$ to vary with $n$. Consequently, we restrict $\Pi_n$ to be VC-subgraph class; see vaart2023empirical, gine2021mathematical for further details.
Various machine learning models can be used as policy classes. Below, we present several examples of such models along with their corresponding VC dimensions, covering both deterministic and randomized policy classes.
Next, we provide several examples of randomized policy classes.
In this section, we present a key lemma demonstrating that the estimation error of the nuisance parameters becomes asymptotically negligible when the objective functions $W_{\Lambda}(\cdot)$ and $\Delta_{\Lambda}(\cdot)$ are estimated using the doubly robust score with cross-fitting.
(ref) is straightforward to verify.(ref) is the standard strict overlap condition in the literature, which is essential for establishing the asymptotical regret upper bounds. Moreover, we adopt an agnostic stance on how the nuisance components, $\widehat{e}$, $ \widehat{q}_{\Lambda}$ and $\widehat{\rho}_{t, \Lambda}$, are estimated. Rather than specifying particular estimation procedures, we impose high-level assumptions on their convergence ratio, as stated below.
The use of doubly robust scores combined with cross-fitting ensures that errors from nuisance parameter estimation become asymptotically negligible for policy learning, provided $\mathrm{VC}(\Pi_n)$ does not grow too quickly with $n$. To formalize this, suppose the true nuisance functions $\eta_{W, \Lambda}$ and $\eta_{\Delta, \Lambda}$ defined in (ref) are known, the oracle estimators for $W_\Lambda(\pi)$ and $\Delta_\Lambda(\pi)$ are give by \[
\] We conclude this subsection by demonstrating that $\widehat{W}_{\Lambda,n}(\pi)$ is a valid approximation to $W_{\Lambda,n}(\pi)$, and similarly, $\widehat{\Delta}_{\Lambda,n}(\pi)$ approximates $\Delta_{\Lambda,n}(\pi)$, both with a convergence rate faster than $n^{-1/2}$.
In this subsection, we will demonstrate that, under appropriate bounds on potential hidden confounding, the CRW-regret and CRI-regret of the learned MMW and MMI policies decay at a rate that is upper bounded by $\sqrt{\mathrm{VC}(\Pi_n)/n}$. Recall that the CRW-regret and CRI-regret of the learned optimal policies are defined as \[
\] where $\widehat{\pi}_{W,n}$ and $ \widehat{\pi}_{\Delta,n}$ are learned from (ref) and (ref).
(ref) improves upon the result in kallus2021minimax, where Proposition 6 shows that estimating the nominal propensity score leads to a regret bound that exceeds the oracle regret bound (i.e., the regret bound assuming a known nominal propensity score) by an additional term of the order \[ \frac{1}{n} \sum_{i=1}^n \left| 1/\widehat{e}\left(X_i\right) - 1/e(X_i) \right|. \] With the application of the doubly robust score and cross-fitting, the regrets of the learned MMW and MMI policies decay at the rate $\sqrt{\mathrm{VC}(\Pi_n)/n}$, and the estimation errors of the nuisance parameters have no asymptotic effect on the regret bound.
In this section, we conduct simulation studies to evaluate the performance of the MMW and MMI policies, which are learned using Algorithm (ref) and (ref), respectively. To facilitate comparison with the results in kallus2021minimax, we use the logistic policies given in (ref).
The data-generating process (DGP) is specified as follows. Let $\log \Lambda^* = 1.5$, which implies that the true sensitivity parameter is $\Lambda^* = 4.482$. The observed covariates $X \in \mathbb{R}^2$, treatment assignment $A \in \{0,1\}$, and outcome $Y \in \mathbb{R}$ are generated according to: \[
\] where $\mu_X = [-1, 1]'$ and $I_2\in \mathbb{R}^{2\times 2}$ is the identity matrix. The nominal propensity score $e(X)$ is defined as: \[ e(X) = \mathbb{P}(A=1 \mid X) = \sigma\left( \zeta(X)^\prime\theta \right), \quad \text{with} \quad \sigma(z) = \frac{1}{1 + e^{-z}}, \] where the parameter vector $\theta$ and the nonlinear feature map $\zeta(X)$ are given by \[ \theta = \left[0.2, 0.4, 0.1, -0.1, 0.5, -0.5\right], \quad \zeta(X) = \left[\max(x_1, 0), \, \frac{x_1 x_2^2}{10}, \, \sin(x_2^2), \, x_1, \, x_2, \, 1\right]'. \]
It can be readily verified that $U$ perturbs the nominal odds ratio $e(X)/(1 - e(X))$ by a multiplicative factor of ${\Lambda^*}^{\pm 1}$, ensuring the DGP satisfies the MSM with $\Lambda^* = 4.482$. Finally, the potential outcome is generated as: \[ Y(A) = \beta_{\text{cons}} + \beta_A A + X' \beta_X + (A \cdot X)'\beta_{X,A} + \beta_U U + \epsilon, \quad \epsilon \sim \mathcal{N}(0,1), \] where: \[ \beta_{\text{cons}} = -0.2, \quad \beta_A = -0.1, \quad \beta_X = [1, -1]', \quad \beta_{X,A} = [0.2, 0.4]', \quad \beta_U = 1.5. \]
We generate $\mathrm{i.i.d.}$ samples $\{ X_i, A_i, U_i, Y_i(1), Y_i(0)\}_{i=1}^{n}$ according to the DGP described above, with $n=2{,}000$. The observed data $\{ X_i, A_i, Y_i \}_{i=1}^{n}$ are used to learn the policies. The conditional quantile functions $q_{\Lambda^*}^{\pm}$ are estimated using gradient boosted trees and the nominal propensity score $e$ and CVaR $\rho_{t, \Lambda^*}^{\pm}$ are estimated using random forests.
We evaluate the performance of estimated policy $\widehat{\pi}$ using three metrics: (i) the expected welfare, defined in (ref); (ii) the worst-case welfare $W_\Lambda(\widehat{\pi})$, as given in Theorem (ref); and (iii) the worst-case policy improvement $\Delta_\Lambda(\widehat{\pi})$ given in Theorem (ref). All three metrics are estimated using sample averages computed on an independent out-of-sample dataset of 100,000 randomly drawn observations. To assess the stability of our method, we evaluate each metric over 100 repeated experiments. We report the mean from these repetitions and visualize the variability using shaded bands that represent the 95% confidence bands, calculated from the standard deviation across the 100 experiments.
We report the performance of the MMW and MMI policies, as well as the robust IPW-based policy of kallus2021minimax (KZ for short). As a benchmark, we include the policy proposed by athey2021policy (AW), which assumes no unobserved confounding $(\Lambda =1)$. To evaluate the MMW, MMI, and KZ policies, we conduct a sensitivity analysis by varying the sensitivity parameter over $\log \Lambda \in \{ 0.1,0.2,\ldots, 3.5 \}$. This assesses their robustness to miscalibration relative to the true level of unobserved confounding in the true DGP with $\log \Lambda^* = 1.5$. By design, the MMW and MMI policies reduce to the AW policy at the $\Lambda=1$ setting. In contrast, the KZ policy remains different from AW even when $\Lambda=1$, as its objective function does not include a bias-correction term for the estimated propensity scores.
(ref) shows how each policy adjusts its treatment probability in response to the assumed level of unobserved confounding, and helps explain their relative performance in terms of expected welfare, presented in (ref). The AW policy, assuming no unobserved confounding($\Lambda = 1$), yields an aggressive strategy, assigning treatment to 95% individuals. In contrast, MMI, MMW and KZ are sensitivity-aware, adopting a conservative approach that is responsive to $\Lambda$. As $\log \Lambda$ increases, they systematically reduce treatment rates to hedge against selection bias. However, their conservative dynamics differ significantly. The KZ policy is overly conservative, exhibiting a precipitous drop in assignments due to its reliance on the loose lower bound on policy improvement from zhao2019sensitivity. In contrast, the MMI policy is precisely conservative; by precisely identifying heterogeneous treatment effects, it establishes a sharper bound on policy improvement, leading to a more calibrated and gradual reduction in treatment. Finally, the MMW policy's aggressiveness under high uncertainty is a direct consequence of its decision rule, illustrated in (ref), which favors treatment whenever $\mu_{\Lambda}^{-}(x,1)$ exceeds $\mu^{-}_\Lambda(x,0)$.
(ref) shows the resulting expected welfare. The MMI policy achieves its peak expected welfare at $\log \Lambda = 1$. It outperforms the MMW policy for assumed confounding levels below the true value of $\log \Lambda^* = 1.5$. In regimes of greater assumed uncertainty $\log \Lambda > 1.5$, however, the MMW policy's performance is superior, yielding significantly higher welfare. Their performance crossover occurs precisely at this true value, confirming their complementary nature. This divergence occurs because, in high-uncertainty regimes, the MMI policy becomes overly cautious, whereas the MMW policy adopts a more risk-tolerant strategy.
Finally, (ref) and (ref) validate that the observed behaviors are consequences of each policy's specific design. As expected, (ref) shows the MMW policy achieving the highest worst-case welfare, and (ref) confirms the MMI policy's success in maximizing the worst-case policy improvement. This confirms that each policy effectively optimizes its intended objective, leading to the distinct and complementary performance profiles we identified.
In this section, we apply our method to two empirical applications. (ref) revisits the JTPA study and examines assignment to job training programs, while (ref) explores enrollment decisions in the Head Start program. In both settings, unobserved confounders likely give rise to endogenous selection.
We apply the MMW and MMI policy learning methods to the National Job Training Partnership Act (JTPA) Study, a large-scale randomized controlled trial evaluating training services for disadvantaged adults bloom1997benefits. In the experiment, eligibility for services was randomly assigned, and participants’ earnings were tracked over the 30-month period after the assignment. This dataset was notably used by kitagawa2018a to estimate an intent-to-treat optimal policy for 9,223 applicants using their Empirical Welfare Maximization (EWM) approach, based on covariates such as education and prior earnings.
While eligibility was randomly assigned, compliance in the JTPA study was imperfect: approximately 23% of individuals did not adhere to their assigned eligibility status, as shown in (ref). This imperfect compliance introduces self-selection into actual participation decisions, raising concerns about unobserved confounding. For instance, eligible individuals who opt into the program may exhibit stronger job search motivation or face fewer personal constraints, such as childcare responsibilities, health issues, or transportation barriers, than those who decline participation. These unobserved factors likely influence both program participation and post-program earnings, violating the unconfoundedness assumption required by conventional policy learning methods.
We focus on actual program participation, as the policy-relevant objective is to target individuals who are most likely to benefit from job training and to allocate participation accordingly. Building on the perspective of d2021orthogonal, who frame policy learning as deriving optimal treatment rules under potential unobserved confounding, we apply the MMW and MMI methods, both of which are explicitly designed to address such confounding. Unlike d2021orthogonal, who use a binary instrumental variable (eligibility) to construct partial identification intervals for the CATE, we adopt a complementary strategy by employing with a pre-specified sensitivity parameter $\Lambda$.
We estimate the optimal MMW and MMI policies using Algorithms (ref) and (ref), with $K = 10$ folds for cross-fitting. The conditional quantile functions $q_{\Lambda^*}^{\pm}$ are estimated using gradient boosted trees, and the nominal propensity score $e$ and CVaR $\rho_{t, \Lambda^*}^{\pm}$ are estimated using random forests. To account for program costs, we subtract \$1,216 from the outcomes of treated individuals. This amount corresponds to the average cost of services per actual participant, as reported in Table 5 of bloom1997benefits. Following the approach of kitagawa2018a and d2021orthogonal, we consider the class of quadrant treatment policies due to its simplicity and interpretability: \[ \Pi \equiv \left\{ \mathds{1} \left\{ s_1(\text{edu} - t_1) > 0, \; s_2(\text{earnings} - t_2) > 0 \right\} : s_1, s_2 \in \{-1, 1\}, \; t_1, t_2 \in \mathbb{R} \right\}. \]
(ref) presents the estimated upper and lower bounds of the CATE under the MSM framework, evaluated across different values of the sensitivity parameter $\Lambda \in \{1, 1.5, 2\}$. When $\Lambda = 1$, the estimated upper and lower bounds coincide. Choosing a larger value of $\Lambda$ can yield more robust policy targeting by ensuring that the identified set contains the true causal effect. However, this robustness comes at the cost of producing less informative estimates and more conservative policy intervention. Therefore, we recommend that practitioners examine results across a range of $\Lambda$ values to assess the sensitivity of their conclusions to unobserved confounding.
(ref) illustrates the estimated optimal quadrant policies for MMW and MMI. When $\Lambda = 1$, both the MMW and MMI policies is exactly the AW policy, but with a treated fraction approaching one. As sensitivity parameter $\Lambda$ increases, the MMI policy becomes increasingly conservative, with the treatment rate dropping rapidly to near zero when $\Lambda \geq 1.5$. In contrast, the MMW policy shows a more gradual decline in the treatment ratio, which remains relatively high at 0.75 when $\Lambda = 2$, and drops sharply to around 0.05 for $\Lambda \geq 2.5$. The extremely low treatment rates under the MMI policy arise because the estimated lower bounds of the CATE fall below zero for nearly all individuals when $\Lambda \geq 1.5$. Consequently, our subsequent analysis focuses on the MMW policy. When $\Lambda = 1$, the quadrant policy targets individuals with fewer than 17.5 years of education and pre-program earnings below \$40{,}800. In contrast, at $\Lambda = 2$, the MMW quadrant policy assigns treatment to those with at least 10.5 years of education and earnings below \$35{,}144. As the sensitivity parameter $\Lambda$ increases, the estimated MMW policy becomes more selective, prioritizing individuals with higher educational level and lower earnings.
We further demonstrate our confounding-robust policy learning method by applying it to enrollment decisions for the Head Start program. The objective is to develop a policy that improves academic outcomes by targeting admission to children most likely to benefit. This application presents a multi-valued treatment setting with three alternatives, Head Start, other preschools, or no preschool, necessitating the use of our MMW approach extended for multiple treatments as detailed in (ref).
Head Start is a federal matching grant initiative designed to foster the cognitive, social, and health development of children from low-income families, with the aim of preparing them to enter school on a more equal footing with their more advantaged peers. A substantial body of research has evaluated the impact of Head Start participation on outcomes such as academic achievement, physical health, and long-term socioeconomic status; see, e.g., currie1993does, deming2009early, ludwig2007does, walters2015inputs.
A key challenge in evaluating Head Start's impact is that enrollment is not randomly assigned. First, families self-select into Head Start or other preschool programs based on eligibility and perceived benefits. Those who value early education may also make unobserved investments in their children, such as fostering enriched home environments or providing supplemental learning. Second, limited program capacity in many areas gives administrators discretion over admissions, leading to selection based on both observed and unobserved characteristics, some of which also affect child outcomes. As a result, failing to account for such confounding can lead to targeting policies that misidentify the children most likely to benefit.
We draw on data from the National Longitudinal Survey of Youth (NLSY) and its child supplement, the National Longitudinal Survey of Youth 1979 Children and Young Adults (NLSCYA). The NLSY began in 1979 with a nationally representative sample of young women, and the NLSCYA has tracked their children since 1986. Our analysis sample is constructed by linking child-level outcomes from the NLSCYA with maternal background characteristics from the original NLSY cohort.
We use the child's standardized score on the Peabody Picture Vocabulary Test (PPVT), a widely used measure of early verbal ability, as a proxy for our policy objective of improving academic achievement. The pre-treatment covariates used in our analysis capture a range of child, maternal, and household characteristics. These include: (1) the percentile rank of average net family income between ages 0 and 4 (Income Pctl); (2) child’s gender; (3) birth weight (Birth Wt); (4) weight at preschool entry (Preschool Entry Wt); (5) firstborn status; (6) mother’s highest completed grade; (7) mother’s AFQT percentile score (Mother AFQT); (8) number of the mother's biological siblings; (9–10) number of household members with less than 12 years and at least 16 years of education; and (11) race (White, Hispanic, or Black). After excluding observations with missing data, the final sample consists of 3,826 children: 755 attended Head Start, 2,020 enrolled in other preschool programs, and 1,051 did not attend any preschool.
While the full set of pre-treatment covariates is used to estimate doubly robust scores, we restrict the variables used for policy optimization to a small, interpretable subset: percentile rank of family income, mother’s AFQT score, birth weight, and weight at preschool entry. These are selected based on fairness and ethical considerations. To ensure interpretability, we restrict the policy class to depth-2 decision trees. Figure (ref) shows the estimated policies under four values of the sensitivity parameter $\Lambda = 1, 2, 3, 4$.
The policy tree learned under $\Lambda = 1$ serves as our benchmark and exactly replicates the policy obtained by approaches such as athey2021policy,zhou2023offline, which do not account for unobserved confounding. An analysis of this benchmark policy reveals several features that present challenges for direct implementation. First, the policy assigns Head Start to a very limited subgroup, only 1.23% of the population. Second, it prioritizes children from higher-income families (albeit with low preschool-entry weight), which deviates from the program’s original mission of serving economically disadvantaged household. At the same time, the policy assigns the lowest-income children to either other preschools or no preschool at all. This motivates our subsequent sensitivity analysis to assess how the policy recommendations change when accounting for unobserved confounding.
As the sensitivity parameter $\Lambda$ increases, the optimal policy undergoes a clear evolution. Initially, at $\Lambda=2$, the policy is still primarily driven by income, but becomes more stringent by reducing the Head Start allocation to just 0.78% of the sample. A more dramatic structural break occurs at $\Lambda = 3$, where the model abandons income-based logic in favor of health indicators, such as birth weight. The policy under $\Lambda = 4$ advocates for a targeted intervention for a clearly defined disadvantaged group while adopting a conservative, non-interventionist stance for the majority.
Formulating effective policies from observational data is challenged by unobserved confounding, which can lead to biased and potentially harmful policy intervention. This paper proposes a robust and efficient framework for policy learning that explicitly accounts for such endogeneity. Our contributions are twofold: first, we derive sharp, closed-form identification results for robust welfare criteria under the Marginal Sensitivity Model; second, we develop doubly robust scores for their estimation. This innovation enables the use of flexible machine learning methods for nuisance estimation within our framework. Building on these components, we establish regret bounds for the resulting confounding-robust policies.
Our work also offers insights into the practical challenges of applying policy learning methods. While powerful, ML-based approaches that optimize a single objective can yield policies that are difficult to trust and interpret, particularly when built on fragile assumptions. Our work underscores that sensitivity analysis is a critical step for examining a policy's foundational reliability.
Finally, several directions remain open for future research. First, while our framework primarily focuses on binary treatments, extending it to multivalued or continuous treatments is a promising direction that would broaden its applicability and introduce new methodological challenges. Second, the choice of the sensitivity parameter $\Lambda$ remains an important yet unresolved issue. Although we suggest that practitioners may specify $\Lambda$ based on domain knowledge, this approach is inherently ad hoc. While hsu2013calibrating offers some guidance, the development of systematic, data-driven methods for selecting $\Lambda$ remains a key avenue for future investigation.