EconBase
← Back to paper

Semiparametric Efficiency in Policy Learning with General Treatments

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,725 characters · 13 sections · 54 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Semiparametric Efficiency in Policy Learning with General Treatments

abstractRecent literature on policy learning has primarily focused on regret bounds of the learned policy. We provide a new perspective by developing a unified semiparametric efficiency framework for policy learning, allowing for general treatments that are discrete, continuous, or mixed. We provide a characterization of the failure of pathwise differentiability for parameters arising from deterministic policies. We then establish efficiency bounds for pathwise differentiable parameters in randomized policies, both when the propensity score is known and when it must be estimated. Building on the convolution theorem, we introduce a notion of efficiency for the asymptotic distribution of welfare regret, showing that inefficient policy estimators not only inflate the variance of the asymptotic regret but also shift its mean upward. We derive the asymptotic theory of several common policy estimators, with a key contribution being a policy-learning analogue of the Hirano–Imbens–Ridder (HIR) phenomenon: the inverse propensity weighting estimator with an estimated propensity is efficient, whereas the same estimator using the true propensity is not. We illustrate the theoretical results with an empirically calibrated simulation study based on data from a job training program and an empirical application to a commitment savings program. Keywords: General Treatment, Policy Learning, Semiparametric Efficiency, Welfare Regret.

Introduction

Policy learning has become a central topic in econometrics, statistics, and machine learning, offering a formal framework for designing data-driven decision rules in settings such as targeted social programs, pricing and revenue management, and personalized medicine. The primary objective is to construct a policy, understood as a mapping from individual characteristics to treatments, that maximizes expected welfare, typically defined as the population mean of the counterfactual outcomes induced by the policy. A widely used approach is empirical welfare maximization (EWM), which selects a policy from a specified class by maximizing an estimator of social welfare built from experimental or observational data under unconfoundedness; see, for example, kitagawa2018should and athey2021policy. The performance of the resulting policy is usually assessed by establishing bounds on its regret—the difference between optimal welfare and the welfare achieved by the learned policy.

In this paper, we offer a new theoretical perspective on policy learning and evaluation that complements the predominant regret-bounding paradigm. Our approach develops a semiparametric efficiency framework for policy learning with general treatment variables, including discrete, continuous, and mixed cases, based on modern semiparametric estimation theory. We first show, building on recent work such as crippa2025regret, that deterministic policies, which assign treatments as deterministic functions of covariates, are not pathwise differentiable objects and therefore cannot be estimated at the root-$n$ rate. This motivates our focus on randomized policies, which assign treatment probabilities rather than treatment levels.

Beyond their statistical regularity, randomized policies also offer several practical advantages. They soften hard thresholds and often lead to greater acceptance among individuals who are subject to the policy. From the perspective of fairness, a randomized allocation rule ensures that individuals with similar characteristics are treated symmetrically rather than being separated by small or noisy differences in the data. Randomized outcomes also discourage gaming and strategic behavior, since the assignment is never fully predictable. Finally, randomization provides a simple and transparent mechanism for assigning scarce resources in settings where deterministic rules may raise legal and ethical concerns about the basis of the allocation.

Due to these reasons, randomized policies have been a standard tool for resource allocation in practice. For example, the U.S. H-1B visa system incorporates a lottery in which groups of applicants face different selection probabilities depending on whether they hold an advanced degree glennon2024restrictions. Vehicle license allocation schemes in cities such as Beijing and Shanghai operate through weighted lotteries, where applicants receive higher or lower chances based on pre-specified criteria li2018better,barwick2024efficiency. Similar structured lotteries appear in public health policy, such as the Oregon Health Plan lottery, which granted Medicaid access through random selection among eligible low-income adults finkelstein2012oregon. School admissions provide another example: many charter schools rely on admissions lotteries in which priority status or sibling preferences influence the probability of receiving an offer abdulkadirouglu2011accountability. Housing policy in urban India also uses randomized allocation rules, where subsidized units are assigned through lotteries that incorporate eligibility categories and priority groups kumar2021housing.

To study the statistical properties of randomized policies, we build on the literature on semiparametric efficiency for treatment effect estimation hahn1998role,chen2008semiparametric,ai2021unified and develop a semiparametric efficiency framework for policy parameters in the setting of randomized policies. Using the H\'ajek–Le Cam convolution theorem, we then introduce a notion of efficiency for the asymptotic distribution of welfare regret. Regret is nonnegative, and its efficient limiting distribution takes a $\chi^2$ form. The analysis reveals an unexpected implication: an inefficient policy estimator not only increases the variance of the asymptotic regret but also shifts its mean upward. This provides a substantive motivation for pursuing efficient policy learning methods, as efficiency has a direct impact on the asymptotic mean of regret.

We then derive the asymptotic behavior of several common policy estimators, which differ in how they estimate social welfare. These include inverse propensity weighting (IPW) with propensity score either known or estimated (through stabilized weights), as well as the doubly robust (DR) estimator. A central finding is a version of the hirano2003efficient (HIR) phenomenon in the policy learning setting: the IPW policy estimator that uses an estimated propensity score attains efficient regret, whereas the counterpart that uses the true propensity score does not. This result is unexpected in light of the existing literature. In kitagawa2018should and mbakop2021model, the regret bounds for EWM with estimated propensities in the binary-treatment deterministic-policy setting are dominated by the rate of convergence of the nuisance estimators. In contrast, we show that for randomized policies the regret can be sharpened to the root-$n$ rate, and in fact becomes more efficient when the propensity score is estimated.

Lastly, we show that the HIR phenomenon may extend beyond the class of randomized policies. In the existing literature, regret bounds for deterministic policies are typically obtained by applying a basic inequality that relates regret to the supremum loss of the welfare estimator kitagawa2018should. Using Gaussian process theory, we demonstrate that this supremum welfare loss also exhibits an HIR phenomenon. The key step is to compare the covariance structures of the limiting Gaussian processes that characterize the welfare process under different estimators. This suggests that the HIR phenomenon is not tied to randomization per se, but reflects deeper efficiency issues for policy learning.

The theoretical results have important implications for policy learning practice. In experimental settings where the propensity score is known, researchers may be inclined to use an IPW policy estimator that relies on the true propensity score when learning an optimal policy. Our findings indicate that this choice is not desirable. To achieve regret efficiency, one should instead use the IPW estimator with an estimated propensity score or the doubly robust estimator. Using an empirically calibrated simulation study based on the National Job Training Partnership Act (JTPA) dataset, we demonstrate that IPW with the true propensity score leads to a larger mean and standard deviation of regret than the efficient policy estimators. In our empirical application, we revisit the commitment savings product experiment of ashraf2006tying and study the optimal assignment of the product across bank clients with rich demographic and financial information. The results indicate that policies learned using inverse propensity weighting with the true propensity score can exhibit substantially larger standard errors than those obtained from efficient estimators.

\paragraph{Related literature.} Our paper contributes to two strands of literature. The first is the policy learning literature for different types of treatment variables. Prior work includes kitagawa2018should,athey2021policy for binary treatments, zhou2023offline,fang2025model for multi valued discrete treatments, and kallus2018policy,ai2024data for continuous treatments. These studies focus on searching for optimal deterministic policies within a specified policy class and establish bounds on welfare regret.\footnote{Another approach to optimal decision making is to characterize the unrestricted optimal (“first-best”) policy. See, for example, manski2004statistical,stoye2009minimax,bhattacharya2012inferring,feng2024statistical,chen2025inference.} crippa2025regret examines the asymptotic distribution of regret for deterministic threshold policies, showing that the resulting rate follows the cube root asymptotics developed by kim1990cube. Using kernel smoothing, crippa2025regret further shows that the rate can be improved to something faster than cube root but still strictly slower than the root $n$ rate. chernozhukov2019semi study efficient policy learning with continuous treatments in a semiparametric framework under a specific functional form for the welfare. Our work provides a complementary perspective by introducing a general treatment framework for randomized policies that can be parametrized through pathwise differentiable functionals, thereby allowing for root n regret efficiency.

Our paper also relates to the literature on semiparametric efficiency for average treatment effects hahn1998role,hirano2003efficient and for more general causal effect functionals chen2008semiparametric,ai2021unified. Although our analysis draws on the same foundational tools, the objects of interest in policy learning fall outside the scope of these existing frameworks. Standard semiparametric theory does not cover policy parameters, nor does it characterize their regret distributions. Our results show that policy learning can indeed be brought into the semiparametric efficiency framework, and doing so leads to several unexpected insights that are absent from the existing policy learning literature.

\paragraph{Organization of the paper.} Section (ref) introduces the general-treatment policy learning model and the corresponding assumptions. Section (ref) discusses the semiparametric efficiency theory for learning randomized policies. Section (ref) examines the asymptotic theory for IPW and DR estimators, demonstrating the HIR phenomenon in the policy learning context. Section (ref) discusses a regret-upper-bound approach for comparing policies beyond randomized settings. Section (ref) and (ref) present the simulation and empirical studies. Section (ref) concludes.

Policy Learning Model

We begin by introducing the policy learning framework. The treatment variable $T$ take values in a Borel set $\mathcal{T} \subset \mathbb{R}$, allowing for discrete, continuous, or mixed support. Let $Y(t)$ denote the potential outcome when treatment is assigned at level $t \in \mathcal{T}$. The observed outcome is $Y = Y(T)$. Let $X$ denote a vector of covariates, taking values in $\mathcal{X} \subset \mathbb{R}^d$. We collect the observed variables as $Z=(Y,T,X)$.

Policies are parameterized by $\theta\in\Theta\subset\mathbb R^p$. We endow the spaces $\mathcal T$, $\mathcal X$, and $\Theta$ with their respective Borel $\sigma$-algebras, denoted $\mathcal B_{\mathcal T}$, $\mathcal B_{\mathcal X}$, and $\mathcal B_\Theta$. A policy is defined as a probability kernel from $(\mathcal X\times\Theta,\ \mathcal B_{\mathcal X}\otimes \mathcal B_\Theta)$ to $(\mathcal T,\ \mathcal B_{\mathcal T})$; that is, a measurable map

align*[align* omitted — 132 chars of source]

such that (i) for each $(x,\theta)$, the mapping $B\mapsto \Pi_\theta(B| x)$ defines a probability measure on $\mathcal B_{\mathcal T}$; and (ii) for each $B\in\mathcal B_{\mathcal T}$, the function $(x,\theta)\mapsto \Pi_\theta(B| x)$ is measurable with respect to $\mathcal B_{\mathcal X}\otimes\mathcal B_\Theta$. This setup accommodates both randomized and deterministic policies, with the latter arising as the special case where the probability measure degenerates to a point mass.

The welfare achieved by a policy $\Pi_\theta$ is defined as the expected outcome under that policy,

align[align omitted — 115 chars of source]

This integral is well-defined whenever $\mathbb{E}[\sup_{t \in \mathcal T}|Y(t)|] < \infty$. The identification of the welfare function typically relies on the following two assumptions.

assumption[Unconfoundedness] $Y(t)\perp T| X$, $\forall t\in\mathcal{T}$.
assumption[Overlap] For all $\theta \in \Theta$ and $x \in \mathcal{X}$, the support of $\Pi_\theta(\cdot|x)$ is contained in the support of $T|X=x$.

Under these assumptions, the welfare for a policy parameter $\theta$ is

align*[align* omitted — 64 chars of source]

where $\mu_\theta(x)=\int m(t, x) \Pi_\theta(dt| x)$ is the conditional value of the policy given covariates $x$, and $m(t, x)=\mathbb E[Y| T=t,X=x]$ is the conditional mean outcome.

The optimal policy is then given by $\Pi_{\theta^*}$, where

align*[align* omitted — 72 chars of source]

so that $\Pi_{\theta^*}$ attains the maximal welfare among all policies in the class $\{\Pi_\theta:\theta\in\Theta\}$. At this stage the argmax need not exist or be unique, but in the subsequent efficiency analysis we impose conditions ensuring both existence and uniqueness.

We observe an independent and identically distributed sample $(Z_i)_{i=1}^n$ from the distribution of $Z=(Y,T,X)$. An estimator $\hat\theta$ maps the sample into the parameter space, and the corresponding welfare of the estimated policy is $W(\hat\theta)$. The performance of a policy estimator is typically evaluated by its regret, defined as the difference between the optimal and achieved welfare, $$ R(\hat\theta)=W(\theta^*)-W(\hat\theta). $$ Existing work has largely focused on nonasymptotic or high-probability bounds for the regret. In contrast, we study the asymptotic behavior of the policy estimator, with particular emphasis on its efficiency and the limiting distribution of the regret.

ai2021unified develop an efficient estimation framework using stabilized weights for general treatment effect functionals: \[ \theta(P) =\operatorname{argmin}_{\theta\in\mathbb{R}^{d_{\theta}}} \mathbb{E}\left[\frac{f_T(T)}{f(T|X)} L\big(Y-g(T,\theta)\big)\right], \] for a function $g$ parametrized by \(\theta\) and a generic loss function $L$. Their estimands are defined by integrating with respect to the marginal distribution of the observed treatment, while our target integrates with respect to the counterfactual distribution of the treatment induced by the policy itself.

Semiparametric Efficiency Theory

Let \( f(\cdot| X) \) denote the conditional density of \( T \) given \( X \) with respect to a dominating measure. We use \( \pi_\theta(t | x) \) to denote the density function of a generic policy measure $\Pi_\theta(dt | x)$.

Pathwise differentiability

We briefly introduce the notion of pathwise differentiability. Let $\mathcal{P}$ denote the statistical model implied by Assumptions (ref) and (ref), i.e., the collection of all distributions $P$ of $Z=(Y,T,X)$ that satisfy Assumptions (ref) and (ref). For any given $P\in\mathcal{P}$, a regular parametric submodel (or path) through $P$ is a collection of $\{P_\varepsilon:\varepsilon\in(-r,r)\}\subset\mathcal{P}$, where $r$ denotes the radius of a parametric path domain, such that: (i) $P_0=P$; (ii) all $P_\varepsilon$ are dominated by a common $\sigma$-finite measure $\nu$; and (iii) the map $\varepsilon\mapsto \log\frac{\partial P_\varepsilon}{\partial\nu}(Z)$ is differentiable at $0$ in $L^2(P)$. The score of the path is $s(Z)=\frac{\partial}{\partial\varepsilon}\log\frac{\partial P_\varepsilon}{\partial \nu}(Z)\big|_{\varepsilon=0}$ in $L^2_0(P)$, where $L^2_0(P)=\{f\in L^2(P):\mathbb{E}_P[f(Z)]=0\}$. The mean-zero tangent space $\mathcal{S}$ is the $L^2(P)$-closure of all such scores.

A generic parameter $\psi:\mathcal{P}\to\mathbb{R}^d$ is said to be pathwise differentiable at $P$ relative to $\mathcal{S}$ if there exists $\phi\in{\mathcal{S}}\subset L^2_0(P)$ such that, for every regular submodel $\{P_\varepsilon\}$ with score $s$, \[ \frac{\partial\psi(P_\varepsilon)}{\partial\varepsilon}\Big|_{\varepsilon=0} = \mathbb{E}_P \big[\phi(Z) s(Z)\big]. \] Pathwise differentiability is necessary for the existence of $\sqrt{n}$-regular estimators of $\psi$ in $\mathcal{P}$ van1991differentiable, van1998asymptotic.\footnote{An estimator is called regular at the true law \(P\) if it is asymptotically linear and its first‑order limiting distribution is stable under all $1/\sqrt{n}$-local (contiguous) parametric perturbations of $P$.}

The (non-)pathwise differentiability of the welfare functional \( W(\theta) \) and the associated optimal policy \( \theta^* \) is not ex ante obvious. In the binary-treatment setting, one might expect \( W(\theta) \) to behave analogously to a treatment effect parameter and hence be pathwise differentiable under standard regularity conditions. However, the corresponding optimal deterministic policy is not pathwise differentiable due to the indicator function's inherent nonsmoothness. crippa2025regret recently demonstrates that in such cases, the optimal policy exhibits cube-root asymptotics, as characterized by kim1990cube. Thus, achieving pathwise differentiability requires excluding deterministic policies and instead considering sufficiently smooth randomized policies.

For continuous treatments, the welfare functional under a deterministic policy \(\Pi_\theta(dt | x) = \delta_{g_\theta(x)}(dt) = \mathbf{1}\{ g_\theta(x) \in dt \}\), where $g_\theta(x)$ is an arbitrary function parametrized by $\theta$, may superficially appear to be \(\sqrt{n}\)-estimable, since it can be expressed as an average/integral of a nonparametric function, $W(\theta) = \mathbb{E}[m(g_\theta(X), X)]$.

However, the function \( m(t,x) \) depends on both \( T \) and \( X \), while the integration in \( W(\theta) \) is taken only along the one-dimensional manifold \( \{(g_\theta(x), x)\} \), which has zero measure in the joint \((T,X)\)-space. Consequently, \( W(\theta) \) is not pathwise differentiable. The following theorem formalizes this intuition.

assumption$\Theta\subset \mathbb{R}^d$ is compact. The welfare $W:\Theta\to\mathbb{R}$ is twice differentiable on a neighborhood of $\theta^*\in\operatorname{int}(\Theta)$. Moreover, $\theta^*$ is the unique maximizer of $W$ on $\Theta$, $\frac{\partial W(\theta)}{\partial\theta}|_{\theta=\theta^*}=0$ and $H\equiv-\frac{\partial^2 W(\theta)}{\partial\theta\partial\theta'}|_{\theta=\theta^*}$ is invertible.
theoremIf the conditional law of $T|X=x$ admits a Lebesgue density $f(\cdot|x)$, then the welfare $W(\theta)$ of a deterministic policy $\Pi_\theta(dt|x)$ is not pathwise differentiable. Consequently, under Assumptions (ref) - (ref), the optimal policy parameter $\theta^*\in\operatorname{int}(\Theta)$ with $H$ is also not pathwise differentiable.

The preceding discussion and Theorem (ref) indicate that, in general treatment models, achieving pathwise differentiability (and consequently \(\sqrt{n}\)-estimability) of the optimal policy necessitates focusing on classes of randomized policies. This issue can be understood from a unified perspective for general treatment model in terms of the richness of the treatment space. In discrete treatment settings, the treatment space is too sparse, rendering policy mappings inherently discontinuous and difficult to estimate. Conversely, when the treatment space is continuous, deterministic policies can be smooth functions, yet the corresponding welfare functionals become harder to estimate.

These observations provide a rigorous theoretical justification for the focus on randomized policy classes, complementing the earlier fairness and resource-scarcity motivations. In line with this rationale, Sections (ref) and (ref) develop the semiparametric efficiency theory and examine the (in)efficiency of estimators for these pathwise-differentiable policy parameters.

Efficiency theory for regret

Throughout this subsection, define $\sigma^2(t,x) = \operatorname{Var}(Y | T = t, X = x)$ as the conditional variance of the outcome. Denote $\int dt$ as integration with respect to this fixed $\sigma$-finite measure on $\mathcal{T}$: Lebesgue for continuous $\mathcal{T}$, counting measure for discrete $\mathcal{T}$, and the mixture for mixed $\mathcal{T}$. For discrete $\mathcal{T}$, $\int dt$ should read as $\sum_{t\in\mathcal{T}} $.

In the next lemma, we derive the efficient influence function (EIF) for the welfare functional. As discussed earlier, if the conditional law of $T|X$ is absolutely continuous with respect to Lebesgue measure, then $W(\theta)$ under a deterministic policy is not pathwise differentiable. Consequently, an EIF and the usual semiparametric efficiency bound do not apply. We therefore focus on the regular case in this and the next section by restricting attention to randomized policies. Deterministic policies are reintroduced in Section (ref), where they are handled within our regret analysis framework.

lemmaLet Assumptions (ref) and (ref) hold. Assume for each $\theta$, \begin{enumerate}[label=(\arabic*)] • $\mathbb{E}\left[\frac{\pi_\theta(T| X)^2}{f(T| X)} \sigma^2(T,X)\right]<\infty$ and $\mathbb{E}\left[\mu_\theta(X)^2\right]<\infty$; • $m(\cdot,\cdot)$ and $\sigma^2(\cdot,\cdot)$ are measurable and finite, \end{enumerate} then the welfare $W(\theta)$ is pathwise differentiable in $\mathcal{P}$ with efficient influence function \begin{equation} \varphi_\theta(Z) =\frac{\pi_\theta(T|X)(Y-m(T,X))}{f(T|X)} + \mu_\theta(X)-W(\theta), \end{equation} and efficiency bound \begin{align*} \operatorname{Var}(\varphi_\theta) &= \mathbb{E}\left[\frac{\pi_\theta(T|X)^2\sigma^2(T,X)}{f(T|X)}\right]+\operatorname{Var}\left(\mu_\theta(X)\right). \end{align*} If the conditional density \( f(\cdot | \cdot) \) is known, the tangent space becomes smaller, while the form of the efficient influence functions remains unchanged.\footnote{As noted in chen2025local, the model is locally just identified when the propensity score is known, and becomes locally overidentified when the propensity score must be estimated.}

For the binary treatment case where $\mathcal{T}=\{0,1\}$, denote $p(X) = P(T=1|X)$, $\pi_\theta(1|x) = \pi_\theta(x)$, and $\pi_\theta(0|x)=1-\pi_\theta(x)$, the EIF becomes

align[align omitted — 222 chars of source]

This is the standard doubly robust policy value score used in the policy learning literature. In particular, it coincides with the influence function based estimand studied in athey2021policy on efficient policy evaluation and learning.

To analyze the efficiency of the optimal policy parameter \( \theta^* \), we impose mild smoothness conditions in the following assumption, ensuring that (ref) is differentiable in $\theta$, and hence restricts our attention to randomized policies. For any vector $v\in\mathbb{R}^d$, $\|v\|$ denotes the Euclidean norm.

assumptionThere exists a neighborhood $\mathcal N$ of $\theta^*$ such that, for almost every \ $(t,x)$, $\theta\mapsto \pi_\theta(t| x)$ is differentiable on $\mathcal N$ with derivative $\partial_\theta\pi_\theta(t| x)\in\mathbb R^d$, and, uniformly in $\theta\in\mathcal N$, \begin{align*} \mathbb{E}\left[\left\|\frac{\partial\pi_\theta(T| X)}{\partial \theta}\right\|^2\frac{\sigma^2(T,X)}{f(T| X)} \right]<\infty, \qquad \mathbb{E}\left[\Big\|\int \frac{\partial\pi_\theta(t| X)}{\partial\theta} m(t,X) dt\Big\|^2\right]<\infty. \end{align*} Moreover, for almost every $x$, differentiation and integration are interchangeable for $\mu_\theta(x)$, i.e. $\frac{\partial\mu_\theta(x)}{\partial\theta}=\int m(t,x) \frac{\partial\pi_\theta(t| x)}{\partial\theta} dt$. The policy has uniformly bounded densities: $\sup_{\theta \in \Theta} \sup_{t,x} \pi_\theta(t \mid x) < \infty$.

The next theorem derives the efficient influence function for the policy parameter, and then applies the H\'ajek–Le Cam convolution theorem to characterize the limiting distribution of any regular estimator.

theoremLet Assumptions (ref) - (ref) and the conditions in Lemma (ref) hold. The policy parameter $\theta^*$ is pathwise differentiable with EIF \begin{align*} \mathrm{EIF}_{\theta^*}(Z) & = H^{-1}\frac{\partial{\varphi}_\theta(Z)}{\partial\theta}\Big|_{\theta=\theta^*}, \end{align*} where \begin{align*} \frac{\partial{\varphi}_\theta(Z)}{\partial\theta}\Big|_{\theta=\theta^*}&=\left(\int m(t,X)\frac{\partial\pi_\theta(t|X)}{\partial\theta}\Big|_{\theta=\theta^*}dt + \frac{Y-m(T,X)}{f(T|X)}\cdot\frac{\partial\pi_\theta(T|X)}{\partial\theta}\Big|_{\theta=\theta^*}\right). \end{align*} Consequently, the semiparametric efficiency bound for $\theta^*$ is $V_{\mathrm{eff}} \equiv H^{-1}\operatorname{Var}(\frac{\partial{\varphi}_\theta(Z)}{\partial\theta}\big|_{\theta=\theta^*})H^{-1}$. If the conditional density \( f(\cdot | \cdot) \) is known, the form of the efficient influence functions remains unchanged. Moreover, for any regular estimator $\hat{\theta}$, \begin{align} \sqrt{n}\big(\hat{\theta}-\theta^*\big) \Rightarrow G + U, \end{align} where $G\sim\mathcal{N}(0,V_{\mathrm{eff}})$, $U$ is mean-zero, independent of $G$, and has covariance $\Sigma_U$, and the symbol “$\Rightarrow$” denotes convergence in distribution.

In the binary treatment case, the efficient influence function simplifies to the following familiar form:

align*[align* omitted — 181 chars of source]
theoremLet the assumptions of Theorem (ref) hold. The regret $R(\hat{\theta})$ of any regular estimator $\hat{\theta}$ has the following asymptotic distribution: \begin{align*} nR(\hat{\theta}) \Rightarrow \frac{1}{2}(G+U)'H(G+U). \end{align*} The asymptotic distribution has mean \begin{align*} \frac{1}{2}\operatorname{tr}\left(H^{1/2}\left(V_{\mathrm{eff}}+\Sigma_U\right)H^{1/2}\right) \end{align*} and variance \begin{align*}\frac{1}{2}\sum_{j=1}^d \lambda_j^2 + \operatorname{tr}\left(H^{1/2}V_{\mathrm{eff}}H\Sigma_U H^{1/2}\right) + \frac{1}{4}\operatorname{Var}\left(U'HU\right), \end{align*} where $\lambda_1\ge\cdots\ge\lambda_d>0$ are the eigenvalues of $H^{1/2} V_{\mathrm{eff}} H^{1/2}$. If $U$ is Gaussian, the variance is $\frac{1}{2}\sum_{j=1}^d\tilde{\lambda}_j^2$, where $\tilde{\lambda}_1\ge\cdots\ge\tilde{\lambda}_d>0$ are the eigenvalues of $H^{1/2}\left(V_{\mathrm{eff}}+\Sigma_U\right)H^{1/2}$.

Theorem (ref) is the central result for describing regret efficiency. The limiting regret takes a quadratic form because welfare is locally quadratic in the policy parameters. This structure implies that the mean of the limiting regret depends directly on the covariance of the estimation error for the policy parameters.\footnote{More specifically, if $V_1\preceq V_2$, then $\tfrac{1}{2}\operatorname{tr} \big(H^{1/2}V_1H^{1/2}\big)\ \le\ \tfrac{1}{2}\operatorname{tr} \big(H^{1/2}V_2H^{1/2}\big)$.} As a consequence, any inefficiency in the estimation of the policy parameters leads not only to larger dispersion in regret but also to a strictly higher mean. This upward shift in the expected regret is a rather surprising theoretical implication, revealing a phenomenon that is not captured by the predominant regret bounding paradigm in the policy learning literature.

Regret Performance of Policy Learning Methods

The previous section develops a formal notion of efficiency for regret. In this section, we study the asymptotic normality of three specific classes of policy estimators: inverse probability weighting (IPW) using the true propensity, IPW using an estimated propensity, and the doubly robust (DR) estimator. Once asymptotic normality is established, standard analytic or bootstrap methods can be used to conduct inference on the policy parameters.

IPW estimator with true propensity

By the law of iterated expectations, the welfare $W(\theta)$ in ((ref)) admits the following IPW expression:

align*[align* omitted — 87 chars of source]

When the true propensity $f(\cdot | \cdot)$ is known, the welfare can be estimated by the sample analogue, and the policy estimator is obtained by maximizing it:

align*[align* omitted — 195 chars of source]

Here, the superscript “tp” stands for true propensity. This estimator serves as the analogue of the empirical welfare maximization estimator of kitagawa2018should in the setting of randomized policies and general treatments.

assumption$|Y(t)|\leq M<\infty$, $\forall t$. There exists $\underline{f}>0$, such that $f(t|x)\geq \underline{f}$ for all $(t,x)$.

Assumption (ref) imposes boundedness on the potential outcomes and requires that the true density be bounded away from zero. These are standard assumptions in the policy learning literature.

theoremUnder Assumptions (ref) -(ref), the estimator $\hat{\theta}^{tp}$ satisfies \begin{align*} \sqrt{n}\left(\hat{\theta}^{tp}-\theta^*\right)\Rightarrow G+ U^{tp}, \end{align*} where $G\sim N(0,V_{\mathrm{eff}})$ and $U^{tp}\sim N(0,\Sigma_U^{tp})$ are independent, and \begin{align*} \Sigma_U^{tp} = H^{-1}\mathbb{E}\left[\left(\frac{m(T,X)}{f(T|X)}\frac{\partial \pi_{\theta}(T|X)}{\partial\theta}\Big|_{\theta=\theta^*}-\frac{\partial_\theta \mu_{\theta}(X)}{\partial\theta}\Big|_{\theta=\theta^*}\right)^2\right]H^{-1}. \end{align*} The corresponding regret has the following asymptotic distribution \[nR(\hat{\theta}^{tp})\Rightarrow \frac{1}{2}(G+U^{tp})'H(G+U^{tp}),\] with mean $\frac{1}{2}\operatorname{tr}\left(H^{1/2}(V_{\mathrm{eff}}+\Sigma_U^{tp})H^{1/2}\right)$ and variance $\frac{1}{2}\sum_{j=1}^d (\tilde{\lambda}_j^{tp})^2$, where $\tilde{\lambda}_1^{tp}\geq \dots \geq \tilde{\lambda}_d^{tp}$ are the eigenvalues of $H^{1/2}(V_{\mathrm{eff}}+\Sigma_U^{tp})H^{1/2}$.

The matrix $\Sigma_U^{tp}$ is positive definite except in the degenerate case where $\frac{m(T,X)}{f(T,X)}\cdot\frac{\partial\pi_{\theta}(T|X)}{\partial\theta}\Big|_{\theta=\theta^*}$ is almost surely constant in $T$ given $X$. Hence, $\Sigma_U^{tp}\neq 0$ and IPW using true propensity leads to inefficiency. This arises because its influence function contains a component orthogonal to the tangent space and therefore cannot coincide with the EIF.

The inefficiency of IPW-type estimators with true propensity in treatment effect estimation is well documented in the literature hirano2003efficient,chen2008semiparametric,ai2021unified. In policy learning, the consequences are stronger: regret is not only more variable but also has a strictly higher mean.

This yields a clear practical message. Even in experimental settings where the true propensity $f(\cdot|\cdot)$ is known, plugging it into IPW for policy learning is suboptimal. Instead, as we discuss next, the efficient alternative uses an estimated propensity, which removes the $\Sigma_U^{tp}$ term and attains the semiparametric efficiency bound for both the estimator and the regret.

IPW with estimated propensity

Let $\omega(t,x)=1/f(t|x)$ be the true weight in IPW. Then the IPW identification condition can be written as $W(\theta) = \mathbb{E}[\pi_\theta(T|X)\omega(T,X)Y]$. We now consider a two-step approach that first estimates the weighting function $\omega$ and then plugs it into the welfare estimator for policy learning.

The estimation of $\omega$ builds on the stabilized weighting approach of ai2021unified. In their setting, the weighting function takes the form \( \frac{f_T(t)}{f(t | x)} \), where $f_T$ denotes the marginal density of the treatment. In contrast, our framework differs in that integration is performed with respect to the policy measure $\Pi(dt |x)$, rather than the observed treatment distribution. The identification of $\omega(T, X)$ follows from the balancing condition that, for all integrable functions $u(T)$ and $v(X)$,

align*[align* omitted — 111 chars of source]

which uniquely characterizes $\omega(t, x)$. Since \(\int u(t)\, dt\) is directly computable, the only remaining component to estimate is \(\mathbb{E}[v(X)]\).

We approximate the function space of $u$ and $v$ with finite-dimensional sieves. Let $u_{K_1}(T) = (u_{K_1,1}(T),\dots,u_{K_1,K_1}(T))'$ and $v_{K_2}(X)=(v_{K_2,1}(X),\dots,v_{K_2,K_2}(X))'$ denote the basis functions with dimension $K_1, K_2\in\mathbb{N}$, and $K=K_1K_2$. Note that, when the treatment $T$ is discrete, the space of $u$ is finite-dimensional and only $v$ requires approximation. Consider the following entropy-tilting program:

align*[align* omitted — 245 chars of source]

Let $\rho(v)=-\exp(-v-1)$ with $\rho'(v) = \exp(-v-1)>0$. Since the dual problem is smooth and strictly concave, the estimated weight can be solved as

align*[align* omitted — 318 chars of source]
assumption\ \begin{enumerate}[label = (\arabic*)] • The supports $\mathcal{X}$ and $\mathcal{T}$ are compact. • There exist $\Lambda_{K_1\times K_2}\in\mathbb{R}^{K_1\times K_2}$ and a positive constant $\alpha > 0$, such that \begin{align*} \left\|(\rho')^{-1}(\omega(t,x))-u_{K_1}(t)'\Lambda_{K_1\times K_2}v_{K_2}(x)\right\|_\infty=O(K^{-\alpha}). \end{align*} • For every $K_1$ and $K_2$, the smallest eigenvalues of $\int u_{K_1}(t)u_{K_1}(t)'dt$ and $\mathbb{E}[v_{K_2}(X)v_{K_2}(X)']$ are bounded away from zero uniformly in $K_1$ and $K_2$. • There exist two sequences of constants $\zeta_1(K_1)$ and $\zeta_2(K_2)$ satisfying $\|u_{K_1}(t)\|_\infty\leq \zeta_1(K_1)$ and $\|v_{K_2}(x)\|_\infty\leq \zeta_2(K_2)$. Let $(K_1, K_2)$ be chosen with $K = K_1K_2$, $\zeta(K) = \zeta_1(K_1)\zeta_2(K_2)$, such that $\zeta(K)\sqrt{K^2/n}\rightarrow 0$, and $\sqrt{n}K^{-\alpha}\rightarrow 0$. \end{enumerate}

The above assumption adapts the conditions of ai2021unified to our setting and yields the following asymptotic properties of the estimated weighting function.

lemmaUnder Assumption (ref)(ii) and (ref), we obtain the following rate regarding the estimated weight function: \begin{align*} &\int_{\mathcal{T}\times\mathcal{X}} |\hat{\omega}_K(t,x)-\omega(t,x)|^2 dF_{T,X}(t,x) = O_p\left(K/n\right),\\ & \frac{1}{n}\sum_{i=1}^n |\hat{\omega}_K(T_i,X_i)-\omega(T_i,X_i)|^2 = O_p\left(K/n\right). \end{align*}

After obtaining \( \hat{\omega}_K \), we substitute it into the construction of the welfare estimator and obtain the corresponding policy estimator as follows:

align*[align* omitted — 188 chars of source]

The next theorem presents the asymptotic distribution of the IPW welfare policy estimator with estimated weights and the corresponding regret.

theoremUnder Assumption (ref)-(ref), the estimator \( \hat{\theta}^{\mathrm{ep}} \) satisfies \begin{align*} \sqrt{n}(\hat{\theta}^{ep}-\theta^*)\Rightarrow G, \end{align*} with $G\sim N(0,V_{\mathrm{eff}})$. The corresponding regret has asymptotic distribution $nR(\hat{\theta}^{ep})\Rightarrow \frac{1}{2} G'HG$, and therefore this policy estimator attains the efficient regret.

As established in Theorem (ref), the IPW estimator based on the estimated weighting function attains the efficient regret, irrespective of whether the propensity is known. This finding can be viewed as the policy-learning analogue of the HIR efficiency phenomenon originally found by hirano2003efficient. To the best of our knowledge, this insight has not been explicitly recognized in the policy learning literature. In particular, studies such as kitagawa2018should and mbakop2021model derive regret bounds for empirical welfare maximization under unknown propensity that depend on the convergence rate of the propensity estimator, which is strictly slower than the \(\sqrt{n}\)-rate in nonparametric settings. In contrast, our result demonstrates that employing an estimated propensity not only preserves the \(\sqrt{n}\)-rate for the regret but also yields a smaller asymptotic mean and variance.

In observational studies, even when certain covariates are known to affect only the outcome and not the treatment assignment process, it is still desirable to include them in the propensity score estimation. From a theoretical perspective, the propensity score effectively accounts for both the treatment selection mechanism and the outcome process.

It is worth noting that our analysis thus far does not cover deterministic policies, which is the focus of kitagawa2018should and mbakop2021model. We extend the discussion to this case in Section (ref).

Doubly robust policy estimator

The EIF in Lemma (ref) yields the following doubly robust welfare expression:

align*[align* omitted — 334 chars of source]

This expression corresponds to the IPW welfare discussed in the previous section, augmented by the following mean-zero adjustment term:

align*[align* omitted — 210 chars of source]

Based on the DR identification equation, we propose to estimate welfare using a doubly robust estimator. This requires first-stage estimators of the nuisance functions \(\omega(t,x)\) and \(m(t,x)\). The nuisance estimators are required to satisfy the following conditions.

assumption\ \begin{enumerate}[label = (\arabic*)] • There exists $\rho_\omega, \rho_m>0$, with $\rho_\omega + \rho_m\geq \frac{1}{2}$ such that $\left\|\hat{\omega}-\omega\right\|_{L_2} = o_p(n^{-\rho_\omega})$, and $\left\|\hat{m}-m\right\|_{L_2} = o_p(n^{-\rho_m})$.\footnote{Here, $\|\cdot\|_{L_2}$ is the $L_2$-norm under the distribution of $(T,X)$.} • With probability approaching one, $\hat{m}$ is bounded. \end{enumerate}

Assumption (ref)(1) concerns the mean squared convergence rate of $\hat{\omega}$ and $\hat{m}$ in the $L_2$ space. Assumption (ref)(2) requires the estimator to have the same property of boundedness as its estimand. We can construct an estimator of $m$ by employing sieve-based methods chen2007large, local polynomial methods calonico2018effect, linear and nonlinear partitioning-based methods cattaneo2024uniform, or machine learning techniques chernozhukov2018double. The estimator of \(\omega\) may be the one studied in Section (ref), or first estimate the conditional density $f$ and then take its inverse, applying the techniques developed by cattaneo2024boundary or ColangeloLee2025.

We implement the following cross-fitting procedure to construct the doubly robust estimator of the welfare. Let the sample be partitioned into $L\geq 2$ folds, $\{\mathcal{I}_{\ell}\}_{\ell=1}^L$, of equal size. For each $\ell$-th fold, we estimate the nuisances on the complement $\mathcal{I}_{\ell}^c$, which include all the observations that are not in $\mathcal{I}_\ell$, obtaining $\hat{m}^{(-\ell)}$ and $\hat{\omega}^{(-\ell)}$, and evaluate them on $\mathcal{I}_\ell$. The superscript $(-\ell)$ signifies that the estimators are constructed using data not in $I_\ell$. Then the doubly robust welfare estimator is constructed as follows, and the doubly robust policy estimator is obtained by maximizing the welfare estimator:

align*[align* omitted — 309 chars of source]

where $\hat{\mu}^{(-\ell)}_\theta(x)=\int \hat{m}^{(-\ell)}(t,x)\pi_\theta(t|x)dt$. The doubly robust policy estimator and its regret have the following asymptotic properties.

theoremUnder Assumptions (ref)-(ref), and (ref), the estimator \( \hat{\theta}^{dr} \) satisfies \begin{align*} \sqrt{n}(\hat{\theta}^{dr}-\theta^*)\Rightarrow G, \end{align*} with $G\sim N(0,V_{\mathrm{eff}})$. The corresponding regret has asymptotic distribution $nR(\hat{\theta}^{dr})\Rightarrow \frac{1}{2} G'HG$, and therefore this policy estimator attains the efficient regret..

Neyman orthogonality eliminates first-order sensitivity to both nuisance estimators, so that estimation errors contribute only through a second-order remainder term. Together, orthogonality, uniform smoothness of the policy class, suitable nuisance rate conditions, and cross-fitting ensure this uniform control over $\theta$. Consequently, the estimator $\hat{\theta}^{dr}$ admits an influence function expansion with asymptotic variance $V_{\mathrm{eff}}$, achieving efficient regret.

Comparing with the literature, our asymptotic results apply to randomized policies, where the policy function $\pi_\theta(t|x)$ is continuous in $t$ and smoothly parametrized by $\theta$. Deterministic policies of the form $\pi_\theta(t|x)=\delta_{g_\theta(x)}(t)$ fall outside this framework and require a different analysis. For deterministic policies, doubly robust estimators have been developed in several settings, including binary treatment athey2021policy, multivalued discrete treatment zhou2023offline,fang2025model, and continuous treatment ai2024data. These works establish uniform convergence of the doubly robust welfare estimators over policy classes, an essential property for deriving regret bounds, but rely on distinct techniques tailored to deterministic policies.

From the proof of Theorem (ref), it follows directly that when the true propensity score is known, and the estimated propensity score is replaced by its true value, the nuisance estimator \(m\) is only required to be \(L_2\)-consistent, with no additional rate conditions. This is another instance of efficient estimator in experimental settings.

Regret Bounds via Welfare Processes

Up to this point, our theoretical results have focused on randomized policies. This creates a gap relative to the existing policy learning literature, where attention is largely centered on deterministic policies. In this section, we provide a complementary perspective to bridge this gap. Rather than studying the asymptotic distribution of regret, we analyze the regret through the supremum of the welfare process. This approach shows that the HIR phenomenon is not specific to randomized policies; it also arises more broadly for deterministic policies. For clarity of exposition, we focus on the binary treatment case.

To connect with the existing literature, we represent a policy \(\pi\) as a mapping from \(\mathcal{X}\) to the binary treatment space \(\{0,1\}\). Let \(\mathbf{\Pi}\) be a given policy class. In this setting, policies are not indexed by a finite-dimensional parameter \(\theta\); instead, the policy class is assumed to have finite VC dimension.\footnote{The analysis extends naturally to randomized policies \(\pi\colon \mathcal{X} \to [0,1]\). In that case, the VC subgraph dimension replaces the VC dimension.}

assumptionThe policy class $\mathbf{\Pi}$ has finite VC dimension $v<\infty$.

Let $\widehat{W}(\pi)$ be a regular welfare estimator of $W(\pi)=\mathbb{E}\left[Y(1)\pi(X)+Y(0)(1-\pi(X))\right]$, and

align*[align* omitted — 72 chars of source]

the estimated welfare-maximizing policy. The regret of $\hat{\pi}$ is defined as $R(\hat{\pi}) = W(\pi^*) - W(\hat{\pi})$, where $\pi^* = \arg\max_{\pi\in\mathbf{\Pi}} W(\pi)$. $R(\hat{\pi})$ satisfies the basic inequality kitagawa2018should:

align[align omitted — 295 chars of source]

Hence, the asymptotic behavior of the regret is governed by the supremum process \[ \sup_{\pi\in \mathbf{\Pi}} \left| \sqrt{n}\,\big(\widehat{W}(\pi) - W(\pi)\big) \right|. \] If the welfare process \(\sqrt{n}\,\big(\widehat{W}(\pi) - W(\pi)\big)\) weakly converges to a centered, tight Gaussian process \(G(\pi)\), then the continuous mapping theorem implies that \[ \sup_{\pi\in \mathbf{\Pi}} \left| \sqrt{n}\,\big(\widehat{W}(\pi) - W(\pi)\big) \right| \Rightarrow \sup_{\pi\in \mathbf{\Pi}} |G(\pi)|. \] Therefore, we may study the mean of \(\sup_{\pi\in \mathbf{\Pi}} |G(\pi)|\) as an indication of regret performance for different policy estimators.

We discuss the performance of the IPW estimators using true and estimated propensity. In the binary treatment setting, the weights are defined as $\omega_{1}(X) = 1/p(X)$ and $\omega_{0}(X) = 1/(1-p(X))$. We compare two IPW estimators of the welfare function $W(\theta)$: one that uses the true propensity score $p(X)$ to construct the weights, and another that replaces the true weights with estimated counterparts, denoted by $\hat{\omega}_1(X)$ and $\hat{\omega}_0(X)$. These two estimators are defined, respectively, as

align*[align* omitted — 288 chars of source]
theoremUnder Assumptions (ref), (ref), and (ref), \begin{align*} \sup_{\pi\in\mathbf{\Pi}}\Big|\sqrt{n}\left(\widehat{W}^{tp}(\pi)-W(\pi)\right)\Big| &=\sup_{\pi\in\mathbf{\Pi}}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^n \varphi^{tp}_\pi(Z_i)\right| \Rightarrow \sup_{\pi\in\mathbf{\Pi}}\big|G_{tp}(\pi)\big|, \end{align*} and if, in addition, Assumption (ref) also holds, then \begin{align*} \sup_{\pi\in\mathbf{\Pi}}\Big|\sqrt{n}\left(\widehat{W}^{ep}(\pi)-W(\pi)\right)\Big| &=\sup_{\pi\in\mathbf{\Pi}}\left|\frac{1}{\sqrt{n}}\sum_{i=1}^n \varphi^{ep}_\pi(Z_i)\right|+o_p(1) \Rightarrow \sup_{\pi\in\mathbf{\Pi}}\big|G_{ep}(\pi)\big|, \end{align*} where $G_{tp}$ and $G_{ep}$ are centered, tight Gaussian processes on $\Pi$ with covariance functions induced by the influence representations \begin{align} \varphi^{tp}_\pi(Z) &=\underbrace{\left(\omega_1(X)TY-\omega_0(X)(1-T)Y\right)}_{\equiv \tau_{tp}(Z)} \pi(X) +{\omega_0(X)(1-T)Y}-W(\pi), \nonumber\\ \varphi^{ep}_\pi(Z) &=\underbrace{\left(\omega_1(X)T(Y-m_1(X))-\omega_0(X)(1-T)(Y-m_0(X))+m_1(X)-m_0(X)\right)}_{\equiv \tau_{ep}(Z)} \pi(X) \nonumber \\ &+{\left(\omega_0(X)(1-T)(Y-m_0(X))+m_0(X)\right)}-W(\pi). \end{align} Moreover, the two expected suprema satisfy the following relationship \begin{align*} \mathbb{E}\left[\sup_{\pi\in\mathbf{\Pi}}\big|G_{ep}(\pi)\big|\right] \le \mathbb{E}\left[\sup_{\pi\in\mathbf{\Pi}}\big|G_{tp}(\pi)\big|\right]. \end{align*}

Theorem (ref) shows that the welfare processes associated with the IPW estimators based on the true and estimated propensity scores converge weakly to Gaussian processes with covariance structures determined by the influence functions \(\varphi^{tp}_{\pi}\) and \(\varphi^{ep}_{\pi}\), respectively. In particular, \(\varphi^{ep}_{\pi}\) coincides with the efficient influence function in equation (ref). This implies that the IPW welfare estimator constructed with the estimated propensity score admits a uniform linear expansion with the efficient influence function over the entire policy class.

The comparison of the expected suprema relies on the Sudakov-Fernique inequality sudakov1971gaussian, fernique1975regularite,chatterjee2005error. The terms \(\tau_{tp}\) and \(\tau_{ep}\) in (ref) satisfy \( \mathbb{E}\!\left[ \tau_{ep}(Z)^{2} | X \right] \le \mathbb{E}\!\left[ \tau_{tp}(Z)^{2} | X \right], \) which yields \[ \mathbb{E}\!\left[ (G_{ep}(\pi)-G_{ep}(\pi'))^{2} \right] \le \mathbb{E}\!\left[ (G_{tp}(\pi)-G_{tp}(\pi'))^{2} \right], \qquad \forall \pi,\pi' \in \mathbf{\Pi}. \] The Sudakov--Fernique inequality then implies the desired comparison result. This result shows that using the estimated rather than the true propensity score weakly decreases the asymptotic mean of the regret bound arising from the basic inequality. The additional first-stage estimation step improves efficiency at the level of the welfare process, providing another instance of the HIR phenomenon in policy learning.

Empirically Calibrated Simulations

To evaluate finite sample regret under realistic heterogeneity, we conduct empirically calibrated simulations based on the National Job Training Partnership Act (JTPA) dataset, which has been widely used in influential policy learning studies kitagawa2018should,kitagawa2021equality,mbakop2021model,crippa2025regret. The JTPA data contain rich baseline covariates and exhibit substantial treatment effect heterogeneity, providing a familiar and empirically relevant setting for comparing policy estimators. Our simulation design preserves the empirical distribution of covariates while using flexible, data-driven estimates of the conditional outcome models as the true nuisance functions, generating complex and nonlinear heterogeneity that closely mirrors the structure of the original data while retaining oracle access to welfare and regret for benchmarking our theoretical results.

The data-generating process (DGP) is as follows. We use variables \((Y^*, T^*, X^*)\) with superscript “\(*\)” to denote variables in the JTPA data, and variables \((Y, T, X)\) without superscripts to denote the corresponding variables in the DGP. In the JTPA data, the outcome \(Y^*\) is the post-treatment earnings, \(T^*\) the treatment indicator, and \(X^*\) the baseline covariates observed prior to treatment assignment, including education, pre-treatment earnings, age, race, gender, and other demographic characteristics. For each Monte Carlo replication, we draw a sample of size \(n \in \{500, 1000, 1500\}\) by sampling with replacement from the empirical distribution of $X^*$ data to construct calibrated covariates $X$.

Using the full JTPA dataset with sample size 8192, we first estimate the conditional mean outcome functions and treat the fitted functions \(m_0^*(\cdot)\) and \(m_1^*(\cdot)\) as the true nuisance functions in the simulations. This approach preserves the nonlinearities and interaction patterns present in the original data, generating heterogeneous treatment effects \(\tau^*(x) = m_1^*(x) - m_0^*(x)\) that are not well approximated by simple parametric models. Potential outcomes are generated according to: $Y(t) = m_t^*(X) + \varepsilon_t, t\in\{0,1\}$, where the error terms $\varepsilon_t$ are drawn from symmetric uniform distributions with variances calibrated to match the estimated residual variance of $Y^*$.

Although treatment assignment in the JTPA study is randomized with a constant probability, we introduce a nontrivial assignment mechanism to reflect stratified experimental designs while remaining distinct from the policy class under consideration. Specifically, treatment is assigned according to the propensity score: $p(x) = \mathbb{P}(T=1|X=x)=\frac{\exp(0.5 - 0.5\texttt{edu})}{1+\exp(0.5 - 0.5\texttt{edu})}$, which ensures overlap and delivers similar marginal distribution as that of $T^*$. For each Monte Carlo replication, treatment assignments $T$ are drawn according to this propensity model, and observed outcomes are constructed as $Y=TY(1)+(1-T)Y(0)$.

Policies are specified as functions of education level and pre-program earnings, denoted by edu and prevearn. We consider the logistic policy class \[ \pi_\theta(x) = \frac{\exp\!\left(\theta_0 + \theta_1\,\texttt{edu} + \theta_2\,\texttt{prevearn}\right)}{1 + \exp\!\left(\theta_0 + \theta_1\,\texttt{edu} + \theta_2\,\texttt{prevearn}\right)}. \] Within each replication, we estimate conditional mean models with the random forest and the inverse of propensity score with the stabilized weights procedure as illustrated in Section (ref), using the simulated sample. We then construct the three welfare estimators discussed in Section (ref): $\hat{\theta}^{tp}$, $\hat{\theta}^{ep}$, and $\hat{\theta}^{dr}$.

To evaluate regret, we generate an independent large test sample of size $10^6$ from the same empirically calibrated DGP and compute \[ \theta^* = \arg\max_{\theta\in\Theta} W(\theta), \qquad R(\hat\theta) = W(\theta^\star) - W(\hat\theta), \] where $W(\cdot)$ is approximated by the test-sample average using the {true} potential outcomes. In our implementation, the oracle welfare on the test sample is $W(\theta^\star)=15.987$.

We report results over 1000 Monte Carlo replications for sample sizes $n=500$, $1000$, and $1500$. Table (ref) summarizes the mean and standard deviation of regret for each estimator across replications. Figure (ref) plots the empirical density of $nR(\hat{\theta})$ for each estimator and sample size. The distributions are consistent with the weighted $\chi^2$ limit in Theorem (ref).

table[table omitted — 771 chars of source]
figure[figure omitted — 679 chars of source]

Across all sample sizes, the IPW estimator using the estimated propensity score and the doubly robust estimator achieve smaller average regret and lower variability than the IPW estimator using the true propensity score. This is consistent with the HIR phenomenon discovered in the previous sections.

commentIn this section, we use Monte Carlo simulations to compare the regret performance of different policy estimators derived from the welfare estimators introduced in Section (ref). We evaluate three approaches: the IPW estimator with the true weight, denoted $\widehat{W}^{tp}(\theta)$; the IPW estimator with the estimated weight, denoted as $\widehat{W}^{ep}(\theta)$; the doubly robust estimator $\widehat{W}^{dr}(\theta)$. We examine these estimators in binary treatment contexts. We begin with the binary treatment case and construct three data-generating processes (DGPs) that capture varying degrees of heterogeneity, nonlinearity, and covariate overlap. Each DGP uses five independent covariates. Let $X_1, X_2,X_3,X_3,X_5\sim U[0,1]$, independently. The propensity score takes simple constant as $p(X)=0.5$. DGP1 admits a linear treatment effect, and DGP2 describes the setting of nonlinear treatment effect. The potential outcomes $Y_i(0)$ and $Y_i(1)$ of the two DGPs are defined as follows, with $e_0,e_1\sim U[-10,10]$, \begin{enumerate}[ leftmargin=*, label={DGP \arabic*:}, ref={DGP \arabic*}] • $Y_i(0) =10 + e_{0i}$, and $Y_i(1) = 10 + 5(X_{1i}-X_{2i}) + e_{1i}$. • $Y_i(0)=10 + e_{0i}$, and $Y_i(1)=10 + 5(X_{1i}-X_{2i}) - 10 (X_{1i}-X_{2i})^2 +e_{1i}$. \end{enumerate} For every DGP, we simulate 5000 Monte Carlo replications with three sample sizes $n=1000, 2000, 5000$. The oracle optimal policy parameter $\theta^*$ and the regret $R(\hat{\theta})=W(\theta^*)-W(\hat{\theta})$ are computed on an independent large testing sample. Results are summarized in Table (ref). \begin{table}[H] \begin{tabular}{cc p{0.01mm} ccp{0.01mm}ccp{0.01mm}cc} \hline\hline \multirow{2}{*}{DGP}& \multirow{2}{*}{n}& & \multicolumn{2}{c}{$R(\hat{\theta}^{tp})$}& &\multicolumn{2}{c}{$R(\hat{\theta}^{ep})$}& &\multicolumn{2}{c}{$R(\hat{\theta}^{dr})$}\\ & & & mean& sd& &mean& sd& &mean& sd\\ \hline \multirow{3}{*}{DGP1} & 1000 && 0.146 & 0.210 & & 0.099 & 0.147 & & 0.037 & 0.055 \\ & 2000 && 0.077 & 0.122 & & 0.056 & 0.077 & & 0.019 & 0.027 \\ & 5000 & & 0.036 & 0.049 & & 0.030 & 0.041 & & 0.009 & 0.013 \\ \hline \multirow{3}{*}{DGP2} & 1000 && 0.035 & 0.060 && 0.028 & 0.055 & & 0.013 & 0.026 \\ & 2000 & & 0.024 & 0.041 & & 0.016 & 0.034 && 0.006 & 0.015 \\ & 5000 && 0.012 & 0.025 && 0.007 & 0.018 && 0.002 & 0.005 \\ \hline \hline \end{tabular} \caption{Comparison of different estimators under a binary treatment} \caption*{“mean” and “sd” report, respectively, the average and standard deviation of the regret for each policy estimator across 5,000 replications.} \end{table} Across all DGPs, the IPW estimator with the estimated weight and the doubly robust estimator exhibit smaller average regret and lower variability compared to the IPW estimator with the true weight. The simulations thus provide a concrete manifestation of the Hirano–Imbens–Ridder phenomenon, showing that estimating the propensity score can reduce regret variance and improve finite-sample performance.

Empirical Study

In this section, we apply the policy learning methods developed in Section (ref) to data from the field experiment of ashraf2006tying. The experiment was implemented in collaboration with a rural bank in the Philippines to evaluate whether offering a voluntary commitment savings product could help individuals overcome self-control problems and increase savings.

The treatment variable \(T\) takes three values corresponding to the interventions offered by the bank. A quarter of the sample was assigned to a pure control group that received no visit or new product offer, which we code as \(T=0\). Half of the sample was assigned to the commitment intervention, under which a trained bank officer visited the client and offered a specially designed commitment savings account that restricted withdrawals until a self chosen target date or goal amount was reached; we code this group as \(T=1\). The remaining quarter received a marketing visit promoting the bank’s standard savings accounts but were not offered the commitment product; we code this group as \(T=2\). Thus, the treatment variable captures three distinct levels of exposure: no visit, a visit with an offer of the commitment product, and a visit with information about standard savings products only.

The outcome variable \(Y\) is the individual’s total savings balance twelve months after the intervention, aggregated across all accounts held at the partner bank. The covariates \(X\) include detailed demographic and socioeconomic information—such as age, marital status, education, and number of dependents—as well as income, occupation, and measures of prior financial behavior, including pre-existing savings, loan history, and deposit frequency. These covariates provide rich sources of heterogeneity for studying differential responses to the interventions and for learning individualized policy rules.

We restrict the sample to individuals whose income per capita falls between the 2.5th and 97.5th percentiles. The final analysis sample contains 1687 individuals: 444 in the pure control group, 785 in the commitment treatment group, and 439 in the marketing visit group. The median household monthly income is \$292, and the median income per capita is \$61.78. In our sample, 35.91% of individuals hold an active account, and 58.63% are female. For the analysis, we use the log of income per capita (measured in hundreds of dollars).

We aim to learn a policy that maps individual characteristics $x$ to treatment probabilities $\pi_\theta(t|x)$, in order to maximize the expected savings balance. We implement the three welfare estimators described in Section (ref). The estimated weights follow the stabilized weights procedure as in Section (ref), and the conditional mean functions are estimated using random forest regressions. The policy is parameterized through a softmax specification in which each treatment arm $t\in\{0,1,2\}$ is associated with a vector of coefficients $\theta_t$, with the control group serving as the reference category (i.e., $\theta_0 = 0$):

align*[align* omitted — 119 chars of source]

This formulation produces smooth assignment probabilities, yields an interpretable multinomial logit structure, and allows flexible dependence of the policy on observable covariates.

We consider two specifications for the policy covariates. In the first specification, the policy depends on two variables: the logarithm of household monthly income per capita (measured in hundreds of U.S.\ dollars, denoted \(X_1\)) and an indicator for recent account activity (\(X_2 = 1\) if the individual made any transaction in the past six months). In the second specification, we replace the activity indicator with an indicator for female (\(X_3 = 1\) if the individual is a woman), motivated by the fact that many development interventions explicitly target women with the aim of improving their economic outcomes and, in turn, household welfare.

Table (ref) reports the policy parameter estimates obtained from the three policy estimators across two policy sets. As evident from the results, the IPW estimator based on the true propensity score yields larger standard errors than the other two efficient estimators.

table[table omitted — 2,566 chars of source]

Under the first policy specification, all three estimators exhibit a clear increase in the probability of assignment to the commitment treatment as income rises, together with a higher probability of assignment to the marketing visit among lower-income households, particularly among individuals with inactive accounts. This pattern aligns with the economic intuition that commitment products are more beneficial for higher-income households with greater saving capacity, whereas marketing interventions are more effective for low-income or inactive clients. Among active users, the probability of assignment to the commitment treatment also increases with income. However, individuals with middle incomes display a relatively high probability of being assigned to the control group, suggesting that middle-income clients with active accounts may already have stronger saving habits and therefore benefit less from additional interventions.

Under the second policy specification, similar income-related patterns emerge across genders. For both genders, the probability of assignment to the commitment treatment increases monotonically with income. All three estimators tend to assign a high probability to the commitment treatment when the policy in this case. In contrast, the marketing treatment is assigned with high probability only among individuals with very low income.

Conclusion

This paper develops a semiparametric efficiency framework for policy learning that accommodates general treatment variables. The semiparametric efficiency approach reveals insights that do not appear in the existing regret-bounding literature. The simulation and empirical studies support the theoretical conclusions and demonstrate their relevance for applied work. The findings indicate that researchers should use efficient policy learning methods, such as inverse propensity weighting with an estimated propensity score or doubly robust estimation, even in experimental settings where the true propensity score is known. Although our analysis of randomized policies is motivated primarily by their favorable econometric properties, such policies may also offer practical and economic benefits in the real world.