Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
148,759 characters · 23 sections · 203 citation commands
Producing Policy Recommendations: from Statistical Decision Theory to Empirical Practice
Around the world, public institutions draw on economic research to guide real-world policy choices with empirical evidence.
In the United States, the \hyperlink{http://whitehouse.gov/cea/}{Council of Economic Advisers} is `charged with offering the President objective economic advice on the formulation of [...] economic policy [...] based [on the] analysis [of] economic research and empirical evidence'. In the European Union, `[the main objective of the] economic research activities of the \hyperlink{https://economy-finance.ec.europa.eu/economic-research-and-databases/economic-research_en?prefLang=de}{European Commission}, [...] is to support policy making [...] by developing research tools and analysing data'. In Japan, the \hyperlink{https://www.esri.cao.go.jp/en/esri/about/menu-e.html?utm_source=chatgpt.com}{Economic and Social Research Institute} mainly engages in [...] conducting empirical research related to economic and social activities [...] [to promote] evidence-based policy making.
Research institutions, academic associations, and researchers acknowledge this role.
The \hyperlink{https://www.nber.org/research}{National Bureau of Economic Research} identifies as its main objective `conducting and disseminating independent, cutting-edge, non-partisan research that advances economic knowledge and informs policy makers and the business community'. The \hyperlink{https://cepr.org/about}{Center for Economic Policy Research} claims its main objective is `to enhance the quality of economic policy-making within Europe and beyond, by fostering high quality, policy-relevant economic research'. The \hyperlink{https://www.aeaweb.org/about-aea}{American Economic Association} defines its members as `committed to the advancement of economics and its enduring contributions to society'. Indeed, around 70% of the empirical papers published in an AEA journal in the last 10 years provide at least one policy recommendation. The total number of recommendations is 517 across 471 empirical papers, 321 of which have at least one recommendation. As a result, the average empirical paper carries between one and two policy recommendations.\footnote{Author's calculations based on the full text of the universe of papers published in AER, AER:Insights, AEJ:Applied, AEJ:EconPolicy, and AER:P&P from 2015 to 2025. See Appendix (ref) for details.}
Therefore, at least in intention, academic papers in economics are motivated (and often funded) by broader normative objectives.
By the nature of many of the questions we ask across fields of economics, such objectives can be framed as the goal of understanding how to intervene on a given status quo to increase welfare. Examples include understanding whether introducing a minimum wage reduces or increases unemployment card_krueger_2000; understanding whether providing free health insurance improves access to health care services medic_aid_2016; understanding whether exposing poor families to better neighborhoods improves children’s long-run outcomes moving_to_opp.
The success of such ambitious objectives rests on the premise that economists' research design choices are aligned with them.
Providing answers to policymakers, that is, producing policy recommendations on a given set of interventions of interest, requires making design choices along at least three dimensions. (i) Planning data collections (e.g. what individual characteristics are relevant, how many individuals should be sampled, what measures should be considered, etc.), (ii) choosing an estimator to provide sufficient statistics to compute the value of each policy, (iii) defining a decision rule, that is, a function that ranks different policies according to their estimated value. It is not obvious a priori how to formalize the problem of optimally learning policy recommendations, what criterion a research design should maximize to be considered optimal or aligned with the broader normative objective, and what criterion the mainstream practice is implicitly maximizing.
The problem of making such choices becomes ever more challenging if they need to be made ex-ante, meaning when only minimal or no information about (i) the probability laws of the population of interest and (ii) the causal law of a given intervention is available. Ideally, researchers would like to commit to a choice of the three elements listed above such that they are certain that, even in the worst case admitted by those minimal assumptions, such choices will not be too far from being optimal.
For example, for a given estimator and choice rule, would it be preferable to collect data from a lab or a field experiment? At what level should the intervention be stratified, and how many units should be collected for each stratum? What is the minimum experimental sample size that would guarantee that our final recommendation is at most as far from optimal? Similar questions can be framed while keeping the collection plan fixed and searching for an optimal estimator or choice rule.
A recent literature initiated by manski_statistical_2004 has examined these problems through the lens of statistical decision theory wald1950statistical,Savage01031951. As the literature has arguably achieved theoretical maturity, this paper contributes to it by reviewing the main theoretical contributions under a common general framework, identifying the directions in which recent research is moving and questions on the frontier.\footnote{An important remark is that in this review I focus on the econometrics literature on the topic and exclude the computer science and biostatistics literatures from the scope. Nevertheless, I think it would be of great value to reconcile these different branches.} Moreover, I provide guidelines for practitioners on how to leverage such results to make ex-ante choices about research designs when the ultimate goal is to provide policy recommendations, and operational tools to communicate the performance of the resulting recommendation once the data are realized.
\paragraph{Structure of the Review.} The review is organized in two main parts. The first is theoretical. Section (ref) introduces a common decision-theoretic framework and formalizes the research design problem for policy choice. Section (ref) then reviews the main theoretical results, first for unconstrained policy spaces (Section (ref)) and then for constrained policy spaces (Section (ref)). Section (ref) concludes the theory part by discussing frontier directions, with particular attention to evidence aggregation and implementation.
The second part is applied. Section (ref) provides (i) a map (Figures (ref) and (ref)) for applied researchers to navigate the theory starting from concrete concerns when deciding among research designs (Section (ref)), and (ii) a workflow to solve the ex-post problem of communicating the performance of different recommendations once the data are realized (Section (ref)). For the latter purpose, I introduce (the beta version of) a new R package named policytargetr that practitioners can use to evaluate the performance of their recommendations before recommending them. Both objectives are illustrated in the setting of hussam_2022.
Given a state $Q \in \mathcal{Q}$, a data-collection design $\omega \in \Omega$, and design parameters $\theta \in \Theta$, let $Q_{\omega,\theta}$ denote the probability law over the sample space $\mathcal{S}$ induced by implementing $(\omega,\theta)$ under state $Q$. The observed sample therefore satisfies
An example is the collection of a panel dataset ($\omega$) that takes the cross-sectional sample size (e.g. \# individuals) and the longitudinal dimension (e.g. \# time periods) as inputs, $\theta = (n,t)$. Another example is the collection of data from a randomized controlled trial ($\omega$) that takes the share of units that receive the intervention and the sample size as inputs: $\theta = (n, p)$. Data collections are costly. Costs vary with the choice of the collection design ($\omega$), the design's parameters ($\theta$) and the probability law $Q$. Define a cost function $c(\omega,\theta,Q)$:
As an example, if $\omega$ is a survey experiment, and the intervention is costly to provide, different choices of $\theta = (n,p)$ map into different costs. Next, define an estimator $m(S,d)$ as:
where $\mathcal{D}$ denotes the set of policies. As an example, consider a panel data collection (defining $\omega$) with $n$ cross-sectional units and $t$ time periods (defining $\theta$). Let $\mathcal{M}$ denote the set of Difference in Differences estimators and $m(\cdot)$ one element of that set. Let $\mathcal{D} = \{0,1\}$ denote the $1$-dimensional choice of whether to introduce an intervention or not (therefore $\ell = 1$). Then, we estimate the counterfactual outcome using DiD and want to leverage the estimate to learn the optimal policy.
Given the estimator, denote a set of choice rules $\delta(m(S,d))$ as:
Following up on the panel data example, one simple choice for this rule is to assign the intervention if the DiD-estimated counterfactual under full adoption is higher than the one under the status quo. I report in Table (ref) other examples for each margin of choice.
In the following definition I state formally the general research design problem for policy choice.
\paragraph{Restrictions on the State Space.} Depending on the context, the researcher may have different information about $\mathcal{Q}$, commonly called the state space manski_2021. Various approaches have been proposed to leverage the knowledge available about $\mathcal{Q}$ to solve optimally the researcher's problem.
For example, a researcher may know a probability law $\pi(Q)$ that assigns a probability to each $Q \in \mathcal{Q}$ and then solve:
This is often referred to as the Bayesian approach rubin_78,DEHEJIA2005141.
A different approach is agnostic about $\pi(Q)$ and focuses on delivering uniform guarantees over the state space by minimizing the worst-case. Formally, the researcher solves:
This is known as the minimax approach, introduced as a criterion for policy choice problems by manski_statistical_2004 and first theorized in generality by wald1950statistical and Savage01031951.
Another approach derives $Q$-dependent guarantees by setting parametric constraints on $\mathcal{Q}$ via a parameter space $\beta \in \mathcal{B}$ such that $\{Q_\beta : \beta \in \mathcal{B}\} \subseteq \mathcal{Q}$ and solves:
In this review, I focus on the minimax approach, as it is the approach most widely used by the theoretical literature on policy choice following manski_statistical_2004.
\paragraph{Choice of the Loss Function.} The choice of $r$ is kept exogenous, or rather decided by the econometrician to provide theoretical results. The mainstream approach in the literature is to consider expected regret. Defining regret requires two additional definitions. First, a welfare functional
which is typically evaluated at the expected value $W_Q(d) = \mathbb{E}_Q[W(d)]$. Next, we need to define an oracle that knows $W_Q(d)$, and therefore can directly solve:
Then, we can define regret as:
where $\hat{d}$ implicitly carries forward the choice of $(\omega,\theta, m, \delta)$, since they are necessary to compute it. For expositional convenience, I use $\hat{d}$ as a shortcut for $(\hat{\omega}, \hat{\theta}, \hat{m}, \hat{\delta})$. However, it must be clear that what I really mean by $\hat{d}$ is the whole research design, and not the actual element $d \in \mathcal{D}$ (see Remark (ref) for further details). Regret measures the average welfare loss from recommending the estimated policy $\hat d$ instead of the oracle policy $d^*$, over draws of the estimating sample $S^n$. The outer expectation is taken with respect to $Q_{\omega,\theta}$, the sampling law induced jointly by the state $Q$ and the collection design $(\omega,\theta)$.
\paragraph{Types of Theoretical Guarantees.} Given a choice of $(\omega,\theta, m, \delta)$, computing a (non-trivial) closed-form expression for $\sup_{Q \in \mathcal{Q}} R_Q(\hat{d})$ under minimal assumptions on $\mathcal{Q}$ is often not possible.\footnote{The formal discussion of these minimal assumptions is deferred. They often entail a uniform bound on the outcome, conditional independence of the intervention assignment, and strict overlap kitagawa_who_2018.} As a result, it is mainstream theoretical practice to derive upper and lower bounds for worst-case regret:
Deriving upper bounds can be easy. The theoretical crux comes with proving that the derived upper bound is sharp, meaning that it is attained with equality for an admissible DGP. For a given research design $\hat d$, this can be shown by finding a specific $Q' \in \mathcal{Q}$ such that
Sharpness of the upper bound identifies the exact worst-case regret of that research design. To additionally establish minimax optimality, one needs to show that no alternative research design can attain a lower worst-case regret. In particular, letting $\hat d^*$ denote a minimizer of $\bar{\mu}(\hat d)$, it is sufficient to establish the minimax lower bound
Together with the upper bound for $\hat d^*$, this implies that $\hat d^*$ is minimax-optimal. Such lower bounds can often be established by constructing specific DGPs, or families of DGPs, for which no alternative algorithm $(m',\delta')$, given $(\Omega,\Theta)$, can achieve lower worst-case regret.
Unfortunately, in many cases, it is not possible to establish minimax optimality exactly. This could be because the upper bound is not sharp, or because a matching minimax lower bound cannot be derived. The second-best result, then, is to prove that the rate of the upper bound is minimax-sharp. In particular, for some sequence $a_n$, suppose that
Then no alternative research design can improve on the rate $a_n$, and $\hat d^*$ is minimax rate-optimal. One way to achieve this result is to prove that the upper and minimax lower bounds coincide up to constants.
Another result that can be stated when both upper and lower bounds are available is minimax dominance. In particular, if the lower bound for a choice $\hat{d}'$ is higher than the upper bound for a choice $\hat{d}''$, then it must be that $\hat{d}''$ is minimax-dominant compared to $\hat{d}'$. Formally,
This is a sufficient condition for $\hat d''$ to be preferred to $\hat d'$ according to the minimax criterion.
In what follows, I first specialize the general framework proposed in Section (ref) to the potential outcome framework, as this is the causal language considered in the literature, and then review the main theoretical contributions.
Let $\mathcal{T}$ denote a finite set of interventions and let $\mathcal{X}$ denote the $s$-dimensional space of observed covariates. Then, let $d : \mathcal{X} \to \mathcal{T}$ denote the assignment rule. Denote the set of assignment rules considered by $\mathcal{D}$. A state $Q \in \mathcal{Q}$ specifies the joint distribution of $(T_i, X_i,\{Y_i(t)\}_{t \in \mathcal{T}})$. For a given design $\theta$, the data collection process $(\omega,\theta)$ generates a sample
where $T_i$ denotes intervention assignment and $Y_i = Y_i(T_i)$ the observed outcome.
manski_statistical_2004 was the first to apply statistical decision theory Savage01031951 to problems of policy choice. He studies the problem of targeting interventions across subgroups of the covariate space when experimental data from the target population can be collected. He provides worst-case regret guarantees for conditional empirical success (CES) rules, namely rules that target interventions to those subgroups that attain a higher empirical mean outcome when receiving the intervention. The definition of subgroups plays a central role in the paper. These are defined through a coarsening $v:\mathcal{X} \to \mathcal{V}$, which aggregates covariate values into subgroups. A first conceptual point in manski_statistical_2004 is that, although an oracle policymaker would always prefer the finest coarsening to get as close as possible to individual variation in intervention effects, a statistical policymaker should choose the coarsening carefully to balance how well the assignment rule fits individual variation and how hard it is to learn from data. This point highlights that standard bias-variance tradeoffs in statistical learning can be studied analytically via worst-case regret in policy learning. A second conceptual point is that data collection interacts with this tradeoff. Referring to the general formulation in Definition (ref), manski_statistical_2004 solves the joint problem of deciding subgroup definition (whether through the coarsening $v$ or the entire covariate space) and sample sizes ($\theta$ in Def. (ref)'s notation) to attain the smallest worst-case regret possible. The key result is that there exist finite sample sizes such that it is minimax dominant (in the sense of Eq. (ref)) to consider the finest coarsening rule.
manski_statistical_2004 constrains the state space $\mathcal{Q}$ with the following assumptions.
Assumption BO (Bounded Outcomes) states that the outcome variable lies in a known interval of length $M$. Assumption CS (Covariate Space) states that there are finitely many covariate values and that the probability distribution of covariates is known. Assumption AM (Assignment Mechanism) states that the intervention is binary (status quo and innovation) and the researcher can conduct a randomized experiment, either with stratified randomization, or with complete randomization, to estimate the optimal assignment rule. Denote the set of states $Q$ that satisfy Assumption (ref) by $\mathcal{Q}_\mathrm{MK}$.
Let $\mathcal{D}_v=\{d_v:\mathcal{V}\to\mathcal{T}\}$ denote the set of assignment rules measurable with respect to the coarsening $v$. Using the notation of Section (ref), the set of interventions is binary $\mathcal{T}=\{0,1\}$ and we want to choose an assignment rule $d_v$ so that $W_Q(d)=\mathbb{E}_Q[Y_i(d_v)]$ is maximized. The corresponding oracle rule among assignment rules measurable with respect to $v$ is then:
with value
Define
where
For each intervention $t \in \{0,1\}$ and coarse cell $\nu \in \mathcal{V}$, define the sample analog
Then, for a given coarsening $v$ and policy $d_v \in \mathcal{D}_v$, define the estimated welfare
The CES rule based on $v$ can then be written as
manski_statistical_2004 compares $\hat d_v$ to the oracle rule that conditions on the full covariate vector $X_i$, namely
with value
Instrumental to the comparison with such an oracle, manski_statistical_2004 defines regret as:
that is, regret relative to the oracle that conditions on the full covariate vector $X_i$.
Before stating the main result, let for each $\nu \in \mathcal{V}$,
that is, the set of fine covariate values pooled into coarse cell $\nu$. Define
That is, $\rho(\nu)$ is the largest probability mass that can be placed on the smaller side of a partition of the fine covariate values pooled into cell $\nu$.
The first term bounds the difference between the oracles $d^*_x$ (Eq. (ref)) and $d^*_v$ (Eq. (ref)) and can be defined as approximation error. The second term bounds the difference between the estimated policy $\hat{d}_v$ and the oracle $d^*_v$ and can be defined as estimation error. This terminology was introduced later by mbakop_model_2021 building an analogy with model selection problems. Note that, for infinite sample sizes, the estimation error is zero, and for the finest coarsening function (i.e. the identity function), the approximation error is zero.\footnote{The second point follows from the fact that when $v$ is the identity, $\mathcal{X}_\nu=x$ and $\mathbb{P}_Q(X_i = x | X_i = x) = 1$. Therefore, $\rho(\nu) = 0$.}
The approximation error term in Proposition (ref) is sharp. As a result, it serves as a lower bound on maximum regret for $\hat{d}_v$: even if we had infinite data, we would incur that welfare loss. Moreover, the estimation error term can also be derived for rules that learn from the finest coarsening, that is, the covariate space itself $\hat{d}_x$. As a result, one can use these two to understand under which conditions on $\theta = (N_{1x},N_{0x})$ it is minimax dominant (see Eq. (ref) for a definition) to learn $\hat{d}_x$ rather than $\hat{d}_v$.
Proposition (ref) provides a sufficient condition for $\theta=(N_{1,x},N_{0,x})$ to make it minimax dominant to learn policies in the finest coarsening, i.e. using the full covariate space.
manski_statistical_2004 opened up a new thread of contributions that studied further the properties of CES rules and the policy choice problem more generally.
hirano_asymptotics_2009 formalized the policy choice problem from a local asymptotic perspective using the limit of experiments framework lecam1986. Relative to manski_statistical_2004, the main novelty is that optimality is not studied uniformly over a fixed state space, but locally around knife-edge states where the policymaker is nearly indifferent between implementing and not implementing the intervention.
For simplicity, fix a covariate value and suppress it from the notation, as in hirano_asymptotics_2009. Consider deciding whether or not to implement one binary intervention under full adoption, so $\mathcal{D}=\{0,1\}$. Define the welfare contrast
and the corresponding oracle
In the notation of Definition (ref), hirano_asymptotics_2009 leaves $(\omega,\theta,m)$ abstract, except for requiring that $m(S^n,1)$ be an estimator of $\tau_Q$. Writing $\hat{\tau}_n:=m(S^n,1)$, they consider local sequences of DGPs $\{Q_{n,h}\}_{h\in\mathbb{R}}\subseteq\mathcal{Q}$ around a reference state $Q_0$ such that $\tau_{Q_0}=0$ and
Hence, $h$ indexes local departures from the knife-edge state in units of $\sqrt{n}$-scaled welfare contrast. The asymptotic approximation is then
under $Q_{n,h}$. Therefore, we can restrict the set of feasible $\Omega \times \Theta \times \mathcal{M}$ to the one that guarantees this condition to hold.
hirano_asymptotics_2009 considers a general loss function that can be asymmetric tetenov_asymmetric_2012 and evaluated at the mean (see Eq. (ref)), rather than at the worst case (see Eq. (ref)). In this review we consider the case of symmetric regret evaluated at the worst case. Then, we can write
which coincides with regret, pointwise in the fixed covariate value. The main result is that, in the Gaussian limit experiment lecam1986, the minimax rule under symmetric loss is the zero-cutoff rule. Mapping it back to the original sample problem yields the plug-in choice rule
Proposition (ref) states that, once the welfare contrast can be estimated efficiently, no alternative sequence of data-dependent choice rules can uniformly improve on the simple sign rule in a local neighborhood of a knife-edge DGP. Unlike Eq. (ref), this is a local asymptotic criterion: it compares rules only along sequences of DGPs that approach the knife-edge state at the $1/\sqrt{n}$ rate.
\paragraph{Exact Finite Sample Behaviour and No-Data Rules.} A complementary contribution by stoye_minimax_2009 studies the same policy choice problem from the opposite angle: instead of passing to a limit experiment, it derives exact finite-sample minimax regret rules. For the binary policy choice problem without covariates, binary outcomes $Y_i(t)\in\{0,1\}$, and random assignment in the estimating sample, let
with the convention that $N_t(\bar Y_t-1/2)=0$ when $N_t=0$. Then, setting $m(S^n,1)=(N_0,\bar Y_0,N_1,\bar Y_1)$, the exact minimax rule assigns intervention $1$ with probability
Under matched pairs (i.e. $N_0 = N_1 = n/2$), this collapses to the empirical success rule with symmetric tie-breaking, while if the status quo is known and only the innovation is sampled, the exact minimax rule becomes a threshold rule in the number of observed successes.
The second main result is about covariates. If the state space imposes no cross-covariate restrictions, so that intervention effects can vary arbitrarily across cells of $\mathcal{X}$, then the global finite-sample minimax problem separates into one conditional minimax problem for each covariate value. In the notation of this review, the minimax rule takes the form $\hat d(x)=\delta_x^*\{m(S_x^n,1)\}$, where $S_x^n$ denotes the subsample with $X_i=x$ and each $\delta_x^*$ is minimax for the corresponding conditional problem. This result sharpens manski_statistical_2004: absent restrictions on the intervention effect function over covariates, pooling information across covariate values (e.g. via the coarsening map $v(x)$ defined above) is never minimax optimal. The most striking case is when $X$ is continuously distributed. Then, for any fixed target value $x$, a finite sample contains no repeated observations with exactly that covariate value with probability one, so the conditional subsample $S_x^n$ is generically empty. If the outcomes under both the status quo and the intervention are unknown, the exact minimax rule assigns the intervention with probability $1/2$ for every $x$, independently of the realized data. If instead the status quo welfare is known, the minimax rule is still a no-data rule, but it assigns intervention $1$ with probability $1-W_Q(0)$, where $W_Q(0)$ is conditional on $x$, at each covariate value. The intuition is that, without smoothness restrictions on the potential outcomes function, observations with similar covariate values cannot be used to learn about the welfare contrast at $x$ uniformly over the state space. Hence, with continuous covariates and a sufficiently rich $\mathcal{Q}$, finite-sample minimax regret may force the researcher to ignore the sample altogether, even though such rules can be pointwise dominated once stronger structure is imposed. stoye_covariates_2012 refines this result and shows that if the state space is restricted so that covariates have bounded influence on welfare, for example
then there exists a strictly positive region in which the exact minimax rule pools information across covariate values. Thus, pooling information across covariate values can be minimax optimal when their importance is bounded.
\paragraph{Shrinkage for Policy Choice.} The value of using covariates is revisited again by Ishihara_2025, who focus on the intermediate region between complete pooling and cell-by-cell empirical success rules. Working with finitely many covariate cells, Gaussian approximations for cell-specific welfare-contrast estimators, and assuming $|\tau_Q(x)-\bar{\tau}_Q|\leq \kappa$ (or equivalently that the intervention effect function has a bounded second moment), they propose shrinkage rules that replace each raw estimate with a convex combination of the cell-specific estimate and the pooled estimate. The shrinkage parameter is chosen by minimizing a tractable upper bound on maximum regret, so the resulting rule interpolates between CES and full pooling. The main result says that when heterogeneity in intervention effects across covariate values is neither negligible nor unbounded, shrinkage can strictly improve on both extremes in worst-case regret terms.
A related use of shrinkage appears in moon2026, who studies a different policy problem in which the planner chooses local changes across many policies rather than a binary assignment rule for one policy. For each policy $j$, the planner faces noisy estimates of benefit and net cost and wishes to choose a vector of small policy changes that maximizes expected welfare. The oracle rule depends on posterior mean benefit and posterior mean net cost for each policy. moon2026 proposes an empirical Bayes approach that shrinks policy-specific estimates toward a common distribution learned from the full collection of policies and shows that this approach can approximate the oracle allocation, while a raw plug-in rule may fail. Relative to Ishihara_2025, the paper therefore applies the same broad shrinkage logic to a different dimension of the problem: pooling information across policies rather than across covariate cells.
The same broad lesson appears in yamin2026, who studies poverty targeting when the policymaker must allocate nonnegative transfers across households subject to a fixed budget. The planner does not observe true poverty gaps, but only noisy estimates of them, so the policy problem is how to translate noisy signals into a feasible transfer allocation. The paper shows that the plug-in rule that treats estimated poverty gaps as true is inadmissible on a compact parameter space. This implies there exists some feasible rule that weakly lowers regret uniformly in the state space and strictly lowers it in some cases. They propose a nonparametric empirical Bayes rule that replaces each poverty estimate with its posterior mean and then applies the same constrained allocation logic as under full information. Finally, they characterize the oracle Bayes allocation and derive regret guarantees relative to that oracle. This procedure is shown to have sizeable gains relative to the plug-in rule in simulations. Relative to moon2026, the shrinkage now pools information across households rather than across policies, but the message is similar: posterior shrinkage can improve downstream policy allocation when the underlying signals are noisy.
\paragraph{Asymmetric Regret.} tetenov_asymmetric_2012 extends manski_statistical_2004 to allow for planners who are differently averse to type I and type II errors. Keep the one-intervention, full-adoption problem with known status quo welfare $W_Q(0)$, innovation welfare $W_Q(1)$ learned from a finite sample, and welfare contrast $\tau_Q:=W_Q(1)-W_Q(0)$. Let $\hat\tau_n:=m(S^n,1)$ denote an estimator of $\tau_Q$. For a threshold rule $\hat d_T:=\mathbf{1}\{\hat\tau_n>T\}$, regret takes the statewise form
The first line is the welfare loss from adopting an inferior innovation, while the second is the welfare loss from rejecting a superior one. This motivates defining, for a generic data-dependent rule $\hat d$,
Therefore $\bar R^{\mathrm{I}}(\hat d)$ is maximum Type I regret and $\bar R^{\mathrm{II}}(\hat d)$ is maximum Type II regret. tetenov_asymmetric_2012 studies the weighted minimax problem
with $K>0$. The symmetric minimax criterion is the special case $K=1$, which recovers the cases studied in manski_statistical_2004, stoye_minimax_2009, stoye_covariates_2012.
tetenov_asymmetric_2012 shows that in the Gaussian experiment where $\hat\tau_n\sim\mathcal{N}(\tau_Q,\sigma^2)$, it is sufficient to search for the minimax choice rule among threshold rules of the form $\hat d_T=\mathbf{1}\{\hat\tau_n>T\}$. Hence, although the original choice variable is the decision rule $\delta\in\Delta$, after this reduction the problem can be re-written as the choice of a scalar threshold $T\in\mathbb{R}$. Using the previous display and writing $h:=\tau_Q/\sigma$, their maximum Type I and Type II regrets can be written as
Therefore the asymmetric minimax rule is equivalently characterized by the threshold choice
where $T_K$ denotes the unique minimizer in the normalized problem with $\sigma=1$. Equivalently, $T_K$ is the unique normalized cutoff satisfying
The corresponding threshold in the original Gaussian problem is $T=\sigma T_K$. Hence the optimal threshold moves to the right with $K$: compared to symmetric minimax, the planner requires stronger evidence before recommending the innovation. For Bernoulli outcomes with known status quo mean $p_0$, the exact finite-sample solution has the same structure: it is a threshold rule in the number of observed successes, defined by the unique cutoff that equalizes weighted maximum Type I and Type II regret. Its large-sample approximation is
which collapses to the plug-in rule when $K=1$. An especially insightful interpretation is that one-sided hypothesis-test rules correspond to asymmetric minimax rules for particular values of $K$: in the normal model, a level $\alpha=0.05$ rule is equivalent to $K\approx 102$, while $\alpha=0.01$ corresponds to $K\approx 970$. This makes precise how strongly conventional testing procedures privilege avoiding mistaken adoption relative to missing a beneficial innovation.
\paragraph{Non-linear Regret.} kitagawa_26_bio consider the policy choice problem of manski_statistical_2004 under non-linear regret. For a possibly fractional rule $\hat d:=\delta\{m(S^n,1)\}\in[0,1]$, they evaluate choice rules through the nonlinear regret risk
where $g:\mathbb{R}_+\to\mathbb{R}_+$ is increasing and nonlinear. Their benchmark case is mean-square regret, $g(r)=r^2$, for which
so the planner penalizes not only average regret but also the volatility of regret induced by sampling uncertainty. This change has a first-order implication: singleton rules are no longer essentially complete (in the sense that restricting the policy space to such rules is without loss), so optimal rules are generally fractional. In the Gaussian limit experiment hirano_asymptotics_2009 where $\hat\tau_n:=m(S^n,1)\sim\mathcal{N}(\tau_Q,\sigma^2)$, the finite-sample minimax mean-square regret rule takes the simple logistic form
and the same structure survives in regular parametric models as a logistic transform of the $t$-statistic.
\paragraph{Certified Decisions.} andrews_chen_2025 study how inference can be turned into decision recommendations with explicit guarantees. In the notation of this review, suppose the researcher constructs a confidence set $\hat{\mathcal{Q}}_n\subseteq\mathcal{Q}$ for the unknown state $Q$. They then consider as-if decisions of the form
where $d_{Q'}^*:=\arg\max_{d\in\mathcal{D}}W_{Q'}(d)$. Thus $\hat r_n$ is a high-probability upper bound on the regret of the recommended decision. If $\hat{\mathcal{Q}}_n$ covers the true state with probability $1-\alpha$, then
Their main result is that, among rules paired with such high-probability regret bounds, there is essentially no loss in restricting attention to these confidence-set-based decisions. The paper therefore gives a general foundation for using inferential objects to support ambiguity-averse downstream decision-makers.
\paragraph{Policy Learning with Confidence.} chernozhukov_lee_rosen_sun_2026 provide a concrete application of this logic to policy choice when, for each feasible rule $d\in\mathcal{D}$, the researcher has an estimate $\hat W_n(d)$ of welfare together with a standard error $\hat s_n(d)$. Rather than choosing the rule with the largest estimated welfare, they consider risk-aware rules that trade off performance and estimation risk and therefore move along an efficient frontier of feasible welfare-risk pairs. Their proposed PoLeCe rule chooses
where $\hat c_{1-\alpha}$ is calibrated so that the objective is a one-sided lower confidence bound for welfare. Hence the rule selects the policy with the highest guaranteed welfare at confidence level $1-\alpha$, rather than the policy with the highest point estimate. Relative to the minimax contributions above, the paper does not modify the loss function directly; instead it changes the decision criterion so that sampling uncertainty enters policy choice through a reporting guarantee.
The conceptual distinction between making decisions under uncertainty, i.e. the problem of solving a population-wide problem with data that provide point identification, and making decisions under ambiguity, i.e. solving a population problem where data only provide partial identification in the sense of manski2003partial, was first noted by MANSKI2000415. Up to this point, we focused on the first kind of problem, but a thread following MANSKI2000415, manski_statistical_2004 has considered settings where both sources of difficulty arise.
\paragraph{Univariate Partial Identification.} A first contribution in this direction is stoye_covariates_2012, who studies the binary full-adoption problem when the experiment does not point-identify the target-population welfare contrast $\tau_Q:=W_Q(1)-W_Q(0)$, but only an estimating-population contrast, say $\tau_Q^{S}$, linked to $\tau_Q$ by
with $a\in(0,1]$ and $b>0$ summarizing the severity of the failure of internal or external validity. This representation nests, among others, selective noncompliance and selective sampling. In the Gaussian experiment, letting $\hat\tau_n:=m(S^n,1)\sim \mathcal{N}(\tau_Q^{S},\sigma^2)$, stoye_covariates_2012 shows that the minimax regret rule is
In the middle case, $\Phi(x;0,v)$ denotes the c.d.f. of a centered Gaussian variable with variance $v$ evaluated at $x$, so the rule assigns the intervention with probability $\Phi(\hat\tau_n;0,(2/\pi)(b/a)^2-\sigma^2)$. Equivalently, the planner adds a centered Gaussian noise with variance $(2/\pi)(b/a)^2-\sigma^2$ to the experimental signal and then applies the sign rule to the perturbed signal. The condition $(\pi/2)^{1/2}\sigma a<b<1$ indexes an intermediate region: ambiguity is already too large for the pure plug-in rule $\mathbf{1}\{\hat\tau_n>0\}$ to remain minimax-optimal, but not yet so large that it is optimal to ignore the data altogether and set $\hat d_{\mathrm{ST}}=1/2$. Hence, when ambiguity is small relative to sampling precision, the planner should follow the sign of the experimental signal exactly as in the point-identified benchmark; once the wedge between $\tau_Q$ and $\tau_Q^{S}$ becomes large enough, minimax regret optimally attenuates the signal through randomization, and for sufficiently severe validity failures collapses to the no-data rule.
\paragraph{General Partial Identification.} yata_2025 shows how to solve finite-sample minimax regret problems in a more general partially identified environment. The researcher observes a Gaussian statistic $Z^n\sim \mathcal{N}(g(\vartheta),\sigma^2 I_n)$, where $\vartheta$ may be high-dimensional or infinite-dimensional, $g:\mathcal{H}\to\mathbb{R}^n$ is linear, and $\mathcal{H}$ is a nonempty, convex, and centrosymmetric parameter space. Welfare is summarized by the linear contrast $\Lambda(\vartheta)=W_1(\vartheta)-W_0(\vartheta)$. For a fixed reduced-form value $\mu$, the identified set
collects all welfare contrasts that remain feasible once the observables are fixed and identify $\mu$. Hence partial identification arises whenever $I(\mu)$ contains both positive and negative values: the data identify the reduced form, but not whether introducing the intervention maximizes welfare. This setup nests the case considered by stoye_covariates_2012, as well as other partial-identification problems that are multivariate in nature, such as evidence aggregation for policy choice ishihara_evidence_2026.
The first-order contribution of the paper is the theory used to solve this problem exactly in finite samples. For each centrosymmetric line $[-\bar{\vartheta},\bar{\vartheta}]\subseteq\mathcal{H}$, Yata first shows that the minimax rule is the sign rule $\mathbf{1}\{g(\bar{\vartheta})'Z^n\geq 0\}$ when the sample is informative on that line, and a $1/2$-randomization rule when it is not. Then, he defines
which measures the largest welfare contrast compatible with signal strength $\epsilon$. The hardest one-dimensional subproblem is then indexed by
Thus, Nature chooses the signal strength that maximizes welfare contrast times error probability. Let $\vartheta_{\epsilon^*}$ attain the modulus at $\epsilon^*$, and $w^*$ denote the local direction along which the upper bound $\bar I(\mu):=\sup I(\mu)$ expands fastest at $\mu=0$. Then, yata_2025 shows that a minimax regret rule depends only on one scalar least-favorable index:
where $\phi$ and $\Phi$ denote the standard normal p.d.f. and c.d.f., respectively.\footnote{There is a knife-edge threshold rule when $2\phi(0)\psi(0)/\psi'(0)=\sigma$, and the minimax risk is $R(\mathcal{H})=\psi(\epsilon^*)\Phi(-\epsilon^*/\sigma)$. } The key intuition is simple: once the least favorable one-dimensional subproblem is found, the original multivariate problem collapses to deciding on the sign of a single index, possibly after adding noise when local ambiguity is too severe.
\paragraph{Local Asymptotic Partial Identification.} A complementary route is taken by christensen_payoffs_2025, who study a closely related binary choice problem under partial identification, but from the local asymptotic perspective of hirano_asymptotics_2009. The researcher observes data $X^n\sim P_{n,\mu}$ informative about a point-identified reduced-form parameter $\mu\in\mathcal{M}$, while policy payoffs depend on a structural object $\vartheta\in\Theta_0(\mu)$ that remains only set identified conditional on $\mu$. For each action $d\in\mathcal{D}$, they define the conditional worst-case risk
and the corresponding oracle
which is infeasible because $\mu$ must be estimated. Localizing around a reference value $\mu_0$ through $\mu=\mu_0+h/\sqrt{n}$, they rank sequences of rules $(\delta_n)_{n\geq 1}$ by the integrated local excess risk
If $\pi_n(\cdot)=\pi(\cdot\mid X^n)$ denotes a posterior, or quasi-posterior, for $\mu$, the main result is that the feasible rule
is asymptotically optimal under this criterion, and so is any bootstrap or quasi-Bayes implementation that is asymptotically equivalent to it. The key technical point is that partial identification typically makes $R(d,\mu)$ only directionally differentiable in $\mu$, so the usual plug-in rule $\delta^o(\hat\mu_n)$ need not be asymptotically equivalent to $\hat d_{\mathrm{CMS}}$ and may therefore be suboptimal. Hence, relative to yata_2025, the paper does not solve the finite-sample minimax regret problem directly; rather, it extends the Hirano-Porter local asymptotic framework to environments in which ambiguity is first profiled through $\Theta_0(\mu)$ and then integrated over local uncertainty in $\mu$.
\paragraph{Multiplicity and Least Randomization.} Olea_2026 revisit a closely related class of partially identified Gaussian problems from a different angle. Rather than constructing a minimax regret rule, their focus is on the structure of the minimax regret solution set. Their main message is that minimax regret may be highly non-unique: in general, there exist many minimax-regret-optimal rules, and fractional rules arise as a general feature of the minimax regret criterion itself. To recover uniqueness of the minimax rule, they propose a refinement based on least randomization, characterizing the minimax regret rule that randomizes on the smallest set of data realizations. In this sense, the paper is complementary to yata_2025: Yata proves existence and derives a formula for a minimax regret rule based on a scalar least-favorable index, while Olea_2026 show that, once the problem is in the randomized region, such rules need not be unique and may be ranked further using least-randomization as a criterion.
\paragraph{External Validity as Distributional Robustness.} A conceptually distinct contribution is adjaho_external_2022, who study policy choice when the target population may differ from the experimental one. Let $P$ denote the experimental distribution of $(X,Y_0,Y_1)$ and let $d:\mathcal{X}\to\{0,1\}$ be a policy rule. Instead of maximizing experimental-population welfare
they propose the robust welfare criterion
where $W(P,Q)$ is a Wasserstein distance and $\epsilon$ has the interpretable meaning of the largest admissible shift in the average intervention effect between the experimental and target populations. Their sharpest result concerns shifts in potential outcomes only, holding the covariate distribution fixed: in that case,
when outcomes are unbounded below, up to the obvious lower-support truncation when they are bounded. Hence the ordering of rules is unchanged, so any policy rule that is optimal, or has small regret, in the experimental population remains optimal, or has the same regret guarantee, under this notion of external validity. A related extension allows the covariate distribution to shift from density $p$ in the estimating population to density $q$ in the target population. Writing $\rho(x):=q(x)/p(x)$, the relevant benchmark becomes the reweighted welfare $\mathbb{E}_P[(Y_1d(X)+Y_0(1-d(X)))\rho(X)]$, and the same invariance result holds: once outcome drift is modeled through a Wasserstein neighborhood, the robust counterpart preserves the ranking of rules up to the same $\epsilon$ penalty. If instead both covariates and potential outcomes are allowed to shift adversarially in unknown ways, robust welfare depends on the distance of each covariate value from the intervention status dictated by the rule.
A closely related contribution is kido_distributionally_2022, who studies policy choice when the source and target populations may differ not only in covariate composition but also in the conditional distribution of potential outcomes. Relative to adjaho_external_2022, the paper adopts a different Wasserstein ambiguity set, centered pointwise on the source conditional distribution and combined with knowledge of the target covariate distribution. This leads to a more constructive policy-learning message: robustness can be implemented by evaluating each rule under a systematically pessimistic version of the estimating population payoffs. Under suitable conditions, the resulting distributionally robust rule coincides with the naive reweighted rule, but in general the ranking of policies can change, so external-validity concerns may alter the optimal policy rather than simply add a common penalty adjaho_external_2022. The paper also derives regret guarantees for the estimated robust rule.
\paragraph{Geometric Control under Donor-Target Mismatch.} opocher_geom_2026 studies a setting in which the estimating sample comes from an innovated donor population, while the policy must be applied to a distinct target population. Relative to the previous literature on external validity and partial identification, the paper fixes a threshold decision rule and focuses on the choice of estimator $m$. It introduces certification: an estimator yields certified decisions if, whenever $|\tau(x)|$ is sufficiently large, the probability of recommending the wrong sign is uniformly controlled. The paper shows that certification implies a bound on worst-case compensation loss. Restricting attention to matching estimators with positive weights, it then derives a certification radius composed of a geometric approximation term and stochastic terms, and shows that in a large-sample regime the problem becomes purely geometric. Within this class, the Delaunay interpolant delivers the best asymptotic affordability guarantee. This geometric perspective allows one to identify the exact point in the covariate space that maximizes the worst-case loss and therefore guides data collection processes capable of bringing the loss down to a target level in as few collection steps as possible.
kato_adaptive_2025 shifts attention from the terminal policy rule to the experimental design itself. He studies the binary policy choice problem under full adoption when the experimenter can adaptively update intervention assignment probabilities during data collection, rather than fixing them ex-ante. In the notation of Definition (ref), this is a problem where the main design object is the data-collection process $\omega$, while the terminal rule $\delta$ remains the simple recommendation of the intervention with the highest estimated mean outcome.
Formally, the experiment lasts for $T$ rounds. At each round $t$, the experimenter chooses an intervention $D_t\in\{0,1\}$ based on past observations $\{(D_s,Y_s)\}_{s=1}^{t-1}$ and then, after the allocation phase, recommends
where $\mu_d:=\mathbb{E}[Y_d]$. Performance is evaluated by simple regret,
that is, the welfare loss from recommending the wrong intervention after the adaptive experiment.
The proposed design is a two-stage Neyman allocation. The first stage allocates both interventions uniformly to estimate their standard deviations. The second stage allocates intervention $d$ proportionally to the estimated standard deviation $\sigma_d$. At the end of the experiment, the experimenter recommends the intervention with the larger sample mean. The main result is that, over a class of mean-parameterized canonical exponential families, this adaptive design is both minimax and Bayes optimal in the sense that its regret upper bounds match the corresponding lower bounds exactly. In particular, if $\bar{\sigma}_d:=\sup_{\mu\in\mathcal{M}}\sigma_d(\mu)$, then
and kato_adaptive_2025 proves a matching minimax lower bound. Thus, relative to manski_statistical_2004, the paper endogenizes sequential data collection and shows that adaptive Neyman allocation is first-best when the researcher's goal is an ex-post assignment rule rather than average effect estimation.
A second wave of research started from the observation that restrictions on the complexity of the policy space can be leveraged together with restrictions on the state space to control worst-case regret.
This fact was first noted by kitagawa_who_2018. They introduce Empirical Welfare Maximization (EWM), a pair $(m, \delta)$ that fixes $m$ to the empirical average of welfare produced by an intervention and sets $\delta$ to select the maximum across the policy space. The main theoretical contribution lies in proving that one can leverage (i) a constraint on the complexity of the policy space and (ii) a within policy space definition of regret to derive rate-sharp regret bounds. Moreover, EWM is shown to be minimax-rate optimal (see Eq. (ref) for a definition). The mechanics of the main results build on well-known results in statistical learning van_der_vaart_weak_2023. This powerful connection renewed the interest in the policy learning problem and empowered a sizeable second wave of theoretical research that adapted the general results in specific settings.
The metric of complexity considered is the VC-dimension van_der_vaart_weak_2023. In particular, kitagawa_who_2018 consider collections of assignment rules $\mathcal{D}_v = \{d : \mathcal{X} \to \{0,1\} : \mathrm{VC}(d)=v\}$. In general, the VC dimension measures the number of points a function can shatter. In the context of assignment rules (i.e. functions that determine who should be assigned to an intervention), the VC dimension counts how many distinct points of $\mathcal{X}$ can be labeled as either assigned or not while respecting the rule. Intuitively, this measure bounds how flexible or complex an assignment rule can be, and therefore how hard the statistical problem of finding the true optimum within that space is.
By within policy class regret we mean that the oracle is defined as:
rather than
Therefore, the performance of any choice of $(m, \delta)$ is evaluated relative to an oracle that is imposed to make choices just as complex as the policy space. As a result, the first-best policy choice $d^{\mathrm{FB}}(x) = \mathbf{1}\{\tau(x)>0\}$ may not be available to the oracle, unless $d^{\mathrm{FB}}\in \mathcal{D}_v$. If that is not the case, the oracle would instead need to choose the best approximation of $d^\mathrm{FB}$ given the policy class.
kitagawa_who_2018 place the following restrictions on the state space.
Assumption AM (Assignment Mechanism) characterizes a quasi-experimental environment in which intervention assignment is independent of potential outcomes conditional on covariates, and the potential outcome of each unit $i$ depends only on their own intervention status. Assumption BO (Bounded Outcomes) implies that both potential outcomes, and thus intervention effects, are uniformly bounded in absolute value by $M/2$. Boundedness is a standard condition in the statistical learning literature as it enables the use of uniform concentration inequalities hoeffding_probability_1963,van_der_vaart_weak_2023. Assumption SO (Strict Overlap) is standard in the causal inference literature and guarantees that all units have a positive probability of receiving the intervention or not. It holds by design in randomized controlled trials, while it can be violated in observational studies. Assumption VC (VC Class) assumes that the policy class cannot shatter infinitely many points or, more intuitively, it cannot be infinitely complex. Define $\mathcal{Q}_\mathrm{KT}$ as the state space that satisfies Assumption (ref).
kitagawa_who_2018 study the properties of EWM. This is defined as:
Notice that EWM can be easily rewritten in the notation of Definition (ref) by setting $\theta= n$, $\Omega$ as the set of designs that respect Assumption (ref).AM, $m(S^n, d)= \hat{W}_n(d)$, and $\delta(\cdot) = \operatorname*{arg\,max}_{d \in \mathcal{D}_v}\{\cdot\}$. If the researcher is collecting data from an experiment she is designing, $\kappa$, the propensity score bound, also enters as a margin of choice: $\theta = (n,\kappa)$.
Proposition (ref) can be summarized into two main points. First, the worst-case regret of $\hat{d}_\mathrm{EWM}$ is bounded by (i) a constant that depends on the choice of $n$ and $v$ (and possibly $\kappa$), (ii) state space-dependent constants $M$ and $\kappa$ (respectively, the upper bound on the outcome and on the propensity score), and (iii) a universal constant $C_1$. Importantly, the rate depends on the ratio between the complexity of the learning problem and the sample size available to solve it. Second, the regret upper bound is rate-sharp, as there exists a DGP $Q'\in \mathcal{Q}_\mathrm{KT}$ such that no algorithm could ever achieve a faster rate than $\sqrt{v/n}$. As a consequence, given $\mathcal{Q}_\mathrm{KT}$, the rate of decay of worst-case regret when choosing $\hat{d}_{\mathrm{EWM}}$ cannot be uniformly improved by any other choice of $(m,\delta)$, keeping fixed $\omega$ and $\theta$ (see Def. (ref) for notation).
We can group different threads of research that followed EWM depending on the design object taken into consideration.
A first thread of literature has focused on alternative assignment mechanisms and endogenous selection issues. \paragraph{Policy Learning with Observational Data.} athey_policy_2021 asks whether similar regret guarantees can still be obtained when policy learning is based on observational data where propensity scores are unknown, and intervention assignment may be endogenous. The main contribution is showing that this is possible whenever the welfare of assigning interventions can be estimated through semiparametrically efficient doubly robust scores. Let $Z_i$ denote the observables used to identify the intervention effect, with $Z_i=T_i$ under selection on observables and $Z_i$ a valid instrument under endogenous assignment.
Assumption ID (Identification) is the key departure from Assumption (ref): instead of requiring a known randomized assignment rule, the paper defines a state space through semiparametric identification conditions that allow the intervention effect to be recovered from observational data. Assumption NE (Nuisance Estimation) requires the researcher to estimate first-step objects, such as outcome regressions, propensity scores, or compliance scores, accurately enough for the final score to be asymptotically well behaved. Assumption VC (VC Class) plays the same role as in kitagawa_who_2018, namely controlling the complexity of the policy space. Finally, Assumption BW (Bounded Weights) generalizes strict overlap: under selection on observables it reduces to bounded inverse propensity weights, while under instrumental variables it rules out weak instruments. Define as $\mathcal{Q}_\mathrm{AW}$ the set of $Q$s that satisfy Assumption (ref).
Under Assumption (ref), athey_policy_2021 define a cross-fitted doubly robust welfare estimator:
where $\hat{\Gamma}_i$ denotes a cross-fitted doubly robust score for the conditional gain from treating unit $i$. Then the corresponding policy rule is:
Notice that, relative to EWM, the choice rule $\delta(\cdot)$ is unchanged: the novelty lies in the construction of $m(S^n,d)$, which now depends on first-step machine-learning estimators of nuisance components. In the notation of Definition (ref), this corresponds to setting $\theta = n$, letting $\omega$ denote an observational data collection process satisfying Assumption (ref).ID, setting $m(S^n,d)=\hat{W}_n(d)$, and choosing $\delta(\cdot)=\arg\max_{d\in\mathcal{D}_v}\{\cdot\}$. Define moreover
Under standard regularity conditions, $S_{Q_n}^*$ is the semiparametric efficiency bound for evaluating the best policy in the class.
Proposition (ref) shows the same leading dependence on the ratio between policy complexity and sample size shown in (ref): regret decays at the $\sqrt{v_n/n}$ rate, up to state-space-dependent constants and logarithmic terms.
The results in athey_policy_2021 were very influential in subsequent literature studying how specific extensions of EWM would behave when propensity scores were unknown.
\paragraph{Policy Learning with Network Interference.} viviano_policy_2024 extends empirical welfare maximization to settings with network spillovers. Relative to kitagawa_who_2018 and athey_policy_2021, the main difference is that a unit's outcome may depend not only on her own intervention, but also on the interventions assigned to her neighbours. A policy therefore determines both each unit's intervention and her exposure to treated neighbours. Consequently, welfare depends on the policy profile over the entire network, and observations are no longer independent.
The researcher may observe only a sample of units, together with information about their neighbours. Identification requires sampling to be exogenous conditional on observables, intervention assignment to be conditionally unconfounded, and sufficient overlap for both own intervention status and network exposure. The policy class is assumed to have finite VC dimension. In addition, the network cannot become too dense relative to the effective sample size. Formally, if $N_n$ measures the maximum neighbourhood size, $\delta_n$ is the smallest probability of observing a relevant exposure, and $n_e$ is the expected number of sampled units, the paper requires
for some $\xi\in(0,1/2]$.
viviano_policy_2024 proposes Network Empirical Welfare Maximization (NEWM), which maximizes an augmented inverse-probability-weighted estimate of welfare. The estimator reweights each sampled unit by the probability of observing the intervention and network exposure induced by a candidate policy, while allowing for a regression adjustment. For many policy classes, the resulting optimization problem can be formulated as a mixed-integer linear program.
With known propensity scores and a bounded regression adjustment, the regret of NEWM is of order
where $v$ is the VC dimension of the policy class. Under the network sparsity condition, regret is therefore $O(n_e^{-\xi})$ and, when the maximum degree is uniformly bounded, recovers the usual $1/\sqrt{n_e}$ rate. Thus, network interference affects policy learning through both network density and exposure overlap. The paper also establishes a corresponding lower bound and proposes a network cross-fitting procedure for estimated nuisance components.
A second thread of literature has focused on how to choose the policy space $\mathcal{D}_v$ optimally, and on the properties of special cases.
\paragraph{Model Selection for Policy Learning.} mbakop_model_2021 studies the problem of selecting optimally the complexity of the policy space. Formally, they consider sieves of assignment rules
of increasing VC dimension. One key insight is that, given this sieve characterization, we can decompose regret as
Here, $d^*$ denotes the oracle policy in the collection and $d^*_k$ the oracle policy in the sieve at level $k$. This decomposition highlights an estimation-approximation error trade-off. On the one hand, the more complex the policy space is, the closer the relative oracle $d^*_k$ will be to the unconstrained optimum $d^*$, and therefore the lower the approximation error. On the other hand, higher complexity leads to higher estimation error both pointwise and in the worst case. To trade these two errors efficiently, mbakop_model_2021 define $\hat{d}_\mathrm{PWM}$ as:
where $C_n(k)$ denotes a cost function that penalizes the policy space complexity, and $\sqrt{t_k/n}$ is a technical device used to ensure that, if $K$ is not finite, the classes get penalized at a sufficiently fast rate as $k$ increases.
Let's now map this framework back onto Definition (ref). In particular, PWM is defined by the estimator
and the choice rule
The margin of choice for data collection is $\theta = n$ and $\omega$ is required to respect unconfoundedness. The state space is the one considered in kitagawa_who_2018\footnote{See Assumption (ref) for the formal definition.}, with additional structure constraining the complexity penalizer $C_n(k)$ and the sieves. Formally, $\mathcal{Q}_\mathrm{MT}$, compared to $\mathcal{Q}_\mathrm{KT}$, further imposes that there exist positive constants $c_0$ and $c_1$ such that $C_n(k)$ satisfies the following tail inequality for every $n, k$, and for every $\epsilon>0$
and that there exists a universal constant $C_1$ such that, for every $n$, $C_n(k)$ satisfies
mbakop_model_2021 show that, provided that $d^* \in \mathcal{D}$, for a finite $K$,
where $v_k$ denotes the VC dimension of $\mathcal{D}_k$ that contains $d^*$. This is a powerful result as it shows that, under the additional conditions stated above, PWM achieves the minimax rate proved in kitagawa_who_2018. As a result, a policymaker who does not have a preference for a specific policy space can let the data identify the best among competing spaces, optimally trading off approximation and estimation error. The optimality is implied by the fact that PWM achieves the minimax rate, which no alternative algorithm can uniformly improve.
\paragraph{Extensions of PEWM.} Recent papers extend the rationale of mbakop_model_2021 to richer observational settings within the doubly robust framework of athey_policy_2021. fang_2025 extends this logic to observational settings with multivalued interventions and unknown propensity scores. Relative to mbakop_model_2021, the main innovation is replacing the empirical welfare criterion with a cross-fitted doubly robust welfare estimator athey_policy_2021, and then selecting among sieve policy classes using either Rademacher or holdout penalties. They derive oracle inequalities showing that the resulting data-driven rule trades off approximation and estimation error as if these quantities were known, and illustrate how to implement the sieve selection with monotone single-index rules and discretizations of smooth policy functions based on linear sieves or deep neural networks.
ai_2026 extends this logic to continuous interventions. In that setting, welfare can no longer be evaluated by a simple sample average, so $m(S^n,d)$ must approximate the value of each continuous level by smoothing across nearby intervention levels. The main contribution is therefore to show how the PEWM can be extended when one must jointly choose policy complexity and a bandwidth parameter. The paper develops a penalized procedure that automates both choices using the data. When propensity scores are unknown, a double-debiased version athey_policy_2021 yields a comparable guarantee, so the paper can be read as the continuous-intervention counterpart to the multivalued-intervention extension in fang_2025.
Ponomarev2026 studies a closely related adaptive choice of policy complexity in observational settings, while keeping focus on utilitarian welfare and binary choice rules. Compared to mbakop_model_2021, their contribution is mainly theoretical: combining doubly robust welfare estimation with sample splitting, they derive sharper finite-sample upper bounds on expected regret, a matching lower bound up to constants, and therefore show that Adaptive Welfare Maximization is minimax-rate optimal while achieving nearly-oracle performance in selecting the relevant policy class.
\paragraph{Policy Learning with Unobserved Heterogeneity.} opocher_2026 studies the case in which intervention effects vary with a latent characteristic $A_i$ that is not directly observed, but only through a noisy proxy $\hat{A}_i$. Relative to kitagawa_who_2018, the assignment mechanism and utilitarian welfare are unchanged, but the policy space can either ignore this source of heterogeneity and use rules $d(X_i)$, or augment targeting with the proxy and use rules $d(X_i,\hat{A}_i)$. The paper evaluates both classes relative to an oracle that observes $A_i$ and derives rate-sharp regret bounds for each. For covariate-based rules, worst-case regret equals the usual statistical term plus an approximation term reflecting the residual intervention-effect heterogeneity left unexplained by $X_i$. For proxy-augmented rules, worst-case regret equals the usual statistical term plus a noise term proportional to the root mean-squared error of $\hat{A}_i$. Therefore, richer targeting variables improve policy learning only if the welfare gain from capturing latent heterogeneity outweighs both the increase in policy-class complexity and the noise introduced by measuring the proxy. In the notation of Def. (ref), this means that the collection parameter space $\Theta$ collects the sample size of the experiment (which defines $\omega$) and the quality of information $t$ that the researcher has about $A_i$. Examples of $t$ include the number of measurements available in a repeated-measurement setting, the sample size used to learn $\hat{A}_i$ when it is estimated with auxiliary data, or the image quality when $\hat{A}_i$ is measured with satellite images. The paper then uses these bounds to study the data-collection problem over $\Theta = (n,t)$: under a budget constraint that specifies the cost of $n$ and $t$, and the available budget, how should the policymaker trade off improving the precision of $\hat{A}_i$ and increasing the policy-learning sample size? It derives sufficient conditions under which either the covariate-based design or an augmented design with information level $t$ is minimax-dominant in the sense of Eq. (ref).
\paragraph{Choosing Who Chooses.} kitagawa_choosing_who_choses expands the policy space in a different direction. Relative to the binary assignment problem in kitagawa_who_2018, the planner now chooses among three policies for each covariate value: assigning the intervention, assigning the status quo, or letting the unit self-select into intervention. The paper compares paternalistic (top-down assignment) and laissez-faire (bottom-up self-selection) approaches to policymaking. Its main conceptual point is that mixing the two approaches across the population (i.e. deciding separately for each covariate value which approach works best) outperforms both approaches when considered separately. The optimal policy is estimated via empirical welfare maximization, implemented over policy trees, on this richer action space. The paper further uses LATEs for takers and non-takers to interpret why self-selection raises welfare in some values of the covariate space and lowers it in others.
\paragraph{Dynamic Assignment.} Shosei_2025 extends EWM to allow for dynamic policy choices that map units' histories into intervention assignments over multiple stages. Relative to kitagawa_who_2018, the data collection process $\omega$ is required to satisfy sequential unconfoundedness, while the rest of the state space preserves the same boundedness, overlap, and complexity assumptions (see Assumption (ref)). The paper proposes two dynamic EWM procedures, one based on backward induction and one based on simultaneous maximization over the whole regime. The first is computationally lighter, but is guaranteed to recover the optimal constrained regime only when the first-best continuation rules belong to the feasible class at later stages; the second is computationally harder, but remains consistent for the best regime within the constrained class without that requirement. For both procedures the paper derives finite-sample, worst-case regret bounds and shows that, in the experimental setting, worst-case regret decays at the minimax-optimal rate of kitagawa_who_2018. It further shows how the simultaneous procedure can accommodate intertemporal budget or capacity constraints.
\paragraph{Matching Policies.} hazard_kitagawa_2025 studies a different policy object: rather than assigning one intervention to each unit, the planner matches units on one side of the market to units on the other side. Relative to kitagawa_who_2018, the policy is therefore not a rule $d:\mathcal{X}\to\mathcal{T}$ but a feasible matching policy $\pi_n$ over pairs $(X_i,X_j)$ satisfying one-to-one constraints. The paper estimates an average match cost $c(X_i,X_j)$ from training data and then chooses $\hat{\pi}_n$ by entropy-regularized empirical optimal transport. Its main result is a non-asymptotic upper bound on regret relative to the oracle unregularized matching policy. Under bounded costs, expected regret is bounded by the estimation error of $\hat{c}$ plus a regularization-bias term of order $\log n/\eta$, where $n$ is the size of the matching market and $\eta$ indexes the strength of regularization. The paper emphasises the computational tractability of the method but does not provide a matching lower bound on regret.
A third thread of literature has considered different welfare objectives that may interest a real-world policymaker.
\paragraph{Equality-Minded EWM.} kitagawa_equality-minded_2021 extended EWM to allow for an equality-minded objective function. In particular, they consider a welfare function that satisfies the Pigou-Dalton principle of transfers: a transfer of (potential) outcome from a higher-ranked individual to a lower-ranked individual is always desirable when it does not change their ranks in the status quo. While satisfying this constraint, a policymaker is allowed to prioritize (meaning place more weight on) lower ranks of the outcome's distribution.
Formally, they redefine welfare as:
where $F(y,d)$ is the quantile function of the outcome under policy $d$ and $\Delta: [0,1] \to [0,1]$ is a nonincreasing, nonnegative function with $\Delta(0)= 1$ and $\Delta(1) = 0$. $W_Q^\mathrm{eq}(d)$ satisfies the Pigou-Dalton principle if and only if $\Delta(\cdot)$ is convex. Regret is then defined as
They define a new welfare estimator:
where
Finally, they define Equality-Minded EWM (EM-EWM) as:
The main result of the paper mimics the main result in kitagawa_who_2018. They show that, under the assumptions that (i) $\Delta$ is nonincreasing and convex, with a finite first derivative at the origin; (ii) the policy class $\mathcal{D}_v$ has finite VC dimension; and (iii) strict overlap, unconfoundedness, and concentration of the outcome hold, $(m_\mathrm{EM}, \delta_\mathrm{EM})$ attains the minimax rate. Interestingly, the minimax rate is the same as the utilitarian welfare case, $\sqrt{v/n}$, as all the welfare losses that come from being equality-minded decay exponentially with $n$.
Two recent papers push this logic further. fan_qi_xu_2025 keeps the group-agnostic flavor of EM-EWM but sharpens its distributional concern by replacing the rank-dependent welfare functional with an $\alpha$-expected welfare criterion that maximizes the average welfare of the worst-off $\alpha$-fraction of the population, thereby interpolating between utilitarian EWM and a Rawlsian objective; they propose a debiased estimator and derive asymptotic regret bounds and inference for optimal welfare. terschuur_2025 instead abstracts from the rank-dependent class used in EM-EWM and studies policy learning with general semiparametric social welfare functions estimable by locally robust orthogonal moments and U-statistics, a framework that accommodates inequality-aware, inequality-of-opportunity-aware, and intergenerational-mobility objectives while still delivering asymptotic regret guarantees.
\paragraph{Fair Policy Learning.} viviano_fair_2024 extends EWM to policy makers who are willing to sacrifice some aggregate welfare to protect groups defined by a sensitive attribute $S_i$. Rather than maximizing a fixed weighted average of group welfare, the policy maker first restricts attention to the Pareto frontier: policies for which one group's welfare cannot be improved without reducing that of another group. With two groups, this frontier can be traced by maximizing
for different values of $\alpha\in(0,1)$. Among the resulting Pareto-efficient policies, the policy maker selects the one that minimizes a chosen measure of unfairness, denoted by $U_Q(d)$.
The framework accommodates several notions of unfairness. These include differences across groups in the probability of receiving the intervention, disparities in the welfare generated by the policy, and violations of incentive compatibility measured by the benefits that individuals could obtain by misreporting their sensitive attribute.
Because the Pareto frontier is unknown, it must be estimated from the data. The authors propose an algorithm that approximates this frontier and selects the estimated policy with the lowest unfairness. Under unconfoundedness, strict overlap, bounded outcomes, a policy class with finite VC dimension, and sufficiently accurate estimation of the conditional means and propensity scores, the resulting unfairness regret converges at rate $1/\sqrt{n}$. A corresponding lower bound shows that no data-dependent procedure can achieve a uniformly faster rate. The proposed policy rule is therefore minimax-rate optimal. \paragraph{Policy Learning with Random Constraints.} sun_2026 considers the case where the budget constraint of Definition (ref) is uncertain and therefore needs to be estimated.
Formally, the cost of collecting data is not considered (in the notation of Definition (ref), $c(\omega,\theta,Q)=0$), and the cost of implementing a policy $d$ is random: $\sigma(d,Q)=\sigma_i$. Then, the decision problem can be re-written as:
Define $\Sigma_Q(d) := \mathbb{E}_Q[\sigma_i\cdot d(X_i)]$. sun_2026 focuses on uniform asymptotic efficiency:
for any $\epsilon > 0$, and uniform asymptotic feasibility:
The first important result in sun_2026 is that, for a sufficiently rich state space (such as $\mathcal{Q}_\mathrm{KT}$), it is impossible for any $(m,\delta)$ to achieve both uniform asymptotic efficiency and feasibility. A second result is that the trivial extension of EWM that just replaces $\Sigma_Q(d)$ with the sample analogue is neither uniformly asymptotically welfare-efficient nor uniformly asymptotically feasible. Finally, sun_2026 introduces the new welfare function $W_Q^\mathrm{TO}$ that trades off units of welfare for budget overrun:
The intuition is that the PM can decide to borrow the quantity $(\Sigma_Q(d) - B_0)_+$ at a cost $\bar{\lambda}$, which can be measured as a loss in units of welfare with the scaling factor $r$. Then, the paper defines a trade-off rule as:
The trade-off rule is shown to achieve uniform asymptotic efficiency and to guarantee a finite bound on the budget overrun. However, $\hat{d}_\mathrm{TO}$ is not proven to be minimax optimal or minimax-rate optimal.
In this section, I highlight some areas at the frontier of the literature that I personally find interesting.
ishihara_evidence_2026 study policy choice when the planner has no direct sample from the target population, but only a collection of external studies reporting effect estimates $\hat\tau_k\sim\mathcal{N}(\tau_k,\sigma_k^2)$ for related populations, $k=1,\dots,K$. Writing $\tau_0$ for the target-population welfare effect, they restrict attention to non-randomized linear aggregation rules
and assume that the feasible set for $(\tau_0,\tau_1,\dots,\tau_K)$ is symmetric and invariant to common shifts. Under these conditions, maximum regret depends on the weights only through the worst-case bias
and the sampling standard deviation
Their main theorem shows that the minimax-regret aggregation weights solve
The key point is that evidence aggregation becomes a bias-variance problem tailored to regret rather than estimation error: relative to the minimax-MSE rule, minimax regret places more emphasis on controlling worst-case bias in extrapolating from study populations to the target one. For meta-regression and Lipschitz-type parameter spaces, the bias term $b(w)$ can be computed by linear programming, which makes the rule operational in applications.
\paragraph{Follow-up Questions.} As the authors acknowledge, one follow-up question arising from this study is how to account for publication bias. I divide publication bias into two layers. First, I consider the editor as an agent who assigns different publication probabilities depending on the statistical significance of results, e.g. whether $\hat{\tau}_k/\sigma_k > 1.96$. This would make the actual distribution of observed evidence a truncated normal and would keep the problem tractable. Second, I consider the researcher as an agent who can design protocols (e.g. data-cleaning algorithms) such that $\hat{\tau}_k/\sigma_k > 1.96$ by design. This would maximize the probability of publishing the paper by backward induction. In practice this could happen in many ways. For example, if the algorithm chosen by the researcher to estimate $\tau$ is sufficiently unstable bousquet, they could trim the data until the result becomes statistically significant.
This scenario opens several questions. How can optimal evidence-aggregation weights account for this distortion? Should the policymaker ignore some studies if there is evidence of manipulation? Should the policymaker set the weights ex-ante and also be minimax against manipulation, or should they instead let the weights adapt to manipulation in the spirit of armstrong_kline_sun_25? At what cost would it be optimal for the policymaker to request the full data from the researcher and independently reassess the results?
It is often tempting to think about complex interventions as binary treatments. This is because such a simple framework allows us to provide solid theoretical guarantees and intuitive guidelines for applied researchers. As an example, consider an information campaign. We describe as treated an individual who knows the information sent in the campaign and as untreated an individual who does not know it. Then, the implementation is the exact way such information is conveyed to individuals. Examples include a flyer in the mailbox, a text message on WhatsApp, or a video on YouTube. Simple ways to think of implementation include considering it as an independent dimension, that is, the way we intervene on the status quo while keeping the intervention fixed, or as a feature of the intervention itself. A more sophisticated way would be to consider the map we adopt to transport an individual from the distribution of their potential outcome under the status quo to the distribution under the intervention. Either way, the policy recommendation a researcher may produce about the information campaign is informative only for a policymaker who plans to intervene in the full population in exactly the same way. In other words, implementation is fixed across the learning and implementation stages. However, it is often not possible for the policymaker to replicate the implementation the researcher adopted in the learning sample in the full population. This could be the case because of implementation costs or ethical concerns.
This scenario opens several questions. Should this mismatch between how a researcher and a policymaker can intervene in the status quo be taken into account at the research design stage? What's the normative value of studies that cannot be implemented? How can those contribute to policy choice? Consider the case in which a policymaker has an implementation in mind and can aggregate evidence across studies that have different implementations as in ishihara_evidence_2026. Under what conditions is it optimal to aggregate existing evidence while keeping the original target implementation? When is it optimal to change the target implementation to be closer to those already tested? And when should a policymaker collect their own data and directly test the target implementation while ignoring existing evidence?
This section is tailored to a more applied audience and has the objective of guiding the steps an applied researcher should take to produce and communicate a policy recommendation.
In the first part, I provide two diagrams that position each theoretical paper in relation to concrete questions and concerns that an applied researcher may consider when solving the research design problem. This part takes the ex-ante perspective described in the rest of the review and can be considered as a map that practitioners can use to navigate the theoretical literature.
In the second part, I consider the ex-post perspective of a researcher who has solved the research design problem and collected the data, and now wants to communicate the value of their recommendation in a simple and intuitive way. I introduce a beta version of policytargetr, a new R package that takes the data and research design as inputs and produces one table and two graphs that the researcher can plug into the `policy implications' section of their paper.
I present the setting of hussam_2022 as a working example to show a use case for the diagrams in Section (ref) and the package in Section (ref).
At this stage, the researcher is writing what one may call a normative pre-analysis plan. This would fix a choice for $(m,\delta)$ before the data collection takes place, following the guidance that the theory provides on the interplay with $(\omega,\theta)$.\footnote{See Definition (ref) for reference on notation.} The data-collection process $\omega$ and its design margins $\theta$ are not fully fixed yet, but the researcher usually has a tentative environment in mind dictated by practical feasibility constraints and wants to understand which choice of $(m,\delta)$ best fits their setting. Therefore, conditional on this tentative $(\omega,\theta)$, the ex-ante question is which estimator $(m)$ and decision rule $(\delta)$ the theory suggests.
A normative pre-analysis plan for policy choice should make explicit at least the following objects:
To mimic the structure of the theoretical review, Figures (ref) and (ref) help navigate the literature, starting with the conditions that may interest an applied researcher across the four dimensions above. Figure (ref) refers to an unconstrained policy space (details in Section (ref)). Figure (ref) refers to a constrained policy space (details in Section (ref)).
A researcher can start from the most binding feature of the environment she expects to face in terms of $(\omega, \Theta)$ (e.g. the assignment is not random, there may be spillovers, there is partial internal validity) and its interplay with $m$, and then ask what the theory suggests to do when choosing $(m,\delta)$. Each leaf then points to the branch of the theory that is most informative for choosing $(m,\delta)$ for that specific concern.
As a working example, consider the setting in hussam_2022. The authors study the effect of providing a cash grant to micro-entrepreneurs on their profits with a randomized controlled trial in rural India. This experiment is motivated by the normative objective of fostering economic development in developing countries by investing in local entrepreneurs. The trial was conducted in the city of Amravati, India, between 2016 and 2018. The sample consists of 1,345 micro-entrepreneurs operating informal businesses in retail and services. First, participants were assigned to peer groups of five based on geographic proximity. Within these groups, individuals were asked to rank their peers on future business outcomes, including future profits and marginal returns to capital. Then, one-third of the sample was randomly assigned to receive an unconditional cash grant of 6,000 INR (roughly \$100). The authors study the effect of the cash grant on various outcomes. Here we focus on entrepreneurs' profits.
I now highlight three challenges for research design for policy choice that could arise in this setting.\footnote{Unfortunately, it is not possible to confirm that these were actual concerns for the authors, since the pre-analysis plan is not available.}
First, spillovers may be a first-order concern in this environment. Entrepreneurs are grouped by geographic proximity and operate in local retail and service markets, so one unit's grant may affect the profits of other units through competition, demand spillovers, imitation, or informal insurance within the peer group. If this is a serious ex-ante concern, the relevant branch in Figure (ref) is the assignment leaf on network spillovers, which points to viviano_policy_2024. In practice, this means that before collecting data the researcher should decide how to measure the relevant network, what variation in neighbours' intervention exposure the experimental design must generate, and then choose an estimator $m$ and decision rule $\delta$ that are robust to interference rather than only to i.i.d. sampling uncertainty, as suggested in viviano_policy_2024.
Second, the PM may care about a target population different from the one observed in Amravati. A lender or ministry may be interested in applying the same grant program in another district, another Indian state, or even another country. If this is a first-order concern for the researcher, Figure (ref) points to adjaho_external_2022 and kido_distributionally_2022. In such cases, the researcher may want their recommendations to be robust over a set of plausible target populations. As a result, they may follow adjaho_external_2022 and kido_distributionally_2022 to define the welfare function and the empirical analogue to maximize accordingly. Moreover, this choice affects what covariates and contextual information should be collected to quantify the distance between populations.
Third, the PM may care not only about average profits, but also about how those gains are distributed across the target population. A targeting rule based on community rankings could maximize utilitarian welfare and still be unattractive if it systematically disadvantages poorer entrepreneurs, younger entrepreneurs, or protected groups. Figure (ref) highlights two possibilities. If the relevant normative concern is inequality in the distribution of outcomes, the appropriate references are kitagawa_equality-minded_2021, fan_qi_xu_2025, and terschuur_2025; these papers would guide the researcher in defining equality-minded welfare objectives and appropriate empirical analogues. If instead the PM wants explicit fairness constraints across protected groups, viviano_fair_2024 suggests maximizing among Pareto-optimal allocations over protected groups.
In this section, we focus on the stage where data collection and the main analysis have already been completed. Therefore, the main focus here switches from making research design choices to communicating a policy recommendation to the policymaker. In particular, the main objective of this section is to provide applied researchers with two figures and a table they can plug in the `policy implications' section of their applied paper to provide evidence on the expected performance of their recommendation.
Let's suppose a researcher has collected data from the randomized controlled trial in hussam_2022, which was funded with the objective of producing recommendations about whether to distribute the intervention through full adoption, no adoption, or targeting based on prespecified covariates with a rectangular rule. For simplicity, let us abstract from the design challenges highlighted in the previous section and focus on the simplest setting, which is the one the authors considered.
The rankings entrepreneurs give one another for future profits, henceforth community rankings, play a central role in the paper. hussam_2022 write: `[...] lending institutions or other organizations aiming to target capital to entrepreneurs with productive opportunities would have good reason to leverage community information'. hussam_2022 motivate this conclusion with the targeting results in Table 2. In the discussion immediately below Table 2, they note that bottom-tercile entrepreneurs have returns statistically indistinguishable from zero, middle-tercile effects are not statistically significant, and the strongest intervention effects are concentrated in the top tercile. They also emphasize that these estimates are stable after controlling for a rich set of observables. Consistent with this, Table 4 shows that adding community rankings to observables strengthens prediction. Although these findings provide valid evidence for discovery objectives alone, they do not provide direct evidence in support of the normative objective: should the PM roll out the cash transfer under full adoption? Should the PM shut down the policy and assign it to nobody? Or is targeting based on the community rankings the best option?
To answer these questions, I introduce (the beta version of) a new R package called policytargetr that takes the outcome variable, the intervention dummy, the propensity score, and two targeting covariates as inputs and produces (i) a plot of the estimated rectangular frontier, (ii) a plot of the welfare produced by each recommendation, and (iii) a companion table.
Listing (ref) reports the workflow used in the context of hussam_2022. The first step is to clean the application data so that it matches the package's expected format: one row per unit and four required columns called outcome, intervention status, x1, and x2. In this application, outcome is follow-up profits, treatment is the treatment-assignment dummy, and \texttt{x1} and \texttt{x2} are, respectively, age and community-rank percentile, the two targeting variables.
In a nutshell, the package uses split_seed and train_share to split the original data set into training and test sets. One can also specify a policy cost function (the $\sigma$ function in Definition (ref)) with the argument cost_fun. For this specific application we know each grant costs 6,000 INR. To ease interpretation, we normalize the cost of treating the whole test set to 1. On the training sample, analyze_rectangular_policy() takes propensity and the formatted data object as inputs to learn the EWM kitagawa_who_2018 rectangular targeting rule. It then computes the empirical welfare under no adoption, full adoption, and targeting on the test sample, returning the fitted object fit. Because the training and test samples are independent, we can perform valid holdout inference and provide standard errors and confidence intervals together with out-of-sample empirical welfare. With \texttt{save_policy_outputs}, the practitioner saves Table (ref) and Figures (ref) and (ref) in the \texttt{output_dir}, with prespecified labels for graphs and a prefix for file names.
Of course, one can extend this beta version, which can be considered a minimum viable product, in many ways by following the literature reviewed in this paper. The same output could then be produced under various departures from this simple setting that follow the directions highlighted in Figures (ref) and (ref).
Table (ref) provides a summary of the results. Panel A reports general information: the total sample size, the experimental propensity score, and, most importantly, the estimated targeting rule, which assigns the cash grant to entrepreneurs older than 29.5 years and above the 71st percentile of the community ranking. Panel B then reports, for each recommendation, both empirical welfare and normalized implementation cost. Full adoption attains the highest point estimate, but targeting is close in welfare while costing one-fourth as much under the specified cost function. Panel C is the key decision panel: targeting dominates no adoption by about 1,021 INR in holdout welfare and this difference is statistically significant, whereas the difference between targeting and full adoption is small and imprecise. The implication is that using community information to target the grant improves meaningfully on the status quo and preserves most of the gains from universal rollout at a much lower cost.
Figure (ref) visualizes the comparison in Panels B and C of Table (ref). The dots report test-set empirical welfare for the three candidate recommendations, the vertical bars report 95% confidence intervals, and the brackets add the pairwise welfare differences together with their p-values. First, the targeting rule lies clearly above no adoption, and the lower bracket makes it transparent that this gain is statistically significant. Second, targeting and full adoption are visually very close, with overlapping confidence intervals and a small, insignificant bracketed difference. Thus, the figure conveys the main practical message of the application immediately: targeting appears to capture most of the welfare gains from expansion while avoiding the cost of treating everyone.
Figure (ref) visualizes who should receive the transfer, as also reported in Table (ref), Panel A. The blue vertical and horizontal cutoffs trace the estimated rectangular policy frontier, while the shaded upper-right region identifies the entrepreneurs who receive the grant under the recommended policy.
Economists' role is recognized worldwide as providing empirical evidence to guide real-world policy choices. Public institutions, research organizations, and the profession itself acknowledge this role. As a result, much empirical work in economics is motivated, and often funded, by broader normative objectives. The success in achieving such objectives is tied to economists' research design choices.
In this review, I first surveyed the theoretical literature initiated by manski_statistical_2004 that studies this problem through the lens of statistical decision theory wald1950statistical, Savage01031951. I proposed a common framework that nests choices over data collection, design margins, estimators, and decision rules, and used it to organize the literature into unconstrained and constrained policy-choice problems. This made it possible to clarify under a common framework what different papers hold fixed, what object they optimize, what type of regret guarantee they deliver, and how recent work extends the earlier contributions.
Then, I turned to the applied side and provided two tools that researchers can use when the goal is to produce and communicate policy recommendations. First, I introduced a practitioner-oriented guide for the ex-ante stage, meant to help researchers write a normative pre-analysis plan by mapping concrete design concerns into the relevant branch of the theory. Second, I introduced a simple ex-post workflow, operationalized through the policytargetr package, to help communicate to policymakers not only what the recommended policy is, but also how it compares with alternatives in terms of welfare and implementation costs.
This review ultimately aims to organize recent theoretical results under a common framework and bring them closer to mainstream empirical practice.