Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
58,355 characters · 19 sections · 105 citation commands
Nonparametric Bayesian Policy Learning
\thispagestyle{empty}
\noindentKeywords: Bayesian bootstrap, Dirichlet process, program evaluation, risk bound, statistical treatment rules. \\ JEL codes: C11, C44, D60.
\setcounter{page}{1}
In many economic settings, policymakers must allocate scarce interventions across heterogeneous populations. This motivates a growing literature on statistical treatment rules, or policy learning, pioneered by manski2004statistical, which studies how to design treatment rules using observable individual characteristics to maximize social welfare; see, for example, hirano2009asymptotics, stoye2009minimax, bhattacharya2012inferring, kitagawa2018should, and athey2021policy, among others.\footnote{Throughout the paper, I use the terms “treatment rule” and “policy” interchangeably.} A growing empirical literature shows that such individualized rules can deliver substantial welfare gains across a wide range of settings, including labor market and education programs athey2025machine, goller2025active, anti-poverty interventions in developing countries haushofer2025targeting, higbee2025policy, and health and precision medicine kosorok2019precision, inoue2023machine, athey2025targeted.
In this paper, I consider a decision-maker (DM) seeking to select an expected welfare-maximizing treatment rule using observable characteristics. A key observation is that, for a given welfare criterion and policy class, the expected welfare function is fully determined by a reduced-form data distribution. Consequently, the sole source of statistical uncertainty in the DM's problem is uncertainty about this distribution. This blends seamlessly with a Bayesian approach in which the DM assumes a nonparametric prior for the reduced-form parameter and uses the resulting posterior to conduct simultaneous inference on optimal treatment assignments, optimal welfare, and comparisons across policy classes. I call this framework Nonparametric Bayesian Policy Learning (NBPL).
NBPL has several appealing features. First, it is grounded in a well-defined posterior, so uncertainty enters the decision problem explicitly. A growing literature shows that ignoring uncertainty in treatment or policy choice can lead to suboptimal decisions dehejia2005program, christensen2025optimal, moon2026optimal. Second, to the best of my knowledge, NBPL is the first unified inferential framework for optimal treatment assignments, optimal welfare, and comparisons across policy classes. The last component incorporates uncertainty about the policy class and is difficult to handle with frequentist methods. Third, NBPL is flexible: the prior for the reduced-form data distribution is a Dirichlet process ferguson1973bayesian, which offers a level of flexibility similar to that of frequentist nonparametric approaches to policy learning kitagawa2018should. Fourth, posterior computation is highly tractable: the recommended Bayesian-bootstrap implementation rubin1981bayesian simply reweights the data using standard exponential weights, reuses off-the-shelf welfare maximization algorithms, and is naturally parallelizable.
I also establish three important theoretical properties of NBPL. First, I show that posterior welfare regret under NBPL converges to zero at the minimax-optimal rate (Theorem (ref)), matching the rate for EWM in kitagawa2018should. This shows that explicitly accounting for uncertainty in the decision problem does not slow the rate of learning; to the best of my knowledge, it is the first posterior welfare-regret guarantee for policy learning. Second, I establish pointwise consistency of posterior model comparison across policy classes (Theorem (ref)); in particular, when two classes are strictly separated at the truth, the posterior asymptotically assigns vanishing probability to the inferior class. Third, I show that EWM arises as a Bayes rule within the NBPL framework (Proposition (ref)). Specifically, EWM first averages welfare over posterior uncertainty and then optimizes, whereas decisions based on the posterior over optimal treatment assignments instead optimize first and then aggregate. The two procedures need not coincide.
I illustrate the NBPL framework using two empirical applications: the Job Training Partnership Act (JTPA) experiment studied by kitagawa2018should and the anti-malaria bednet subsidy experiment analyzed by bhattacharya2012inferring. In both applications, EWM welfare estimates tend to fall below the posterior median of welfare under NBPL. NBPL also yields credible intervals for welfare that are tighter than the corresponding EWM confidence intervals. In addition, the posterior model comparison procedure provides strong evidence that decision-tree rules achieve higher welfare than linear rules.
This paper connects to three literatures in economics. The first is the literature on statistical treatment rules, or policy learning. A leading approach is the Empirical Welfare Maximization (EWM) framework of kitagawa2018should, which selects treatment rules as if the empirical distribution were the true data-generating process and achieves the minimax-optimal regret convergence rate. By contrast, NBPL places a Dirichlet process prior on the reduced-form data distribution, explicitly incorporates uncertainty into the decision problem through the posterior, and delivers the same minimax-optimal posterior regret rate. chamberlain2011bayesian is closely related in spirit, using Dirichlet priors on outcome distributions to derive posterior-expected-utility treatment decisions in a limited-information Bayesian framework. NBPL builds on this approach by accommodating continuous distributions and constrained policy classes, while also delivering asymptotic regret guarantees. NBPL is also related to the broader literature on uncertainty-aware decision-making dehejia2005program, chernozhukov2025policy, christensen2025optimal, moon2026optimal. Other contributions to policy learning include hirano2009asymptotics, stoye2009minimax, bhattacharya2012inferring, kallus2018confounding, athey2021policy, kitagawa2021equality, mbakop2021model, zhou2023offline, adjaho2025externally, kitagawa2025leave, olea2025decision, viviano2025policy, and ida2026choosing.
Another related literature studies Bayesian inference through reduced-form parameters. In many economic models, an identifiable reduced-form parameter maps into a structural parameter of interest, so a posterior on the former induces a posterior on the latter. The key distinction is whether the structural object is point or partially identified.\footnote{In partially identified settings, moon2012bayesian show that placing a prior directly on the structural parameter can lead to asymptotic divergence between Bayesian credible sets and frequentist confidence sets because the prior is not revisable by data. giacomini2021robust address this issue through a robust prior specification.} In NBPL, the optimal welfare value is point identified, whereas optimal treatment rules may be set-valued and thus only partially identified. Point-identified examples include chamberlain2003nonparametric and walker2026semiparametric, who study structural parameters defined by unconditional and conditional moment restrictions, respectively. NBPL is also closely related to the loss-likelihood bootstrap of lyddon2019general, with welfare interpreted as a negative loss. For the partially identified component, my inference procedure is closely related to kline2016bayesian and florens2021revisiting, which place a prior only on the reduced-form parameter and conduct inference on the structural object through the induced posterior; among these, florens2021revisiting is especially close in spirit because it places a Dirichlet process prior on the population distribution. Additional related contributions include norets2014semiparametric and liao2019bayesian. More broadly, andrews2025communicating develop a regret-based framework for evaluating approximate posterior reports and justify bootstrap distributions, including the Bayesian bootstrap, as devices for communicating uncertainty.
A final related literature studies frequentist inference for optimal treatment rules and optimal welfare values. Inference on the optimal welfare value is important because it quantifies the maximum gains from policy learning, but frequentist inference is challenging because this parameter is irregular; see, for example, the impossibility results in hirano2012impossibility. luedtke2016statistical, ponomarev2025lower, and whitehouse2025inference develop, respectively, confidence intervals, lower confidence bands, and smoothing-based inference procedures for this parameter. Inference on optimal treatment assignments is also crucial because it quantifies the strength of the evidence in favor of treating individuals selected for treatment by the rule, but it has received comparatively less attention; two frequentist examples are rai2018statistical and armstrong2023inference. Relatedly, kitagawa2023stochastic develop a quasi-Bayesian inference framework for linear rules that places a prior directly on the policy class rather than on the data distribution. The NBPL framework provides a Bayesian complement to this literature.
The rest of the paper is structured as follows. Section (ref) introduces the NBPL framework and its implementation. Section (ref) presents the empirical applications, and Section (ref) develops the theoretical results. Section (ref) concludes. Appendices (ref)--(ref) contain additional results and proofs.
A decision-maker (DM) seeks to choose a treatment rule in order to maximize expected welfare in a target population. Individual $i$ in the target population is characterized by the random vector $(Y_i(1), Y_i(0), X_i) \sim P_0^{\star}$, where $Y_i(1), Y_i(0) \in \mathcal{Y} \subseteq \mathbb{R}$ denote the potential outcomes under treatment and control, respectively, and $X_i \in \mathcal{X} \subseteq \mathbb{R}^{d_x}$ denote a vector of baseline covariates. The DM's problem is to choose a treatment rule $G \subseteq \mathcal{X}$ that maximizes utilitarian welfare,\footnote{Hereafter, I normalize welfare relative to the status quo of assigning no treatment, so that $W(P_0^{\star}; G) = \mathbb{E}_{P_0^{\star}}[(Y(1)-Y(0)) \mathds{1}\{X \in G \}]$. This normalization leaves the optimal treatment assignment unchanged, but welfare values should henceforth be interpreted as gains relative to the status quo.}
The DM chooses a treatment rule from a feasible policy class $\mathcal{G}\subseteq 2^{\mathcal X}$, so treatment assignment may depend only on observable individual characteristics $X$.
Suppose the DM observes a random sample $\mathcal{D}_n \coloneqq \{D_i\}_{i=1}^n$ drawn from an experimental population. Individual $i$ in this population is characterized by the random vector $(Y_i(1), Y_i(0), X_i) \sim Q_0^{\star}$. The observed data take the form $D_i \coloneqq (Y_i, T_i, X_i)$, where $T_i \in \{0,1\}$ denotes the realized treatment assignment and $Y_i = Y_i(1)T_i + Y_i(0)(1-T_i)$ denotes the realized outcome. Let $Q_0$ denote the joint distribution of $(Y,T,X)$ and define the propensity score $e(x) \coloneqq \mathbb{E}_{Q_0}[T_i \mid X_i=x]$.
Assumption (ref) is standard in the literature.\footnote{Assumption (ref)(c) weakens Assumption 2.1(BO) of kitagawa2018should and Assumption 3 of kitagawa2023stochastic by allowing unbounded outcomes. It serves as a technical moment condition for controlling empirical process terms in the proofs.} Assumption (ref)(a) ensures that the experimental data are informative about the target population distribution; in particular, it holds when the sample is drawn directly from the target population.\footnote{See, e.g., mo2021learning, kido2022distributionally, qi2023robustness, adjaho2025externally, ben2025safe for approaches that relax external validity in policy learning.} Assumptions (ref)(b) and (d) are standard in causal inference. Moreover, the assumption of a known propensity score is not restrictive in many applications: it holds by design in randomized controlled trials, including the two empirical applications considered in this paper bhattacharya2012inferring, kitagawa2018should, and can sometimes be reasonable in quasi-experimental settings angrist1990lifetime, angrist1999using, katz2001moving, abdulkadirouglu2011accountability, abdulkadirouglu2022breaking.
Under Assumption (ref)(a), (b) and (d), welfare in the target population under any treatment rule $G$ is point-identified as kitagawa2018should:
Here, $P_0$ denotes the reduced-form distribution of $\widetilde{D}$, where $\widetilde{D} \coloneqq (Z_1, Z_0, X)$ with $Z_1 \coloneqq YT/e(X)$ and $Z_0 \coloneqq Y(1-T)/[1-e(X)]$ (see Remark (ref)). The key observation is that, for any treatment rule $G$, population welfare is a deterministic function of $P_0$ only.
Appendix (ref) discusses alternative welfare criteria that can be handled by modifying the reduced-form object, and Appendix (ref) extends the setup to multi-valued discrete treatments.
Define the (possibly non-unique) set of optimal treatment rules under $P_0$ as
and the corresponding optimal welfare level as
The key challenge in welfare maximization (ref) is that the DM does not observe the true reduced-form distribution $P_0$. For any given policy class $\mathcal{G}$, both the set of optimal treatment rules and the optimal welfare level, $(G^{\star}(P_0), W_{\mathcal{G}}^{\star}(P_0))$, are deterministic functions of $P_0$ only. Consequently, all uncertainty about these welfare-relevant objects is induced entirely by uncertainty about $P_0$.
To account for this uncertainty, I assume the DM is Bayesian and places a nonparametric prior on the population distribution, $P \sim \Pi$. Specifically, $\Pi$ is taken to be a Dirichlet process (DP) prior ferguson1973bayesian, a standard nonparametric prior on probability distributions.
The Dirichlet process has two appealing features. First, it is highly flexible: it imposes no parametric restrictions on the underlying distribution and, provided the base measure $\alpha$ has full support, assigns positive prior mass to neighborhoods of any probability distribution.\footnote{ The Dirichlet process has large support under the weak topology (i.e., the topology of weak convergence of probability measures). Let $\mathcal{M}$ denote the space of Borel probability measures on $\mathcal{D}$, endowed with the weak topology. The weak support of $\mathrm{DP}(\alpha)$ is $\mathrm{supp}_w\big(\mathrm{DP}(\alpha)\big) \coloneqq \{ P \in \mathcal{M} : \mathrm{supp}(P) \subseteq \mathrm{supp}(\alpha) \}$.
In particular, if $\alpha$ has full support on $\mathcal{D}$, then $\mathrm{DP}(\alpha)$ assigns positive prior mass to every weak neighborhood of any $P \in \mathcal{M}$. This is appealing because large support under the weak topology induces a sufficiently rich prior over functionals of $P$, such as the optimal welfare $W_{\mathcal{G}}^{\star}(P)$. } Second, the Dirichlet process is conjugate ferguson1973bayesian: if $P \sim \mathrm{DP}(\alpha)$, then the posterior distribution satisfies $P \mid \mathcal{D}_n \sim \mathrm{DP}(\alpha + n \mathbb{P}_n)$, where $\mathbb{P}_n \coloneqq n^{-1} \sum_{i=1}^n \delta_{D_i}$ denotes the empirical measure. This conjugacy implies that posterior computation remains highly tractable, mitigating the possible concern that inference with nonparametric priors poses computational challenges.
The posterior over $P$ induces posterior uncertainty about both the optimal treatment rules and the optimal welfare level. In particular, the marginal posterior distributions of $G^{\star}(P)$ and $W_{\mathcal{G}}^{\star}(P)$ are given by the pushforward\footnote{ Given a measurable mapping $T: \mathcal{M} \to \mathcal{T}$ and a probability measure $\mu$ on $\mathcal{M}$, the pushforward measure $\mu \circ T^{-1}$ on $\mathcal{T}$ is defined by $(\mu \circ T^{-1})(A)=\mu(T^{-1}(A))$ for any measurable set $A \subseteq \mathcal{T}$. } of $\Pi(P \in \cdot \mid \mathcal{D}_n)$ under the mappings $G^{\star}$ and $W_{\mathcal{G}}^{\star}$, respectively:
I refer to (ref) and (ref) as the optimal treatment posterior\footnote{Equation (ref) is a posterior distribution over the identified set of optimal treatment rules, in the sense of kline2016bayesian. For example, if $\mathcal G^{\mathrm{fair}}\subseteq\mathcal G$ denotes the collection of rules satisfying a fairness constraint, then $\Pi\bigl(G^{\star}(P)\cap \mathcal G^{\mathrm{fair}}\neq\varnothing \mid \mathcal{D}_n\bigr)$ and $\Pi\bigl(G^{\star}(P)\subseteq \mathcal G^{\mathrm{fair}} \mid \mathcal{D}_n\bigr)$ are the posterior probabilities that at least one optimal rule, and that all optimal rules, satisfy the fairness constraint, respectively. Both are obtained by evaluating under $\Pi(P\in\cdot\mid\mathcal{D}_n)$ the set of reduced-form distributions $P$ for which the stated property holds, subject to the requisite measurability of the corresponding inverse image in the space of reduced-form distributions.} and the optimal welfare posterior, respectively. Together, they fully characterize posterior uncertainty about $(G^{\star}(P), W_{\mathcal{G}}^{\star}(P))$ and support coherent Bayesian inference for any functions thereof.
The posterior distributions (ref) and (ref) introduced above naturally imply a set of reportable summaries for optimal welfare, optimal treatment rules, and comparisons across policy classes.
Bayesian inference on the optimal welfare is straightforward, since (ref) defines a posterior distribution over a scalar parameter. One may therefore use the posterior median as a point estimate and report a $(1-\alpha)$ equal-tailed credible interval, with endpoints given by the $(\alpha/2)$ and $(1-\alpha/2)$ quantiles of (ref). This offers a Bayesian complement to existing frequentist approaches to inference on optimal population welfare luedtke2016statistical, ponomarev2025lower, whitehouse2025inference.
Bayesian inference on optimal treatment rules is more subtle, since (ref) is a posterior distribution over sets. Such inference can be conducted within the framework of kline2016bayesian, where each treatment rule is viewed as a parameter and $G^{\star}(P)$ is interpreted as the identified set of optimal rules.\footnote{ My prior specification fits naturally into the framework of kline2016bayesian, in which a point-identified reduced-form parameter $\mu$ (here $P$) maps into the identified set of a partially identified structural parameter $\theta$ (here the optimal treatment rule) through a known set-valued mapping (here $G^{\star}(\cdot)$). } The posterior in (ref) then enables posterior probability statements about the identified set, such as whether a candidate rule belongs to $G^{\star}(P)$, or whether all optimal rules satisfy a given property (e.g., a capacity constraint or fairness criterion). Likewise, quantities induced by the set of optimal rules, such as the treatment share, may themselves be set-identified.
In addition to inference on $(G^{\star}(P), W_{\mathcal{G}}^{\star}(P))$ for a fixed policy class $\mathcal{G}$, NBPL also supports comparison across classes and thus quantifies uncertainty about the policy class itself. Consider two classes $\mathcal{G}_1$ and $\mathcal{G}_2$; for example, a policymaker may wish to compare linear rules qian2011performance, kitagawa2018should and decision-tree rules athey2021policy, ida2026choosing while accounting for uncertainty about $P$. Since the posterior for $P$ induces a joint posterior for $(W_{\mathcal{G}_1}^{\star}(P), W_{\mathcal{G}_2}^{\star}(P))$, the posterior probability $\Pi(W_{\mathcal{G}_1}^{\star}(P) > W_{\mathcal{G}_2}^{\star}(P) \mid \mathcal{D}_n)$ provides a natural measure of the evidence that $\mathcal{G}_1$ delivers higher optimal welfare values. Theorem (ref) in Section (ref) establishes the pointwise consistency of this posterior comparison procedure.
In this section, I develop a tractable algorithm to generate $S \geq 1$ draws $\{(G^{\star,[s]}, W_{\mathcal{G}}^{\star,[s]})\}_{s=1}^S$ from the posterior distribution of $(G^{\star}(P), W_{\mathcal{G}}^{\star}(P))$ using the Bayesian bootstrap rubin1981bayesian.
Importantly, existing welfare maximization algorithms can be readily adapted by reweighing each observation with normalized independent $\mathrm{Exp}(1)$ weights in place of the uniform weights $n^{-1}$.\footnote{ Formally, draws from the Bayesian bootstrap posterior $\mathrm{DP}(n\mathbb P_n)$ admit the representation $P \stackrel{d}{=} \sum_{i=1}^n \widetilde{\omega}_i\,\delta_{D_i}$, where $\widetilde{\omega}_i = \omega_i/\sum_{j=1}^n\omega_j$ and $\omega_i\stackrel{\mathrm{i.i.d.}}{\sim}\mathrm{Exp}(1)$ rubin1981bayesian. Thus, Bayesian bootstrap draws can be viewed as randomly reweighted (and hence smoothed) versions of the empirical distribution. } The $S$ posterior draws can be computed in parallel.
Given posterior draws $\{G^{\star,[s]}, W_{\mathcal{G}}^{\star,[s]}, \pi^{[s]}\}_{s=1}^S$, one may estimate optimal welfare by the empirical median of $\{W_{\mathcal{G}}^{\star,[s]}\}_{s=1}^S$, and quantify uncertainty using a $(1-\alpha)$ equal-tailed credible interval based on the $(\alpha/2)$ and $(1-\alpha/2)$ empirical quantiles. Model comparison between two classes $\mathcal{G}_1$ and $\mathcal{G}_2$ can likewise be based on the empirical fraction of draws for which $W_{\mathcal{G}_1}^{\star,[s]} > W_{\mathcal{G}_2}^{\star,[s]}$.
I illustrate the Nonparametric Bayesian Policy Learning (NBPL) framework in two empirical applications and compare its performance with empirical welfare maximization. In both applications, I consider two policy classes---linear rules and decision-tree rules of depth at most two, denoted by $\mathcal{G}_{\mathrm{lin}}$ and $\mathcal{G}_{\mathrm{tree},2}$, respectively---and also illustrate posterior model selection, a novel feature of the NBPL framework.
The first application considers the Job Training Participation Act (JTPA) experiment. In this experiment, applicants' eligibility for job training and related services was randomly assigned for an 18-month period. The treatment indicator $T$ denotes program eligibility, and the outcome $Y$ is earnings measured 30 months after assignment. Following kitagawa2018should, I consider two outcome measures: one without accounting for treatment costs, and one that subtracts a treatment cost of \$774 per participant. Baseline covariates are given by $X = (\texttt{PreEarn},\texttt{Educ})^{\top}$, recording pre-treatment earnings and years of education. Random assignment ensures unconditional unconfoundedness, $(Y(1),Y(0),X)\perp \!\!\! \perp T$. In the dataset studied by kitagawa2018should, the sample size is $n=9{,}223$, and the propensity score is known and constant at $2/3$.
Figure (ref) reports the posterior distribution of optimal welfare under NBPL in the absence of treatment costs. Relative to kitagawa2018should, the analysis additionally considers decision-tree rules and supports posterior model comparison across policy classes.
The left panel plots the empirical cumulative distribution functions (CDFs) of $W^\star_{\mathcal{G}_{\mathrm{lin}}}(P)$ and $W^\star_{\mathcal{G}_{\mathrm{tree},2}}(P)$ across posterior draws. The distribution for $\mathcal{G}_{\mathrm{tree},2}$ lies uniformly to the right of that for $\mathcal{G}_{\mathrm{lin}}$, indicating first-order stochastic dominance. The EWM estimates correspond to relatively low posterior quantiles---approximately the 31st percentile for $\mathcal{G}_{\mathrm{lin}}$ and the 23rd percentile for $\mathcal{G}_{\mathrm{tree},2}$---suggesting that plug-in estimates may understate optimal welfare relative to the posterior distribution. The right panel plots the posterior distribution of the welfare difference $W^\star_{\mathcal{G}_{\mathrm{tree},2}} - W^\star_{\mathcal{G}_{\mathrm{lin}}}$. The posterior probability that $\mathcal{G}_{\mathrm{tree},2}$ yields higher optimal welfare than $\mathcal{G}_{\mathrm{lin}}$ is 94.8%, providing strong evidence in favor of the tree-based policy class.
Table (ref) in Appendix (ref) shows that the 95% equal-tailed Bayesian credible intervals are typically tighter than the 95% bootstrap confidence intervals reported by kitagawa2018should. Appendix (ref) provides additional empirical results.
The second example considers the randomized subsidy experiment for insecticide-treated bednets (ITNs) analyzed in bhattacharya2012inferring. In this study, rural households in Western Kenya were randomly offered bednets at varying subsidy levels through vouchers redeemable at local retailers dupas2009what. The treatment indicator $T$ equals one if household $i$ was offered a highly subsidized net and zero otherwise. The outcome $Y$ is a binary indicator for bednet coverage, equal to one if the household both redeemed the voucher and was observed using the net. Baseline covariates are given by $X=(\texttt{YoungChild},\texttt{BankAccount},\texttt{Wealth})^{\top}$, where $\texttt{YoungChild}$ indicates the presence of a child under age ten, $\texttt{BankAccount}$ indicates bank account ownership, and $\texttt{Wealth}$ denotes log household wealth per capita. Treatment was randomly assigned at the household level, implying unconditional unconfoundedness. In the dataset studied by bhattacharya2012inferring, the sample size is $n=1{,}098$, and the propensity score is constant at $0.16$.\footnote{The experiment features complete random assignment at the household level dupas2009what. Since the propensity score is not reported explicitly, I approximate it by the empirical treatment share in the sample.}
I consider two settings: one without a capacity constraint and one imposing a constraint that at most 70% of the population can be treated.\footnote{I impose a 70% capacity constraint to reflect empirically plausible ITN coverage in Kenya. Large-scale campaigns reached about 67.5% of households with children under five hightower2010bed, while recent survey data show universal coverage of 37% nationally and about 63% in high-risk endemic regions dhs2022kenya. A 70% constraint therefore represents a high-coverage but still resource-constrained setting.} The latter reflects the policy environment studied by bhattacharya2012inferring, where subsidizing bednets imposes budget constraints on local governments, rendering universal treatment infeasible.
Figures (ref) and (ref) report posterior distributions of optimal welfare under NBPL without and with a 70% capacity constraint. In both figures, the left panels plot the empirical CDFs of $W^\star_{\mathcal{G}_{\mathrm{lin}}}(P)$ and $W^\star_{\mathcal{G}_{\mathrm{tree},2}}(P)$ across posterior draws. Without capacity constraints, the distribution for $\mathcal{G}_{\mathrm{tree},2}$ lies only marginally to the right of that for $\mathcal{G}_{\mathrm{lin}}$, indicating limited gains from more adaptive tree-based policies. Under the 70% capacity constraint, however, the shift becomes more pronounced, suggesting larger welfare gains and first-order stochastic dominance. In the unconstrained case, the EWM estimates lie close to the posterior medians for both classes. Under the capacity constraint, by contrast, they correspond to the 29th and 41st percentiles of the posterior distributions under $\mathcal{G}_{\mathrm{lin}}$ and $\mathcal{G}_{\mathrm{tree},2}$, respectively.
The right panels plot the posterior distribution of the welfare difference $W^\star_{\mathcal{G}_{\mathrm{tree},2}} - W^\star_{\mathcal{G}_{\mathrm{lin}}}$. In both cases, the posterior probability that $\mathcal{G}_{\mathrm{tree},2}$ outperforms $\mathcal{G}_{\mathrm{lin}}$ is approximately 86.1%, indicating moderate but consistent evidence in favor of the tree-based class. Although this probability is similar across the two settings, the magnitude of the welfare gains is larger under the capacity constraint, suggesting that constraints increase the value of more adaptive policy classes.
Appendix (ref) provides additional empirical results.
This section presents two theoretical results for NBPL. First, I establish a minimax-optimal convergence rate for posterior welfare regret (Theorem (ref)). Second, I show that this regret convergence implies pointwise consistency for posterior model selection across policy classes ranked by optimal welfare (Theorem (ref)). I conclude by briefly comparing NBPL with empirical welfare maximization (EWM) kitagawa2018should.
Both theoretical results rely on controlling the complexity of the policy class $\mathcal{G}$. Following kitagawa2018should, I assume that $\mathcal{G}$ has finite Vapnik--Chervonenkis (VC) dimension.\footnote{ The Vapnik--Chervonenkis (VC) dimension of a class of sets $\mathcal G \subseteq 2^{\mathcal X}$ measures its complexity in a combinatorial sense. For any finite collection of points $\{x_1,\ldots,x_\ell\} \subseteq \mathcal X$, consider the subsets obtained by intersecting this set with elements of $\mathcal G$. The class $\mathcal G$ is said to shatter $\{x_1,\ldots,x_\ell\}$ if it can pick out all $2^\ell$ possible subsets. The VC dimension is the largest $\ell$ for which some collection of $\ell$ points is shattered by $\mathcal G$ vaart2023weak. }
This assumption covers many tractable policy classes, including linear rules and decision-tree rules, both of which yield interpretable policy prescriptions. At the same time, it does not cover all classes; for example, it excludes monotone rules mbakop2021model.
Following manski2004statistical, manski2009identification, I evaluate a treatment rule by its welfare loss relative to the best feasible policy in the class $\mathcal{G}$.
The regret $R(P_0;P)$ measures the welfare loss from not knowing the true reduced-form distribution $P_0$, and can be viewed as a decision-relevant discrepancy between $P$ and $P_0$. Posterior regret asymptotic analysis asks whether the posterior concentrates on distributions that incur small regret. This parallels posterior contraction in Bayesian models ghosal2017fundamentals, where one fixes a true data-generating parameter and studies whether the posterior concentrates around it.
Theorem (ref) shows that posterior regret contracts to zero at the minimax-optimal rate. To the best of my knowledge, this is the first posterior welfare-regret guarantee for policy learning.\footnote{Although $R(P_0;P)$ is infeasible because $P_0$ is unknown, it nevertheless admits a posterior distribution induced by the posterior for $P$, namely the pushforward measure $\Pi(R \in \cdot \mid \mathcal{D}_n) = \Pi(P \in \cdot \mid \mathcal{D}_n) \circ R^{-1}$.}
Theorem (ref) has two key implications. First, posterior welfare regret under NBPL attains the minimax-optimal rate of empirical welfare maximization (EWM) in kitagawa2018should, up to an arbitrarily slowly diverging factor $C_n$.\footnote{ Such slowly diverging factors are standard in posterior contraction theory and typically arise from testing and entropy arguments ghosal2000convergence, ghosal2007convergence, vaart2008rates, ghosal2017fundamentals. They are generally regarded as inessential and do not affect comparisons of optimal rates; see also rockova2020posterior and zhang2020variational for recent developments. } Thus, incorporating posterior uncertainty does not compromise the rate of learning relative to frequentist approaches. Second, NBPL controls welfare regret over the full posterior distribution: posterior mass concentrates on distributions $P$ for which $R(P_0;P)$ is small. In this sense, NBPL yields a finer form of asymptotic control than EWM, whose regret guarantees concern the expected performance of the single rule selected from the sample.
To derive the rate in (ref), I control the terms of the upper bound \[ R(P_0;P) \le \rho_{\mathcal{G}}(P_0, \mathbb{P}_n)+\rho_{\mathcal{G}}(\mathbb{P}_n, P), \] where $\rho_{\mathcal{G}}(Q_1,Q_2)\coloneqq \sup_{G\in\mathcal{G}}\lvert W(Q_1;G)-W(Q_2;G)\rvert$ measures the maximal welfare discrepancy between $Q_1$ and $Q_2$ over the policy class. The first term, $\rho_{\mathcal{G}}(P_0, \mathbb{P}_n)$, is the usual sampling error. Since $\mathcal{G}$ has finite VC dimension, closeness of the empirical distribution $\mathbb{P}_n$ to the truth $P_0$ yields uniform $n^{-1/2}$ control of welfare over the class. The second term, $\rho_{\mathcal{G}}(\mathbb{P}_n, P)$, is specific to NBPL and reflects posterior uncertainty. Under the Dirichlet process posterior, a draw $P$ is a shrinkage estimator of the empirical distribution: it is a convex combination of a base-measure component and a randomly reweighted empirical distribution. The shrinkage weight on the base-measure component is $O_P(n^{-1})$ and is therefore asymptotically negligible. The remaining empirical component uses normalized exponential weights, which fluctuate around the uniform weight $1/n$ at order $O_P(n^{-1})$. Hence the reweighted empirical distribution remains close to $\mathbb{P}_n$, and finite VC dimension again yields uniform $n^{-1/2}$ control of welfare over $\mathcal{G}$.
More generally, Assumption (ref) can be relaxed to allow a sequence of policy classes $\{\mathcal{G}_n\}_{n \ge 1}$ whose VC dimension grows with the sample size. Provided that $\mathrm{VC}(\mathcal{G}_n) = o(n^{\zeta})$ for some $0 < \zeta < 1/2$, posterior regret still converges to zero at the minimax optimal rate athey2021policy, as shown in the next result.\footnote{For example, the regret bound still converges to zero for linear rules when the covariate dimension satisfies $p_n \le n^{\zeta}$ for some $\zeta < 1/2$, and for decision-tree rules when the maximum depth satisfies $L_n = \lfloor \zeta \log_2 n \rfloor$ for some $\zeta < 1/2$.
Minimax rate optimality follows by combining the upper bound in Theorem 1 with the lower bound in Theorem 5 of athey2021policy.}
The NBPL framework can also quantify uncertainty about the policy class through posterior comparison across classes. Consider two non-nested classes $\mathcal{G}_1$ and $\mathcal{G}_2$, such as linear rules and decision-tree rules, and suppose the goal is to compare them by their optimal population welfare. Under NBPL, the posterior for $P$ induces a joint posterior for $(W_{\mathcal{G}_1}^{\star}(P),W_{\mathcal{G}_2}^{\star}(P))$, so the posterior probability $\Pi(W_{\mathcal{G}_1}^{\star}>W_{\mathcal{G}_2}^{\star}\mid\mathcal{D}_n)$ provides a natural measure of the evidence that $\mathcal{G}_1$ yields higher optimal welfare. To the best of my knowledge, NBPL is the first framework for model comparison across non-nested policy classes in policy learning.
Theorem (ref) provides a pointwise consistency guarantee for this comparison: when the two classes are strictly separated at the truth, the posterior probability of assigning the correct sign to the welfare difference converges to one, and hence the posterior probability of selecting the inferior class converges to zero. In the equality case, the posterior instead concentrates on arbitrarily small welfare differences, ruling out any fixed non-negligible gap asymptotically.
The proof of Theorem (ref) draws on intermediate steps from the proof of the regret convergence rate. If the two classes are strictly separated at the truth, an incorrect posterior ranking can occur only if, for at least one class $\mathcal{G}_i$, the optimal welfare under $P$ deviates sufficiently from its value under $P_0$. Since $|W_{\mathcal{G}_i}^{\star}(P)-W_{\mathcal{G}_i}^{\star}(P_0)| \le \rho_{\mathcal{G}_i}(P,P_0)$, such a mistake requires at least one discrepancy term $\rho_{\mathcal{G}_i}(P,P_0)$ to be non-negligible. The control established in the proof of Theorem (ref) then implies that posterior mass on such draws vanishes asymptotically. The knife-edge case is handled similarly.
As a corollary of Theorem (ref), the posterior probability of correctly ranking any two fixed treatment rules converges to one in $P_0$-probability as $n\to\infty$. This applies, for example, when a policymaker compares competing rules proposed by different analysts, or benchmarks a proposed targeting rule against the status quo of assigning no treatment. The result follows by applying Theorem (ref) to the singleton classes $\mathcal{G}_1=\{G_1\}$ and $\mathcal{G}_2=\{G_2\}$.
More generally, both Theorem (ref) and Corollary (ref) remain valid---hence the Bayesian model selection procedure is consistent---when $\mathrm{VC}(\mathcal{G}_n) \le n^{\zeta}$ for some $0 < \zeta < 1/2$.
This section relates empirical welfare maximization kitagawa2018should to the Nonparametric Bayesian Policy Learning (NBPL) framework. EWM is an as-if optimization rule manski2021econometrics: it treats the empirical distribution as the true population distribution and selects
By contrast, NBPL first represents uncertainty about the reduced-form distribution through a posterior. This posterior can then be propagated through the optimization map $P\mapsto G^\star(P)$, yielding a posterior distribution over optimal treatment assignments and optimal welfare.
For any treatment rule $G$, define the welfare-regret loss at distribution $P$ by \[ L(P;G)\coloneqq W_{\mathcal G}^\star(P)-W(P;G). \] kitagawa2018should study the frequentist risk \[ \mathcal R_n(P_0) \coloneqq \mathbb E_{P_0}[ L(P_0;\widehat G_{\mathrm{EWM}})], \] and show that it converges to zero at the minimax-optimal rate $\sqrt{v/n}$ uniformly over a class of data-generating processes.
The next result gives EWM a Bayesian decision-theoretic interpretation.
Proposition (ref) places EWM within the NBPL framework. Under the limited-prior-informativeness Dirichlet process posterior $\mathrm{DP}(n\mathbb P_n)$, EWM is exactly the Bayes rule that minimizes posterior expected welfare regret. It is therefore an average-then-optimize rule: EWM first averages welfare-regret over posterior uncertainty and then selects a treatment rule. By contrast, NBPL is an uncertainty-aware treatment choice framework that retains the full posterior distribution over optimal treatment rules and welfare. Decisions obtained by first optimizing under each posterior draw and then aggregating the induced posterior over optimal treatment assignments need not coincide with EWM.
This paper studies a decision-maker (DM) seeking to select an expected welfare-maximizing treatment rule using observable characteristics. The key observation is that, for a given welfare criterion and policy class, the sole source of uncertainty in the DM's problem is statistical uncertainty about a reduced-form distribution. I therefore propose Nonparametric Bayesian Policy Learning (NBPL), which places a nonparametric Dirichlet process prior on this reduced-form distribution and use the resulting posterior to conduct simultaneous inference on optimal treatment assignments, optimal welfare, and comparisons across policy classes. NBPL provides a natural, flexible, and computationally tractable framework for uncertainty-aware treatment choice. I show that posterior welfare regret converges at the minimax-optimal rate and that posterior model comparison across policy classes is pointwise consistent.
More generally, transparent uncertainty communication in policy analysis must account for both statistical uncertainty and ambiguity manski2013public,manski2019communicating,manski2025discourse, arising from limited data and unsupported assumptions, respectively. Under ambiguity, expected welfare may be only partially identified. NBPL provides a nonparametric Bayesian approach to statistical uncertainty in policy learning, complementing existing work on uncertainty-aware treatment choice chamberlain2011bayesian,chernozhukov2025policy, christensen2025optimal, moon2026optimal. Developing a framework that balances statistical uncertainty and ambiguity is an important direction for future research. Another important extension is to allow for unknown propensity scores.