Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
43,326 characters · 9 sections · 66 citation commands
Policy Learning with Confidence$^$
\noindentKeywords: budget allocation, risk-aware policy learning, statistical decision theory
\noindentJEL classification codes: C14, C44, C52.
Consider a decision maker (DM) faced with choosing from a menu of policies to maximize expected welfare. The DM could be a planner deciding how to allocate treatments to heterogeneous individuals, a firm choosing different potential innovations in which to invest resources, an auctioneer choosing an optimal auction design or reserve price, or a policy-maker choosing expenditure shares to allocate to government programs to maximize a measure of public welfare. If the DM knew the welfare that would be obtained from implementing each policy, the choice would be easy; she would then simply choose the policy that would yield the highest welfare. However, the DM lacks such precise knowledge, but instead has data from one or possibly multiple studies that can be used to consistently estimate the welfare that would be achieved by each policy. The estimates are measured with varying degrees of precision, as reflected by their standard errors. How should the DM use the available information for policy selection?
An intuitively appealing choice is to select the policy with the highest estimated welfare, the so-called plug-in rule, which simply replaces the population objective with its sample analog in determining policy choice. In the literature on policy learning, this is referred to as empirical welfare maximization (EWM).\footnote{See for example KT2018 and AW2021. andrews2024inference alternatively refers to this as the “natural rule” or “picking the winners”.} The same logical task applies more broadly to a variety of economic contexts, for example when researchers use structural models, they obtain estimates of the performance of different policies that can then be used to inform policy choice. Manski2021 calls this “as-if” optimization, described as, “specification of a model, point estimation of its parameters, and use of the point estimate to make a decision that would be optimal if the estimate were accurate.” Here we refer to such rules interchangeably as plug-in rules or empirical welfare maximization (EWM) rules.
Unfortunately, the available estimates for the welfare of different choices are generally not identical to their population values. Estimation error could result in the EWM rule delivering suboptimal policy choice. It may be observed, for example, that among a menu of options policy A has the highest estimated welfare with policy B coming in a close second, but that the standard error of policy B is much smaller than that of policy A, reflecting that the welfare of policy B is estimated much more precisely. Could it be that policy B is actually better than policy A, and that the higher estimated welfare of policy A is simply down to a lack of precision in the estimates? Should the DM consider choosing policy B since it is estimated more precisely to hedge against estimation error?
This paper proposes a rule for selecting policies with the goal of maximizing expected welfare in the presence of estimation uncertainty, explicitly accounting for estimation risk in the decision rule. As inputs to her decision the DM has consistent estimates for the welfare of each available policy, and their corresponding sample variance matrix, with sample estimates approximately normally distributed around their population means.\footnote{If the estimates are from independent samples the sample variance matrix will simply be a diagonal matrix with each policy's sample variance along the diagonal. } We define and analyze a class of risk-aware policy rules that provide a principled manner for explicitly balancing the estimated welfare of each policy against its estimation risk, and which we show have favorable regret properties. Risk-aware rules select a policy from the (Pareto) efficient decision frontier that balances performance and precision as measured by trading off higher sample estimates against smaller sample variance, where the exact tradeoff is governed by the DM's choice of policy rule from the class of risk-aware rules. Our newly proposed rule is determined from within this class by balancing this tradeoff using the tangent line to the frontier whose slope is determined by the critical value of a one-sided upper confidence band for welfare following the approach for intersection bound inference developed in CLR2013. Consequently, this rule delivers a reporting guarantee, ensuring with high confidence that the actual welfare delivered exceeds a lower threshold, while also having the favorable regret properties of all risk-aware rules. The proposed rule thus ensures Policy Learning with Confidence, and is subsequently referred to as the PoLeCe rule.
To illustrate, consider a setting in which a budget-conscious DM needs to allocate funds across several social programs that affect the same group of individuals, either in response to a slight budget increase or for incremental cost-cutting. Suppose the DM wishes to use the Marginal Value of Public Funds (MVPF), spearheaded by Hendren/Sprung-Keyser:2020, for this purpose. The MVPF of each social program is the ratio of its marginal benefit to its net marginal cost to the government, inclusive of the impact of any behavioral responses on the government budget. For instance, an MVPF of 1.5 indicates that each \$1 of net government spending generates \$1.50 in benefits for beneficiaries. For a utilitarian DM, the MVPF provides a metric to compare the “bang for the buck” of different policies that target the same beneficiaries. A DM interested in allocating additional marginal expenditure will prioritize programs with the highest MVPF to maximize welfare.\footnote{Here we abstract from distributional incidence and assume the DM values dollars equally across the beneficiaries of each policy. If welfare weights vary across beneficiaries, as in the original framework of Hendren/Sprung-Keyser:2020, our analysis can accommodate this by adjusting each MVPF in the value function using the DM's designated weights. Example welfare weights have been derived in hendren2020measuring. Other metrics for measuring programs' marginal benefits and costs have also been discussed in garcia2022three.}
In practice the DM must make decisions based on estimated MVPFs. The Policy Impacts Library, policy_impacts_library, provides easily accessible MVPF estimates online while also noting significant estimation uncertainty in some cases. Using the EWM rule to allocate funds would direct the entire budget to the program with the highest MVPF estimate. However, the DM may be reluctant to direct the entire budget to a single social program if this MVPF estimate has a large standard error relative to other programs whose MVPF estimates are nearly as high, and which have smaller standard errors. How should the DM balance the tradeoff between the size of the estimated MVPF and the precision with which it is estimated?
This tradeoff is illustrated in Figure (ref) using the MVPF estimates from the Policy Impacts Library, which comprises 14 US programs after we focus on estimates based on randomized control trials (RCTs) with finite standard errors. To mimic a DM who applies constant welfare weights within certain age groups but not necessarily across them due to distributional concerns, we use the Policy Impacts Library classification of programs based on intended beneficiaries' age. We then consider budget allocations separately for two distinct age groups. Six of these programs target beneficiaries age 25 and under, covering early childhood education programs, college financial aid, and job training.\footnote{As explained in the Policy Impacts Library, alternative specifications have led to larger MVPF estimates for early childhood education depending on how lifetime earnings are forecasted and whether one accounts for the transfer value of preschool subsidies to parents.} Eight of these programs target beneficiaries age 25 and above, covering federal social assistance, health insurance, and housing vouchers. In both cases, the MVPF of the programs selected by EWM are not estimated as precisely as the welfare delivered by diversifying the budget across several programs whose MVPFs are estimated with significantly higher precision. If a DM values both welfare and precision, how should these goals be balanced?
The PoLeCe rule is determined by the intersection of the green line in each panel of Figure (ref) with the decision frontier for RW-rules. The rule maximizes estimated welfare offset by a data-dependent penalization factor such that the resulting objective value automatically provides the lower bound of a one-sided confidence interval for the welfare achieved by the selected rule. Thus, the rule delivers a reporting guarantee, ensuring with high confidence that the actual welfare exceeds a lower threshold; no adjustment for post-selection inference is required. Importantly, policies that allocate fractional shares to different treatments or programs can be allowed, producing a richer frontier than that available from singleton allocations such as the individual programs indicated by hollow squares in Figure (ref).
In the above example, all that was required to map out the efficient decision frontier, and indeed all that is needed to determine the optimal allocation following our proposal, are point estimates for the welfare of each available policy and their joint variance-covariance.\footnote{Our proposed PoLeCe rule uses the sample correlation of the estimates to calibrate the precision penalty to achieve the aforementioned reporting guarantee. The general class of risk-aware decision rules described in Section (ref) allows for different choices for the penalization factor, which corresponds to the green slope highlighted in Figure (ref). Other risk-aware rules may thus only require standard errors rather than estimates of the entire correlation structure, depending on their choice of penalization factor.} This can be obtained in applications that feature a variety of different models and sampling processes. An important special case to which our framework applies is the analysis of optimal treatment assignment using data on individual-specific treatments and allocations, such as from an RCT. This is demonstrated in online Appendix (ref) with three additional applications spanning the study of treatments that encourage immunization, enable informal savings technologies, and encourage sobriety among low-wage workers. In a real-world setting of treatment assignment, while the nonprofit GiveWell uses EWM based on RCTs, they have recently highlighted the importance of accounting for estimation uncertainty in decision-making for the sake of “transparency" salisbury_how_2024. Our proposal offers a practical solution to this concern, while also applying to settings in which individual-specific treatment responses may not be directly observable, and when estimates may be obtained from different sources.
Prescriptions for decision making with sample data include conditional Bayes rules, maximin rules, and minimax regret rules. These are summarized in Manski2021, which advocates the use of statistical decision theory from the frequentist perspective of Wald:50. Minimax regret rules in particular have recently received renewed attention, starting with the pioneering work of Manski2004. Follow-up work includes a decision theoretic framework introduced by dehejia_program_2005, finite-sample bounds considered by stoye_minimax_2009, and an asymptotic framework introduced by hirano_asymptotics_2009,Hirano/Porter:Handbook. While the use of statistical decision theory for selecting policy rules summarized in Manski2021,Manski:2023 is conceptually appealing, computational challenges remain, particularly in settings involving samples from multiple data sources as in the preceding example. Recent papers on policy learning provide approaches for treatment rule estimates with favorable asymptotic regret properties. KT2018 prove the optimality of EWM rules in the sense that as the sample size increases expected regret converges to zero at the minimax rate. mbakop_model_2021 proposes a penalized welfare maximization rule that penalizes the complexity of each policy, and establishes an oracle property for model selection. AW2021 and DSLC2019 propose doubly-robust estimation of average welfare, which leads to an optimal rule even with quasi-experimental data. The rule proposed here provides a practical approach to providing a decision rule with favorable regret properties while also yielding a confidence interval for welfare. As such, the PoLeCe rule fits within the class of P-certified decision rules studied by andrews2025certifieddecisions, a point we come back to in Section (ref) where we discuss the comparison to other recently proposed rules and inference approaches.
The rest of the paper proceeds as follows. Section (ref) sets out the framework and defines the class of risk-aware (RW) decision rules. Section (ref) establishes high probability regret bounds for RW-rules in the spirit of regret properties analyzed for EWM rules by e.g. KT2018 and AW2021. Section (ref) sets forth the PoLeCe rule among the class of RW decision rules and formally provides its dual regret and coverage guarantees. Section (ref) discusses the general class of P-certified decisions and other rules and inference approaches proposed in the recent literature. Section (ref) returns to the context of a planner deciding on the allocation of additional public funds across different social programs and presents the full application of our approach to this empirical setting. Section (ref) concludes. Proofs of propositions and details of the Gaussian bootstrap used to compute the PoLeCe penalization factor are provided in the appendices. Additional results including alternative motivations for risk-aware decisions, details for implementation and construction of the efficient decision frontier, three empirical applications to treatment choice using data from RCTs, and computational experiments calibrated to those applications are included in the online supplementary appendices.
We begin with the following key problem as motivation. A public or private agency is considering programs \(\{1,...,J\}\) and must decide on investment shares \[ \pi = (\pi_1,..., \pi_J), \] subject to \[ \pi \in \Pi := \Bigl\{\,(\pi_1,\dots,\pi_J): \ \sum_{j=1}^J \pi_j =1,\ 0 \leq a_j \leq \pi_j \leq b_j \leq 1\Bigr\}, \] which is the unit simplex intersected with a rectangular set of constraints. These additional constraints may reflect diversity or other requirements. The welfare of allocation \(\pi\) is \[ V(\pi) = \pi' R, \] where \(R\) is a vector of the rate of return measures, for instance the ratios of marginal benefit to net cost of funds used in public finance. The agency is given estimates \(\widehat R\) that are approximately Gaussian\footnote{In the empirical application we consider, \(\Omega\) is block-diagonal, and \(n\) is the notional sample size used to study the behavior of rules as more information is acquired.} \[ \widehat R \;\overset{a}{\sim}\; N(R, \Omega/n), \] and the estimated welfare of the allocations are also approximately Gaussian by the continuous mapping theorem:
where \(\{Z_\pi\}_{\pi \in \Pi}\) is a Gaussian vector with standard normal marginals. Here \(\overset{a}{\sim}\) means “approximately distributed as” formalized in condition (ref) below.
Hence, for each fund allocation \(\pi\), the agency is given \(\widehat V(\pi)\) along with its associated estimation risk \(s(\pi)\). What should the agency do? Our proposal is to use the risk-aware rule that maximizes empirical welfare offset by the estimation risk times a critical value:
where \(k\) is a critical value and \(\widehat{s}(\pi)\) is a consistent estimator of the estimation risk \(s(\pi)\). By varying \(k\), we trace out the efficient decision frontier of treatment policies, as illustrated by Figure 1 in the introduction. Each point on the frontier corresponds to a particular \(k>0\). Appendix (ref) in the online supplement provides an efficient algorithm to compute the frontier for any application.
Our leading proposal to choose \(k\) is to meet certain reporting guarantees with high confidence, thereby conducting “policy learning with confidence” (PoLeCe). The resulting choice generally differs from the empirical welfare maximizer (EWM):
In the budget allocation problem, the EWM rule would simply allocate all funds to the program with the highest estimated return. This is generally unappealing on intuitive grounds. In contrast, the risk-aware approach ((ref)) would spread out funds over an “efficient” portfolio of programs, similar to Markowitz portfolio allocation in finance. However, the motivation and formulation of our approach are distinct from the Markowitz model.
A focal aim in the related literature on treatment choice has been providing decision rules with favorable regret properties. Adopting regret as a benchmark, we now provide regret bounds for any risk-aware decision rule \(\widehat \pi_{\text{RW}}(\widehat k)\) that solves the risk-adjusted empirical welfare problem: \[ \widehat \pi_{\text{RW}}(\widehat k) \;\in\; \arg \max_{\pi \in \Pi}\bigl\{\widehat V(\pi) \;-\; \widehat k \,\widehat s(\pi)\bigr\}, \] where \(\widehat k \ge 0\) may depend on both the decision maker's preferences and the data. Note that \(\widehat k = 0\) corresponds to the EWM rule.
Define \[ \widehat{Z}_\pi \;:=\; \frac{\widehat{V}(\pi) - V(\pi)}{\widehat{s}(\pi)} \] to be the normalized estimation error process. We use the following Gaussian approximation condition, denoted by (G). Suppose the policy class $\Pi$ can be well-approximated by a $p$-dimensional discretization. Let \(\mathcal{A}\) represent the collection of rectangular sets in \(\mathbb{R}^p\).
These conditions are known to hold under mild assumptions.\footnote{See CCKK-AOS,CCK-AOP for sharp forms of Gaussian approximations, and CCKK-Review for a review. Additional references, e.g. belloni2018uniformly,quintas2022finite discuss Gaussian approximations for many or a continuum of target parameters, including those learned via debiased machine learning.} Furthermore, (ref) implicitly requires consistency of risk estimates: \(\max_{\pi \in \Pi}\lvert \widehat s(\pi)/s(\pi) - 1\rvert \to_P 0\).
Quantiles of the Maximal Estimation Error. For a subset \(K \subseteq \Pi\), let \[ q_{1-\beta, K} \;:=\; (1-\beta) \text{-Quantile} \Bigl(\sup_{\pi \in K} Z_{\pi}\Bigr) \] denote the \((1-\beta)\)-quantile of the estimation error over \(K\) under the Gaussian approximation. Under (ref), with probability at least \(1 - \beta - r_n\), \(\max_{\pi \in K}\widehat{Z}_\pi \le q_{1-\beta,K}\). Let \(V_{\max} = \sup_{\pi\in\Pi}V(\pi)\) be the maximal true welfare and \(\Pi_0 = \{\pi \in \Pi : V(\pi)= V_{\max}\}\) be the set of policies with maximal true welfare, with cardinality \(p_0 := |\Pi_0|\). Define \[ \bar{\sigma}_{\Pi} \;:=\; \sqrt{n}\,\max_{\pi \in \Pi} \widehat s(\pi), \quad \underline{\sigma}_{\Pi_0} \;:=\; \sqrt{n}\,\min_{\pi \in \Pi_0} \widehat s(\pi). \] These represent, respectively, an upper bound on the estimation risk of all policies in \(\Pi\) and a lower bound on the estimation risk of the best policies \(\Pi_0\).
As \(\beta\) decreases, the probability that the regret bound (ref) of Proposition (ref) holds increases, since the failure probability is at most \(2\beta + 2r_n\). However, this comes at the cost of a looser bound because the quantile \(q_{1-\beta, \Pi}\) increases as \(\beta\) gets smaller. This reflects a basic trade-off: a smaller \(\beta\) provides a more conservative bound that holds with higher probability, while a larger \(\beta\) results in a tighter bound that is less certain to be valid.
Proposition (ref) provides dimension-free bounds. To obtain simpler dimension-based bounds, observe that by standard concentration inequalities,\footnote{This follows from Lemma A.8 in BCCHK:18. Talagrand-style upper bounds can also be used, but for this these simpler concentration bounds suffice.} \[ q_{1-\beta,K} \;\;\le\;\; \mathrm{E}\Bigl(\sup_{\pi \in K} Z_\pi\Bigr) \;+\; \sqrt{2\log(1/\beta)} \;\;\le\;\; \sqrt{2\log |K|} \;+\; \sqrt{2\log(1/\beta)} \;=\; u_{1-\beta, |K|}. \] We now see that the broad class of RW-rules has regret bounded by \[ \bar{\sigma}_{\Pi}\,\sqrt{\frac{\log p}{n}}, \] as \(p \to \infty\) (and \(n \to \infty\)). This class includes the EWM rule (\(k = 0\)) and matches known optimal minimax rates when available (see AW2021).
A key observation of Proposition (ref) is that choosing \(\widehat{k} > 0\) sufficiently large can yield a regret bound of \[ \underline{\sigma}_{\Pi_0}\,\sqrt{\frac{\log p}{n}}, \] which has the same rate but may have a much smaller multiplicative constant -- $$\underline{\sigma}_{\Pi_0} < \bar \sigma_{\Pi}$$ if there is significant heterogeneity in estimation risk across policies. The leading constant in the regret bound is governed by the least risky among the best policies, and can therefore be significantly smaller than that of the EWM rule. The empirical risk minimization literature has also used penalties on estimation risk to improve error bounds, albeit in different contexts; see, for example, maurer2009empirical, swaminathan2015counterfactual, foster2023orthogonal.
These ideas naturally lead to the next section, where we propose a main policy rule that carefully constructs a data-driven \(\widehat{k} \approx q_{1-\alpha, \Pi}\) to achieve small regret and provide additional reporting guarantees.
In practice, to obtain reporting guarantees for RW decisions, one needs to approximate \(q_{1-\alpha,\Pi}\). We do so via the bootstrap, where we set: \[ \widehat{q}_{1-\alpha,\Pi} \;:=\; (1-\alpha)\text{-Quantile} \Bigl(\max_{\pi \in \Pi} \widehat{Z}_\pi^*\Bigr), \] where \((\widehat{Z}_\pi^*)_{\pi \in \Pi} \sim N\bigl(0,\widehat{C}\bigr)\), and \(\widehat{C}\) is a consistent estimator of the covariance matrix \(C = \mathrm{Cov}((Z_\pi)_{\pi \in \Pi})\). Denote by \(\Pr^*\) the bootstrap-induced probability measure computed conditional on \(\widehat{C}\). This construction relies on the following condition:
Like (ref), condition (ref) is satisfied under mild assumptions even when \(p\) is much larger than \(n\), and has been verified for a variety of estimation methods, including debiased machine learning.\footnote{See CCKK-AOS,CCK-AOP and CCKK-Review for more details, along with other references on “approximate means” settings such as debiased ML (belloni2018uniformly,quintas2022finite).}
A key consequence is that, letting \(r_n' := 2r_n + \delta_n\), we have
(see Lemma (ref) for a proof). Inequality (ref) then directly implies a uniform lower confidence bound (LCB) on the performance of all policies:
We now introduce policy learning with confidence (PoLeCe), a risk-aware rule that provides a key reporting guarantee by maximizing the LCB on policy welfare:
Notably, PoLeCe is an RW decision that uses the bootstrap quantile $$\widehat{k} = \widehat{q}_{1-\alpha,\Pi} $$ as the penalty for estimation risk. Furthermore, the maximized LCB \[ LV_{1-\alpha,\Pi} \;:=\; \max_{\pi \in \Pi} LV_{1-\alpha}(\pi) \;=\; \max_{\pi \in \Pi} \Bigl\{ \widehat{V}(\pi) \;-\; \widehat{q}_{1-\alpha,\Pi} \,\widehat{s}(\pi) \Bigr\} \] provides a natural performance guarantee since $V\left(\widehat{\pi}_{PoLeCe}\right) \geq LV_{1-\alpha,\Pi}$ with probability at least $1-\alpha - r_n^{\prime}$ by Proposition (ref).
The following result establishes formal properties of the PoLeCe rule.
The first result in Proposition (ref) shows that $LV_{1-\alpha, \Pi}$ provides a valid high-confidence lower bound on the true welfare of the PoLeCe rule. The second result sharpens the general regret bounds presented in Proposition (ref). As noted following Proposition (ref), the parameter $\beta$ governs the trade-off between failure probability and bound tightness, while the constant $\alpha$, chosen by the researcher, determines the confidence level of the lower bound $LV_{1-\alpha, \Pi}$. The refinement in Proposition (ref) depends on the relationship between $\alpha$ and $\beta$. When $\alpha = \beta$, Propositions (ref) and (ref) yield qualitatively identical guarantees. When $\alpha < \beta$, the bounds in both propositions remain qualitatively identical, but the bound in Proposition (ref) holds with higher probability. When $\alpha > \beta$, the bound in Proposition (ref) is tighter, but it may hold with lower probability. For a large $\alpha$, the bound may fail to hold; in such cases, the general bound from Proposition (ref) remains valid and informative, as it applies to all risk-aware decision rules.
The PoLeCe rule maximizes $LV_{1-\alpha}(\pi)$, the welfare estimate of each policy $\pi$, penalized by the critical value $\widehat{q}_{1-\alpha,\Pi}$ times its standard error $\widehat{s}(\pi)$. The function $LV_{1-\alpha}(\cdot)$ provides a $1-\alpha$ lower confidence band for $V(\cdot)$ following the construction of CLR2013. Thus, the PoLeCe rule can be obtained by minimizing worst-case loss over a $1-\alpha$ lower confidence band for welfare. As such, it falls within the class of rules Manski2021 calls “as-if optimization with set estimates”. Section 3.2 of Manski2021 suggested that one might consider decision making using confidence sets, but left this to future research.
The contemporaneously developed working paper andrews2025certifieddecisions introduces the term P-certified decisions for the class of rules that ensure an upper bound on loss -- equivalently a lower bound on welfare -- with probability at least $1 - \alpha$. andrews2025certifieddecisions argue that such decisions are useful for risk-averse decision makers and show that “as-if” decisions using confidence sets form an essentially complete class of P-certified decisions. A similar finding is made in Kiyani/Pappas/Roth/Hassani:25, which studies rules that maximize worst-case loss over conformal prediction sets. Within this class, those that comprise upper contour sets of conditional quantiles of utility are found to be optimal. Use of these sets for as-if optimization with set estimates differs from the use of upper confidence bands that corresponds to the PoLeCe rule. Ben-Michael/Greiner/Imai/Jiang:25 also propose a P-certified decision rule for criminal release recommendations. Their rule again differs from ours, as it is based on a two-sided confidence set for expected utility.\footnote{The setting studied by Ben-Michael/Greiner/Imai/Jiang:25 has other distinguishing features, for example using data based on deterministic functions of pre-trial risk assessment scores, and the need to take on the challenge of partially identified policy welfare.} While the PoLeCe rule falls within the class of P-certified decisions studied in andrews2025certifieddecisions, its P-certificate is based on confidence bands of the type studied in CLR2013 and it thus differs from the other policy rules studied within this class. A convenient feature of the PoLeCe rule is that it is delivered directly by maximizing $LV_{1-\alpha}(\pi)$ in a single stage without requiring explicitly computing a maximin rule over a set estimate. Furthermore, it provides a rule that lies on the efficient decision frontier with the balance between estimated performance and sample variation pinned down by the critical value corresponding to the planner's desired confidence level.
P-certified decisions are also connected to the recent literature on inference on winners, e.g. BenjaminiSelected, andrews2024inference, Zrnic/Fithian:24,Zrnic/Fithian:25. These approaches provide confidence sets for the welfare of the selected choice that has the highest empirical welfare in-sample. Inference approaches for optimal assignment and/or welfare in the population include those in the online supplement of KT2018, Rai2019, Armstrong/Shen:2023, and ponomarev2024lowerconfidencebandoptimal. Proposition (ref) shows that the PoLeCe rule provides a coverage guarantee for both the welfare of the selected policy and the optimal rule.
Recent alternatives that account for sampling uncertainty in policy selection include Sun2021 and Moon:25. Sun2021 considers policy selection when the decision maker faces a budget constraint and there is uncertainty regarding the cost of the policy. Moon:25 proposes an empirical Bayes approach to deal with statistical uncertainty in the welfare of different policies. The approach in this paper differs from each of these.
This section continues the discussion from the introduction, and illustrates how to solve the problem of investment allocations for public programs described in Section (ref). We obtain MVPF estimates $\widehat{R}$ from the Policy Impacts Library, policy_impacts_library. There are 172 programs in total, and we focus on those in the United States whose MVPF is estimated by an RCT. We assume the reported upper and lower bounds correspond to 95% confidence intervals and infer the standard errors for each $\widehat{R}$ from these bounds, yielding $14$ programs.\footnote{We exclude programs for which the upper and lower bounds are infinite.} We also assume the estimation errors across programs are independent, so their covariance matrix is diagonal, with the squared standard errors on the diagonal. Note that here we only account for estimation error, holding the specification in the original papers that produced these MVPF estimates fixed. Alternative specifications can lead to different MVPF estimates for the same program. For example, accounting for transfers to parents could lead to larger MVPF estimates for early childhood programs.
In the following illustrations, we implement PoLeCe by reformulating the problem as a root-finding task that involves solving second-order cone programs, which has superior computational efficiency relative to grid search as detailed in Appendix (ref). We set $\alpha=0.05$ throughout.
Table (ref) reports the allocation selected by PoLeCe and EWM among programs targeting individuals older than 25. As expected, the EWM rule is a corner solution that places all investment in the highest MVPF program according to the empirical estimates, Holistic Wrap-around Services Can Improve Employment Rates. The PoLeCe rule, in contrast to EWM, places weight on Medicaid for single adults (OHIE, Single Adults) and job training for adults (JTPA) as their MVPF estimates are among the highest, and are highly precise.
Table (ref) reports the allocation selected by PoLeCe and EWM among programs targeting individuals aged 25 and under. Compared to allocations for programs for those over age 25 in Table (ref), the PoLeCe solutions here allocate a large share towards programs whose MVPF estimates are less precise. This is because $\widehat{q}_{0.95, \Pi}$ is smaller, resulting in less aversion to estimation uncertainty.
In this paper we have focused on the problem faced by a DM who wishes to select from a menu of policies to maximize welfare using imperfect sample estimates. In order to balance the estimated performance of each policy with its associated statistical uncertainty, we proposed and analyzed the properties of a class of risk-aware policies that make the precision/performance tradeoff explicit. Such policies achieve favorable regret properties. We proposed a specific rule from the class of risk-aware policies, namely the PoLeCe rule, which uses a data-dependent construction for balancing the inherent tradeoff between estimated performance and sample uncertainty, such that a lower confidence bound on the welfare obtained by the chosen policy is automatically provided.
A large body of work synthesized in Manski (2013, 2019, 2024) has advocated for greater acknowledgement and incorporation of uncertainty in planning problems.\nocite{Manski:PubPolUncertain}\nocite{Manski:2019}\nocite{Manski:Discourse} This paper contributes by proposing a principled way to acknowledge and incorporate statistical uncertainty into decision making. Statistical uncertainty is however only one of the many different types of uncertainty that may be present.\footnote{For example, Manski:2019 discusses transitory uncertainty, permanent uncertainty, and conceptual uncertainty. The statistical uncertainty considered here is one type of permanent uncertainty.} We have focused on settings in which the welfare of each policy is point identified, such that consistent estimates of the welfare and sample variance of the policies are available. Application of the concepts developed here to settings that feature so-called “deep uncertainty”, or ambiguity, would seem an important direction for further development given the large number of settings in which the mean performance of various policies is only credibly partially identified.\footnote{The literature on treatment choice when mean treatment performance is partially identified goes back to Manski:2000. Recent research with references to the broader literature includes Russell:20, Ishihara/Kitagawa:21, Yata2021, Christensen/Moon/Schorfheide:23, Kido:23, Olea/Qiu/Stoye:23, and Ben-Michael/Greiner/Imai/Jiang:25.} This would require balancing statistical uncertainty with a collection of interval estimates for each policy's performance.
While the nuance required to formally develop such an approach is beyond the scope of the present paper, a rough prescription could be made based on an extension of the reporting guarantee established here. To see how, consider the definition of the PoLeCe rule in (ref). The rule selects the policy that maximizes a $1-\alpha$ lower confidence band for the maximal achievable welfare. Construction of such a confidence band does not require that the optimal level of welfare or the optimal rule be point identified, and could be implemented using techniques developed in e.g. CLR2013 under partial identification. This is of course not the only way to balance statistical uncertainty and ambiguity. A more thorough study of the performance of such an approach, and the implicit tradeoff between statistical uncertainty and ambiguity due to partial identification would seem a useful direction for future research.