EconBase
← All papers

Stable Policy Learning

Harvey Barnhard, Giacomo Opocher, Rahul Singh

arXiv 16 Sep 2026 · Econometrics

arXiv:2609.19418 · PDF · Extracted main text

Abstract

In evidence-based policymaking, typically one experimental sample is observed, then a learned policy recommendation is implemented at scale. Policies learned from the experimental data can perform well in expected welfare, yet random sampling in the experiment can produce recommendations with poor welfare outcomes. In this paper, we ask: how should policy learning algorithms balance expected welfare against sampling risk? Our main contribution is to show that algorithmic stability plays a central role in characterizing and navigating the tradeoff. Intuitively, if a policy learning algorithm's recommendation remains stable when one experimental unit is replaced, then that algorithm has limited sampling risk. We propose a method for policy learning called policy-vote bagging, which learns treatment decisions on many subsamples then averages their votes into treatment probabilities. Relative to using one subsample, averaging across subsamples preserves expected welfare and improves expected utility for a risk-averse researcher. We derive sharp bounds linking estimation accuracy, subsample size, and welfare variation, including an exact guarantee under CARA utility.

Citation extraction

54
references
69
in-text mentions
50
distinct cited
2
self-citations
10,555
main-text words

appendix boundary found by appendix_command · 40% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1L. Breiman (1996) Bagging predictors0.64422100%
2Q. Chen, V. Syrgkanis, and M. Austern (2022) Debiased machine learning without sample-splitting for stable estimators0.64422100%
3V. Chernozhukov, W. K. Newey, R. Singh, and V. Syrgkanis (2026) Adversarial estimation of Riesz representers self0.64422100%
4T. Denti and L. Pomatto (2022) Model and predictive uncertainty: A foundation for smooth ambiguity preferences0.64422100%
5A. Elisseeff, T. Evgeniou, and M. Pontil (2005) Stability of randomized learning algorithms0.64422100%
6T. Kitagawa, S. Lee, and C. Qiu (2026) Treatment choice with nonlinear regret0.64422100%
7T. Kitagawa and A. Tetenov (2018) Who should be treated? Empirical welfare maximization methods for treatment choice0.64422100%
8T. Kitagawa and A. Tetenov (2021) Equality-minded treatment choice0.64422100%
9P. Klibanoff, M. Marinacci, and S. Mukerji (2005) A smooth model of decision making under ambiguity0.64422100%
10F. Maccheroni, M. Marinacci, and D. Ruffino (2013) Alpha as ambiguity: Robust mean-variance portfolio analysis0.64422100%

Showing the top 10 of 50 scored citations.