EconBase
← All papers

General Bayesian Policy Learning

Masahiro Kato

arXiv 27 Feb 2026 · Statistics — Machine Learning

arXiv:2602.23672 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This study proposes the General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action from an action set to maximize its expected welfare. Typical examples include treatment choice and portfolio selection. In such problems, the statistical target is a decision rule, and the prediction of each outcome $Y(a)$ is not necessarily of primary interest. We formulate this policy learning problem by loss-based Bayesian updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class is equivalent to minimizing a scaled squared error in the outcome difference, up to a quadratic regularization controlled by a tuning parameter $ζ>0$. This rewriting yields a General Bayes posterior over decision rules that admits a Gaussian pseudo-likelihood interpretation. We clarify two Bayesian interpretations of the resulting generalized posterior, a working Gaussian view and a decision-theoretic loss-based view. As one implementation example, we introduce neural networks with tanh-squashed outputs. Finally, we provide theoretical guarantees in a PAC-Bayes style.

Citation extraction

54
references
90
in-text mentions
54
distinct cited
2
self-citations
7,844
main-text words

appendix boundary found by appendix_command · 48% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Pier Giovanni Bissiri, Chris C. Holmes, and Stephen G. Walker (2016) A general framework for updating belief distributions0.7948450%
2O. Catoni, Project Euclid, Cornell University. Library, and Duke Uni… (2008) PAC-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning0.7375340%
3Zhengyuan Zhou, Susan Athey, and Stefan Wager (2023) Offline multi-action policy learning: Generalization and optimization0.7374350%
4Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.6443267%
5Pierre Alquier (2024) User-friendly introduction to pac-bayes bounds0.5853333%
6Maxime Haddouche, Benjamin Guedj, Omar Rivasplata, and John Shawe-Ta… (2021) Pac-bayes unleashed: Generalisation bounds with unbounded losses0.5853333%
7Masahiro Kato (2024) General bayesian predictive synthesis, 2024 self0.5113233%
8Adith Swaminathan and Thorsten Joachims (2015) Batch learning from logged bandit feedback through counterfactual risk minimization0.5113233%
9Emily Tallman and Mike West (2024) Predictive decision synthesis for portfolios: Betting on better models, 20240.5113233%
10Susan Athey and Stefan Wager (2021) Policy learning with observational data0.5112250%

Showing the top 10 of 54 scored citations.