arXiv 27 Feb 2026 · Statistics — Machine Learning
arXiv:2602.23672 · PDF · DOI · OpenAlex · Extracted main text
This study proposes the General Bayes framework for policy learning. We consider decision problems in which a decision-maker chooses an action from an action set to maximize its expected welfare. Typical examples include treatment choice and portfolio selection. In such problems, the statistical target is a decision rule, and the prediction of each outcome $Y(a)$ is not necessarily of primary interest. We formulate this policy learning problem by loss-based Bayesian updating. Our main technical device is a squared-loss surrogate for welfare maximization. We show that maximizing empirical welfare over a policy class is equivalent to minimizing a scaled squared error in the outcome difference, up to a quadratic regularization controlled by a tuning parameter $ζ>0$. This rewriting yields a General Bayes posterior over decision rules that admits a Gaussian pseudo-likelihood interpretation. We clarify two Bayesian interpretations of the resulting generalized posterior, a working Gaussian view and a decision-theoretic loss-based view. As one implementation example, we introduce neural networks with tanh-squashed outputs. Finally, we provide theoretical guarantees in a PAC-Bayes style.
appendix boundary found by appendix_command · 48% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Pier Giovanni Bissiri, Chris C. Holmes, and Stephen G. Walker (2016) A general framework for updating belief distributions | 0.794 | 8 | 4 | 50% |
| 2 | O. Catoni, Project Euclid, Cornell University. Library, and Duke Uni… (2008) PAC-Bayesian Supervised Classification: The Thermodynamics of Statistical Learning | 0.737 | 5 | 3 | 40% |
| 3 | Zhengyuan Zhou, Susan Athey, and Stefan Wager (2023) Offline multi-action policy learning: Generalization and optimization | 0.737 | 4 | 3 | 50% |
| 4 | Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.644 | 3 | 2 | 67% |
| 5 | Pierre Alquier (2024) User-friendly introduction to pac-bayes bounds | 0.585 | 3 | 3 | 33% |
| 6 | Maxime Haddouche, Benjamin Guedj, Omar Rivasplata, and John Shawe-Ta… (2021) Pac-bayes unleashed: Generalisation bounds with unbounded losses | 0.585 | 3 | 3 | 33% |
| 7 | Masahiro Kato (2024) General bayesian predictive synthesis, 2024 self | 0.511 | 3 | 2 | 33% |
| 8 | Adith Swaminathan and Thorsten Joachims (2015) Batch learning from logged bandit feedback through counterfactual risk minimization | 0.511 | 3 | 2 | 33% |
| 9 | Emily Tallman and Mike West (2024) Predictive decision synthesis for portfolios: Betting on better models, 2024 | 0.511 | 3 | 2 | 33% |
| 10 | Susan Athey and Stefan Wager (2021) Policy learning with observational data | 0.511 | 2 | 2 | 50% |
Showing the top 10 of 54 scored citations.