EconBase
← All papers

Beating the Winner's Curse via Inference-Aware Policy Optimization

Hamsa Bastani, Osbert Bastani, Bryce McLaughlin

arXiv 20 Oct 2025 · Statistics — Machine Learning

arXiv:2510.18161 · PDF · DOI · OpenAlex · Extracted main text

Abstract

There has been a surge of recent interest in automatically learning policies to target treatment decisions based on rich individual covariates. In addition, practitioners want confidence that the learned policy has better performance than the incumbent policy according to downstream policy evaluation. However, due to the winner's curse -- an issue where the policy optimization procedure exploits prediction errors rather than finding actual improvements -- predicted performance improvements are often not substantiated by downstream policy evaluation. To address this challenge, we propose a novel strategy called inference-aware policy optimization, which modifies policy optimization to account for how the policy will be evaluated downstream. Specifically, it optimizes not only for the estimated objective value, but also for the chances that the estimate of the policy's improvement passes a significance test during downstream policy evaluation. We mathematically characterize the Pareto frontier of policies according to the tradeoff of these two goals. Based on our characterization, we design a policy optimization algorithm that estimates the Pareto frontier using machine learning models; then, the decision-maker can select the policy that optimizes their desired tradeoff, after which policy evaluation can be performed on the test set as usual. Finally, we perform simulations to illustrate the effectiveness of our methodology.

Citation extraction

58
references
82
in-text mentions
58
distinct cited
0
self-citations
30,282
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bansak K, Ferwerda J, Hainmueller J, Dillon A, Hangartner D, Lawrenc… (2018) Improving refugee integration through data-driven algorithmic assignment0.81142100%
2Rosenbaum PR, Rubin DB (1983) The central role of the propensity score in observational studies for causal effects0.81142100%
3Swaminathan A, Joachims T (2015) b) The self-normalized estimator for counterfactual learning0.73732100%
4Athey S, Tibshirani J, Wager S (2019) Generalized random forests0.73732100%
5Dudḱ M, Langford J, Li L (2011) Doubly robust policy evaluation and learning0.64441100%
6Lu B, Hardin J (2021) A unified framework for random forest prediction error estimation0.64422100%
7Robins JM, Rotnitzky A, Zhao LP (1994) Estimation of regression coefficients when some regressors are not always observed0.64422100%
8Rubin D (1974) Estimating causal effects of treatments in randomized and nonrandomized studies0.64422100%
9Bottou L, Peters J, Quiñonero-Candela J, Charles DX, Chickering DM,… (2013) Counterfactual reasoning and learning systems: The example of computational advertising0.58531100%
10Andrews I, Kitagawa T, McCloskey A (2024) Inference on Winners0.51121100%

Showing the top 10 of 58 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching0.92843
2Valuing Winners: When and How to Correct for Selection Bias in Randomized Experiments0.64422