Hamsa Bastani, Osbert Bastani, Bryce McLaughlin
arXiv 20 Oct 2025 · Statistics — Machine Learning
arXiv:2510.18161 · PDF · DOI · OpenAlex · Extracted main text
There has been a surge of recent interest in automatically learning policies to target treatment decisions based on rich individual covariates. In addition, practitioners want confidence that the learned policy has better performance than the incumbent policy according to downstream policy evaluation. However, due to the winner's curse -- an issue where the policy optimization procedure exploits prediction errors rather than finding actual improvements -- predicted performance improvements are often not substantiated by downstream policy evaluation. To address this challenge, we propose a novel strategy called inference-aware policy optimization, which modifies policy optimization to account for how the policy will be evaluated downstream. Specifically, it optimizes not only for the estimated objective value, but also for the chances that the estimate of the policy's improvement passes a significance test during downstream policy evaluation. We mathematically characterize the Pareto frontier of policies according to the tradeoff of these two goals. Based on our characterization, we design a policy optimization algorithm that estimates the Pareto frontier using machine learning models; then, the decision-maker can select the policy that optimizes their desired tradeoff, after which policy evaluation can be performed on the test set as usual. Finally, we perform simulations to illustrate the effectiveness of our methodology.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bansak K, Ferwerda J, Hainmueller J, Dillon A, Hangartner D, Lawrenc… (2018) Improving refugee integration through data-driven algorithmic assignment | 0.811 | 4 | 2 | 100% |
| 2 | Rosenbaum PR, Rubin DB (1983) The central role of the propensity score in observational studies for causal effects | 0.811 | 4 | 2 | 100% |
| 3 | Swaminathan A, Joachims T (2015) b) The self-normalized estimator for counterfactual learning | 0.737 | 3 | 2 | 100% |
| 4 | Athey S, Tibshirani J, Wager S (2019) Generalized random forests | 0.737 | 3 | 2 | 100% |
| 5 | Dudḱ M, Langford J, Li L (2011) Doubly robust policy evaluation and learning | 0.644 | 4 | 1 | 100% |
| 6 | Lu B, Hardin J (2021) A unified framework for random forest prediction error estimation | 0.644 | 2 | 2 | 100% |
| 7 | Robins JM, Rotnitzky A, Zhao LP (1994) Estimation of regression coefficients when some regressors are not always observed | 0.644 | 2 | 2 | 100% |
| 8 | Rubin D (1974) Estimating causal effects of treatments in randomized and nonrandomized studies | 0.644 | 2 | 2 | 100% |
| 9 | Bottou L, Peters J, Quiñonero-Candela J, Charles DX, Chickering DM,… (2013) Counterfactual reasoning and learning systems: The example of computational advertising | 0.585 | 3 | 1 | 100% |
| 10 | Andrews I, Kitagawa T, McCloskey A (2024) Inference on Winners | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 58 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Winner's Curse Drives False Promises in Data-Driven Decisions: A Case Study in Refugee Matching | 0.928 | 4 | 3 |
| 2 | Valuing Winners: When and How to Correct for Selection Bias in Randomized Experiments | 0.644 | 2 | 2 |