Hamsa Bastani, Osbert Bastani, Bryce McLaughlin
arXiv 9 Feb 2026 · Statistics — Machine Learning
arXiv:2602.08892 · PDF · DOI · OpenAlex · Extracted main text
A major challenge in data-driven decision-making is accurate policy evaluation-i.e., guaranteeing that a learned decision-making policy achieves the promised benefits. A popular strategy is model-based policy evaluation, which estimates a model from data to infer counterfactual outcomes. This strategy is known to produce unwarrantedly optimistic estimates of the true benefit due to the winner's curse. We searched the recent literature on data-driven decision-making, identifying a sample of 55 papers published in the Management Science in the past decade; all but two relied on this flawed methodology. Several common justifications are provided: (1) the estimated models are accurate, stable, and well-calibrated, (2) the historical data uses random treatment assignment, (3) the model family is well-specified, and (4) the evaluation methodology uses sample splitting. Unfortunately, we show that no combination of these justifications avoids the winner's curse. First, we provide a theoretical analysis demonstrating that the winner's curse can cause large, spurious reported benefits even when all these justifications hold. Second, we perform a simulation study based on the recent and consequential data-driven refugee matching problem. We construct a synthetic refugee matching environment (calibrated to closely match the real setting) but designed so that no assignment policy can improve expected employment compared to random assignment. Model-based methods report large, stable gains of around 60% even when the true effect is zero; these gains are on par with improvements of 22-75% reported in the literature. Our results provide strong evidence against model-based evaluation.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bansak, Kirk and Ferwerda, Jeremy and Hainmueller, Jens and Dillon,… (2018) Improving refugee integration through data-driven algorithmic assignment | 1.000 | 25 | 3 | 100% |
| 2 | Ahani, Narges and Andersson, Tommy and Martinello, Alessandro and Te… (2021) Placement optimization in refugee resettlement | 1.000 | 14 | 3 | 100% |
| 3 | Bastani, Hamsa and Bastani, Osbert and McLaughlin, Bryce (2025) Beating the Winner's Curse via Inference-Aware Policy Optimization self | 0.928 | 4 | 3 | 100% |
| 4 | Banerjee, Abhijit and Chandrasekhar, Arun G and Dalpath, Suresh and… (2025) Selecting the most effective nudge: Evidence from a large-scale experiment on immunization | 0.737 | 3 | 2 | 100% |
| 5 | Chernozhukov, Victor and Lee, Sokbae and Rosen, Adam M and Sun, Liyang (2025) Policy Learning with Confidence | 0.737 | 3 | 2 | 100% |
| 6 | Mandyam, Aishwarya and Meng, Jason and Gao, Ge and Sun, Jiankai and… (2025) PERRY: Policy Evaluation with Confidence Intervals using Auxiliary Data | 0.737 | 3 | 2 | 100% |
| 7 | Harrison, J Richard and March, James G (1984) Decision making and postdecision surprises | 0.585 | 3 | 1 | 100% |
| 8 | Smith, James E and Winkler, Robert L (2006) The optimizer’s curse: Skepticism and postdecision surprise in decision analysis | 0.585 | 3 | 1 | 100% |
| 9 | Andrews, Isaiah and Kitagawa, Toru and McCloskey, Adam (2024) Inference on Winners | 0.511 | 2 | 1 | 100% |
| 10 | Horvitz, Daniel G and Thompson, Donovan J (1952) A generalization of sampling without replacement from a finite universe | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 23 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Robustness of Refugee-Matching Gains to Off-Policy Evaluation Choices | 0.511 | 2 | 1 |