Ayush Sawarni, Jikai Jin, Justin Whitehouse, Vasilis Syrgkanis
arXiv 22 Oct 2025 · Machine Learning
arXiv:2510.19672 · PDF · DOI · OpenAlex · Extracted main text
Policy learning algorithms are widely used in areas such as personalized medicine and advertising to develop individualized treatment regimes. However, most methods force a decision even when predictions are uncertain, which is risky in high-stakes settings. We study policy learning with abstention, where a policy may defer to a safe default or an expert. When a policy abstains, it receives a small additive reward on top of the value of a random guess. We propose a two-stage learner that first identifies a set of near-optimal policies and then constructs an abstention rule from their disagreements. We establish fast O(1/n)-type regret guarantees when propensities are known, and extend these guarantees to the unknown-propensity case via a doubly robust (DR) objective. We further show that abstention is a versatile tool with direct applications to other core problems in policy learning: it yields improved guarantees under margin conditions without the common realizability assumption, connects to distributionally robust policy learning by hedging against small data shifts, and supports safe policy improvement by ensuring improvement over a baseline policy with high probability.
appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Bousquet, Olivier and Zhivotovskiy, Nikita (2021) Fast classification rates without standard margin assumptions | 0.928 | 5 | 4 | 80% |
| 2 | Thomas, Philip and Theocharous, Georgios and Ghavamzadeh, Mohammad (2015) High confidence policy improvement | 0.928 | 5 | 3 | 80% |
| 3 | Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.843 | 5 | 4 | 60% |
| 4 | Athey, Susan and Wager, Stefan (2021) Policy learning with observational data | 0.811 | 4 | 2 | 100% |
| 5 | Cho, Brian and Pop, Ana-Roxana and Gan, Kyra and Corbett-Davies, Sam… (2025) CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies | 0.644 | 2 | 2 | 100% |
| 6 | Foster, Dylan J and Syrgkanis, Vasilis (2023) Orthogonal statistical learning self | 0.511 | 3 | 2 | 33% |
| 7 | Alexander Luedtke and Antoine Chambaz (2020) Performance guarantees for policy learning | 0.511 | 2 | 1 | 100% |
| 8 | Bang, Heejung and Robins, James M (2005) Doubly robust estimation in missing data and causal inference models | 0.405 | 1 | 1 | 100% |
| 9 | Bartlett, Peter L. and Wegkamp, Marten H (2008) Classification with a Reject Option Using a Hinge Loss | 0.405 | 1 | 1 | 100% |
| 10 | Ben-David, Shai and Urner, Ruth (2014) The sample complexity of agnostic learning under deterministic labels | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 47 scored citations.