EconBase
← All papers

Policy Learning with Abstention

Ayush Sawarni, Jikai Jin, Justin Whitehouse, Vasilis Syrgkanis

arXiv 22 Oct 2025 · Machine Learning

arXiv:2510.19672 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Policy learning algorithms are widely used in areas such as personalized medicine and advertising to develop individualized treatment regimes. However, most methods force a decision even when predictions are uncertain, which is risky in high-stakes settings. We study policy learning with abstention, where a policy may defer to a safe default or an expert. When a policy abstains, it receives a small additive reward on top of the value of a random guess. We propose a two-stage learner that first identifies a set of near-optimal policies and then constructs an abstention rule from their disagreements. We establish fast O(1/n)-type regret guarantees when propensities are known, and extend these guarantees to the unknown-propensity case via a doubly robust (DR) objective. We further show that abstention is a versatile tool with direct applications to other core problems in policy learning: it yields improved guarantees under margin conditions without the common realizability assumption, connects to distributionally robust policy learning by hedging against small data shifts, and supports safe policy improvement by ensuring improvement over a baseline policy with high probability.

Citation extraction

47
references
67
in-text mentions
47
distinct cited
3
self-citations
6,452
main-text words

appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bousquet, Olivier and Zhivotovskiy, Nikita (2021) Fast classification rates without standard margin assumptions0.9285480%
2Thomas, Philip and Theocharous, Georgios and Ghavamzadeh, Mohammad (2015) High confidence policy improvement0.9285380%
3Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.8435460%
4Athey, Susan and Wager, Stefan (2021) Policy learning with observational data0.81142100%
5Cho, Brian and Pop, Ana-Roxana and Gan, Kyra and Corbett-Davies, Sam… (2025) CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies0.64422100%
6Foster, Dylan J and Syrgkanis, Vasilis (2023) Orthogonal statistical learning self0.5113233%
7Alexander Luedtke and Antoine Chambaz (2020) Performance guarantees for policy learning0.51121100%
8Bang, Heejung and Robins, James M (2005) Doubly robust estimation in missing data and causal inference models0.40511100%
9Bartlett, Peter L. and Wegkamp, Marten H (2008) Classification with a Reject Option Using a Hinge Loss0.40511100%
10Ben-David, Shai and Urner, Ruth (2014) The sample complexity of agnostic learning under deterministic labels0.40511100%

Showing the top 10 of 47 scored citations.