EconBase
← All papers

Conformal Policy Learning with Distribution-Free Safety Guarantees

Ying Jin, Naoki Egami

arXiv 15 Sep 2026 · Statistics — Methodology

arXiv:2609.17296 · PDF · Extracted main text

Abstract

Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of “do no harm.” In this paper, we propose conformal policy learning (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.

Citation extraction

57
references
126
in-text mentions
57
distinct cited
8
self-citations
16,531
main-text words

appendix boundary found by appendix_command · 38% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Athey, Susan and Wager, Stefan (2021) Policy learning with observational data1.00053100%
2Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? empirical welfare maximization methods for treatment choice1.00053100%
3Manski, Charles F (2004) Statistical Treatment Rules for Heterogeneous Populations1.00053100%
4Jin, Ying and Candès, Emmanuel J (2023) Selection by prediction with conformal p-values self0.9568488%
5Hirano, Keisuke and Porter, Jack R (2009) Asymptotics for Statistical Treatment Rules0.92843100%
6Lei, Lihua and Candès, Emmanuel J (2021) Conformal inference of counterfactuals and individual treatment effects0.92843100%
7Zhao, Yingqi and Zeng, Donglin and Rush, A John and Kosorok, Michael R (2012) Estimating Individualized Treatment Rules using Outcome Weighted Learning0.92843100%
8Kallus, Nathan (2022) What's the harm? sharp bounds on the fraction negatively affected by treatment0.9209578%
9Li, Haoxuan and Zheng, Chunyuan and Cao, Yixiao and Geng, Zhi and Li… (2023) Trustworthy policy learning under the counterfactual no-harm criterion0.8558562%
10Vovk, Vladimir and Gammerman, Alexander and Shafer, Glenn (2005) Algorithmic learning in a random world0.8434475%

Showing the top 10 of 57 scored citations.