Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
arXiv 17 Jun 2026 · Statistics — Methodology
arXiv:2606.19117 · PDF · DOI · OpenAlex · Extracted main text
Offline policy learning has received growing attention in causal inference. The primary objective is to learn a policy (individualized treatment rule) as a mapping from covariates to treatment that maximizes the empirical welfare defined as the mean of scalar-valued potential outcomes. In this paper, we study offline policy learning with distribution-valued outcomes, where each potential outcome is a probability measure on $\mathbb{R}$ and the reward is defined through a utility functional applied to the Wasserstein barycenter of induced outcome distributions. We establish statistical guarantees for the policy learning framework based on both Inverse Probability Weighting (IPW) and Doubly Robust (DR) estimators. By handling the challenging uniform deviation over the product of the combinatorial policy class and the infinite-dimensional quantile domain, we prove that the finite-sample regret has leading dependence $\widetilde{\mathcal{O}}(\sqrt{N-dim(Π)/N})$. In the one-dimensional Wasserstein setting and under the stated regularity conditions, the leading regret rate is still governed by the policy-class complexity. Moreover, we provide a minimax lower bound establishing the sharpness of the leading dependence on $N$ and $N-dim(Π)$.
appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Athey, Susan and Wager, Stefan (2021) Policy learning with observational data | 0.874 | 5 | 2 | 100% |
| 2 | Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.874 | 5 | 2 | 100% |
| 3 | Cui, Yifan and Han, Sukjin (2025) Policy learning with distributional welfare | 0.737 | 3 | 2 | 100% |
| 4 | Kurisu, Daisuke and Zhou, Yidong and Otsu, Taisuke and Müller, Hans-… (2024) Geodesic causal inference | 0.737 | 3 | 2 | 100% |
| 5 | Kallus, Nathan and Zhou, Angela (2021) Minimax-optimal policy learning under unobserved confounding | 0.644 | 2 | 2 | 100% |
| 6 | Lin, Zhenhua and Kong, Dehan and Wang, Linbo (2023) Causal inference on distribution functions | 0.644 | 2 | 2 | 100% |
| 7 | Wang, Lan and Zhou, Yu and Song, Rui and Sherwood, Ben (2018) Quantile-optimal treatment regimes | 0.644 | 2 | 2 | 100% |
| 8 | Aliprantis, Dionissi and Carroll, Daniel and Young, Eric (2022) The Dynamics of the Racial Wealth Gap | 0.405 | 1 | 1 | 100% |
| 9 | Adjaho, Christopher and Christensen, Timothy (2022) Externally valid treatment choice | 0.405 | 1 | 1 | 100% |
| 10 | Ai, Chunrong and Fang, Yue and Xie, Haitian (2026) Data-driven policy learning for continuous treatments | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 41 scored citations.