EconBase
← All papers

Wasserstein Policy Learning for Distributional Outcomes

Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang

arXiv 17 Jun 2026 · Statistics — Methodology

arXiv:2606.19117 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Offline policy learning has received growing attention in causal inference. The primary objective is to learn a policy (individualized treatment rule) as a mapping from covariates to treatment that maximizes the empirical welfare defined as the mean of scalar-valued potential outcomes. In this paper, we study offline policy learning with distribution-valued outcomes, where each potential outcome is a probability measure on $\mathbb{R}$ and the reward is defined through a utility functional applied to the Wasserstein barycenter of induced outcome distributions. We establish statistical guarantees for the policy learning framework based on both Inverse Probability Weighting (IPW) and Doubly Robust (DR) estimators. By handling the challenging uniform deviation over the product of the combinatorial policy class and the infinite-dimensional quantile domain, we prove that the finite-sample regret has leading dependence $\widetilde{\mathcal{O}}(\sqrt{N-dim(Π)/N})$. In the one-dimensional Wasserstein setting and under the stated regularity conditions, the leading regret rate is still governed by the policy-class complexity. Moreover, we provide a minimax lower bound establishing the sharpness of the leading dependence on $N$ and $N-dim(Π)$.

Citation extraction

41
references
56
in-text mentions
41
distinct cited
0
self-citations
6,949
main-text words

appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Athey, Susan and Wager, Stefan (2021) Policy learning with observational data0.87452100%
2Kitagawa, Toru and Tetenov, Aleksey (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.87452100%
3Cui, Yifan and Han, Sukjin (2025) Policy learning with distributional welfare0.73732100%
4Kurisu, Daisuke and Zhou, Yidong and Otsu, Taisuke and Müller, Hans-… (2024) Geodesic causal inference0.73732100%
5Kallus, Nathan and Zhou, Angela (2021) Minimax-optimal policy learning under unobserved confounding0.64422100%
6Lin, Zhenhua and Kong, Dehan and Wang, Linbo (2023) Causal inference on distribution functions0.64422100%
7Wang, Lan and Zhou, Yu and Song, Rui and Sherwood, Ben (2018) Quantile-optimal treatment regimes0.64422100%
8Aliprantis, Dionissi and Carroll, Daniel and Young, Eric (2022) The Dynamics of the Racial Wealth Gap0.40511100%
9Adjaho, Christopher and Christensen, Timothy (2022) Externally valid treatment choice0.40511100%
10Ai, Chunrong and Fang, Yue and Xie, Haitian (2026) Data-driven policy learning for continuous treatments0.40511100%

Showing the top 10 of 41 scored citations.