EconBase
← All papers

Policy Learning with Distributional Welfare

Yifan Cui, Sukjin Han

arXiv 27 Nov 2023 · Statistics — Methodology · publishedJournal of the American Statistical Association (2025) · 1 citations (OpenAlex)

arXiv:2311.15878 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper, we explore optimal treatment allocation policies that target distributional welfare. Most literature on treatment choice has considered utilitarian welfare based on the conditional average treatment effect (ATE). While average welfare is intuitive, it may yield undesirable allocations especially when individuals are heterogeneous (e.g., with outliers) - the very reason individualized treatments were introduced in the first place. This observation motivates us to propose an optimal policy that allocates the treatment based on the conditional quantile of individual treatment effects (QoTE). Depending on the choice of the quantile probability, this criterion can accommodate a policymaker who is either prudent or negligent. The challenge of identifying the QoTE lies in its requirement for knowledge of the joint distribution of the counterfactual outcomes, which is not generally point-identified. We introduce minimax policies that are robust to this model uncertainty. A range of identifying assumptions can be used to yield more informative policies. For both stochastic and deterministic policies, we establish the asymptotic bound on the regret of implementing the proposed policies. The framework can be generalized to any setting where welfare is defined as a functional of the joint distribution of the potential outcomes.

Citation extraction

83
references
116
in-text mentions
83
distinct cited
6
self-citations
10,865
main-text words

appendix boundary found by appendix_command · 63% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zhao, Y., D. Zeng, A. J. Rush, and M. R. Kosorok (2012) Estimating individualized treatment rules using outcome weighted learning0.9285480%
2Leqi, L. and E. H. Kennedy (2021) Median optimal treatment regimes0.87452100%
3Cui, Y (2021) Individualized decision making under partial identification: three perspectives, two optimality results, and one paradox self0.73732100%
4Wang, L., Y. Zhou, R. Song, and B. Sherwood (2018) Quantile-optimal treatment regimes0.73732100%
5Blundell, R., A. Gosling, H. Ichimura, and C. Meghir (2007) Changes in the distribution of male and female wages accounting for employment composition using bounds0.64422100%
6D'Adamo, R (2021) Orthogonal Policy Learning Under Ambiguity0.64422100%
7Frandsen, B. R. and L. J. Lefgren (2021) Partial identification of the distribution of treatment effects with an application to the Knowledge is Power Program (KIPP)0.64422100%
8Han, S (2023) Optimal dynamic treatment regimes and partial welfare ordering self0.64422100%
9Hirano, K. and G. W. Imbens (2001) Estimation of causal effects using propensity score weighting: An application to data on right heart catheterization0.64422100%
10Manski, C. F (2007) Minimax-regret treatment choice with missing outcome data0.64422100%

Showing the top 10 of 83 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Wasserstein Policy Learning for Distributional Outcomes0.73732
2Inference for Interval-Identified Parameters Selected from an Estimated Set0.64422
3The Identification Power of Combining Experimental and Observational Data for Distributional Treatment Effect Parameters0.64422
4Debiased Machine Learning of Aggregated Intersection Bounds and Other Causal Parameters0.51121
5On the Lower Confidence Band for the Optimal Welfare in Policy Learning0.40511
6Counting Defiers: A Design-Based Model of an Experiment Can Reveal Evidence Beyond the Average Effect0.40511
7Locally Robust Policy Learning: Inequality, Inequality of Opportunity and Intergenerational Mobility0.40511