EconBase
← All papers

Bridging the Gap between Empirical Welfare Maximization and Conditional Average Treatment Effect Estimation in Policy Learning

Masahiro Kato

arXiv 30 Oct 2025 · Statistics — Machine Learning

arXiv:2510.26723 · PDF · DOI · OpenAlex · Extracted main text

Abstract

The goal of policy learning is to train a policy function that recommends a treatment given covariates to maximize population welfare. There are two major approaches in policy learning: the empirical welfare maximization (EWM) approach and the plug-in approach. The EWM approach is analogous to a classification problem, where one first builds an estimator of the population welfare, which is a functional of policy functions, and then trains a policy by maximizing the estimated welfare. In contrast, the plug-in approach is based on regression, where one first estimates the conditional average treatment effect (CATE) and then recommends the treatment with the highest estimated outcome. This study bridges the gap between the two approaches by showing that both are based on essentially the same optimization problem. In particular, we prove an exact equivalence between EWM and least squares over a reparameterization of the policy class. As a consequence, the two approaches are interchangeable in several respects and share the same theoretical guarantees under common conditions. Leveraging this equivalence, we propose a regularization method for policy learning. The reduction to least squares yields a smooth surrogate that is typically easier to optimize in practice. At the same time, for many natural policy classes the inherent combinatorial hardness of exact EWM generally remains, so the reduction should be viewed as an optimization aid rather than a universal bypass of NP-hardness.

Citation extraction

13
references
20
in-text mentions
13
distinct cited
2
self-citations
3,572
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice0.87452100%
2Susan Athey and Stefan Wager (2021) Policy learning with observational data0.64422100%
3Adith Swaminathan and Thorsten Joachims Batch learning from logged bandit feedback through counterfactual risk minimization0.51121100%
4Adith Swaminathan and Thorsten Joachims (2015) Counterfactual risk minimization: learning from logged bandit feedback0.51121100%
5Jean-Yves Audibert and Alexandre B. Tsybakov (2007) Fast learning rates for plug-in classifiers0.40511100%
6Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters0.40511100%
7Victor Chernozhukov, Whitney K Newey, and Rahul Singh (2022) Debiased machine learning of global and local parameters using regularized riesz representers0.40511100%
8Victor Chernozhukov, Whitney K. Newey, Victor Quintas-Martinez, and… (2024) Automatic debiased machine learning via riesz regression, 20240.40511100%
9Masahiro Kato (2025) Direct bias-correction term estimation for propensity scores and average treatment effect estimation, 2025a self0.40511100%
10Masahiro Kato (2025) Direct debiased machine learning via bregman divergence minimization, 2025b self0.40511100%

Showing the top 10 of 13 scored citations.