arXiv 30 Oct 2025 · Statistics — Machine Learning
arXiv:2510.26723 · PDF · DOI · OpenAlex · Extracted main text
The goal of policy learning is to train a policy function that recommends a treatment given covariates to maximize population welfare. There are two major approaches in policy learning: the empirical welfare maximization (EWM) approach and the plug-in approach. The EWM approach is analogous to a classification problem, where one first builds an estimator of the population welfare, which is a functional of policy functions, and then trains a policy by maximizing the estimated welfare. In contrast, the plug-in approach is based on regression, where one first estimates the conditional average treatment effect (CATE) and then recommends the treatment with the highest estimated outcome. This study bridges the gap between the two approaches by showing that both are based on essentially the same optimization problem. In particular, we prove an exact equivalence between EWM and least squares over a reparameterization of the policy class. As a consequence, the two approaches are interchangeable in several respects and share the same theoretical guarantees under common conditions. Leveraging this equivalence, we propose a regularization method for policy learning. The reduction to least squares yields a smooth surrogate that is typically easier to optimize in practice. At the same time, for many natural policy classes the inherent combinatorial hardness of exact EWM generally remains, so the reduction should be viewed as an optimization aid rather than a universal bypass of NP-hardness.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 0.874 | 5 | 2 | 100% |
| 2 | Susan Athey and Stefan Wager (2021) Policy learning with observational data | 0.644 | 2 | 2 | 100% |
| 3 | Adith Swaminathan and Thorsten Joachims Batch learning from logged bandit feedback through counterfactual risk minimization | 0.511 | 2 | 1 | 100% |
| 4 | Adith Swaminathan and Thorsten Joachims (2015) Counterfactual risk minimization: learning from logged bandit feedback | 0.511 | 2 | 1 | 100% |
| 5 | Jean-Yves Audibert and Alexandre B. Tsybakov (2007) Fast learning rates for plug-in classifiers | 0.405 | 1 | 1 | 100% |
| 6 | Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo,… (2018) Double/debiased machine learning for treatment and structural parameters | 0.405 | 1 | 1 | 100% |
| 7 | Victor Chernozhukov, Whitney K Newey, and Rahul Singh (2022) Debiased machine learning of global and local parameters using regularized riesz representers | 0.405 | 1 | 1 | 100% |
| 8 | Victor Chernozhukov, Whitney K. Newey, Victor Quintas-Martinez, and… (2024) Automatic debiased machine learning via riesz regression, 2024 | 0.405 | 1 | 1 | 100% |
| 9 | Masahiro Kato (2025) Direct bias-correction term estimation for propensity scores and average treatment effect estimation, 2025a self | 0.405 | 1 | 1 | 100% |
| 10 | Masahiro Kato (2025) Direct debiased machine learning via bregman divergence minimization, 2025b self | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 13 scored citations.