arXiv 28 Dec 2025 · Econometrics
arXiv:2512.22846 · PDF · DOI · OpenAlex · Extracted main text
This study proposes an end-to-end algorithm for policy learning in causal inference. We observe data consisting of covariates, treatment assignments, and outcomes, where only the outcome corresponding to the assigned treatment is observed. The goal of policy learning is to train a policy from the observed data, where a policy is a function that recommends an optimal treatment for each individual, to maximize the policy value. In this study, we first show that maximizing the policy value is equivalent to minimizing the mean squared error for the conditional average treatment effect (CATE) under ${-1, 1}$ restricted regression models. Based on this finding, we modify the causal forest, an end-to-end CATE estimation algorithm, for policy learning. We refer to our algorithm as the causal-policy forest. Our algorithm has three advantages. First, it is a simple modification of an existing, widely used CATE estimation method, therefore, it helps bridge the gap between policy learning and CATE estimation in practice. Second, while existing studies typically estimate nuisance parameters for policy learning as a separate task, our algorithm trains the policy in a more end-to-end manner. Third, as in standard decision trees and random forests, we train the models efficiently, avoiding computational intractability.
appendix boundary found by appendix_command · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice | 1.000 | 5 | 3 | 100% |
| 2 | Zhengyuan Zhou, Susan Athey, and Stefan Wager (2023) Offline multi-action policy learning: Generalization and optimization | 0.737 | 3 | 2 | 100% |
| 3 | Stefan Wager and Susan Athey (2018) Estimation and inference of heterogeneous treatment effects using random forests | 0.644 | 2 | 2 | 100% |
| 4 | Masahiro Kato (2025) Nearest neighbor matching as least squares density ratio estimation and riesz regression, 2025b self | 0.585 | 3 | 1 | 100% |
| 5 | Zhexiao Lin, Peng Ding, and Fang Han (2023) Estimation based on nearest neighbor matching: from density ratio to average treatment effect | 0.585 | 3 | 1 | 100% |
| 6 | Susan Athey and Stefan Wager (2021) Policy learning with observational data | 0.511 | 2 | 1 | 100% |
| 7 | Susan Athey and Guido Imbens (2016) Recursive partitioning for heterogeneous causal effects | 0.405 | 1 | 1 | 100% |
| 8 | Victor Chernozhukov, Whitney K. Newey, Victor Quintas-Martinez, and… (2021) Automatic debiased machine learning via riesz regression, 2021 | 0.405 | 1 | 1 | 100% |
| 9 | Victor Chernozhukov, Whitney K. Newey, and Rahul Singh (2022) Automatic debiased machine learning of causal and structural effects | 0.405 | 1 | 1 | 100% |
| 10 | Masahiro Kato (2025) Bridging the gap between empirical welfare maximization and conditional average treatment effect estimation in policy learning,… self | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 12 scored citations.