EconBase
← All papers

Causal-Policy Forest for End-to-End Policy Learning

Masahiro Kato

arXiv 28 Dec 2025 · Econometrics

arXiv:2512.22846 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This study proposes an end-to-end algorithm for policy learning in causal inference. We observe data consisting of covariates, treatment assignments, and outcomes, where only the outcome corresponding to the assigned treatment is observed. The goal of policy learning is to train a policy from the observed data, where a policy is a function that recommends an optimal treatment for each individual, to maximize the policy value. In this study, we first show that maximizing the policy value is equivalent to minimizing the mean squared error for the conditional average treatment effect (CATE) under ${-1, 1}$ restricted regression models. Based on this finding, we modify the causal forest, an end-to-end CATE estimation algorithm, for policy learning. We refer to our algorithm as the causal-policy forest. Our algorithm has three advantages. First, it is a simple modification of an existing, widely used CATE estimation method, therefore, it helps bridge the gap between policy learning and CATE estimation in practice. Second, while existing studies typically estimate nuisance parameters for policy learning as a separate task, our algorithm trains the policy in a more end-to-end manner. Third, as in standard decision trees and random forests, we train the models efficiently, avoiding computational intractability.

Citation extraction

12
references
24
in-text mentions
12
distinct cited
2
self-citations
4,552
main-text words

appendix boundary found by appendix_command · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Toru Kitagawa and Aleksey Tetenov (2018) Who should be treated? empirical welfare maximization methods for treatment choice1.00053100%
2Zhengyuan Zhou, Susan Athey, and Stefan Wager (2023) Offline multi-action policy learning: Generalization and optimization0.73732100%
3Stefan Wager and Susan Athey (2018) Estimation and inference of heterogeneous treatment effects using random forests0.64422100%
4Masahiro Kato (2025) Nearest neighbor matching as least squares density ratio estimation and riesz regression, 2025b self0.58531100%
5Zhexiao Lin, Peng Ding, and Fang Han (2023) Estimation based on nearest neighbor matching: from density ratio to average treatment effect0.58531100%
6Susan Athey and Stefan Wager (2021) Policy learning with observational data0.51121100%
7Susan Athey and Guido Imbens (2016) Recursive partitioning for heterogeneous causal effects0.40511100%
8Victor Chernozhukov, Whitney K. Newey, Victor Quintas-Martinez, and… (2021) Automatic debiased machine learning via riesz regression, 20210.40511100%
9Victor Chernozhukov, Whitney K. Newey, and Rahul Singh (2022) Automatic debiased machine learning of causal and structural effects0.40511100%
10Masahiro Kato (2025) Bridging the gap between empirical welfare maximization and conditional average treatment effect estimation in policy learning,… self0.40511100%

Showing the top 10 of 12 scored citations.