EconBase
← All papers

Flexible machine learning estimation of conditional average treatment effects: a blessing and a curse

Richard Post, Isabel van den Heuvel, Marko Petkovic, Edwin van den Heuvel

arXiv 29 Oct 2022 · Statistics — Methodology · publishedEpidemiology (2023) · 4 citations (OpenAlex)

arXiv:2210.16547 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Causal inference from observational data requires untestable identification assumptions. If these assumptions apply, machine learning (ML) methods can be used to study complex forms of causal effect heterogeneity. Recently, several ML methods were developed to estimate the conditional average treatment effect (CATE). If the features at hand cannot explain all heterogeneity, the individual treatment effects (ITEs) can seriously deviate from the CATE. In this work, we demonstrate how the distributions of the ITE and the CATE can differ when a causal random forest (CRF) is applied. We extend the CRF to estimate the difference in conditional variance between treated and controls. If the ITE distribution equals the CATE distribution, this estimated difference in variance should be small. If they differ, an additional causal assumption is necessary to quantify the heterogeneity not captured by the CATE distribution. The conditional variance of the ITE can be identified when the individual effect is independent of the outcome under no treatment given the measured features. Then, in the cases where the ITE and CATE distributions differ, the extended CRF can appropriately estimate the variance of the ITE distribution while the CRF fails to do so.

Citation extraction

57
references
82
in-text mentions
57
distinct cited
0
self-citations
8,155
main-text words

appendix boundary found by appendix_command · 59% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Athey S, Tibshirani J, Wager S (2019) Generalized random forests1.00053100%
2Hernán MA, Robins JM (2020) Causal Inference: What If0.87452100%
3Athey S, Imbens GW (2016) Recursive partitioning for heterogeneous causal effects0.81142100%
4Wager S, Athey S (2018) Estimation and Inference of Heterogeneous Treatment Effects using Random Forests0.64441100%
5Chiu LS, Pedley A, Massaro JM, Benjamin EJ, Mitchell GF, McManus DD,… (2020) The association of non-alcoholic fatty liver disease and cardiac structure and function—framingham heart study0.6443267%
6Naimi AI, Mishler AE, Kennedy EH (2021) Challenges in Obtaining Valid Causal Effect Estimates with Machine Learning Algorithms0.64422100%
7Nie X, Wager S (2020) Quasi-oracle estimation of heterogeneous treatment effects0.64422100%
8Chernozhukov V, Chetverikov D, Demirer M, Duflo E, Hansen C, Newey W… (2018) Double/debiased machine learning for treatment and structural parameters0.51121100%
9Dickerman BA, Hernán MA (2020) Counterfactual prediction is not only for causal inference0.51121100%
10Hill JL (2011) Bayesian nonparametric modeling for causal inference0.51121100%

Showing the top 10 of 57 scored citations.