EconBase
← All papers

Recursive Partitioning for Heterogeneous Causal Effects

Susan Athey, Guido Imbens

arXiv 5 Apr 2015 · Statistics — Machine Learning · publishedProceedings of the National Academy of Sciences (2016) · 1,551 citations (OpenAlex)

arXiv:1504.01132 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper we study the problems of estimating heterogeneity in causal effects in experimental or observational studies and conducting inference about the magnitude of the differences in treatment effects across subsets of the population. In applications, our method provides a data-driven approach to determine which subpopulations have large or small treatment effects and to test hypotheses about the differences in these effects. For experiments, our method allows researchers to identify heterogeneity in treatment effects that was not specified in a pre-analysis plan, without concern about invalidating inference due to multiple testing. In most of the literature on supervised machine learning (e.g. regression trees, random forests, LASSO, etc.), the goal is to build a model of the relationship between a unit's attributes and an observed outcome. A prominent role in these methods is played by cross-validation which compares predictions to actual outcomes in test samples, in order to select the level of complexity of the model that provides the best predictive power. Our method is closely related, but it differs in that it is tailored for predicting causal effects of a treatment rather than a unit's outcome. The challenge is that the "ground truth" for a causal effect is not observed for any individual unit: we observe the unit with the treatment, or without the treatment, but not both at the same time. Thus, it is not obvious how to use cross-validation to determine whether a causal effect has been accurately predicted. We propose several novel cross-validation criteria for this problem and demonstrate through simulations the conditions under which they perform better than standard methods for the problem of causal effects. We then apply the method to a large-scale field experiment re-ranking results on a search engine.

Citation extraction

30
references
43
in-text mentions
25
distinct cited
5
self-citations
11,674
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1G. Imbens and D. Rubin, Causal Inference for Statistics, Social, and… (2015) self0.92843100%
2L. Breiman, J. Friedman, R. Olshen, and C. Stone, Classification and… (1984) Wadsworth0.81142100%
3P. Holland, Statistics and Causal Inference (with discussion), Journ… (1986) 945-9700.73732100%
4D. Rubin, Estimating Causal Effects of Treatments in Randomized and… (1974) 688-7010.73732100%
5A. Beygelzimer and J. Langford, The Offset Tree for Learning with Pa… (2009)0.64422100%
6M. Dudik, J. Langford and L. Li, Doubly Robust Policy Evaluation and… (2011)0.64422100%
7T. Hastie, R. Tibshirani, and J. Friedman, The Elements of Statistic… (2011) Springer0.64422100%
8A. Zeileis, T. Hothorn, and K. Hornik, Model-based recursive partiti… (2008) 492-5140.58531100%
9L. Breiman, Random forests, Machine Learning, 45 (2001) 5-320.51121100%
10K. Hirano, G. Imbens and G. Ridder, Efficient Estimation of Average… (2003) 1161-1189 self0.51121100%

Showing the top 10 of 25 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Aggregation Trees1.000164
2On the adaptation of causal forests to manifold data1.00063
3Fisher-Schultz Lecture: Generic Machine Learning Inference on Heterogenous Treatment Effects in Randomized Experiments, with an Application to Immunization in India1.00054
4Machine Learning Methods Economists Should Know About1.00053
50cmFrom interpretability to inference: an estimation framework for universal approximators0.92843
6A Comparison of Methods for Treatment Assignment with an Application to Playlist Generation0.92843
7Heterogeneous Responses to Continuous Treatments: A Cluster-Based Causal Framework0.87452
8On regression-adjusted imputation estimators of the average treatment effect0.84333
9Feature Selection for Personalized Policy Analysis0.81142
10Finding Subgroups with Significant Treatment Effects0.73732