Matias D. Cattaneo, Jason M. Klusowski, Ruiqi Rae Yu
arXiv 14 Sep 2025 · Mathematics — Statistics Theory
arXiv:2509.11381 · PDF · DOI · OpenAlex · Extracted main text
Recursive decision trees have emerged as a leading methodology for heterogeneous causal treatment effect estimation and inference in experimental and observational settings. These procedures are fitted using the celebrated CART (Classification And Regression Tree) algorithm [Breiman et al., 1984], or custom variants thereof, and hence are believed to be "adaptive" to high-dimensional data, sparsity, or other specific features of the underlying data generating process. Athey and Imbens [2016] proposed several "honest" causal decision tree estimators, which have become the standard in both academia and industry. We study their estimators, and variants thereof, and establish lower bounds on their estimation error. We demonstrate that these popular heterogeneous treatment effect estimators cannot achieve a polynomial-in-$n$ convergence rate under basic conditions, where $n$ denotes the sample size. Contrary to common belief, honesty does not resolve these limitations and at best delivers negligible logarithmic improvements in sample size or dimension. As a result, these commonly used estimators can exhibit poor performance in practice, and even be inconsistent in some settings. Our theoretical insights are empirically validated through simulations.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Horváth, Lajos (1993) The maximum likelihood method for testing changes in the parameters of normal observations | 0.737 | 3 | 2 | 100% |
| 2 | Csörgö, M. and Horváth, L (1997) Limit Theorems in Change-Point Analysis | 0.693 | 15 | 1 | 100% |
| 3 | Klusowski, Jason M and Tian, Peter M (2024) Large scale prediction with decision trees self | 0.693 | 7 | 1 | 100% |
| 4 | Victor Chernozhukov and Denis Chetverikov and Kengo Kato (2017) Central limit theorems and bootstrap in high dimensions | 0.693 | 6 | 1 | 100% |
| 5 | F. Eicker (1979) The Asymptotic Distribution of the Suprema of the Standardized Empirical Processes | 0.585 | 3 | 1 | 100% |
| 6 | Shorack, Galen R and Smythe, RT (1976) Inequalities for max| Sk|/bk where k $in$ Nr | 0.585 | 3 | 1 | 100% |
| 7 | László Györfi and Michael Kohler and Adam Krzyżak and Harro Walk (2002) A Distribution-Free Theory of Nonparametric Regression | 0.511 | 2 | 1 | 100% |
| 8 | Chernozhuokov, Victor and Chetverikov, Denis and Kato, Kengo and Koi… (2022) Improved central limit theorem and bootstrap approximations in high dimensions | 0.511 | 2 | 1 | 100% |
| 9 | Csörgö, M. and Révész, P (1981) Strong Approximations in Probability and Statistics | 0.511 | 2 | 1 | 100% |
| 10 | Latała, Rafał and Matlak, Dariusz (2017) Royen's Proof of the Gaussian Correlation Inequality | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 16 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | An Introduction to Double/Debiased Machine Learning | 0.644 | 2 | 2 |
| 2 | Decision Theory for the Archetype Discovery Problem | 0.405 | 1 | 1 |