Harold D. Chiang, Yukitoshi Matsushita, Taisuke Otsu
arXiv 17 Nov 2025 · Statistics — Machine Learning
arXiv:2511.13934 · PDF · Extracted main text
We develop an empirical likelihood (EL) framework for random forests and related ensemble methods, providing a likelihood-based approach to quantify their statistical uncertainty. Exploiting the incomplete $U$-statistic structure inherent in ensemble predictions, we construct an EL statistic that is asymptotically chi-squared when subsampling induced by incompleteness is not overly sparse. Under sparser subsampling regimes, the EL statistic tends to over-cover due to loss of pivotality; we therefore propose a modified EL that restores pivotality through a simple adjustment. Our method retains key properties of EL while remaining computationally efficient. Theory for honest random forests and simulations demonstrate that modified EL achieves accurate coverage and practical reliability relative to existing inference methods.
appendix boundary found by appendix_command · 53% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Peng, Wei and Coleman, Tim and Mentch, Lucas (2022) Rates of convergence for random forests via generalized U-statistics | 0.941 | 6 | 4 | 83% |
| 2 | Wager, Stefan and Athey, Susan (2018) Estimation and inference of heterogeneous treatment effects using random forests | 0.890 | 17 | 4 | 71% |
| 3 | Matsushita, Yukitoshi and Otsu, Taisuke (2021) Jackknife empirical likelihood: small bandwidth, sparse network and high-dimensional asymptotics self | 0.874 | 5 | 2 | 100% |
| 4 | Peng, Wei and Mentch, Lucas and Stefanski, Leonard (2025) Bias, consistency, and alternative perspectives of the infinitesimal jackknife | 0.843 | 5 | 3 | 60% |
| 5 | Owen, Art B (1988) Empirical likelihood ratio confidence intervals for a single functional | 0.811 | 4 | 2 | 100% |
| 6 | Breiman, Leo (2001) Random forests | 0.737 | 3 | 2 | 100% |
| 7 | Efron, Bradley and Stein, Charles (1981) The jackknife estimate of variance | 0.737 | 3 | 2 | 100% |
| 8 | Jing, Bing-Yi and Yuan, Junqing and Zhou, Wang (2009) Jackknife empirical likelihood | 0.737 | 3 | 2 | 100% |
| 9 | Wager, Stefan and Hastie, Trevor and Efron, Bradley (2014) Confidence intervals for random forests: The jackknife and the infinitesimal jackknife | 0.737 | 3 | 2 | 100% |
| 10 | Breiman, Leo (1996) Bagging predictors | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 42 scored citations.