EconBase
← All papers

To Bag is to Prune

Philippe Goulet Coulombe

arXiv 17 Aug 2020 · Statistics — Machine Learning · publishedStudies in Nonlinear Dynamics and Econometrics (2024) · 2 citations (OpenAlex)

arXiv:2008.07063 · PDF · DOI · OpenAlex · Extracted main text

Abstract

It is notoriously difficult to build a bad Random Forest (RF). Concurrently, RF blatantly overfits in-sample without any apparent consequence out-of-sample. Standard arguments, like the classic bias-variance trade-off or double descent, cannot rationalize this paradox. I propose a new explanation: bootstrap aggregation and model perturbation as implemented by RF automatically prune a latent "true" tree. More generally, randomized ensembles of greedily optimized learners implicitly perform optimal early stopping out-of-sample. So there is no need to tune the stopping point. By construction, novel variants of Boosting and MARS are also eligible for automatic tuning. I empirically demonstrate the property, with simulated and real data, by reporting that these new completely overfitting ensembles perform similarly to their tuned counterparts -- or better.

Citation extraction

54
references
98
in-text mentions
54
distinct cited
0
self-citations
15,720
main-text words

appendix boundary found by appendix_command · 88% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Elliott, G., Gargano, A., and Timmermann, A (2013) Complete subset regressions1.00053100%
2Friedman, J., Hastie, T., and Tibshirani, R (2001) The Elements of Statistical Learning, volume 10.87472100%
3Friedman, J. H (1991) Multivariate adaptive regression splines0.8746567%
4Belkin, M., Hsu, D., Ma, S., and Mandal, S (2019) Reconciling modern machine-learning practice and the classical bias–variance trade-off0.87462100%
5Breiman, L (1996) Bagging predictors0.87452100%
6Breiman, L (2001) Random forests0.84333100%
7Milborrow, S (2018) earth: Multivariate Adaptive Regression Splines0.7373367%
8Gu, S., Kelly, B., and Xiu, D (2020) Empirical asset pricing via machine learning0.7373367%
9Goulet Coulombe, P., Leroux, M., Stevanovic, D., and Surprenant, S (2022) How is machine learning useful for macroeconomic forecasting?0.73732100%
10Kotchoni, R., Leroux, M., and Stevanovic, D (2019) Macroeconomic forecast accuracy in a data-rich environment0.64422100%

Showing the top 10 of 54 scored citations.