Ye Luo, Martin Spindler, Jannis Kück
arXiv 29 Feb 2016 · Statistics — Machine Learning · 20 citations (OpenAlex)
arXiv:1602.08927 · PDF · DOI · OpenAlex · Extracted main text
Boosting is one of the most significant developments in machine learning. This paper studies the rate of convergence of $L_2$Boosting, which is tailored for regression, in a high-dimensional setting. Moreover, we introduce so-called \textquotedblleft post-Boosting\textquotedblright. This is a post-selection estimator which applies ordinary least squares to the variables selected in the first stage by $L_2$Boosting. Another variant is \textquotedblleft Orthogonal Boosting\textquotedblright\ where after each step an orthogonal projection is conducted. We show that both post-$L_2$Boosting and the orthogonal boosting achieve the same rate of convergence as LASSO in a sparse, high-dimensional setting. We show that the rate of convergence of the classical $L_2$Boosting depends on the design matrix described by a sparse eigenvalue constant. To show the latter results, we derive new approximation results for the pure greedy algorithm, based on analyzing the revisiting behavior of $L_2$Boosting. We also introduce feasible rules for early stopping, which can be easily implemented and used in applied work. Our results also allow a direct comparison between LASSO and boosting which has been missing from the literature. Finally, we present simulation studies and applications to illustrate the relevance of our theoretical results and to provide insights into the practical aspects of boosting. In these simulation studies, post-$L_2$Boosting clearly outperforms LASSO.
appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Jerome H. Friedman (2001) Greedy function approximation: A gradient boosting machine | 0.874 | 5 | 2 | 100% |
| 2 | Peter Bühlmann (2006) Boosting for high-dimensional linear models | 0.843 | 3 | 3 | 100% |
| 3 | R. A. DeVore and V. N. Temlyakov (1996) Some remarks on greedy algorithms | 0.811 | 4 | 2 | 100% |
| 4 | E.D. Livshitz and V.N. Temlyakov (2003) Two lower estimates in greedy approximation | 0.737 | 3 | 2 | 100% |
| 5 | Peter Bühlmann and Bin Yu (2003) Boosting with the $l_2$ Loss: Regression and classification | 0.737 | 3 | 2 | 100% |
| 6 | R Core Team (2014) R: A Language and Environment for Statistical Computing | 0.644 | 2 | 2 | 100% |
| 7 | Victor Chernozhukov, Christian Hansen, and Martin Spindler (2015) hdm: High-Dimensional Metrics, 2015 self | 0.644 | 2 | 2 | 100% |
| 8 | Leo Breiman (1998) Arcing classifiers | 0.644 | 2 | 2 | 100% |
| 9 | A. Belloni, D. Chen, V. Chernozhukov, and C. Hansen (2012) Sparse models and methods for optimal instruments with an application to eminent domain | 0.511 | 2 | 1 | 100% |
| 10 | Peter Bühlmann, Markus Kalisch, and Lukas Meier (2014) High-dimensional statistics with a view toward applications in biology | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 32 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Double/Debiased Machine Learning for Treatment and Structural Parameters | 0.405 | 1 | 1 |
| 2 | Debiased Machine Learning of Set-Identified Linear Models | 0.405 | 1 | 1 |
| 3 | Regularized Orthogonal Machine Learning for Nonlinear Semiparametric Models | 0.405 | 1 | 1 |
| 4 | Forward-Selected Panel Data Approach for Program Evaluation | 0.405 | 1 | 1 |
| 5 | 1908.08779 | 0.405 | 1 | 1 |
| 6 | 2002.12710 | 0.405 | 1 | 1 |
| 7 | 2012.00370 | 0.405 | 1 | 1 |
| 8 | 2012.00745 | 0.405 | 1 | 1 |
| 9 | Orthogonal Series Estimation for the Ratio of Conditional Expectation Functions | 0.405 | 1 | 1 |
| 10 | ddml: Double/debiased machine learning in Stata | 0.405 | 1 | 1 |