arXiv 10 Feb 2017 · Statistics — Machine Learning · publishedAmerican Economic Review (2017) · 12 citations (OpenAlex)
arXiv:1702.03244 · PDF · DOI · OpenAlex · Extracted main text
In the recent years more and more high-dimensional data sets, where the number of parameters $p$ is high compared to the number of observations $n$ or even larger, are available for applied researchers. Boosting algorithms represent one of the major advances in machine learning and statistics in recent years and are suitable for the analysis of such data sets. While Lasso has been applied very successfully for high-dimensional data sets in Economics, boosting has been underutilized in this field, although it has been proven very powerful in fields like Biostatistics and Pattern Recognition. We attribute this to missing theoretical results for boosting. The goal of this paper is to fill this gap and show that boosting is a competitive method for inference of a treatment effect or instrumental variable (IV) estimation in a high-dimensional setting. First, we present the $L_2$Boosting with componentwise least squares algorithm and variants which are tailored for regression problems which are the workhorse for most Econometric problems. Then we show how $L_2$Boosting can be used for estimation of treatment effects and IV estimation. We highlight the methods and illustrate them with simulations and empirical examples. For further results and technical details we refer to Luo and Spindler (2016, 2017) and to the online supplement of the paper.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Y. Luo \ M. Spindler (2016) High-Dimensional L2Boosting: Rate of Convergence | 0.644 | 2 | 2 | 100% |
| 2 | Y. Luo \ M. Spindler (2017) $L_2$Boosting for Economic Applications | 0.644 | 2 | 2 | 100% |
| 3 | Jerome H. Friedman (2001) Greedy Function Approximation: A Gradient Boosting Machine | 0.511 | 2 | 1 | 100% |
| 4 | Alexandre Belloni, Daniel Chen, Victor Chernozhukov \ Christian Hansen (2012) Sparse Models and Methods for Optimal Instruments with an Application to Eminent Domain | 0.405 | 1 | 1 | 100% |
| 5 | Alexandre Belloni, Victor Chernozhukov \ Christian Hansen (2014) Inference on Treatment Effects After Selection Amongst High-Dimensional Controls | 0.405 | 1 | 1 | 100% |
| 6 | V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen \ a… (2016) Double Machine Learning for Treatment and Causal Parameters | 0.405 | 1 | 1 | 100% |
| 7 | Victor Chernozhukov, Christian Hansen \ Martin Spindler (2015) Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach | 0.405 | 1 | 1 | 100% |
| 8 | Alexandre Belloni \ Victor Chernozhukov (2013) Least squares after model selection in high-dimensional sparse models | 0.405 | 1 | 1 | 100% |
| 9 | Leo Breiman (1998) Arcing Classifiers | 0.405 | 1 | 1 | 100% |
| 10 | Peter Bühlmann \ Bin Yu (2003) Boosting with the $l_2$ Loss: Regression and Classification | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 10 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Semiparametric Estimation of Long-Term Treatment Effects$^*$ | 0.405 | 1 | 1 |