EconBase
← All papers

pystacked: Stacking generalization and machine learning in Stata

Achim Ahrens, Christian B. Hansen, Mark E. Schaffer

arXiv 23 Aug 2022 · Econometrics · publishedThe Stata Journal Promoting communications on statistics and Stata (2023) · 31 citations (OpenAlex)

arXiv:2208.10896 · PDF · DOI · OpenAlex · Extracted main text

Abstract

pystacked implements stacked generalization (Wolpert, 1992) for regression and binary classification via Python's scikit-learn. Stacking combines multiple supervised machine learners -- the "base" or "level-0" learners -- into a single learner. The currently supported base learners include regularized regression, random forest, gradient boosted trees, support vector machines, and feed-forward neural nets (multi-layer perceptron). pystacked can also be used with as a `regular' machine learning program to fit a single base learner and, thus, provides an easy-to-use API for scikit-learn's machine learning algorithms.

Citation extraction

25
references
29
in-text mentions
25
distinct cited
2
self-citations
7,316
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Ahrens, A., C. B. Hansen, and M. E. Schaffer (2020) lassopack: Model selection and prediction with regularized regression in Stata self0.64422100%
2Breiman, L (1996) Stacked regressions0.64422100%
3Hastie, T., R. Tibshirani, and J. Friedman (2009) The Elements of Statistical Learning0.58531100%
4Athey, S., and G. W. Imbens (2019) Machine learning methods that economists should know about0.40511100%
5Chernozhukov, V., D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W… (2018) Double/debiased machine learning for treatment and structural parameters0.40511100%
6Droste, M (2020) pylearn0.40511100%
7Guenther, N., and M. Schonlau (2018) SVMACHINES: Stata module providing Support Vector Machines for both Classification and Regression0.40511100%
8Huntington-Klein, N. C (2021) mlrtime0.40511100%
9Pace, R. K., and R. Barry (1997) Sparse spatial autoregressions0.40511100%
10van der Laan, M. J., E. C. Polley, and A. E. Hubbard (2007) Super Learner0.40511100%

Showing the top 10 of 25 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Nonparametric rich covariateswithout saturation0.64422
2Model Averaging and Double Machine Learning0.40511
3Hyperparameter Tuning for Causal Inference with Double Machine Learning: A Simulation Study0.40511