EconBase
← All papers

Scalable Econometrics on Big Data -- The Logistic Regression on Spark

Aurélien Ouattara, Matthieu Bulté, Wan-Ju Lin, Philipp Scholl, Benedikt Veit, Christos Ziakas, Florian Felice, Julien Virlogeux, George Dikos

arXiv 18 Jun 2021 · Statistics — Computation

arXiv:2106.10341 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Extra-large datasets are becoming increasingly accessible, and computing tools designed to handle huge amount of data efficiently are democratizing rapidly. However, conventional statistical and econometric tools are still lacking fluency when dealing with such large datasets. This paper dives into econometrics on big datasets, specifically focusing on the logistic regression on Spark. We review the robustness of the functions available in Spark to fit logistic regression and introduce a package that we developed in PySpark which returns the statistical summary of the logistic regression, necessary for statistical inference.

Citation extraction

43
references
52
in-text mentions
43
distinct cited
1
self-citations
4,598
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Wooldridge, J. M (2010) Econometric Analysis of Cross Section and Panel Data0.73732100%
2Rodriguez, G Introduction to R0.64422100%
3Spark Documentation Classification and Regression - Spark Documentation, f0.64422100%
4Zaharia, M., Chowdhury, N. M. M., Franklin, M., Shenker, S., and Sto… (2010) Spark: Cluster computing with working sets0.51121100%
5Carroll, R., Wang, S., Simpson, D., Stromberg, A., and Ruppert, D (1998) The sandwich (robust covariance matrix) estimator0.51121100%
6L'Ecuyer, P (2012) Random Number Generation0.51121100%
7Fernández-Villaverde, J. and Zarruk Valencia, D (2018) A practical guide to parallelization in economics0.51121100%
8R Core Team (2017) R core - glm.R, October 20170.51121100%
9Avriel, M (2003) Nonlinear Programming: Analysis and Methods0.40511100%
10Bluhm, B. and Cutura, J (2020) Econometrics at scale: Spark up big data in economics0.40511100%

Showing the top 10 of 43 scored citations.