Aurélien Ouattara, Matthieu Bulté, Wan-Ju Lin, Philipp Scholl, Benedikt Veit, Christos Ziakas, Florian Felice, Julien Virlogeux, George Dikos
arXiv 18 Jun 2021 · Statistics — Computation
arXiv:2106.10341 · PDF · DOI · OpenAlex · Extracted main text
Extra-large datasets are becoming increasingly accessible, and computing tools designed to handle huge amount of data efficiently are democratizing rapidly. However, conventional statistical and econometric tools are still lacking fluency when dealing with such large datasets. This paper dives into econometrics on big datasets, specifically focusing on the logistic regression on Spark. We review the robustness of the functions available in Spark to fit logistic regression and introduce a package that we developed in PySpark which returns the statistical summary of the logistic regression, necessary for statistical inference.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Wooldridge, J. M (2010) Econometric Analysis of Cross Section and Panel Data | 0.737 | 3 | 2 | 100% |
| 2 | Rodriguez, G Introduction to R | 0.644 | 2 | 2 | 100% |
| 3 | Spark Documentation Classification and Regression - Spark Documentation, f | 0.644 | 2 | 2 | 100% |
| 4 | Zaharia, M., Chowdhury, N. M. M., Franklin, M., Shenker, S., and Sto… (2010) Spark: Cluster computing with working sets | 0.511 | 2 | 1 | 100% |
| 5 | Carroll, R., Wang, S., Simpson, D., Stromberg, A., and Ruppert, D (1998) The sandwich (robust covariance matrix) estimator | 0.511 | 2 | 1 | 100% |
| 6 | L'Ecuyer, P (2012) Random Number Generation | 0.511 | 2 | 1 | 100% |
| 7 | Fernández-Villaverde, J. and Zarruk Valencia, D (2018) A practical guide to parallelization in economics | 0.511 | 2 | 1 | 100% |
| 8 | R Core Team (2017) R core - glm.R, October 2017 | 0.511 | 2 | 1 | 100% |
| 9 | Avriel, M (2003) Nonlinear Programming: Analysis and Methods | 0.405 | 1 | 1 | 100% |
| 10 | Bluhm, B. and Cutura, J (2020) Econometrics at scale: Spark up big data in economics | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 43 scored citations.