EconBase
← All papers

Fast Inference for Quantile Regression with Tens of Millions of Observations

Sokbae Lee, Yuan Liao, Myung Hwan Seo, Youngki Shin

arXiv 29 Sep 2022 · Econometrics · publishedJournal of Econometrics (2024) · 4 citations (OpenAlex)

arXiv:2209.14502 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Big data analytics has opened new avenues in economic research, but the challenge of analyzing datasets with tens of millions of observations is substantial. Conventional econometric methods based on extreme estimators require large amounts of computing resources and memory, which are often not readily available. In this paper, we focus on linear quantile regression applied to "ultra-large" datasets, such as U.S. decennial censuses. A fast inference framework is presented, utilizing stochastic subgradient descent (S-subGD) updates. The inference procedure handles cross-sectional data sequentially: (i) updating the parameter estimate with each incoming "new observation", (ii) aggregating it as a $Polyak-Ruppert$ average, and (iii) computing a pivotal statistic for inference using only a solution path. The methodology draws from time-series regression to create an asymptotically pivotal statistic through random scaling. Our proposed test statistic is calculated in a fully online fashion and critical values are calculated without resampling. We conduct extensive numerical studies to showcase the computational merits of our proposed inference. For inference problems as large as $(n, d) \sim (10^7, 10^3)$, where $n$ is the sample size and $d$ is the number of regressors, our method generates new insights, surpassing current inference methods in computation. Our method specifically reveals trends in the gender gap in the U.S. college wage premium using millions of observations, while controlling over $10^3$ covariates to mitigate confounding effects.

Citation extraction

56
references
112
in-text mentions
57
distinct cited
3
self-citations
10,695
main-text words

appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1He, X., X. Pan, K. M. Tan, and W.-X. Zhou (2023) Smoothed quantile regression with large-scale inference1.00093100%
2Lee, S., Y. Liao, M. H. Seo, and Y. Shin (2022) Fast and robust online inference with stochastic gradient descent via random scaling self0.9416383%
3Portnoy, S. and R. Koenker (1997) The Gaussian hare and the Laplacian tortoise: computability of squared-error versus absolute-error estimators0.87452100%
4Yang, J., X. Meng, and M. Mahoney (2013) Quantile regression for large-scale applications0.73732100%
5Kiefer, N. M., T. J. Vogelsang, and H. Bunzel (2000) Simple robust testing of regression hypotheses0.73732100%
6Gadat, S. and F. Panloup (2023) Optimal non-asymptotic analysis of the Ruppert-Polyak averaging stochastic algorithm0.6597329%
7Koenker, R. and G. Bassett (1978) Regression quantiles0.64441100%
8Fernandes, M., E. Guerre, and E. Horta (2021) Smoothing quantile regressions0.64422100%
9Ruggles, S., S. Flood, R. Goeken, M. Schouweiler, and M. Sobek (2022) IPUMS USA: Version 12.0 [dataset]. Minneapolis, MN: IPUMS0.64422100%
10Fang, Y., J. Xu, and L. Yang (2018) Online bootstrap confidence intervals for the stochastic gradient descent estimator0.64422100%

Showing the top 10 of 57 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Noisy, Non-Smooth, Non-Convex Estimation of Moment Condition Models0.40511
2SLIM: Stochastic Learning and Inference in Overidentified Models0.40511