Sokbae Lee, Yuan Liao, Myung Hwan Seo, Youngki Shin
arXiv 29 Sep 2022 · Econometrics · publishedJournal of Econometrics (2024) · 4 citations (OpenAlex)
arXiv:2209.14502 · PDF · DOI · OpenAlex · Extracted main text
Big data analytics has opened new avenues in economic research, but the challenge of analyzing datasets with tens of millions of observations is substantial. Conventional econometric methods based on extreme estimators require large amounts of computing resources and memory, which are often not readily available. In this paper, we focus on linear quantile regression applied to "ultra-large" datasets, such as U.S. decennial censuses. A fast inference framework is presented, utilizing stochastic subgradient descent (S-subGD) updates. The inference procedure handles cross-sectional data sequentially: (i) updating the parameter estimate with each incoming "new observation", (ii) aggregating it as a $Polyak-Ruppert$ average, and (iii) computing a pivotal statistic for inference using only a solution path. The methodology draws from time-series regression to create an asymptotically pivotal statistic through random scaling. Our proposed test statistic is calculated in a fully online fashion and critical values are calculated without resampling. We conduct extensive numerical studies to showcase the computational merits of our proposed inference. For inference problems as large as $(n, d) \sim (10^7, 10^3)$, where $n$ is the sample size and $d$ is the number of regressors, our method generates new insights, surpassing current inference methods in computation. Our method specifically reveals trends in the gender gap in the U.S. college wage premium using millions of observations, while controlling over $10^3$ covariates to mitigate confounding effects.
appendix boundary found by appendix_command · 62% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | He, X., X. Pan, K. M. Tan, and W.-X. Zhou (2023) Smoothed quantile regression with large-scale inference | 1.000 | 9 | 3 | 100% |
| 2 | Lee, S., Y. Liao, M. H. Seo, and Y. Shin (2022) Fast and robust online inference with stochastic gradient descent via random scaling self | 0.941 | 6 | 3 | 83% |
| 3 | Portnoy, S. and R. Koenker (1997) The Gaussian hare and the Laplacian tortoise: computability of squared-error versus absolute-error estimators | 0.874 | 5 | 2 | 100% |
| 4 | Yang, J., X. Meng, and M. Mahoney (2013) Quantile regression for large-scale applications | 0.737 | 3 | 2 | 100% |
| 5 | Kiefer, N. M., T. J. Vogelsang, and H. Bunzel (2000) Simple robust testing of regression hypotheses | 0.737 | 3 | 2 | 100% |
| 6 | Gadat, S. and F. Panloup (2023) Optimal non-asymptotic analysis of the Ruppert-Polyak averaging stochastic algorithm | 0.659 | 7 | 3 | 29% |
| 7 | Koenker, R. and G. Bassett (1978) Regression quantiles | 0.644 | 4 | 1 | 100% |
| 8 | Fernandes, M., E. Guerre, and E. Horta (2021) Smoothing quantile regressions | 0.644 | 2 | 2 | 100% |
| 9 | Ruggles, S., S. Flood, R. Goeken, M. Schouweiler, and M. Sobek (2022) IPUMS USA: Version 12.0 [dataset]. Minneapolis, MN: IPUMS | 0.644 | 2 | 2 | 100% |
| 10 | Fang, Y., J. Xu, and L. Yang (2018) Online bootstrap confidence intervals for the stochastic gradient descent estimator | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 57 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Noisy, Non-Smooth, Non-Convex Estimation of Moment Condition Models | 0.405 | 1 | 1 |
| 2 | SLIM: Stochastic Learning and Inference in Overidentified Models | 0.405 | 1 | 1 |