Sokbae Lee, Yuan Liao, Myung Hwan Seo, Youngki Shin
arXiv 6 Jun 2021 · Statistics — Machine Learning · publishedProceedings of the AAAI Conference on Artificial Intelligence (2022) · 19 citations (OpenAlex)
arXiv:2106.03156 · PDF · DOI · OpenAlex · Extracted main text
We develop a new method of online inference for a vector of parameters estimated by the Polyak-Ruppert averaging procedure of stochastic gradient descent (SGD) algorithms. We leverage insights from time series regression in econometrics and construct asymptotically pivotal statistics via random scaling. Our approach is fully operational with online data and is rigorously underpinned by a functional central limit theorem. Our proposed inference method has a couple of key advantages over the existing methods. First, the test statistic is computed in an online fashion with only SGD iterates and the critical values can be obtained without any resampling methods, thereby allowing for efficient implementation suitable for massive online data. Second, there is no need to estimate the asymptotic variance and our inference method is shown to be robust to changes in the tuning parameters for SGD algorithms in simulation experiments with synthetic data.
appendix boundary found by appendix_command · 76% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Zhu, W., X. Chen, and W. B. Wu (2021) Online covariance matrix estimation in stochastic gradient descent | 1.000 | 10 | 3 | 100% |
| 2 | Polyak, B. T. and A. B. Juditsky (1992) Acceleration of stochastic approximation by averaging | 0.874 | 10 | 2 | 100% |
| 3 | Chen, X., J. D. Lee, X. T. Tong, and Y. Zhang (2020) Statistical inference for model parameters in stochastic gradient descent | 0.811 | 4 | 2 | 100% |
| 4 | Kiefer, N. M., T. J. Vogelsang, and H. Bunzel (2000) Simple robust testing of regression hypotheses | 0.811 | 4 | 2 | 100% |
| 5 | Lazarus, E., D. J. Lewis, J. H. Stock, and M. W. Watson (2018) Har inference: Recommendations for practice | 0.644 | 2 | 2 | 100% |
| 6 | Zhu, Y. and J. Dong (2020) On constructing confidence region for model parameters in stochastic gradient descent via batch means | 0.585 | 3 | 1 | 100% |
| 7 | Abadir, K. M. and P. Paruolo (1997) Two mixed normal densities from cointegration analysis | 0.511 | 2 | 1 | 100% |
| 8 | Godichon-Baggioni, A (2017) A central limit theorem for averaged stochastic gradient algorithms in hilbert spaces and online estimation of the asymptotic va… | 0.511 | 2 | 1 | 100% |
| 9 | Kushner, H. J. and J. Yang (1993) Stochastic approximation with averaging of the iterates: Optimal asymptotic rate of convergence for general processes | 0.511 | 2 | 1 | 100% |
| 10 | Su, W. J. and Y. Zhu (2018) Uncertainty quantification for online learning and stochastic approximation via hierarchical incremental gradient descent | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 30 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.