arXiv 15 Jul 2020 · Statistics — Machine Learning
arXiv:2007.07781 · PDF · DOI · OpenAlex · Extracted main text
Researchers may perform regressions using a sketch of data of size $m$ instead of the full sample of size $n$ for a variety of reasons. This paper considers the case when the regression errors do not have constant variance and heteroskedasticity robust standard errors would normally be needed for test statistics to provide accurate inference. We show that estimates using data sketched by random projections will behave `as if' the errors were homoskedastic. Estimation by random sampling would not have this property. The result arises because the sketched estimates in the case of random projections can be expressed as degenerate $U$-statistics, and under certain conditions, these statistics are asymptotically normal with homoskedastic variance. We verify that the conditions hold not only in the case of least squares regression when the covariates are exogenous, but also in instrumental variables estimation when the covariates are endogenous. The result implies that inference, including first-stage F tests for instrument relevance, can be simpler than the full sample case if the sketching scheme is appropriately chosen.
appendix boundary found by appendix_command · 54% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Ma, Zhang, Xing, Ma, and Mahoney (2020) Asymptotic Analysis of Sampling Estimators for Randomized Numerical Linear Algebra Algorithms | 1.000 | 5 | 3 | 100% |
| 2 | Lee and Ng (2020) An Econometric Perspective on Algorithmic Subsampling | 0.644 | 2 | 2 | 100% |
| 3 | Ahfock, Astle, and Richardson (2020) Statistical properties of sketching algorithms | 0.644 | 2 | 2 | 100% |
| 4 | Hall (1984) Central limit theorem for integrated square error of multivariate nonparametric density estimators | 0.511 | 3 | 2 | 33% |
| 5 | Woodruff (2014) Sketching as a tool for numerical linear algebra | 0.511 | 2 | 2 | 50% |
| 6 | Angrist and Krueger (1991) Does compulsory school attendance affect schooling and earnings? | 0.511 | 2 | 1 | 100% |
| 7 | Angrist and Krueger (1992) The Effect of Age at School Entry on Educational Attainment: An Application of Instrumental Variables with Moments from Two Samp… | 0.405 | 1 | 1 | 100% |
| 8 | Angrist and Krueger (1995) Split-Sample Instrumental Variables Estimates of the Return to Schooling | 0.405 | 1 | 1 | 100% |
| 9 | Inoue and Solon (2010) Two-Sample Instrumental Variables Estimators | 0.405 | 1 | 1 | 100% |
| 10 | Liu and Dobriban (2020) Ridge Regression: Structure, Cross-Validation, and Sketching | 0.405 | 1 | 1 | 100% |
Showing the top 10 of 28 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Fast Inference for Quantile Regression with Tens of Millions of Observations | 0.405 | 1 | 1 |
| 2 | On Using The Two-Way Cluster-Robust Standard Errors | 0.405 | 1 | 1 |
| 3 | SLIM: Stochastic Learning and Inference in Overidentified Models | 0.000 | 1 | 1 |