arXiv 18 Oct 2022 · Econometrics · 2 citations (OpenAlex)
arXiv:2210.10024 · PDF · DOI · OpenAlex · Extracted main text
This paper studies the properties of linear regression on centrality measures when network data is sparse -- that is, when there are many more agents than links per agent -- and when they are measured with error. We make three contributions in this setting: (1) We show that OLS estimators can become inconsistent under sparsity and characterize the threshold at which this occurs, with and without measurement error. This threshold depends on the centrality measure used. Specifically, regression on eigenvector is less robust to sparsity than on degree and diffusion. (2) We develop distributional theory for OLS estimators under measurement error and sparsity, finding that OLS estimators are subject to asymptotic bias even when they are consistent. Moreover, bias can be large relative to their variances, so that bias correction is necessary for inference. (3) We propose novel bias correction and inference methods for OLS with sparse noisy networks. Simulation evidence suggests that our theory and methods perform well, particularly in settings where the usual OLS estimators and heteroskedasticity-consistent/robust t-tests are deficient. Finally, we demonstrate the utility of our results in an application inspired by De Weerdt and Deacon (2006), in which we consider consumption smoothing and social insurance in Nyakatoke, Tanzania.
appendix boundary found by appendix_command · 54% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Cai, J., D. Yang, W. Zhu, H. Shen, and L. Zhao (2021) Network regression and supervised centrality estimation | 1.000 | 5 | 3 | 100% |
| 2 | Le, C. M. and T. Li (2020) Linear regression and its inference on noisy network-linked data | 1.000 | 5 | 3 | 100% |
| 3 | Avella-Medina, M., F. Parise, M. T. Schaub, and S. Segarra (2020) Centrality measures for graphons: Accounting for uncertainty in networks | 0.928 | 5 | 4 | 80% |
| 4 | De Weerdt, J. and S. Dercon (2006) Risk-sharing networks and insurance against illness | 0.928 | 4 | 3 | 100% |
| 5 | Graham, B. S (2020) Sparse network asymptotics for logistic regression | 0.843 | 3 | 3 | 100% |
| 6 | Banerjee, A., A. G. Chandrasekhar, E. Duflo, and M. O. Jackson (2013) The diffusion of microfinance | 0.737 | 3 | 2 | 100% |
| 7 | Bickel, P. J. and A. Chen (2009) A nonparametric view of network models and Newman–Girvan and other modularities | 0.737 | 3 | 2 | 100% |
| 8 | Alt, J., R. Ducatez, and A. Knowles (2021) Poisson statistics and localization at the spectral edge of sparse Erdos–Rényi graphs | 0.644 | 3 | 2 | 67% |
| 9 | Bickel, P. J., A. Chen, and E. Levina (2011) The method of moments and degree distributions for network models | 0.644 | 2 | 2 | 100% |
| 10 | Crippa, F (2025) Identification, estimation, and inference in two-sided interaction models | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 68 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Endogenous Interference in Randomized Experiments | 0.843 | 3 | 3 |
| 2 | Flexible Imputation of Incomplete Network Data | 0.644 | 2 | 2 |
| 3 | On the Asymptotic Properties of Debiased Machine Learning Estimators | 0.405 | 1 | 1 |
| 4 | Robust Market Interventions | 0.405 | 1 | 1 |