Mingli Chen, Kengo Kato, Chenlei Leng
arXiv 8 Aug 2019 · Mathematics — Statistics Theory · publishedJournal of the Royal Statistical Society Series B (Statistical Methodology) (2021) · 23 citations (OpenAlex)
arXiv:1908.03152 · PDF · DOI · OpenAlex · Extracted main text
Data in the form of networks are increasingly available in a variety of areas, yet statistical models allowing for parameter estimates with desirable statistical properties for sparse networks remain scarce. To address this, we propose the Sparse $\beta$-Model (S$\beta$M), a new network model that interpolates the celebrated Erd\H{o}s-R\'enyi model and the $\beta$-model that assigns one different parameter to each node. By a novel reparameterization of the $\beta$-model to distinguish global and local parameters, our S$\beta$M can drastically reduce the dimensionality of the $\beta$-model by requiring some of the local parameters to be zero. We derive the asymptotic distribution of the maximum likelihood estimator of the S$\beta$M when the support of the parameter vector is known. When the support is unknown, we formulate a penalized likelihood approach with the $\ell_0$-penalty. Remarkably, we show via a monotonicity lemma that the seemingly combinatorial computational problem due to the $\ell_0$-penalty can be overcome by assigning nonzero parameters to those nodes with the largest degrees. We further show that a $\beta$-min condition guarantees our method to identify the true model and provide excess risk bounds for the estimated parameters. The estimation procedure enjoys good finite sample properties as shown by simulation studies. The usefulness of the S$\beta$M is further illustrated via the analysis of a microfinance take-up example.
appendix boundary found by appendix_command · 70% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Krivitsky, P. N. and E. D. Kolaczyk (2015) On the question of effective sample size in network modeling: An asymptotic inquiry | 1.000 | 5 | 3 | 100% |
| 2 | Banerjee, A., A. G. Chandrasekhar, E. Duflo, and M. O. Jackson (2013) The diffusion of microfinance | 0.874 | 5 | 2 | 100% |
| 3 | Chatterjee, S., P. Diaconis, and A. Sly (2011) Random graphs with a given degree sequence | 0.874 | 5 | 2 | 100% |
| 4 | Mukherjee, R., S. Mukherjee, and S. Sen (2019) Detection thresholds for the $$-model on sparse graphs | 0.811 | 4 | 2 | 100% |
| 5 | Greenshtein, E. and Y. Ritov (2004) Persistence in high-dimensional linear predictor selection and the virtue of overparametrization | 0.737 | 3 | 2 | 100% |
| 6 | Yan, T. and J. Xu (2013) A central limit theorem in the $$-model for undirected random graphs with a diverging number of vertices | 0.737 | 3 | 2 | 100% |
| 7 | Fan, Y. and C. T. Tang (2013) Tuning parameter selection in high dimensional penalized likelihood | 0.644 | 4 | 1 | 100% |
| 8 | Holland, P. W. and S. Leinhardt (1981) An exponential family of probability distributions for directed graphs | 0.644 | 2 | 2 | 100% |
| 9 | Newman, M (2018) Networks (2nd Edition) | 0.585 | 3 | 1 | 100% |
| 10 | Chen, J. and Z. Chen (2008) Extended bayesian information criterion for model selection with large model space | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 56 scored citations.