EconBase
← All papers

Estimating Stochastic Block Models in the Presence of Covariates

Yuichi Kitamura, Louise Laage

arXiv 26 Feb 2024 · Econometrics

arXiv:2402.16322 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In the standard stochastic block model for networks, the probability of a connection between two nodes, often referred to as the edge probability, depends on the unobserved communities each of these nodes belongs to. We consider a flexible framework in which each edge probability, together with the probability of community assignment, are also impacted by observed covariates. We propose a computationally tractable two-step procedure to estimate the conditional edge probabilities as well as the community assignment probabilities. The first step relies on a spectral clustering algorithm applied to a localized adjacency matrix of the network. In the second step, k-nearest neighbor regression estimates are computed on the extracted communities. We study the statistical properties of these estimators by providing non-asymptotic bounds.

Citation extraction

19
references
58
in-text mentions
19
distinct cited
0
self-citations
16,735
main-text words

appendix boundary found by appendix_titled_section at “Supplement: Some Useful results” · 78% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Rohe, Qin, and Yu (2016) Co-clustering directed graphs to discover asymmetries and directional communities1.00083100%
2Lei and Rinaldo (2015) Consistency of spectral clustering in stochastic block models1.00073100%
3Jiang (2019) Non-asymptotic uniform rates of consistency for k-nn regression0.87472100%
4Vershynin (2018) High-dimensional probability: An introduction with applications in data science0.81142100%
5Portier (2021) Nearest neighbor process: weak convergence and non-asymptotic bound0.6597243%
6Vu and Lei (2013) Minimax sparse principal subspace estimation in high dimensions0.58531100%
7Bhatia (2013) Matrix analysis0.40511100%
8Bickel, Chen, and Levina (2011) The method of moments and degree distributions for network models0.40511100%
9Chung and Lu (2006) Complex graphs and networks0.40511100%
10Horn and Johnson (1990) Matrix Analysis0.40511100%

Showing the top 10 of 19 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Flexible Imputation of Incomplete Network Data0.40511
2Post-selection inference for network structure 10.40511