EconBase
← All papers

Sparse network asymptotics for logistic regression

Bryan S. Graham

arXiv 9 Oct 2020 · Econometrics · 4 citations (OpenAlex)

arXiv:2010.04703 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Consider a bipartite network where $N$ consumers choose to buy or not to buy $M$ different products. This paper considers the properties of the logistic regression of the $N\times M$ array of i-buys-j purchase decisions, $\left[Y_{ij}\right]_{1\leq i\leq N,1\leq j\leq M}$, onto known functions of consumer and product attributes under asymptotic sequences where (i) both $N$ and $M$ grow large and (ii) the average number of products purchased per consumer is finite in the limit. This latter assumption implies that the network of purchases is sparse: only a (very) small fraction of all possible purchases are actually made (concordant with many real-world settings). Under sparse network asymptotics, the first and last terms in an extended Hoeffding-type variance decomposition of the score of the logit composite log-likelihood are of equal order. In contrast, under dense network asymptotics, the last term is asymptotically negligible. Asymptotic normality of the logistic regression coefficients is shown using a martingale central limit theorem (CLT) for triangular arrays. Unlike in the dense case, the normality result derived here also holds under degeneracy of the network graphon. Relatedly, when there happens to be no dyadic dependence in the dataset in hand, it specializes to recently derived results on the behavior of logistic regression with rare events and iid data. Sparse network asymptotics may lead to better inference in practice since they suggest variance estimators which (i) incorporate additional sources of sampling variation and (ii) are valid under varying degrees of dyadic dependence.

Citation extraction

37
references
78
in-text mentions
37
distinct cited
5
self-citations
8,958
main-text words

appendix boundary found by appendix_command · 77% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Wang, H (2020) Pmlr1.00083100%
2Graham, B. S (2020) Handbook of Econometrics, volume 7, chapter Network data self1.00054100%
3Menzel, K (2017) Bootstrap with clustering in two or more dimensions1.00053100%
4Bickel, P. J., Chen, A., and Levina, E (2011) The method of moments and degree distributions for network models0.92843100%
5Davezies, L., d'Haultfoeuille, X., and Guyonvarch, Y (2020) Empirical process results for exchangeable arrayes0.84333100%
6Graham, B. S (2020) The Econometrics of Social and Economic Networks, chapter Dyadic regression, pages 25 – 41 self0.84333100%
7King, G. and Zeng, L (2001) Logistic regression in rare events data0.81142100%
8Aldous, D. J (1981) Representations for partially exchangeable arrays of random variables0.73732100%
9Aronow, P. M., Samii, C., and Assenova, V. A (2017) Cluster-robust variance estimation for dyadic data0.73732100%
10Cameron, A. C. and Miller, D. L (2014) Robust inference for dyadic data0.73732100%

Showing the top 10 of 37 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Robust Inference in Locally Misspecified Bipartite Networks0.965105
2Multiway empirical likelihood0.85584
31 Linear Regression with Centrality Measures0.84333
4Detecting Latent Communities in Network Formation Models0.40511
5Minimax Risk and Uniform Convergence Rates for Nonparametric Dyadic Regression0.40511
6Asymptotic Theory for Two-Way Clustering0.40511
7Functional Differencing in Networks0.40511
8A NEYMAN-ORTHOGONALIZATION APPROACH TO THE INCIDENTAL PARAMETER PROBLEM0.40511