Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
134,254 characters · 20 sections · 78 citation commands
1 Linear Regression with Centrality Measures
\noindentKeywords: networks, diffusion centrality, eigenvector centrality.
\noindentJEL Classification Codes: C18, C21, C81.
Linear regression with centrality measures is a common approach for studying the relationship between network position and economic outcomes. In such analyses, researchers regress an outcome of interest on centrality measures, which are node-level statistics that capture an agent’s importance in a network. For example, hochberg2007whom studies the network of venture capital firms and finds that better-connected firms successfully exit a greater proportion of their investments. cruz2017politician examines social networks in the Philippines and shows that more central families are disproportionately represented in political office. Similarly, banerjee2013diffusion studies the diffusion of microfinance in India and shows that seeding information to more central agents leads to greater program participation.
Across such papers, a variety of centrality measures are used to capture different aspects of network positions. For example, the degree centrality of an agent reflects the number or intensity of their direct links, while eigenvector centrality is designed so that the influence of agents is proportional to that of their connections. The correlation between an outcome variable and a particular centrality measure may therefore reveal the types of interactions that drive a given economic phenomenon: an outcome that is well-predicted solely by degree is likely to be determined in a highly local manner, whereas outcomes that correlate more strongly with eigenvector centrality may reflect interactions that propagate through longer network paths. As such, when researchers estimate these correlations and test their statistical significance, they frequently do so with the goal of drawing conclusions about the economic significance of various centrality measures and the implied mechanisms for outcome determination. Such an exercise is credible provided that standard OLS asymptotic theory provides good approximations in finite samples.
However, network data have two features that may threaten the statistical validity of OLS. First, networks may be sparse, with many more agents than links per agent. This could happen because interactions are observed with low frequency, or because the interactions in question are rare. chandrasekhar2016formation argues that many economic networks are sparse, providing evidence from commonly used social network datasets (e.g. AddHealth; Karnataka Villages (banerjee2013diffusion); Harvard social network (leider2009directed)). Sparsity poses a challenge to estimation and inference: intuitively, a largely empty network contains little information about the parameters of interest.
Second, the observed network may differ from the true network of interest. Centrality measures are often calculated on data which are obtained by survey or constructed using some proxy for interaction between agents, though subsequent analysis would frequently treat the true network as known. Ignoring observation error may lead to estimates that perform poorly. A growing literature works with networks that are assumed to be observed with error, and increasingly, in combination with sparsity. This is important since the two issues are mutually reinforcing: sparser networks contain weaker signals, which are in turn more difficult to pick out from noisy measurements. The upshot is that OLS estimators computed on sparse, noisy networks may have particularly poor properties. Asymptotic theory that ignores these features will provide similarly poor approximations to their finite sample behavior. Estimation and inference procedures based on these theories may, in turn, lead to invalid conclusions about the economic significance of centrality measures.
To understand the effect of observation error and sparsity, this paper adopts the asymptotic framework of bickel2009nonparametric, featuring a single growing network. Here, agents are endowed with latent types and a graphon (graph function) maps pairs of latent types into link intensities. This is the true network that is unobserved. Instead, researchers observe a single draw from the graphon, in which each link is an independent Bernoulli random variable with success probability given by the link intensity. The difference between the two networks is our observation error. To reflect sparsity, the Bickel–Chen model introduces a scale parameter for the graphon that can shrink it towards zero as sample size increases. As such, the expected number of links per agent can grow much more slowly than $n$ as it tends to infinity.
We study OLS estimation with centrality measures in the Bickel-Chen model, focusing on degree, diffusion and eigenvector centralities. We make the following contributions: (1) We characterize the amount of sparsity at which OLS estimators become inconsistent under observation error, finding that this threshold varies depending on the centrality measure used. Specifically, regression on eigenvector centrality is less robust to sparsity than regressions on degree and diffusion. This suggests that researchers should be cautious about comparing regressions on different centrality measures, since they may differ in statistical properties in addition to economic significance. (2) We develop distributional theory for OLS estimators under observation error and sparsity. We restrict ourselves to sparsity regimes under which OLS is consistent and find that it has an asymptotic bias that can be large when the target coefficient is non-zero. We obtain sharp rates for the asymptotic bias and show that they can be large compared to the variance of the OLS estimators. This implies that the usual confidence intervals do not have the desired coverage probabilities and the $t$-test is invalid for hypothesis of the form $H_0: \beta = b$ when $b \neq 0$. (3) We propose bias-corrected estimators for degree and diffusion that are competitive with the no observation error benchmark. For eigenvector, we propose to conduct inference by widening confidence intervals using an asymptotic bound for the bias. The procedure is conservative when the network is sparse, but exact once the network is moderately dense.
Because our statistical model captures important features of real-world data, we expect our theory and methods to be more appropriate for settings with observation error and sparsity. We provide simulation evidence supporting this view. Finally, we demonstrate the empirical relevance of our results in an application inspired by deweerdt2006risk, in which we conduct a stylized study of consumption smoothing and social insurance in Nyakatoke, Tanzania.
This work is most closely related to papers that study linear regression with centrality measures. thirkettle2019identification constructs proposes bounds for linear regression parameters with degree and diffusion centrality under partial sampling of nodes and edges. Several papers consider linear regression with eigenvector centrality. le2020linear studies regression on multiple eigenvectors also in the Bickel-Chen framework. Their conditions are denser than necessary for consistency. We obtain sharp thresholds. cai2021network focuses on rank-1 networks that are observed with an additive, Gaussian error. In contrast, the Bickel-Chen model features Bernoulli noise and networks with arbitrary rank. However, we will assume finite rank for some results (more details in Remark (ref)). cheng2021tackling considers inference on linear functionals of eigenvectors. The uncertainty that drives their asymptotic results is dominated by the regression error in our setting.
Our paper also speaks to the literature on the estimation of centrality measures in the presence of noise. Early work provided simulation evidence that centrality measures on noisy networks become less accurate as sparsity increases costenbader2003stability, borgatti2006robustness. segarra2015stability theoretically studies the stability of network statistics under perturbations, finding that degree and eigenvector centralities are stable, while betweenness is not. avella2020centrality and dasaratha2020distributions consider settings similar to ours and provide concentration results for degree and eigenvector centralities among others, but not for diffusion centrality. Additionally, they accommodate less sparsity than us, in part because we are not concerned with estimation of centrality measures, only their subsequent use in linear regression.
There is also a large literature involving other types of observation error in network data. In a setting with partial sampling of nodes and edges, chandrasekhar2016econometrics characterizes the bias in linear regression on network-level statistics in a setting with many networks. thirkettle2019identification provides bounds for centrality statistics and node-level regression parameters in a similar setting. For peer effects regressions, griffith2022name examines the effects of selective censoring, while lewbel2025estimating considers for link misclassification. lewbel2021social examines more general forms of error under the assumption that they vanish asymptotically. A separate body of work assumes that the network is completely unobserved manresa2016estimating, rose2016identification,lewbel2023social,depaula2025identifying. With the exception of thirkettle2019identification, the aforementioned papers do not consider centrality regressions.
Finally, this work is also connected to papers that emphasizes the statistical challenges arising from sparse networks, whether with network formation models (e.g. dePaula2018identifying, jochmans2018semiparametric, graham2020sparse,menzel2022strategic, leung2025normal) or models of two-sided interactions (e.g. jochmans2019fixed,verdier2020estimation, crippa2025identification). We derive the implications of sparsity for centrality regressions.
The rest of this paper is organized as follows. Section (ref) describes the set-up of our paper. Section (ref) presents the theoretical results. Simulation results are contained in Section (ref). In Section (ref), we apply our results to the social insurance network in Nyakatoke, Tanzania. Section (ref) concludes. Proofs are contained in the appendix.
When $X$ is a square matrix, $\lambda_j(X)$ denotes the $j^\text{th}$ eigenvalue of $X$ while $v_j(X)$ denotes the corresponding eigenvector. When $f \in L^2([0,1]^2)$ is a symmetric real function, $\lambda_j(f)$ denotes the $j^{\text{th}}$ eigenvalue of the corresponding Hilbert-Schmidt integral operator, $T_f(g) = \int f(x,y) g(y) dy$, while $\phi_j$ is the corresponding eigenfunction. For deterministic, monotone sequences $x_n$ and $y_n$, we write $x_n \gg z_n$ if $x_n/z_n \to \infty$ and $x_n \ll z_n$ if $x_n/z_n \to 0$. $x_n \approx z_n$ indicates that $x_n/z_n \to k$, where $0 < k <\infty$. We write $x_n \gtrsim z_n$ to mean $\neg (x_n \ll z_n)$ and similarly for $x_n \lesssim z_n$. Let $\iota_n$ be the $n \times 1$ vector of $1$'s. For two $m\times n$ matrices $X$ and $Z$, let $X \circ Z$ denote their entrywise (Hadamard) product. Finally, $[n]$ denotes the set of integers from $1$ to $n$.
This section outlines the framework for our analysis. Section (ref) states the linear regression model. Section (ref) describes the data generating process for the true and observed networks. Section (ref) motivates the sparse network asymptotics of bickel2009nonparametric. Section (ref) defines the centrality measures.
Consider the regression:
where $Y_i$ is the outcome of interest and $C^{(d)}_i$ is a network centrality measure of type $d$. $W_i \in \mathbb{R}^L$ is a vector of other covariates. As will become clear in the next section, $W_i$ is allowed to be correlated with the true network but not with the observation errors. We assume that researchers observe $\{(Y_i, W_i)\}_{i=1}^n$ and either the adjacency matrix $A$ or a noisy version $\hat{A}$. $A$ is an $n \times n$ matrix whose $(i,j)^\text{th}$ entry, $A_{ij}$, records the link intensity between agents $i$ and $j$. $\hat{A}$ is some estimate of $A$. $C^{(d)}_i$ is not directly observed but it can be computed exactly using $A$, or estimated using $\hat{A}$. The parameter of interest is $\beta^{(d)}$, which we interpret as the slope coefficient in the linear conditional expectation function of $Y_i$ on $C_i^{(d)}$ (see Assumption (ref)).
We assume that the data-generating process yields $\{(\varepsilon_i, W_i, U_i)\}_{i=1}^n$ which are independent and identically distributed. $\varepsilon_i$ is the regression error and $U_i$ is an unobserved latent type that will be used to construct the network. Let $\mathbf{Y}, \mathbf{C}^{(d)}, \bm{\varepsilon}$ and $\mathbf{U}$ be vectors which stack the corresponding variables. Similarly for the matrix $\mathbf{W}$.
In the following sections, we describe (i) the data-generating process for $A$ and $\hat{A}$ via the $U_i$'s, (ii) the Bickel-Chen framework for sparse network asymptotics, and (iii) the use of $A$ and $\hat{A}$ in computing/estimating centrality statistics for OLS estimation. Throughout our discussion, we motivate the econometric framework through the example of consumption smoothing via informal insurance:
Let $A$ be an $n \times n$ symmetric adjacency matrix. We assume that the relationship between two agents in a network is solely determined by their unobserved latent types $U_i$ through the graphon $f$:
In this model, any two agents have a relationship that is between $0$ and $1$, which we interpret as the underlying intensity of interaction between agents. It reflects factors such as the duration of friendship, the frequency of communication, or similarity in characteristics, and is the quantity that is relevant for the outcome variable. The parameter $p_n$ governs the overall density of the network and will be allowed to converge to zero at various rates. As we explain in the next section, it serves as a theoretical device for studying the behavior of OLS estimators when networks are sparse. Finally, we restrict attention to symmetric adjacency matrices because eigenvector centrality is not generally well-defined otherwise. Note that $f$ is not assumed to be known.
When $A$ is observed, we say that there is no observation error. However, researchers do not observe the latent interaction intensity directly in many empirical settings. Instead, they observe some proxy for $A$. We model the observed network, $\hat{A}$, as a matrix of Bernoulli random variables whose success probabilities are given by the latent intensities:
Conditional on $\mathbf{W}$ and $\mathbf{U}$, observation errors are additive, have mean $0$, and are independent across agent pairs. Since we assume conditional independence between $\varepsilon_i$ and $\xi_{ij}$, our model of observation error is akin to classical measurement error. However, additive errors on $\hat{A}_{ij}$ can become non-linear errors on the centrality measures. The covariates $W_i$ are uncorrelated with $\xi_{jk}$ for all $i,j,k$. However, they are allowed to be correlated with $U_i$ and hence the network.
To better capture the behavior of estimators when agents in networks have few relationships with one another, we study their properties under sparse network asymptotics. Following bickel2009nonparametric, this entails letting $p_n \to 0$ as $n \to \infty$.
A vector or matrix is typically said to be sparse if many of its entries are zero. In our setting, we say that $A$ and $\hat{A}$ are sparse if their row sums -- that is, total interactions of agents, actual or observed -- are small. Because the entries of $\hat{A}$ are restricted to be binary, this is the same as having many entries which are $0$. Without such a restriction, the row sums of $A$ could be small even if no entry takes value $0$, as long as non-zero entries are small. In our setting, sparsity should therefore be understood as low interaction intensities in the true network, $A$, but which gives rise to observed networks, $\hat{A}$, that are sparse in the conventional sense.
To see how $p_n \to 0$ gives rise to sparsity, suppose $p_n \to c > 0$. Then $\sum_{j} \mathbb{E}[A_{ij}] = O(n)$. The network is said to be dense since each agent is linked to a non-vanishing fraction of the sample. In practice, however, researchers may observe sparse networks, in which each agent has few or weak relationships. This is captured by allowing $p_n $ to approach $0$ as $n$ increases. For example, if we set $p_n = k/n$ for some $k > 0$, then $\sum_{j} \mathbb{E}[A_{ij}] = O(1)$. Each agent then has a bounded number of relationships in expectation even as $n \to \infty$. We say that the network is sparse whenever $p_n \to 0$. The rate at which this occurs corresponds to the amount of sparsity in the data. We will not impose any particular rate on $p_n$, but instead compare centrality measures across sparsity regimes to shed light on the differences between them.
Since network data is often sparse, the Bickel-Chen framework leads to asymptotic approximations that capture more features of observed data. Theory developed in such a framework should therefore better reflect the statistical properties of OLS estimators applied to sparse data. Within econometrics, related approaches have been used to study network formation models (e.g. dePaula2018identifying, jochmans2018semiparametric, graham2020sparse,menzel2022strategic, leung2025normal) and models of two-sided matching (e.g. jochmans2019fixed,verdier2020estimation, crippa2025identification) among others.
The Bickel-Chen model has become a standard data-generating process for network data in nonparametric statistics and machine learning (graham2020network; also see discussion in de2017econometrics). While papers working in this framework are sometimes agnostic about the interpretation of $A$, we interpret it as the matrix of true relationships. By “true", we mean that these intensities capture the underlying interaction that is relevant for determining the outcome variable $Y_i$. This perspective is consistent with avella2020centrality, which argues that centrality measures should be defined with respect to the latent network rather than its noisy realization. Related work studying regressions with eigenvector centrality, such as le2020linear and cai2021network, also adopts a similar interpretation. More generally, this view is consistent with a broad class of models in which outcomes depend on the network only through the latent network types of agents. In a stochastic block model (see Example (ref)), for instance, it implies that outcomes depend solely on block membership, in line with the grouped fixed-effects framework of bonhomme2015grouped. This interpretation also fits models in which network data is used to infer unobserved heterogeneity, as in auerbach2022identification and xu2025networks.
In principle, there are multiple ways of modeling observation error and sparsity. A key motivation for using the Bickel-Chen framework is analytical tractability. Because $\hat{A}$ is a sparse inhomogeneous Erdos-Renyi graph conditional on $\mathbf{W}$ and $\mathbf{U}$, we are able to benefit from a large literature on random graph theory.
Nonetheless, the model is a reasonable description of network data. In particular, it is consistent with the highly influential weak ties theory of social networks granovetter1973strength. This theory posits that low intensity links, which constitute most of any given person's relationships, are the key drivers of many important social and economic outcomes.\footnote{In tracing the network of job referrals, granovetter1973strength finds that 83% of recent job changers in a Boston suburb found their new jobs through friends whom they saw fewer than twice a week, and who were only “marginally included in the current network of contacts". The author further notes: “It is remarkable that people receive crucial information from individuals whose very existence they have forgotten." A series of empirical work has found evidence in favor of the weak ties theory across diverse applications such as innovation (reagans2001networks), economic development (eagle2010network) and job referrals (rajkumar2022causal).} Such a view aligns with our model, in which the economically meaningful interactions are numerous and weak. Researchers observe sparse reported connections, which may correspond to strong ties, but are useful only to the extent that they are informative about the weak ties.
Although our proposed model may not describe every type of economic interaction on networks, it is a plausible model that researchers may not want to rule out ex ante. In that case, the issues that we highlight in the following sections remain relevant.
We next define centrality statistics and the OLS estimators that are based on them. Centrality measures are agent-level measures of importance in a network. Many centrality measures exist, each capturing a different aspect of network position. However, they are all functions of $A$ and can be computed exactly when $A$ is observed. We focus on three popular measures: degree, diffusion and eigenvector centralities. While they are most intuitive when $A$ is binary, centrality measures should be understood as functions of general weighted (symmetric) adjacency matrices. Our definitions are standard up to scaling (see e.g. jackson2010social; bloch2021centrality).
Suppose for now that $a_n^{(1)} = 1$. Then agent $i$'s degree centrality is the sum of row $i$ in $A$. If $A$ is binary, degree centrality is simply the number of agents with whom $i$ has a relationship. Under the Bickel-Chen model, it is possible that $\text{Var}[(A\iota_n)_i] \to \infty$. The scaling parameter, $a_n^{(1)}$, that we take to be $a_*^{(1)}/np_n$, is introduced to stabilize the limit of $\text{Var}[C^{(1)}_i]$. This ensures that the variance of $\varepsilon_i$ relative to $Y_i$ remains non-trivial as $n \to \infty$. Although $p_n$ is unknown, we can achieve the desired rate if we divide by the mean expected degree $\overline{\text{Deg}} := \frac{1}{n} \sum_{i=1}^n \mathbb{E}\left[\left(A\iota_n\right)\right]_i$ or the standard deviation of unscaled degree, $s.d.(A\iota_n)$. In the latter case, we can interpret $\beta^{(1)}$ as the effect of increasing the number of friends by one standard deviation.
Proposed by banerjee2013diffusion, diffusion centrality captures the influence of agent $i$ in terms of how many agents they can reach over $T$ periods. Consider again the case of binary $A$. Then the $(i,j)^\text{th}$ entry of $A^t$ is the number of walks from $i$ to $j$ that are of length $t$, which can be thought of as the influence of $i$ on $j$ over $t$ periods of message passing. Without the scaling factor $a_n^{(T)}$, diffusion centrality for agent $i$ is the sum of their influence on all other agents in the network over time up to period $T$, with a decay of $\delta_n$ per period. bramoulle2018diffusion provides further discussion on the theoretical foundations of diffusion centrality.
We treat $T$ as fixed because, as argued by banerjee2019using, diffusion centrality tends to perform best in practice when $T$ is small. Accordingly, it is “best thought of as finite and not very large.” The decay parameter $\delta_n$, usually taken to be $<1$, captures the idea that an agent’s influence weakens when their messages take longer to reach others. We choose $\delta_n \propto 1/np_n$ because in our framework, this is the unique rate at which all terms remain non-negligible asymptotically. Outside of this rate, either $A\iota_n$ or $A^T\iota_n$ dominates the other terms. Setting $\delta_n \propto 1/np_n$ therefore leads to a richer asymptotic approximation, but makes analysis more challenging. Our chosen rate is consistent with empirical practice. For example, banerjee2019using sets $\delta_n$ to be the inverse of the leading eigenvalue of the adjacency matrix. In our set up, this term would be of order $1/np_n$. Finally, $a_n^{(T)}$ serves the same function as $a_n^{(1)}$ in ensuring that the variance of $\varepsilon_i$ is non-negligible asymptotically. As before, this rate can be achieved, for example, by setting $a_n^{(T)} = 1/\overline{\text{Deg}}$.
Eigenvector centrality is based on the idea that an individual's influence should be proportional to the influence of their friends. That is, for some $k > 0$, we seek the following property:
The eigenvectors of $A$ solve the above equations, with $k$ being the corresponding eigenvalue. By the Perron-Frobenius Theorem, the leading eigenvector is the unique eigenvector that can be chosen so that every entry is non-negative, motivating its use as a centrality measure.
The leading eigenvector is well-defined only if the leading eigenvalue of $A$ has multiplicity 1, that is, if $\lambda_1(A) \neq \lambda_2(A)$. To ensure that this occurs with high probability, we will make the following assumption when analyzing eigenvector centrality:
Eigenvectors are defined only up to scale: if $\mathbf{C}$ satisfies Equation (ref), so will $a^{(\infty)}_n \cdot \mathbf{C}$ for any $a^{(\infty)}_n \in \mathbb{R}$. We set $a_n^{(\infty)} \propto \sqrt{n}$ to stabilize the variance of $C_i^{(\infty)}$ as $n \to \infty$. This is in line with cai2021network. le2020linear does not fix the length of the eigenvector, but their goal is to recover the projection $\mathbf{C}^{(d)}\beta^{(d)}$ and not $\beta^{(d)}$ itself.
This paper focuses on the above three centrality measures, which are closely related (bloch2021centrality). When $T = 1$, $\mathbf{C}^{(1)} \propto \mathbf{C}^{(T)}$. Furthermore, as shown by banerjee2019using, if $\delta_n \geq 1/\lambda_1(A)$, then $\lim_{T \to \infty} \mathbf{C}^{(T)} \propto \mathbf{C}^{(\infty)}$, where the limit is taken with $n$ fixed. We can therefore understand the centrality measures as sums of $A^t$ up to different $T$'s, motivating our notational choice. Notably, we do not discuss Katz-Bonacich centrality which is the limit of diffusion as $T \to \infty$ when $\delta_n < 1/\lambda_1(A)$. Its properties are challenging to analyze in sparse regimes.
When $A$ is observed,we have access to the following estimators.
When networks are observed with errors, we assume that network centralities are estimated using $\hat{A}$ in place of $A$:
We assume that $a_n^{(d)}$ and $\delta_n$ are known even when $\hat{A}$ is unknown. Estimation of these constants will not affect our result if their rates of convergence is faster than $1/\sqrt{n}$. By Theorem 1 in bickel2011method, this is true for a large class of statistics that they call network moments, including mean degree. However, statistics such as standard deviation may require bias correction.
In the case with observation error, we consider plug-in estimators based on noisy centrality measures:
For the two procedures, we also define the corresponding regression residuals:
We conclude this section with two assumptions. The first concerns the moments of $\varepsilon_i$ and $W_i$:
Assumption (a) above implies that $E[ \varepsilon^{(d)}_{i} \, \big\lvert \, W_i, C_i^{(d)} ] = 0$, justifying linear regression. Meanwhile, (c) ensures that $\varepsilon_i^{(d)}$ converges sufficiently quickly and implies the upper bound in (b). The lower bound in (b) prevents degeneracy. Finally (d) requires $W_i$ to have finite second moment, which is standard.
Our second assumption rules out perfect collinearity between $W_i$ and the centrality measures. We state our assumption in terms of the following:
Here, $\phi_1$ is the leading eigenfunction of the operator $T(g) = \int f(x, y) g(y) \; dy$. It is well known that $C^{(d)}_i \mid U_i \overset{p}{\to} C^{*,(d)}(U_i)$ for $d \in \{1, \infty\}$ (see e.g. avella2020centrality). The case of $C^{*,(T)}(U_i)$ follows straightforwardly from the law of large numbers for U-statistics. We can now state our final assumption:
Assumption (ref) rules out $f(u,v) = 0$ since otherwise $C_i^{*, (d)} = 0$. If $W_i$ includes a constant, it rules out the constant graphon $f(u,v) = k$ since all agents would have identical centrality. Similarly, if $W_i$ includes fixed effects based on some cluster membership $c$, then $f$ cannot be a stochastic block model based on $c$ only.
This section studies the consistency of OLS estimators under varying degrees of network sparsity. Section (ref) presents the no observation error benchmark. Section (ref) characterizes the level of sparsity at which consistency of each $\hat{\beta}^{(d)}$ fails. The upshot is that $\hat{\beta}^{(\infty)}$ is able to accommodate less sparsity than $\hat{\beta}^{(1)}$ and $\hat{\beta}^{(T)}$.
In the absence of observation error, $\tilde{\beta}^{(d)}$ is consistent for $\beta^{(d)}$ and converges to a standard normal distribution at rate $\sqrt{n}$. The three centrality measures perform equally well in the sense that consistency holds regardless of network sparsity, since $\tilde{\beta}^{(d)}$ is invariant to $p_n$. As usual, the plug-in estimator for $V^{(d)}$ is also consistent. Specifically, let $\tilde{V}^{(d)}$ be the estimator obtained from $V^{(d)}$ by replacing $C_i^{*,(d)}$ with $C_i^{(d)}$ and $\varepsilon_{i}$ with the regression residual $\tilde{\varepsilon}_i$. Then, replacing $V^{(d)}$ with $\tilde{V}^{(d)}$ does not change the result above.
In the presence of observation error, consistency of the OLS estimator depends on $p_n$.
We draw two lessons from the consistency thresholds in Theorem (ref). First, whereas $p_n$ did not affect consistency in the absence of observation error, it is relevant for the consistency of $\hat{\beta}^{(d)}$. Sparsity is therefore an issue only under observation error. Likewise, observation error is not an issue when $p_n$ is sufficiently dense. As such, it is the interaction of sparsity and observation error that leads to the poor properties of the plug-in estimators $\hat{\beta}^{(d)}$.
Second, $\hat{\beta}^{(\infty)}$ requires greater $p_n$ for consistency than $\hat{\beta}^{(1)}$ and $\hat{\beta}^{(T)}$. In other words, $\hat{\beta}^{(\infty)}$ is less robust to sparsity than $\hat{\beta}^{(1)}$ and $\hat{\beta}^{(T)}$. The threshold for consistency for degree and diffusion is the threshold at which the expected number of reported links diverges to infinity. Since degree and diffusion are mean-type statistics, this is sufficient for them to concentrate to their true values. On the other hand, when $p_n \approx 1/n$, it is well-known that $\sum_{j} \hat{A}_{ij} \approx \text{Poisson}\left(\sum_{j} A_{ij}\right)$, giving rise to classical measurement error, which translates into the usual attenuation bias in the absence of covariates.
For eigenvector centrality, our inconsistency results include a lower bound on $p_n$. Note, however, that $n^{-1} \sqrt{\frac{\log n}{\log \log n}}$ is a sharp consistency threshold: $\hat{\beta}^{(d)}$ is consistent right above that rate and inconsistency at and below it. To understand how the threshold arises, observe that the eigenvalues of A are of order $np_n$. The matrix $\hat{A}$ has informative eigenvectors which are close to those of $A$, and with similar eigenvalues. However, it also contains junk eigenvectors. Specifically, each high degree node gives rise to an eigenvector with eigenvalue approximately equal to the square root of the degree. It is well known that maximum degree $\hat{d}_\text{max}$ has order:
When $\sqrt{\hat{d}_\text{max}} \gg np_n$, the leading value belongs to an uninformative eigenvector with high probability, and we cannot estimate the eigenvectors of $A$ using the eigenvectors of $\hat{A}$. This is the result of benaych2020spectral and alt2021poisson.
This section studies the asymptotic distribution of $\hat{\beta}^{(1)}$, $\hat{\beta}^{(T)}$ and $\hat{\beta}^{(\infty)}$, focusing on the regimes in which they are consistent. Our main result is that observation error leads to asymptotic bias that can be of larger order than the variance when $\beta^{(d)} \neq 0$. To better characterize the bias of $\hat{\beta}^{(\infty)}$, we introduce the following assumption:
In Equation (ref), we express $f$ in terms of its eigenfunctions $\left\{\phi_r\right\}_{r=1}^R$. Assumption (ref) implies that the true network has low-dimensional structure and is satisfied by many popular network models, such as the stochastic block model (holland1983stochastic, see also Example (ref) below) and random dot product graphs (young2007random). This assumption is also commonly found in the network (e.g. levin2019bootstrapping; li2020network) and matrix completion literatures (e.g. negahban2012restricted; chatterjee2015matrix; athey2021matrix). Importantly, existing papers on inference with eigenvectors (le2020linear; cai2021network) also make this assumption. Note that Assumption (ref) implies Assumption (ref).
We can now describe the asymptotic distribution of the estimators:
Under the conditions for consistency, observation error does not affect the asymptotic variance of the OLS estimator. However, it gives rise to an asymptotic bias that is of order $1/np_n$ when $\beta^{(d)} \neq 0$. Equivalently, $\hat{\beta}^{(d)}$ has an “attenuation" factor of $\frac{1}{1-B^{(d)}}$, where:
where
and $\hat{\pi}^{(d)}$ are the OLS estimates from the regression of $\hat{\mathbf{C}}^{(d)}$ on $\mathbf{W}$.
The above display shows that there are two sources of bias. The first is the attenuation bias term, $B_1^{(d)}$, which is approximately $\text{Var}[\hat{C}^{(d)}_i] - \text{Var}[C^{(d)}_i] > 0$. In our setup, the variance inflation vanishes asymptotically, but potentially too slowly for inference. The second source of bias, $B_2^{(d)}$, arises only in the presence of covariates. Under observation error, it is possible that $\text{Cov}[W_{i}, \hat{C}_i] \neq \text{Cov}[W_{i}, C_i]$, and as a result, variation due to $W_i$ can be misattributed to $\hat{C}_i$ and vice versa. This happens when the entry-wise bias of centrality statistics, $\mathbb{E}[\hat{C}^{(d)}_i - C^{(d)}_i \mid U_i]$, is correlated with $C_i^{(d)}$ and hence $W_i$. $B_2^{(d)}$ can be positive or negative depending on the relationship between $W_i$ and $U_i$. Note that this second channel is absent for degree centrality since $\hat{C}_i^{(1)}$ is unbiased for $C_i^{(1)}$.
Asymptotic bias is exacerbated when the networks are sparse. When $p_n \lesssim 1/\sqrt{n}$, bias is of the same or larger order compared to the variance ($\approx 1/\sqrt{n}$). The plug-in $t$-statistic is therefore incorrectly centered when $\beta^{(d)} \neq 0$, rendering the $t$-test and the associated confidence intervals invalid. When the asymptotic bias dominates the variance, $\hat{\beta}^{(d)}$ does not have a non-degenerate asymptotic distribution. That is, there are no scaling factors $\alpha, \beta$ for which $n^{\alpha}p_n^{\beta}(\hat{\beta}^{(d)}-\beta^{(d)})$ converges to a distribution with non-zero variance. Valid inference therefore requires alternative procedures, which we develop in the next section. We note an important exception to the above discussion: the plug-in $t$-test is valid for the null hypothesis $H_0: {\beta}^{(d)} = 0$.
Theorem (ref) provides upper bounds on the rates at which the asymptotic bias goes to $0$. The next result shows that these rates are attained by a homogeneous network with no covariates. The implication is that practitioners should be concerned about asymptotic bias even for relatively dense networks $(p_n \approx 1/\sqrt{n})$.
Degree and diffusion attain the stated bounds across the full range of consistency. For eigenvector centrality, we require slightly more density. It is possible to extend our arguments into slightly sparser regimes but this would not change the conclusion of the present exercise.
This section proposes strategies for addressing the asymptotic bias of $\hat{\beta}^{(d)}$. Section (ref) considers analytical bias correction for degree and diffusion. Section (ref) proposes a conservative form of bias-aware inference for eigenvector centrality.
This section develops analytical bias corrections for the degree and diffusion regressions. We first express the bias terms as sums of terms we call mixed products, defined below. We then show that the asymptotically non-negligible components of mixed products can be expressed as counts of particular trees. This yields plug-in estimators for mixed products that can be combined to construct estimators for $B_1^{(d)}$ and $B_2^{(d)}$.
For degree and diffusion, $B_1^{(d)}$ can be written as:
The final sum is over $\tilde{\mathcal{B}}(t,\tau)$, which is defined to be the multiset obtained by collecting terms of the same order. This set is straightforward to enumerate.
The same reasoning applies to $B_2^{(d)}$:
Here, $\mathcal{B}(t,\tau)$ is the full set of mixed products as in Definition (ref) and $\mathbf{W}_{\cdot, j}$ refers to the $j$-th column of $\mathbf{W}$. We have seen that $B_1^{(d)}$ and $B_2^{(d)}$ are finite sums of terms of the form $w'Bv/n^{t+1}p_n^t$. It is therefore sufficient to find good estimators for such terms.
It turns out that the asymptotically non-negligible components of the mixed product $B$ have a tree structure. Algorithm (ref) describes how the relevant trees can be constructed for a mixed product of order $(t, \tau)$. Figure (ref) provides an example. Intuitively, the algorithm replaces contiguous blocks of $A$ with path graphs and contiguous blocks of $\boldsymbol{\xi}$ with appropriate trees from the set $\mathcal{G}(p_s)$. We keep track of the number of each tree using the function $\gamma: \mathcal{G}(p_s) \mapsto \mathbb{N}$. To describe the construction of $\mathcal{G}(p_s)$ and $\gamma(g)$, we introduce the following definitions.
Note that the edge set of $g^{\text{simp}}(\mathbf{i})$, denoted $E(\mathbf{i})$, is the set of unique edges in $g(\mathbf{i})$. Similarly, $V(\mathbf{i})$ is the set of unique nodes in $g(\mathbf{i})$. Recall that a tree is a connected, acyclic graph. Now, let $\mathcal{I}$ be the set of walks of length $s$ on the complete graph on $[s+1]$, with terminal node $i_{s+1}$ having label $m$. From $\mathcal{I}$, delete every $\mathbf{i}$ for which $g(\mathbf{i})$ has at least one edge with multiplicity exactly equal to 1. Next, delete every $\mathbf{i}$ for which $g^{\text{simp}}(\mathbf{i})$ is not a tree (i.e. contains cycles), denote this final set $\tilde{\mathcal{I}}$. Let $\mathcal{G}(s,m)$ be the set of homorphism classes of $\{g^{\text{simp}}(\mathbf{i}) : \mathbf{i} \in \tilde{\mathcal{I}}\}$. Node labels are identical for all graphs in the same homomorphism class. We define the representative element of each class to be the graph for which node identity is equal to the node label. Let $\gamma(g)$ record the number of homorphisms for each $g \in \mathcal{G}(s,m)$. This set is straightforward to enumerate for $s$ of moderate sizes. Appendix (ref) presents $\mathcal{G}(s,m)$ for $s \leq 10$. Finally, we define $\mathcal{G}(p_s) = \bigcup_{m} \mathcal{G}(p_s,m)$.
Using Algorithm (ref), we obtain $\mathcal{T}(B)$, defined to be the set of pairs $(g, k)$. With some abuse of notation, we will also treat $\mathcal{T}(B)$ as the set of trees $\{g: (g,k) \in \mathcal{T}(B)\}$ and $k_B(g) = k \mbox{ for } (g,k) \in \mathcal{T}(B)$ as the function that keeps track of the multiplicity of $g$. Let $E(g)$ denote the edge set of $g$. Recall that $n(g)$ is the number of nodes in $g$. Then, the asymptotically non-negligible components of $w'Bv$ are
Here, $m = m(g)$ is the label on the terminal node of $g$. Note that the index of summation should not repeat i.e. $i_j \neq i_k$ for all $j \neq k$. See the caption of Figure (ref) for an example.
This motivates the plug-in estimator, $\check{Q}(B,w,v)$, which is obtained from $Q(B,w,v)$ by replacing $A_{i_j,i_k}$ with $\hat{A}_{i_j, i_k}$. When $B$ is a mixed product for which $t$ is large, computation of $\check{Q}(B,w,v)$ can be expensive. Appendix (ref) discusses the computation of $\check{Q}(B,w,v)$ on a random subset of indices. Finally, using $\check{Q}(B,w,v)$, we can construct the following estimators:
In turn, we define:
and our debiased-OLS estimator to be $\check{\beta}^{(d)} := \frac{\hat{\beta}^{(d)}}{1-\check{B}^{(d)}}$. Then,
Our debiased estimators for degree and diffusion converge to $\beta^{(d)}$ at the same rate as $\tilde{\beta}^{(d)}$. $\check{\beta}^{(d)}$ and $\tilde{\beta}^{(d)}$ have the same asymptotic variance since $\check{B}^{(d)} = O_p(1/np_n)$. Using $\check{\beta}^{(d)}$ in place of $\hat{\beta}^{(d)}$ therefore makes the plug-in $t$-test and confidence intervals valid. Although the adjustment in the denominator is not technically necessary, our simulations suggest that it can lead to substantial differences in size control and coverage probability in finite sample, particularly when the network is sparse. As such, we recommend rescaling the variance estimator by the same factor.
Theorem (ref) holds for any sequence $p_n \gg 1/n$. This requires $\check{B}^{(d)}_1$ and $\check{B}^{(d)}_2$ to have errors that are of strictly smaller order than $1/\sqrt{n}$, even when $p_n$ is close to $n^{-1}$. To achieve this, we need a precise characterization of the asymptotic bias as well as sufficiently accurate estimators, necessitating the tree characterization. In particular, a first order bias correction will not lead to valid inference, since that reduces the asymptotic bias only to $1/n^2p_n^2$, which would dominate $1/\sqrt{n}$ if $p_n \ll n^{-3/4}$. On the other hand, if we are willing to assume that $(np_n)^{-k} \ll 1/\sqrt{n}$ for some $k$, then we would only need to perform bias correction for lower order terms.
The above approach does not apply to eigenvector centrality. Although the asymptotic bias of $\hat{\beta}^{(\infty)}$ also involves mixed products, they are pre-multiplied by terms of the form:
We show that this term is $O_p(1/np_n)$ when $f$ is has finite rank, which allows us to form improved bounds in the next section (see discussion after Corollary (ref)). However, our characterization is not precise enough for analytical bias correction.
This section considers inference based on asymptotically conservative bounds for the biases $B_1^{(d)}$ and $B_2^{(d)}$. Let the (unnormalized) average degree and its plug-in estimator be:
It follows from Theorem 1 in bickel2011method that
Under the conditions in Theorem (ref), we have, for $\eta > 0$, that $\left(\widehat{\text{Deg}} \right)^{-1+\eta} > |B^{(d)}|$ w.p.a. 1 as $n \to \infty$. This motivates the following $(1-\alpha)$ confidence intervals for $\beta^{(d)}$:
where
The preceding discussion yields:
Relative to the plug-in confidence intervals, we essentially replace the point estimate $\hat{\beta}^{(d)}$ with the interval $\left[\underline{\beta}^{(d)}, \overline{\beta}^{(d)}\right]$, chosen to cover the true parameter even in the presence of bias. Our approach is in the spirit of bias-aware inference (see e.g. armstrong2018optimal, armstrong2020simple). These confidence intervals are adaptive to sparsity in that they do not require the users to specify $p_n$. They are conservative when $(np_n)^{1-\eta} \ll \sqrt{n}$, with asymptotic coverage probability of 1. However, they are exact when $(np_n)^{1-\eta} \gg \sqrt{n}$. To bound the bias of eigenvector centrality using the same rate as degree and diffusion, we need the nuisance term in (ref) to be $O_p(1/np_n)$. We show that this is the case under Assumption (ref). Otherwise, the Davis-Kahan inequality only implies a rate that is $O_p(1/\sqrt{np_n})$, leading to much more conservative bounds.
Our bounds involve a tuning parameter $\eta$. Simulations in the next section suggest that it is reasonable to set it to 0.1, and we defer a data-driven choice to future work. Finally, we emphasize that while bias correction may be challenging, it is not necessary for valid tests of the null hypothesis $H_0: \beta^{(d)} = 0$.
In this section, we present simulation evidence to support our theory. We generate the network from a 2-block stochastic block model with connection matrix:
We can think of block 1 as the extroverts and group 2 as the introverts. Let $c(i)$ indicate the block membership of $i$. Then $c(i) = 1$ if $U_i \leq 0.5$ and $c(i) = 2$ otherwise. As such, blocks are of equal sizes and $f(U_i, U_j) = B_{c(i), c(j)}$. Note that because of $p_n$ (described below), the actual connection probabilities between any two groups are always strictly below $1$. Next, let
where $C^{(d)}_i$ are centrality measures calculated on $A$. We will focus on $d \in \{1,2,3,\infty\}$. That is, degree, diffusion with $T \in \{2,3\}$, and eigenvector centrality. We set $\gamma = 1$ and $\beta = 1$. The covariate is $W_i = \sqrt{12} \cdot U_i$, where $U_i \sim \text{Uniform}[0,1]$ is the latent type. It is scaled to have variance $1$. For comparability, we also set $\delta_* = 1$ and choose $a_*^{(d)}$ so that $\text{Var}[C_i^{*,(d)}] =1$. Finally, $\varepsilon^{(d)}_{i} \overset{\text{i.i.d.}}{\sim} \text{N}(0,1)$, where $\varepsilon^{(d)}_i \perp\!\!\!\perp U_i$ and $\varepsilon^{(d)}_i \perp\!\!\!\perp \hat{A}_{jk}$ for all $i,j,k \in [n]$. We vary $n$ and $p_n$ across simulation designs, setting $p_n = k \cdot r(n)$ with $k := 2/\int f du dv$ so that mean degree is 2 in the sparsest regime we consider $(r(n) = n^{-1})$. For brevity, the tables report only the rate component of $p_n$, suppressing the multiplicative constant $k$. Simulations are based on 2,500 draws for each set of parameter values.
Table (ref) presents the root mean-squared error (RMSE) of the OLS estimator when $A$ is known ($\tilde{\beta}$), when $\hat{A}$ is used as a plug-in ($\hat{\beta}$) and of our bias-corrected estimator ($\check{\beta}$). As $n$ increases from 250 to 4000, the RMSE of $\tilde{\beta}$ converges quickly to $0$. This holds for all centrality measures and the rate of convergence does not depend on $p_n$, consistent with Theorem (ref).
In contrast, the RMSE of $\hat{\beta}$ depends on $p_n$. We first observe that there is a sharp degradation in the RMSE when going from the first three regimes of $p_n$, which are moderately sparse, to the last three regimes, which are extremely sparse. For degree and diffusion, RMSE decreases in $n$ for all regimes except $p_n \propto n^{-1}$, where it converges to a constant. For eigenvector, RMSE decreases in $n$ when $p_n \gg n^{-1}\sqrt{\frac{\log n}{\log \log n}}$, the threshold for consistency, but diverges when $p_n \propto n^{-1}$. At the threshold, the RMSE of eigenvector is slowly decreasing. Based on Theorem (ref), we would expect the RMSE to stabilize at a value away from $0$, though this might require larger samples to become apparent. Overall, the behavior of $\hat{\beta}$ is in line with threshold provided in Theorem (ref). Comparing RMSE across centrality measures also underscores the fact that eigenvector is less robust to sparsity than degree and diffusion centralities.
Finally, the column labelled $\check{\beta}$ present the RMSE of our bias corrected estimators for degree and diffusion. We speed up computation by computing the bias estimator on a randomly sampled subset of indices, as described in Appendix (ref). To reduce variance, we average over 20 random partitions. As the table shows, bias correction is effective. Even in the moderately sparse regime of $p_n \propto n^{-2/3}$, $\hat{\beta}$ has RMSE that is 1.5 times that of $\check{\beta}^{(d)}$, even when $n = 250$. This ratio is greater than 3 by the time $n = 4000$ and can exceed 10 in sparser regimes.
Table (ref) presents the coverage probabilities of 95% confidence intervals (CIs) under regimes in which $\tilde{\beta}$, $\hat{\beta}$ and $\check{\beta}$ are consistent. The column headings indicate the estimator on which the CIs are based. BA refers to the bias-aware CIs described in Section (ref).
As expected, CIs based on $\tilde{\beta}$ achieve close to 95% coverage across centrality measures and sparsity regimes, in line with Theorem (ref).
Across centrality measures, CIs based on $\hat{\beta}$ achieve the nominal coverage probability when $p_n \propto n^{-1/3}$. At $p_n \propto n^{-1/2}$, the coverage probability appears to converge to a value below the nominal level as $n$ increases. At $p_n \propto n^{-2/3}$ and below, coverage probability converges to $0$. This behavior is consistent with an asymptotic bias that is $\Theta_p(1/np_n)$, as stated in Corollary (ref).
We address the asymptotic bias using CIs based on $\check{\beta}$ for degree and diffusion. When $p_n$ is $n^{-2/3}$ or above, our proposed CIs have close to nominal coverage. At lower $p_n$'s, coverage probability can be as low as 70%. This is because $np_n$ ranges from 4 to 6 in these regimes and our asymptotic results are based on $np_n \to \infty$. Nonetheless, they perform substantially better than the plug-in CIs based on $\hat{\beta}$. For eigenvector centrality, we propose to construct CIs by bounding $B^{(d)}$ using the mean degree. Their coverage probabilities when $\eta = 0.1$ are presented in the column denoted BA. These CIs are conservative across the board, but the conservativeness is slowly vanishing in $n$ for $p_n \propto n^{-1/3}$.
Figure (ref) presents power curves for two-sided tests of the hypothesis $H_0: \beta = 1$ when $n=2000$. The first panel presents power of the $t$-tests based on $\tilde{\beta}^{(d)}$. Its power is close to 1 when $\beta = 1.05$. By construction, the power of these tests do not depend on $p_n$. The remaining panels presents the power of our procedures, which are $t$-tests based on $\check{\beta}$ for degree and diffusion, and tests based on the bias-aware confidence intervals for eigenvector. The second panel focuses on the moderately sparse case when $p_n = n^{-1/2}$. We see that tests based on $\check{\beta}$ are competitive with those based on $\tilde{\beta}$. The bias-aware procedure is conservative, but has power 1 once $\beta = 1.1$. In the third panel, $p_n = n^{-2/3}$. Tests based on $\check{\beta}$ continue to do well, even though there is some size distortion when $\beta = 1$. The bias-aware procedure becomes more conservative, since the bounds have to accommodate a bias term that is potentially much bigger. However, its behavior is reasonable, achieving a power of 1 around $\beta = 1.15$. Finally, note that we do not plot the performance of tests based on $\hat{\beta}$ since as Table (ref) shows, their Type I error can be as high as 20% when $p_n = n^{-1/2}$ and 90% when $p_n = n^{-2/3}$.
Overall, our simulations confirm the behavior of $\tilde{\beta}^{(d)}$ and $\hat{\beta}^{(d)}$ as described in Theorem (ref) and (ref). Our proposed estimators and confidence intervals perform substantially better than those based on $\hat{\beta}^{(d)}$ and can even be competitive with those based on $\tilde{\beta}^{(d)}$.
In this section, we demonstrate the utility of our theoretical results via an application inspired by deweerdt2006risk.\footnote{The data is obtained from Joachim De Weerdt's website: \url{https://www.uantwerpen.be/en/staff/joachim-deweerdt/public-data-sets/nyakatoke-network/}.} When access to formal credit markets is limited, social insurance is an important mechanism for smoothing consumption. deweerdt2006risk studies Nyakatoke, a village with 119 households in rural Tanzania, and finds that social insurance helps households to smooth consumption following health shocks. The data they use comprises five rounds of panel data on household consumption, illness, and other covariates, collected from February to December 2000. The authors also had access to social network data collected during the first round of the survey, in which households were asked to identify those whom they depend on or who depend on them for help. The authors then regress a household's change in consumption following illness on the mean consumption of its network neighbors, finding evidence of positive co-movements.
Another way to study the role of social insurance in consumption smoothing is to regress the standard deviation of consumption expenditure on network centrality measures. Specifically, consider the regression:
where $Y_i$ is the standard deviation in log food expenditure over the five survey rounds. $C^{(d)}_i$ denotes degree ($d=1$), diffusion with 2 periods ($d=2$), diffusion with 3 periods ($d=3$), or eigenvector centrality ($d=\infty$). The vector of covariates $W_i$ includes a constant and the number of major illnesses the household experienced. Because wealthier households may have smoother consumption and higher centralities, we also include mean log food expenditure as a control variable.
Centrality regression could be preferable to the authors' specification. This might be the case if we do not know which covariates best capture the channels through which social insurance operates. For example, it might be a household's stock of savings that co-moves with the decision to lend to their friends, rather than their own consumption. However, in either case, we might expect households that are more central in an appropriate sense to benefit from more risk-sharing. Centrality regressions may also capture more complex patterns of assistance. For example, there might be a second-order form of social insurance in which friends receive support after lending to a household with illness. Such activity could be captured by an appropriate centrality measure, but might be less tractable to model explicitly.
The above regression requires information on the network of social insurance. We might think of this as the matrix $A$, where each entry $A_{ij}$ records the propensity or probability that $i$ lends money to $j$, or vice versa, over the survey period. Following the authors, we consider the use of the following proxy networks:
The networks are plotted in Figure (ref) and the degree distributions are described in Table (ref). UF is the densest network, with a mean degree of 16, followed by US with a mean degree of 8. By construction, BS is extremely sparse, with a mean degree of 2.
Regression results are presented in Table (ref). In this exercise, we set $a^{(d)}_n$ and $\delta_n$ equal to the inverse of the mean degree. This scaling achieves the stated rates in Section (ref) without introducing further asymptotic bias. The column labelled $\hat{\beta}$ presents the OLS estimators based on $\hat{A}$, while $p$-values are for two-sided tests of the hypotheses $H_0: \beta^{(d)} = 0$. Across the board, ${\beta}^{(d)}$ is estimated to be weakly negative, suggesting that more central households exhibit lower consumption volatility, in line with improved access to informal insurance. In terms of statistical significance, degree and diffusion centralities yield estimates qualitatively similar to those based on eigenvector centrality in the denser networks US and UF. For BS, however, we see that degree and diffusion are substantially more predictive than eigenvector centrality. Although this could be taken to mean that the latter is not economically meaningful, our results suggest that the observed difference could simply be due to the eigenvector centrality's relatively worse statistical properties.
Table (ref) also presents our bias corrected estimators for degree and diffusion, as well as the estimated asymptotic bias, $\check{\beta} - \hat{\beta}$. For UF, the bias is between -0.003 to -0.029, so that bias correction makes a substantial difference to the estimates, particularly for diffusion centrality. US is sparser and consequently, the bias is of larger absolute value, ranging from 0.003 to -0.103. For BS, bias ranges between -0.02 to 0.042. However, with a mean degree of 2.3, BS is so sparse that we do not expect our estimators to be reliable. These estimates should therefore not be overinterpreted. Finally, the last column presents bias-corrected CIs for degree and diffusion, and bias-aware type CIs for eigenvector. As expected, these CIs are generally wider than the CIs based on $\hat{\beta}$.
In sum, centrality regressions on the Nyakatoke networks exhibit behavior in line with the predictions of our theory. Apart from the sparsest network, our methods appear to work well in this setting.
This paper studies the properties of linear regression on degree, diffusion and eigenvector centrality when networks are sparse and observed with error. We show that these issues threaten the consistency of OLS estimators and characterize the amount of sparsity at which inconsistency occurs. In doing so, we find that eigenvector centrality is less robust to sparsity than the others and that the statistical properties of the corresponding regression are sensitive to the scaling.
Additionally, we show that an asymptotic bias arises whenever the true slope parameter is not zero and that the bias can be of larger order than the variance, invalidating plug-in $t$-tests and confidence intervals even when the estimators are consistent. For degree and diffusion, we obtain a precise characterization of the asymptotic bias by relating them to counts of particular trees, and we provide algorithms for enumerating such trees. This allows us to construct bias corrected estimator with substantially better performance compared to the plug-in estimator. For eigenvector, we propose to widen the confidence intervals using a bound of the asymptotic bias. Although the procedure is conservative when networks are sparse, they become exact once the network is moderately dense.
Our results suggest that applied researchers should view their estimates and confidence intervals with caution when applying OLS to sparse, noisy networks. Specifically, comparing the statistical significance of eigenvector centrality with degree or diffusion may yield misleading conclusions since they differ not only in economic significance but also in their statistical properties. Provided that the networks are not too sparse, the usual $t$-test is valid for the null hypothesis that the slope parameter is $0$. However, alternative inference procedures will be necessary for non-zero null hypotheses and for constructing valid confidence intervals. Additionally, there may be scope for improving estimation through the use of bias-corrected estimators. Estimation and inference under bounded degrees remain an open question, though as le2017concentration and graham2020sparse show, parametric models may point to a way forward.