Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
74,557 characters · 16 sections · 18 citation commands
Network-Adjusted GMM Estimation under Network Uncertainty
Social interactions play a crucial role in human behavior. Knowledge diffusion among workers, peer effects among schoolmates, and the transmission of infectious diseases across communities are only a few examples in which accounting for social interactions is essential. However, in most empirical applications, researchers do not have precise information about the underlying interaction structure. For example, friendship networks among students are often collected directly through surveys, but individuals may differ in how they report friendships, and some links may be missing. Moreover, even when such a survey-based network is available, it is not clear whether it is the relevant network for the social interaction effect under study. In spatial interaction models, researchers often construct a network from administrative adjacency or geographic distance. However, different choices of distance measures, cutoffs, or normalizations can lead to different networks, and it is often difficult to justify that a particular choice is appropriate. More generally, even if the true network were observed at a particular time point, there may exist unobserved heterogeneity in how units perceive the strength of each link across different situations.
Since incomplete network data and network errors are common and often unavoidable in practice, a growing literature has developed statistical inference methods that are robust to such uncertainty. One approach is to estimate a specific network formation model using partially observed network data and simulate relevant moment conditions to estimate the parameters of interest (e.g., chandrasekhar2011econometrics, reeves2024model, boucher2025estimating). Another approach, which is available even when no network information is observed, is to directly estimate the interaction matrix nonparametrically using large panel data (e.g., de2025identifying). lewbel2023social also do not require prior knowledge of the network structure, but their methodology relies on the availability of many small networks. By contrast, lewbel2024ignoring show that conventional two-stage least squares (2SLS) estimators can remain valid even when network errors are ignored, provided that the observed network contains only a sufficiently small amount of measurement error. Although we do not discuss it here, there is also a growing literature in causal inference that studies treatment spillovers under incomplete network data.
As demonstrated by these studies, estimating social interaction models under network uncertainty typically requires substantial additional sources of identification and restrictions, such as specific network formation models, large panel data, many small networks, or sufficiently small network errors. However, in many empirical situations, such additional information is not readily available. Thus, the key question of this paper is: what can researchers do when such information is unavailable but network uncertainty is still a concern, other than naively applying conventional estimators that ignore network errors? To address this question, we propose a novel penalized generalized method of moments (GMM) approach that allows the observed network to be modified through optimization.
If the model specification is correct and the moment conditions are valid under the true interaction network, network errors may result in violations of the moment conditions. For example, in standard linear spatial (or network) autoregressive (SAR) models, researchers often employ a 2SLS estimator with spatially lagged covariates as instrumental variables (IVs) for the endogenous interaction effect term. In this case, the validity of the moment conditions can be checked empirically, for example, by using classical Sargan's over-identification test. Then, if we reject the over-identification test, we may suspect the presence of network errors, provided that all other components of the model are correctly specified. This naturally leads to a new class of moment-based estimators, which we call network-adjusted GMM (NA-GMM), that adjust the network itself simultaneously with the estimation of model parameters. Since each network is typically a high dimensional object, allowing it to be modified freely leads to an identification problem due to too many degrees of freedom. Thus, the NA-GMM method introduces a penalty term for the size of the network correction.
Conceptually, the proposed NA-GMM method is closely related to the generalized empirical likelihood (GEL) method and the optimally-transported GMM (OT-GMM) developed by schennach2026optimally. Similar to this paper, these studies consider situations in which the model may fail to pass over-identification tests. As for the source of this moment mismatch, GEL considers biased sampling from the target population, and OT-GMM considers general measurement errors in the data. These approaches aim to improve the moment fit by allowing such distortions to be corrected in the estimation process. NA-GMM builds on the same motivation as theirs. However, we attribute the moment mismatch to an imprecisely observed network.
Another closely related literature is GMM estimation with misspecified moment conditions (e.g., hall2003large, hansen2021inference, kleibergen2025double, kleibergen2025testing). Although the NA-GMM procedure can update the observed network to improve the moment fit, the resulting modified network is generally different from the true network. Consequently, NA-GMM estimates of the model parameters are also inconsistent for the true parameter, and they generally converge to corresponding pseudo-true values. The aforementioned studies on misspecified GMM investigate the convergence of estimators to pseudo-true values and asymptotic normality around those values. We establish similar results for NA-GMM.
This paper proceeds as follows. Section (ref) presents the basic framework of NA-GMM. We first define network errors as differences between the true interaction matrix and the matrix used by the researcher, and focus on models in which these errors enter the moment conditions linearly. This class of models includes standard linear SAR models and linear treatment spillover models. We then introduce an uncertainty set, a set of entries of the interaction matrix that are suspected to be affected by network errors, and define the NA-GMM estimator as the solution to a joint optimization problem over the model parameters and corrections to the entries in the uncertainty set. Because unlimited network modifications can create an identification problem when the uncertainty set is large, the NA-GMM criterion includes a quadratic penalty for the size of the network modification. Thanks to the linearity of network errors and the quadratic cost function, we can show that the optimal network correction has a closed-form solution, and the NA-GMM criterion reduces to a simple continuous-updating (CU) GMM-type criterion in which the model parameters enter the GMM weight matrix.
In Section (ref), we focus on the bias reduction property of NA-GMM in linear SAR models. We show that the NA-GMM weight matrix downweights particular directions of the moment conditions that are more severely affected by network errors. In addition, as a simplified alternative to the original NA-GMM estimator, we consider a fixed-weight version of NA-GMM. For this fixed-weight estimator, we prove that the estimator has a smaller worst-case bias than the naive GMM estimator without network adjustment.
We next turn to the asymptotic properties of NA-GMM in Section (ref). We first establish consistency of the NA-GMM estimator for the pseudo-true parameter under high-level conditions and derive asymptotic normality under local misspecification, where the moment misspecification caused by network errors is only local. Meanwhile, as pointed out by hall2003large, under general moment misspecification, the convergence of the GMM weight matrix can affect the limiting distribution of the GMM estimator, and this significantly complicates the asymptotic analysis. Therefore, to investigate the distribution of the NA-GMM estimator under non-local misspecification, we narrow our attention only to linear SAR models and prove its asymptotic normality under more primitive conditions. This result also shows that performing statistical inference based directly on the asymptotic distribution is difficult in practice, because the asymptotic variance depends on the unknown true parameter and true interaction matrix. We therefore suggest using NA-GMM as a diagnostic tool for assessing the robustness of GMM estimates to network uncertainty and provide several diagnostic analyses.
Sections (ref) and (ref) provide numerical illustrations of NA-GMM. Section (ref) reports Monte Carlo simulations. The simulation results show that, across various experimental setups, NA-GMM generally achieves a smaller average bias than the naive GMM estimator that ignores network errors, confirming its desirable bias reduction property. The fixed-weight NA-GMM estimator also has a slightly smaller bias than the naive GMM estimator, as predicted by the theory, although it does not perform as good as the original NA-GMM estimator. Importantly, when there is no network error, all GMM estimators are accurate. This suggests that NA-GMM can be used as a robust alternative to conventional GMM. The empirical application in Section (ref) analyzes U.S. county-level COVID-19 infection data. The naive GMM estimate indicates a large positive spatial spillover effect of COVID-19 infection. Applying NA-GMM to the same model substantially improves the moment fit, while the estimate of the spatial effect changes only slightly. This finding suggests that the estimated spatial effect is relatively robust to the specification of the interaction network.
Section (ref) concludes.
For an integer $a$, we write $[a] \coloneqq \{1,2,\ldots,a\}$. For a vector $x$, $\|x\|$ denotes the Euclidean norm. For a symmetric matrix $A$, $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ denote its minimum and maximum eigenvalues, respectively. For a $p \times q$ matrix $M = (m_{ij})$, define $\|M\| \coloneqq \left(\sum_{i=1}^p \sum_{j=1}^q m_{ij}^2\right)^{1/2}$, $\|M\|_{\mathrm{op}} \coloneqq \left\{\lambda_{\max}(M^\top M)\right\}^{1/2}$, $\|M\|_\infty \coloneqq \max_{1 \le i \le p} \sum_{j=1}^q |m_{ij}|$, and $\|M\|_1 \coloneqq \max_{1 \le j \le q} \sum_{i=1}^p |m_{ij}|$. Finally, $c$, possibly with subscripts, denotes a generic positive constant whose value may vary from line to line.
Suppose that we observe data of size $n$: $\{(W_i, A_{i1}^{\mathrm{obs}}, \ldots, A_{in}^{\mathrm{obs}}) : i \in [n]\}$. Here, $W_i$ is a vector of observed variables, including response variables and covariates, and $A_{ij}^{\mathrm{obs}}$ is the $(i,j)$-th element of the adjacency matrix $\bm A_n^{\mathrm{obs}}$ that represents the observed social network. The true social network is denoted by $\bm A_n = (A_{ij})_{i,j \in [n]}$, and it may or may not be fully observed.
Typically, network effects are determined not directly through $\bm A_n$, but through its transformed version $\bm G_n = (G_{ij})_{i,j \in [n]}$; for example, $G_{ij} = \left(\sum_{k \neq i} A_{ik}\right)^{-1} A_{ij}$. The observed counterpart of $\bm G_n$ is denoted by $\bm G_n^{\mathrm{obs}} = (G_{ij}^{\mathrm{obs}})_{i,j \in [n]}$. The transformation from $\bm A_n$ to $\bm G_n$ may be unknown and heterogeneous across individuals, whereas that from $\bm A_n^{\mathrm{obs}}$ to $\bm G_n^{\mathrm{obs}}$ is predetermined by the researcher. For both $\bm G_n$ and $\bm G_n^{\mathrm{obs}}$, we assume that the diagonal elements are zero. In what follows, both of these matrices are treated as non-random.\footnote{ This treatment does not necessarily rule out stochastic network formation or random network errors. Rather, all subsequent arguments should be interpreted as being conditional on the realized true and observed interaction matrices. }
Under this setup, the target parameter of interest is $\theta_0 \in \mathbb{R}^{d_\theta}$, which satisfies
$\bm W_n = (W_1, \ldots, W_n)^\top$, and $\mu_i(\bm W_n, \bm G_n; \theta)$ is an $\mathbb{R}^{d_\mu}$-valued moment function. In what follows, we focus on over-identified cases, where $d_\mu > d_\theta$.
If $\bm G_n$ is observable, we can construct a GMM estimator based on (ref) to estimate $\theta_0$. However, in our setting, $\bm G_n$ is not available; instead, only $\bm G_n^{\mathrm{obs}}$ is available, so that, in general,
We define the network error matrix as
whose $(i,j)$-th element is given by $D_{ij} = G_{ij} - G_{ij}^{\mathrm{obs}}$ for $i \neq j$ and $D_{ii} = 0$ for all $i \in [n]$. For the sources of network errors, we can consider two types: measurement error in the observed network $\bm A_n^{\mathrm{obs}}$ and misspecification of the transformation used to construct $\bm G_n^{\mathrm{obs}}$. We cover both types of errors and refer to them simply as network errors.
We often have prior knowledge that $G_{ij}^{\mathrm{obs}} = G_{ij}$, or equivalently $D_{ij} = 0$, holds for certain pairs $(i,j)$. For example, if students in different grades do not interact, or if residents in distant neighborhoods have no opportunity to communicate, then it may be reasonable to set $G_{ij}^\text{obs} = G_{ij} = 0$ when $i$ and $j$ belong to different communities. We define the set of such pairs as $\mathcal{F}_n \coloneqq \{(i,j) \in [n]^2: \: D_{ij} = 0\}$. Since $\mathrm{diag}(\bm D_n) = 0$ is always satisfied, $(i,i) \in \mathcal{F}_n$ for all $i \in [n]$. In addition, we define the uncertainty set $\mathcal{U}_n \coloneqq \{(i,j): \: i,j \in [n]\} \setminus \mathcal{F}_n$ as the set of pairs $(i,j)$ for which the researcher is uncertain about the value of $G_{ij}$. Hereinafter, when there is no confusion, we drop the subscript $n$ from $\mathcal{F}_n$ and $\mathcal{U}_n$.
Let $\mathcal{U}(i) \coloneqq \{ j \in [n] : \: (i,j) \in \mathcal{U} \}$ be the $i$-th uncertainty set and let $m_i \coloneqq |\mathcal{U}(i)|$ be its size. To facilitate the analysis, we assume throughout that the moment function is linear in the network error matrix in the following sense.
Two specific examples that fit into our framework are given below.
Under Assumption (ref), we can write $\overline\mu_n(\bm G_n; \theta) = \overline q_n(\theta) + \overline u_n(\vec{\bm D}_\mathcal{U}; \theta)$, where
and $\vec{\bm D}_\mathcal{U} \coloneqq (\bm D_{\mathcal{U}(1)}^\top, \ldots, \bm D_{\mathcal{U}(n)}^\top)^\top$. Here, $m \coloneqq \sum_{i \in [n]} m_i$, $\vec{\bm d}_m \coloneqq (\bm d_{m,1}^\top, \ldots, \bm d_{m,n}^\top)^\top \in \mathbb{R}^m$, and $\bm d_{m,i} = (d_{ij})_{j \in \mathcal{U}(i)}$ is a generic $m_i \times 1$ vector. Throughout, we assume that $m > 0$.
We now introduce our network-adjusted GMM (NA-GMM) estimator. Define the NA-GMM criterion function as follows:
where $\|x \|_{\Omega_n}^2 \coloneqq x^\top \Omega_n x$ and $\Omega_n \in \mathbb{R}^{d_\mu \times d_\mu}$ is a non-stochastic, symmetric, positive definite GMM weight matrix. Then, the NA-GMM estimator is defined by
where $\Theta \subset \mathbb{R}^{d_\theta}$ is a compact parameter set. The criterion function is designed to account for moment mismatch caused by network errors. The sample moments based on the observed network, $\overline q_n(\theta)$, may not be close to zero because $\bm G_n^{\mathrm{obs}}$ is mismeasured, and $\overline u_n(\vec{\bm d}_m; \theta)$ represents an adjustment term to improve the fit of moment conditions.
Introducing a penalty term $(\rho/m) \| \vec{\bm d}_m \|^2$ is crucial. If the network adjustments were unrestricted (i.e., $\rho = 0$), the moment conditions could be made nearly satisfied for arbitrary $\theta$, which leads to a weak identification issue for the NA-GMM estimator. In our criterion, the penalty term measures the quadratic cost of adjusting the network. Hence, $Q_{n,\rho}(\theta)$ becomes small only when the observed-network moment conditions can be corrected by a relatively small amount of network adjustments.
Note that the adjusted network obtained through the inner optimization should not be interpreted as an estimate of the true interaction network $\bm G_n$. It is chosen only to improve the moment fit under the penalty. Therefore, under general network errors, the NA-GMM estimator should be interpreted as estimating a pseudo-true parameter. That said, in many empirically relevant situations, the NA-GMM method is expected to produce an estimate with smaller bias than the naive GMM estimator without network adjustment:
For more specific discussions on this topic in the context of SAR models, see Section (ref).
Recall that we have assumed that $d_\mu > d_\theta$. Under just-identification in which the sample moment equation $\overline q_n(\theta) = \bm 0$ has a solution, the inner minimization can be achieved by $\vec{\bm d}_m = 0$. In this case, the estimator reduces to the just-identified naive GMM estimator based on $\bm G_n^\text{obs}$, and the network adjustment plays no role. The over-identification assumption rules out this trivial situation.
To compute the NA-GMM estimator, we need to solve a profiled minimization problem in which the inner minimization is with respect to a potentially high-dimensional vector $\vec{\bm d}_m$. Fortunately, this inner minimization problem has a closed-form solution. To see this, for each $i \in [n]$, define
and let $\bm H_m(\theta) = \left( \bm H_{m,1}(\theta), \ldots, \bm H_{m,n}(\theta) \right) \in \mathbb{R}^{d_\mu \times m}$. Then, we have
Hence, the NA-GMM criterion function can be written as $Q_{n,\rho}(\theta) = \inf_{\vec{\bm d}_m} \{\| \overline q_n(\theta) + \bm H_m(\theta)\vec{\bm d}_m \|_{\Omega_n}^2 + (\rho/m) \| \vec{\bm d}_m \|^2 \}$. This shows that the network adjustment term $\vec{\bm d}_m$ enters the criterion function quadratically. Due to this quadratic structure, the minimization problem with respect to $\vec{\bm d}_m$ can be solved exactly. We formally state this result in the following lemma.
The proof is by direct calculation; see Appendix (ref).
Using the Woodbury matrix identity, we obtain
which is positive semidefinite. Hence, compared with the naive GMM in (ref), the NA-GMM criterion modifies the weight matrix by downweighting components associated with $\bm H_m(\theta)$. In view of (ref), the amount of downweighting is governed by the reciprocal of $\rho$ and the eigenvalues of $\bm H_m(\theta)\bm H_m(\theta)^\top$. That is, directions associated with larger eigenvalues of this matrix are downweighted more heavily, and a smaller value of $\rho$ strengthens this downweighting. Note that, under our setup,
where the approximation becomes equality in expectation. Thus, when the interaction matrix is misspecified, the resulting violation of the moment conditions approximately lies in the column space of $\bm H_m(\theta_0)$. The NA-GMM downweights such violations, and therefore may be less sensitive than the naive GMM estimator to network uncertainty.
As $\rho$ becomes smaller, the downweighting becomes stronger. At the same time, one should not choose $\rho$ to be too small. In an extreme case, if one sets $\rho = 0$, then the criterion function is reduced to $Q_{n,0}(\theta) = \inf_{\vec{\bm d}_m} \| \overline q_n(\theta) + \bm H_m(\theta)\vec{\bm d}_m \|_{\Omega_n}^2$. Because $\vec{\bm d}_m$ typically has a large number of degrees of freedom, this may lead to a weak identification issue for $\theta$. Thus, there is a trade-off in the choice of $\rho$: a smaller $\rho$ makes the estimator less sensitive to violations of the moment conditions, whereas a larger $\rho$ leads to stronger identification.
In this section, with a particular focus on linear SAR models, we investigate the bias reduction property of the NA-GMM estimators. Specifically, we consider the same model as in Example (ref):
with $Z_i = (\sum_{j \neq i} G_{ij}^{\mathrm{obs}} X_j^\top, X_i^\top)^\top$. As we have seen above, the moment function can be written as
where $\theta = (\alpha, \beta^\top)^\top$, $q_i(\bm W_n, \bm G_n^\text{obs}; \theta) = Z_i ( Y_i - \alpha \sum_{j \neq i} G_{ij}^\text{obs} Y_j - X_i^\top \beta )$, $\bm V_i(\bm W_n, \mathcal{U}; \theta) = -\alpha \bm Y_{\mathcal{U}(i)}$, and $\bm Y_{\mathcal{U}(i)} = (Y_j)_{j \in \mathcal{U}(i)}$. Hence, $\bm H_{m,i}(\theta) = n^{-1} Z_i \bm V_i(\bm W_n,\mathcal{U};\theta)^\top = -(\alpha/n) Z_i \bm Y_{\mathcal{U}(i)}^\top$. Moreover, let
so that we can write $\bm H_m(\theta) = - \alpha \Gamma_n$ and $\overline u_n(\vec{\bm d}_m; \theta) = - \alpha \Gamma_n \vec{\bm d}_m$. Then, the NA-GMM criterion function can be written as
where
and $\Phi_n \coloneqq \Gamma_n \Gamma_n^\top$.
To simplify the discussion, let us assume $\Omega_n = I_{d_\mu}$ hereinafter in this section. As remarked in the previous section, this criterion function is a downweighted GMM criterion in certain directions, compared with the naive GMM. To see this more clearly, let the eigenvalue decomposition of $\Phi_n$ be
where $\Lambda_n = \mathrm{diag}(\lambda_{n,1},\ldots,\lambda_{n,d_\mu})$, and $\lambda_{n,k} \ge 0$ for all $k \in [d_\mu]$. Then
and hence
Thus, $\Psi_{n,\rho}(\alpha)$ and $\Phi_n$ share the same eigenvectors, and each eigen-direction is shrunk by the factor $\rho/(\rho + m\alpha^2\lambda_{n,k})$. Directions associated with larger eigenvalues of $\Phi_n$ are therefore downweighted more strongly.
Recall that, in the SAR model, moment misspecification arises through the term $-\bm H_m(\theta_0) \vec{\bm D}_{\mathcal{U}} = \alpha_0 \Gamma_n \vec{\bm D}_{\mathcal{U}}$. Thus, directions associated with larger eigenvalues of $\Phi_n = \Gamma_n\Gamma_n^\top$ are precisely those in which network errors can generate larger moment violations. The NA-GMM weight matrix $\Psi_{n,\rho}(\alpha)$ downweights these directions, whereas the naive GMM with $\Omega_n=I_{d_\mu}$ assigns equal weight to all directions.
It is difficult to analyze the bias reduction property of the original CU-type estimator in depth because it does not have a closed-form expression. Therefore, here we instead investigate the fixed-weight version of the NA-GMM estimator.
Write
where $\bm Y_n = (Y_1, \ldots, Y_n)^\top$, $\bm X_n = (X_1, \ldots, X_n)^\top$, $\bm Z_n = (Z_1, \ldots, Z_n)^\top$, and $\bm R_n^\text{obs} = (\bm G_n^\text{obs} \bm Y_n,\bm X_n)$. Let $\bm R_n = (\bm G_n \bm Y_n, \bm X_n)$ and $\bm \varepsilon_n = (\varepsilon_1, \ldots, \varepsilon_n)^\top$ so that $\bm Y_n = \bm R_n \theta_0 + \bm \varepsilon_n$. Then,
where $\Pi_n \coloneqq n^{-1} \bm Z_n^\top \bm R_n^\text{obs}$, and $\bm b_n \coloneqq n^{-1} \bm Z_n^\top (\alpha_0 \bm D_n \bm Y_n + \bm \varepsilon_n) = \alpha_0 \Gamma_n \vec{\bm D}_\mathcal{U} + n^{-1} \bm Z_n^\top \bm \varepsilon_n$.
Define, for some $\kappa > 0$,
and
assuming that $\Pi_n^\top \Psi_{n, \kappa} \Pi_n$ is non-singular. As $\kappa \to \infty$, the weight matrix $\Psi_{n, \kappa}$ collapses to an identity matrix, leading to the naive GMM estimator; that is, $\widehat \theta_{n}^{\text{naive}} = \widehat \theta_{n,\infty}^{\text{fw}}$. Write
where
From (ref), we can clearly see that the bias component of the fixed-weight NA-GMM comes from the network errors transformed through the matrix $L_n(\kappa)$. Since the second term of (ref) is typically of order $O_P(n^{-1/2})$ for any $\kappa > 0$, the relative performance of the fixed-weight NA-GMM compared with the naive GMM estimator is essentially governed by the first term. The next proposition shows that $L_n(\kappa)$ generally becomes larger in operator norm as $\kappa$ increases.
In this section, we investigate the asymptotic properties of the NA-GMM estimator. We first consider a general NA-GMM estimator in a local misspecification setting under some high-level conditions. We then study the asymptotic properties of the estimator in SAR models under general moment misspecification and provide more primitive conditions.
\setcounter{assumption}{0}
Since the NA-GMM estimator is defined based on misspecified moment conditions, its probability limit generally does not coincide with the true $\theta_0$ (hall2003large,hansen2021inference). To characterize the limiting behavior of the NA-GMM estimator, we define the pseudo-true parameter value as follows:
where $Q_{n,\rho}^*(\theta) \coloneqq \| \varphi_n(\theta) \|^2_{\Psi_{n, \rho}^*(\theta)}$, $\varphi_n(\theta) \coloneqq \mathbb{E}[\overline q_n(\theta)]$, and
We first show that, under the following set of conditions, the NA-GMM estimator $\widehat \theta_{n, \rho}$ converges in probability to $\theta_{n,\rho}^*$.
Assumptions (ref) and (ref) are high-level conditions that need to be verified case by case. In the next subsection, we provide more primitive alternative conditions to these in the case of SAR models.
In Assumption (ref), we assume that the IV vector $Z_i$ is fixed and bounded. $Z_i$ typically comprises the covariate vector $X_i$ and its network-lagged version constructed using the observed interaction matrix $\bm G_n^\text{obs}$, and, as mentioned above, $\bm G_n^\text{obs}$ is treated as non-random. Thus, the assumption essentially requires the covariates to be non-stochastic, which is a common setup in the literature on spatial and network econometrics. Assumption (ref) assumes that the GMM weight matrix is fixed, which precludes two-step optimal GMM-type estimators. Combined with Assumption (ref), typical choices of $\Omega_n$ include $\Omega_n = I_{d_\mu}$ and $\Omega_n = (n^{-1} \bm Z_n^\top \bm Z_n)^{-1}$. These assumptions are introduced mainly to simplify the theoretical presentation and to highlight the main properties of NA-GMM. Relaxing them could be possible, but is beyond the scope of this paper.
Next, we derive the asymptotic distribution of the NA-GMM estimator. As pointed out in hall2003large, under general moment misspecification, the form of the asymptotic distribution critically depends on the convergence rate and limiting distribution of the weight matrix $\Psi_{n,\rho}(\theta)$. Thus, a fully general discussion would complicate the analysis and would not necessarily be informative for practical use. To avoid this complexity, we first consider a local misspecification case in this subsection. The case of general moment misspecification will be discussed in the next subsection.
Let $\mathcal N_0 \subset \Theta$ be a given fixed neighborhood of $\theta_0$.
Assumption (ref)(i) is a local misspecification condition, which can be equivalently written as $\mathbb{E}[\overline u_n(\vec{\bm D}_{\mathcal{U}}; \theta_0)] = -\delta_n/\sqrt{n}$. In the SAR model in Section (ref), we have
provided that $\mathbb{E}[Y_j]$'s and $\vec{\bm D}_{\mathcal{U}}$ are uniformly bounded. Thus, Assumption (ref)(i) is satisfied if $m = O(\sqrt{n})$. This condition covers empirical situations in which only a small part of the network is uncertain. For instance, in survey-based network data, respondents may be allowed to nominate only a fixed number of peers. Most respondents may not reach this nomination limit, but a few respondents may have more peers than the limit allows. The condition also covers cases in which the survey is incomplete only for a small number of schools, villages, or firms.
Assumptions (ref) and (ref) are mild regularity conditions. By contrast, Assumptions (ref) and (ref) directly impose asymptotic properties on the sample moments and the weight matrix. In the SAR model considered below, these conditions can be verified under relatively primitive conditions.
The second result of Theorem (ref) implies that, as a special case of Assumption (ref)(i), if $\delta_n = o(1)$, then the limiting distribution of $\sqrt{n}(\widehat\theta_{n,\rho} - \theta_0)$ is centered at zero. A closely related result can be found in Proposition 3.2 of lewbel2024ignoring.
\setcounter{assumption}{0}
We next study the asymptotic properties of the NA-GMM estimator for SAR models in a general network error setting. The model and the choice of IVs are the same as those in Example (ref) and Section (ref). The GMM weight matrix $\Omega_n$ is allowed to be any matrix satisfying Assumption (ref).
We first show that the high-level conditions introduced in the previous subsection can be verified under the following set of primitive assumptions.
In Assumption (ref), we assume that the error terms are IID. This assumption is stronger than necessary, but is imposed for simplicity and to demonstrate the properties of the NA-GMM estimator more transparently. Noting that $Z_i = (\sum_{j \neq i} G_{ij}^{\mathrm{obs}} X_j^\top, X_i^\top)^\top$, Assumption (ref) parallels Assumption (ref). Assumption (ref) is standard in the literature and limits the magnitude of social interactions.
Assumption (ref) is essential. It requires that the size of the uncertainty set for each unit, $m_i$, is uniformly bounded, and also that each unit does not appear in other units' uncertainty sets too many times. The assumption implies that $m = O(n)$, and therefore accommodates a wider range of network uncertainty than the local misspecification setting discussed above. Roughly speaking, the uncertainty set should have a locally clustered structure. This condition should be natural in many empirical applications where the network consists of multiple clusters, such as schools, grades, villages, or firms, and units in different clusters typically do not have close interactions.
Assumption (ref) is the main identification condition in the context of SAR models, where condition (i) ensures the identification of the pseudo-true parameter $\beta_{n,\rho}^*$ for $\beta$, and condition (ii) assumes the identifiability of $\alpha^*_{n,\rho}$, the pseudo-true parameter for $\alpha$. Assumption (ref) is a technical condition that rules out variance degeneracy of the NA-GMM estimator. In Lemma (ref), we prove that $\mathcal A_n(v) \bm\varepsilon_n + \bm \varepsilon_n^\top \mathcal B^s_n(v) \bm \varepsilon_n - \mathbb{E}[\bm\varepsilon_n^\top \mathcal B^s_n(v) \bm\varepsilon_n]$ is asymptotically distributed as $N(0, \omega_\rho^2(v))$ under the assumptions made here. The explicit expression for $\omega_\rho^2(v)$ can be obtained by applying the definitions of $\mathcal A_n(v)$ and $\mathcal B_n^s(v)$ to (3.2) of kelejian2001asymptotic.
We first show in the next lemma that the assumptions given above are sufficient to verify the main high-level conditions imposed in Theorems (ref) and (ref).
Lemma (ref) implies that, under Assumptions (ref)--(ref), the NA-GMM estimator for the SAR model is consistent for the pseudo-true parameter and asymptotically normal in the sense of Theorem (ref), if the misspecification of the moment conditions is mild. In fact, under the same conditions, we can show that, even when the moment misspecification is not local, the estimator is asymptotically normally distributed around the pseudo-true parameter.
As shown in Theorem (ref), the NA-GMM estimator for the SAR model is asymptotically normal around the pseudo-true parameter under fairly general network error settings. However, statistical inference based on this result appears challenging, not only because the asymptotic variance takes an extremely complicated form, but also, more fundamentally, because it depends on the true parameter $\theta_0$ and the true interaction matrix $\bm G_n$, which are fully unknown in our context. One might consider alternative simulation-based inference procedures; however, these generally require prior knowledge of the true (or, at least, approximately true) interaction structure (e.g., kojevnikov2021bootstrap, conley2023bootstrap). Moreover, there is some debate about the empirical usefulness of statistical inference for pseudo-true parameters (e.g., hansen2021inference, andrews2026true), despite the desirable bias reduction property of NA-GMM estimators discussed in Proposition (ref).
Considering these issues, we suggest using the NA-GMM method as a diagnostic tool for assessing the robustness of the naive GMM estimator to network uncertainty. In the next subsection, we provide several diagnostic analyses that may be useful in empirical applications.
In NA-GMM, $\rho$ should be interpreted more as a sensitivity parameter to network errors than as a regularization parameter governing the prediction performance. Thus, the NA-GMM method can be used most informatively when we examine how the estimate $\widehat\theta_{n,\rho}$ changes along the path of $\rho$. If the estimate remains stable even for relatively small values of $\rho$, this suggests that the naive GMM estimate is not very sensitive to the type of network uncertainty captured by the network adjustment procedure. By contrast, if the estimate changes substantially as $\rho$ decreases, this indicates that the empirical conclusion may be sensitive to potential network errors. Thus, such a path serves as a direct diagnostic for the robustness of the naive GMM estimator to network uncertainty.
As a diagnostic statistic for quantifying the improvement in moment fit due to the network adjustment, we consider
where $\vec{\bm d}_m^\circ(\theta)$ is the optimal network correction vector, whose definition is given in (ref) in Appendix. This statistic measures the relative improvement in moment fit achieved by the NA-GMM adjustment.
Note that $0 \le r_n(\rho) \le 1$ holds as long as $\left\| \overline q_n(\widehat \theta_n^\text{naive}) \right\|_{\Omega_n} > 0$. Indeed, we have
Hence, a value of $r_n(\rho)$ close to one indicates a large improvement in moment fit, whereas a value close to zero indicates that the network adjustment yields little improvement for this $\rho$.
As another diagnostic statistic for measuring the extent of moment fit improvement, we consider a $J$-type statistic. In SAR models, define
where $\widehat\sigma^2_{n,\rho} \coloneqq n^{-1} \left\| \bm Y_n - \bm R_n^\text{obs}\widehat \theta_{n,\rho} \right\|^2$. When $\rho=\infty$, the adjustment term disappears and $\widehat \theta_{n,\infty}$ reduces to the naive GMM estimator. In particular, if $\Omega_n = (n^{-1} \bm Z_n^\top \bm Z_n)^{-1}$, we have
which coincides with the classical Sargan $J$-statistic. Similar to the preceding diagnostic statistic, $J_n(\rho)$ summarizes the improvement in the fit of the moment conditions.
In this section, we examine the finite sample performance of NA-GMM in linear SAR models. The purpose of the simulation is to evaluate the bias reduction property of the fixed-weight NA-GMM and the CU-type original NA-GMM estimators.
The simulation design is as follows. For each sample size $n \in \{500,1000\}$, we first randomly allocate $n$ units on a square lattice of size $\sqrt{1.2n} \times \sqrt{1.2n}$. The true adjacency matrix $\bm A_n$ is then constructed based on these locations. For each unit $i$, we first define its peer candidates as units located within distance three from $i$, and then randomly select a number of candidates from $\{0,1,\ldots,5\}$ to form links with $i$. The true interaction matrix $\bm G_n$ is obtained by row-normalizing this adjacency matrix. For each $i$, the uncertainty set $\mathcal{U}(i)$ consists of all units excluding $i$ located within distance five from $i$. The outcome is then generated from the linear SAR model
where $\alpha_0 = 0.6$, $(\beta_{00}, \beta_{10}, \beta_{20}) = (1.2,-1,1.4)$, $X_{1i} \sim N(0,1)$, $X_{2i} \sim |N(0,1)|$, and $\varepsilon_i \sim N(0,1)$ independently across $i$.
After the data are generated, we create the imprecisely observed network $\bm A_n^{\mathrm{obs}}$ by deleting true links and adding false links to $\bm A_n$. Specifically, each true link in the uncertainty set is deleted with probability $p_{\mathrm{drop}} \in \{0,0.2,0.3\}$, and each absent link in the uncertainty set is added as a false link with probability $p_{\mathrm{add}} \in \{0,0.05,0.1\}$. The observed interaction matrix $\bm G_n^{\mathrm{obs}}$ is obtained by row-normalizing $\bm A_n^{\mathrm{obs}}$. For each design, we compare the performance of the naive GMM estimator based on $\bm G_n^{\mathrm{obs}}$, the fixed-weight NA-GMM estimator, and the original NA-GMM estimator. The IVs for the spatial autoregressive term are constructed from the first and second spatial lags of $(X_1, X_2)$ based on $\bm G_n^{\mathrm{obs}}$. For NA-GMM and fixed-weight NA-GMM, we consider two GMM weight matrices: the identity weight $\Omega_n = I_{d_\mu}$ and the 2SLS weight $\Omega_n = (n^{-1}\bm Z_n^\top \bm Z_n)^{-1}$. Note that the naive GMM estimator with the 2SLS weight corresponds to the usual 2SLS estimator. For the penalty parameter in the NA-GMM estimators, we consider $\rho \in \{2^{-6},2^{-5},\ldots,2^{14}\}$, where we set $\kappa = \rho/\alpha_0^2$ for fixed-weight NA-GMM. Each simulation setup is repeated 500 times, and the performance of each estimator is evaluated by the average bias and root mean squared error (RMSE) for estimating $\alpha_0$ and $\beta_{10}$.
To save space, the summary tables for the simulation results with $\rho = 2^{-3}$, $2^8$, and $2^{12}$ are provided in Appendix (ref). The main findings from these tables are as follows. First, for the estimation of the coefficient $\beta_{10}$, the three GMM estimators perform very similarly, and all of them produce accurate estimates of $\beta_{10}$ in all scenarios. This is not surprising because network errors are generated independently of the covariates.
We now focus on the estimation of $\alpha_0$. When there are no network errors, not only the naive GMM estimator but also the two NA-GMM estimators estimate $\alpha_0$ very accurately. This is an expected result because, when there are no errors in the observed interaction matrix, the population optimal network correction is $\vec{\bm d}_m = \bm 0$, which reduces the NA-GMM criterion to the standard naive GMM criterion. In the presence of network errors, for both types of GMM weights, the original NA-GMM estimator with small $\rho$ generally outperforms the other estimators in terms of bias reduction. This corroborates the desirable bias reduction property of the NA-GMM method.
Interestingly, although Proposition (ref) shows that the fixed-weight NA-GMM estimator also has a bias reduction property, the magnitude of bias reduction is much smaller than that of the original NA-GMM estimator. In particular, almost no bias reduction is observed when we set $\Omega_n = (n^{-1}\bm Z_n^\top \bm Z_n)^{-1}$. This finding may be explained by the structure of the GMM weight matrix. As shown in (ref), when $\Omega_n = (n^{-1}\bm Z_n^\top \bm Z_n)^{-1}$, the fixed-weight NA-GMM estimator uses the weight matrix $\Psi_{n,\kappa} = (n^{-1}\bm Z_n^\top \bm Z_n + (m/\kappa)\Phi_n)^{-1}$. Noting that $\Phi_n = n^{-2} \sum_{i \in [n]} \sum_{j \in \mathcal{U}(i)} Z_i Z_i^\top Y_j^2$, if $n^{-1}\sum_{j \in \mathcal{U}(i)} Y_j^2$ does not vary much across $i$, then the network adjustment component $(m/\kappa)\Phi_n$ becomes approximately proportional to $n^{-1}\bm Z_n^\top \bm Z_n$ for any $\kappa$. In such cases, fixed-weight NA-GMM would behave very similarly to 2SLS, as observed here. These findings indicate that allowing the GMM weight matrix to depend on the parameter plays an important role in reducing the bias. Another interesting finding is that, even for the original NA-GMM estimator, we do not observe any bias reduction when network errors arise only from dropping existing links. This indicates that, while NA-GMM can effectively reduce some types of bias, it may be less effective for other types. These points deserve further investigation in future work.
Figures (ref) and (ref) summarize the simulation results for estimating $\alpha_0$. The gray area in each panel corresponds to the simulated pointwise 95 percent interval at each value of $\rho$. The results discussed above can also be visually confirmed in these figures.
In this section, we illustrate the NA-GMM method by applying it to U.S. county-level COVID-19 infection rate data in 2022. County-level COVID-19 case data for 2022 were obtained from the data archive of the Centers for Disease Control and Prevention (CDC).\footnote{\url{https://data.cdc.gov/}} For county characteristics, we use the median household income, the share of bachelor's degree (or higher) holders, and the unemployment rate. These variables, together with the county population data, were obtained from the U.S. Department of Agriculture (USDA) Economic Research Service county-level data.\footnote{\url{https://www.ers.usda.gov/data-products/county-level-data-sets}}
After excluding non-contiguous regions, isolated counties, and one county with an extreme outlying infection rate, the final sample consists of 3,099 contiguous U.S. counties. Using the annual infection rate, defined as total cases in 2022 divided by the population, as the outcome variable, we consider the following SAR model:
Here, the observed counterpart of $G_{ij}$, $G^\text{obs}_{ij}$, is the row-normalized version of $A^\text{obs}_{ij}$, where $A^\text{obs}_{ij}$ is equal to one if counties $i$ and $j$ are adjacent and zero otherwise. Figure (ref) shows a map of $\texttt{infection\ rate}_i$ across counties. This pattern indicates that COVID-19 infection rates in adjacent counties are strongly related, suggesting that $\alpha_0$ is relatively large and positive in the SAR model.
We first estimated the model by 2SLS, using the first and second spatial lags of the regressors as IVs. The estimate of $\alpha_0$ was about 0.8 and significantly positive. However, the computed Sargan $J$-statistic was 12.17, suggesting possible misspecification of the moment conditions. We next estimated the same model using the NA-GMM method. Figure (ref) reports the path of the NA-GMM estimate $\widehat \alpha_{n,\rho}$ over different values of $\rho$. As shown in the figure, as $\rho$ tends to infinity, the estimate converges to the 2SLS estimate, which is consistent with our theory. Moreover, even for very small values of $\rho$, the estimate deviates only slightly from the 2SLS estimate.
Interestingly, the two diagnostic statistics for the improvement in moment fit introduced in Subsection (ref) show that the moment fit is substantially improved for small values of $\rho$, as reported in Figures (ref) and (ref). Taken together, these findings suggest that, although the moment conditions appear to be misspecified to some extent, the moment correction induced by the network adjustment does not lead to a substantial change in the estimate of the spatial parameter $\alpha_0$. Thus, the estimate of $\alpha_0$ appears relatively robust to the type of network uncertainty in this dataset.
In this paper, we proposed a network-adjusted generalized method of moments (NA-GMM) framework for IV-based linear social interaction models under network uncertainty. The proposed method modifies the standard GMM approach by allowing the observed interaction matrix to be adjusted while penalizing the size of the adjustment. For linear SAR models, we showed that the NA-GMM criterion downweights moment directions that are sensitive to network errors and established a bias reduction property for a fixed-weight version of the estimator. We also established consistency for the pseudo-true parameter and asymptotic normality under general moment misspecification. Finally, we proposed diagnostic analyses for evaluating the robustness of naive GMM estimates to network errors. An empirical application to U.S. county-level COVID-19 infection data demonstrated the usefulness of the proposed method.
The present analysis leaves several open questions and potential extensions. First, it would be important to further investigate the bias reduction property of NA-GMM; this paper established a related result only for a fixed-weight version in linear SAR models. Second, developing feasible statistical inference procedures for NA-GMM estimators would be important for improving the applicability of the proposed framework. The present framework could also be extended to more general models, such as models in which network errors enter nonlinearly and models with endogenous network errors. Finally, while we adopted a quadratic penalty for the network adjustment for convenience, it would be worth studying more general penalty structures. These topics are left for future research.