Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
144,382 characters · 27 sections · 64 citation commands
Estimating Network Spillovers under Dense Measurement Error
{JEL Classification:} C21, C23, D57
{Keywords:} network analysis, measurement error, LASSO, nuclear norm penalty, penalized GMM
{
Network analysis provides powerful tools to quantify policy spillovers and identify pivotal agents for targeted interventions, particularly when policymakers face binding resource constraints Ballester2006, Zenou2016, Galeotti2020. A large literature shows that economic outcomes often depend on interactions among interconnected units, giving rise to amplification, diffusion, and multiplier effects that are central to policy design Acemoglu2012, Giorgi2020, Araujo2023. These insights have motivated the widespread use of spatial and social interaction models to study peer effects, diffusion dynamics, and network-based policy interventions across a broad range of economic settings.
Methodologically, the literature on spatial and social interaction models differs mainly in how the network structure is handled. A dominant strand assumes that the adjacency matrix is correctly specified and directly observed, and studies spillover effects conditional on this object Lee2007a, Bramoulle2009, Lee2010a, Banerjee2013, Elhorst2014, Lee2014, Yang2017, Zhu2020a, Kuersteiner2020,\\ Shin2021, chernozhukov2021uniform. An alternative and growing line of research relaxes this assumption by estimating the network jointly with model parameters, typically imposing dimension reduction constraints such as sparsity to address the curse of dimensionality in high-dimensional settings Manresa2016, Lam2020, Paula2023.
Sparsity-based methods work well when the latent network is dominated by a few strong links, but they fail in the dense networks that arise in production, trade, and financial settings, where the observed adjacency matrix is itself contaminated by measurement error FismanWei2004, Upper2011, Anand2015. In such settings, two distinct forms of latent structure coexist beyond idiosyncratic noise. First, outcomes may be driven by a small number of pervasive (though individually weak) latent forces such as common shocks, shared technologies, or institutional linkages, naturally represented by a low-rank component Chudik2013, Kapetanios2021. Second, networks may contain a small number of strong bilateral or idiosyncratic interactions, naturally represented by a sparse component Lam2020, Paula2023. Imposing sparsity alone discards the pervasive signal, while ignoring measurement error produces non-vanishing bias as elementwise small errors accumulate along network paths Lewbeletal2024EJ. The stakes are practical: in our cross-state tax-competition application, the standard spatial GMM estimator with the raw inverse-distance weight matrix yields $\widehat\lambda > 1$, violating the Leontief stability condition and implying explosive multipliers, while our estimator returns $\widehat\lambda < 1$ within the admissible region. We therefore propose a regularization approach for the observed adjacency matrix that filters measurement error while preserving low-rank or sparse signal, developed formally in Section (ref).
Throughout, by regularization we mean imposing restrictions: low-rank, sparse, or low-rank-plus-sparse, directly on the latent adjacency matrix $W_0$ rather than on the regression coefficients, as identifying restrictions that make denoising of the observed network feasible.
The problem is non-trivial because it combines matrix-valued measurement error with a spillover estimator that depends nonlinearly on the network matrix through the Leontief inverse $(I_n - \lambda W)^{-1}$, so errors in $W$ propagate through the estimator in ways classical errors-in-variables corrections do not handle. Low-rank and sparse structures have deep foundations in statistics and econometrics, appearing in approximate factor models, latent-variable graphical models, and high-dimensional covariance estimation Bai2002, Chandrasekaran2010, Fan2011. A large methodological literature develops penalized procedures to recover low-rank or sparse patterns in covariance and precision matrices Cai2011, Candes2011, Agarwal2012, Fan2018. Despite their success in these settings, regularization of this kind has not, to our knowledge, been adopted in the econometrics literature to address measurement error in observed network adjacency matrices, or to improve the estimation of spillover effects and other policy-relevant network statistics. We build on this literature by developing two complementary estimation approaches. The first is a two-stage procedure that denoises the observed adjacency matrix by separating low-rank or sparse signal from measurement error before estimating spillover effects. The second is a regularized Generalized Method of Moments (GMM) estimator that jointly estimates the network and the model parameters in a single step. We establish that our estimator is consistent, in contrast to naive estimation that ignores measurement error and remains inconsistent in dense regimes. Simulation results show that the proposed method achieves a $50\text{--}80\%$ reduction in the root mean squared error of spillover estimates under noisy weight matrices, with measurement noise ranging from moderate to high.
Our main theoretical contribution is to characterize how the bias of standard spillover estimators depends on the structure of the latent network and the measurement error, and to show that this bias does not vanish in the dense regime that characterizes economic networks. Lewbeletal2024EJ show that modest link misclassification bias in the estimation of linear social-interaction models may be negligible as long as misclassification is not pervasive; we extend their analysis to settings in which the latent network has many but individually weak connections, and in which the measurement error itself need not be sparse. We further allow the measurement error to be correlated with unobservables in the outcome equation Candelaria2023, lewbel2024estimating, which can introduce non-vanishing bias if ignored. We derive convergence rates for both the spillover-parameter estimator and our estimated adjacency matrix. Beyond these formal results, the framework also delivers an interpretable decomposition of network connections (a low-rank component capturing pervasive relationships and a sparse component isolating localized interactions, with identification following standard strategies in the literature chandrasekaran2011rank) and we use this decomposition to provide a novel expression of the network eigencentrality that quantifies the contributions of each component.
We illustrate the practical relevance of our framework through two empirical applications: international spillovers in GDP growth across 23 economies and tax competition among 48 U.S.\ states. In both settings, our estimated network produces spillover estimates that differ markedly from those obtained using the raw adjacency matrix, highlighting the empirical importance of network denoising. In the cross-economy GDP-growth application, regularization affects not only the scale of spillovers but also the inferred structure of interdependence. The network decomposition further reveals distinct regional and global components of cross-economy interdependence. These empirical findings demonstrate how our framework can uncover the structure of interdependence in economic networks and support more reliable inference in spatial econometric analysis. In the U.S.\ tax-competition application, the raw inverse-distance weight matrices yield explosive estimates $\widehat\lambda > 1$ that violate the Leontief stability condition, while our estimator returns stable estimates.
The remainder of our paper is organized as follows. Section (ref) motivates the framework by explaining why measurement error matters in dense networks, illustrating four possible network structures (sparse, low-rank, low-rank-plus-sparse), and demonstrating with the U.S.\ Input-Output Table that real-world dense networks exhibit exploitable regularization structure. Section (ref) presents the spatial interaction model and introduces the plug-in and regularized GMM estimators under low-rank, sparse, or low-rank-plus-sparse structure. Section (ref) provides assumptions and establishes consistency and asymptotic normality results for both estimators. Section (ref) presents simulation evidence. Section (ref) applies our methods to estimate spillover effects using noisy network data. Section (ref) concludes. Proofs, algorithms, and additional technical details are provided in the appendices. }
This section sets up the notation and previews the regularization framework used throughout the paper. Let $\mathcal{N} \coloneqq \{1, \ldots, n\}$ denote the set of $n$ units or nodes and $\mathcal{E}$ the set of edges between them, so that $(\mathcal{N}, \mathcal{E})$ represents the graph. Let $W_0$ be the $n \times n$ adjacency matrix encoding the true connections, with $(i,j)$-th element $W_{0,ij}$; we write $W_{0,ij} \neq 0$ when unit $i$ is connected to unit $j$ and $W_{0,ij} = 0$ otherwise.\footnote{We allow the connection from $i$ to $j$ to differ from that from $j$ to $i$, so the graph may be directed and $W_0$ need not be symmetric.} We further assume no self-loops, so that $W_{0,ii} = 0$ for all $i = 1, \ldots, n$.
The econometrician observes a noisy version of $W_0$, \[ W = W_0 + E, \] where $E$ is a measurement-error matrix whose sources are discussed in Section (ref). Before formally motivating the regularization framework, we briefly preview the structural decomposition of $W_0$ that plays a central role throughout the paper.
To reduce estimation complexity for large $n$, it is natural to impose structural assumptions on $W_0$. The most common choice (sparsity) is plausible for social networks dominated by a few strong links, but fails to describe networks in which a large number of units share common interests or characteristics: a highly sparse network typically splits into several disconnected components. We therefore propose to model $W_0$ as the sum of a low-rank and a sparse component,
where $L_0$ captures pervasive, common connections and $S_0$ captures sporadic or idiosyncratic connections.
The components $L_0$ and $S_0$ need not individually satisfy the zero-diagonal constraint. When $W_0$ has a zero diagonal, we can define $L_0^* = L_0 - \operatorname{diag}(L_0)$ and $S_0^* = S_0 + \operatorname{diag}(L_0)$, which absorb the diagonal corrections so that both represent networks without self-loops:
While $L_0^*$ is not exactly low-rank, it inherits the low-rank structure of $L_0$ and differs only on the diagonal. Section (ref) illustrates the usefulness of this decomposition through concrete examples. We further develop our theoretical analysis (Section (ref)) in terms of $L_0$ and $S_0$ and note that the results extend directly to $L_0^*$ and $S_0^*$. Both theoretical and empirical evidence in the following sections show that this decomposition captures several meaningful network structures encountered in practice Diebold2014, Barranca2015, Gamble2016, newman2010networks.
A central feature of the applications mentioned in the introduction is that the observed network is dense and measured with error. We document four such sources of error in Section (ref); here we develop the consequences for spillover estimation. The existing literature, such as Lewbeletal2024EJ, assumes that measurement error in the network is sufficiently sparse that it does not affect the consistency of the spillover estimator. In particular, consistency of $\widehat{\lambda}$ can be preserved when the aggregate magnitude of the error vanishes asymptotically, for example when \[ \frac{1}{n}\sum_{i,j} |E_{ij}| = o_p(1). \] This condition is naturally satisfied when the total error budget grows slowly relative to the size of the network, as in settings where the underlying adjacency matrix itself is sparse. Intuitively, when only a small fraction of links are contaminated and the magnitude of each contamination is small, the resulting perturbation to the network propagation mechanism becomes negligible in large samples.
The situation is fundamentally different when measurement errors are dense. Even when each entry of $E$ is elementwise small, the errors accumulate across the $O(n^2)$ entries of the matrix and across the propagation paths generated by the spatial multiplier. The aggregate distortion may not vanish asymptotically, and the estimator $\widehat{\lambda}_{\mathrm{GMM}}$ may fail to be consistent.
To recover $W_0$ from $W$ in the dense regime, some structural restriction on $W_0$ is needed. Depending on the application, the appropriate structure is:
Misspecifying this structure leaves residual bias: assuming $W_0 = S_0^*$ when a non-negligible $L_0^*$ is present discards the pervasive systematic connections that drive spillovers; assuming $W_0 = L_0^*$ when strong idiosyncratic links are present absorbs those links into the residual. Moreover, the low-rank and sparse decomposition introduced above is the device we use to operationalize this regularization. We emphasize that it is a flexible modelling device, not an economic interpretation. The interactions-model estimator $\widehat\theta_p$ depends only on $\widehat W = \widehat L^* + \widehat S^*$, not on the individual components. We thus do not rely on any economic interpretation of $L_0^*$ and $S_0^*$ as separate network objects. Nevertheless, Appendix (ref) discusses the economic interpretation of $L_0^*$ and provides identification conditions for the $L_0^* + S_0^*$ decomposition.
The practical importance of regularizing $W_0$ rests on three observations. First, measurement error in dense networks is non-negligible: it may not vanish with sample size, so standard asymptotic consistency arguments fail. Second, it is endogenous: errors share common sources and are typically correlated with the true network entries, violating the classical measurement-error assumptions that justify standard instrumental-variable corrections. Third, it is economically consequential: in our empirical illustration with the U.S.\ Input--Output Table (Section (ref)), ignoring measurement error entirely distorts the estimated propagation matrix by $9\%$, and misspecifying the structure of the error (treating it as purely sparse) compounds the distortion to $68\%$. A method that filters dense, structured noise while preserving the signal in $W_0$ is therefore not a refinement but a prerequisite for credible inference about network spillovers.
The economic literature documents several mechanisms through which the observed adjacency matrix $W$ may depart from the latent network $W_0$. Many policy-relevant production, trade, and reconstructed financial networks contain numerous weak links and can be dense relative to conventional social networks. Consequently, measurement error may affect a large fraction of their entries. A related source of potentially dense measurement error arises when the adjacency matrix is constructed from geographical or economic distance LeSage2009, Getis2004: errors in the underlying distance measurem (arising, for example, from centroid approximations, omitted transport links, or bandwidth choices) can propagate across entire rows or columns of $W$. We discuss several additional sources of measurement error below.
A widely studied dense economic network is the input--output table, where the $(i,j)$ entry records the share of commodity $i$ used as an intermediate input in producing commodity $j$. Acemoglu2012 establish that propagation multipliers are highly sensitive to the precise pattern of bilateral input shares across all $n^2$ entries of this matrix. The observed Bureau of Economic Analysis table aggregates transactions across hundreds of heterogeneous firms and products into sector-level flows, and is further subject to reporting inaccuracies, imputation of missing entries, periodic revisions, and reconciliation between the supply and use tables. Given these, the analyst observes a sector-level $W$ that departs from any meaningful sector-level $W_0$, and the resulting errors propagate along multi-step Leontief chains.\footnote{The literature on aggregation in production networks treats this as a separate methodological problem; see e.g.\ barrot2016input and boehm2019input. Our framework is complementary: it addresses reporting and revision error at whatever level the network is observed, and is silent on the firm-to-sector aggregation step.} Our method takes the sector-level network as primitive and denoises around a sector-level truth. We show in Section (ref) that even at the sector level, denoising restores Leontief stability and materially changes estimated cross-sector multipliers.
International trade flows provide a second canonical example of a dense network where measurement error affects every bilateral entry. Trade flows are recorded twice: by the exporting country as exports and by the importing country as imports. These two independent records of the same physical shipment should agree, but in practice they differ substantially and systematically. FismanWei2004 exploit this bilateral mirror-statistic structure to measure the evasion gap, i.e., the log difference between Hong Kong's reported exports to China and China's reported imports from Hong Kong at the six-digit product level and show that a one-percentage-point increase in the tariff rate is associated with a three-percent increase in the evasion gap. This tariff-related underreporting may affect many product categories and can be amplified by product misclassification, as importers reclassify high-tariff goods into lower-tariff categories. The resulting measurement error need not be confined to a small number of trade links and may affect a substantial fraction of the bilateral trade network in a manner correlated with tariff rates. Because tariff schedules are systematically related to trade volumes and revealed comparative advantage, the measurement error $E_{ij}$ tends to correlate with the true $W_{0,ij}$, violating classical errors-in-variables assumptions and biasing estimates of trade-network spillovers.
Bilateral interbank exposures form a naturally dense network for studying systemic risk and financial contagion. Each bank's balance sheet records its total interbank lending and borrowing, but the bilateral breakdown, who lent to whom and in what amount, is typically confidential. Regulatory data under the Basel II framework require reporting of exposures only above a threshold of 10% of regulatory capital, so the majority of bilateral links are unobserved Anand2015. Researchers and regulators therefore reconstruct the full $n \times n$ bilateral exposure matrix from the observable marginals (total lending and borrowing of each bank) using maximum-entropy algorithms. As Upper2011 and the systematic evaluation of Anand2015 document, the maximum-entropy solution fills in the matrix as uniformly as possible, implicitly treating the network as if every bank lent to every other bank in proportion to its total activity. This assumption yields a fully dense reconstructed matrix, but one that systematically underestimates concentration and overestimates diversification: in reality, interbank lending is relationship-based and concentrated among a small set of counterparties. The reconstruction error $E = W - W_0$ is may therefore be dense. It affects every entry of the matrix, and is potentially correlated with the true bilateral exposures, since banks with large total lending are disproportionately misrepresented. Stress-testing exercises that rely on the reconstructed network underestimate contagion risk precisely because the entropy assumption smooths away the concentrated exposures that drive cascading failure.
Even when network data are collected via direct reports rather than algorithmic reconstruction, the observed adjacency matrix is subject to recall error, misunderstanding, and measurement timing mismatches. Lewbeletal2024EJ study economic network models in which some reported binary links are misclassified: true connections are recorded as absent (false negatives) and absent connections are recorded as present (false positives). They establish that this misclassification introduces new sources of endogeneity beyond the standard simultaneity problem in spatial models: the mismeasured adjacency matrix contaminates the spatial lag $WY$, creating correlation between regressors and structural errors, and invalidating standard instrumental variables based on powers of $W$. The result is that conventional Two-Stage Least Squares (2SLS) estimators that ignore misclassification are inconsistent even when the fraction of misclassified links is small relative to $n$, because in a dense network a small misclassification rate still corresponds to a large absolute number of erroneous entries. Panel data settings compound this problem: if the network is observed at lower frequency than the outcome variable, links created or dissolved between survey waves appear as persistent misclassifications correlated with individual fixed effects that cannot be differenced away lewbel2024estimating.
The four mechanisms above are heterogeneous in origin but share the three features highlighted at the end of Section (ref): the measurement error is dense, correlated across entries through shared sources (aggregation, tariff incentives, reconstruction algorithm, or recall bias), and asymptotically non-negligible. These properties together explain why classical errors-in-variables corrections may be insufficient: they presume sparse or vanishing noise that can be handled by instrumental variable (IV) or measurement-error bias corrections, whereas the dense, correlated errors documented here require a different approach. Imposing a regularization on $W_0$ (as a low-rank matrix, a sparse matrix, or their combination) provides the identifying restrictions needed to disentangle the systematic signal from the noise, motivating the framework developed in the remainder of the paper.
To build intuition, we describe four common economic network structures and show how each fits into the low-rank, sparse, or low-rank plus sparse framework. For simplicity, throughout this subsection, we assume the network is observed without measurement error, so $W = W_0$. \footnote{Note that the link matrix or adjacency matrix could be used to describe the connection strength between individuals. If it does not satisfy the bounded row sum assumption, it could be further scaled by its $\ell_1$ norm.}
We begin by considering one of the simplest examples: a fully connected network. In this setup, the link matrix $W_0$ is an $n \times n$ matrix of nonzero values, where it is common to set $W_{0, ii} = 0$ to avoid self-loops. The link matrix corresponding to a complete network generated by nonzero constants $\bm{c}=(c_1, \cdots, c_n)^{\top}$ is defined as
{where $c_i$'s are positive constants within the bounded support $[c_{\min}, c_{\max}]$, and they could be viewed as individual characteristics that vary across $i$. Let $L_0= \bm{c} \bm{c}^{\top}$, which is a symmetric low-rank matrix with rank $1$. Its singular value is $\|\bm{c}\|_2^2$, and the associated singular vector is $\bm{c} /\|\bm{c}\|_2$. Then we could also express $W_0=L_0+S_0=L^*_0$, where the sparse component $S_0$ is a diagonal matrix whose $(i,i)$th element is $- c_{i}^2$ to guarantee no self-loops in $W_0$. Such a network could be used to capture the common interaction structure driven by individuals or agents who interact through a common market or platform. }
Motivated by the structure of the latent factor, we consider the symmetric network in which the strength of the connection is determined by $W_{0,ij}=a_i+z_iz_j$ if $i\neq j$, { where $a_{i}$'s are individual characteristics and $z_i$'s are some latent variables. \footnote{In a more general term, the low rank component can be understood as network effects related to a principal agent, see Galeotti2020 . The principal agent is formed by connections with respect to many nodes within the network, and the effects are dense.} Denote $\bm a$ as a vector consisting of $a_i$. In this example, $W_0=L_0+S_0=L_0^*$, where $L_0={\bm a} {\bf 1}_n^\top + {\bm z} {\bm z}^\top$ is a matrix of rank no more than 2 with ${\bm a}=(a_1, \cdots, a_n)$ and ${\bm z}=(z_1, \cdots, z_n)$, and $S_0$ is a diagonal matrix whose $(i,i)$th element is $-L_{0,ii}$ to remove self-loops. Networks of this kind arise in production networks and financial systems, where common shocks or institutional linkages generate pervasive, dense interactions.}
Another example of interest is the pure group network, where there are two or more communities, potentially of distinct sizes to avoid eigenvalue multiplicity in the link matrix. Moreover, individuals within the same community are assumed to be connected with each other, while individuals do not connect to others outside their community. For example, suppose $n=20$, and $6$ individuals belong to community $1$ and the rest belong to community $2$. We define
where $L_1=(c_1, c_2, \cdots, c_6)^{\top}(c_1, c_2, \cdots c_6)$ and $L_2=(c_7, \cdots, c_{20})^{\top}(c_7, \cdots, c_{20})$, and $S_0$ is a diagonal matrix such that $S_{0,ii}=-c_i^2$. Note that $W_0$ still admits the low-rank and sparse decomposition. The low-rank component admits a block structure and its rank equals the number of communities. One of its two singular vectors is proportional to $(0, \cdots, 0, c_7, \cdots, c_{20})$, while the other is proportional to $(c_1, \cdots, c_6, 0, \cdots, 0)$. See Figure (ref) for the network graph and the heatmap of $W_0$. {This setting captures trade blocs, regional banking clusters, and sectoral linkages, and the rank of $L_0$ equals the number of communities. We could also extend it to allow sporadic between-group connections, which are described by few nonzero off-diagonal elements in $S_0$. Then the general model could be written as $W_0=L_0+S_0=L_0^*+S_0^*$, where the low-rank component captures the within-group connections and the sparse component represents the between-group connections.}
An interesting example comes from considering networks containing a set of dominant units also known as key players, see CalvoArmengol2009, Zenou2016. One definition of dominant units used in the network literature identifies them as the set of units whose actions have large and persistent effects on the other agents.\footnote{We follow the definition of dominant units as in Pesaran2020, Pesaran2021, Lee2022. } One way to measure the effect of the unit $j$ on others is to use its weighted out-degree $d_j$, defined simply by the $j$-th column sum of $W_0$: $d_j = \sum_{i = 1}^n W_{0,ij}$. We then call the unit $j$ a dominant unit if its weighted out-degree is large or grows as the sample size does. For example, suppose $n=20$ and there is only one dominant unit, namely the first individual, who is connected to $13$ individuals. We assume that each nondominant unit is also connected to its immediate neighbors. Then the link matrix $W_0$ is specified as below:
See Figure (ref) for the network graph and the heat map of $W_0$.
This network is naturally sparse. The number of nonzero entries in each row is bounded, although the column corresponding to the dominant unit may contain a growing number of nonzero entries. While this column can be represented as a sparse rank-one matrix, the full adjacency matrix need not be low-rank because it also contains links among neighboring units. We therefore treat $W_0$ as a sparse matrix. Appendix (ref) provides further discussion of alternative low-rank-plus-sparse representations and their identification.
{
To demonstrate that measurement error in dense economic networks has economically meaningful consequences, and that the $L+S$ decomposition provides an effective device for recovering the underlying true network structure, we revisit the 2002 US Input-Output table from graham2020econometric, compiled by the Bureau of Economic Analysis and studied by Acemoglu2012. It contains an $n \times n$ network of $n=417$ sectors (after dropping rows and columns with zero direct input requirements), where entry $(i,j)$ indicates the share of sector $i$'s output used in sector's $j$ production. This network is dense by construction(nearly all sectors interact) making it a canonical example where sparsity cannot be assumed and measurement error is pervasive. We treat the observed $W$ as a noisy version of the true network $W_0$ and apply our Algorithm (ref) (Appendix (ref)) to recover $W_0$.
To assess the quality of the denoising, we compare three candidate representations of $W_0$, a purely sparse estimate $\widehat{S}$, a purely low-rank estimate $\widehat{L}$, and the combined version $\widehat{W} = \widehat{L}+\widehat{S}$. The purely sparse estimate satisfies $\|\widehat{S}-W\|_{\max}=0.025$ and $\|\widehat{S}-W\|_2=0.470$, indicating that the tiny elements might have non-negligible aggregation effects. The purely low-rank estimate satisfies $\|\widehat{L}-W\|_{\max}=0.999$ and $\|\widehat{L}-W\|_2=1.214$, indicating that the low-rank structure alone might fail to capture occasionally large idiosyncratic links. The combined estimate achieves $\|\widehat{W}-W\|_{\max}=0.024$ and $\|\widehat{W}-W\|_2=0.183$, materially outperforming either alternative. This superiority is further confirmed by eigenvector alignment. The inner product between the leading eigenvectors of $W$ and $\widehat{W}$ is $0.981$, compared with $0.526$ for $\widehat{S}$ and $0.222$ for $\widehat{L}$, indicating that the low-rank and sparse decomposition might best preserve the spectral structure of the observed network. From Figure (ref), we could see that the sparse component $\widehat{S}$ captures large idiosyncratic bilateral links, while the low-rank component $\widehat{L}$ exhibits a distinctive stripe pattern, revealing that pervasive but individually weak linkages aggregate into systematic structure that a purely sparse representation would miss.
To quantify the econometric consequences of measurement error, we compare the Leontief diffusion matrices $(I_n - \lambda \widehat{W})^{-1}$, $(I_n - \lambda W)^{-1}$, and $(I_n - \lambda \widehat{S})^{-1}$ following acemoglu2016networks with $\lambda = 0.7$. Leaving measurement errors uncorrected produces a $9\%$ deviation in the spectral norm relative to our estimators: \[ \frac{\|(I_n-0.7\widehat{W})^{-1}-(I_n-0.7W)^{-1}\|_2} {\|(I_n-0.7\widehat{W})^{-1}\|_2}=9\%. \] Misspecifying the structure of measurement errors by treating it as a purely sparse discarding the low-rank component compounds the distortion substantially, producing a $68\%$ deviation: \[ \frac{\|(I_n-0.7\widehat{W})^{-1}-(I_n-0.7\widehat{S})^{-1}\|_2} {\|(I_n-0.7\widehat{W})^{-1}\|_2}=68\%. \] These results demonstrate that measurement error in dense networks is not merely a theoretical concern: it materially distorts estimated propagation channels. Hence it is essential for reliable inference to correctly specify the error structure accounting for both its sparse idiosyncratic component and its low-rank systematic component. Figure (ref) illustrates this point visually: the stripe pattern present in the full diffusion matrix $(I_n - 0.7\widehat{W})^{-1}$ is eliminated when the low-rank component is omitted, confirming that misspecification of network structure biases spillover estimates in economically meaningful ways.
}
We present the general model in which the latent network $W_0 = L_0 + S_0$ admits both a low-rank and a sparse component, with the purely low-rank and purely sparse settings arising as special cases. We emphasize that this decomposition need not be unique or economically interpretable on its own; rather, the joint regularization imposed on $L_0$ and $S_0$ provides sufficient structure to separate $W_0$ from the noise $E$. We provide the modeling definitions and estimation procedures we propose to tackle this problem. We offer additional assumptions and asymptotic results in Section (ref).
We formally describe here the social/spatial interaction model of interest. We consider a panel data scenario, but assume the true underlying network does not vary with time. Denote data points as $\{Y_t, X_t\}_{t=1, \cdots, T}$, where $Y_t$ is the $n \times 1$ outcome vector and $X_t$ is the $n \times K$ design matrix of exogenous covariates. We assume that the true model is generated by the unobserved $n \times n$ adjacency matrix $W_0$ encoding the true connection, and that we have a noisy or contaminated observation, such that we can write
where $\lambda_0$ represents the scalar (social / spatial) interaction or spillover effect, $\beta_0 \in \mathbb{R}^K$ are standard slope coefficients, $\gamma_0 \in \mathbb{R}^K$ are contextual effects (spillover effects of the observable characteristics $X$), $\alpha = (\alpha_1, \ldots, \alpha_n)^\top$ is the vector of individual fixed effects, $\iota_t$ is a time fixed effect at period $t$, ${\bf 1}_{n\times 1}$ is an $n \times 1$ vector of ones, and $\varepsilon_t$ is an $n$-dimensional vector of idiosyncratic disturbances, assumed for now to be independent and identically distributed over units $i$ and time $t$. We assume the covariates $X_t$ to be exogenous to the outcome unobservables and network formation process such that $\operatorname{\mathbb{E}}[\varepsilon_{it} \mid X_t, W_0X_t, \cdots, W_0^dX_t] = 0$, for a small integer $d>0$, $i = 1,\cdots, n$ and $t = 1, \cdots, T$. Crucially, we allow for the measurement errors to be endogenous in the sense that $\operatorname{\mathbb{E}}[\varepsilon_{it} \mid E] \neq 0$, which introduces additional bias into conventional estimates when $W_0$ is replaced by $W$ boucher2025estimating, hardy2019estimating, Lewbeletal2024EJ.
Collect all parameters of interest into $\theta_0 \coloneqq (\lambda_0, \beta^{^{\top}}_0, \gamma^{^{\top}}_0)^{^{\top}} \in \mathbb{R}^{2K+1}$. When covariates $X_t$ are exogenous and the adjacency matrix $W_0$ is correctly observed, the literature proposes to select $m$ linearly independent columns from $[X_t, W_0 X_t, \ldots, W_0^d X_t]$, for a small positive integer $d$, as instruments for the regressors $\bar{X}_t(W_0) = [W_0 Y_t, X_t, W_0 X_t]$ in estimating $\theta_0$ Kelejian1998, Lee2003, Lee2007.
The degree $d$ should be chosen such that the number of retained instruments is at least $2K+1$ to satisfy the order condition, i.e. $2K+1 \leq m \leq (d+1)K$ (see Assumption (ref)). For an arbitrarily pre-specified $n \times n$ adjacency matrix $M$, let $Z_t(M)$ be the $n \times m$ matrix that collects the linearly independent columns from $[X_t, M X_t, \ldots, M^d X_t]$ and let ${X}_t(W)$ be the $n \times (2K+1)$ matrix that collects the regressors $[W Y_t, X_t, W X_t]$. Note that ${X}_t(W)$ uses the observed adjacency matrix $W$ as opposed to the true $W_0$ showing up in model (ref).
By defining $\varepsilon_t(\theta, W) \coloneqq Y_t - {X}_t(W) \theta$ and thus $\varepsilon_t =\varepsilon_t(\theta_0, W_0)$ , we obtain the population moment condition $\operatorname{\mathbb{E}}[Z_t^{\top}(M)\varepsilon_{t}] = \mathbf{0}_{m\times 1}$ for any exogenous $M$. Thus we have the empirical moment conditions:
The distinction made between the matrix used to construct the regressors ${X}_t(W)$ and the instruments $Z_t(M)$ is necessary to define our supervised estimator below. When $M=W$ such that no distinction is necessary, we simply write $\bar{g}_{nT}(\theta, W) = \bar{g}_{nT}(\theta, W, W)$. For a given $m \times m$ moment-weighting matrix $\widehat{\Lambda}_{nT}^{-1}$, we define the sample GMM objective function as
When fixed effects are present, within transformation {\color{red}\footnote{The within transformation eliminates the individual effects $\alpha$ by subtracting unit-specific time averages: writing $Y_{\cdot} \coloneqq \frac{1}{T}\sum_{t=1}^{T} Y_t$, the transformed outcome is $\ddot{Y}_t \coloneqq Y_t - Y_{\cdot}$, and $X_t(W)$ and $Z_t(M)$ are transformed analogously. Since $\alpha$ is constant over $t$, it is differenced out by this demeaning; the time effects $\iota_t$ are removed by the additional cross-sectional demeaning discussed in Remark (ref). The moment conditions (ref) constructed from the demeaned variables remain valid.}} is applied to $Y_t$, $X_t(W)$, and $Z_t(M)$ before constructing the GMM objective; we suppress this transformation in the notation that follows. The extension to two-way fixed effects is discussed in Remark (ref).
Finally, we stack all time series observations into the $nT$-dimensional outcome vector $Y \coloneqq [Y_1^{\top}, \ldots, Y_T^{\top}]^{\top}$, the $nT \times (2K+1)$ design matrix ${X}(W) \coloneqq [{X}_1(W)^{\top}, \ldots, {X}_T(W)^{\top}]^{\top}$, and the $nT \times m$ matrix of instruments ${Z}(M) \coloneqq [{Z}_1(M)^{\top}, \ldots,{Z}_T(M)^{\top}]^{\top}$. Using this notation, we can express the GMM estimator as
We define $\widehat{\theta}_\text{GMM} = \widehat{\theta}_\text{GMM}(W,W)$. In practice, we can obtain a two-step GMM estimator by using the weight matrix $\widehat{\Lambda}_{nT}^{-1}$. {\color{red}\footnote{In practice, we obtain a feasible two-step GMM estimator by first computing $\widetilde{\theta}$ using the identity weighting matrix and then using $\widehat{\Lambda}_{nT}^{-1}$, where \[ \widehat{\Lambda}_{nT} = \frac{1}{nT}\sum_{t=1}^T Z_t(W)^\top \operatorname{diag}\!\left( \widehat{\varepsilon}_{1t}^{\,2},\ldots, \widehat{\varepsilon}_{nt}^{\,2} \right) Z_t(W), \qquad \widehat{\varepsilon}_t = Y_t-X_t(W)\widetilde{\theta}. \] } }representing an initial consistent estimator obtained by using the identity $\widehat{\Lambda}_{nT}^{-1} = I_m$ as the first-step GMM weight matrix.
The rates of our estimators along with all formal assumptions and theoretical results, are presented in Section (ref). If $W_0$ was observed without error, Kelejian1998 and Lee2007, for example, guarantee that the standard GMM estimator $\widehat{\theta}_\text{GMM}$ is consistent for $\theta_0$ in (ref). However, when measurement error is present, small elementwise noises in $E$ can accumulate, leading to inconsistent estimation of the regression coefficients and spillover effects.
Therefore, when measurement error exists, we need a way to correct for this additional noise in our estimation. We make use of the information afforded by the low-rank plus sparse plus error decomposition as in (ref) to obtain an estimate $\widehat W$. Theoretically, we quantify the estimation accuracy of $\widehat W$ and its asymptotic influence when used as an input for estimating $\theta_0$. We now describe the details leading to our two proposed estimators. Our first option is the “plug-in estimator,” and the other is the “supervised estimator,” both defined in the following subsection.
Our proposed estimators make use of the structure for the adjacency $W$ in (ref) by using penalized convex low-rank and sparse estimation. Our first proposal is simply a plug-in estimator that first produces a denoised estimate for $W_0$ and uses this estimate in place of the observed version when calculating the GMM estimator. Recall that $A$ is an $n\times n$ matrix and $A_{ij}$ denotes its $(i,j)$th element. We first minimize the following loss function to de-noise the adjacency matrix { by} taking advantage of its low-rank plus sparse structure:
where $\|A\|_F$ is the Frobenius norm of $A$, $\|L\|_*$ is the nuclear norm of $L$, $\nu_n$ and $\tau_n$ are positive tuning parameters. Similar to Candes2011, Cao2017, we propose to use Algorithm (ref) to solve the optimization problem in (ref). Defining $\widehat{W} \coloneqq \widehat{L} + \widehat{S}$, our plug-in estimator is then simply given as in (ref) {to replace} $W$ with $\widehat{W}$.
Our second proposal, the supervised estimator, {minimizes the combination of the GMM objective function and the penalized square loss based on the low-rank plus sparse decomposition}. The estimator $ (\widehat{\theta}_{s}, \widehat{L}_s, \widehat{S}_s) $ is defined as
where $\xi_{nT}$ is an additional nonnegative tuning parameter necessary to balance the {{variations contained in the observed adjacency and the outcome variables}. {It is worth noting that cai2021network also study a supervised method based on a purely low-rank structure for analyzing network centrality. Their approach, which focuses on the leading singular vectors, differs from ours in that it does not incorporate a spatial regression.}
A full procedure for solving (ref) is given in Algorithm (ref) in Appendix (ref). Note that, as a by-product of the algorithm, the estimator produces supervised {low-rank and sparse} decompositions that incorporate information from the outcome and covariates.
In this section, we discuss the theoretical properties of our estimation procedures. Our main interest is to quantify the impact of using an estimated adjacency matrix on recovering both the true adjacency and the spillover parameters of interest. We shall use the following notation: Let $A$ and $B$ be matrices of dimensions $n\times n$. For $i = 1, \ldots, n$, we use $\lambda_i(A)$ and $\sigma_i(A)$ denote the $i$-th largest eigenvalue and the $i$-th largest singular value of the matrix $A$. Let $\lambda_{\min}(A)$ denote the minimum eigenvalue of $A$. Let $\|.\|_2$ denote the Euclidean-norm of a vector. We use $\|A\|_1$, $\|A\|_2$, $\|A\|_{\infty}$, $\|A\|_{\max}$ to denote the $l_1$-norm, $l_2$-norm, $l_{\infty}$ norm, max norm respectively. {$|A|_{1,1} =\sum_{i=1}^n\sum_{j=1}^n |A_{ij}|$.} Let $1\leq r\leq n$ be the rank of $A$. For a matrix $A \in \mathbb{R}^{n \times n}$, let:
We use $e_i$ to denote the standard basis vector, i.e., the $i$th column of the identity matrix, and $\operatorname{vec}(A)$ to denote the column vectorization of matrix $A$. For sequences $\{a_n\}_{n = 1}^{\infty}$ and $\{b_n\}_{n = 1}^{\infty}$, we denote $a_n\lesssim b_n$ if there exists a positive constant $C$ such that $a_n/b_n\le C$ for all $n$, and denote $a_n=o(b_n)$ (resp. $a_n \asymp b_n$) if $a_n/b_n\to 0$ or $a_n \ll b_n$ {(resp. $a_n\lesssim b_n$ and $b_n\lesssim a_n$)}. We use $\overset{d}{\to}$ and ${\to_p}$ to indicate convergence in distribution and in probability, respectively. We use $X_n = O_p(a_n) (\lesssim_p a_n)$ to indicate $X_n/a_n$ is bounded in probability, i.e. for any $\varepsilon > 0$, there exists $M(\varepsilon) > 0$ such that $\mathbb{P}(|X_n/a_n| > M(\varepsilon) ) < \varepsilon$ for all $n \in \mathbb{N}$. We assume $\mbox{min}{(n,T)}\to \infty$ or $n \to \infty$. In what follows, $Y_t$, $X_t$, $Z_t(M)$, and $\varepsilon_t$ denote the within-transformed versions of the model objects, with individual fixed effects removed. The asymptotic results carry through under the two-way transformation that additionally removes time fixed effects; see Remark (ref).
Assumption (ref) collects conditions on $W_0$. Part (i) requires the low-rank signal to be pervasive but individually weak: the condition $\|L_0\|_{\max} = O(1/\sqrt{m_r(L_0)m_c(L_0)})$ allows $L_0$ to be distinguished from $S_0$; it fails when $L_0$ is spiked (as in the dominant-units example), in which case the low-rank part is better absorbed into $S_0$. The rank $r$ may diverge with $n$ provided the relevant rates still vanish. Part (ii) requires $S_0$ to be bounded and Part (iii) allows $S_0$ to {have relatively few nonzero links compared with the dense part.}
Hence the decomposition $W_0 = L_0 + S_0$ is uniquely determined when $n$ is large, making the rank parameters $r$ and $m_s(S_0)$ well-defined. We stress this is a mathematical requirement only: $\widehat\theta_p$ uses $\widehat W = \widehat L + \widehat S$ as a single object, so we rely on no economic interpretation of the components, and inference about $\theta$ is robust to weak identification of the decomposition. Part (iii) holds trivially when $L_0 = \mathbf{0}_{n\times n}$ or $S_0 = \mathbf{0}_{n\times n}$, and as a side implication limits the sum over individual columns of $W_0$, yielding $\|W_0\|_1 = o(\sqrt{n})$, which is used in the rate analysis of $\widehat\theta_p(\widehat W)$. The following assumption imposes conditions on the measurement error $E$ in the observed matrix $W$, required for the large-sample analysis as $n \to \infty$.
Assumption (ref) is mild. It does not require the entries of $E$ to be independent or identically distributed (a useful feature in spatial and social network contexts), where measurement errors often exhibit dependence. It also does not require $w_{2n}$ or $w_{\max n}$ to vanish; these are sequences that bound the spectral and entrywise norms of $E$, and may grow with $n$. Whether the bounds delivered by Theorems (ref)--(ref) are informative depends on joint conditions involving $(w_{2n}, w_{\max n}, r, m_s(S_0), n)$, discussed after each theorem.
The roles of $w_{2n}$ and $w_{\max n}$ are complementary. The spectral norm bound $w_{2n}$ controls the aggregate noise level and governs how well the low-rank component $L_0$ can be separated from $E$. The entrywise maximum $w_{\max n}$ controls the largest individual perturbation and governs the recovery of $S_0$. We illustrate how $E$ looks like in the following examples.
We next show the properties $\widehat L$ and $\widehat S$ corresponding to (ref). Define \[ a_n \coloneqq
\] Define the rates
Let \[ B_0=(I_n-\lambda_0W_0)^{-1}, \qquad \rho_n=w_{2n}+\mathcal R_s(W_0,E), \] \[ s_n=\mathcal R_s(W_0,E)\sqrt{m_s(S_0)}, \qquad \kappa_n=1+\|W_0\|_2+\rho_n, \] and define the dimension-discounted rate \[ \mathcal R_{nT}^* \coloneqq (1+\|B_0\|_2)\kappa_n^{d+1} \frac{(r\vee1)\rho_n+s_n}{n}. \]
The rates $\mathcal{R}_F$, $\mathcal{R}_2$, and $\mathcal{R}_s$ share a common structure with three error sources: a term in $w_{2n}$ about low-rank recovery; a term $m_s(S_0)\,w_{\max n}^2$ related to $m_s(S_0)$ nonzero sparse entries, and a term involving $\|L_0\|_{\max}$. The factor $r$ in $\mathcal{R}_F$ is replaced by unity in $\mathcal{R}_2$, reflecting the rank discount of Theorem 9.24 in wainwright2019high. Both rates parallel Corollary 1 of Agarwal2012 and Theorem 9.19 of wainwright2019high. Note that estimator of \( L_0^* \) and \( S_0^* \) inherits the theoretical properties of the corresponding estimators of \( L_0 \) and \( S_0 \). For instance, since \( \widehat{L}^* = \widehat{L} - \operatorname{diag}(\widehat{L}) \), we have $\| \widehat{L}^* - L_0^* \|_2 = O_p( \| \widehat{L} - L_0 \|_2)$, which implies that estimation accuracy is preserved under the diagonal adjustment.
The conditions required for first-step and second-step consistency differ in an essential way. For the L+S decomposition $(\widehat L, \widehat S)$ to be consistent in operator norm using the bounds in Assumption (ref), i.e., $\|\widehat L - L_0\|_2 \to_p 0$ and $\|\widehat S - S_0\|_2 \to_p 0$, we require \[ w_{2n}^2\to0, \qquad m_s(S_0)w_{\max n}^2\to0, \qquad m_s(S_0)a_n^2\to0. \] in keeping with the standard $L+S$ literature. Frobenius consistency, $\|\widehat L-L_0\|_F\to_p0$ and $\|\widehat S-S_0\|_F\to_p0$, requires the stronger condition $\max\{r,1\}w_{2n}^2\to0$ in place of $w_{2n}^2\to0$, while the other two conditions remain unchanged. Consistency of the spillover estimator $\widehat\theta_p(\widehat W)$, $\|\widehat\theta_p(\widehat W) - \theta_0\|_2 \to_p 0$, requires only weaker dimension-discounted conditions, given in Theorem (ref) below.
Next, define $\widehat W\coloneqq\widehat L+\widehat S$, where $(\widehat L,\widehat S)$ is the estimator in Theorem (ref). We use $\widehat W$ as the working matrix for estimating the model in (ref). We now impose the conditions used to analyze the spillover-parameter estimators.
Assumption (ref) imposes standard conditions on the true parameters. The assumption $\lambda_0 \beta_0 + \gamma_0 \neq \mathbf{0}_{K\times 1}$ is imposed to guarantee the identification of all the parameters in the model (see, e.g., Bramoulle2009). Assumption (ref) (i) requires that the true adjacency matrix \(W_{0}\) have zero diagonal, which rules out self‐loops. Regarding Assumption (ref) (ii), the matrices \(I_{n}\), \(W_{0}\), and \(W_{0}^{2}\) are linearly independent, i.e.\ there do not exist any nontrivial scalars \(\alpha_{0},\alpha_{1},\alpha_{2}\) such that \[ \alpha_{0}I_{n} + \alpha_{1}W_{0} + \alpha_{2}W_{0}^{2} = \mathbf{0}_{n\times n}. \] Assumption (ref)(iii) restricts the number of dominant units to a fixed finite value and imposes uniform stability. The first stability condition guarantees that $I_n-\lambda W_0$ is invertible uniformly over $\lambda\in\mathcal C$.
For each $i$, let $x_{t,i}=(x_{t,i1},\ldots,x_{t,iK})^\top$.
Assumption (ref) and (ref) include standard moment assumptions. The independence \ requirement across $i$ in Assumption (ref) and Assumption (ref)(i) might be strong in network settings, but it can be relaxed to weak cross-sectional dependence with all rates in Theorems (ref) and (ref) unchanged in order. We maintain independence for expositional simplicity. { Assumptions (ref) (ii) are regularity conditions on the instrumental variables. Among other things, the requirement that $\mathbb{E}[\varepsilon_t \mid Z_t(M)] = \mathbf{0}_{n \times 1}$ precludes $M=W$ if the measurement error $E$ is endogenous. Assumptions (ref) (iii) and (iv) impose nonsingularity on the variance covariance matrix of the instrumental variables, and rank conditions of gradient of the moment functions.} The full column rank condition ensures that we only consider linearly independent columns of the instrument matrix, a common practice in the literature Lee2003. Take $M=W_0$, this is linked to the rank conditions of $W_0$, the variance covariance structure of $Z_t(W_0)$ and the correlation between $X_t$ and $Z_t(W_0)$. To further understand the implications of the above assumption, let {$Z_t(W_0)$ be the $n\times m$ matrix}. Note that $(nT)^{-1}\sum_{t=1}^T Z_t(W_0)^{\top} Z_t(W_0)$ is a {partitioned} matrix whose elements have the form $(nT)^{-1}\sum_{t=1}^T X_t^{\top} [W_0^{l_1}]^{\top} W_0^{l_2}X_t,$ for any $l_1$ and $l_2$ between 0 and $d$. This condition can be implied by $\lambda_{\min}(\lim_{n,T\to \infty}\{nT\}^{-1}\sum_t\mathbb{E}\{Z_t^{\top}Z_t\})>0 ,$ and $\lambda_{2K+1}(W_0W^{\top}_0)> 0$. Then $\Sigma_{Z_0Q_0}$ is of full column rank could be implied by a condition that $\lambda_{2K+1}(W_0W_{0}^{\top})>0$, $\lim_{T\to \infty}T^{-1}\sum_t[X_t,W_0X_t,\cdots, W_0^dX_t]$ has column rank greater than or equal to $2K+1$, and all eigenvalues of $I_n - \lambda_0 W_0$ are bounded away from $0$ and $\infty$. Assumption (ref) is a standard assumption imposed on the GMM weight matrix. Because $M$ is fixed and time invariant, $Z_t(M)$ is a measurable function of $X_t$. Therefore, $\{Z_t(M)\}$ inherits the strong-mixing property, with $\alpha_{Z(M),n}(h)\leq\alpha_{X,n}(h)$.
Before establishing the consistency of our spillover-parameter estimator, we first examine the bias introduced by $E$ when using the classical GMM estimator. We provide an intuitive argument illustrating how the measurement error matrix \(E\) biases the moment functions and the importance of its removal, particularly when the instrument matrix \(Z_t(M)\) also depends on \(W\), as in standard spatial regression. For simplicity here, assume no contextual effects are present.When $Z_t(M)$ consists of exogenous instruments satisfying $\mathbb{E}[Z_t(M)^\top \varepsilon_t] = \mathbf{0}_{m \times 1}$, the moment function is given by \[ \bar g_{nT}(\theta, W, M) = \frac{1}{nT} \sum_{t=1}^{T} Z_t(M)^\top (Y_t - \lambda W Y_t - X_t \beta). \] Substituting $W = W_0 + E$, \[ \bar g_{nT}(\theta, W, M) = \frac{1}{nT} \sum_{t=1}^{T} Z_t(M)^\top (Y_t - \lambda W_0 Y_t - X_t \beta) - \frac{\lambda}{nT} \sum_{t=1}^{T} Z_t(M)^\top E Y_t. \]
Thus, if $E$ is independent of $(Z_t(M),Y_t)$ and has mean zero, then \[ \mathbb{E}\!\left[Z_t(M)^\top E Y_t\right] =\mathbf{0}_{m\times 1}, \] and hence \[ \mathbb{E}\!\left[\bar g_{nT}(\theta_0,W,M)\right] =\mathbf{0}_{m\times 1}. \] However, when $Z_t(M)$ is a plug-in version of the true instrument with $M = W$, i.e., \[ Z_t(W) = [X_t, W X_t] = [X_t, (W_0 + E) X_t], \] the moment function becomes \[ \bar g_{nT}(\theta, W) = \frac{1}{nT} \sum_{t=1}^{T} [X_t, (W_0 + E) X_t]^\top \big[(Y_t - \lambda W_0 Y_t - X_t \beta) - \lambda E Y_t\big]. \]
Expanding the moment function $\Bar{g}_{nT}(\theta, W)$ with the instrument matrix $[X_t, (W_0+E)X_t]$, we get: {\[ \left[ \frac{1}{nT} \sum_{t=1}^{T}(-\lambda X_t^{\top} EY_t)^{\top}, \frac{1}{nT} \sum_{t=1}^{T}[X_t^\top E^\top (Y_t-\lambda W_0 Y_t-X_t\beta) - \lambda X_t^\top W_0^\top E Y_t - \lambda X_t^\top E^\top E Y_t]^{\top} \right]^{\top}. \]} Even if \(E_{ij}\)'s are i.i.d. mean zero, with bounded second moment and independent of $(X_t,\varepsilon_t)$, there will be moment condition misspecification. Specifically, \[ \mathbb{E}[\Bar{g}_{nT}(\theta_0, W)] \neq { \mathbf{0}_{m\times 1}}, \] since the expectation of the last term yields: $n^{-1}\mathbb{E}[X_t^\top E^\top E Y_t] = n^{-1} \mathbb{E}\big[X_t^\top \mathbb{E}[E^\top E]\,(I-\lambda_0 W_0)^{-1} X_t\big]\beta_0 \neq \mathbf{0}_{K\times 1},$ which might be non-negligible under our assumptions. Recall that \(\widehat{\theta}_{GMM}\) denotes the efficient two-step GMM estimator based on a given adjacency matrix \(W\), obtained in closed form in (ref). From the preceding results, it follows that \(\widehat{\theta}_{GMM}\) is likely to be inconsistent. We now analyze the convergence rate of \(\widehat{\theta}_p(\widehat{W})\).
Since \(W_0\) is not directly available for constructing instruments, one approach is to use \(Z_t(\widehat{W})\) as plug-in instruments, where \(\widehat{W}\) is obtained from Algorithms (ref) and satisfies Theorem (ref). When \(W_0\) is correctly observed, the minimizer \(\widehat{\theta}_p(W_0)\) satisfies the standard panel-data estimator convergence rate: \[ \|\widehat{\theta}_p(W_0) - \theta_0\|_2 = O_p(1/\sqrt{nT}). \] However, when \(W_0\) is replaced by a denoised estimate \(\widehat{W}\), the associated minimizer \(\widehat{\theta}_p(\widehat{W})\) depends on the deviation \(\widehat{W} - W_0\). The corresponding convergence properties are detailed in the following results.
Under the assumptions of Theorem (ref)(a), the plug-in estimator $\widehat\theta_p(\widehat W)$ is consistent provided that \[ \mathcal R_{nT}^*=o(1). \]
The factor $1/n$ discounts the noise contribution by the network size, so consistency holds even when $w_{2n}$ and $w_{\max n}$ do not vanish individually, for example, when $r$ and $m_s(S_0)$ remain bounded and $w_{2n}$ is of constant order. By contrast, the available bound for the GMM estimator based on the noisy matrix does not guarantee consistency when $w_{2n}$ is bounded away from zero; the moment-misspecification example above shows that inconsistency can occur in this case. The noise-induced component of the plug-in estimator's rate is strictly smaller whenever \[ \mathcal R_{nT}^* =o(w_{2n}+w_{2n}^2). \]
The improvement reflects how the $L+S$ decomposition enters the analysis. Denote $\Delta L := \widehat L - L_0$ and $\Delta S := \widehat S - S_0$ as the recovery errors from Theorem (ref). The recovery error $\widehat W - W_0 = \Delta L + \Delta S$ affects the spillover estimator only through bilinear forms (averages over the network) with the instruments and residual design, and these averages discount the error according to its structure. The low-rank error enters through its nuclear norm, controlled by $\sqrt r\,\|\Delta L\|_F$, contributing $O(r)$ effective directions rather than all $n$; the sparse error enters through $|\Delta S|_{1,1}$, controlled by its $m_s(S_0)$ nonzero entries rather than all $n^2$. The measurement error can therefore be large in aggregate while the spillover estimator still converges.
Example (ref) constructs regimes in which $|E|_{1,1}/n$ does not vanish, while the dimension-discounted condition $\mathcal R_{nT}^*=o(1)$ may still hold. This condition is sharp for sparse measurement error such as link misclassification in $0$--$1$ networks, but fails in the dense case that motivates our analysis: Example (ref) constructs the case in which $|E|_{1,1}$ is of order $n^2$ yet the plug-in estimator remains consistent. The two analyses cover disjoint measurement-error regimes: sparse errors with Lewbeletal2024EJ, dense errors with structured latent adjacency in our framework. When $L_0 = \mathbf{0}_{n \times n}$, the rate specializes to $O_p(m_s(S_0)w_{\max n}/n)$ (Corollary (ref)).
The following corollary specializes Theorem (ref) to the pure-sparse case $L_0=\mathbf 0_{n\times n}$. Within Assumption (ref), only part (ii) on $S_0$ is needed, and the coherence terms drop out. The rate is driven entirely by the sparse recovery error $n^{-1}|\widehat S-S_0|_{1,1}$, the same $\ell_1$-type quantity as the $|E|_{1,1}/n$ condition of Lewbeletal2024EJ, but applied to the recovery error after denoising rather than to the raw measurement error.
The rate in Corollary (ref) follows from Lemma (ref)(b): the sparse recovery error enters the spillover estimator only through $n^{-1}|\widehat S-S_0|_{1,1}$. In the pure-sparse case, the sparse-recovery argument used in the proof of Theorem (ref) gives \[ \frac{1}{n}|\widehat S-S_0|_{1,1} = O_p\!\left( \frac{m_s(S_0)w_{\max n}}{n} \right). \] Although Corollary (ref) covers the pure-sparse case, the gain from our general analysis is most pronounced when the low-rank component $L_0$ is genuinely present and $r$ and $m_s(S_0)$ remain controlled relative to $n$.
{
}
As mentioned previously, we emphasize that our results are also suitable for the case when $E$ is endogenous with respect to $\varepsilon_t$. As an example, in social networks this can arise when misclassification errors are not generated at random, instead being correlated with unobservables driving the network formation process, such as the framework considered by Candelaria2023. To see this, we provide a rate analysis to compare the rate when $E$ is endogenous or not of both our estimator and the GMM estimator. For simplicity, we adopt the following assumption on the relationship between these two sources of error, see similar assumptions as in Lee2014:
Assumption (ref) is imposed together with Assumptions (ref) and (ref). Thus, the disturbances remain marginally i.i.d.\ across $i$ and $t$, while dependence between $E$ and $\varepsilon_t$ is allowed. The Gaussian specification is imposed solely to illustrate the rate in the presence of endogenous measurement error.
Denote $\widehat{\theta}_{GMM}(W,M)$ and $\widehat{\theta}_p(\widehat{W},M)$ the GMM and plug-in estimator with instrument $Z_t(M)$. In comparison to the exogenous measurement error case, the bias of $\widehat{\theta}_{GMM}(W,M) = \widehat{\theta}_p(W,M)$ could be larger than our proposed estimators. It is because the deviation between sample moments $\Bar{g}_{nT}(\theta, W, M)$ and $\Bar{g}_{nT}(\theta, W_0, M)$ might be more severe. This is reflected in our simulation exercises, see Section (ref). For simplicity, we shall illustrate without the contextual-effects regressor $W_0X_t$ (i.e., setting $\gamma_0 = \mathbf{0}_K$), so the regressors are $[W_0Y_t, X_t]$. Define ${G}(M) = -n^{-1}\mathbb{E}[Z^{\top}_t(M)[W_0 Y_t, X_t]]$. To explain this, take $M$ as exogenous; the leading term for the moment estimator $\widehat{\theta}_p(\widehat{W},M) - \theta_0$ is $-(G^{\top}(M)\Lambda^{-1}G(M))^{-1}G^{\top}(M)\Lambda^{-1} \Bar{g}_{nT}(\theta_0, W_0, M)$, according to Theorem (ref). We shall suppress the dependence of $M$ on $W_0$ and $E$. The following proposition characterizes a rate bound for both the GMM estimator and our estimator.
Recall that $\bar g_{nT}(\theta,W,M) = \frac1{nT}\sum_{t=1}^T Z_t(M)^\top \left( Y_t-\lambda WY_t-X_t\beta-WX_t\gamma \right)$.
Three features of this comparison are worth highlighting.
First, the endogeneity bias $n^{1/2-s}\|\rho_{\varepsilon E}\|_2$ appears only in the GMM rate (ref). It originates from a bilinear form that is quadratic in $E$: the correlation between $E$ and $\varepsilon_t$ contributes a non-vanishing mean to $Z_t^\top E B_0\varepsilon_t$, whose magnitude is governed by $\rho_{\varepsilon E}$. In the plug-in case, the analogous bilinear form involves $(\Delta L+\Delta S)B_0\varepsilon_t$. Its low-rank component is controlled by \[ \|\Delta L\|_* = O_p\!\left((r\vee1)(w_{2n}+\mathcal R_s)\right), \] whereas its off-diagonal sparse component is controlled by \[ |\Delta S|_1 = O_p\!\left(\mathcal R_s\sqrt{m_s(S_0)}\right). \] Under the conditional-moment conditions, the plug-in bound therefore contains no additional explicit term involving $\rho_{\varepsilon E}$. Second, even in the absence of endogeneity ($\rho_{\varepsilon E} = 0$), the denoised rate improves upon the GMM rate: the noise-level term $w_{2n}$ in (ref) is replaced by the rate $\mathcal{R}^*_{nT}$ in (ref), which is strictly smaller whenever $\mathcal{R}^*_{nT} = o(w_{2n})$.Third, the denoised rate attenuates the dependence on the noise magnitude $w_{2n}$ rather than eliminating it: $w_{2n}$ enters through the $L+S$ coupling $\|\Delta L\|_2 = O_p(w_{2n}+\mathcal{R}_s)$, but its effect is multiplied by the factor $r/n$ rather than by 1. As a result, the rate scales not with the raw noise level $w_{2n}$ but with the ratio of the structural complexity ($r$, $m_s(S_0)$) to the sample size $n$. This reflects the benefit of our approach. Finally, we consider the rate of the supervised estimator as defined in Equation ((ref)). Denote $\widehat{W}_s = \widehat{L}_s+\widehat{S}_s.$
The supervised estimator achieves the same rate as the plug-in estimator in Theorem (ref) and it is strictly smaller than that of the noisy-$W$ GMM estimator whenever $\mathcal R_{nT}^*=o(w_{2n}+w_{2n}^2)$. The rate still depends on $w_{2n}$, but only through the dimension-discounted recovery rate $\mathcal R_{nT}^*$. It is worth noting that the rate in (ref) depends on the parameters $(r, m_s(S_0), n, T)$ and on the spectral characteristics of the underlying network ($\|B_0\|_2, \|W_0\|_2$), rather than directly on the noise level $w_{2n}$ compared to the rate of the GMM estimator. The condition on $\xi_{nT}$ ensures that the generated error from the first-iteration estimator $\widehat{\theta}^{[1]}$ does not inflate either the sparse recovery in Step b) or the low-rank recovery in Step c) beyond the unsupervised noise level, so that the supervised decomposition matches the plug-in decomposition rates.
The following corollary specializes Theorem (ref) to the pure-sparse case $L_0 = \mathbf{0}_{n\times n}$, paralleling Corollary (ref): only Assumption (ref)(ii) on $S_0$ is needed, the coherence terms drop out, and the rate is driven entirely by the sparse recovery error $n^{-1}|\widehat S_s - S_0|_{1,1}$, making the connection to the misclassification framework of Lewbeletal2024EJ explicit.
The theoretical performance of the supervised estimator is similar to that of the plug-in estimator, as confirmed by the simulation and application results in Sections (ref) and (ref). We shall also note that both Theorems (ref) and (ref) do not require $E$ to be exogenous, as long as $Z_t(M)$ does not directly involve $W$. Furthermore, both theorems yield rates in which the noise level $w_{2n}$ enters only after dimension-discounting by the factor $r/n$ (rather than at the raw spectral rate $w_{2n}$ delivered by the noisy-$W$ baseline), reflecting the fundamental advantage of the $L+S$ decomposition combined with the within transformation.
In this section, we present simulation exercises to evaluate the performance of our proposed estimators. We first introduce four data-generating processes (DGPs) for the network adjacency matrix $W_0$. We then describe how it is contaminated with noise, and discuss the Monte Carlo simulation results.
The diagonal elements of $W_0$ are assumed to be 0. Hence we set all diagonal elements of $E$ to be $0$ and do not penalize $S_0$ on its diagonal. The other entries of $E$ are generated as explained in Section (ref).
For our simulation exercises, we set $\lambda_0 = 0.25$, $\beta_0 = (-1, 2)$. We generate two covariates for $X_t$ independently from a $\mathbb{N}(0, 1)$ and a $\mathbb{N}(5, 2)$, respectively. We set sample sizes $n \in \{40,80, 120\}$ and time dimensions $T \in \{1, 5, 15, 50\}$, considering specific combinations of panel sizes due to computational constraints. We take as given a true network $W_0$ generated using one of the DGPs introduced in the previous subsection, and keep it constant across 100 Monte Carlo simulations.
Our exercises additionally consider the possibility that measurement errors are endogenous. That is, we allow for correlated disturbances between the outcome equation ($\varepsilon_t$) and the errors that contaminate the adjacency matrix $(E)$, as outlined in our Assumption (ref). Let $\sigma_E $ and $\sigma_\varepsilon$ be two positive constants. Specifically, for $i,j = 1, \cdots, n$, we generate $E_{ij} \sim i.i.d. \quad \mathbb{N}(0, \sigma^2_E)$ for $i\neq j$, and set $E_{ii}=0$. We generate $\varepsilon_t$ such that its $i$th component $\varepsilon_{it}$ satisfies $\varepsilon_{it}=\rho \sigma_\varepsilon\sum_{j=1}^n E_{ij}/\sqrt{n} +v_{it}$, and $v_{it}$ are i.i.d. $ \mathbb{N} (0, (1-\rho^2)\sigma_\varepsilon^2)$ that are independent of $E_{ij}$. Note that $\varepsilon_{t}|E$ is normally distributed and we could simply obtain its conditional mean and variance. We define $\mbox{Vec}(E)$ as the vector obtained by stacking the off-diagonal elements of $E$.
where $\Sigma_{\varepsilon E}$ is the $n\times (n^2-n)$ matrix that captures the correlation between $\varepsilon_{it}$ and $ \mbox{Vec}(E)$ stacks all $E_{ij}$ for $i\neq j$ into a long vector. Note that it is implied from the previous sections that $\mbox{cov}(\varepsilon_{it}, E_{ij})=\rho\sigma_E^2\sigma_\varepsilon/\sqrt{n}$ for $i\neq j$ and $\mbox{cov}(\varepsilon_{it}, E_{kj})=0$ if $k\neq i$. In the simulation, we set $\sigma_\varepsilon=0.15$, $\sigma_E = 0.3/n^{0.7}$ and $\rho=0$ or $0.7$. By setting $\rho = 0$, we recover the exogenous measurement error case from Assumption (ref).
Given values for the adjacency $W_0$, covariates $X_t$, parameters $\theta_0$, and disturbances $\varepsilon_t$, we use the reduced-form representation of model (ref) to generate our outcome $Y_t$ at every time period:
Throughout the simulation we set $W = W_0 + E$, which is the observed noisy version of the adjacency matrix, to construct estimates of parameters $\theta_0$. Once the adjacency matrix is generated, we use it to construct a rough estimate of the variance of the errors and set the tuning parameters for estimation. Specifically, we fix $\xi_{nT} = 1$ and set $\mbox{Int}(W)$ as the interquartile range of $W_{ij}$. Then we choose the sparse penalty as $\tau_n=\tau_{nT} = 2\mbox{Int}(W)\log(n)$ and $\nu_n=\nu_{nT}=\mbox{Int}(W) \sqrt{n}$.
{\scriptsize
}
Tables (ref) - (ref) present the simulation results using the setting outlined in the previous subsections. We highlight several key takeaways in this setup. First, Tables (ref) reports relative root mean squared errors (RMSE) for estimates of the spillover scalar parameter $\lambda_0$, where the benchmark is the standard GMM estimator. Observe that both our plug-in and supervised estimators achieve lower RMSEs than the GMM benchmark uniformly across all DGPs and sample sizes. Moreover, the GMM estimates tend to perform worse in the endogenous case and hence there is more improvement in the endogenous case. Our proposed methods perform better than GMM by approximately 50--80% for small $n$ and $T$. Our methods achieve at least a 75% improvement in RMSE for $n=120$, $T=50$. The relative performance of our methods generally improves as the number of nodes in the network $n$ and the number of time periods $T$ increase. This behavior is as predicted by the theory in the cases where the adjacency matrix is assumed to be constant across time.
In terms of the recovery performance of the denoised adjacency matrix, as evidenced in Table (ref), the plug-in method and the supervised method perform similarly. Moreover, there is little difference between the exogenous and endogenous cases because we use the same observed adjacency matrix.
Finally, we explore how our methods perform in recovering not just the scalar spillover effect $\lambda_0$, but the full diffusion multiplier matrix given by $(I_n - \lambda_0 W_0)^{-1}$. This quantity is highly important as a policy tool, as it summarizes the cumulative spillover effects of a given economic shock. As shown in Table (ref), by providing reliable estimates for both $\lambda_0$ and $W_0$, our proposed methods achieve better performance in terms of Frobenius norm across DGPs and sample sizes. Although the estimates of $W_0$ might be similar in both the exogenous and the endogenous cases, our method could obtain more reliable estimates of $\lambda_0$ as well as better recovery accuracy in the endogenous case.
This section illustrates the use of our estimators for social/spatial interactions. We consider two examples, applying our methods and examining the spillover effects. In each example, we first construct a raw spatial weight matrix and then apply Algorithms (ref) and (ref) described in the Appendix to obtain our estimates. To select the penalties, we use the interquartile range (IQR) of the raw weight-matrix entries as a robust scale estimate, denoted by $\widehat\sigma$. The sparse and low-rank penalties are parameterized as $C_S\widehat\sigma\log(n)$ and $C_L\widehat\sigma\sqrt{n}$, respectively, with the constants chosen by the Bayesian Information Criterion (BIC). We estimate the spillover coefficient $\widehat\lambda$ and refer the column sum of the Leontief inverse $(I_n-\widehat\lambda\widehat W)^{-1}$ as the total network influence.
In our first example, we examine an application to annual Gross Domestic Product (GDP) growth as our outcome of interest. Following MRW1992, the explanatory variables include the annual growth rate of the working-age population defined as individuals aged 15 to 64, denoted as \(X_1\), and the log investment share, \(X_2\), which serves as a proxy for the saving rate. We also include individual and time fixed effects in our specifications. The dataset is downloaded from the World Development Indicators (WDI) and covers a balanced panel of 23 economies from 1970 to 2023. \footnote{These economies are: Australia (AU), Canada (CA), Mainland China (CN), Denmark (DK), Finland (FI), France (FR), Germany (DE), Hong Kong SAR China (HK), Indonesia (ID), Ireland (IE), Italy (IT), Japan (JP), Korea (KR), Malaysia (MY), Mexico (MX), Netherlands (NL), New Zealand (NZ), Singapore (SG), Spain (ES), Sweden (SE), Thailand (TH), United Kingdom (GB), and United States (US).}
Economic activity is well known to exhibit substantial spatial dependence Krugman1991. Neighboring economies tend to trade more, share regional demand and supply shocks, and compete for mobile factors of production. For this reason, we use a geographic‑proximity weight matrix as a natural and predetermined benchmark to study spillovers. The raw spatial weight matrix is constructed as the row-normalized inverse squared distance matrix based on Haversine distances and denoted by $W_2$. \footnote{ For economies \(i\neq j\), we compute the Haversine great-circle distance \(D_{ij}\) between capital cities (in millions of meters) and then form the raw weight matrix by setting the $(i,j)$ entry as $1/D_{ij}^{2}$ for $i\neq j$, and the diagonal entries as 0, and then apply row-normalization. } Note that row-normalization on the weight matrix is standard in the spatial econometrics literature, as it facilitates a clear interpretation of the spatial autoregressive coefficient as the strength of the average influence. All off-diagonal entries of $W_2$ are strictly positive, which contrasts sharply with the sparsity assumptions sometimes invoked in the literature. This dense structure makes $W_2$ a natural candidate for our estimating approach, since many of the small off-diagonal weights likely reflect weak, noise‑contaminated links rather than economically meaningful connections.
We estimate Model ((ref)) with two-way fixed effects under two scenarios. Panel A includes only \(X_1\) and \(X_2\) as covariates. We assume that those variables are contemporaneously exogenous and use the spatial lags \(WX_1\) and \(WX_2\) as instruments. Panel B includes both the original covariates and the contextual variables \(WX_1\) and \(WX_2\) as regressors, using the second-order spatial lags \(W^2X_1\) and \(W^2X_2\) as instruments. Table (ref) reports the deviations of the estimated weight matrices from the raw matrix $W_2$. We consider three alternative specifications: a purely sparse structure with $\tau=0.0443$, a purely low-rank structure with $\nu=0.2709$, and a combination of both with $\tau=0.0443$ and $\nu=0.2709$, where the penalties are selected according to BIC respectively. We first examine the differences between estimated adjacency matrices and the raw matrix $W_2$. We find that the purely low-rank estimator $\widehat L$ exhibits substantial departure from $W_2$, particularly in terms of the matrix maximum norm. Most of the elements in $\widehat L$ are small but not exactly 0, suggesting that a low-rank approximation alone may fail to preserve several important links in the original network. By contrast, the purely sparse estimator $\widehat S$ retains only 96 nonzero entries. Although this produces a more parsimonious network representation, it has the largest deviation from $W_2$ under $l_2$ norm. The low-rank-plus-sparse estimators provide a more balanced representation and both $\|\widehat{W}_2-W_2\|_{\max}$ and $\|\widehat{W}_2-W_2\|_{2}$ remain small relative to the corresponding deviations of the competitors. A further calculation shows that $\|\widehat{W}_2-\widehat S\|_{\max}=0.02$ and $\|\widehat{W}_2-\widehat S\|_{2}=0.11$, indicating that it may capture diffuse connections that may be discarded by pure sparsification. Notably, the supervised low-rank and sparse estimates based on two panel specifications remain very similar to the unsupervised one, which is consistent with our theoretical findings. A further comparison of the heat maps of the raw $W_{2}$, the estimated $\widehat{W}_{2}$, and $\widehat{S}$ is also shown in Figure (ref). We observe that the low-rank plus sparse estimate $\widehat{W}_{2}$ appears to preserve the dominant network structure while allowing for pervasive weak link effects.
In the regression analysis, we evaluate the impact of the investment rate and working-age population growth, and report the estimated spillover coefficient $\widehat{\lambda}$ based on different weight matrices in Table (ref). For each candidate weight matrix $W$, we also compute the column sums of the Leontief inverse $(I_n-\widehat{\lambda}\widehat W)^{-1}$ and designate the economy with the largest column sum as the key player. Relative changes compared to the raw $W_2$ benchmark are also reported.
In Panel A (without contextual effects), the purely sparse estimator $\widehat{S}_2$ only retains the strong connections and reduces $\widehat{\lambda}$ by 16.7% and total influence (i.e., the corresponding column sum from the Leontief matrix, see above) by 11.3% compared to those obtained from the raw $W_2$. Conversely, the purely low-rank estimator $\widehat{L}_2$, amplifies network connections via aggregating pervasive weak links and increases $\widehat{\lambda}$ by 11.3% and total influence by 12.0% compared to those obtained from the raw $W_2$. These results indicate that the sparse and low-rank components might capture distinct aspects of the network. The sparse component filters out scattered weak links, whereas the low-rank component summarizes pervasive dependence. Correspondingly, the low-rank-plus-sparse estimator $\widehat{W}_2$ produced more moderate deviations from the raw benchmark.
The benefits of our approach become more pronounced in Panel B with contextual effects. The raw $W_2$ produces a negative estimate $\widehat{\lambda}$ that is much lower than the our estimates. The sparse estimator and the low-rank-plus-sparse estimator produce moderate positive values that might be more reasonable. As higher-order spatial lags can magnify noise in the raw matrix, this reversal suggests that our approach can improve the plausibility and stability of the estimated spillover effects.
Moreover, the decomposition also yields interesting practical insights. Compared with $\widehat S_2$, we find that including the low-rank component in $\widehat W_2$ primarily rescales the overall magnitude of the spillover influence vector rather than altering the key player. This pattern indicates that, while the identity of the key player remains robust to the decomposition strategy, the estimated strength of network spillovers depends materially on whether the low-rank component is retained, underscoring the practical relevance of the full low-rank plus sparse decomposition over the purely sparse estimator. Figure (ref) presents the graphical representations of $\widehat{W}_{2}$, $\widehat{L}_{2}$, and $\widehat{S}_{2}$, respectively. $\widehat W_2$ admits the decomposition $\widehat W_2=\widehat S_{\widehat W_2}+\widehat L_{\widehat W_2}$ and comparing the leading eigenvectors of $\widehat W_2$, $\widehat L_{\widehat W_2}$ and $\widehat S_{\widehat W_2}$, we find that ${\bf v}_{\widehat{W}_2}\approx 0.56\,{\bf v}_{\widehat{L}_{\widehat W_2}}+0.48\,{\bf v}_{\widehat{S}_{\widehat W_2}}$ from a regression perspective as discussed in equation ((ref)). Both the low-rank and the sparse components contribute substantially in this application, indicating that the identification of influential nodes depends not only on a few strong connections, but also on the cumulative effect of widespread weak links. Notably, as shown in Figure (ref), ${\bf v}_{\widehat{S}_{\widehat W_2}}$ has many zero entries, and its nonzero entries concentrate on the Asian economies, whereas ${\bf v}_{\widehat{L}_{\widehat W_2}}$ loads on all economies but assigns slightly smaller weights to Western economies. One possible interpretation is that the sparse eigenvector captures localized regional clustering centered on Asia, while the low-rank eigenvector reflects a global common factor that is slightly attenuated for Western economies in accounting for the broader geographic diffusion inherent in the low-rank structure.
In our second example, we apply our methods to examine tax competition among the 48 contiguous U.S.\ states. This application was originally studied by Besley1995 using data from 1962 to 1988, with a pre‑specified geographic weight matrix $W_g$ defined as the row‑normalized 0‑1 neighborhood adjacency matrix setting 1 to state pairs that are geographically adjacent and zero, otherwise. Besley1995 emphasize a political‑economy channel for tax‑setting spillovers, as policymakers respond to voters who benchmark their policies against those of peer states. Consequently, states may face both political and economic incentives to align their fiscal policies with those of their neighbors. This dataset has recently been extended to cover $1962$--$2015$ ($T = 53$) and analyzed by Paula2023. We obtain the spatial weight matrix $W_2$, which is a row-normalized inverse squared distance matrix based on Haversine distances between state capitals, as the observed weight matrix. In addition to endogenous spatial effects captured by $WY$, the covariates include state-level characteristics such as per-capita income, the unemployment rate, and the proportions of young and elderly populations, without and with their contextual covariates $WX$ in Panel A and B respectively. We also include individual and time fixed effects in our specifications. Lagged neighbor income and lagged neighbor unemployment are used as instruments as in the original paper by Besley1995.
We consider three specifications, which yield a purely sparse matrix estimate $\widehat{S}_2$, a purely low‑rank matrix estimate $\widehat{L}_2$, and a low‑rank plus sparse estimate $\widehat{W}_2$. The purely sparse estimate $\widehat{S}_2$ captures the strong links and retains 90 nonzero entries. The estimate $\widehat{L}_2$ is a zero matrix under the selected penalty, suggesting that the data do not provide enough evidence of a low-rank component. The combined estimator $\widehat{W}_2$ coincides with the sparse‑only estimator $\widehat{S}_2$, and hence we only report results about $\widehat{W}_2$ in the subsequent analysis. The supervised variants $W_s$ in Panels A and B are very close to those of the unsupervised estimate $\widehat{W}_2$.
Figures (ref) and (ref) provide heatmaps and graphical representations of the raw geographic weight matrices $W_g$, $W_2$ and the estimate $\widehat W_2$, respectively. We find that the raw inverse geographic distance matrix $W_2$ is much denser than the raw geographic adjacency matrix $W_g$. It may be more vulnerable to measurement errors, and thus exaggerate the overall degree of network dependence and obscure the dominant transmission channels. By contrast, the estimated network discards many weak, noise‑dominated connections, thus providing a clearer visualization of the significant links and improving interpretability.
Table (ref) provides a more detailed characterization of the state-level network structure. States are sorted according to their column sums, which measure the aggregate strength of their outward connections. Massachusetts is identified as the most influential state under all three weight matrices. Its out-degree is 1.70 under the row-normalized adjacency matrix $W_g$, 1.61 under the row normalized inverse squared distance matrix $W_2$, and 0.72 under the estimate $\widehat W_2$. The estimated network therefore preserves the leading position of Massachusetts while substantially reducing the magnitude of its estimated outward influence. This suggests that our estimate does not fundamentally alter the ranking of the major network hubs, but removes potentially noisy links that may inflate the overall strength of network connections rather than incorporating them into the core neighborhood topology. By contrast, the ranking based on $W_g$ differs more noticeably because the geographical adjacency matrix only captures direct neighboring relationships, whereas $W_2$ allows for distance-decaying connections between all states.
Table (ref) reports the estimated spillover coefficient $\widehat{\lambda}$ and the regional influence analysis using different weight matrices. We consider two scenarios, Panel A (without $WX$) and Panel B (with $WX$), respectively. In both panels, two raw weight matrices $W_g$ and $W_2$ produce quite similar results, and they both exhibit much larger or even explosive values of $\widehat{\lambda}$ in Panel B. Such large values are likely driven by measurement error or noise in the raw adjacency matrices, yielding exaggerated spillover estimates that could be unreliable and potentially misleading. In contrast, the estimators, including the low-rank-plus-sparse estimator $\widehat{W}$, and its supervised variants, produce substantially smaller and more stable $\widehat{\lambda}$ values, and consistently identify the South Atlantic as the key region. Recall that each weight matrix is row-normalized prior to estimation. By the Perron–Frobenius theorem, the spectral radius of a row-normalized non‑negative matrix is 1, and hence the standard Leontief-stability condition reduces to $|\widehat{\lambda}|<1$ LeSage2009. Explosive spillover estimates can cause some elements of the Leontief inverse to become negative and hard to interpret. These results indicate that our approach might yield more plausible and interpretable spillover effects, avoiding the inflated influence suggested by raw network matrices.
Table (ref) further examines the propagation of shocks originating from Massachusetts. We calculate the Leontief Inverse and examine the four largest elements in the column that corresponding to Massachusetts. The two states most strongly influenced by a shock originating from Massachusetts are Massachusetts itself and Rhode Island. However, the magnitude of the estimated network influence is attenuated. When contextual effects $WX$ are incorporated, the Leontief influence values increase substantially. Under the estimate $\widehat W_2$, the influence measure for Massachusetts rises from 1.04 to 2.51, while the corresponding values for Rhode Island, New Hampshire, and Maine increase to 1.88, 1.82, and 1.48, respectively. Overall, the results consistently identify Massachusetts as an important regional hub and indicate that contextual effects could considerably amplify the propagation of state-level shocks. Our approach retains stronger local transmission channels while suppressing weaker long-range associations that may be sensitive to measurement error.
This paper introduces a robust framework for studying interaction effects in social and spatial networks under noisy adjacency matrices. By leveraging the low-rank and sparse structure of real-world networks, our de-noising approach, combining LASSO and nuclear norm penalization, significantly improves estimation accuracy. We propose two estimation procedures—a two-step and a one-step supervised GMM estimator—that outperform standard GMM by $50-80\%$ in RMSE under noise and endogeneity.
Applying our proposed methodology to examine the international spillover of economic growth and the tax competition across U.S. states, we find pronounced discrepancy in estimating the spillover effects, which highlights the essential role of accurate network denoising. Failing to distinguish genuine connectivity from measurement noise can systematically distort spillover coefficients as well as the strength and direction of cross-unit interactions. Moreover, our decomposition framework might provide a new perspective on complex spatial dependencies by disentangling global and local components. This decomposition not only sharpens the economic interpretation but also improve econometric inference, thereby contributing to credible spatial econometric analysis.
During the preparation of this work the authors used AI to suggest edits to some lengthy passages for concision and clarity. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.