EconBase
← Back to paper

Estimating Network Spillovers under Dense Measurement Error

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

144,382 characters · 27 sections · 64 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimating Network Spillovers under Dense Measurement Error

abstractThis paper analyzes spillover effects in spatial (network) models when the neighborhood (adjacency) matrix is contaminated by measurement error from reporting, aggregation, or disclosure imperfections, leading to inconsistent estimation of network effects. We introduce a regularization framework for the latent network that allows for sparse and/or low-rank structure and accommodates potential correlation between measurement errors and outcomes. We propose two estimators: (i) a two-stage procedure that first denoises the adjacency matrix and then incorporates the purified network into a regression analysis, and (ii) a Generalized Method of Moments (GMM) estimator that jointly estimates regression parameters and refines the network structure. We then establish strictly improved consistency rates for the spillover effect estimator relative to naive estimation ignoring measurement error. Simulations demonstrate that, in the presence of noisy networks, our approach reduces the root mean squared error of spillover estimates relative to conventional methods by approximately $50-80\%$. We apply our framework to examine the international spillover of economic growth, and the tax competition across U.S. states, illustrating that denoising might restore Leontief stability and yields improved estimates of spillovers.

{JEL Classification:} C21, C23, D57

{Keywords:} network analysis, measurement error, LASSO, nuclear norm penalty, penalized GMM

{

Introduction

Network analysis provides powerful tools to quantify policy spillovers and identify pivotal agents for targeted interventions, particularly when policymakers face binding resource constraints Ballester2006, Zenou2016, Galeotti2020. A large literature shows that economic outcomes often depend on interactions among interconnected units, giving rise to amplification, diffusion, and multiplier effects that are central to policy design Acemoglu2012, Giorgi2020, Araujo2023. These insights have motivated the widespread use of spatial and social interaction models to study peer effects, diffusion dynamics, and network-based policy interventions across a broad range of economic settings.

Methodologically, the literature on spatial and social interaction models differs mainly in how the network structure is handled. A dominant strand assumes that the adjacency matrix is correctly specified and directly observed, and studies spillover effects conditional on this object Lee2007a, Bramoulle2009, Lee2010a, Banerjee2013, Elhorst2014, Lee2014, Yang2017, Zhu2020a, Kuersteiner2020,\\ Shin2021, chernozhukov2021uniform. An alternative and growing line of research relaxes this assumption by estimating the network jointly with model parameters, typically imposing dimension reduction constraints such as sparsity to address the curse of dimensionality in high-dimensional settings Manresa2016, Lam2020, Paula2023.

Sparsity-based methods work well when the latent network is dominated by a few strong links, but they fail in the dense networks that arise in production, trade, and financial settings, where the observed adjacency matrix is itself contaminated by measurement error FismanWei2004, Upper2011, Anand2015. In such settings, two distinct forms of latent structure coexist beyond idiosyncratic noise. First, outcomes may be driven by a small number of pervasive (though individually weak) latent forces such as common shocks, shared technologies, or institutional linkages, naturally represented by a low-rank component Chudik2013, Kapetanios2021. Second, networks may contain a small number of strong bilateral or idiosyncratic interactions, naturally represented by a sparse component Lam2020, Paula2023. Imposing sparsity alone discards the pervasive signal, while ignoring measurement error produces non-vanishing bias as elementwise small errors accumulate along network paths Lewbeletal2024EJ. The stakes are practical: in our cross-state tax-competition application, the standard spatial GMM estimator with the raw inverse-distance weight matrix yields $\widehat\lambda > 1$, violating the Leontief stability condition and implying explosive multipliers, while our estimator returns $\widehat\lambda < 1$ within the admissible region. We therefore propose a regularization approach for the observed adjacency matrix that filters measurement error while preserving low-rank or sparse signal, developed formally in Section (ref).

Throughout, by regularization we mean imposing restrictions: low-rank, sparse, or low-rank-plus-sparse, directly on the latent adjacency matrix $W_0$ rather than on the regression coefficients, as identifying restrictions that make denoising of the observed network feasible.

The problem is non-trivial because it combines matrix-valued measurement error with a spillover estimator that depends nonlinearly on the network matrix through the Leontief inverse $(I_n - \lambda W)^{-1}$, so errors in $W$ propagate through the estimator in ways classical errors-in-variables corrections do not handle. Low-rank and sparse structures have deep foundations in statistics and econometrics, appearing in approximate factor models, latent-variable graphical models, and high-dimensional covariance estimation Bai2002, Chandrasekaran2010, Fan2011. A large methodological literature develops penalized procedures to recover low-rank or sparse patterns in covariance and precision matrices Cai2011, Candes2011, Agarwal2012, Fan2018. Despite their success in these settings, regularization of this kind has not, to our knowledge, been adopted in the econometrics literature to address measurement error in observed network adjacency matrices, or to improve the estimation of spillover effects and other policy-relevant network statistics. We build on this literature by developing two complementary estimation approaches. The first is a two-stage procedure that denoises the observed adjacency matrix by separating low-rank or sparse signal from measurement error before estimating spillover effects. The second is a regularized Generalized Method of Moments (GMM) estimator that jointly estimates the network and the model parameters in a single step. We establish that our estimator is consistent, in contrast to naive estimation that ignores measurement error and remains inconsistent in dense regimes. Simulation results show that the proposed method achieves a $50\text{--}80\%$ reduction in the root mean squared error of spillover estimates under noisy weight matrices, with measurement noise ranging from moderate to high.

Our main theoretical contribution is to characterize how the bias of standard spillover estimators depends on the structure of the latent network and the measurement error, and to show that this bias does not vanish in the dense regime that characterizes economic networks. Lewbeletal2024EJ show that modest link misclassification bias in the estimation of linear social-interaction models may be negligible as long as misclassification is not pervasive; we extend their analysis to settings in which the latent network has many but individually weak connections, and in which the measurement error itself need not be sparse. We further allow the measurement error to be correlated with unobservables in the outcome equation Candelaria2023, lewbel2024estimating, which can introduce non-vanishing bias if ignored. We derive convergence rates for both the spillover-parameter estimator and our estimated adjacency matrix. Beyond these formal results, the framework also delivers an interpretable decomposition of network connections (a low-rank component capturing pervasive relationships and a sparse component isolating localized interactions, with identification following standard strategies in the literature chandrasekaran2011rank) and we use this decomposition to provide a novel expression of the network eigencentrality that quantifies the contributions of each component.

We illustrate the practical relevance of our framework through two empirical applications: international spillovers in GDP growth across 23 economies and tax competition among 48 U.S.\ states. In both settings, our estimated network produces spillover estimates that differ markedly from those obtained using the raw adjacency matrix, highlighting the empirical importance of network denoising. In the cross-economy GDP-growth application, regularization affects not only the scale of spillovers but also the inferred structure of interdependence. The network decomposition further reveals distinct regional and global components of cross-economy interdependence. These empirical findings demonstrate how our framework can uncover the structure of interdependence in economic networks and support more reliable inference in spatial econometric analysis. In the U.S.\ tax-competition application, the raw inverse-distance weight matrices yield explosive estimates $\widehat\lambda > 1$ that violate the Leontief stability condition, while our estimator returns stable estimates.

The remainder of our paper is organized as follows. Section (ref) motivates the framework by explaining why measurement error matters in dense networks, illustrating four possible network structures (sparse, low-rank, low-rank-plus-sparse), and demonstrating with the U.S.\ Input-Output Table that real-world dense networks exhibit exploitable regularization structure. Section (ref) presents the spatial interaction model and introduces the plug-in and regularized GMM estimators under low-rank, sparse, or low-rank-plus-sparse structure. Section (ref) provides assumptions and establishes consistency and asymptotic normality results for both estimators. Section (ref) presents simulation evidence. Section (ref) applies our methods to estimate spillover effects using noisy network data. Section (ref) concludes. Proofs, algorithms, and additional technical details are provided in the appendices. }

Motivation

This section sets up the notation and previews the regularization framework used throughout the paper. Let $\mathcal{N} \coloneqq \{1, \ldots, n\}$ denote the set of $n$ units or nodes and $\mathcal{E}$ the set of edges between them, so that $(\mathcal{N}, \mathcal{E})$ represents the graph. Let $W_0$ be the $n \times n$ adjacency matrix encoding the true connections, with $(i,j)$-th element $W_{0,ij}$; we write $W_{0,ij} \neq 0$ when unit $i$ is connected to unit $j$ and $W_{0,ij} = 0$ otherwise.\footnote{We allow the connection from $i$ to $j$ to differ from that from $j$ to $i$, so the graph may be directed and $W_0$ need not be symmetric.} We further assume no self-loops, so that $W_{0,ii} = 0$ for all $i = 1, \ldots, n$.

The econometrician observes a noisy version of $W_0$, \[ W = W_0 + E, \] where $E$ is a measurement-error matrix whose sources are discussed in Section (ref). Before formally motivating the regularization framework, we briefly preview the structural decomposition of $W_0$ that plays a central role throughout the paper.

To reduce estimation complexity for large $n$, it is natural to impose structural assumptions on $W_0$. The most common choice (sparsity) is plausible for social networks dominated by a few strong links, but fails to describe networks in which a large number of units share common interests or characteristics: a highly sparse network typically splits into several disconnected components. We therefore propose to model $W_0$ as the sum of a low-rank and a sparse component,

equation[equation omitted — 101 chars of source]

where $L_0$ captures pervasive, common connections and $S_0$ captures sporadic or idiosyncratic connections.

The components $L_0$ and $S_0$ need not individually satisfy the zero-diagonal constraint. When $W_0$ has a zero diagonal, we can define $L_0^* = L_0 - \operatorname{diag}(L_0)$ and $S_0^* = S_0 + \operatorname{diag}(L_0)$, which absorb the diagonal corrections so that both represent networks without self-loops:

equation[equation omitted — 105 chars of source]

While $L_0^*$ is not exactly low-rank, it inherits the low-rank structure of $L_0$ and differs only on the diagonal. Section (ref) illustrates the usefulness of this decomposition through concrete examples. We further develop our theoretical analysis (Section (ref)) in terms of $L_0$ and $S_0$ and note that the results extend directly to $L_0^*$ and $S_0^*$. Both theoretical and empirical evidence in the following sections show that this decomposition captures several meaningful network structures encountered in practice Diebold2014, Barranca2015, Gamble2016, newman2010networks.

Why dense measurement errors matter

A central feature of the applications mentioned in the introduction is that the observed network is dense and measured with error. We document four such sources of error in Section (ref); here we develop the consequences for spillover estimation. The existing literature, such as Lewbeletal2024EJ, assumes that measurement error in the network is sufficiently sparse that it does not affect the consistency of the spillover estimator. In particular, consistency of $\widehat{\lambda}$ can be preserved when the aggregate magnitude of the error vanishes asymptotically, for example when \[ \frac{1}{n}\sum_{i,j} |E_{ij}| = o_p(1). \] This condition is naturally satisfied when the total error budget grows slowly relative to the size of the network, as in settings where the underlying adjacency matrix itself is sparse. Intuitively, when only a small fraction of links are contaminated and the magnitude of each contamination is small, the resulting perturbation to the network propagation mechanism becomes negligible in large samples.

The situation is fundamentally different when measurement errors are dense. Even when each entry of $E$ is elementwise small, the errors accumulate across the $O(n^2)$ entries of the matrix and across the propagation paths generated by the spatial multiplier. The aggregate distortion may not vanish asymptotically, and the estimator $\widehat{\lambda}_{\mathrm{GMM}}$ may fail to be consistent.

To recover $W_0$ from $W$ in the dense regime, some structural restriction on $W_0$ is needed. Depending on the application, the appropriate structure is:

itemize• Sparse: $W_0 = S_0^*$, where most links are absent. Suitable for social networks. • Low-rank: $W_0 = L_0^*$, where connections arise from a small number of common factors or shared institutional features. The primary case for dense networks such as input--output tables and financial exposure matrices. • Low-rank plus sparse: $W_0 = L_0^* + S_0^*$, the general case combining pervasive common connections with idiosyncratic strong links.

Misspecifying this structure leaves residual bias: assuming $W_0 = S_0^*$ when a non-negligible $L_0^*$ is present discards the pervasive systematic connections that drive spillovers; assuming $W_0 = L_0^*$ when strong idiosyncratic links are present absorbs those links into the residual. Moreover, the low-rank and sparse decomposition introduced above is the device we use to operationalize this regularization. We emphasize that it is a flexible modelling device, not an economic interpretation. The interactions-model estimator $\widehat\theta_p$ depends only on $\widehat W = \widehat L^* + \widehat S^*$, not on the individual components. We thus do not rely on any economic interpretation of $L_0^*$ and $S_0^*$ as separate network objects. Nevertheless, Appendix (ref) discusses the economic interpretation of $L_0^*$ and provides identification conditions for the $L_0^* + S_0^*$ decomposition.

The practical importance of regularizing $W_0$ rests on three observations. First, measurement error in dense networks is non-negligible: it may not vanish with sample size, so standard asymptotic consistency arguments fail. Second, it is endogenous: errors share common sources and are typically correlated with the true network entries, violating the classical measurement-error assumptions that justify standard instrumental-variable corrections. Third, it is economically consequential: in our empirical illustration with the U.S.\ Input--Output Table (Section (ref)), ignoring measurement error entirely distorts the estimated propagation matrix by $9\%$, and misspecifying the structure of the error (treating it as purely sparse) compounds the distortion to $68\%$. A method that filters dense, structured noise while preserving the signal in $W_0$ is therefore not a refinement but a prerequisite for credible inference about network spillovers.

Sources of measurement error in economic networks

The economic literature documents several mechanisms through which the observed adjacency matrix $W$ may depart from the latent network $W_0$. Many policy-relevant production, trade, and reconstructed financial networks contain numerous weak links and can be dense relative to conventional social networks. Consequently, measurement error may affect a large fraction of their entries. A related source of potentially dense measurement error arises when the adjacency matrix is constructed from geographical or economic distance LeSage2009, Getis2004: errors in the underlying distance measurem (arising, for example, from centroid approximations, omitted transport links, or bandwidth choices) can propagate across entire rows or columns of $W$. We discuss several additional sources of measurement error below.

(i) Aggregation and reporting error in input--output networks

A widely studied dense economic network is the input--output table, where the $(i,j)$ entry records the share of commodity $i$ used as an intermediate input in producing commodity $j$. Acemoglu2012 establish that propagation multipliers are highly sensitive to the precise pattern of bilateral input shares across all $n^2$ entries of this matrix. The observed Bureau of Economic Analysis table aggregates transactions across hundreds of heterogeneous firms and products into sector-level flows, and is further subject to reporting inaccuracies, imputation of missing entries, periodic revisions, and reconciliation between the supply and use tables. Given these, the analyst observes a sector-level $W$ that departs from any meaningful sector-level $W_0$, and the resulting errors propagate along multi-step Leontief chains.\footnote{The literature on aggregation in production networks treats this as a separate methodological problem; see e.g.\ barrot2016input and boehm2019input. Our framework is complementary: it addresses reporting and revision error at whatever level the network is observed, and is silent on the firm-to-sector aggregation step.} Our method takes the sector-level network as primitive and denoises around a sector-level truth. We show in Section (ref) that even at the sector level, denoising restores Leontief stability and materially changes estimated cross-sector multipliers.

(ii) Mirror-statistic discrepancies in bilateral trade networks

International trade flows provide a second canonical example of a dense network where measurement error affects every bilateral entry. Trade flows are recorded twice: by the exporting country as exports and by the importing country as imports. These two independent records of the same physical shipment should agree, but in practice they differ substantially and systematically. FismanWei2004 exploit this bilateral mirror-statistic structure to measure the evasion gap, i.e., the log difference between Hong Kong's reported exports to China and China's reported imports from Hong Kong at the six-digit product level and show that a one-percentage-point increase in the tariff rate is associated with a three-percent increase in the evasion gap. This tariff-related underreporting may affect many product categories and can be amplified by product misclassification, as importers reclassify high-tariff goods into lower-tariff categories. The resulting measurement error need not be confined to a small number of trade links and may affect a substantial fraction of the bilateral trade network in a manner correlated with tariff rates. Because tariff schedules are systematically related to trade volumes and revealed comparative advantage, the measurement error $E_{ij}$ tends to correlate with the true $W_{0,ij}$, violating classical errors-in-variables assumptions and biasing estimates of trade-network spillovers.

(iii) Entropy-based reconstruction error in interbank exposure networks

Bilateral interbank exposures form a naturally dense network for studying systemic risk and financial contagion. Each bank's balance sheet records its total interbank lending and borrowing, but the bilateral breakdown, who lent to whom and in what amount, is typically confidential. Regulatory data under the Basel II framework require reporting of exposures only above a threshold of 10% of regulatory capital, so the majority of bilateral links are unobserved Anand2015. Researchers and regulators therefore reconstruct the full $n \times n$ bilateral exposure matrix from the observable marginals (total lending and borrowing of each bank) using maximum-entropy algorithms. As Upper2011 and the systematic evaluation of Anand2015 document, the maximum-entropy solution fills in the matrix as uniformly as possible, implicitly treating the network as if every bank lent to every other bank in proportion to its total activity. This assumption yields a fully dense reconstructed matrix, but one that systematically underestimates concentration and overestimates diversification: in reality, interbank lending is relationship-based and concentrated among a small set of counterparties. The reconstruction error $E = W - W_0$ is may therefore be dense. It affects every entry of the matrix, and is potentially correlated with the true bilateral exposures, since banks with large total lending are disproportionately misrepresented. Stress-testing exercises that rely on the reconstructed network underestimate contagion risk precisely because the entropy assumption smooths away the concentrated exposures that drive cascading failure.

(iv) Recall error and endogeneity in reported network links

Even when network data are collected via direct reports rather than algorithmic reconstruction, the observed adjacency matrix is subject to recall error, misunderstanding, and measurement timing mismatches. Lewbeletal2024EJ study economic network models in which some reported binary links are misclassified: true connections are recorded as absent (false negatives) and absent connections are recorded as present (false positives). They establish that this misclassification introduces new sources of endogeneity beyond the standard simultaneity problem in spatial models: the mismeasured adjacency matrix contaminates the spatial lag $WY$, creating correlation between regressors and structural errors, and invalidating standard instrumental variables based on powers of $W$. The result is that conventional Two-Stage Least Squares (2SLS) estimators that ignore misclassification are inconsistent even when the fraction of misclassified links is small relative to $n$, because in a dense network a small misclassification rate still corresponds to a large absolute number of erroneous entries. Panel data settings compound this problem: if the network is observed at lower frequency than the outcome variable, links created or dissolved between survey waves appear as persistent misclassifications correlated with individual fixed effects that cannot be differenced away lewbel2024estimating.

The four mechanisms above are heterogeneous in origin but share the three features highlighted at the end of Section (ref): the measurement error is dense, correlated across entries through shared sources (aggregation, tariff incentives, reconstruction algorithm, or recall bias), and asymptotically non-negligible. These properties together explain why classical errors-in-variables corrections may be insufficient: they presume sparse or vanishing noise that can be handled by instrumental variable (IV) or measurement-error bias corrections, whereas the dense, correlated errors documented here require a different approach. Imposing a regularization on $W_0$ (as a low-rank matrix, a sparse matrix, or their combination) provides the identifying restrictions needed to disentangle the systematic signal from the noise, motivating the framework developed in the remainder of the paper.

Examples of network structures

To build intuition, we describe four common economic network structures and show how each fits into the low-rank, sparse, or low-rank plus sparse framework. For simplicity, throughout this subsection, we assume the network is observed without measurement error, so $W = W_0$. \footnote{Note that the link matrix or adjacency matrix could be used to describe the connection strength between individuals. If it does not satisfy the bounded row sum assumption, it could be further scaled by its $\ell_1$ norm.}

Complete network

We begin by considering one of the simplest examples: a fully connected network. In this setup, the link matrix $W_0$ is an $n \times n$ matrix of nonzero values, where it is common to set $W_{0, ii} = 0$ to avoid self-loops. The link matrix corresponding to a complete network generated by nonzero constants $\bm{c}=(c_1, \cdots, c_n)^{\top}$ is defined as

equation*[equation* omitted — 268 chars of source]

{where $c_i$'s are positive constants within the bounded support $[c_{\min}, c_{\max}]$, and they could be viewed as individual characteristics that vary across $i$. Let $L_0= \bm{c} \bm{c}^{\top}$, which is a symmetric low-rank matrix with rank $1$. Its singular value is $\|\bm{c}\|_2^2$, and the associated singular vector is $\bm{c} /\|\bm{c}\|_2$. Then we could also express $W_0=L_0+S_0=L^*_0$, where the sparse component $S_0$ is a diagonal matrix whose $(i,i)$th element is $- c_{i}^2$ to guarantee no self-loops in $W_0$. Such a network could be used to capture the common interaction structure driven by individuals or agents who interact through a common market or platform. }

Low-rank network

Motivated by the structure of the latent factor, we consider the symmetric network in which the strength of the connection is determined by $W_{0,ij}=a_i+z_iz_j$ if $i\neq j$, { where $a_{i}$'s are individual characteristics and $z_i$'s are some latent variables. \footnote{In a more general term, the low rank component can be understood as network effects related to a principal agent, see Galeotti2020 . The principal agent is formed by connections with respect to many nodes within the network, and the effects are dense.} Denote $\bm a$ as a vector consisting of $a_i$. In this example, $W_0=L_0+S_0=L_0^*$, where $L_0={\bm a} {\bf 1}_n^\top + {\bm z} {\bm z}^\top$ is a matrix of rank no more than 2 with ${\bm a}=(a_1, \cdots, a_n)$ and ${\bm z}=(z_1, \cdots, z_n)$, and $S_0$ is a diagonal matrix whose $(i,i)$th element is $-L_{0,ii}$ to remove self-loops. Networks of this kind arise in production networks and financial systems, where common shocks or institutional linkages generate pervasive, dense interactions.}

Group network

Another example of interest is the pure group network, where there are two or more communities, potentially of distinct sizes to avoid eigenvalue multiplicity in the link matrix. Moreover, individuals within the same community are assumed to be connected with each other, while individuals do not connect to others outside their community. For example, suppose $n=20$, and $6$ individuals belong to community $1$ and the rest belong to community $2$. We define

equation*[equation* omitted — 146 chars of source]

where $L_1=(c_1, c_2, \cdots, c_6)^{\top}(c_1, c_2, \cdots c_6)$ and $L_2=(c_7, \cdots, c_{20})^{\top}(c_7, \cdots, c_{20})$, and $S_0$ is a diagonal matrix such that $S_{0,ii}=-c_i^2$. Note that $W_0$ still admits the low-rank and sparse decomposition. The low-rank component admits a block structure and its rank equals the number of communities. One of its two singular vectors is proportional to $(0, \cdots, 0, c_7, \cdots, c_{20})$, while the other is proportional to $(c_1, \cdots, c_6, 0, \cdots, 0)$. See Figure (ref) for the network graph and the heatmap of $W_0$. {This setting captures trade blocs, regional banking clusters, and sectoral linkages, and the rank of $L_0$ equals the number of communities. We could also extend it to allow sporadic between-group connections, which are described by few nonzero off-diagonal elements in $S_0$. Then the general model could be written as $W_0=L_0+S_0=L_0^*+S_0^*$, where the low-rank component captures the within-group connections and the sparse component represents the between-group connections.}

Dominant units

An interesting example comes from considering networks containing a set of dominant units also known as key players, see CalvoArmengol2009, Zenou2016. One definition of dominant units used in the network literature identifies them as the set of units whose actions have large and persistent effects on the other agents.\footnote{We follow the definition of dominant units as in Pesaran2020, Pesaran2021, Lee2022. } One way to measure the effect of the unit $j$ on others is to use its weighted out-degree $d_j$, defined simply by the $j$-th column sum of $W_0$: $d_j = \sum_{i = 1}^n W_{0,ij}$. We then call the unit $j$ a dominant unit if its weighted out-degree is large or grows as the sample size does. For example, suppose $n=20$ and there is only one dominant unit, namely the first individual, who is connected to $13$ individuals. We assume that each nondominant unit is also connected to its immediate neighbors. Then the link matrix $W_0$ is specified as below:

equation*[equation* omitted — 641 chars of source]

See Figure (ref) for the network graph and the heat map of $W_0$.

This network is naturally sparse. The number of nonzero entries in each row is bounded, although the column corresponding to the dominant unit may contain a growing number of nonzero entries. While this column can be represented as a sparse rank-one matrix, the full adjacency matrix need not be low-rank because it also contains links among neighboring units. We therefore treat $W_0$ as a sparse matrix. Appendix (ref) provides further discussion of alternative low-rank-plus-sparse representations and their identification.

figure[figure omitted — 762 chars of source]
figure[figure omitted — 809 chars of source]

{

Empirical illustration: the US Input-Output Table

To demonstrate that measurement error in dense economic networks has economically meaningful consequences, and that the $L+S$ decomposition provides an effective device for recovering the underlying true network structure, we revisit the 2002 US Input-Output table from graham2020econometric, compiled by the Bureau of Economic Analysis and studied by Acemoglu2012. It contains an $n \times n$ network of $n=417$ sectors (after dropping rows and columns with zero direct input requirements), where entry $(i,j)$ indicates the share of sector $i$'s output used in sector's $j$ production. This network is dense by construction(nearly all sectors interact) making it a canonical example where sparsity cannot be assumed and measurement error is pervasive. We treat the observed $W$ as a noisy version of the true network $W_0$ and apply our Algorithm (ref) (Appendix (ref)) to recover $W_0$.

To assess the quality of the denoising, we compare three candidate representations of $W_0$, a purely sparse estimate $\widehat{S}$, a purely low-rank estimate $\widehat{L}$, and the combined version $\widehat{W} = \widehat{L}+\widehat{S}$. The purely sparse estimate satisfies $\|\widehat{S}-W\|_{\max}=0.025$ and $\|\widehat{S}-W\|_2=0.470$, indicating that the tiny elements might have non-negligible aggregation effects. The purely low-rank estimate satisfies $\|\widehat{L}-W\|_{\max}=0.999$ and $\|\widehat{L}-W\|_2=1.214$, indicating that the low-rank structure alone might fail to capture occasionally large idiosyncratic links. The combined estimate achieves $\|\widehat{W}-W\|_{\max}=0.024$ and $\|\widehat{W}-W\|_2=0.183$, materially outperforming either alternative. This superiority is further confirmed by eigenvector alignment. The inner product between the leading eigenvectors of $W$ and $\widehat{W}$ is $0.981$, compared with $0.526$ for $\widehat{S}$ and $0.222$ for $\widehat{L}$, indicating that the low-rank and sparse decomposition might best preserve the spectral structure of the observed network. From Figure (ref), we could see that the sparse component $\widehat{S}$ captures large idiosyncratic bilateral links, while the low-rank component $\widehat{L}$ exhibits a distinctive stripe pattern, revealing that pervasive but individually weak linkages aggregate into systematic structure that a purely sparse representation would miss.

To quantify the econometric consequences of measurement error, we compare the Leontief diffusion matrices $(I_n - \lambda \widehat{W})^{-1}$, $(I_n - \lambda W)^{-1}$, and $(I_n - \lambda \widehat{S})^{-1}$ following acemoglu2016networks with $\lambda = 0.7$. Leaving measurement errors uncorrected produces a $9\%$ deviation in the spectral norm relative to our estimators: \[ \frac{\|(I_n-0.7\widehat{W})^{-1}-(I_n-0.7W)^{-1}\|_2} {\|(I_n-0.7\widehat{W})^{-1}\|_2}=9\%. \] Misspecifying the structure of measurement errors by treating it as a purely sparse discarding the low-rank component compounds the distortion substantially, producing a $68\%$ deviation: \[ \frac{\|(I_n-0.7\widehat{W})^{-1}-(I_n-0.7\widehat{S})^{-1}\|_2} {\|(I_n-0.7\widehat{W})^{-1}\|_2}=68\%. \] These results demonstrate that measurement error in dense networks is not merely a theoretical concern: it materially distorts estimated propagation channels. Hence it is essential for reliable inference to correctly specify the error structure accounting for both its sparse idiosyncratic component and its low-rank systematic component. Figure (ref) illustrates this point visually: the stripe pattern present in the full diffusion matrix $(I_n - 0.7\widehat{W})^{-1}$ is eliminated when the low-rank component is omitted, confirming that misspecification of network structure biases spillover estimates in economically meaningful ways.

figure[figure omitted — 662 chars of source]
figure[figure omitted — 641 chars of source]

}

Methodology

We present the general model in which the latent network $W_0 = L_0 + S_0$ admits both a low-rank and a sparse component, with the purely low-rank and purely sparse settings arising as special cases. We emphasize that this decomposition need not be unique or economically interpretable on its own; rather, the joint regularization imposed on $L_0$ and $S_0$ provides sufficient structure to separate $W_0$ from the noise $E$. We provide the modeling definitions and estimation procedures we propose to tackle this problem. We offer additional assumptions and asymptotic results in Section (ref).

Estimation

We formally describe here the social/spatial interaction model of interest. We consider a panel data scenario, but assume the true underlying network does not vary with time. Denote data points as $\{Y_t, X_t\}_{t=1, \cdots, T}$, where $Y_t$ is the $n \times 1$ outcome vector and $X_t$ is the $n \times K$ design matrix of exogenous covariates. We assume that the true model is generated by the unobserved $n \times n$ adjacency matrix $W_0$ encoding the true connection, and that we have a noisy or contaminated observation, such that we can write

eqnarray[eqnarray omitted — 174 chars of source]

where $\lambda_0$ represents the scalar (social / spatial) interaction or spillover effect, $\beta_0 \in \mathbb{R}^K$ are standard slope coefficients, $\gamma_0 \in \mathbb{R}^K$ are contextual effects (spillover effects of the observable characteristics $X$), $\alpha = (\alpha_1, \ldots, \alpha_n)^\top$ is the vector of individual fixed effects, $\iota_t$ is a time fixed effect at period $t$, ${\bf 1}_{n\times 1}$ is an $n \times 1$ vector of ones, and $\varepsilon_t$ is an $n$-dimensional vector of idiosyncratic disturbances, assumed for now to be independent and identically distributed over units $i$ and time $t$. We assume the covariates $X_t$ to be exogenous to the outcome unobservables and network formation process such that $\operatorname{\mathbb{E}}[\varepsilon_{it} \mid X_t, W_0X_t, \cdots, W_0^dX_t] = 0$, for a small integer $d>0$, $i = 1,\cdots, n$ and $t = 1, \cdots, T$. Crucially, we allow for the measurement errors to be endogenous in the sense that $\operatorname{\mathbb{E}}[\varepsilon_{it} \mid E] \neq 0$, which introduces additional bias into conventional estimates when $W_0$ is replaced by $W$ boucher2025estimating, hardy2019estimating, Lewbeletal2024EJ.

Collect all parameters of interest into $\theta_0 \coloneqq (\lambda_0, \beta^{^{\top}}_0, \gamma^{^{\top}}_0)^{^{\top}} \in \mathbb{R}^{2K+1}$. When covariates $X_t$ are exogenous and the adjacency matrix $W_0$ is correctly observed, the literature proposes to select $m$ linearly independent columns from $[X_t, W_0 X_t, \ldots, W_0^d X_t]$, for a small positive integer $d$, as instruments for the regressors $\bar{X}_t(W_0) = [W_0 Y_t, X_t, W_0 X_t]$ in estimating $\theta_0$ Kelejian1998, Lee2003, Lee2007.

The degree $d$ should be chosen such that the number of retained instruments is at least $2K+1$ to satisfy the order condition, i.e. $2K+1 \leq m \leq (d+1)K$ (see Assumption (ref)). For an arbitrarily pre-specified $n \times n$ adjacency matrix $M$, let $Z_t(M)$ be the $n \times m$ matrix that collects the linearly independent columns from $[X_t, M X_t, \ldots, M^d X_t]$ and let ${X}_t(W)$ be the $n \times (2K+1)$ matrix that collects the regressors $[W Y_t, X_t, W X_t]$. Note that ${X}_t(W)$ uses the observed adjacency matrix $W$ as opposed to the true $W_0$ showing up in model (ref).

By defining $\varepsilon_t(\theta, W) \coloneqq Y_t - {X}_t(W) \theta$ and thus $\varepsilon_t =\varepsilon_t(\theta_0, W_0)$ , we obtain the population moment condition $\operatorname{\mathbb{E}}[Z_t^{\top}(M)\varepsilon_{t}] = \mathbf{0}_{m\times 1}$ for any exogenous $M$. Thus we have the empirical moment conditions:

align[align omitted — 155 chars of source]

The distinction made between the matrix used to construct the regressors ${X}_t(W)$ and the instruments $Z_t(M)$ is necessary to define our supervised estimator below. When $M=W$ such that no distinction is necessary, we simply write $\bar{g}_{nT}(\theta, W) = \bar{g}_{nT}(\theta, W, W)$. For a given $m \times m$ moment-weighting matrix $\widehat{\Lambda}_{nT}^{-1}$, we define the sample GMM objective function as

equation*[equation* omitted — 176 chars of source]

When fixed effects are present, within transformation {\color{red}\footnote{The within transformation eliminates the individual effects $\alpha$ by subtracting unit-specific time averages: writing $Y_{\cdot} \coloneqq \frac{1}{T}\sum_{t=1}^{T} Y_t$, the transformed outcome is $\ddot{Y}_t \coloneqq Y_t - Y_{\cdot}$, and $X_t(W)$ and $Z_t(M)$ are transformed analogously. Since $\alpha$ is constant over $t$, it is differenced out by this demeaning; the time effects $\iota_t$ are removed by the additional cross-sectional demeaning discussed in Remark (ref). The moment conditions (ref) constructed from the demeaned variables remain valid.}} is applied to $Y_t$, $X_t(W)$, and $Z_t(M)$ before constructing the GMM objective; we suppress this transformation in the notation that follows. The extension to two-way fixed effects is discussed in Remark (ref).

Finally, we stack all time series observations into the $nT$-dimensional outcome vector $Y \coloneqq [Y_1^{\top}, \ldots, Y_T^{\top}]^{\top}$, the $nT \times (2K+1)$ design matrix ${X}(W) \coloneqq [{X}_1(W)^{\top}, \ldots, {X}_T(W)^{\top}]^{\top}$, and the $nT \times m$ matrix of instruments ${Z}(M) \coloneqq [{Z}_1(M)^{\top}, \ldots,{Z}_T(M)^{\top}]^{\top}$. Using this notation, we can express the GMM estimator as

eqnarray[eqnarray omitted — 334 chars of source]

We define $\widehat{\theta}_\text{GMM} = \widehat{\theta}_\text{GMM}(W,W)$. In practice, we can obtain a two-step GMM estimator by using the weight matrix $\widehat{\Lambda}_{nT}^{-1}$. {\color{red}\footnote{In practice, we obtain a feasible two-step GMM estimator by first computing $\widetilde{\theta}$ using the identity weighting matrix and then using $\widehat{\Lambda}_{nT}^{-1}$, where \[ \widehat{\Lambda}_{nT} = \frac{1}{nT}\sum_{t=1}^T Z_t(W)^\top \operatorname{diag}\!\left( \widehat{\varepsilon}_{1t}^{\,2},\ldots, \widehat{\varepsilon}_{nt}^{\,2} \right) Z_t(W), \qquad \widehat{\varepsilon}_t = Y_t-X_t(W)\widetilde{\theta}. \] } }representing an initial consistent estimator obtained by using the identity $\widehat{\Lambda}_{nT}^{-1} = I_m$ as the first-step GMM weight matrix.

The rates of our estimators along with all formal assumptions and theoretical results, are presented in Section (ref). If $W_0$ was observed without error, Kelejian1998 and Lee2007, for example, guarantee that the standard GMM estimator $\widehat{\theta}_\text{GMM}$ is consistent for $\theta_0$ in (ref). However, when measurement error is present, small elementwise noises in $E$ can accumulate, leading to inconsistent estimation of the regression coefficients and spillover effects.

Therefore, when measurement error exists, we need a way to correct for this additional noise in our estimation. We make use of the information afforded by the low-rank plus sparse plus error decomposition as in (ref) to obtain an estimate $\widehat W$. Theoretically, we quantify the estimation accuracy of $\widehat W$ and its asymptotic influence when used as an input for estimating $\theta_0$. We now describe the details leading to our two proposed estimators. Our first option is the “plug-in estimator,” and the other is the “supervised estimator,” both defined in the following subsection.

The plug-in estimator and the supervised estimator

Our proposed estimators make use of the structure for the adjacency $W$ in (ref) by using penalized convex low-rank and sparse estimation. Our first proposal is simply a plug-in estimator that first produces a denoised estimate for $W_0$ and uses this estimate in place of the observed version when calculating the GMM estimator. Recall that $A$ is an $n\times n$ matrix and $A_{ij}$ denotes its $(i,j)$th element. We first minimize the following loss function to de-noise the adjacency matrix { by} taking advantage of its low-rank plus sparse structure:

align[align omitted — 252 chars of source]

where $\|A\|_F$ is the Frobenius norm of $A$, $\|L\|_*$ is the nuclear norm of $L$, $\nu_n$ and $\tau_n$ are positive tuning parameters. Similar to Candes2011, Cao2017, we propose to use Algorithm (ref) to solve the optimization problem in (ref). Defining $\widehat{W} \coloneqq \widehat{L} + \widehat{S}$, our plug-in estimator is then simply given as in (ref) {to replace} $W$ with $\widehat{W}$.

eqnarray[eqnarray omitted — 436 chars of source]

Our second proposal, the supervised estimator, {minimizes the combination of the GMM objective function and the penalized square loss based on the low-rank plus sparse decomposition}. The estimator $ (\widehat{\theta}_{s}, \widehat{L}_s, \widehat{S}_s) $ is defined as

eqnarray[eqnarray omitted — 345 chars of source]

where $\xi_{nT}$ is an additional nonnegative tuning parameter necessary to balance the {{variations contained in the observed adjacency and the outcome variables}. {It is worth noting that cai2021network also study a supervised method based on a purely low-rank structure for analyzing network centrality. Their approach, which focuses on the leading singular vectors, differs from ours in that it does not incorporate a spatial regression.}

A full procedure for solving (ref) is given in Algorithm (ref) in Appendix (ref). Note that, as a by-product of the algorithm, the estimator produces supervised {low-rank and sparse} decompositions that incorporate information from the outcome and covariates.

Asymptotic theory

In this section, we discuss the theoretical properties of our estimation procedures. Our main interest is to quantify the impact of using an estimated adjacency matrix on recovering both the true adjacency and the spillover parameters of interest. We shall use the following notation: Let $A$ and $B$ be matrices of dimensions $n\times n$. For $i = 1, \ldots, n$, we use $\lambda_i(A)$ and $\sigma_i(A)$ denote the $i$-th largest eigenvalue and the $i$-th largest singular value of the matrix $A$. Let $\lambda_{\min}(A)$ denote the minimum eigenvalue of $A$. Let $\|.\|_2$ denote the Euclidean-norm of a vector. We use $\|A\|_1$, $\|A\|_2$, $\|A\|_{\infty}$, $\|A\|_{\max}$ to denote the $l_1$-norm, $l_2$-norm, $l_{\infty}$ norm, max norm respectively. {$|A|_{1,1} =\sum_{i=1}^n\sum_{j=1}^n |A_{ij}|$.} Let $1\leq r\leq n$ be the rank of $A$. For a matrix $A \in \mathbb{R}^{n \times n}$, let:

align*[align* omitted — 259 chars of source]

We use $e_i$ to denote the standard basis vector, i.e., the $i$th column of the identity matrix, and $\operatorname{vec}(A)$ to denote the column vectorization of matrix $A$. For sequences $\{a_n\}_{n = 1}^{\infty}$ and $\{b_n\}_{n = 1}^{\infty}$, we denote $a_n\lesssim b_n$ if there exists a positive constant $C$ such that $a_n/b_n\le C$ for all $n$, and denote $a_n=o(b_n)$ (resp. $a_n \asymp b_n$) if $a_n/b_n\to 0$ or $a_n \ll b_n$ {(resp. $a_n\lesssim b_n$ and $b_n\lesssim a_n$)}. We use $\overset{d}{\to}$ and ${\to_p}$ to indicate convergence in distribution and in probability, respectively. We use $X_n = O_p(a_n) (\lesssim_p a_n)$ to indicate $X_n/a_n$ is bounded in probability, i.e. for any $\varepsilon > 0$, there exists $M(\varepsilon) > 0$ such that $\mathbb{P}(|X_n/a_n| > M(\varepsilon) ) < \varepsilon$ for all $n \in \mathbb{N}$. We assume $\mbox{min}{(n,T)}\to \infty$ or $n \to \infty$. In what follows, $Y_t$, $X_t$, $Z_t(M)$, and $\varepsilon_t$ denote the within-transformed versions of the model objects, with individual fixed effects removed. The asymptotic results carry through under the two-way transformation that additionally removes time fixed effects; see Remark (ref).

assumption[True weight] \begin{enumerate}[(i)] • (Low-rank component) $L_0$ is an $n\times n$ low-rank matrix with $\mathrm{rank}(L_0) = r$, $\min(m_r(L_0), m_c(L_0)) \to \infty$, and $\|L_0\|_{\max} = O(1/\sqrt{m_r(L_0)m_c(L_0)})$. • (Sparse component) $S_0$ is an $n\times n$ sparse matrix with $\|S_0\|_{\max} = O(1)$. • (Joint identification) $\frac{m_s(S_0)}{m_r(L_0)m_c(L_0)} \to 0$ and $\frac{\deg_{\max}(S_0)^2}{\min\{m_r(L_0),\,m_c(L_0)\}} \to 0$. \end{enumerate}

Assumption (ref) collects conditions on $W_0$. Part (i) requires the low-rank signal to be pervasive but individually weak: the condition $\|L_0\|_{\max} = O(1/\sqrt{m_r(L_0)m_c(L_0)})$ allows $L_0$ to be distinguished from $S_0$; it fails when $L_0$ is spiked (as in the dominant-units example), in which case the low-rank part is better absorbed into $S_0$. The rank $r$ may diverge with $n$ provided the relevant rates still vanish. Part (ii) requires $S_0$ to be bounded and Part (iii) allows $S_0$ to {have relatively few nonzero links compared with the dense part.}

Hence the decomposition $W_0 = L_0 + S_0$ is uniquely determined when $n$ is large, making the rank parameters $r$ and $m_s(S_0)$ well-defined. We stress this is a mathematical requirement only: $\widehat\theta_p$ uses $\widehat W = \widehat L + \widehat S$ as a single object, so we rely on no economic interpretation of the components, and inference about $\theta$ is robust to weak identification of the decomposition. Part (iii) holds trivially when $L_0 = \mathbf{0}_{n\times n}$ or $S_0 = \mathbf{0}_{n\times n}$, and as a side implication limits the sum over individual columns of $W_0$, yielding $\|W_0\|_1 = o(\sqrt{n})$, which is used in the rate analysis of $\widehat\theta_p(\widehat W)$. The following assumption imposes conditions on the measurement error $E$ in the observed matrix $W$, required for the large-sample analysis as $n \to \infty$.

assumption[Noise] \begin{enumerate}[(i)] • $E$ is an $n \times n$ noise matrix such that $E_{ii}=0$ for $i=1, \cdots, n$. • $E$ satisfies $\|E\|_{2}\leq w_{2n}$ with probability greater than $1-p_{n}$, where $p_{n}=o(1)$. • $E$ satisfies $\|E\|_{\max}\leq w_{\max n}$ with probability greater than $1-p_{n}$, where $p_{n}=o(1)$. \end{enumerate}

Assumption (ref) is mild. It does not require the entries of $E$ to be independent or identically distributed (a useful feature in spatial and social network contexts), where measurement errors often exhibit dependence. It also does not require $w_{2n}$ or $w_{\max n}$ to vanish; these are sequences that bound the spectral and entrywise norms of $E$, and may grow with $n$. Whether the bounds delivered by Theorems (ref)--(ref) are informative depends on joint conditions involving $(w_{2n}, w_{\max n}, r, m_s(S_0), n)$, discussed after each theorem.

The roles of $w_{2n}$ and $w_{\max n}$ are complementary. The spectral norm bound $w_{2n}$ controls the aggregate noise level and governs how well the low-rank component $L_0$ can be separated from $E$. The entrywise maximum $w_{\max n}$ controls the largest individual perturbation and governs the recovery of $S_0$. We illustrate how $E$ looks like in the following examples.

example[A sample-covariance error matrix] Suppose the error arises from estimating a covariance matrix: let $z_1, \ldots, z_T$ be independent mean-zero Gaussian vectors in $\mathbb{R}^n$ with covariance $W_z$, whose eigenvalues are positive and bounded away from 0 and $\infty$. Let $\widetilde E = T^{-1}\sum_{t=1}^T z_t z_t^\top - W_z$ be the sampling error. To respect the no-self-loop convention we take $E = \widetilde E - \mathrm{diag}(\widetilde E)$, so $E_{ii} = 0$ and Assumption (ref)(i) holds. Since $\|\mathrm{diag}(\widetilde E)\|_2 = \max_i|\widetilde E_{ii}| = O_p(\sqrt{\log n/T})$ is dominated by $\|\widetilde E\|_2$, removing the diagonal leaves the rates unchanged. By Theorem 6.5 of wainwright2019high, taking $\eta=\sqrt{n/T}$ gives $\|E\|_2\leq C\|W_z\|_2\{\sqrt{n/T}+n/T\}=:w_{2n}$ with probability $\geq1-c_1\exp(-c_2n)$; taking $\eta=\sqrt{\log(n)/T}$ and applying a union bound gives $\|E\|_{\max}\leq C(\max_iW_{z,ii})\{\sqrt{\log(n)/T}+\log(n)/T\} =:w_{\max n}$ with probability $\geq1-2n^{-c_3}$. Thus, Assumption (ref) holds with $w_{2n}\asymp\sqrt{n/T}+n/T$ and $w_{\max n}\asymp\sqrt{\log(n)/T}+\log(n)/T$. If $n/T\to0$, these rates simplify to $w_{2n}\asymp\sqrt{n/T}$ and $w_{\max n}\asymp\sqrt{\log(n)/T}$, and both vanish, with the spectral bound dominating the entrywise one --- the aggregate noise exceeds the worst individual entry, as is typical in dense networks. $\blacksquare$
example[Dense Gaussian noise: comparison with Lewbeletal2024EJ] Suppose $E_{ij}$ are i.i.d.\ $\mathbb{N}(0, \sigma^2_{n})$ for $i \neq j$ and $E_{ii} = 0$, with $\sigma_n$ possibly diminishing in $n$. As in Example (ref), applying the operator-norm bound of wainwright2019high and a union bound over the entries gives, with probability approaching $1$, $\|E\|_2 \leq 4\sigma_n\sqrt n =: w_{2n}$ and $\|E\|_{\max} \leq 8\sigma_n\sqrt{\log n} =: w_{\max n}$. Assumption (ref) holds with $w_{\max n} = 8\sigma_n\sqrt{\log n}$ and $w_{\max n}\to 0$ whenever $\sigma_n = o(1/\sqrt{\log n})$. By contrast, $\mathbb{E}|E|_{1,1} \asymp n^2 \sigma_n$. The condition $|E|_{1,1} = o_p(n)$ used by Lewbeletal2024EJ requires $\sigma_n = o(1/n)$. For $1/n \ll \sigma_n \ll 1/\sqrt n$, the spectral-norm condition $w_{2n} \to 0$ ensures the GMM estimator $\widehat\theta_{GMM}(W)$ is consistent (cf. Theorem (ref) b), while their $|E|_{1,1} = o_p(n)$ condition fails. Our estimator $\widehat\theta_p(\widehat W)$ further sharpens the rate via dimension discounting. For $\sigma_n\geq c>0$, the available bound for the GMM estimator does not guarantee consistency. More generally, the plug-in estimator is consistent whenever the dimension-discounted rate condition stated in Theorem (ref)(a) holds. $\blacksquare$

We next show the properties $\widehat L$ and $\widehat S$ corresponding to (ref). Define \[ a_n \coloneqq

cases\{m_r(L_0)m_c(L_0)\}^{-1/2}, & L_0\neq\mathbf 0_{n\times n},\\ 0,&L_0=\mathbf 0_{n\times n}.

\] Define the rates

align*[align* omitted — 224 chars of source]

Let \[ B_0=(I_n-\lambda_0W_0)^{-1}, \qquad \rho_n=w_{2n}+\mathcal R_s(W_0,E), \] \[ s_n=\mathcal R_s(W_0,E)\sqrt{m_s(S_0)}, \qquad \kappa_n=1+\|W_0\|_2+\rho_n, \] and define the dimension-discounted rate \[ \mathcal R_{nT}^* \coloneqq (1+\|B_0\|_2)\kappa_n^{d+1} \frac{(r\vee1)\rho_n+s_n}{n}. \]

theoremSuppose Assumptions (ref)--(ref) hold. Let the tuning parameters satisfy \[ \nu_n=C_\nu w_{2n}, \qquad \tau_n=C_\tau(w_{\max n}+a_n), \] for sufficiently large constants $C_\nu,C_\tau>0$. (a) Frobenius bound. The two-step estimator satisfies \begin{align} \|\widehat L-L_0\|_F^2+\|\widehat S-S_0\|_F^2 &\lesssim_p\mathcal R_F(W_0,E),\\ |\widehat S-S_0|_{1,1} &\lesssim_p \mathcal R_s(W_0,E)\sqrt{m_s(S_0)} +n\{w_{2n}+\mathcal R_s(W_0,E)\}. \end{align} (b) Sharpened spectral-norm bound. The two-step estimator satisfies \[ \|\widehat L-L_0\|_2^2+\|\widehat S-S_0\|_2^2 \lesssim_p \mathcal R_2(W_0,E). \]

The rates $\mathcal{R}_F$, $\mathcal{R}_2$, and $\mathcal{R}_s$ share a common structure with three error sources: a term in $w_{2n}$ about low-rank recovery; a term $m_s(S_0)\,w_{\max n}^2$ related to $m_s(S_0)$ nonzero sparse entries, and a term involving $\|L_0\|_{\max}$. The factor $r$ in $\mathcal{R}_F$ is replaced by unity in $\mathcal{R}_2$, reflecting the rank discount of Theorem 9.24 in wainwright2019high. Both rates parallel Corollary 1 of Agarwal2012 and Theorem 9.19 of wainwright2019high. Note that estimator of \( L_0^* \) and \( S_0^* \) inherits the theoretical properties of the corresponding estimators of \( L_0 \) and \( S_0 \). For instance, since \( \widehat{L}^* = \widehat{L} - \operatorname{diag}(\widehat{L}) \), we have $\| \widehat{L}^* - L_0^* \|_2 = O_p( \| \widehat{L} - L_0 \|_2)$, which implies that estimation accuracy is preserved under the diagonal adjustment.

The conditions required for first-step and second-step consistency differ in an essential way. For the L+S decomposition $(\widehat L, \widehat S)$ to be consistent in operator norm using the bounds in Assumption (ref), i.e., $\|\widehat L - L_0\|_2 \to_p 0$ and $\|\widehat S - S_0\|_2 \to_p 0$, we require \[ w_{2n}^2\to0, \qquad m_s(S_0)w_{\max n}^2\to0, \qquad m_s(S_0)a_n^2\to0. \] in keeping with the standard $L+S$ literature. Frobenius consistency, $\|\widehat L-L_0\|_F\to_p0$ and $\|\widehat S-S_0\|_F\to_p0$, requires the stronger condition $\max\{r,1\}w_{2n}^2\to0$ in place of $w_{2n}^2\to0$, while the other two conditions remain unchanged. Consistency of the spillover estimator $\widehat\theta_p(\widehat W)$, $\|\widehat\theta_p(\widehat W) - \theta_0\|_2 \to_p 0$, requires only weaker dimension-discounted conditions, given in Theorem (ref) below.

Next, define $\widehat W\coloneqq\widehat L+\widehat S$, where $(\widehat L,\widehat S)$ is the estimator in Theorem (ref). We use $\widehat W$ as the working matrix for estimating the model in (ref). We now impose the conditions used to analyze the spillover-parameter estimators.

assumption[True parameters] \begin{enumerate}[(i)] • The parameter space $\Theta = \mathcal{C} \times \mathcal{B}\times {\Gamma}$ is a compact subset of $\mathbb{R} \times {\mathbb{R}^K} \times \mathbb{R}^K$ for some finite dimension $K$. • The true value, denoted as $\theta_0 = (\lambda_0, \beta_0^{\top}, \gamma_0 ^{\top})^{\top}$, lies in the interior of the parameter space $\Theta$. Moreover, $\lambda_0\beta_0+\gamma_0\neq \mathbf{0}_{K\times 1}$. \end{enumerate}
assumption[Adjacency matrix] $W_0$ is a non-stochastic spatial weight matrix satisfying \begin{enumerate}[(i)] • The diagonal elements of $W_0$ are 0, such that $W_{0, ii} = 0$ for all $i = 1, \ldots, n$. • $I_n, W_0, W_0^2$ are linearly independent. • Suppose that $J$ is fixed and that the absolute column sums of $W_0$ are uniformly bounded in $n$, except possibly for those of its first $J$ columns. Let $W_{0,u}$ contain the first $J$ columns of $W_0$ and zeros elsewhere, and let $W_{0,b}=W_0-W_{0,u}$. When $J=0$, set $W_{0,u}=\mathbf 0_{n\times n}$ and $W_{0,b}=W_0$. Assume that there exist constants $c_\infty,c_1>0$, independent of $n$, such that \[ \sup_{\lambda\in\mathcal C} |\lambda|\|W_0\|_\infty\leq1-c_\infty, \qquad \sup_{\lambda\in\mathcal C} |\lambda|\|W_{0,b}\|_1\leq1-c_1. \] \end{enumerate}

Assumption (ref) imposes standard conditions on the true parameters. The assumption $\lambda_0 \beta_0 + \gamma_0 \neq \mathbf{0}_{K\times 1}$ is imposed to guarantee the identification of all the parameters in the model (see, e.g., Bramoulle2009). Assumption (ref) (i) requires that the true adjacency matrix \(W_{0}\) have zero diagonal, which rules out self‐loops. Regarding Assumption (ref) (ii), the matrices \(I_{n}\), \(W_{0}\), and \(W_{0}^{2}\) are linearly independent, i.e.\ there do not exist any nontrivial scalars \(\alpha_{0},\alpha_{1},\alpha_{2}\) such that \[ \alpha_{0}I_{n} + \alpha_{1}W_{0} + \alpha_{2}W_{0}^{2} = \mathbf{0}_{n\times n}. \] Assumption (ref)(iii) restricts the number of dominant units to a fixed finite value and imposes uniform stability. The first stability condition guarantees that $I_n-\lambda W_0$ is invertible uniformly over $\lambda\in\mathcal C$.

assumption[Spatial errors] The primitive disturbances $\varepsilon_{it}$ are independently and identically distributed across $i$ and $t$ with mean $0$, variance $\sigma_0^2>0$ and $ \operatorname{\mathbb{E}}|\varepsilon_{it}|^{2+\delta}<\infty$ for a constant $\delta>2$.

For each $i$, let $x_{t,i}=(x_{t,i1},\ldots,x_{t,iK})^\top$.

assumption[Instrumental variables] \begin{enumerate}[(i)] • The covariates $X_t$ are stochastic $n \times K$ matrices with elements $x_{t,il}$ satisfying $\sup_{n,t,i,\,1\leq l\leq K} \mathbb{E}|x_{t,il}|^{2+\delta}<\infty$ for some $\delta>2$. For each $i$, $\{x_{t,i}:t\in\mathbb Z\}$ is strictly stationary and strongly mixing over $t$. The mixing coefficients are uniform in $n$ and $i$: there exist constants $C<\infty$ and $\rho\in(0,1)$ such that \[ \sup_{n,i}\alpha_{n,i}(h)\leq C\rho^h, \qquad h\geq1. \] The cross-sectional processes $\{x_{t,i}:t\in\mathbb Z\}$ are independent across $i$. Dependence between $E$ and $X_t$ is allowed, subject to the conditional moment restrictions stated in Lemma (ref)(c) in the Appendix. • $M$ is a pre-specified adjacency matrix. The instrumental variable matrix $Z_t(M)$ has columns chosen from $M^{\ell} X_{jt}$ for $\ell = 0, \ldots, d$ and $j = 1, \ldots, K$. We require $\mathbb{E}[\varepsilon_t \mid Z_t(M)] = \mathbf{0}_{n \times 1}$. Suppose $Z_t(M)$ has $m$ linearly independent columns, where $m$ is a fixed constant satisfying $m \geq 2K+1$. \[ \sup_{n,t,i,\,1\leq l\leq m} \mathbb E|z_{t,il}|^{2+\delta}<\infty, \qquad \delta>2. \]$\lim_{\min(n,T) \to \infty} (nT)^{-1} \sum_{t=1}^T \mathbb{E}\bigl[Z_t(M)^\top Z_t(M)\bigr]$ exists and is nonsingular, denoted as $\Sigma_{Z_0 Z_0}$. • Let $Q_{0t} \coloneqq \bigl[W_0(I_n - \lambda_0 W_0)^{-1}(X_t \beta_0 + W_0 X_t \gamma_0),\, X_t,\, W_0 X_t\bigr]$. We assume the $m \times (2K+1)$ matrix $\lim_{\min(n,T) \to \infty} (nT)^{-1} \sum_{t=1}^T \mathbb{E}\bigl[Z_t(M)^\top Q_{0t}\bigr]$ exists, denoted as $G_0$, and is of full column rank. \end{enumerate}
assumption[Moment condition and GMM weight matrix] The GMM weight matrix $\widehat{\Lambda}_{nT}^{-1} \in \mathbb{R}^{m \times m}$ converges in probability to some symmetric positive-definite matrix $\Lambda^{-1} \in \mathbb{R}^{m \times m}$, whose eigenvalues are bounded away from 0 and $\infty$.

Assumption (ref) and (ref) include standard moment assumptions. The independence \ requirement across $i$ in Assumption (ref) and Assumption (ref)(i) might be strong in network settings, but it can be relaxed to weak cross-sectional dependence with all rates in Theorems (ref) and (ref) unchanged in order. We maintain independence for expositional simplicity. { Assumptions (ref) (ii) are regularity conditions on the instrumental variables. Among other things, the requirement that $\mathbb{E}[\varepsilon_t \mid Z_t(M)] = \mathbf{0}_{n \times 1}$ precludes $M=W$ if the measurement error $E$ is endogenous. Assumptions (ref) (iii) and (iv) impose nonsingularity on the variance covariance matrix of the instrumental variables, and rank conditions of gradient of the moment functions.} The full column rank condition ensures that we only consider linearly independent columns of the instrument matrix, a common practice in the literature Lee2003. Take $M=W_0$, this is linked to the rank conditions of $W_0$, the variance covariance structure of $Z_t(W_0)$ and the correlation between $X_t$ and $Z_t(W_0)$. To further understand the implications of the above assumption, let {$Z_t(W_0)$ be the $n\times m$ matrix}. Note that $(nT)^{-1}\sum_{t=1}^T Z_t(W_0)^{\top} Z_t(W_0)$ is a {partitioned} matrix whose elements have the form $(nT)^{-1}\sum_{t=1}^T X_t^{\top} [W_0^{l_1}]^{\top} W_0^{l_2}X_t,$ for any $l_1$ and $l_2$ between 0 and $d$. This condition can be implied by $\lambda_{\min}(\lim_{n,T\to \infty}\{nT\}^{-1}\sum_t\mathbb{E}\{Z_t^{\top}Z_t\})>0 ,$ and $\lambda_{2K+1}(W_0W^{\top}_0)> 0$. Then $\Sigma_{Z_0Q_0}$ is of full column rank could be implied by a condition that $\lambda_{2K+1}(W_0W_{0}^{\top})>0$, $\lim_{T\to \infty}T^{-1}\sum_t[X_t,W_0X_t,\cdots, W_0^dX_t]$ has column rank greater than or equal to $2K+1$, and all eigenvalues of $I_n - \lambda_0 W_0$ are bounded away from $0$ and $\infty$. Assumption (ref) is a standard assumption imposed on the GMM weight matrix. Because $M$ is fixed and time invariant, $Z_t(M)$ is a measurable function of $X_t$. Therefore, $\{Z_t(M)\}$ inherits the strong-mixing property, with $\alpha_{Z(M),n}(h)\leq\alpha_{X,n}(h)$.

Before establishing the consistency of our spillover-parameter estimator, we first examine the bias introduced by $E$ when using the classical GMM estimator. We provide an intuitive argument illustrating how the measurement error matrix \(E\) biases the moment functions and the importance of its removal, particularly when the instrument matrix \(Z_t(M)\) also depends on \(W\), as in standard spatial regression. For simplicity here, assume no contextual effects are present.When $Z_t(M)$ consists of exogenous instruments satisfying $\mathbb{E}[Z_t(M)^\top \varepsilon_t] = \mathbf{0}_{m \times 1}$, the moment function is given by \[ \bar g_{nT}(\theta, W, M) = \frac{1}{nT} \sum_{t=1}^{T} Z_t(M)^\top (Y_t - \lambda W Y_t - X_t \beta). \] Substituting $W = W_0 + E$, \[ \bar g_{nT}(\theta, W, M) = \frac{1}{nT} \sum_{t=1}^{T} Z_t(M)^\top (Y_t - \lambda W_0 Y_t - X_t \beta) - \frac{\lambda}{nT} \sum_{t=1}^{T} Z_t(M)^\top E Y_t. \]

Thus, if $E$ is independent of $(Z_t(M),Y_t)$ and has mean zero, then \[ \mathbb{E}\!\left[Z_t(M)^\top E Y_t\right] =\mathbf{0}_{m\times 1}, \] and hence \[ \mathbb{E}\!\left[\bar g_{nT}(\theta_0,W,M)\right] =\mathbf{0}_{m\times 1}. \] However, when $Z_t(M)$ is a plug-in version of the true instrument with $M = W$, i.e., \[ Z_t(W) = [X_t, W X_t] = [X_t, (W_0 + E) X_t], \] the moment function becomes \[ \bar g_{nT}(\theta, W) = \frac{1}{nT} \sum_{t=1}^{T} [X_t, (W_0 + E) X_t]^\top \big[(Y_t - \lambda W_0 Y_t - X_t \beta) - \lambda E Y_t\big]. \]

Expanding the moment function $\Bar{g}_{nT}(\theta, W)$ with the instrument matrix $[X_t, (W_0+E)X_t]$, we get: {\[ \left[ \frac{1}{nT} \sum_{t=1}^{T}(-\lambda X_t^{\top} EY_t)^{\top}, \frac{1}{nT} \sum_{t=1}^{T}[X_t^\top E^\top (Y_t-\lambda W_0 Y_t-X_t\beta) - \lambda X_t^\top W_0^\top E Y_t - \lambda X_t^\top E^\top E Y_t]^{\top} \right]^{\top}. \]} Even if \(E_{ij}\)'s are i.i.d. mean zero, with bounded second moment and independent of $(X_t,\varepsilon_t)$, there will be moment condition misspecification. Specifically, \[ \mathbb{E}[\Bar{g}_{nT}(\theta_0, W)] \neq { \mathbf{0}_{m\times 1}}, \] since the expectation of the last term yields: $n^{-1}\mathbb{E}[X_t^\top E^\top E Y_t] = n^{-1} \mathbb{E}\big[X_t^\top \mathbb{E}[E^\top E]\,(I-\lambda_0 W_0)^{-1} X_t\big]\beta_0 \neq \mathbf{0}_{K\times 1},$ which might be non-negligible under our assumptions. Recall that \(\widehat{\theta}_{GMM}\) denotes the efficient two-step GMM estimator based on a given adjacency matrix \(W\), obtained in closed form in (ref). From the preceding results, it follows that \(\widehat{\theta}_{GMM}\) is likely to be inconsistent. We now analyze the convergence rate of \(\widehat{\theta}_p(\widehat{W})\).

Since \(W_0\) is not directly available for constructing instruments, one approach is to use \(Z_t(\widehat{W})\) as plug-in instruments, where \(\widehat{W}\) is obtained from Algorithms (ref) and satisfies Theorem (ref). When \(W_0\) is correctly observed, the minimizer \(\widehat{\theta}_p(W_0)\) satisfies the standard panel-data estimator convergence rate: \[ \|\widehat{\theta}_p(W_0) - \theta_0\|_2 = O_p(1/\sqrt{nT}). \] However, when \(W_0\) is replaced by a denoised estimate \(\widehat{W}\), the associated minimizer \(\widehat{\theta}_p(\widehat{W})\) depends on the deviation \(\widehat{W} - W_0\). The corresponding convergence properties are detailed in the following results.

theorem[Performance of the plug-in estimator] Recall the definitions of $B_0$, $\rho_n$, $s_n$, $\kappa_n$, and $\mathcal R_{nT}^*$ given above. Consider the model described in (ref). Suppose Assumptions (ref)--(ref) hold. (a) The plug-in estimator. If $\mathcal R_{nT}^*=o(1)$, then \[ \|\widehat\theta_p(\widehat W)-\theta_0\|_2 = O_p(\mathcal R_{nT}^*) + O_p\!\left(\frac{1}{\sqrt{nT}}\right). \] (b)(The GMM estimator). Using the noisy matrix $W = W_0 + E$ directly, \begin{align} \|\widehat{\theta}_{GMM} - \theta_0\|_2 = O_p(w_{2n} + w_{2n}^2) + O_p\!\left(\frac{1}{\sqrt{nT}}\right). \end{align} (c) Asymptotic normality. If \[ \sqrt{nT}\,\mathcal R_{nT}^*=o(1), \] then \[ \sqrt{nT}\, \bigl(\widehat\theta_p(\widehat W)-\theta_0\bigr) \overset{d}{\longrightarrow} \mathcal N(\mathbf 0,\Sigma_{GMM}). \] where \[ \Sigma_{GMM} =\left(G_0^\top \Lambda^{-1}G_0\right)^{-1} G_0^\top \Lambda^{-1}\Omega_0\Lambda^{-1}G_0 \left(G_0^\top \Lambda^{-1}G_0\right)^{-1}, \] and \[ \Omega_0 = \frac{1}{n} \mathbb{E}(Z_t(W_0)^\top \varepsilon_t\varepsilon_t^\top Z_t(W_0)). \]

Under the assumptions of Theorem (ref)(a), the plug-in estimator $\widehat\theta_p(\widehat W)$ is consistent provided that \[ \mathcal R_{nT}^*=o(1). \]

The factor $1/n$ discounts the noise contribution by the network size, so consistency holds even when $w_{2n}$ and $w_{\max n}$ do not vanish individually, for example, when $r$ and $m_s(S_0)$ remain bounded and $w_{2n}$ is of constant order. By contrast, the available bound for the GMM estimator based on the noisy matrix does not guarantee consistency when $w_{2n}$ is bounded away from zero; the moment-misspecification example above shows that inconsistency can occur in this case. The noise-induced component of the plug-in estimator's rate is strictly smaller whenever \[ \mathcal R_{nT}^* =o(w_{2n}+w_{2n}^2). \]

The improvement reflects how the $L+S$ decomposition enters the analysis. Denote $\Delta L := \widehat L - L_0$ and $\Delta S := \widehat S - S_0$ as the recovery errors from Theorem (ref). The recovery error $\widehat W - W_0 = \Delta L + \Delta S$ affects the spillover estimator only through bilinear forms (averages over the network) with the instruments and residual design, and these averages discount the error according to its structure. The low-rank error enters through its nuclear norm, controlled by $\sqrt r\,\|\Delta L\|_F$, contributing $O(r)$ effective directions rather than all $n$; the sparse error enters through $|\Delta S|_{1,1}$, controlled by its $m_s(S_0)$ nonzero entries rather than all $n^2$. The measurement error can therefore be large in aggregate while the spillover estimator still converges.

Example (ref) constructs regimes in which $|E|_{1,1}/n$ does not vanish, while the dimension-discounted condition $\mathcal R_{nT}^*=o(1)$ may still hold. This condition is sharp for sparse measurement error such as link misclassification in $0$--$1$ networks, but fails in the dense case that motivates our analysis: Example (ref) constructs the case in which $|E|_{1,1}$ is of order $n^2$ yet the plug-in estimator remains consistent. The two analyses cover disjoint measurement-error regimes: sparse errors with Lewbeletal2024EJ, dense errors with structured latent adjacency in our framework. When $L_0 = \mathbf{0}_{n \times n}$, the rate specializes to $O_p(m_s(S_0)w_{\max n}/n)$ (Corollary (ref)).

The following corollary specializes Theorem (ref) to the pure-sparse case $L_0=\mathbf 0_{n\times n}$. Within Assumption (ref), only part (ii) on $S_0$ is needed, and the coherence terms drop out. The rate is driven entirely by the sparse recovery error $n^{-1}|\widehat S-S_0|_{1,1}$, the same $\ell_1$-type quantity as the $|E|_{1,1}/n$ condition of Lewbeletal2024EJ, but applied to the recovery error after denoising rather than to the raw measurement error.

corollarySuppose $L_0=\mathbf 0_{n\times n}$ and let $\widehat S$ denote the pure-sparse estimator obtained by fixing $\widehat L=\mathbf 0_{n\times n}$. Suppose Assumption (ref)(ii), Assumption (ref)(i),(iii), and Assumptions (ref)--(ref) hold. Assume additionally that \[ \|S_0\|_1=O(1), \qquad \|E\|_\infty=O_p(1), \] and \[ \max_{1\leq j\leq K} \frac{1}{T}\sum_{t=1}^T\|X_{jt}\|_{\max}^2=O_p(1), \qquad \frac{1}{T}\sum_{t=1}^T \|\varepsilon_t\|_{\max}^2=O_p(1). \] Then \begin{align*} \|\widehat\theta_p(\widehat S)-\theta_0\|_2 &= O_p\!\left( \frac{|\widehat S-S_0|_{1,1}}{n} \right) + O_p\!\left(\frac{1}{\sqrt{nT}}\right)\\ &= O_p\!\left( \frac{m_s(S_0)w_{\max n}}{n} \right) + O_p\!\left(\frac{1}{\sqrt{nT}}\right). \end{align*}

The rate in Corollary (ref) follows from Lemma (ref)(b): the sparse recovery error enters the spillover estimator only through $n^{-1}|\widehat S-S_0|_{1,1}$. In the pure-sparse case, the sparse-recovery argument used in the proof of Theorem (ref) gives \[ \frac{1}{n}|\widehat S-S_0|_{1,1} = O_p\!\left( \frac{m_s(S_0)w_{\max n}}{n} \right). \] Although Corollary (ref) covers the pure-sparse case, the gain from our general analysis is most pronounced when the low-rank component $L_0$ is genuinely present and $r$ and $m_s(S_0)$ remain controlled relative to $n$.

{

remark[Asymptotic collinearity] It should be noted that variation in peer connections plays an important role in identification and estimation as noted, for example, in Manski1993 (p.535) and Paula2017. In a single, dense and large network, the leading eigenvector of a row-normalized and positive adjacency matrix may become (asymptotically) proportional to a vector of ones. In this case, peer averages become increasingly similar across individuals, complicating identification and estimation of spillover effects. A recent formulation of this issue in the context of peer effects is provided by hayes2024minimax, who show that this collinearity arises when $W_0$ is row-normalized, positive, and exhibits unbounded minimum degrees. Our setting differs in two respects. For a nonnegative row-normalized matrix, $W_0\mathbf 1_n=\mathbf 1_n$ holds exactly. When nodal covariates are independent of the network and the minimum degree diverges, peer averages may concentrate around a common value, producing asymptotic collinearity; see hayes2024minimax. Our assumptions do not generally impose row normalization or diverging minimum degree, so this conclusion does not follow automatically in our setting. The presence of a sparse component $S_0$ does not by itself eliminate this possibility. Without row normalization, $\mathbf 1_n$ need not be an eigenvector of $W_0$, and the argument requires separate verification. $\blacksquare$

}

As mentioned previously, we emphasize that our results are also suitable for the case when $E$ is endogenous with respect to $\varepsilon_t$. As an example, in social networks this can arise when misclassification errors are not generated at random, instead being correlated with unobservables driving the network formation process, such as the framework considered by Candelaria2023. To see this, we provide a rate analysis to compare the rate when $E$ is endogenous or not of both our estimator and the GMM estimator. For simplicity, we adopt the following assumption on the relationship between these two sources of error, see similar assumptions as in Lee2014:

assumption[Endogenous measurement error] Suppose Model (ref) holds and $W_0$ is symmetric and positive semidefinite. For each $i$, let \[ \varepsilon_i = (\varepsilon_{i1},\ldots,\varepsilon_{iT})^\top \in\mathbb R^T, \] and let \[ E_{i,-i} = (E_{ij}:j\neq i)^\top \in\mathbb R^{n-1} \] collect the off-diagonal elements of the $i$th row of $E$, with $E_{ii}=0$. Suppose that the vectors $(\varepsilon_i^\top,E_{i,-i}^\top)^\top$ are independent across $i$ and jointly Gaussian: \[ \begin{pmatrix} \varepsilon_i\\ E_{i,-i} \end{pmatrix} \sim \mathcal N\!\left( \mathbf 0,\, \Sigma_{\varepsilon E} \right), \] where \[ \Sigma_{\varepsilon E} = \begin{bmatrix} \sigma_0^2 I_T & \sigma_0\sigma_E n^{-s}\rho_{\varepsilon E}\\ \sigma_0\sigma_E n^{-s}\rho_{\varepsilon E}^\top & \sigma_E^2n^{-2s}\Sigma_E \end{bmatrix}, \qquad s>\frac12. \] Here, $\rho_{\varepsilon E}\in\mathbb R^{T\times(n-1)}$ governs the dependence between the outcome disturbances and the measurement error, and $\Sigma_E\in\mathbb R^{(n-1)\times(n-1)}$ is a correlation matrix. Assume that there exist constants $c_E,C_E,C_\rho>0$, independent of $n$ and $T$, such that \[ c_E \leq \lambda_{\min}(\Sigma_E) \leq \lambda_{\max}(\Sigma_E) \leq C_E, \qquad \|\rho_{\varepsilon E}\|_2\leq C_\rho, \] and that $\Sigma_{\varepsilon E}$ is positive semidefinite.

Assumption (ref) is imposed together with Assumptions (ref) and (ref). Thus, the disturbances remain marginally i.i.d.\ across $i$ and $t$, while dependence between $E$ and $\varepsilon_t$ is allowed. The Gaussian specification is imposed solely to illustrate the rate in the presence of endogenous measurement error.

Denote $\widehat{\theta}_{GMM}(W,M)$ and $\widehat{\theta}_p(\widehat{W},M)$ the GMM and plug-in estimator with instrument $Z_t(M)$. In comparison to the exogenous measurement error case, the bias of $\widehat{\theta}_{GMM}(W,M) = \widehat{\theta}_p(W,M)$ could be larger than our proposed estimators. It is because the deviation between sample moments $\Bar{g}_{nT}(\theta, W, M)$ and $\Bar{g}_{nT}(\theta, W_0, M)$ might be more severe. This is reflected in our simulation exercises, see Section (ref). For simplicity, we shall illustrate without the contextual-effects regressor $W_0X_t$ (i.e., setting $\gamma_0 = \mathbf{0}_K$), so the regressors are $[W_0Y_t, X_t]$. Define ${G}(M) = -n^{-1}\mathbb{E}[Z^{\top}_t(M)[W_0 Y_t, X_t]]$. To explain this, take $M$ as exogenous; the leading term for the moment estimator $\widehat{\theta}_p(\widehat{W},M) - \theta_0$ is $-(G^{\top}(M)\Lambda^{-1}G(M))^{-1}G^{\top}(M)\Lambda^{-1} \Bar{g}_{nT}(\theta_0, W_0, M)$, according to Theorem (ref). We shall suppress the dependence of $M$ on $W_0$ and $E$. The following proposition characterizes a rate bound for both the GMM estimator and our estimator.

Recall that $\bar g_{nT}(\theta,W,M) = \frac1{nT}\sum_{t=1}^T Z_t(M)^\top \left( Y_t-\lambda WY_t-X_t\beta-WX_t\gamma \right)$.

proposition[Endogenous error bias] Suppose that $M$ is fixed or exogenous and that the conditional-moment conditions in Lemma (ref)(c) hold. Under Assumptions (ref)--(ref), with $\|B_0\|_2(1+\|W_0\|_2)=O(1)$ and $\mathcal R_{nT}^*=o(1)$, we have \begin{align} \|\widehat\theta_p(\widehat W,M)-\theta_0\|_2 &\lesssim_p \mathcal R_{nT}^* +\frac1{\sqrt{nT}}, \\ \|\widehat\theta_{GMM}(W,M)-\theta_0\|_2 &\lesssim_p n^{1/2-s}\|\rho_{\varepsilon E}\|_2 +w_{2n} +\frac1{\sqrt{nT}}. \end{align}

Three features of this comparison are worth highlighting.

First, the endogeneity bias $n^{1/2-s}\|\rho_{\varepsilon E}\|_2$ appears only in the GMM rate (ref). It originates from a bilinear form that is quadratic in $E$: the correlation between $E$ and $\varepsilon_t$ contributes a non-vanishing mean to $Z_t^\top E B_0\varepsilon_t$, whose magnitude is governed by $\rho_{\varepsilon E}$. In the plug-in case, the analogous bilinear form involves $(\Delta L+\Delta S)B_0\varepsilon_t$. Its low-rank component is controlled by \[ \|\Delta L\|_* = O_p\!\left((r\vee1)(w_{2n}+\mathcal R_s)\right), \] whereas its off-diagonal sparse component is controlled by \[ |\Delta S|_1 = O_p\!\left(\mathcal R_s\sqrt{m_s(S_0)}\right). \] Under the conditional-moment conditions, the plug-in bound therefore contains no additional explicit term involving $\rho_{\varepsilon E}$. Second, even in the absence of endogeneity ($\rho_{\varepsilon E} = 0$), the denoised rate improves upon the GMM rate: the noise-level term $w_{2n}$ in (ref) is replaced by the rate $\mathcal{R}^*_{nT}$ in (ref), which is strictly smaller whenever $\mathcal{R}^*_{nT} = o(w_{2n})$.Third, the denoised rate attenuates the dependence on the noise magnitude $w_{2n}$ rather than eliminating it: $w_{2n}$ enters through the $L+S$ coupling $\|\Delta L\|_2 = O_p(w_{2n}+\mathcal{R}_s)$, but its effect is multiplied by the factor $r/n$ rather than by 1. As a result, the rate scales not with the raw noise level $w_{2n}$ but with the ratio of the structural complexity ($r$, $m_s(S_0)$) to the sample size $n$. This reflects the benefit of our approach. Finally, we consider the rate of the supervised estimator as defined in Equation ((ref)). Denote $\widehat{W}_s = \widehat{L}_s+\widehat{S}_s.$

theorem[Performance of the supervised estimator] Suppose Assumptions (ref)--(ref) hold. Let $M$ be fixed, exogenous, or $\sigma(E)$-measurable and satisfy the stated instrument and design conditions. Let $\widehat\theta_s$ denote the supervised estimator defined in (ref). \[ \nu_{nT}=C_\nu w_{2n}, \qquad \tau_{nT}=C_\tau(w_{\max n}+a_n), \] for sufficiently large constants $C_\nu,C_\tau>0$. Suppose \[ \mathcal R^*_{nT}=o(1) \qquad\text{and}\qquad \xi_{nT}=O(1). \] Then \begin{equation} \|\widehat\theta_s-\theta_0\|_2 = O_p\left( \frac1{\sqrt{nT}}+\mathcal R^*_{nT} \right). \end{equation} If, in addition, \[ \sqrt{nT}\,\mathcal R^*_{nT}=o(1) \qquad\text{and}\qquad \xi_{nT}=o(1), \] then \[ \sqrt{nT} (\widehat\theta_s-\theta_0) \overset{d}{\longrightarrow} \mathcal N(\mathbf 0,\Sigma_{GMM}). \]

The supervised estimator achieves the same rate as the plug-in estimator in Theorem (ref) and it is strictly smaller than that of the noisy-$W$ GMM estimator whenever $\mathcal R_{nT}^*=o(w_{2n}+w_{2n}^2)$. The rate still depends on $w_{2n}$, but only through the dimension-discounted recovery rate $\mathcal R_{nT}^*$. It is worth noting that the rate in (ref) depends on the parameters $(r, m_s(S_0), n, T)$ and on the spectral characteristics of the underlying network ($\|B_0\|_2, \|W_0\|_2$), rather than directly on the noise level $w_{2n}$ compared to the rate of the GMM estimator. The condition on $\xi_{nT}$ ensures that the generated error from the first-iteration estimator $\widehat{\theta}^{[1]}$ does not inflate either the sparse recovery in Step b) or the low-rank recovery in Step c) beyond the unsupervised noise level, so that the supervised decomposition matches the plug-in decomposition rates.

The following corollary specializes Theorem (ref) to the pure-sparse case $L_0 = \mathbf{0}_{n\times n}$, paralleling Corollary (ref): only Assumption (ref)(ii) on $S_0$ is needed, the coherence terms drop out, and the rate is driven entirely by the sparse recovery error $n^{-1}|\widehat S_s - S_0|_{1,1}$, making the connection to the misclassification framework of Lewbeletal2024EJ explicit.

corollarySuppose $L_0=\mathbf 0_{n\times n}$ and consider the pure-sparse supervised estimator obtained by imposing $L=\mathbf 0_{n\times n}$, so that $\widehat W_s=\widehat S_s$. Suppose Assumption (ref)(ii), Assumption (ref)(i),(iii), and Assumptions (ref)--(ref) hold. Assume, in addition, that \[ \|W_0\|_1=O(1), \qquad \|E\|_\infty=O_p(1), \qquad \xi_{nT}=O(1), \] and \[ \max_{1\leq j\leq K} \frac{1}{T}\sum_{t=1}^T\|X_{jt}\|_{\max}^2=O_p(1), \qquad \frac{1}{T}\sum_{t=1}^T \|\varepsilon_t\|_{\max}^2=O_p(1). \] Then \begin{align} \|\widehat\theta_s-\theta_0\|_2 &= O_p\!\left( \frac{1}{n}|\widehat S_s-S_0|_{1,1} \right) + O_p\!\left(\frac{1}{\sqrt{nT}}\right). \end{align} Moreover, Lemma (ref)(b), together with the sparse-recovery bound for $\widehat S_s$, gives \begin{align} \|\widehat\theta_s-\theta_0\|_2 &= O_p\!\left( \frac{m_s(S_0)w_{\max n}}{n} \right) + O_p\!\left(\frac{1}{\sqrt{nT}}\right). \end{align}

The theoretical performance of the supervised estimator is similar to that of the plug-in estimator, as confirmed by the simulation and application results in Sections (ref) and (ref). We shall also note that both Theorems (ref) and (ref) do not require $E$ to be exogenous, as long as $Z_t(M)$ does not directly involve $W$. Furthermore, both theorems yield rates in which the noise level $w_{2n}$ enters only after dimension-discounting by the factor $r/n$ (rather than at the raw spectral rate $w_{2n}$ delivered by the noisy-$W$ baseline), reflecting the fundamental advantage of the $L+S$ decomposition combined with the within transformation.

Simulations

In this section, we present simulation exercises to evaluate the performance of our proposed estimators. We first introduce four data-generating processes (DGPs) for the network adjacency matrix $W_0$. We then describe how it is contaminated with noise, and discuss the Monte Carlo simulation results.

Generating $\mathbf{W}_0$

The diagonal elements of $W_0$ are assumed to be 0. Hence we set all diagonal elements of $E$ to be $0$ and do not penalize $S_0$ on its diagonal. The other entries of $E$ are generated as explained in Section (ref).

itemize• DGP 1(Low rank) In this setup, $L_0$ is generated by a pure low-rank matrix as: \begin{equation} L_0 = U DV^{\top}, \quad U, V \in \mathbb{R}^{n \times r} \, , and \quad U^{\top} U = I_r, \quad V^{\top} V = I_r, \end{equation} where $D$ is the $r\times r$ diagonal matrix $D = \operatorname{diag}\{0.8^0, 0.8^1, \ldots, 0.8^{r-1}\}$, and $U$ and $V$ ---the left and right singular vectors of $L_0$, respectively--- are the first $r$ columns of two random $n \times n$ orthonormal matrices generated according to standard algorithms Stewart1980, Mezzadri2007. We then set $W_0=L_0^*$, where $L_0^*$ has zero diagonal elements and its off-diagonal elements are the same as $L_0$. • DGP 2 (Low rank+ sparse) Our second DGP still uses the same $L_0$ as in DGP 1, but adds a sparse $n \times n$ component $S_0$, where we randomly pick $m_r=2$ elements in each row and draw their values from a $\text{Uniform}(0.5, 1)$ distribution. Note that $W_0$ could be written as \begin{equation} W_0 =L_0^*+ S_0^*, \end{equation} where $L_0^*$ is defined as in DGP 1 using $U$, $V$ and $D$ generated as in (ref), and $S_0^*$ has zero diagonal elements and its off-diagonal elements are the same as $S_0$. • DGP 3 (Dominant units) Our third DGP generates $W_0$ as a network with dominant units as presented in Subsection (ref). Suppose the number of dominant units is $J=2$. The first $J$ units are dominant players in the network, whereas the remaining $n-J$ units are ordinary players. We could partition $W_0$ as \begin{eqnarray} W_0=\begin{bmatrix} W_{11} & W_{12} \\ W_{21} & W_{22} \end{bmatrix} . \end{eqnarray} Let $W_D = [W_{11}^{\top}, W_{21}^{\top}]^{\top}$ be the $n \times J$ dominant block of the adjacency matrix corresponding to the dominant units, and $W_{R}= [W_{12}^{\top}, W_{22}^{\top}]^{\top}$, the right block of the adjacency matrix corresponding to ordinary units. For each column of $W_D$, the first $\lfloor n^{0.9} \rfloor$ entries, except the diagonal elements, are drawn independently from a $\text{Uniform}(0, 1)$ distribution, and the remaining entries are set equal to $0$. For $W_{R}$, we induce sparsity by imposing that each individual is connected with equal weights to the immediate predecessor and successor. Note that $W_0$ could be viewed as a purely sparse matrix $S_0^*$ as each of its row has at most $2+J$ nonzero elements, and we treat $L_0^*$ as a zero matrix. • DGP 4 (Group) Our fourth DGP divides the individuals into two groups. The first $n/4$ members form group 1, the remaining $3n/4$ form group 2, and only members within the same group have connections. For each group, the connection strength between any two different individuals $i$ and $j$ within the same group is defined by $d_g u_{g,i}u_{g,j}/\sum_i u_{g,i}^2$, where $d_{1,1}=1$ and $d_{2,2}=0.9$, $u_{g,i}$'s are latent individual-specific characteristic generated independently from standard normal random variables. Note that DGP 4 admits the low-rank form such that $L_0=UDU^\top$, where $U=[U_1, U_2]$, $D$ is a $2\times 2$ diagonal matrix with $d_{1,1}=1$ and $d_{2,2}=0.9$, $U_1=[u_{1,1}, \cdots, u_{1,n/4}, 0, \cdots, 0]/\sqrt{\sum_i u_{1,i}^2}$ and $U_2=[0, \cdots, 0, u_{2,1}, \cdots, u_{2,3n/4}]/\sqrt{\sum_i u_{2,i}^2} $, and $W_0=L_0^*$, where $L_0^*$ has zero diagonal entries and its off diagonal elements are the same as $L_0$.

Monte Carlo setup

For our simulation exercises, we set $\lambda_0 = 0.25$, $\beta_0 = (-1, 2)$. We generate two covariates for $X_t$ independently from a $\mathbb{N}(0, 1)$ and a $\mathbb{N}(5, 2)$, respectively. We set sample sizes $n \in \{40,80, 120\}$ and time dimensions $T \in \{1, 5, 15, 50\}$, considering specific combinations of panel sizes due to computational constraints. We take as given a true network $W_0$ generated using one of the DGPs introduced in the previous subsection, and keep it constant across 100 Monte Carlo simulations.

Our exercises additionally consider the possibility that measurement errors are endogenous. That is, we allow for correlated disturbances between the outcome equation ($\varepsilon_t$) and the errors that contaminate the adjacency matrix $(E)$, as outlined in our Assumption (ref). Let $\sigma_E $ and $\sigma_\varepsilon$ be two positive constants. Specifically, for $i,j = 1, \cdots, n$, we generate $E_{ij} \sim i.i.d. \quad \mathbb{N}(0, \sigma^2_E)$ for $i\neq j$, and set $E_{ii}=0$. We generate $\varepsilon_t$ such that its $i$th component $\varepsilon_{it}$ satisfies $\varepsilon_{it}=\rho \sigma_\varepsilon\sum_{j=1}^n E_{ij}/\sqrt{n} +v_{it}$, and $v_{it}$ are i.i.d. $ \mathbb{N} (0, (1-\rho^2)\sigma_\varepsilon^2)$ that are independent of $E_{ij}$. Note that $\varepsilon_{t}|E$ is normally distributed and we could simply obtain its conditional mean and variance. We define $\mbox{Vec}(E)$ as the vector obtained by stacking the off-diagonal elements of $E$.

eqnarray*[eqnarray* omitted — 399 chars of source]

where $\Sigma_{\varepsilon E}$ is the $n\times (n^2-n)$ matrix that captures the correlation between $\varepsilon_{it}$ and $ \mbox{Vec}(E)$ stacks all $E_{ij}$ for $i\neq j$ into a long vector. Note that it is implied from the previous sections that $\mbox{cov}(\varepsilon_{it}, E_{ij})=\rho\sigma_E^2\sigma_\varepsilon/\sqrt{n}$ for $i\neq j$ and $\mbox{cov}(\varepsilon_{it}, E_{kj})=0$ if $k\neq i$. In the simulation, we set $\sigma_\varepsilon=0.15$, $\sigma_E = 0.3/n^{0.7}$ and $\rho=0$ or $0.7$. By setting $\rho = 0$, we recover the exogenous measurement error case from Assumption (ref).

Given values for the adjacency $W_0$, covariates $X_t$, parameters $\theta_0$, and disturbances $\varepsilon_t$, we use the reduced-form representation of model (ref) to generate our outcome $Y_t$ at every time period:

align[align omitted — 131 chars of source]

Throughout the simulation we set $W = W_0 + E$, which is the observed noisy version of the adjacency matrix, to construct estimates of parameters $\theta_0$. Once the adjacency matrix is generated, we use it to construct a rough estimate of the variance of the errors and set the tuning parameters for estimation. Specifically, we fix $\xi_{nT} = 1$ and set $\mbox{Int}(W)$ as the interquartile range of $W_{ij}$. Then we choose the sparse penalty as $\tau_n=\tau_{nT} = 2\mbox{Int}(W)\log(n)$ and $\nu_n=\nu_{nT}=\mbox{Int}(W) \sqrt{n}$.

Simulation results

table[table omitted — 4,519 chars of source]

{\scriptsize

table[table omitted — 4,317 chars of source]

}

table[table omitted — 4,589 chars of source]

Tables (ref) - (ref) present the simulation results using the setting outlined in the previous subsections. We highlight several key takeaways in this setup. First, Tables (ref) reports relative root mean squared errors (RMSE) for estimates of the spillover scalar parameter $\lambda_0$, where the benchmark is the standard GMM estimator. Observe that both our plug-in and supervised estimators achieve lower RMSEs than the GMM benchmark uniformly across all DGPs and sample sizes. Moreover, the GMM estimates tend to perform worse in the endogenous case and hence there is more improvement in the endogenous case. Our proposed methods perform better than GMM by approximately 50--80% for small $n$ and $T$. Our methods achieve at least a 75% improvement in RMSE for $n=120$, $T=50$. The relative performance of our methods generally improves as the number of nodes in the network $n$ and the number of time periods $T$ increase. This behavior is as predicted by the theory in the cases where the adjacency matrix is assumed to be constant across time.

In terms of the recovery performance of the denoised adjacency matrix, as evidenced in Table (ref), the plug-in method and the supervised method perform similarly. Moreover, there is little difference between the exogenous and endogenous cases because we use the same observed adjacency matrix.

Finally, we explore how our methods perform in recovering not just the scalar spillover effect $\lambda_0$, but the full diffusion multiplier matrix given by $(I_n - \lambda_0 W_0)^{-1}$. This quantity is highly important as a policy tool, as it summarizes the cumulative spillover effects of a given economic shock. As shown in Table (ref), by providing reliable estimates for both $\lambda_0$ and $W_0$, our proposed methods achieve better performance in terms of Frobenius norm across DGPs and sample sizes. Although the estimates of $W_0$ might be similar in both the exogenous and the endogenous cases, our method could obtain more reliable estimates of $\lambda_0$ as well as better recovery accuracy in the endogenous case.

Empirical Applications

This section illustrates the use of our estimators for social/spatial interactions. We consider two examples, applying our methods and examining the spillover effects. In each example, we first construct a raw spatial weight matrix and then apply Algorithms (ref) and (ref) described in the Appendix to obtain our estimates. To select the penalties, we use the interquartile range (IQR) of the raw weight-matrix entries as a robust scale estimate, denoted by $\widehat\sigma$. The sparse and low-rank penalties are parameterized as $C_S\widehat\sigma\log(n)$ and $C_L\widehat\sigma\sqrt{n}$, respectively, with the constants chosen by the Bayesian Information Criterion (BIC). We estimate the spillover coefficient $\widehat\lambda$ and refer the column sum of the Leontief inverse $(I_n-\widehat\lambda\widehat W)^{-1}$ as the total network influence.

Example 1

In our first example, we examine an application to annual Gross Domestic Product (GDP) growth as our outcome of interest. Following MRW1992, the explanatory variables include the annual growth rate of the working-age population defined as individuals aged 15 to 64, denoted as \(X_1\), and the log investment share, \(X_2\), which serves as a proxy for the saving rate. We also include individual and time fixed effects in our specifications. The dataset is downloaded from the World Development Indicators (WDI) and covers a balanced panel of 23 economies from 1970 to 2023. \footnote{These economies are: Australia (AU), Canada (CA), Mainland China (CN), Denmark (DK), Finland (FI), France (FR), Germany (DE), Hong Kong SAR China (HK), Indonesia (ID), Ireland (IE), Italy (IT), Japan (JP), Korea (KR), Malaysia (MY), Mexico (MX), Netherlands (NL), New Zealand (NZ), Singapore (SG), Spain (ES), Sweden (SE), Thailand (TH), United Kingdom (GB), and United States (US).}

Economic activity is well known to exhibit substantial spatial dependence Krugman1991. Neighboring economies tend to trade more, share regional demand and supply shocks, and compete for mobile factors of production. For this reason, we use a geographic‑proximity weight matrix as a natural and predetermined benchmark to study spillovers. The raw spatial weight matrix is constructed as the row-normalized inverse squared distance matrix based on Haversine distances and denoted by $W_2$. \footnote{ For economies \(i\neq j\), we compute the Haversine great-circle distance \(D_{ij}\) between capital cities (in millions of meters) and then form the raw weight matrix by setting the $(i,j)$ entry as $1/D_{ij}^{2}$ for $i\neq j$, and the diagonal entries as 0, and then apply row-normalization. } Note that row-normalization on the weight matrix is standard in the spatial econometrics literature, as it facilitates a clear interpretation of the spatial autoregressive coefficient as the strength of the average influence. All off-diagonal entries of $W_2$ are strictly positive, which contrasts sharply with the sparsity assumptions sometimes invoked in the literature. This dense structure makes $W_2$ a natural candidate for our estimating approach, since many of the small off-diagonal weights likely reflect weak, noise‑contaminated links rather than economically meaningful connections.

We estimate Model ((ref)) with two-way fixed effects under two scenarios. Panel A includes only \(X_1\) and \(X_2\) as covariates. We assume that those variables are contemporaneously exogenous and use the spatial lags \(WX_1\) and \(WX_2\) as instruments. Panel B includes both the original covariates and the contextual variables \(WX_1\) and \(WX_2\) as regressors, using the second-order spatial lags \(W^2X_1\) and \(W^2X_2\) as instruments. Table (ref) reports the deviations of the estimated weight matrices from the raw matrix $W_2$. We consider three alternative specifications: a purely sparse structure with $\tau=0.0443$, a purely low-rank structure with $\nu=0.2709$, and a combination of both with $\tau=0.0443$ and $\nu=0.2709$, where the penalties are selected according to BIC respectively. We first examine the differences between estimated adjacency matrices and the raw matrix $W_2$. We find that the purely low-rank estimator $\widehat L$ exhibits substantial departure from $W_2$, particularly in terms of the matrix maximum norm. Most of the elements in $\widehat L$ are small but not exactly 0, suggesting that a low-rank approximation alone may fail to preserve several important links in the original network. By contrast, the purely sparse estimator $\widehat S$ retains only 96 nonzero entries. Although this produces a more parsimonious network representation, it has the largest deviation from $W_2$ under $l_2$ norm. The low-rank-plus-sparse estimators provide a more balanced representation and both $\|\widehat{W}_2-W_2\|_{\max}$ and $\|\widehat{W}_2-W_2\|_{2}$ remain small relative to the corresponding deviations of the competitors. A further calculation shows that $\|\widehat{W}_2-\widehat S\|_{\max}=0.02$ and $\|\widehat{W}_2-\widehat S\|_{2}=0.11$, indicating that it may capture diffuse connections that may be discarded by pure sparsification. Notably, the supervised low-rank and sparse estimates based on two panel specifications remain very similar to the unsupervised one, which is consistent with our theoretical findings. A further comparison of the heat maps of the raw $W_{2}$, the estimated $\widehat{W}_{2}$, and $\widehat{S}$ is also shown in Figure (ref). We observe that the low-rank plus sparse estimate $\widehat{W}_{2}$ appears to preserve the dominant network structure while allowing for pervasive weak link effects.

table[table omitted — 1,219 chars of source]
figure[figure omitted — 387 chars of source]

In the regression analysis, we evaluate the impact of the investment rate and working-age population growth, and report the estimated spillover coefficient $\widehat{\lambda}$ based on different weight matrices in Table (ref). For each candidate weight matrix $W$, we also compute the column sums of the Leontief inverse $(I_n-\widehat{\lambda}\widehat W)^{-1}$ and designate the economy with the largest column sum as the key player. Relative changes compared to the raw $W_2$ benchmark are also reported.

In Panel A (without contextual effects), the purely sparse estimator $\widehat{S}_2$ only retains the strong connections and reduces $\widehat{\lambda}$ by 16.7% and total influence (i.e., the corresponding column sum from the Leontief matrix, see above) by 11.3% compared to those obtained from the raw $W_2$. Conversely, the purely low-rank estimator $\widehat{L}_2$, amplifies network connections via aggregating pervasive weak links and increases $\widehat{\lambda}$ by 11.3% and total influence by 12.0% compared to those obtained from the raw $W_2$. These results indicate that the sparse and low-rank components might capture distinct aspects of the network. The sparse component filters out scattered weak links, whereas the low-rank component summarizes pervasive dependence. Correspondingly, the low-rank-plus-sparse estimator $\widehat{W}_2$ produced more moderate deviations from the raw benchmark.

The benefits of our approach become more pronounced in Panel B with contextual effects. The raw $W_2$ produces a negative estimate $\widehat{\lambda}$ that is much lower than the our estimates. The sparse estimator and the low-rank-plus-sparse estimator produce moderate positive values that might be more reasonable. As higher-order spatial lags can magnify noise in the raw matrix, this reversal suggests that our approach can improve the plausibility and stability of the estimated spillover effects.

table[table omitted — 1,881 chars of source]

Moreover, the decomposition also yields interesting practical insights. Compared with $\widehat S_2$, we find that including the low-rank component in $\widehat W_2$ primarily rescales the overall magnitude of the spillover influence vector rather than altering the key player. This pattern indicates that, while the identity of the key player remains robust to the decomposition strategy, the estimated strength of network spillovers depends materially on whether the low-rank component is retained, underscoring the practical relevance of the full low-rank plus sparse decomposition over the purely sparse estimator. Figure (ref) presents the graphical representations of $\widehat{W}_{2}$, $\widehat{L}_{2}$, and $\widehat{S}_{2}$, respectively. $\widehat W_2$ admits the decomposition $\widehat W_2=\widehat S_{\widehat W_2}+\widehat L_{\widehat W_2}$ and comparing the leading eigenvectors of $\widehat W_2$, $\widehat L_{\widehat W_2}$ and $\widehat S_{\widehat W_2}$, we find that ${\bf v}_{\widehat{W}_2}\approx 0.56\,{\bf v}_{\widehat{L}_{\widehat W_2}}+0.48\,{\bf v}_{\widehat{S}_{\widehat W_2}}$ from a regression perspective as discussed in equation ((ref)). Both the low-rank and the sparse components contribute substantially in this application, indicating that the identification of influential nodes depends not only on a few strong connections, but also on the cumulative effect of widespread weak links. Notably, as shown in Figure (ref), ${\bf v}_{\widehat{S}_{\widehat W_2}}$ has many zero entries, and its nonzero entries concentrate on the Asian economies, whereas ${\bf v}_{\widehat{L}_{\widehat W_2}}$ loads on all economies but assigns slightly smaller weights to Western economies. One possible interpretation is that the sparse eigenvector captures localized regional clustering centered on Asia, while the low-rank eigenvector reflects a global common factor that is slightly attenuated for Western economies in accounting for the broader geographic diffusion inherent in the low-rank structure.

figure[figure omitted — 398 chars of source]
figure[figure omitted — 1,123 chars of source]

Example 2

In our second example, we apply our methods to examine tax competition among the 48 contiguous U.S.\ states. This application was originally studied by Besley1995 using data from 1962 to 1988, with a pre‑specified geographic weight matrix $W_g$ defined as the row‑normalized 0‑1 neighborhood adjacency matrix setting 1 to state pairs that are geographically adjacent and zero, otherwise. Besley1995 emphasize a political‑economy channel for tax‑setting spillovers, as policymakers respond to voters who benchmark their policies against those of peer states. Consequently, states may face both political and economic incentives to align their fiscal policies with those of their neighbors. This dataset has recently been extended to cover $1962$--$2015$ ($T = 53$) and analyzed by Paula2023. We obtain the spatial weight matrix $W_2$, which is a row-normalized inverse squared distance matrix based on Haversine distances between state capitals, as the observed weight matrix. In addition to endogenous spatial effects captured by $WY$, the covariates include state-level characteristics such as per-capita income, the unemployment rate, and the proportions of young and elderly populations, without and with their contextual covariates $WX$ in Panel A and B respectively. We also include individual and time fixed effects in our specifications. Lagged neighbor income and lagged neighbor unemployment are used as instruments as in the original paper by Besley1995.

We consider three specifications, which yield a purely sparse matrix estimate $\widehat{S}_2$, a purely low‑rank matrix estimate $\widehat{L}_2$, and a low‑rank plus sparse estimate $\widehat{W}_2$. The purely sparse estimate $\widehat{S}_2$ captures the strong links and retains 90 nonzero entries. The estimate $\widehat{L}_2$ is a zero matrix under the selected penalty, suggesting that the data do not provide enough evidence of a low-rank component. The combined estimator $\widehat{W}_2$ coincides with the sparse‑only estimator $\widehat{S}_2$, and hence we only report results about $\widehat{W}_2$ in the subsequent analysis. The supervised variants $W_s$ in Panels A and B are very close to those of the unsupervised estimate $\widehat{W}_2$.

figure[figure omitted — 441 chars of source]
figure[figure omitted — 453 chars of source]

Figures (ref) and (ref) provide heatmaps and graphical representations of the raw geographic weight matrices $W_g$, $W_2$ and the estimate $\widehat W_2$, respectively. We find that the raw inverse geographic distance matrix $W_2$ is much denser than the raw geographic adjacency matrix $W_g$. It may be more vulnerable to measurement errors, and thus exaggerate the overall degree of network dependence and obscure the dominant transmission channels. By contrast, the estimated network discards many weak, noise‑dominated connections, thus providing a clearer visualization of the significant links and improving interpretability.

Table (ref) provides a more detailed characterization of the state-level network structure. States are sorted according to their column sums, which measure the aggregate strength of their outward connections. Massachusetts is identified as the most influential state under all three weight matrices. Its out-degree is 1.70 under the row-normalized adjacency matrix $W_g$, 1.61 under the row normalized inverse squared distance matrix $W_2$, and 0.72 under the estimate $\widehat W_2$. The estimated network therefore preserves the leading position of Massachusetts while substantially reducing the magnitude of its estimated outward influence. This suggests that our estimate does not fundamentally alter the ranking of the major network hubs, but removes potentially noisy links that may inflate the overall strength of network connections rather than incorporating them into the core neighborhood topology. By contrast, the ranking based on $W_g$ differs more noticeably because the geographical adjacency matrix only captures direct neighboring relationships, whereas $W_2$ allows for distance-decaying connections between all states.

table[table omitted — 1,066 chars of source]
table[table omitted — 1,741 chars of source]

Table (ref) reports the estimated spillover coefficient $\widehat{\lambda}$ and the regional influence analysis using different weight matrices. We consider two scenarios, Panel A (without $WX$) and Panel B (with $WX$), respectively. In both panels, two raw weight matrices $W_g$ and $W_2$ produce quite similar results, and they both exhibit much larger or even explosive values of $\widehat{\lambda}$ in Panel B. Such large values are likely driven by measurement error or noise in the raw adjacency matrices, yielding exaggerated spillover estimates that could be unreliable and potentially misleading. In contrast, the estimators, including the low-rank-plus-sparse estimator $\widehat{W}$, and its supervised variants, produce substantially smaller and more stable $\widehat{\lambda}$ values, and consistently identify the South Atlantic as the key region. Recall that each weight matrix is row-normalized prior to estimation. By the Perron–Frobenius theorem, the spectral radius of a row-normalized non‑negative matrix is 1, and hence the standard Leontief-stability condition reduces to $|\widehat{\lambda}|<1$ LeSage2009. Explosive spillover estimates can cause some elements of the Leontief inverse to become negative and hard to interpret. These results indicate that our approach might yield more plausible and interpretable spillover effects, avoiding the inflated influence suggested by raw network matrices.

table[table omitted — 1,621 chars of source]

Table (ref) further examines the propagation of shocks originating from Massachusetts. We calculate the Leontief Inverse and examine the four largest elements in the column that corresponding to Massachusetts. The two states most strongly influenced by a shock originating from Massachusetts are Massachusetts itself and Rhode Island. However, the magnitude of the estimated network influence is attenuated. When contextual effects $WX$ are incorporated, the Leontief influence values increase substantially. Under the estimate $\widehat W_2$, the influence measure for Massachusetts rises from 1.04 to 2.51, while the corresponding values for Rhode Island, New Hampshire, and Maine increase to 1.88, 1.82, and 1.48, respectively. Overall, the results consistently identify Massachusetts as an important regional hub and indicate that contextual effects could considerably amplify the propagation of state-level shocks. Our approach retains stronger local transmission channels while suppressing weaker long-range associations that may be sensitive to measurement error.

Conclusions

This paper introduces a robust framework for studying interaction effects in social and spatial networks under noisy adjacency matrices. By leveraging the low-rank and sparse structure of real-world networks, our de-noising approach, combining LASSO and nuclear norm penalization, significantly improves estimation accuracy. We propose two estimation procedures—a two-step and a one-step supervised GMM estimator—that outperform standard GMM by $50-80\%$ in RMSE under noise and endogeneity.

Applying our proposed methodology to examine the international spillover of economic growth and the tax competition across U.S. states, we find pronounced discrepancy in estimating the spillover effects, which highlights the essential role of accurate network denoising. Failing to distinguish genuine connectivity from measurement noise can systematically distort spillover coefficients as well as the strength and direction of cross-unit interactions. Moreover, our decomposition framework might provide a new perspective on complex spatial dependencies by disentangling global and local components. This decomposition not only sharpens the economic interpretation but also improve econometric inference, thereby contributing to credible spatial econometric analysis.

Declaration of generative AI and AI-assisted technologies in the writing process

During the preparation of this work the authors used AI to suggest edits to some lengthy passages for concision and clarity. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.