The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
144,382 characters
Estimating Network Spillovers under Dense Measurement Error
\maketitle
\begin{abstract}
This paper analyzes spillover effects in spatial (network) models when the neighborhood (adjacency) matrix is contaminated by measurement error from reporting, aggregation, or disclosure imperfections, leading to inconsistent estimation of network effects. We introduce a regularization framework for the latent network that allows for sparse and/or low-rank structure and accommodates potential correlation between measurement errors and outcomes.
We propose two estimators: (i) a two-stage procedure that first denoises the adjacency matrix and then incorporates the purified network into a regression analysis, and (ii) a Generalized Method of Moments (GMM) estimator that jointly estimates regression parameters and refines the network structure. We then establish strictly improved consistency rates for the spillover effect estimator relative to naive estimation ignoring measurement error. Simulations demonstrate that, in the presence of noisy networks,
our approach reduces the root mean squared error of spillover estimates relative to conventional methods by approximately $50-80\%$.
We apply our framework to examine the international spillover of economic growth, and the tax competition across U.S. states, illustrating that denoising might restore Leontief stability and yields improved estimates of spillovers.
\end{abstract}
\vspace{5mm}
\noindent{JEL Classification:} C21, C23, D57
\vspace{1mm} \noindent{Keywords:} network analysis, measurement error, LASSO, nuclear norm penalty, penalized GMM
\newpage
{
\section{Introduction}
\label{Sec:Introduction}
Network analysis provides powerful tools to quantify policy spillovers
and identify pivotal agents for targeted interventions, particularly
when policymakers face binding resource constraints
\citep{Ballester2006, Zenou2016, Galeotti2020}. A large literature
shows that economic outcomes often depend on interactions among
interconnected units, giving rise to amplification, diffusion, and
multiplier effects that are central to policy design
\citep{Acemoglu2012, Giorgi2020, Araujo2023}. These insights have
motivated the widespread use of spatial and social interaction models
to study peer effects, diffusion dynamics, and network-based policy
interventions across a broad range of economic settings.
Methodologically, the literature on spatial and social interaction models
differs mainly in how the network structure is handled. A dominant strand
assumes that the adjacency matrix is correctly specified and directly
observed, and studies spillover effects conditional on this object
\citep{Lee2007a, Bramoulle2009, Lee2010a, Banerjee2013, Elhorst2014,
Lee2014, Yang2017, Zhu2020a, Kuersteiner2020},\\ \citep{ Shin2021,
chernozhukov2021uniform}. An alternative and growing line of research
relaxes this assumption by estimating the network jointly with model
parameters, typically imposing dimension reduction constraints such as
sparsity to address the curse of dimensionality in high-dimensional
settings \citep{Manresa2016, Lam2020, Paula2023}.
Sparsity-based methods work well when the latent network is dominated
by a few strong links, but they fail in the dense networks that arise
in production, trade, and financial settings, where the observed
adjacency matrix is itself contaminated by measurement error
\citep{FismanWei2004, Upper2011, Anand2015}. In such settings, two
distinct forms of latent structure coexist beyond idiosyncratic noise.
First, outcomes may be driven by a small number of pervasive (though
individually weak) latent forces such as common shocks, shared
technologies, or institutional linkages, naturally represented by a
low-rank component \citep{Chudik2013, Kapetanios2021}. Second,
networks may contain a small number of strong bilateral or
idiosyncratic interactions, naturally represented by a sparse
component \citep{Lam2020, Paula2023}. Imposing sparsity alone discards
the pervasive signal, while ignoring measurement error produces
non-vanishing bias as elementwise small errors accumulate along
network paths \citep{Lewbeletal2024EJ}. The stakes are practical: in
our cross-state tax-competition application, the standard spatial GMM
estimator with the raw inverse-distance weight matrix yields
$\widehat\lambda > 1$, violating the Leontief stability condition and
implying explosive multipliers, while our estimator returns
$\widehat\lambda < 1$ within the admissible region. We therefore
propose a regularization approach for the observed adjacency matrix
that filters measurement error while preserving low-rank or sparse
signal, developed formally in Section~\ref{Sec:Methodology}.
Throughout, by \emph{regularization} we mean imposing restrictions: low-rank, sparse, or low-rank-plus-sparse, directly on the
latent adjacency matrix $W_0$ rather than on the regression
coefficients, as identifying restrictions that make denoising of
the observed network feasible.
The problem is non-trivial because it combines matrix-valued
measurement error with a spillover estimator that depends nonlinearly
on the network matrix through the Leontief inverse
$(I_n - \lambda W)^{-1}$, so errors in $W$ propagate through the
estimator in ways classical errors-in-variables corrections do not
handle.
Low-rank and sparse structures have deep foundations in statistics
and econometrics, appearing in approximate factor models,
latent-variable graphical models, and high-dimensional covariance
estimation \citep{Bai2002, Chandrasekaran2010, Fan2011}. A large
methodological literature develops penalized procedures to recover
low-rank or sparse patterns in covariance and precision matrices
\citep{Cai2011, Candes2011, Agarwal2012, Fan2018}. Despite their
success in these settings, regularization of this kind has not, to
our knowledge, been adopted in the econometrics literature to address
measurement error in observed network adjacency matrices, or to
improve the estimation of spillover effects and other policy-relevant
network statistics. We build on this literature by developing two
complementary estimation approaches. The first is a two-stage
procedure that denoises the observed adjacency matrix by separating
low-rank or sparse signal from measurement error before estimating
spillover effects. The second is a regularized Generalized Method of
Moments (GMM) estimator that jointly estimates the network and the
model parameters in a single step.
We establish that our estimator is consistent, in contrast to naive estimation that ignores measurement error and remains inconsistent in dense regimes.
Simulation results show that the proposed method achieves a $50\text{--}80\%$ reduction in the root mean squared error of spillover estimates under noisy weight matrices, with measurement noise ranging from moderate to high.
Our main theoretical contribution is to characterize how the bias
of standard spillover estimators depends on the structure of the
latent network and the measurement error, and to show that this
bias does not vanish in the dense regime that characterizes economic
networks. \cite{Lewbeletal2024EJ} show that modest link
misclassification bias in the estimation of linear social-interaction
models may be negligible as long as misclassification is not
pervasive; we extend their analysis to settings in which the latent
network has many but individually weak connections, and in which the
measurement error itself need not be sparse. We further allow the
measurement error to be correlated with unobservables in the outcome
equation \citep{Candelaria2023, lewbel2024estimating}, which can
introduce non-vanishing bias if ignored. We derive convergence rates
for both the spillover-parameter estimator and our estimated
adjacency matrix. Beyond these formal results, the framework also
delivers an interpretable decomposition of network connections (a
low-rank component capturing pervasive relationships and a sparse
component isolating localized interactions, with identification
following standard strategies in the literature
\citep{chandrasekaran2011rank}) and we use this decomposition to
provide a novel expression of the network eigencentrality that
quantifies the contributions of each component.
We illustrate the practical relevance of our framework through two
empirical applications: international spillovers in GDP growth across
23 economies and tax competition among 48 U.S.\ states. In both
settings, our estimated network produces spillover estimates that
differ markedly from those obtained using the raw adjacency matrix,
highlighting the empirical importance of network denoising. In the
cross-economy GDP-growth application, regularization affects not only the scale of spillovers but also the inferred structure of interdependence.
The network decomposition further reveals
distinct regional and global components of cross-economy
interdependence. These empirical findings demonstrate how our
framework can uncover the structure of interdependence in economic
networks and support more reliable inference in spatial econometric
analysis. In the
U.S.\ tax-competition application, the raw inverse-distance weight
matrices yield explosive estimates $\widehat\lambda > 1$ that violate
the Leontief stability condition, while our estimator
returns stable estimates.
The remainder of our paper is organized as follows.
Section~\ref{Sec:Motivation} motivates the framework by explaining
why measurement error matters in dense networks, illustrating four
possible network structures (sparse, low-rank, low-rank-plus-sparse),
and demonstrating with the U.S.\ Input-Output Table that real-world
dense networks exhibit exploitable regularization structure.
Section~\ref{Sec:Methodology} presents the spatial interaction model
and introduces the plug-in and regularized GMM estimators under
low-rank, sparse, or low-rank-plus-sparse structure.
Section~\ref{Sec:theory} provides assumptions and establishes
consistency and asymptotic normality results for both estimators.
Section~\ref{Sec:Simulation} presents simulation evidence.
Section~\ref{Sec:Application} applies our methods to estimate
spillover effects using noisy network data.
Section~\ref{Sec:Conclusion} concludes. Proofs, algorithms, and
additional technical details are provided in the appendices.
}
\section{Motivation} \label{Sec:Motivation}
This section sets up the notation and previews the regularization
framework used throughout the paper. Let $\mathcal{N} \coloneqq
\{1, \ldots, n\}$ denote the set of $n$ units or nodes and
$\mathcal{E}$ the set of edges between them, so that
$(\mathcal{N}, \mathcal{E})$ represents the graph. Let $W_0$ be the
$n \times n$ adjacency matrix encoding the true connections, with
$(i,j)$-th element $W_{0,ij}$; we write $W_{0,ij} \neq 0$ when unit
$i$ is connected to unit $j$ and $W_{0,ij} = 0$ otherwise.\footnote{We
allow the connection from $i$ to $j$ to differ from that from $j$
to $i$, so the graph may be directed and $W_0$ need not be symmetric.}
We further assume no self-loops, so that $W_{0,ii} = 0$ for all
$i = 1, \ldots, n$.
The econometrician observes a noisy version of $W_0$,
\[
W = W_0 + E,
\]
where $E$ is a measurement-error matrix whose sources are discussed
in Section~\ref{sec:densemeasurement}. Before formally motivating
the regularization framework, we briefly preview the structural
decomposition of $W_0$ that plays a central role throughout the paper.
To reduce estimation complexity for large $n$, it is natural to
impose structural assumptions on $W_0$. The most common choice (sparsity) is plausible for social networks dominated by a few
strong links, but fails to describe networks in which a large
number of units share common interests or characteristics: a
highly sparse network typically splits into several disconnected
components. We therefore propose to model $W_0$ as the sum of a
low-rank and a sparse component,
\begin{equation}
\label{Eq:Std_Matrix_Structure1}
W = \underbrace{L_0 + S_0}_{\equiv W_0} + E,
\end{equation}
where $L_0$ captures pervasive, common connections and $S_0$
captures sporadic or idiosyncratic connections.
The components $L_0$ and $S_0$ need not individually satisfy the
zero-diagonal constraint. When $W_0$ has a zero diagonal, we can
define $L_0^* = L_0 - \operatorname{diag}(L_0)$ and $S_0^* = S_0 + \operatorname{diag}(L_0)$,
which absorb the diagonal corrections so that both represent
networks without self-loops:
\begin{equation}
\label{Eq:Std_Matrix_Structure2}
W = \underbrace{L_0^* + S_0^*}_{\equiv W_0} + E.
\end{equation}
While $L_0^*$ is not exactly low-rank, it inherits the low-rank
structure of $L_0$ and differs only on the diagonal.
Section~\ref{set:Iden} illustrates the usefulness of this decomposition through
concrete examples. We further develop our theoretical analysis (Section \ref{Sec:theory}) in terms of $L_0$ and $S_0$ and
note that the results extend directly to $L_0^*$ and $S_0^*$. Both theoretical and empirical evidence in
the following sections show that this decomposition captures
several meaningful network structures encountered in practice
\citep{Diebold2014, Barranca2015, Gamble2016, newman2010networks}.
\subsection{Why dense measurement errors matter}
\label{sec:densemeasurement}
A central feature of the applications mentioned in the introduction
is that the observed network is dense and measured with error.
We document four such sources of error in Section~\ref{Sec:Sources};
here we develop the consequences for spillover estimation.
The existing literature, such as \citet{Lewbeletal2024EJ}, assumes
that measurement error in the network is sufficiently sparse that
it does not affect the consistency of the spillover estimator. In
particular, consistency of $\widehat{\lambda}$ can be preserved
when the aggregate magnitude of the error vanishes asymptotically,
for example when
\[
\frac{1}{n}\sum_{i,j} |E_{ij}| = o_p(1).
\]
This condition is naturally satisfied when the total error budget
grows slowly relative to the size of the network, as in settings
where the underlying adjacency matrix itself is sparse.
Intuitively, when only a small fraction of links are contaminated
\emph{and the magnitude of each contamination is small}, the
resulting perturbation to the network propagation mechanism
becomes negligible in large samples.
The situation is fundamentally different when measurement errors
are \emph{dense}. Even when each entry of $E$ is elementwise small,
the errors accumulate across the $O(n^2)$ entries of the matrix
and across the propagation paths generated by the spatial
multiplier. The aggregate distortion may not vanish asymptotically, and the
estimator $\widehat{\lambda}_{\mathrm{GMM}}$ may fail to be consistent.
To recover $W_0$ from $W$ in the dense regime, some structural
restriction on $W_0$ is needed. Depending on the application,
the appropriate structure is:
\begin{itemize}
\item[i)] \textbf{Sparse:} $W_0 = S_0^*$, where most links are
absent. Suitable for social networks.
\item[ii)] \textbf{Low-rank:} $W_0 = L_0^*$, where connections
arise from a small number of common factors or shared
institutional features. The primary case for dense networks
such as input--output tables and financial exposure matrices.
\item[iii)] \textbf{Low-rank plus sparse:} $W_0 = L_0^* + S_0^*$,
the general case combining pervasive common connections with
idiosyncratic strong links.
\end{itemize}
Misspecifying this structure leaves residual bias: assuming
$W_0 = S_0^*$ when a non-negligible $L_0^*$ is present discards the
pervasive systematic connections that drive spillovers; assuming
$W_0 = L_0^*$ when strong idiosyncratic links are present absorbs
those links into the residual. Moreover, the low-rank and sparse decomposition introduced above is the device we use to
operationalize this regularization. We emphasize that it is a
flexible modelling device, not an economic interpretation. The
interactions-model estimator $\widehat\theta_p$ depends only on
$\widehat W = \widehat L^* + \widehat S^*$, not on the individual
components. We thus do not rely on any economic interpretation of
$L_0^*$ and $S_0^*$ as separate network objects.
Nevertheless, Appendix~\ref{Appendix:decomp} discusses the economic
interpretation of $L_0^*$ and provides identification conditions
for the $L_0^* + S_0^*$ decomposition.
The practical importance of regularizing $W_0$ rests on three
observations. First, measurement error in dense networks is
\emph{non-negligible}: it may not vanish with sample size, so
standard asymptotic consistency arguments fail. Second, it is
\emph{endogenous}: errors share common sources and are typically
correlated with the true network entries, violating the classical
measurement-error assumptions that justify standard
instrumental-variable corrections. Third, it is \emph{economically
consequential}: in our empirical illustration with the U.S.\
Input--Output Table (Section~\ref{empirical}),
ignoring measurement error entirely distorts the estimated
propagation matrix by $9\%$, and misspecifying the structure of
the error (treating it as purely sparse) compounds the distortion
to $68\%$. A method that filters dense, structured noise while
preserving the signal in $W_0$ is therefore not a refinement but a
prerequisite for credible inference about network spillovers.
\subsection{Sources of measurement error in economic networks}\label{Sec:Sources}
The economic literature documents several mechanisms through which
the observed adjacency matrix $W$ may depart from the latent network
$W_0$. Many policy-relevant production, trade, and reconstructed
financial networks contain numerous weak links and can be dense
relative to conventional social networks. Consequently, measurement
error may affect a large fraction of their entries. A related source
of potentially dense measurement error arises when the adjacency
matrix is constructed from geographical or economic distance
\citep{LeSage2009, Getis2004}: errors in the underlying distance
measurem (arising, for example, from centroid approximations, omitted
transport links, or bandwidth choices) can propagate across entire
rows or columns of $W$. We discuss several additional sources of
measurement error below.
\subsubsection*{(i) Aggregation and reporting error in input--output networks}
A widely studied dense economic network is the input--output table, where
the $(i,j)$ entry records the share of commodity $i$ used as an intermediate
input in producing commodity $j$. \citet{Acemoglu2012} establish that
propagation multipliers
are highly sensitive to
the precise pattern of bilateral input shares across all $n^2$ entries of
this matrix. The observed Bureau of Economic Analysis table aggregates
transactions across hundreds of heterogeneous firms and products into
sector-level flows, and is further subject to reporting inaccuracies,
imputation of missing entries, periodic revisions, and reconciliation
between the supply and use tables. Given these, the analyst observes a sector-level $W$ that departs from any
meaningful sector-level $W_0$, and the resulting errors propagate along
multi-step Leontief chains.\footnote{The literature on aggregation in
production networks treats this as a separate methodological problem;
see e.g.\ \citet{barrot2016input} and \citet{boehm2019input}. Our framework
is complementary: it addresses reporting and revision error at whatever
level the network is observed, and is silent on the firm-to-sector
aggregation step.} Our method takes the sector-level network as primitive
and denoises around a sector-level truth. We show in Section~\ref{Sec:Application} that
even at the sector level, denoising restores Leontief stability and
materially changes estimated cross-sector multipliers.
\subsubsection*{(ii) Mirror-statistic discrepancies in bilateral trade networks}
International trade flows provide a second canonical example of a dense network where measurement error affects every bilateral entry. Trade flows are recorded twice: by the exporting country as exports and by the importing country as imports. These two independent records of the same physical shipment should agree, but in practice they differ substantially and systematically. \citet{FismanWei2004} exploit this bilateral mirror-statistic structure to measure the \emph{evasion gap}, i.e., the log difference between Hong Kong's reported exports to China and China's reported imports from Hong Kong at the six-digit product level and show that a one-percentage-point increase in the tariff rate is associated with a three-percent increase in the evasion gap.
This tariff-related underreporting may affect many product categories
and can be amplified by product misclassification, as importers
reclassify high-tariff goods into lower-tariff categories. The
resulting measurement error need not be confined to a small number
of trade links and may affect a substantial fraction of the bilateral
trade network in a manner correlated with tariff rates.
Because tariff schedules are systematically related to trade volumes
and revealed comparative advantage, the measurement error $E_{ij}$
tends to correlate with the true $W_{0,ij}$, violating classical
errors-in-variables assumptions and biasing estimates of
trade-network spillovers.
\subsubsection*{(iii) Entropy-based reconstruction error in interbank exposure networks}
Bilateral interbank exposures form a naturally dense network for studying systemic risk and financial contagion. Each bank's balance sheet records its total interbank lending and borrowing, but the bilateral breakdown, who lent to whom and in what amount, is typically confidential. Regulatory data under the Basel II framework require reporting of exposures only above a threshold of 10\% of regulatory capital, so the majority of bilateral links are unobserved \citep{Anand2015}. Researchers and regulators therefore reconstruct the full $n \times n$ bilateral exposure matrix from the observable marginals (total lending and borrowing of each bank) using maximum-entropy algorithms. As \citet{Upper2011} and the systematic evaluation of \citet{Anand2015} document, the maximum-entropy solution fills in the matrix as uniformly as possible, implicitly treating the network as if every bank lent to every other bank in proportion to its total activity. This assumption yields a fully dense reconstructed matrix, but one that systematically underestimates concentration and overestimates diversification: in reality, interbank lending is relationship-based and concentrated among a small set of counterparties. The reconstruction error $E = W - W_0$ is may therefore be dense. It affects every entry of the matrix, and is potentially correlated with the true bilateral exposures, since banks with large total lending are disproportionately misrepresented. Stress-testing exercises that rely on the reconstructed network underestimate contagion risk precisely because the entropy assumption smooths away the concentrated exposures that drive cascading failure.
\subsubsection*{(iv) Recall error and endogeneity in reported network links}
Even when network data are collected via direct reports rather than algorithmic reconstruction, the observed adjacency matrix is subject to recall error, misunderstanding, and measurement timing mismatches. \citet{Lewbeletal2024EJ} study economic network models in which some reported binary links are misclassified: true connections are recorded as absent (false negatives) and absent connections are recorded as present (false positives). They establish that this misclassification introduces new sources of endogeneity beyond the standard simultaneity problem in spatial models: the mismeasured adjacency matrix contaminates the spatial lag $WY$, creating correlation between regressors and structural errors, and invalidating standard instrumental variables based on powers of $W$. The result is that conventional Two-Stage Least Squares (2SLS) estimators that ignore misclassification are inconsistent even when the fraction of misclassified links is small relative to $n$, because in a dense network a small misclassification rate still corresponds to a large absolute number of erroneous entries. Panel data settings compound this problem: if the network is observed at lower frequency than the outcome variable, links created or dissolved between survey waves appear as persistent misclassifications correlated with individual fixed effects that cannot be differenced away \citep{lewbel2024estimating}.
The four mechanisms above are heterogeneous in origin but share the
three features highlighted at the end of Section~\ref{sec:densemeasurement}:
the measurement error is dense, correlated across entries through
shared sources (aggregation, tariff incentives, reconstruction
algorithm, or recall bias), and asymptotically non-negligible. These
properties together explain why classical errors-in-variables
corrections may be insufficient: they presume sparse or vanishing
noise that can be handled by instrumental variable (IV) or
measurement-error bias corrections, whereas the dense, correlated
errors documented here require a different approach. Imposing a
regularization on $W_0$ (as a low-rank matrix, a sparse matrix,
or their combination) provides the identifying restrictions needed
to disentangle the systematic signal from the noise, motivating the
framework developed in the remainder of the paper.
\subsection{Examples of network structures}\label{set:Iden}
To build intuition, we describe four common economic network structures and show how each fits into the low-rank, sparse, or low-rank plus sparse framework.
For simplicity, throughout this subsection, we assume
the network is observed without measurement error, so $W = W_0$.
\footnote{Note that the link matrix or adjacency matrix could be used to describe the connection strength between individuals. If it does not satisfy the bounded row sum assumption, it could be further scaled by its $\ell_1$ norm.}
\subsubsection{Complete network}
We begin by considering one of the simplest examples: a fully connected network. In this setup, the link matrix $W_0$ is an $n \times n$ matrix of nonzero values, where it is common to set $W_{0, ii} = 0$ to avoid self-loops.
The link matrix corresponding to a complete network generated by nonzero constants $\bm{c}=(c_1, \cdots, c_n)^{\top}$ is defined as
\begin{equation*}
W_0 = \begin{bmatrix}
0 & c_{1}c_2 & c_{1}c_3 &\cdots & c_{1}c_n \\
c_{2}c_1 & 0 & c_{2}c_3 &\cdots & c_2c_n \\
\vdots & \vdots & \vdots & \ddots & \vdots \\
c_nc_1 & c_nc_2 & c_nc_3 & \cdots & 0
\end{bmatrix} ,
\end{equation*}
{where $c_i$'s are positive constants within the bounded support $[c_{\min}, c_{\max}]$, and they could be viewed as individual characteristics that vary across $i$. Let $L_0= \bm{c} \bm{c}^{\top}$, which is a symmetric low-rank matrix with rank $1$. Its singular value is $\|\bm{c}\|_2^2$, and the associated singular vector is
$\bm{c} /\|\bm{c}\|_2$. Then we could also express $W_0=L_0+S_0=L^*_0$, where the sparse component $S_0$ is a diagonal matrix whose $(i,i)$th element is $- c_{i}^2$ to guarantee no self-loops in $W_0$. Such a network could be used to capture the common interaction structure driven by individuals or agents who interact through a common market or platform. }
\subsubsection{Low-rank network}
Motivated by the structure of the latent factor, we consider the symmetric network in which the strength of the connection is determined by $W_{0,ij}=a_i+z_iz_j$ if $i\neq j$, { where $a_{i}$'s are individual characteristics and $z_i$'s are some latent variables. \footnote{In a more general term, the low rank component can be understood as network effects related to a principal agent, see \citet{Galeotti2020} . The principal agent is formed by connections with respect to many nodes within the network, and the effects are dense.}
Denote $\bm a$ as a vector consisting of $a_i$. In this example, $W_0=L_0+S_0=L_0^*$, where $L_0={\bm a} {\bf 1}_n^\top + {\bm z} {\bm z}^\top$ is a matrix of rank no more than 2 with ${\bm a}=(a_1, \cdots, a_n)$ and ${\bm z}=(z_1, \cdots, z_n)$, and $S_0$ is a diagonal matrix whose $(i,i)$th element is $-L_{0,ii}$ to remove self-loops. Networks of this kind arise in production networks and financial systems, where common shocks or institutional linkages generate pervasive, dense interactions.}
\subsubsection{Group network}
Another example of interest is the pure group network, where there are two or more communities, potentially of distinct sizes to avoid eigenvalue multiplicity in the link matrix. Moreover, individuals within the same community are assumed to be connected with each other, while individuals do not connect to others outside their community.
For example, suppose $n=20$, and $6$ individuals belong to community $1$ and the rest belong to community $2$. We define
\begin{equation*}
W_0 = \begin{bmatrix}
L_1 & \bm{0}_{6 \times 14} \\
\bm{0}_{14 \times 6} & L_2
\end{bmatrix}+S_0=L_0^*\, ,
\end{equation*}
where $L_1=(c_1, c_2, \cdots, c_6)^{\top}(c_1, c_2, \cdots c_6)$ and $L_2=(c_7, \cdots, c_{20})^{\top}(c_7, \cdots, c_{20})$,
and $S_0$ is a diagonal matrix such that $S_{0,ii}=-c_i^2$.
Note that $W_0$ still admits the low-rank and sparse decomposition. The low-rank component admits a block structure and its rank equals
the number of communities. One of its two singular vectors
is proportional to $(0, \cdots, 0, c_7, \cdots, c_{20})$, while the other is
proportional to $(c_1, \cdots, c_6, 0, \cdots, 0)$.
See Figure \ref{fig:group} for the network graph and the heatmap of $W_0$.
{This setting captures trade blocs, regional banking clusters, and sectoral linkages, and the rank of $L_0$ equals the number of communities. We could also extend it to allow sporadic between-group connections, which are described by few nonzero off-diagonal elements in $S_0$. Then the general model could be written as $W_0=L_0+S_0=L_0^*+S_0^*$, where the low-rank component captures the within-group connections and the sparse component represents the between-group connections.}
\subsubsection{Dominant units}
\label{Subsec:Dominant_Units_Simulation}
An interesting example comes from considering networks containing a set of dominant units also known as key players, see \citet{CalvoArmengol2009, Zenou2016}. One definition of dominant units used in the network literature identifies them as the set of units whose actions have large and persistent effects on the other agents.\footnote{We follow the definition of dominant units as in \citet{Pesaran2020, Pesaran2021, Lee2022}. }
One way to measure the effect of the unit $j$ on others is to use its weighted out-degree $d_j$,
defined simply by the $j$-th column sum of $W_0$:
$d_j = \sum_{i = 1}^n W_{0,ij}$.
We then call the unit $j$ a dominant unit if its weighted out-degree is large or grows as the sample size does.
For example, suppose $n=20$ and there is only one dominant unit, namely the first individual, who is connected to $13$ individuals.
We assume that each nondominant unit is also connected to its
immediate neighbors. Then the link matrix $W_0$ is specified as below:
\begin{equation*}
W_0 = \begin{bmatrix}
0 & w_{1,2} & 0 & \cdots & \cdots & \cdots &\cdots&\cdots& 0 \\
w_{2,1} & 0 & w_{2,3} & \cdots & \cdots &\cdots&\cdots& \cdots & 0 \\
\vdots & \vdots & \vdots & \vdots & \ddots & \vdots & \vdots &\vdots &\vdots\\
w_{14,1} & 0 & \cdots & w_{14,13} & 0 & w_{14,15} & \cdots & 0 & 0 \\
0 & \cdots & \cdots & 0 & w_{15,14} & 0 & w_{15,16} & \cdots & 0 \\
\vdots & \vdots & \vdots & \vdots & \vdots & \vdots & \vdots &\vdots &\vdots\\
0 & 0 & 0 & 0 & \cdots &\cdots&\cdots& w_{20,19} & 0 \\
\end{bmatrix} \, .
\end{equation*}
See Figure \ref{fig:dominant} for the network graph and the heat map of $W_0$.
This network is naturally sparse. The number of nonzero entries in
each row is bounded, although the column corresponding to the dominant
unit may contain a growing number of nonzero entries. While this
column can be represented as a sparse rank-one matrix, the full
adjacency matrix need not be low-rank because it also contains links
among neighboring units. We therefore treat $W_0$ as a sparse matrix.
Appendix~\ref{Appendix:iden} provides further discussion of alternative
low-rank-plus-sparse representations and their identification.
\begin{figure}[htbp]
\centering
\caption{Graphical representations of simulated adjacency matrices associated with different groups}
\begin{subfigure}{0.4\textwidth}
\includegraphics[width = \textwidth]{GroupWnet.pdf}
\end{subfigure}
\hfill
\begin{subfigure}{0.4\textwidth}
\includegraphics[width = \textwidth]{GroupWh.pdf}
\end{subfigure}
\hfill
\caption*{\footnotesize Notes: Examples for the weight matrix formed by individuals who belong to different groups. The left one is the network graph and the right one is the heat map of the adjacency matrix. There are two groups, one of which has $6$ members and the other has $14$ members. Only members within the same group have connections, and the connection strength is randomly simulated. }
\label{fig:group}
\end{figure}
\begin{figure}[htbp]
\centering
\caption{Graphical representations of the simulated adjacency matrix with dominant units. }
\begin{subfigure}{0.4\textwidth}
\includegraphics[width = \textwidth]{DominantWnet.pdf}
\end{subfigure}
\hfill
\begin{subfigure}{0.4\textwidth}
\includegraphics[width = \textwidth]{DominantWh.pdf}
\end{subfigure}
\hfill
\caption*{\footnotesize Notes: Examples for the weight matrix with dominant units. The left one is the network graph and the right one is the heat map of the adjacency matrix. There are $20$ individuals and each individual connects to its neighbors with strength set by $0.25$. The first one is a key player that influences 13 other individuals, and the influence strengths are generated by $\text{Uniform}(0, 1)$ distributed random variables. }
\label{fig:dominant}
\end{figure}
{\subsection{Empirical illustration: the US Input-Output Table}
\label{empirical}
To demonstrate that measurement error in dense economic networks has economically
meaningful consequences, and that the $L+S$ decomposition provides an effective device for
recovering the underlying true network structure, we revisit the 2002 US Input-Output
table from \cite{graham2020econometric}, compiled by the Bureau of Economic Analysis
and studied by \cite{Acemoglu2012}. It contains an $n \times n$ network of $n=417$
sectors (after dropping rows and columns with zero direct input requirements), where
entry $(i,j)$ indicates the share of sector $i$'s output used in sector's $j$ production.
This network is dense by construction(nearly all sectors interact) making it a
canonical example where sparsity cannot be assumed and measurement error is pervasive.
We treat the observed $W$ as a noisy version of the true network $W_0$ and apply our
Algorithm~\ref{Algo:CD_LRSED} (Appendix~\ref{algo}) to recover $W_0$.
To assess the quality of the denoising, we compare three candidate representations of
$W_0$, a purely sparse estimate $\widehat{S}$, a purely low-rank estimate $\widehat{L}$,
and the combined version $\widehat{W} = \widehat{L}+\widehat{S}$. The purely
sparse estimate satisfies $\|\widehat{S}-W\|_{\max}=0.025$ and
$\|\widehat{S}-W\|_2=0.470$, indicating that the tiny elements might have non-negligible aggregation effects.
The purely low-rank estimate satisfies
$\|\widehat{L}-W\|_{\max}=0.999$ and $\|\widehat{L}-W\|_2=1.214$, indicating that the low-rank
structure alone might fail to capture occasionally large idiosyncratic links. The combined
estimate achieves $\|\widehat{W}-W\|_{\max}=0.024$ and $\|\widehat{W}-W\|_2=0.183$,
materially outperforming either alternative. This superiority is further confirmed by
eigenvector alignment. The inner product between the leading eigenvectors of $W$ and
$\widehat{W}$ is $0.981$, compared with $0.526$ for $\widehat{S}$ and $0.222$ for
$\widehat{L}$, indicating that the low-rank and sparse decomposition might best preserve the spectral structure
of the observed network. From Figure~\ref{heatmap}, we could see that
the sparse component $\widehat{S}$ captures large idiosyncratic bilateral
links, while the low-rank component $\widehat{L}$ exhibits a distinctive stripe
pattern, revealing that pervasive but individually weak linkages
aggregate into systematic structure that a purely sparse representation would miss.
To quantify the econometric consequences of measurement error, we compare the
Leontief diffusion matrices $(I_n - \lambda \widehat{W})^{-1}$,
$(I_n - \lambda W)^{-1}$, and $(I_n - \lambda \widehat{S})^{-1}$ following
\cite{acemoglu2016networks} with $\lambda = 0.7$. Leaving measurement errors uncorrected
produces a $9\%$ deviation in the spectral norm relative to our estimators:
\[
\frac{\|(I_n-0.7\widehat{W})^{-1}-(I_n-0.7W)^{-1}\|_2}
{\|(I_n-0.7\widehat{W})^{-1}\|_2}=9\%.
\]
Misspecifying the structure of measurement errors by treating it as a purely sparse
discarding the low-rank
component compounds the distortion substantially, producing a
$68\%$ deviation:
\[
\frac{\|(I_n-0.7\widehat{W})^{-1}-(I_n-0.7\widehat{S})^{-1}\|_2}
{\|(I_n-0.7\widehat{W})^{-1}\|_2}=68\%.
\]
These results demonstrate that measurement error in dense networks is not merely a theoretical concern: it materially distorts estimated propagation channels. Hence it is essential for reliable inference to correctly specify the error structure accounting for both its sparse idiosyncratic component and its low-rank systematic component. Figure~\ref{heatmap2} illustrates this point visually: the stripe pattern present in
the full diffusion matrix $(I_n - 0.7\widehat{W})^{-1}$ is eliminated when the
low-rank component is omitted, confirming that misspecification of network structure
biases spillover estimates in economically meaningful ways.
\begin{figure}[htbp]
\centering
\begin{subfigure}{0.3\textwidth}
\includegraphics[width=\textwidth]{sW.pdf}
\end{subfigure}
\hfill
\begin{subfigure}{0.3\textwidth}
\includegraphics[width=\textwidth]{sS.pdf}
\end{subfigure}
\hfill
\begin{subfigure}{0.3\textwidth} \includegraphics[width=\textwidth]{sL.pdf}
\end{subfigure}
\caption{\footnotesize Heatmaps of 20 sectors from the 2002 Input-Output Table $W$ (left),
estimated sparse component $\widehat{S}$ (centre), and estimated low-rank component
$\widehat{L}$ (right). The stripe pattern in $\widehat{L}$ reveals pervasive
systematic structure missed by a purely sparse
representation.\label{heatmap}}
\end{figure}
\begin{figure}[htbp]
\centering
\begin{subfigure}{0.3\textwidth}
\includegraphics[width=\textwidth]{iW.pdf}
\end{subfigure}
\hfill
\begin{subfigure}{0.3\textwidth}
\includegraphics[width=\textwidth]{iS.pdf}
\end{subfigure}
\hfill
\begin{subfigure}{0.3\textwidth}
\includegraphics[width=\textwidth]{iL.pdf}
\end{subfigure}
\caption{\footnotesize Heatmaps of 20 sectors from the diffusion matrix
$(I_n-0.7\widehat{W})^{-1}$ (left), $(I_n-0.7\widehat{S})^{-1}$ (centre), and
the difference (right). Omitting the low-rank component eliminates the stripe pattern
and distorts estimated spillover propagation by $68\%$.\label{heatmap2}}
\end{figure}
}
\section{Methodology} \label{Sec:Methodology}
We present the general model in which the latent network $W_0 = L_0 + S_0$
admits both a low-rank and a sparse component, with the purely low-rank and purely sparse settings arising as special cases.
We emphasize that this decomposition need not be unique or economically interpretable on its own; rather, the joint regularization imposed on $L_0$ and $S_0$ provides sufficient structure to separate $W_0$ from the noise $E$. We provide the modeling definitions and estimation procedures we propose to tackle this problem. We offer additional assumptions and asymptotic results in Section~\ref{Sec:theory}.
\subsection{Estimation}\label{Sec:Estimators}
We formally describe here the social/spatial interaction model of interest. We consider a panel data scenario, but
assume the true underlying network does not vary with time. Denote data points as $\{Y_t, X_t\}_{t=1, \cdots, T}$, where $Y_t$ is the $n \times 1$ outcome vector and $X_t$ is the $n \times K$ design matrix of exogenous covariates. We assume that the true model is generated by the unobserved $n \times n$ adjacency matrix $W_0$ encoding the true connection, and that we have a noisy or contaminated observation, such that we can write
\begin{eqnarray}
Y_t &=& \lambda_0 W_0 Y_t + \alpha + \iota_t {\bf 1}_{n\times 1} + X_t\beta_0 + W_0 X_t \gamma_0 + \varepsilon_t \nonumber \\
W &=& W_0 + E, \label{maineq}
\end{eqnarray}
where $\lambda_0$ represents the scalar (social / spatial) interaction or spillover effect, $\beta_0 \in \mathbb{R}^K$ are standard slope coefficients, $\gamma_0 \in \mathbb{R}^K$ are contextual effects (spillover effects of the observable characteristics $X$), $\alpha = (\alpha_1, \ldots, \alpha_n)^\top$ is the vector of individual fixed effects, $\iota_t$ is a time fixed effect at period $t$, ${\bf 1}_{n\times 1}$ is an $n \times 1$ vector of ones, and $\varepsilon_t$ is an $n$-dimensional vector of idiosyncratic disturbances, assumed for now to be independent and identically distributed over units $i$ and time $t$. We assume the covariates $X_t$ to be exogenous to the outcome unobservables and network formation process such that $\operatorname{\mathbb{E}}[\varepsilon_{it} \mid X_t, W_0X_t, \cdots, W_0^dX_t] = 0$, for a small integer $d>0$, $i = 1,\cdots, n$ and $t = 1, \cdots, T$. Crucially, we allow for the measurement errors to be endogenous in the sense that $\operatorname{\mathbb{E}}[\varepsilon_{it} \mid E] \neq 0$,
which introduces additional bias into conventional estimates when $W_0$ is replaced by $W$ \citep{boucher2025estimating, hardy2019estimating, Lewbeletal2024EJ}.
Collect all parameters of interest into $\theta_0 \coloneqq (\lambda_0, \beta^{^{\top}}_0, \gamma^{^{\top}}_0)^{^{\top}} \in \mathbb{R}^{2K+1}$. When covariates $X_t$ are exogenous and the adjacency matrix $W_0$ is correctly observed,
the literature proposes to select $m$ linearly independent columns
from $[X_t, W_0 X_t, \ldots, W_0^d X_t]$, for a small positive
integer $d$, as instruments for the regressors
$\bar{X}_t(W_0) = [W_0 Y_t, X_t, W_0 X_t]$ in estimating $\theta_0$
\citep{Kelejian1998, Lee2003, Lee2007}.
The degree $d$ should be chosen such that the number of retained instruments
is at least $2K+1$ to satisfy the order condition, i.e. $2K+1 \leq m \leq (d+1)K$ (see Assumption \ref{Assump:Instruments}). For an arbitrarily pre-specified $n \times n$ adjacency matrix $M$, let $Z_t(M)$ be the $n \times m$ matrix that collects the linearly independent columns from $[X_t, M X_t, \ldots, M^d X_t]$ and let ${X}_t(W)$ be the $n \times (2K+1)$ matrix that collects the regressors $[W Y_t, X_t, W X_t]$. Note that ${X}_t(W)$ uses the observed adjacency matrix $W$ as opposed to the true $W_0$ showing up in model \eqref{maineq}.
By defining $\varepsilon_t(\theta, W) \coloneqq Y_t - {X}_t(W) \theta$ and thus $\varepsilon_t =\varepsilon_t(\theta_0, W_0)$ , we obtain the population moment condition $\operatorname{\mathbb{E}}[Z_t^{\top}(M)\varepsilon_{t}] = \mathbf{0}_{m\times 1}$ for any exogenous $M$.
Thus we have the empirical moment conditions:
\begin{align}
\label{Eq:Moment_Conditions}
\bar{g}_{nT}(\theta, W, M) & \coloneqq \frac{1}{nT} \sum_{t=1}^{T} Z_t(M) ^{\top} \varepsilon_t(\theta, W) \, .
\end{align}
The distinction made between the matrix used to construct the regressors ${X}_t(W)$ and the instruments $Z_t(M)$ is necessary to define
our supervised estimator below. When $M=W$ such that no distinction is necessary, we simply write $\bar{g}_{nT}(\theta, W) = \bar{g}_{nT}(\theta, W, W)$. For a given $m \times m$ moment-weighting matrix $\widehat{\Lambda}_{nT}^{-1}$, we define the sample GMM objective function as
\begin{equation*}
J_{nT}(\theta, W, M; \widehat{\Lambda}_{nT}^{-1}) \coloneqq (nT)\bar{g}_{nT}(\theta, W, M)^{\top} \widehat{\Lambda}_{nT}^{-1} \, \bar{g}_{nT}(\theta, W, M) .
\end{equation*}
When fixed effects are present, within transformation {\color{red}\footnote{The within transformation eliminates the
individual effects $\alpha$ by subtracting unit-specific time
averages: writing $Y_{\cdot} \coloneqq \frac{1}{T}\sum_{t=1}^{T} Y_t$,
the transformed outcome is $\ddot{Y}_t \coloneqq Y_t - Y_{\cdot}$, and
$X_t(W)$ and $Z_t(M)$ are transformed analogously. Since $\alpha$ is
constant over $t$, it is differenced out by this demeaning; the time
effects $\iota_t$ are removed by the additional cross-sectional
demeaning discussed in Remark~\ref{rem:twoway_demean}. The moment conditions
\eqref{Eq:Moment_Conditions} constructed from the demeaned variables
remain valid.}} is applied
to $Y_t$, $X_t(W)$, and $Z_t(M)$ before constructing the GMM objective;
we suppress this transformation in the notation that follows. The
extension to two-way fixed effects is discussed in
Remark~\ref{rem:twoway_demean}.
Finally, we stack all time series observations into the $nT$-dimensional outcome vector $Y \coloneqq [Y_1^{\top}, \ldots, Y_T^{\top}]^{\top}$, the $nT \times (2K+1)$ design matrix ${X}(W) \coloneqq [{X}_1(W)^{\top}, \ldots, {X}_T(W)^{\top}]^{\top}$, and the $nT \times m$ matrix of instruments ${Z}(M) \coloneqq [{Z}_1(M)^{\top}, \ldots,{Z}_T(M)^{\top}]^{\top}$. Using this notation,
we can express the GMM estimator as
\begin{eqnarray}
\label{Eq:GMM_estimator}
\widehat{\theta}_\text{GMM}(W,M) &&\coloneqq \arg \min_{\theta \in \Theta} J_{nT}(\theta, W, M; \widehat{\Lambda}_{nT}^{-1}) \\&&=\left[{X}(W)^{\top} Z(M) \widehat{\Lambda}_{nT}^{-1} Z(M)^{\top} {X}(W)\right]^{-1} {X}(W)^{\top} Z(M) \widehat{\Lambda}_{nT}^{-1} Z(M)^{\top} Y. \nonumber
\end{eqnarray}
We define $\widehat{\theta}_\text{GMM} = \widehat{\theta}_\text{GMM}(W,W)$. In practice, we can obtain a two-step GMM estimator by using the weight matrix $\widehat{\Lambda}_{nT}^{-1}$. {\color{red}\footnote{In practice, we obtain a feasible two-step GMM estimator by first computing
$\widetilde{\theta}$ using the identity weighting matrix and then using
$\widehat{\Lambda}_{nT}^{-1}$, where
\[
\widehat{\Lambda}_{nT}
=
\frac{1}{nT}\sum_{t=1}^T
Z_t(W)^\top
\operatorname{diag}\!\left(
\widehat{\varepsilon}_{1t}^{\,2},\ldots,
\widehat{\varepsilon}_{nt}^{\,2}
\right)
Z_t(W),
\qquad
\widehat{\varepsilon}_t
=
Y_t-X_t(W)\widetilde{\theta}.
\] } }representing an initial consistent estimator obtained by using the identity $\widehat{\Lambda}_{nT}^{-1} = I_m$ as the first-step GMM weight matrix.
The rates of our estimators along with all formal assumptions and theoretical results, are presented in Section \ref{Sec:theory}. If $W_0$ was observed without error, \cite{Kelejian1998} and \cite{Lee2007}, for example, guarantee that the standard GMM estimator $\widehat{\theta}_\text{GMM}$ is consistent for $\theta_0$ in \eqref{maineq}. However, when measurement error is present, small elementwise noises in $E$ can accumulate, leading to inconsistent estimation of the regression coefficients and spillover effects.
Therefore, when measurement error exists, we need a way to correct for this additional noise in our estimation. We make use of the information afforded by the low-rank plus sparse plus error decomposition as in \eqref{Eq:Std_Matrix_Structure1} to obtain an estimate $\widehat W$. Theoretically, we quantify the estimation accuracy of $\widehat W$
and its asymptotic influence when used as an input for estimating $\theta_0$. We now describe the details leading to our two proposed estimators. Our first option is the ``plug-in estimator,'' and the other is the ``supervised estimator,'' both defined in the following subsection.
\vspace{-0.2cm}
\subsection{The plug-in estimator and the supervised estimator}
\label{Subsec:Plug-in}
Our proposed estimators make use of
the structure for the adjacency $W$ in \eqref{maineq} by using penalized convex low-rank and sparse estimation.
Our first proposal is simply a plug-in estimator that first produces a denoised estimate for $W_0$ and uses this estimate in place of the observed version when calculating the GMM estimator.
Recall that $A$ is an $n\times n$ matrix and $A_{ij}$ denotes its $(i,j)$th element. We first minimize the following loss function to de-noise the adjacency matrix { by} taking advantage of its low-rank plus sparse structure:
\begin{align}
\label{Eq:LRSED}
(\widehat{L}, \widehat{S}) \coloneqq \arg \min_{L \in \mathbb{R}^{n \times n}, S \in \mathbb{R}^{n \times n}, L_{ii}+S_{ii}=0, 1\leq i\leq n} \frac{1}{2}\|W-L-S\|_F^2 + \nu_{n} \|L\|_* + \tau_{n} \sum_{i\neq j} |S_{ij}|,
\end{align}
where $\|A\|_F$ is the Frobenius norm of $A$, $\|L\|_*$ is the nuclear norm of $L$, $\nu_n$ and $\tau_n$ are positive tuning parameters. Similar to \citep{Candes2011, Cao2017}, we propose to use Algorithm \ref{Algo:CD_LRSED} to solve the optimization problem in \eqref{Eq:LRSED}.
Defining $\widehat{W} \coloneqq \widehat{L} + \widehat{S}$, our plug-in estimator is then simply given as in \eqref{Eq:GMM_estimator} {to replace} $W$ with $\widehat{W}$.
\begin{eqnarray}
&&\widehat{\theta}_{p}(\widehat W) \coloneqq \arg \min_{\theta \in \Theta} J_{nT}(\theta, \widehat{W},\widehat{W}; \widehat{\Lambda}_{nT}^{-1})\nonumber
\\&&= \left[{X}(\widehat{W})^{\top} Z(\widehat{W}) \widehat{\Lambda}_{nT}^{-1} Z(\widehat{W})^{\top} {X}(\widehat{W})\right]^{-1} {X}(\widehat{W})^{\top} Z(\widehat{W}) \widehat{\Lambda}_{nT}^{-1} Z(\widehat{W})^{\top} Y .
\label{Eq:Plugin_Estimator}
\end{eqnarray}
Our second proposal, the supervised estimator,
{minimizes the combination of the GMM objective function and
the penalized square loss based on the low-rank plus sparse decomposition}.
The estimator $ (\widehat{\theta}_{s}, \widehat{L}_s, \widehat{S}_s) $
is defined as
\begin{eqnarray}
\label{Eq:Supervised_Estimator}
&&\arg\min_{\theta \in \Theta, L \in \mathbb{R}^{n \times n}, S \in \mathbb{R}^{n \times n}, L_{ii}+S_{ii}=0, 1\leq i\leq n} \xi_{nT} J_{nT}(\theta, L + S, M; \widehat{\Lambda}_{nT}^{-1}) \nonumber
\\&& + \frac{1}{2}\|W-L-S\|_F^2 + \nu_{nT} \|L\|_* + \tau_{nT} \sum_{ i\neq j}|S_{ij}|,
\end{eqnarray}
where $\xi_{nT}$ is an additional nonnegative tuning parameter necessary to balance the
{{variations contained in the observed adjacency and the outcome variables}. {It is worth noting that \cite{cai2021network} also study a supervised method based on a purely low-rank structure for analyzing network centrality. Their approach, which focuses on the leading singular vectors, differs from ours in that it does not incorporate a spatial regression.}
A full procedure for solving \eqref{Eq:Supervised_Estimator} is given in Algorithm \ref{Algo:ADMM_GMM_FBS} in Appendix \ref{Appendix:Add_Results}. Note that, as a by-product of the algorithm, the estimator produces supervised {low-rank and sparse} decompositions
that incorporate information from the outcome and covariates.
\section{Asymptotic theory}
\label{Sec:theory}
In this section, we discuss the theoretical properties of our estimation procedures. Our main interest is to quantify the impact of using an estimated adjacency matrix on recovering both the true adjacency and the spillover parameters of interest.
We shall use the following notation: Let $A$ and $B$ be matrices of dimensions $n\times n$. For $i = 1, \ldots, n$, we use $\lambda_i(A)$ and $\sigma_i(A)$
denote the $i$-th largest eigenvalue and the $i$-th largest singular value
of the matrix $A$. Let $\lambda_{\min}(A)$ denote the minimum eigenvalue of $A$. Let $\|.\|_2$ denote the Euclidean-norm of a vector. We use $\|A\|_1$, $\|A\|_2$, $\|A\|_{\infty}$, $\|A\|_{\max}$
to denote the $l_1$-norm, $l_2$-norm, $l_{\infty}$ norm, max norm
respectively. {$|A|_{1,1} =\sum_{i=1}^n\sum_{j=1}^n |A_{ij}|$.} Let $1\leq r\leq n$ be the rank of $A$.
For a matrix $A \in \mathbb{R}^{n \times n}$, let:
\begin{align*}
&m_r(A) := \max_{1 \leq i \leq n} \sum_{j=1}^n \mathbb{I}(A_{ij} \neq 0),
&& m_c(A) := \max_{1 \leq j \leq n} \sum_{i=1}^n \mathbb{I}(A_{ij} \neq 0),\\
&m_s(A) := \sum_{i,j} \mathbb{I}(A_{ij} \neq 0),
&& \deg_{\max}(A) := \max\{m_r(A), m_c(A)\}.
\end{align*}
We use $e_i$ to denote the standard basis vector, i.e., the $i$th column of the identity matrix, and $\operatorname{vec}(A)$ to denote the column vectorization of matrix $A$. For sequences $\{a_n\}_{n = 1}^{\infty}$ and $\{b_n\}_{n = 1}^{\infty}$,
we denote $a_n\lesssim b_n$ if there exists a positive constant $C$ such that $a_n/b_n\le C$ for all $n$, and denote $a_n=o(b_n)$ (resp. $a_n \asymp b_n$) if $a_n/b_n\to 0$ or $a_n \ll b_n$ {(resp. $a_n\lesssim b_n$ and $b_n\lesssim a_n$)}.
We use $\overset{d}{\to}$ and ${\to_p}$ to indicate convergence in distribution and in probability, respectively. We use $X_n = O_p(a_n) (\lesssim_p a_n)$ to indicate $X_n/a_n$ is bounded in probability, i.e. for any $\varepsilon > 0$, there exists $M(\varepsilon) > 0$
such that $\mathbb{P}(|X_n/a_n| > M(\varepsilon) ) < \varepsilon$ for all $n \in \mathbb{N}$. We assume $\mbox{min}{(n,T)}\to \infty$ or $n \to \infty$. In what follows, $Y_t$, $X_t$, $Z_t(M)$, and $\varepsilon_t$ denote
the within-transformed versions of the model objects, with individual
fixed effects removed. The asymptotic results carry through under the
two-way transformation that additionally removes time fixed effects;
see Remark~\ref{rem:twoway_demean}.
\begin{assumption}[True weight]
\label{weight}
\begin{enumerate}[(i)]
\item (Low-rank component) $L_0$ is an $n\times n$ low-rank matrix
with $\mathrm{rank}(L_0) = r$,
$\min(m_r(L_0), m_c(L_0)) \to \infty$, and
$\|L_0\|_{\max} = O(1/\sqrt{m_r(L_0)m_c(L_0)})$.
\item (Sparse component) $S_0$ is an $n\times n$ sparse matrix
with $\|S_0\|_{\max} = O(1)$.
\item (Joint identification)
$\frac{m_s(S_0)}{m_r(L_0)m_c(L_0)} \to 0$ and
$\frac{\deg_{\max}(S_0)^2}{\min\{m_r(L_0),\,m_c(L_0)\}} \to 0$.
\end{enumerate}
\end{assumption}
Assumption~\ref{weight} collects conditions on $W_0$. Part~(i)
requires the low-rank signal to be pervasive but individually
weak: the condition $\|L_0\|_{\max} = O(1/\sqrt{m_r(L_0)m_c(L_0)})$
allows $L_0$ to be distinguished from $S_0$;
it fails when $L_0$ is spiked (as in the
dominant-units example), in which case the low-rank part is better absorbed
into $S_0$. The rank $r$ may diverge with $n$ provided the relevant rates still vanish.
Part~(ii) requires $S_0$ to be bounded and Part (iii) allows $S_0$ to
{have relatively few nonzero links compared with the dense part.}
Hence the decomposition $W_0 = L_0 + S_0$ is uniquely determined when $n$ is large, making the
rank parameters $r$ and $m_s(S_0)$ well-defined. We stress this is a
\emph{mathematical} requirement only: $\widehat\theta_p$ uses
$\widehat W = \widehat L + \widehat S$ as a single object, so we rely
on no economic interpretation of the components, and inference about
$\theta$ is robust to weak identification of the decomposition.
Part~(iii) holds trivially when $L_0 = \mathbf{0}_{n\times n}$ or
$S_0 = \mathbf{0}_{n\times n}$, and as a side implication limits the sum over individual columns
of $W_0$, yielding $\|W_0\|_1 = o(\sqrt{n})$, which is used in the
rate analysis of $\widehat\theta_p(\widehat W)$. The following
assumption imposes conditions on the measurement error $E$ in the
observed matrix $W$, required for the large-sample analysis as
$n \to \infty$.
\begin{assumption}[Noise]
\label{Assump:E_noise}
\begin{enumerate}[(i)]
\item $E$ is an $n \times n$ noise matrix such that $E_{ii}=0$ for $i=1, \cdots, n$.
\item $E$ satisfies $\|E\|_{2}\leq w_{2n}$ with probability greater than $1-p_{n}$, where $p_{n}=o(1)$.
\item $E$ satisfies $\|E\|_{\max}\leq w_{\max n}$ with probability greater than $1-p_{n}$, where $p_{n}=o(1)$.
\end{enumerate}
\end{assumption}
Assumption~\ref{Assump:E_noise} is mild. It does not require the
entries of $E$ to be independent or identically distributed (a
useful feature in spatial and social network contexts), where
measurement errors often exhibit dependence. It also does not
require $w_{2n}$ or $w_{\max n}$ to vanish; these are sequences
that bound the spectral and entrywise norms of $E$, and may grow
with $n$. Whether the bounds delivered by
Theorems~\ref{Thm:DK1}--\ref{Thm:DK22} are informative depends on
joint conditions involving $(w_{2n}, w_{\max n}, r, m_s(S_0), n)$,
discussed after each theorem.
The roles of $w_{2n}$ and $w_{\max n}$ are complementary. The
spectral norm bound $w_{2n}$ controls the aggregate noise level and
governs how well the low-rank component $L_0$ can be separated
from $E$. The entrywise maximum $w_{\max n}$ controls the largest
individual perturbation and governs the recovery of $S_0$. We illustrate how $E$ looks like in the following examples.
\begin{example}[A sample-covariance error matrix]
\label{ex:covariance}
Suppose the error arises from estimating a covariance matrix: let
$z_1, \ldots, z_T$ be independent mean-zero Gaussian vectors in
$\mathbb{R}^n$ with covariance $W_z$, whose eigenvalues are positive and bounded away from 0 and $\infty$. Let
$\widetilde E = T^{-1}\sum_{t=1}^T z_t z_t^\top - W_z$ be the
sampling error. To respect the no-self-loop convention we take
$E = \widetilde E - \mathrm{diag}(\widetilde E)$, so
$E_{ii} = 0$ and Assumption~\ref{Assump:E_noise}(i) holds. Since
$\|\mathrm{diag}(\widetilde E)\|_2 = \max_i|\widetilde E_{ii}| =
O_p(\sqrt{\log n/T})$ is dominated by $\|\widetilde E\|_2$, removing
the diagonal leaves the rates unchanged.
By Theorem~6.5 of \citet{wainwright2019high}, taking
$\eta=\sqrt{n/T}$ gives
$\|E\|_2\leq C\|W_z\|_2\{\sqrt{n/T}+n/T\}=:w_{2n}$
with probability $\geq1-c_1\exp(-c_2n)$; taking
$\eta=\sqrt{\log(n)/T}$ and applying a union bound gives
$\|E\|_{\max}\leq
C(\max_iW_{z,ii})\{\sqrt{\log(n)/T}+\log(n)/T\}
=:w_{\max n}$
with probability $\geq1-2n^{-c_3}$. Thus,
Assumption~\ref{Assump:E_noise} holds with
$w_{2n}\asymp\sqrt{n/T}+n/T$ and
$w_{\max n}\asymp\sqrt{\log(n)/T}+\log(n)/T$.
If $n/T\to0$, these rates simplify to
$w_{2n}\asymp\sqrt{n/T}$ and
$w_{\max n}\asymp\sqrt{\log(n)/T}$, and both vanish, with the
spectral bound dominating the entrywise one --- the aggregate noise
exceeds the worst individual entry, as is typical in dense networks.
$\blacksquare$
\end{example}
\begin{example}[Dense Gaussian noise: comparison with \citet{Lewbeletal2024EJ}]
\label{lewbel:e}
Suppose $E_{ij}$ are i.i.d.\ $\mathbb{N}(0, \sigma^2_{n})$ for $i \neq j$
and $E_{ii} = 0$, with $\sigma_n$ possibly diminishing in $n$.
As in Example~\ref{ex:covariance}, applying the operator-norm bound
of \citet{wainwright2019high} and a union bound over the entries
gives, with probability approaching $1$,
$\|E\|_2 \leq 4\sigma_n\sqrt n =: w_{2n}$ and
$\|E\|_{\max} \leq 8\sigma_n\sqrt{\log n} =: w_{\max n}$.
Assumption~\ref{Assump:E_noise} holds
with $w_{\max n} = 8\sigma_n\sqrt{\log n}$ and $w_{\max n}\to 0$ whenever
$\sigma_n = o(1/\sqrt{\log n})$.
By contrast, $\mathbb{E}|E|_{1,1} \asymp n^2 \sigma_n$. The
condition $|E|_{1,1} = o_p(n)$ used by \citet{Lewbeletal2024EJ}
requires $\sigma_n = o(1/n)$.
For $1/n \ll \sigma_n \ll 1/\sqrt n$, the spectral-norm
condition $w_{2n} \to 0$ ensures the GMM estimator
$\widehat\theta_{GMM}(W)$ is consistent (cf. Theorem \ref{Thm:DK22} b), while their
$|E|_{1,1} = o_p(n)$ condition fails. Our estimator
$\widehat\theta_p(\widehat W)$ further sharpens the rate via
dimension discounting.
For $\sigma_n\geq c>0$, the available bound for the GMM estimator
does not guarantee consistency. More generally, the plug-in estimator is consistent whenever the
dimension-discounted rate condition stated in
Theorem~\ref{Thm:DK22}(a) holds. $\blacksquare$
\end{example}
We next show the properties
$\widehat L$ and $\widehat S$ corresponding to \eqref{Eq:LRSED}. Define
\[
a_n
\coloneqq
\begin{cases}
\{m_r(L_0)m_c(L_0)\}^{-1/2},
& L_0\neq\mathbf 0_{n\times n},\\
0,&L_0=\mathbf 0_{n\times n}.
\end{cases}
\]
Define the rates
\begin{align*}
\mathcal R_F(W_0,E)
&=
\max\{r,1\}w_{2n}^2
+m_s(S_0)w_{\max n}^2
+m_s(S_0)a_n^2,\\
\mathcal R_2(W_0,E)
&=
w_{2n}^2
+m_s(S_0)w_{\max n}^2
+m_s(S_0)a_n^2,\\
\mathcal R_s(W_0,E)
&=
(w_{\max n}+a_n)\sqrt{m_s(S_0)}.
\end{align*}
Let
\[
B_0=(I_n-\lambda_0W_0)^{-1},
\qquad
\rho_n=w_{2n}+\mathcal R_s(W_0,E),
\]
\[
s_n=\mathcal R_s(W_0,E)\sqrt{m_s(S_0)},
\qquad
\kappa_n=1+\|W_0\|_2+\rho_n,
\]
and define the dimension-discounted rate
\[
\mathcal R_{nT}^*
\coloneqq
(1+\|B_0\|_2)\kappa_n^{d+1}
\frac{(r\vee1)\rho_n+s_n}{n}.
\]
\begin{theorem}
\label{Thm:DK1}
Suppose Assumptions~\ref{weight}--\ref{Assump:E_noise} hold.
Let the tuning parameters satisfy
\[
\nu_n=C_\nu w_{2n},
\qquad
\tau_n=C_\tau(w_{\max n}+a_n),
\]
for sufficiently large constants $C_\nu,C_\tau>0$.
\medskip
\noindent\textbf{(a) Frobenius bound.}
The two-step estimator satisfies
\begin{align}
\|\widehat L-L_0\|_F^2+\|\widehat S-S_0\|_F^2
&\lesssim_p\mathcal R_F(W_0,E),\\
|\widehat S-S_0|_{1,1}
&\lesssim_p
\mathcal R_s(W_0,E)\sqrt{m_s(S_0)}
+n\{w_{2n}+\mathcal R_s(W_0,E)\}.
\end{align}
\medskip
\noindent\textbf{(b) Sharpened spectral-norm bound.}
The two-step estimator satisfies
\[
\|\widehat L-L_0\|_2^2+\|\widehat S-S_0\|_2^2
\lesssim_p
\mathcal R_2(W_0,E).
\]
\end{theorem}
The rates $\mathcal{R}_F$, $\mathcal{R}_2$, and $\mathcal{R}_s$ share
a common structure with three error sources: a term in $w_{2n}$ about low-rank recovery; a term
$m_s(S_0)\,w_{\max n}^2$ related to $m_s(S_0)$ nonzero sparse
entries, and a term involving $\|L_0\|_{\max}$.
The factor $r$ in $\mathcal{R}_F$ is replaced
by unity in $\mathcal{R}_2$, reflecting the rank discount of
Theorem 9.24 in \citet{wainwright2019high}. Both rates parallel
Corollary 1 of \citet{Agarwal2012} and Theorem 9.19 of
\citet{wainwright2019high}. Note that estimator of \( L_0^* \) and \( S_0^* \) inherits the theoretical properties of the corresponding estimators of \( L_0 \) and \( S_0 \). For instance, since \( \widehat{L}^* = \widehat{L} - \operatorname{diag}(\widehat{L}) \), we have
$\| \widehat{L}^* - L_0^* \|_2 = O_p( \| \widehat{L} - L_0 \|_2)$,
which implies that estimation accuracy is preserved under the diagonal adjustment.
The conditions required for first-step and second-step
consistency differ in an essential way. For the L+S decomposition
$(\widehat L, \widehat S)$ to be consistent in operator norm using the bounds in Assumption \ref{Assump:E_noise},
i.e., $\|\widehat L - L_0\|_2 \to_p 0$ and $\|\widehat S - S_0\|_2
\to_p 0$, we require
\[
w_{2n}^2\to0,
\qquad
m_s(S_0)w_{\max n}^2\to0,
\qquad
m_s(S_0)a_n^2\to0.
\]
in keeping with the standard $L+S$ literature. Frobenius consistency,
$\|\widehat L-L_0\|_F\to_p0$ and
$\|\widehat S-S_0\|_F\to_p0$, requires the stronger condition
$\max\{r,1\}w_{2n}^2\to0$ in place of $w_{2n}^2\to0$, while the
other two conditions remain unchanged. Consistency of the spillover estimator
$\widehat\theta_p(\widehat W)$, $\|\widehat\theta_p(\widehat W) -
\theta_0\|_2 \to_p 0$, requires only weaker dimension-discounted
conditions, given in Theorem~\ref{Thm:DK22} below.
Next, define $\widehat W\coloneqq\widehat L+\widehat S$, where
$(\widehat L,\widehat S)$ is the estimator in
Theorem~\ref{Thm:DK1}. We use $\widehat W$ as the working matrix for
estimating the model in~\eqref{maineq}. We now impose the conditions
used to analyze the spillover-parameter estimators.
\begin{assumption}[True parameters]
\label{Assump:Theta}
\begin{enumerate}[(i)]
\item The parameter space $\Theta = \mathcal{C} \times \mathcal{B}\times {\Gamma}$
is a compact subset of $\mathbb{R} \times {\mathbb{R}^K} \times \mathbb{R}^K$ for some finite dimension $K$.
\item The true value, denoted as $\theta_0 = (\lambda_0, \beta_0^{\top}, \gamma_0 ^{\top})^{\top}$, lies in the interior of the parameter space $\Theta$. Moreover, $\lambda_0\beta_0+\gamma_0\neq \mathbf{0}_{K\times 1}$.
\end{enumerate}
\end{assumption}
\begin{assumption}[Adjacency matrix]
\label{Assump:True_Adjacency}
$W_0$ is a non-stochastic spatial weight matrix satisfying
\begin{enumerate}[(i)]
\item The diagonal elements of $W_0$ are 0, such that $W_{0, ii} = 0$ for all $i = 1, \ldots, n$.
\item $I_n, W_0, W_0^2$ are linearly independent.
\item Suppose that $J$ is fixed and that the absolute column sums of
$W_0$ are uniformly bounded in $n$, except possibly for those of its
first $J$ columns. Let $W_{0,u}$ contain the first $J$ columns of
$W_0$ and zeros elsewhere, and let
$W_{0,b}=W_0-W_{0,u}$. When $J=0$, set
$W_{0,u}=\mathbf 0_{n\times n}$ and $W_{0,b}=W_0$.
Assume that there exist constants $c_\infty,c_1>0$, independent of
$n$, such that
\[
\sup_{\lambda\in\mathcal C}
|\lambda|\|W_0\|_\infty\leq1-c_\infty,
\qquad
\sup_{\lambda\in\mathcal C}
|\lambda|\|W_{0,b}\|_1\leq1-c_1.
\]
\end{enumerate}
\end{assumption}
Assumption \ref{Assump:Theta} imposes standard conditions on the true parameters. The assumption
$\lambda_0 \beta_0 + \gamma_0 \neq \mathbf{0}_{K\times 1}$
is imposed to guarantee the identification of all the parameters in the model (see, e.g., \cite{Bramoulle2009}).
Assumption \ref{Assump:True_Adjacency} (i) requires that the true adjacency matrix \(W_{0}\) have zero diagonal, which rules out self‐loops. Regarding Assumption \ref{Assump:True_Adjacency} (ii), the matrices \(I_{n}\), \(W_{0}\), and \(W_{0}^{2}\) are linearly independent, i.e.\ there do not exist any nontrivial scalars \(\alpha_{0},\alpha_{1},\alpha_{2}\) such that
\[
\alpha_{0}I_{n} + \alpha_{1}W_{0} + \alpha_{2}W_{0}^{2} = \mathbf{0}_{n\times n}.
\]
Assumption~\ref{Assump:True_Adjacency}(iii) restricts the number of
dominant units to a fixed finite value and imposes uniform stability.
The first stability condition guarantees that
$I_n-\lambda W_0$ is invertible uniformly over
$\lambda\in\mathcal C$.
\begin{assumption}[Spatial errors]
\label{Assump:Spatial_Errors}
The primitive disturbances $\varepsilon_{it}$ are independently and identically distributed across $i$ and $t$ with mean $0$, variance $\sigma_0^2>0$ and $ \operatorname{\mathbb{E}}|\varepsilon_{it}|^{2+\delta}<\infty$ for a constant $\delta>2$.
\end{assumption}
For each $i$, let
$x_{t,i}=(x_{t,i1},\ldots,x_{t,iK})^\top$.
\begin{assumption}[Instrumental variables]
\label{Assump:Instruments}
\begin{enumerate}[(i)]
\item The covariates $X_t$
are stochastic $n \times K$ matrices with elements $x_{t,il}$ satisfying
$\sup_{n,t,i,\,1\leq l\leq K}
\mathbb{E}|x_{t,il}|^{2+\delta}<\infty$
for some $\delta>2$.
For each $i$, $\{x_{t,i}:t\in\mathbb Z\}$ is strictly stationary
and strongly mixing over $t$. The mixing coefficients are uniform
in $n$ and $i$: there exist constants $C<\infty$ and
$\rho\in(0,1)$ such that
\[
\sup_{n,i}\alpha_{n,i}(h)\leq C\rho^h,
\qquad h\geq1.
\]
The cross-sectional processes
$\{x_{t,i}:t\in\mathbb Z\}$ are independent across $i$. Dependence between $E$ and $X_t$ is allowed, subject to the
conditional moment restrictions stated in
Lemma~\ref{rateG}(c) in the Appendix.
\item $M$ is a pre-specified adjacency matrix. The instrumental variable matrix $Z_t(M)$
has columns chosen from $M^{\ell} X_{jt}$ for $\ell = 0, \ldots, d$ and $j = 1, \ldots, K$.
We require $\mathbb{E}[\varepsilon_t \mid Z_t(M)] = \mathbf{0}_{n \times 1}$.
Suppose $Z_t(M)$ has $m$ linearly independent columns, where $m$ is a fixed constant
satisfying $m \geq 2K+1$. \[
\sup_{n,t,i,\,1\leq l\leq m}
\mathbb E|z_{t,il}|^{2+\delta}<\infty,
\qquad \delta>2.
\]
\item $\lim_{\min(n,T) \to \infty} (nT)^{-1} \sum_{t=1}^T \mathbb{E}\bigl[Z_t(M)^\top Z_t(M)\bigr]$
exists and is nonsingular, denoted as $\Sigma_{Z_0 Z_0}$.
\item Let $Q_{0t} \coloneqq \bigl[W_0(I_n - \lambda_0 W_0)^{-1}(X_t \beta_0 + W_0 X_t \gamma_0),\,
X_t,\, W_0 X_t\bigr]$. We assume the $m \times (2K+1)$ matrix
$\lim_{\min(n,T) \to \infty} (nT)^{-1} \sum_{t=1}^T \mathbb{E}\bigl[Z_t(M)^\top Q_{0t}\bigr]$
exists, denoted as $G_0$, and is of full column rank.
\end{enumerate}
\end{assumption}
\begin{assumption}[Moment condition and GMM weight matrix]
\label{Assump:Moments_Weight}
The GMM weight matrix $\widehat{\Lambda}_{nT}^{-1} \in \mathbb{R}^{m \times m}$ converges in probability to some symmetric positive-definite matrix $\Lambda^{-1} \in \mathbb{R}^{m \times m}$, whose eigenvalues are bounded away from 0 and $\infty$.
\end{assumption}
Assumption \ref{Assump:Spatial_Errors} and \ref{Assump:Instruments} include standard moment assumptions. The independence \ requirement across $i$ in
Assumption \ref{Assump:Spatial_Errors} and Assumption~\ref{Assump:Instruments}(i) might be strong in network
settings, but it can be relaxed to weak cross-sectional
dependence with all rates in
Theorems~\ref{Thm:DK22} and~\ref{Thm:Supervised_Rate} unchanged
in order. We maintain independence for expositional simplicity. {
Assumptions \ref{Assump:Instruments} (ii) are regularity conditions on the instrumental variables. Among other things, the requirement that $\mathbb{E}[\varepsilon_t \mid Z_t(M)] = \mathbf{0}_{n \times 1}$ precludes $M=W$ if the measurement error $E$ is endogenous. Assumptions \ref{Assump:Instruments} (iii) and (iv) impose nonsingularity on the variance covariance matrix of the instrumental variables, and rank conditions of gradient of the moment functions.} The full column rank condition ensures that we only consider linearly independent columns of the instrument matrix, a common practice in the literature \citep{Lee2003}.
Take $M=W_0$, this is linked to the rank conditions of $W_0$, the variance covariance structure of $Z_t(W_0)$ and the correlation between $X_t$ and $Z_t(W_0)$.
To further understand the implications of the above assumption, let {$Z_t(W_0)$ be the $n\times m$ matrix}. Note that
$(nT)^{-1}\sum_{t=1}^T Z_t(W_0)^{\top} Z_t(W_0)$ is a {partitioned} matrix whose elements have the form
$(nT)^{-1}\sum_{t=1}^T X_t^{\top} [W_0^{l_1}]^{\top} W_0^{l_2}X_t,$ for any $l_1$ and $l_2$ between 0 and $d$.
This condition can be implied by
$\lambda_{\min}(\lim_{n,T\to \infty}\{nT\}^{-1}\sum_t\mathbb{E}\{Z_t^{\top}Z_t\})>0
,$
and
$\lambda_{2K+1}(W_0W^{\top}_0)> 0$. Then $\Sigma_{Z_0Q_0}$ is of full column rank could be implied by a condition that $\lambda_{2K+1}(W_0W_{0}^{\top})>0$, $\lim_{T\to \infty}T^{-1}\sum_t[X_t,W_0X_t,\cdots, W_0^dX_t]$ has column rank greater than or equal to $2K+1$, and all eigenvalues of $I_n - \lambda_0 W_0$ are bounded away from
$0$ and $\infty$.
Assumption \ref{Assump:Moments_Weight} is a standard assumption imposed on the GMM weight matrix. Because $M$ is fixed and time invariant, $Z_t(M)$ is a measurable
function of $X_t$. Therefore, $\{Z_t(M)\}$ inherits the strong-mixing
property, with
$\alpha_{Z(M),n}(h)\leq\alpha_{X,n}(h)$.
Before establishing the consistency of our spillover-parameter estimator, we first examine the bias introduced by
$E$ when using the classical GMM estimator.
We provide an intuitive argument illustrating how the measurement error matrix \(E\) biases the moment functions and the importance of its removal, particularly when the instrument matrix \(Z_t(M)\) also depends on \(W\), as in standard spatial regression. For simplicity here, assume no contextual effects are present.When $Z_t(M)$ consists of exogenous instruments satisfying
$\mathbb{E}[Z_t(M)^\top \varepsilon_t] = \mathbf{0}_{m \times 1}$,
the moment function is given by
\[
\bar g_{nT}(\theta, W, M) = \frac{1}{nT} \sum_{t=1}^{T}
Z_t(M)^\top (Y_t - \lambda W Y_t - X_t \beta).
\]
Substituting $W = W_0 + E$,
\[
\bar g_{nT}(\theta, W, M) = \frac{1}{nT} \sum_{t=1}^{T}
Z_t(M)^\top (Y_t - \lambda W_0 Y_t - X_t \beta)
- \frac{\lambda}{nT} \sum_{t=1}^{T} Z_t(M)^\top E Y_t.
\]
Thus, if $E$ is independent of $(Z_t(M),Y_t)$ and has mean zero, then
\[
\mathbb{E}\!\left[Z_t(M)^\top E Y_t\right]
=\mathbf{0}_{m\times 1},
\]
and hence
\[
\mathbb{E}\!\left[\bar g_{nT}(\theta_0,W,M)\right]
=\mathbf{0}_{m\times 1}.
\]
However, when $Z_t(M)$ is a plug-in version of the true instrument
with $M = W$, i.e.,
\[
Z_t(W) = [X_t, W X_t] = [X_t, (W_0 + E) X_t],
\]
the moment function becomes
\[
\bar g_{nT}(\theta, W) = \frac{1}{nT} \sum_{t=1}^{T}
[X_t, (W_0 + E) X_t]^\top \big[(Y_t - \lambda W_0 Y_t - X_t \beta)
- \lambda E Y_t\big].
\]
Expanding the moment function $\Bar{g}_{nT}(\theta, W)$ with the
instrument matrix $[X_t, (W_0+E)X_t]$, we get:
{\[
\left[ \frac{1}{nT} \sum_{t=1}^{T}(-\lambda X_t^{\top} EY_t)^{\top}, \frac{1}{nT} \sum_{t=1}^{T}[X_t^\top E^\top (Y_t-\lambda W_0 Y_t-X_t\beta) - \lambda X_t^\top W_0^\top E Y_t - \lambda X_t^\top E^\top E Y_t]^{\top} \right]^{\top}.
\]}
Even if \(E_{ij}\)'s are i.i.d. mean zero, with bounded second moment and independent of $(X_t,\varepsilon_t)$, there will be moment condition misspecification. Specifically,
\[
\mathbb{E}[\Bar{g}_{nT}(\theta_0, W)] \neq { \mathbf{0}_{m\times 1}},
\]
since the expectation of the last term yields:
$n^{-1}\mathbb{E}[X_t^\top E^\top E Y_t]
= n^{-1} \mathbb{E}\big[X_t^\top \mathbb{E}[E^\top E]\,(I-\lambda_0 W_0)^{-1} X_t\big]\beta_0
\neq \mathbf{0}_{K\times 1},$
which might be non-negligible under our assumptions.
Recall that \(\widehat{\theta}_{GMM}\) denotes the efficient two-step GMM estimator based on a given adjacency matrix \(W\), obtained in closed form in \eqref{Eq:GMM_estimator}. From the preceding results, it follows that \(\widehat{\theta}_{GMM}\) is likely to be inconsistent. We now analyze the convergence rate of \(\widehat{\theta}_p(\widehat{W})\).
Since \(W_0\) is not directly available for constructing instruments, one approach is to use \(Z_t(\widehat{W})\) as plug-in instruments, where \(\widehat{W}\) is obtained from Algorithms \ref{Algo:CD_LRSED}
and satisfies Theorem \ref{Thm:DK1}. When \(W_0\) is correctly observed, the minimizer \(\widehat{\theta}_p(W_0)\) satisfies the standard panel-data estimator convergence rate:
\[
\|\widehat{\theta}_p(W_0) - \theta_0\|_2 = O_p(1/\sqrt{nT}).
\]
However, when \(W_0\) is replaced by a denoised estimate \(\widehat{W}\), the associated minimizer \(\widehat{\theta}_p(\widehat{W})\) depends on the deviation \(\widehat{W} - W_0\). The corresponding convergence properties are detailed in the following results.
\begin{theorem}[Performance of the plug-in estimator]
\label{Thm:DK22}
Recall the definitions of
$B_0$, $\rho_n$, $s_n$, $\kappa_n$, and
$\mathcal R_{nT}^*$ given above.
Consider the model described in \eqref{maineq}.
Suppose Assumptions \ref{weight}--\ref{Assump:Moments_Weight} hold.
\medskip
\noindent\textbf{(a) The plug-in estimator.}
If $\mathcal R_{nT}^*=o(1)$, then
\[
\|\widehat\theta_p(\widehat W)-\theta_0\|_2
=
O_p(\mathcal R_{nT}^*)
+
O_p\!\left(\frac{1}{\sqrt{nT}}\right).
\]
\medskip
\noindent\textbf{(b)(The GMM estimator).}
Using the noisy matrix $W = W_0 + E$ directly,
\begin{align}
\|\widehat{\theta}_{GMM} - \theta_0\|_2
= O_p(w_{2n} + w_{2n}^2)
+ O_p\!\left(\frac{1}{\sqrt{nT}}\right).
\end{align}
\medskip
\medskip
\noindent\textbf{(c) Asymptotic normality.}
If
\[
\sqrt{nT}\,\mathcal R_{nT}^*=o(1),
\]
then
\[
\sqrt{nT}\,
\bigl(\widehat\theta_p(\widehat W)-\theta_0\bigr)
\overset{d}{\longrightarrow}
\mathcal N(\mathbf 0,\Sigma_{GMM}).
\]
where
\[
\Sigma_{GMM}
=\left(G_0^\top \Lambda^{-1}G_0\right)^{-1}
G_0^\top \Lambda^{-1}\Omega_0\Lambda^{-1}G_0
\left(G_0^\top \Lambda^{-1}G_0\right)^{-1},
\]
and
\[
\Omega_0
=
\frac{1}{n}
\mathbb{E}(Z_t(W_0)^\top
\varepsilon_t\varepsilon_t^\top
Z_t(W_0)).
\]
\end{theorem}
Under the assumptions of Theorem~\ref{Thm:DK22}(a), the plug-in
estimator $\widehat\theta_p(\widehat W)$ is consistent provided that
\[
\mathcal R_{nT}^*=o(1).
\]
The factor $1/n$ discounts the noise contribution by the network
size, so consistency holds even when $w_{2n}$ and $w_{\max n}$ do
not vanish individually, for example, when $r$ and $m_s(S_0)$
remain bounded and $w_{2n}$ is of constant order. By contrast, the available bound for the GMM estimator based on the
noisy matrix does not guarantee consistency when $w_{2n}$ is bounded
away from zero; the moment-misspecification example above shows that
inconsistency can occur in this case. The noise-induced component of the plug-in estimator's rate is
strictly smaller whenever
\[
\mathcal R_{nT}^*
=o(w_{2n}+w_{2n}^2).
\]
The improvement reflects how the $L+S$ decomposition enters the
analysis. Denote
$\Delta L := \widehat L - L_0$ and $\Delta S := \widehat S - S_0$ as
the recovery errors from Theorem~\ref{Thm:DK1}. The recovery error $\widehat W - W_0 = \Delta L + \Delta S$
affects the spillover estimator only through bilinear forms
(averages over the network) with the instruments and residual
design, and these averages discount the error according to its
structure. The low-rank error enters through its nuclear norm,
controlled by $\sqrt r\,\|\Delta L\|_F$, contributing $O(r)$
effective directions rather than all $n$; the sparse error enters
through $|\Delta S|_{1,1}$, controlled by its $m_s(S_0)$ nonzero
entries rather than all $n^2$. The measurement error can therefore
be large in aggregate while the spillover estimator still converges.
Example~\ref{lewbel:e} constructs regimes in which
$|E|_{1,1}/n$ does not vanish, while the dimension-discounted
condition $\mathcal R_{nT}^*=o(1)$ may still hold. This
condition is sharp for sparse measurement error such as link
misclassification in $0$--$1$ networks, but fails in the dense
case that motivates our analysis: Example~\ref{lewbel:e}
constructs the case in which $|E|_{1,1}$ is of order $n^2$ yet
the plug-in estimator remains consistent. The two analyses
cover disjoint measurement-error regimes: sparse errors with
\citet{Lewbeletal2024EJ}, dense errors with structured latent
adjacency in our framework. When $L_0 = \mathbf{0}_{n \times n}$,
the rate specializes to $O_p(m_s(S_0)w_{\max n}/n)$
(Corollary~\ref{cor1}).
The following corollary specializes Theorem~\ref{Thm:DK22} to the
pure-sparse case $L_0=\mathbf 0_{n\times n}$. Within
Assumption~\ref{weight}, only part~(ii) on $S_0$ is needed, and the
coherence terms drop out. The rate is driven entirely by the sparse
recovery error $n^{-1}|\widehat S-S_0|_{1,1}$, the same
$\ell_1$-type quantity as the $|E|_{1,1}/n$ condition of
\citet{Lewbeletal2024EJ}, but applied to the recovery error after
denoising rather than to the raw measurement error.
\begin{corollary}\label{cor1}
Suppose $L_0=\mathbf 0_{n\times n}$ and let $\widehat S$ denote the
pure-sparse estimator obtained by fixing
$\widehat L=\mathbf 0_{n\times n}$. Suppose
Assumption~\ref{weight}(ii),
Assumption~\ref{Assump:E_noise}(i),(iii), and
Assumptions~\ref{Assump:Theta}--\ref{Assump:Moments_Weight} hold.
Assume additionally that
\[
\|S_0\|_1=O(1),
\qquad
\|E\|_\infty=O_p(1),
\]
and
\[
\max_{1\leq j\leq K}
\frac{1}{T}\sum_{t=1}^T\|X_{jt}\|_{\max}^2=O_p(1),
\qquad
\frac{1}{T}\sum_{t=1}^T
\|\varepsilon_t\|_{\max}^2=O_p(1).
\]
Then
\begin{align*}
\|\widehat\theta_p(\widehat S)-\theta_0\|_2
&=
O_p\!\left(
\frac{|\widehat S-S_0|_{1,1}}{n}
\right)
+
O_p\!\left(\frac{1}{\sqrt{nT}}\right)\\
&=
O_p\!\left(
\frac{m_s(S_0)w_{\max n}}{n}
\right)
+
O_p\!\left(\frac{1}{\sqrt{nT}}\right).
\end{align*}
\end{corollary}
The rate in Corollary~\ref{cor1} follows from
Lemma~\ref{rateG}(b): the sparse recovery error enters the spillover
estimator only through
$n^{-1}|\widehat S-S_0|_{1,1}$. In the pure-sparse case, the
sparse-recovery argument used in the proof of
Theorem~\ref{Thm:DK1} gives
\[
\frac{1}{n}|\widehat S-S_0|_{1,1}
=
O_p\!\left(
\frac{m_s(S_0)w_{\max n}}{n}
\right).
\]
Although Corollary~\ref{cor1} covers the pure-sparse case, the gain
from our general analysis is most pronounced when the low-rank
component $L_0$ is genuinely present and $r$ and $m_s(S_0)$ remain
controlled relative to $n$.
{
\begin{remark}[Asymptotic collinearity]
It should be noted that variation in peer connections plays an important role in identification and estimation as noted, for example, in \cite{Manski1993} (p.535) and \cite{Paula2017}. In a \emph{single}, \emph{dense} and \emph{large} network, the leading eigenvector of a row-normalized and positive adjacency matrix may become (asymptotically) proportional to a vector of ones. In this case, peer averages become increasingly similar across individuals, complicating identification and estimation of spillover effects. A recent formulation of this issue in the context of peer effects is provided by \cite{hayes2024minimax}, who show that this collinearity arises when $W_0$ is row-normalized, positive, and exhibits unbounded minimum degrees. Our setting differs in two respects.
For a nonnegative row-normalized matrix,
$W_0\mathbf 1_n=\mathbf 1_n$ holds exactly. When nodal covariates
are independent of the network and the minimum degree diverges, peer
averages may concentrate around a common value, producing asymptotic
collinearity; see \citet{hayes2024minimax}. Our assumptions do not
generally impose row normalization or diverging minimum degree, so
this conclusion does not follow automatically in our setting. The
presence of a sparse component $S_0$ does not by itself eliminate
this possibility. Without row normalization, $\mathbf 1_n$ need not
be an eigenvector of $W_0$, and the argument requires separate
verification.
$\blacksquare$
\end{remark}}
As mentioned previously, we emphasize that our results are also suitable
for the case when $E$ is endogenous with respect to $\varepsilon_t$. As
an example, in social networks this can arise when misclassification
errors are not generated at random, instead being correlated with unobservables
driving the network formation process, such as the framework considered
by \citet{Candelaria2023}. To see this, we provide a rate analysis to
compare the rate when $E$ is endogenous or not of both our estimator and
the GMM estimator. For simplicity, we adopt the following assumption on
the relationship between these two sources of error, see similar
assumptions as in \citep{Lee2014}:
\begin{assumption}[Endogenous measurement error]
\label{Assump:E_noise_endogenous}
Suppose Model~\eqref{maineq} holds and $W_0$ is symmetric and
positive semidefinite. For each $i$, let
\[
\varepsilon_i
=
(\varepsilon_{i1},\ldots,\varepsilon_{iT})^\top
\in\mathbb R^T,
\]
and let
\[
E_{i,-i}
=
(E_{ij}:j\neq i)^\top
\in\mathbb R^{n-1}
\]
collect the off-diagonal elements of the $i$th row of $E$, with
$E_{ii}=0$.
Suppose that the vectors
$(\varepsilon_i^\top,E_{i,-i}^\top)^\top$ are independent across
$i$ and jointly Gaussian:
\[
\begin{pmatrix}
\varepsilon_i\\
E_{i,-i}
\end{pmatrix}
\sim
\mathcal N\!\left(
\mathbf 0,\,
\Sigma_{\varepsilon E}
\right),
\]
where
\[
\Sigma_{\varepsilon E}
=
\begin{bmatrix}
\sigma_0^2 I_T
&
\sigma_0\sigma_E n^{-s}\rho_{\varepsilon E}\\
\sigma_0\sigma_E n^{-s}\rho_{\varepsilon E}^\top
&
\sigma_E^2n^{-2s}\Sigma_E
\end{bmatrix},
\qquad s>\frac12.
\]
Here,
$\rho_{\varepsilon E}\in\mathbb R^{T\times(n-1)}$ governs the
dependence between the outcome disturbances and the measurement
error, and
$\Sigma_E\in\mathbb R^{(n-1)\times(n-1)}$ is a correlation matrix.
Assume that there exist constants $c_E,C_E,C_\rho>0$, independent
of $n$ and $T$, such that
\[
c_E
\leq
\lambda_{\min}(\Sigma_E)
\leq
\lambda_{\max}(\Sigma_E)
\leq
C_E,
\qquad
\|\rho_{\varepsilon E}\|_2\leq C_\rho,
\]
and that $\Sigma_{\varepsilon E}$ is positive semidefinite.
\end{assumption}
Assumption~\ref{Assump:E_noise_endogenous} is imposed together with
Assumptions~\ref{Assump:E_noise} and~\ref{Assump:Spatial_Errors}.
Thus, the disturbances remain marginally i.i.d.\ across $i$ and $t$,
while dependence between $E$ and $\varepsilon_t$ is allowed. The
Gaussian specification is imposed solely to illustrate the rate in
the presence of endogenous measurement error.
Denote $\widehat{\theta}_{GMM}(W,M)$ and $\widehat{\theta}_p(\widehat{W},M)$
the GMM and plug-in estimator with instrument $Z_t(M)$. In comparison to the
exogenous measurement error case, the bias of
$\widehat{\theta}_{GMM}(W,M) = \widehat{\theta}_p(W,M)$ could be larger than
our proposed estimators. It is because the deviation between sample moments
$\Bar{g}_{nT}(\theta, W, M)$ and $\Bar{g}_{nT}(\theta, W_0, M)$ might be
more severe. This is reflected in our simulation exercises, see
Section~\ref{Sec:Simulation}. For simplicity, we shall illustrate without the contextual-effects regressor $W_0X_t$
(i.e., setting $\gamma_0 = \mathbf{0}_K$), so the regressors are
$[W_0Y_t, X_t]$. Define
${G}(M) = -n^{-1}\mathbb{E}[Z^{\top}_t(M)[W_0 Y_t, X_t]]$. To explain this,
take $M$ as exogenous; the leading term for the moment estimator
$\widehat{\theta}_p(\widehat{W},M) - \theta_0$ is
$-(G^{\top}(M)\Lambda^{-1}G(M))^{-1}G^{\top}(M)\Lambda^{-1}
\Bar{g}_{nT}(\theta_0, W_0, M)$,
according to Theorem~\ref{Thm:DK22}. We shall suppress the dependence of
$M$ on $W_0$ and $E$. The following proposition characterizes a rate bound
for both the GMM estimator and our estimator.
Recall that
$\bar g_{nT}(\theta,W,M)
=
\frac1{nT}\sum_{t=1}^T
Z_t(M)^\top
\left(
Y_t-\lambda WY_t-X_t\beta-WX_t\gamma
\right)$.
\begin{proposition}[Endogenous error bias]
\label{lemmabias}
Suppose that $M$ is fixed or exogenous and that the conditional-moment
conditions in Lemma~\ref{rateG}(c) hold. Under
Assumptions~\ref{weight}--\ref{Assump:E_noise_endogenous}, with
$\|B_0\|_2(1+\|W_0\|_2)=O(1)$ and
$\mathcal R_{nT}^*=o(1)$, we have
\begin{align}
\|\widehat\theta_p(\widehat W,M)-\theta_0\|_2
&\lesssim_p
\mathcal R_{nT}^*
+\frac1{\sqrt{nT}},
\label{prop1_denoised}\\
\|\widehat\theta_{GMM}(W,M)-\theta_0\|_2
&\lesssim_p
n^{1/2-s}\|\rho_{\varepsilon E}\|_2
+w_{2n}
+\frac1{\sqrt{nT}}.
\label{prop1_gmm}
\end{align}
\end{proposition}
Three features of this comparison are worth highlighting.
\emph{First}, the endogeneity bias $n^{1/2-s}\|\rho_{\varepsilon E}\|_2$
appears only in the GMM rate~\eqref{prop1_gmm}. It originates from a
bilinear form that is quadratic in $E$: the correlation between $E$ and
$\varepsilon_t$ contributes a non-vanishing mean to
$Z_t^\top E B_0\varepsilon_t$, whose magnitude is governed by
$\rho_{\varepsilon E}$.
In the plug-in case, the analogous bilinear form involves
$(\Delta L+\Delta S)B_0\varepsilon_t$. Its low-rank component is
controlled by
\[
\|\Delta L\|_*
=
O_p\!\left((r\vee1)(w_{2n}+\mathcal R_s)\right),
\]
whereas its off-diagonal sparse component is controlled by
\[
|\Delta S|_1
=
O_p\!\left(\mathcal R_s\sqrt{m_s(S_0)}\right).
\]
Under the conditional-moment conditions, the plug-in bound therefore
contains no additional explicit term involving
$\rho_{\varepsilon E}$.
\emph{Second}, even in the absence of endogeneity
($\rho_{\varepsilon E} = 0$), the denoised rate improves upon the GMM
rate: the noise-level term $w_{2n}$
in~\eqref{prop1_gmm} is replaced by the rate
$\mathcal{R}^*_{nT}$ in~\eqref{prop1_denoised}, which is strictly
smaller whenever
$\mathcal{R}^*_{nT} = o(w_{2n})$.\emph{Third}, the denoised rate \emph{attenuates} the dependence on
the noise magnitude $w_{2n}$ rather than eliminating it: $w_{2n}$
enters through the $L+S$ coupling $\|\Delta L\|_2 = O_p(w_{2n}+\mathcal{R}_s)$,
but its effect is multiplied by the factor $r/n$ rather
than by 1. As a result, the rate scales not with the raw noise level
$w_{2n}$ but with the ratio of the structural complexity ($r$, $m_s(S_0)$)
to the sample size $n$. This reflects the benefit of our approach.
Finally, we consider the rate of the supervised estimator as defined in Equation (\ref{Eq:Supervised_Estimator}).
Denote $\widehat{W}_s = \widehat{L}_s+\widehat{S}_s.$
\begin{theorem}[Performance of the supervised estimator]
\label{Thm:Supervised_Rate}
Suppose Assumptions~\ref{weight}--\ref{Assump:Moments_Weight}
hold. Let $M$ be fixed, exogenous, or $\sigma(E)$-measurable and
satisfy the stated instrument and design conditions.
Let $\widehat\theta_s$ denote the supervised estimator defined
in~\eqref{Eq:Supervised_Estimator}.
\[
\nu_{nT}=C_\nu w_{2n},
\qquad
\tau_{nT}=C_\tau(w_{\max n}+a_n),
\]
for sufficiently large constants $C_\nu,C_\tau>0$.
Suppose
\[
\mathcal R^*_{nT}=o(1)
\qquad\text{and}\qquad
\xi_{nT}=O(1).
\]
Then
\begin{equation}
\|\widehat\theta_s-\theta_0\|_2
=
O_p\left(
\frac1{\sqrt{nT}}+\mathcal R^*_{nT}
\right).
\label{superv8}
\end{equation}
If, in addition,
\[
\sqrt{nT}\,\mathcal R^*_{nT}=o(1)
\qquad\text{and}\qquad
\xi_{nT}=o(1),
\]
then
\[
\sqrt{nT}
(\widehat\theta_s-\theta_0)
\overset{d}{\longrightarrow}
\mathcal N(\mathbf 0,\Sigma_{GMM}).
\]
\end{theorem}
The supervised estimator achieves the same rate as the plug-in
estimator in Theorem~\ref{Thm:DK22} and it is strictly smaller
than that of the noisy-$W$ GMM estimator whenever
$\mathcal R_{nT}^*=o(w_{2n}+w_{2n}^2)$.
The rate still depends on $w_{2n}$, but only through the
dimension-discounted recovery rate $\mathcal R_{nT}^*$.
It is worth noting that the rate in~\eqref{superv8} depends on
the parameters $(r, m_s(S_0), n, T)$ and on the
spectral characteristics of the underlying network
($\|B_0\|_2, \|W_0\|_2$), rather than directly on the noise level
$w_{2n}$ compared to the rate of the GMM estimator. The condition on $\xi_{nT}$
ensures that the generated error from the first-iteration
estimator $\widehat{\theta}^{[1]}$ does not inflate either the
sparse recovery in Step b) or the low-rank recovery in Step c)
beyond the unsupervised noise level, so that the supervised
decomposition matches the plug-in decomposition rates.
The following corollary specializes Theorem~\ref{Thm:Supervised_Rate}
to the pure-sparse case $L_0 = \mathbf{0}_{n\times n}$, paralleling
Corollary~\ref{cor1}: only Assumption~\ref{weight}(ii) on $S_0$ is
needed, the coherence terms drop out, and the rate is driven entirely
by the sparse recovery error $n^{-1}|\widehat S_s - S_0|_{1,1}$,
making the connection to the misclassification framework of
\citet{Lewbeletal2024EJ} explicit.
\begin{corollary}\label{cor2}
Suppose $L_0=\mathbf 0_{n\times n}$ and consider the pure-sparse
supervised estimator obtained by imposing $L=\mathbf 0_{n\times n}$,
so that $\widehat W_s=\widehat S_s$. Suppose
Assumption~\ref{weight}(ii),
Assumption~\ref{Assump:E_noise}(i),(iii), and
Assumptions~\ref{Assump:Theta}--\ref{Assump:Moments_Weight} hold.
Assume, in addition, that
\[
\|W_0\|_1=O(1),
\qquad
\|E\|_\infty=O_p(1),
\qquad
\xi_{nT}=O(1),
\]
and
\[
\max_{1\leq j\leq K}
\frac{1}{T}\sum_{t=1}^T\|X_{jt}\|_{\max}^2=O_p(1),
\qquad
\frac{1}{T}\sum_{t=1}^T
\|\varepsilon_t\|_{\max}^2=O_p(1).
\]
Then
\begin{align}
\|\widehat\theta_s-\theta_0\|_2
&=
O_p\!\left(
\frac{1}{n}|\widehat S_s-S_0|_{1,1}
\right)
+
O_p\!\left(\frac{1}{\sqrt{nT}}\right).
\end{align}
Moreover, Lemma~\ref{rateG}(b), together with the sparse-recovery
bound for $\widehat S_s$, gives
\begin{align}
\|\widehat\theta_s-\theta_0\|_2
&=
O_p\!\left(
\frac{m_s(S_0)w_{\max n}}{n}
\right)
+
O_p\!\left(\frac{1}{\sqrt{nT}}\right).
\end{align}
\end{corollary}
The theoretical performance of the supervised estimator is similar
to that of the plug-in estimator, as confirmed by the simulation
and application results in Sections~\ref{Sec:Simulation}
and~\ref{Sec:Application}. We shall also note that both
Theorems~\ref{Thm:DK22} and~\ref{Thm:Supervised_Rate} do not
require $E$ to be exogenous, as long as $Z_t(M)$ does not directly
involve $W$. Furthermore, both theorems yield rates in which the
noise level $w_{2n}$ enters only after dimension-discounting by
the factor $r/n$ (rather than at the raw spectral rate
$w_{2n}$ delivered by the noisy-$W$ baseline), reflecting the
fundamental advantage of the $L+S$ decomposition combined with the
within transformation.
\section{Simulations}
\label{Sec:Simulation}
In this section, we present simulation exercises to evaluate the performance of our proposed estimators. We first introduce four data-generating processes (DGPs) for the network adjacency matrix $W_0$. We then describe how it is contaminated with noise, and discuss the Monte Carlo simulation results.
\subsection{Generating $\mathbf{W}_0$}
\label{Subsec:Network_DGPs}
The diagonal elements of $W_0$ are assumed to be 0. Hence we set all diagonal elements of $E$ to be $0$ and do not penalize $S_0$ on its diagonal. The other entries of $E$ are generated as explained in Section \ref{dgp}.
\begin{itemize}
\item[i)] \textbf{DGP 1}(Low rank)
In this setup, $L_0$ is generated by a pure low-rank matrix as:
\begin{equation}
\label{Eq:DGP_LowRank}
L_0 = U DV^{\top}, \quad U, V \in \mathbb{R}^{n \times r} \, , \text{ and} \quad U^{\top} U = I_r, \quad V^{\top} V = I_r,
\end{equation}
where $D$ is the $r\times r$ diagonal matrix $D = \operatorname{diag}\{0.8^0, 0.8^1, \ldots, 0.8^{r-1}\}$, and $U$ and $V$ ---the left and right singular vectors of $L_0$, respectively--- are the first $r$ columns of two random $n \times n$ orthonormal matrices generated according to standard algorithms \citep{Stewart1980, Mezzadri2007}. We then set $W_0=L_0^*$, where $L_0^*$ has zero diagonal elements and its off-diagonal elements are the same as $L_0$.
\item[ii)] \textbf{DGP 2} (Low rank+ sparse)
Our second DGP still uses the same $L_0$ as in DGP 1, but adds a sparse $n \times n$ component $S_0$, where we randomly pick $m_r=2$ elements in each row and draw their values from a $\text{Uniform}(0.5, 1)$ distribution. Note that $W_0$ could be written as
\begin{equation}
\label{Eq:DGP_LowRank_Sparse}
W_0 =L_0^*+ S_0^*,
\end{equation}
where $L_0^*$ is defined as in DGP 1 using $U$, $V$ and $D$ generated as in \eqref{Eq:DGP_LowRank}, and $S_0^*$ has zero diagonal elements and its off-diagonal elements are the same as $S_0$.
\item[iii)] \textbf{DGP 3} (Dominant units) Our third DGP generates $W_0$ as a network with dominant units as presented in Subsection \ref{Subsec:Dominant_Units_Simulation}. Suppose the number of dominant units is $J=2$. The first $J$ units are dominant players in the network, whereas the remaining $n-J$ units are ordinary players. We could partition $W_0$ as \begin{eqnarray}
\label{Eq:DGP_Dominant_Network}
W_0=\begin{bmatrix}
W_{11} & W_{12} \\
W_{21} & W_{22}
\end{bmatrix} .
\end{eqnarray}
Let $W_D =
[W_{11}^{\top}, W_{21}^{\top}]^{\top}$ be the $n \times J$ dominant block of the adjacency matrix corresponding to the dominant units, and $W_{R}= [W_{12}^{\top}, W_{22}^{\top}]^{\top}$, the right block of the adjacency matrix corresponding to ordinary units.
For each column of $W_D$, the first $\lfloor n^{0.9} \rfloor$ entries, except the diagonal elements, are drawn independently from a $\text{Uniform}(0, 1)$ distribution, and the remaining entries are set equal to $0$. For $W_{R}$, we induce sparsity by imposing that
each individual is connected with equal weights to the immediate predecessor and successor. Note that $W_0$ could be viewed as a purely sparse matrix $S_0^*$ as each of its row has at most $2+J$ nonzero elements, and we treat $L_0^*$ as a zero matrix.
\item[iv)]\textbf{DGP 4} (Group) Our fourth DGP divides the individuals into two groups. The first $n/4$ members form group 1, the remaining $3n/4$ form group 2, and only members within the same group have connections. For each group, the connection strength between any two different individuals $i$ and $j$ within the same group is defined by $d_g u_{g,i}u_{g,j}/\sum_i u_{g,i}^2$, where $d_{1,1}=1$ and $d_{2,2}=0.9$, $u_{g,i}$'s are latent individual-specific characteristic generated independently from standard normal random variables. Note that DGP 4 admits the low-rank form such that
$L_0=UDU^\top$, where $U=[U_1, U_2]$, $D$ is a $2\times 2$ diagonal matrix with $d_{1,1}=1$ and $d_{2,2}=0.9$, $U_1=[u_{1,1}, \cdots, u_{1,n/4}, 0, \cdots, 0]/\sqrt{\sum_i u_{1,i}^2}$ and $U_2=[0, \cdots, 0, u_{2,1}, \cdots, u_{2,3n/4}]/\sqrt{\sum_i u_{2,i}^2} $, and $W_0=L_0^*$, where $L_0^*$ has zero diagonal entries and its off diagonal elements are the same as $L_0$.
\end{itemize}
\subsection{Monte Carlo setup}\label{dgp}
For our simulation exercises, we set $\lambda_0 = 0.25$, $\beta_0 = (-1, 2)$. We generate two covariates for $X_t$ independently from a $\mathbb{N}(0, 1)$ and a $\mathbb{N}(5, 2)$, respectively. We set sample sizes $n \in \{40,80, 120\}$ and time dimensions $T \in \{1, 5, 15, 50\}$, considering specific combinations of panel sizes due to computational constraints. We take as given a true network $W_0$ generated using one of the DGPs introduced in the previous subsection, and keep it constant across 100 Monte Carlo simulations.
Our exercises additionally consider the possibility that measurement errors are endogenous. That is, we allow for correlated disturbances between the outcome equation ($\varepsilon_t$) and the errors that contaminate the adjacency matrix $(E)$, as outlined in our Assumption \ref{Assump:E_noise_endogenous}. Let $\sigma_E $ and $\sigma_\varepsilon$ be two positive constants. Specifically, for $i,j = 1, \cdots, n$, we generate $E_{ij} \sim i.i.d. \quad \mathbb{N}(0, \sigma^2_E)$ for $i\neq j$, and set $E_{ii}=0$. We generate $\varepsilon_t$ such that its $i$th component $\varepsilon_{it}$ satisfies $\varepsilon_{it}=\rho \sigma_\varepsilon\sum_{j=1}^n E_{ij}/\sqrt{n} +v_{it}$, and $v_{it}$ are i.i.d. $ \mathbb{N} (0, (1-\rho^2)\sigma_\varepsilon^2)$ that are independent of $E_{ij}$. Note that $\varepsilon_{t}|E$ is normally distributed and we could simply obtain its conditional mean and variance. We define $\mbox{Vec}(E)$ as the vector obtained by stacking the off-diagonal elements of $E$.
\begin{eqnarray*}
\label{Eq:Correlated_Errors}
\left ( \begin{matrix}
\varepsilon_t \\
\mbox{Vec}(E)
\end{matrix} \right )
\sim \mathbb{N} \left( \mathbf{0}_{n^2}, \begin{bmatrix}
\sigma_{\varepsilon}^2 \cdot I_n & \Sigma_{\varepsilon E} \\
\Sigma_{\varepsilon E}^{\top} & \sigma_E^2 \cdot I_{n^2-n}
\end{bmatrix} \right) , \quad t = 1, \ldots, T.
\end{eqnarray*}
where
$\Sigma_{\varepsilon E}$ is the $n\times (n^2-n)$ matrix that captures the correlation between $\varepsilon_{it}$ and $ \mbox{Vec}(E)$ stacks all $E_{ij}$ for $i\neq j$ into a long vector. Note that it is implied from the previous sections that $\mbox{cov}(\varepsilon_{it}, E_{ij})=\rho\sigma_E^2\sigma_\varepsilon/\sqrt{n}$ for $i\neq j$ and $\mbox{cov}(\varepsilon_{it}, E_{kj})=0$ if $k\neq i$. In the simulation, we set $\sigma_\varepsilon=0.15$, $\sigma_E = 0.3/n^{0.7}$ and $\rho=0$ or $0.7$. By setting $\rho = 0$, we recover the exogenous measurement error case from Assumption \ref{Assump:E_noise}.
Given values for the adjacency $W_0$, covariates $X_t$, parameters $\theta_0$, and disturbances $\varepsilon_t$, we use the reduced-form representation of model \eqref{maineq} to generate our outcome $Y_t$ at every time period:
\begin{align}
\label{Eq:SAR_reduced}
Y_t = (I_n - \lambda_0 W_0)^{-1}(X_t \beta_0 + \varepsilon_t), \, t = 1, \ldots, T \, .
\end{align}
Throughout the simulation we set $W = W_0 + E$, which is the observed noisy version of the adjacency matrix, to construct estimates of parameters $\theta_0$. Once the adjacency matrix is generated, we use it to construct a rough estimate of the variance of the errors and set the tuning parameters for estimation. Specifically, we fix $\xi_{nT} = 1$ and set $\mbox{Int}(W)$ as the interquartile range of $W_{ij}$. Then we choose the sparse penalty as $\tau_n=\tau_{nT} = 2\mbox{Int}(W)\log(n)$ and $\nu_n=\nu_{nT}=\mbox{Int}(W) \sqrt{n}$.
\subsection{Simulation results}
\begin{table}[htbp]
\centering
\captionsetup{width=0.975\linewidth}
\caption{Relative RMSE for estimators of $\lambda_0$}
\label{Tab:Relative_RMSE_Exogenous}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc|cc|cc|cc}
\toprule
Exogenous & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised \\
\midrule
L (Low-rank) & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.499 & 0.409 & 0.262 & 0.258 & 0.251 & 0.245 & 0.239 & 0.235 \\
n=80 & 0.279 & 0.279 & 0.203 & 0.203 & 0.200 & 0.196 & 0.197 & 0.188 \\
n=120 & 0.361 & 0.353 & 0.146 & 0.140 & 0.131 & 0.124 & 0.120 & 0.135 \\
\midrule
L+Sparse & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.444 & 0.438 & 0.285 & 0.280 & 0.231 & 0.221 & 0.211 & 0.207 \\
n=80 & 0.360 & 0.359 & 0.237 & 0.235 & 0.201 & 0.198 & 0.181 & 0.181 \\
n=120 & 0.297 & 0.296 & 0.146 & 0.145 & 0.101 & 0.102 & 0.071 & 0.071 \\
\midrule
Dominant & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.449 & 0.450 & 0.328 & 0.331 & 0.248 & 0.252 & 0.260 & 0.282\\
n=80 & 0.341 & 0.341 & 0.218 & 0.218 & 0.158 & 0.160 & 0.125 & 0.131 \\
n=120 & 0.285 & 0.285 & 0.185 & 0.186 & 0.164 & 0.166 & 0.128 & 0.133 \\
\midrule
Group & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.364 & 0.356 & 0.243 & 0.228 & 0.214 & 0.200 & 0.207 & 0.204 \\
n=80 & 0.239 & 0.230 & 0.154 & 0.155 & 0.140 & 0.137 & 0.135 & 0.145 \\
n=120 & 0.206 & 0.205 & 0.148 & 0.147 & 0.130 & 0.126 & 0.132 & 0.122 \\
\bottomrule
\end{tabular}
}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc|cc|cc|cc}
\toprule
Endogenous & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised \\
\midrule
L (Low-rank) & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.376 & 0.317 & 0.261 & 0.260 & 0.256 & 0.261 & 0.249 & 0.260 \\
n=80 & 0.223 & 0.224 & 0.207 & 0.206 & 0.207 & 0.201 & 0.204 & 0.196 \\
n=120 & 0.276 & 0.269 & 0.135 & 0.124 & 0.128 & 0.120 & 0.123 & 0.145 \\
\midrule
L+Sparse & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.371 & 0.364 & 0.274 & 0.266 & 0.242 & 0.235 & 0.235 & 0.234 \\
n=80 & 0.327 & 0.325 & 0.218 & 0.214 & 0.190 & 0.190 & 0.181 & 0.183 \\
n=120 & 0.238 & 0.238 & 0.136 & 0.135 & 0.105 & 0.105 & 0.092 & 0.093 \\
\midrule
Dominant & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.383 & 0.384 & 0.314 & 0.318 & 0.290 & 0.299 & 0.280 & 0.305 \\
n=80 & 0.274 & 0.274 & 0.203 & 0.204 & 0.167 & 0.169 & 0.162 & 0.171 \\
n=120 & 0.211 & 0.211 & 0.144 & 0.144 & 0.119 & 0.120 & 0.114 & 0.119 \\
\midrule
Group & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.311 & 0.305 & 0.237 & 0.218 & 0.219 & 0.202 & 0.217 & 0.224 \\
n=80 & 0.183 & 0.183 & 0.152 & 0.153 & 0.144 & 0.141 & 0.142 & 0.152 \\
n=120 & 0.174 & 0.172 & 0.140 & 0.141 & 0.135 & 0.132 & 0.136 & 0.135 \\
\bottomrule
\end{tabular}
}
\caption*{\footnotesize Denote $\widehat \lambda_W$ and $\widehat \lambda_p$ as the baseline estimate and the plugin estimate calculated using equation \eqref{Eq:Plugin_Estimator} with weight matrices $W=W_0+E$ and $\widehat W$ respectively. The supervised estimate is obtained by equation \eqref{Eq:Supervised_Estimator} with $M=\widehat{W}$ to construct the instruments.
All entries represent RMSE values normalized by the RMSE of the baseline $\widehat \lambda_W$. The endogenous measurement error is simulated with a correlation structure parameter $\rho = 0.7$ as described in \eqref{Eq:Correlated_Errors}. The listed four DGPs correspond to those introduced in Subsection \ref{Subsec:Network_DGPs}.}
\end{table}
{\scriptsize
\begin{table}[htbp]
\centering
\captionsetup{width=0.975\linewidth}
\caption{Relative recovery accuracy of $\widehat W$}
\label{Tab:Relative_RecoveryW_Exogenous}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc|cc|cc|cc}
\toprule
Exogenous & $\widehat W$ & Supervised $\widehat W$& $\widehat W$ & Supervised $\widehat W$& $\widehat W$ & Supervised & $\widehat W$ & Supervised $\widehat W$\\
\midrule
L (Low-rank) & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.320 & 0.320 & 0.320 & 0.321 & 0.320 & 0.323 & 0.320 & 0.328 \\
n=80 & 0.226 & 0.226 & 0.226 & 0.227 & 0.226 & 0.227 & 0.226 & 0.228 \\
n=120 & 0.185 & 0.185 & 0.185 & 0.186 & 0.185 & 0.186 & 0.185 & 0.188 \\
\midrule
L+Sparse & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.392 & 0.391 & 0.392 & 0.391 & 0.392 & 0.391 & 0.392 & 0.392 \\
n=80 & 0.275 & 0.274 & 0.275 & 0.274 & 0.275 & 0.274 & 0.275 & 0.274 \\
n=120 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 \\
\midrule
Dominant & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 \\
n=80 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 \\
n=120 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 \\
\midrule
Group & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.340 & 0.362 & 0.340 & 0.363 & 0.340 & 0.364 & 0.340 & 0.365 \\
n=80 & 0.224 & 0.224 & 0.224 & 0.225 & 0.224 & 0.226 & 0.224 & 0.227 \\
n=120 & 0.185 & 0.187 & 0.185 & 0.187 & 0.185 & 0.187 & 0.185 & 0.188 \\
\bottomrule
\end{tabular}
}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc|cc|cc|cc}
\toprule
Endogenous & $\widehat W$ & Supervised $\widehat W$& $\widehat W$ & Supervised $\widehat W$& $\widehat W$ & Supervised & $\widehat W$ & Supervised $\widehat W$\\
\midrule
L (Low-rank) & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.320 & 0.320 & 0.320 & 0.323 & 0.320 & 0.328 & 0.320 & 0.346 \\
n=80 & 0.226 & 0.226 & 0.226 & 0.227 & 0.226 & 0.228 & 0.226 & 0.232 \\
n=120 & 0.185 & 0.186 & 0.185 & 0.186 & 0.185 & 0.188 & 0.185 & 0.194 \\
\midrule
L+Sparse & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\
\midrule
n=40 & 0.392 & 0.391 & 0.392 & 0.391 & 0.392 & 0.391 & 0.392 & 0.392 \\
n=80 & 0.275 & 0.274 & 0.275 & 0.274 & 0.275 & 0.274 & 0.275 & 0.274 \\
n=120 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 & 0.225 \\
\midrule
Dominant & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 & 0.556 \\
n=80 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 & 0.489 \\
n=120 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 & 0.333 \\
\midrule
Group & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.340 & 0.362 & 0.340 & 0.365 & 0.340 & 0.367 & 0.340 & 0.374 \\
n=80 & 0.224 & 0.225 & 0.224 & 0.226 & 0.224 & 0.228 & 0.224 & 0.233 \\
n=120 & 0.185 & 0.187 & 0.185 & 0.187 & 0.185 & 0.188 & 0.185 & 0.190 \\
\bottomrule
\end{tabular}
}
\caption*{\footnotesize
The relative recovery accuracy is defined as $\|\widehat W-W_0\|_F/\|W-W_0\|_F$, with baseline weight matrix $W=W_0+E$. Endogenous measurement error is simulated with a correlation parameter $\rho = 0.7$, as specified in \eqref{Eq:Correlated_Errors}. The listed four DGPs correspond to those introduced in Subsection \ref{Subsec:Network_DGPs}. }
\end{table}
}
\begin{table}[htbp]
\captionsetup{width=0.975\linewidth}
\caption{Relative recovery accuracy of $(I_n-\widehat\lambda_{\widehat W} \widehat W)^{-1}$ }
\label{Tab:Relative_SpilloverMultiplier_Exogenous}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc|cc|cc|cc}
\toprule
Exogenous & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised \\
\midrule
L (Low-rank) & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.541 & 0.478 & 0.420 & 0.418 & 0.415 & 0.411 & 0.414 & 0.407 \\
n=80 & 0.315 & 0.315 & 0.297 & 0.296 & 0.296 & 0.294 & 0.296 & 0.292 \\
n=120 & 0.447 & 0.440 & 0.288 & 0.281 & 0.277 & 0.269 & 0.271 & 0.277 \\
\midrule
L+Sparse & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.417 & 0.416 & 0.398 & 0.397 & 0.394 & 0.392 & 0.392 & 0.390 \\
n=80 & 0.295 & 0.294 & 0.278 & 0.278 & 0.274 & 0.274 & 0.273 & 0.272 \\
n=120 & 0.243 & 0.243 & 0.227 & 0.226 & 0.223 & 0.223 & 0.222 & 0.222 \\
\midrule
Dominant & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.545 & 0.545 & 0.543 & 0.543 & 0.542 & 0.542 & 0.542 & 0.542 \\
n=80 & 0.507 & 0.507 & 0.504 & 0.505 & 0.505 & 0.505 & 0.504 & 0.505 \\
n=120 & 0.332 & 0.332 & 0.324 & 0.324 & 0.324 & 0.324 & 0.324 & 0.323 \\
\midrule
Group & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.539 & 0.556 & 0.475 & 0.486 & 0.468 & 0.480 & 0.469 & 0.480 \\
n=80 & 0.387 & 0.383 & 0.332 & 0.332 & 0.328 & 0.325 & 0.328 & 0.326 \\
n=120 & 0.329 & 0.331 & 0.300 & 0.304 & 0.295 & 0.297 & 0.296 & 0.295 \\
\bottomrule
\end{tabular}
}
\resizebox{\textwidth}{!}{
\begin{tabular}{lcc|cc|cc|cc}
\toprule
Endogenous & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised & Plug-in & Supervised \\
\midrule
L (Low-rank) & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.474 & 0.421 & 0.409 & 0.408 & 0.406 & 0.407 & 0.405 & 0.404 \\
n=80 & 0.324 & 0.324 & 0.325 & 0.324 & 0.328 & 0.323 & 0.328 & 0.320 \\
n=120 & 0.340 & 0.334 & 0.224 & 0.215 & 0.217 & 0.210 & 0.213 & 0.230 \\
\midrule
L+Sparse & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.409 & 0.407 & 0.393 & 0.391 & 0.388 & 0.386 & 0.387 & 0.385 \\
n=80 & 0.297 & 0.296 & 0.276 & 0.275 & 0.273 & 0.272 & 0.272 & 0.271 \\
n=120 & 0.241 & 0.241 & 0.223 & 0.223 & 0.219 & 0.219 & 0.218 & 0.218 \\
\midrule
Dominant & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.533 & 0.532 & 0.535 & 0.535 & 0.535 & 0.535 & 0.536 & 0.537 \\
n=80 & 0.492 & 0.492 & 0.494 & 0.494 & 0.496 & 0.496 & 0.496 & 0.496 \\
n=120 & 0.320 & 0.320 & 0.321 & 0.321 & 0.321 & 0.321 & 0.321 & 0.321 \\
\midrule
Group & \multicolumn{2}{c}{$T = 1$} & \multicolumn{2}{c}{$T = 5$} & \multicolumn{2}{c}{$T = 15$} & \multicolumn{2}{c}{$T = 50$} \\ \midrule
n=40 & 0.492 & 0.507 & 0.442 & 0.442 & 0.440 & 0.441 & 0.444 & 0.450 \\
n=80 & 0.327 & 0.331 & 0.301 & 0.301 & 0.298 & 0.293 & 0.297 & 0.295 \\
n=120 & 0.315 & 0.317 & 0.297 & 0.301 & 0.295 & 0.295 & 0.296 & 0.294 \\
\bottomrule
\end{tabular}
}
\caption*{\footnotesize The recovery accuracy of the Leontief inverse is defined as $\|[I_n - \widehat \lambda_p \widehat W]^{-1} - (I_n - \lambda_0 W_0)^{-1}\|_F$ and $\|[I_n - \widehat \lambda_s (\widehat L_s+\widehat S_s)]^{-1} - (I_n - \lambda_0 W_0)^{-1}\|_F$ for the plug-in and supervised estimates solved from equations (\ref{Eq:Plugin_Estimator}) and (\ref{Eq:Supervised_Estimator}) respectively. Both are then normalized by the baseline $\|(I_n - \widehat \lambda_{W} W)^{-1} - (I_n - \lambda_0 W_0)^{-1}\|_F$ with $W=W_0+E$ to obatin the relative recovery accuracy. Endogenous measurement error is simulated with a correlation parameter $\rho = 0.7$, as specified in \eqref{Eq:Correlated_Errors}. The DGPs correspond to those introduced in Subsection \ref{Subsec:Network_DGPs}.}
\end{table}
Tables \ref{Tab:Relative_RMSE_Exogenous} - \ref{Tab:Relative_SpilloverMultiplier_Exogenous} present the simulation results using the setting outlined in the previous subsections. We highlight several key takeaways in this setup. First, Tables \ref{Tab:Relative_RMSE_Exogenous} reports relative root mean squared errors (RMSE) for estimates of the spillover scalar parameter $\lambda_0$, where the benchmark is the standard GMM estimator. Observe that both our plug-in and supervised estimators achieve lower RMSEs than the GMM benchmark uniformly across all DGPs and sample sizes. Moreover, the GMM estimates tend to perform worse in the endogenous case and hence there is more improvement in the endogenous case.
Our proposed methods perform better than GMM by approximately 50--80\% for small $n$ and $T$. Our methods achieve at least a 75\% improvement in RMSE for $n=120$, $T=50$. The relative performance of our methods generally improves as the number of nodes in the network $n$ and the number of time periods $T$ increase. This behavior is as predicted by the theory in the cases where the adjacency matrix is assumed to be constant across time.
In terms of the recovery performance of the denoised adjacency matrix, as evidenced in Table \ref{Tab:Relative_RecoveryW_Exogenous}, the plug-in method and the supervised method perform similarly. Moreover, there is little difference between the exogenous and endogenous cases because we use the same observed adjacency matrix.
Finally,
we explore how our methods perform in recovering not just the scalar spillover effect $\lambda_0$, but the full diffusion multiplier matrix given by $(I_n - \lambda_0 W_0)^{-1}$. This quantity is highly important as a policy tool, as it summarizes the cumulative spillover effects of a given economic shock. As shown in Table \ref{Tab:Relative_SpilloverMultiplier_Exogenous}, by providing reliable estimates for both $\lambda_0$ and
$W_0$, our proposed methods achieve better performance in terms of Frobenius norm across DGPs and sample sizes. Although the estimates of $W_0$ might be similar in both the exogenous and the endogenous cases, our method could obtain more reliable estimates of $\lambda_0$ as well as better recovery accuracy in the endogenous case.
\section{Empirical Applications}
\label{Sec:Application}
This section illustrates the use of our estimators for social/spatial interactions. We consider two examples, applying our methods and examining the spillover effects. In each example, we first construct a raw spatial weight matrix and then apply Algorithms \ref{Algo:CD_LRSED} and \ref{Algo:ADMM_GMM_FBS} described in the Appendix to obtain our estimates. To select the penalties, we use the interquartile range (IQR) of the raw weight-matrix entries as a robust scale estimate, denoted by $\widehat\sigma$. The sparse and low-rank penalties are parameterized as $C_S\widehat\sigma\log(n)$ and $C_L\widehat\sigma\sqrt{n}$, respectively, with the constants chosen by the Bayesian Information Criterion (BIC). We estimate the spillover coefficient $\widehat\lambda$ and refer the column sum of the Leontief inverse $(I_n-\widehat\lambda\widehat W)^{-1}$ as the total network influence.
\subsection{Example 1}
In our first example, we examine an application to annual Gross Domestic Product (GDP) growth as our outcome of interest. Following \cite{MRW1992}, the explanatory variables include the annual growth rate of the working-age population defined as individuals aged 15 to 64, denoted as \(X_1\), and the log investment share, \(X_2\), which serves as a proxy for the saving rate. We also include individual and time fixed effects in our specifications. The dataset is downloaded from the World Development Indicators (WDI) and covers a balanced panel of 23~economies from 1970 to~2023. \footnote{These economies are: Australia (AU), Canada (CA), Mainland China (CN),
Denmark (DK), Finland (FI), France (FR), Germany (DE), Hong Kong SAR China (HK), Indonesia (ID), Ireland (IE), Italy (IT), Japan (JP), Korea (KR),
Malaysia (MY), Mexico (MX), Netherlands (NL), New Zealand (NZ), Singapore (SG), Spain (ES), Sweden (SE),
Thailand (TH), United Kingdom (GB), and United States (US).}
Economic activity is well known to exhibit substantial spatial dependence \citep[see, e.g.,][]{Krugman1991}. Neighboring economies tend to trade more, share regional demand and supply shocks, and compete for mobile factors of production. For this reason, we use a geographic‑proximity weight matrix as a natural and predetermined benchmark to study spillovers. The raw spatial weight matrix is constructed as the row-normalized inverse squared distance matrix based on Haversine distances and denoted by $W_2$. \footnote{ For economies \(i\neq j\), we compute the Haversine great-circle distance \(D_{ij}\) between capital cities (in millions of meters) and then form the raw weight matrix by setting the $(i,j)$ entry as $1/D_{ij}^{2}$ for $i\neq j$, and the diagonal entries as 0, and then apply row-normalization. } Note that row-normalization on the weight matrix is standard in the spatial econometrics literature, as it facilitates a clear interpretation of the spatial autoregressive coefficient as the strength of the average influence. All off-diagonal entries of $W_2$ are strictly positive, which contrasts sharply with the sparsity assumptions sometimes invoked in the literature. This dense structure makes $W_2$ a natural candidate for our estimating approach, since many of the small off-diagonal weights likely reflect weak, noise‑contaminated links rather than economically meaningful connections.
We estimate Model (\ref{maineq}) with two-way fixed effects under two scenarios. Panel~A includes only \(X_1\) and \(X_2\) as covariates. We assume that those variables are contemporaneously exogenous and use the spatial lags \(WX_1\) and \(WX_2\) as instruments. Panel~B includes both the original covariates and the contextual variables \(WX_1\) and \(WX_2\) as regressors, using the second-order spatial lags \(W^2X_1\) and \(W^2X_2\) as instruments.
Table \ref{tab:app-gdp-denoising} reports the deviations of the estimated weight matrices from the raw matrix $W_2$. We consider three alternative specifications: a purely sparse structure with $\tau=0.0443$, a purely low-rank structure with $\nu=0.2709$, and a combination of both with $\tau=0.0443$ and $\nu=0.2709$, where the penalties are selected according to BIC respectively.
We first examine the differences between estimated adjacency matrices and the raw matrix $W_2$. We find that the purely low-rank estimator $\widehat L$ exhibits substantial departure from $W_2$, particularly in terms of the matrix maximum norm. Most of the elements in $\widehat L$ are small but not exactly 0, suggesting that a low-rank approximation alone may fail to preserve several important links in the original network. By contrast, the purely sparse estimator $\widehat S$ retains only 96 nonzero entries. Although this produces a more parsimonious network representation, it has the largest deviation from $W_2$ under $l_2$ norm. The low-rank-plus-sparse estimators provide a more balanced representation and both $\|\widehat{W}_2-W_2\|_{\max}$ and $\|\widehat{W}_2-W_2\|_{2}$ remain small relative to the corresponding deviations of the competitors. A further calculation shows that $\|\widehat{W}_2-\widehat S\|_{\max}=0.02$ and $\|\widehat{W}_2-\widehat S\|_{2}=0.11$, indicating that it may capture diffuse connections that may be discarded by pure sparsification. Notably, the supervised low-rank and sparse estimates based on two panel specifications remain very similar to the unsupervised one, which is consistent with our theoretical findings. A further comparison of the heat maps of the raw $W_{2}$, the estimated $\widehat{W}_{2}$, and $\widehat{S}$ is also shown in Figure~\ref{heatmapW2}. We observe that the low-rank plus sparse estimate $\widehat{W}_{2}$ appears to preserve the dominant network structure while allowing for pervasive weak link effects.
\begin{table}[htbp]
\centering
\caption{Comparisons of Estimated Weight Matrices}
\label{tab:app-gdp-denoising}
\begin{tabular}{lrrrr}
\toprule
Estimator & $\|\widehat W-W_2\|_{\max}$ & $\|\widehat W-W_2\|_2$ & Rank($\widehat L_2$) & Nonzero entries($\widehat S_2$)\\
\midrule
$\widehat S$ & 0.044 & 0.385 & -- & 96\\
$\widehat L$ & 0.269 & 0.349 & 14 & --\\
$\widehat W_2$ & 0.044 & 0.289 & 2 & 78\\
$\widehat W_s$ (Panel A) & 0.047 & 0.301 & 2 & 115\\
$\widehat W_s$ (Panel B) & 0.045 & 0.299 & 2 & 116\\
\bottomrule
\end{tabular}
\caption*{\footnotesize Note: We calculate the differences between each estimated weight matrix and the raw weight matrix $W_2$. We consider the purely sparse matrix estimate $\widehat S$, the purely low-rank matrix estimate $\widehat L$, the low-rank plus sparse matrix estimate $\widehat W_2$ and its supervised estimates $\widehat W_s$ under model (\ref{maineq}) in Panel A (without contextual effects) or Panel B (with contextual effects). Rank ($\widehat L$) reports the rank of the low-rank component, and nonzero entries ($\widehat S$) refer to the number of nonzero entries in the sparse component when the corresponding decomposition is available.}
\end{table}
\begin{figure}[htbp]
\includegraphics[width=1\textwidth]{matrix_heatmap_raw_w2_est_w2_est_s2.pdf}
\caption{\footnotesize Heatmaps of different spatial weight matrices associated with 23 economies, including the raw weight matrix $W_2$ (left), the estimated low-rank and sparse matrix $\widehat W_2$ (center), and the estimated purely sparse matrix $\widehat S$ (right). \label{heatmapW2}}
\end{figure}
In the regression analysis, we evaluate the impact of the investment rate and working-age population growth,
and report the estimated spillover coefficient $\widehat{\lambda}$ based on different weight matrices in Table~\ref{tab:app-gdp-spillovers}. For each candidate weight matrix $W$, we also compute the column sums of the Leontief inverse $(I_n-\widehat{\lambda}\widehat W)^{-1}$ and designate the economy with the largest column sum as the key player.
Relative changes compared to the raw $W_2$ benchmark are also reported.
In Panel A (without contextual effects), the purely sparse estimator $\widehat{S}_2$ only retains the strong connections and reduces $\widehat{\lambda}$ by 16.7\% and total influence (i.e., the corresponding column sum from the Leontief matrix, see above) by 11.3\% compared to those obtained from the raw $W_2$. Conversely, the purely low-rank estimator $\widehat{L}_2$, amplifies network connections via aggregating pervasive weak links and increases $\widehat{\lambda}$ by 11.3\% and total influence by 12.0\% compared to those obtained from the raw $W_2$. These results indicate that the sparse and low-rank components might capture distinct aspects of the network. The sparse component filters out scattered weak links, whereas the low-rank component summarizes pervasive dependence. Correspondingly, the low-rank-plus-sparse estimator $\widehat{W}_2$ produced more moderate deviations from the raw benchmark.
The benefits of our approach become more pronounced in Panel B with contextual effects. The raw $W_2$ produces a negative estimate $\widehat{\lambda}$ that is much lower than the our estimates. The sparse estimator and the low-rank-plus-sparse estimator produce moderate positive values that might be more reasonable. As higher-order spatial lags can magnify noise in the raw matrix, this reversal suggests that our approach can improve the plausibility and stability of the estimated spillover effects.
\begin{table}[htbp]
\centering
\caption{GDP-Growth GMM Spillover Estimates and Influence Analysis}
\label{tab:app-gdp-spillovers}
\begin{tabular}{lcccccc}
\toprule
Weight matrix & $\widehat\lambda$ & $\Delta \widehat\lambda$ & Key player & Total influence & $\Delta influence$ \\
\midrule
\multicolumn{4}{l}{\textit{Panel A: without contextual effects}}\\\\
Raw $W_2$ & 0.377 & 0.0\% & SGP & 2.211 & 0.0\%\\
$\widehat S_2$ & 0.314& -16.7\% & SGP & 1.961 & -11.3\%\\
$\widehat L_2$ & 0.420 & 11.3\% & SGP & 2.475 &12.0\%\\
$\widehat W_2$ & 0.354 & -6.2\% & SGP & 2.147 &-2.9\%\\
$\widehat W_s$ & 0.356& -5.6\% & SGP & 2.178 & -1.5\%\\
\addlinespace
\midrule
\multicolumn{4}{l}{\textit{Panel B: with contextual effects}}\\\\
Raw $W_2$ & -0.109 & 0.0\% & MEX & 0.985 & 0.0\%\\
$\widehat S_2$ & 0.124 & 214.2\% & SGP & 1.290& 31.0\%\\
$\widehat L_2$ & -0.052 & 52.3\% & MEX & 0.993 &0.8\%\\
$\widehat W_2$ & 0.095 & 187.6\% & SGP & 1.212 &23.0\%\\
$\widehat W_s$ & 0.095 & 186.9\% & SGP & 1.211 & 22.9\%\\
\addlinespace
\bottomrule
\end{tabular}
\caption*{\footnotesize We report the estimated spillover coefficient $\widehat{\lambda}$ obtained using different weight matrices. The raw $W_2$
matrix is the row-normalized inverse squared distance weight matrix. We consider its sparse estimator $\widehat S$, its low-rank estimator $\widehat L$, and its low-rank plus sparse estimator $\widehat W_2$ and its supervised version $\widehat W_s$.
Panels A and B correspond to the scenarios without and with contextual effects respectively. We record the maximum column sum of the Leontief inverse $(I_n-\widehat{\lambda} \widehat W)^{-1}$ and refer to it as the total influence of key player. The relative changes of the spillover effect and the total influence are also computed using the raw matrix $W_2$ as the baseline. }
\end{table}
Moreover, the decomposition also yields interesting practical insights. Compared with $\widehat S_2$, we find that including the low-rank component in $\widehat W_2$ primarily rescales the overall magnitude of the spillover influence vector rather than altering the key player. This pattern indicates that, while the identity of the key player remains robust to the decomposition strategy, the estimated strength of network spillovers depends materially on whether the low-rank component is retained, underscoring the practical relevance of the full low-rank plus sparse decomposition over the purely sparse estimator. Figure~\ref{GR2} presents the graphical representations of $\widehat{W}_{2}$, $\widehat{L}_{2}$, and $\widehat{S}_{2}$, respectively. $\widehat W_2$ admits the decomposition $\widehat W_2=\widehat S_{\widehat W_2}+\widehat L_{\widehat W_2}$ and comparing the leading eigenvectors of $\widehat W_2$, $\widehat L_{\widehat W_2}$ and $\widehat S_{\widehat W_2}$, we find that ${\bf v}_{\widehat{W}_2}\approx 0.56\,{\bf v}_{\widehat{L}_{\widehat W_2}}+0.48\,{\bf v}_{\widehat{S}_{\widehat W_2}}$ from a regression perspective as discussed in equation (\ref{rhoLS}). Both the low-rank and the sparse components contribute substantially in this application, indicating that the identification of influential nodes depends not only on a few strong connections, but also on the cumulative effect of widespread weak links. Notably, as shown in Figure~\ref{fig:networksparseE1}, ${\bf v}_{\widehat{S}_{\widehat W_2}}$ has many zero entries, and its nonzero entries concentrate on the Asian economies, whereas ${\bf v}_{\widehat{L}_{\widehat W_2}}$ loads on all economies but assigns slightly smaller weights to Western economies. One possible interpretation is that the sparse eigenvector captures localized regional clustering centered on Asia, while the low-rank eigenvector reflects a global common factor that is slightly attenuated for Western economies in accounting for the broader geographic diffusion inherent in the low-rank structure.
\begin{figure}[htbp]
\includegraphics[width=1\textwidth]{tenet_qgraph_est_w2_est_l2_est_s2.pdf}
\caption{\footnotesize Graphical representation of different spatial weight matrix associated with 23 economies, including the estimated low-rank and sparse one $\widehat W_2$ (left), the estimated low-rank one $\widehat L$ (center) and the estimated purely sparse one $\widehat S$ (right). \label{GR2}}
\end{figure}
\begin{figure}[htbp]
\includegraphics[width=1\textwidth]{eigen_components_lrs.pdf}
\caption{\footnotesize Plots of eigenvector centrality associated with different weight matrices. Denote the low-rank and sparse components of our estimator $\widehat W_2$ as $\widehat L$ and $\widehat S$, i.e. $\widehat W_2=\widehat L_{\widehat W_2}+\widehat S_{\widehat W_2}$. From left to right, we provide the eigenvector centrality plots of $\widehat W_2$, $\widehat L_{\widehat W_2}$ and $\widehat S_{\widehat W_2}$ respectively. Orange dots mark economies whose entries in the leading eigenvector of $\widehat{S}_{\widehat W_2}$ are nonzero, while red dots mark economies whose entries in the leading eigenvector of $\widehat{L}_{\widehat W_2}$ are smaller than $0.2$. The economies corresponding to red dots are Canada, Denmark, Finland, France, Germany, Ireland, Italy, Mexico, Netherlands, Spain, Sweden, United Kingdom, and United States, and those corresponding to orange dots are Australia, China, Hong Kong SAR China, Indonesia, Japan, Korea Rep., Malaysia, New Zealand, Singapore, Thailand.
}
\label{fig:networksparseE1}
\end{figure}
\subsection{Example 2}
\label{sec:geo_application}
In our second example, we apply our methods to examine tax competition among the 48 contiguous U.S.\ states. This application was originally studied by \cite{Besley1995} using data from 1962 to 1988, with a pre‑specified geographic weight matrix $W_g$ defined as the row‑normalized 0‑1 neighborhood adjacency matrix setting 1 to state pairs that are geographically adjacent and zero, otherwise. \cite{Besley1995} emphasize a political‑economy channel for tax‑setting spillovers, as policymakers respond to voters who benchmark their policies against those of peer states. Consequently, states may face both political and economic incentives to align their fiscal policies with those of their neighbors. This dataset has recently been extended to cover $1962$--$2015$ ($T = 53$) and analyzed by \cite{Paula2023}. We obtain the spatial weight matrix $W_2$, which is a row-normalized inverse squared distance matrix based on Haversine distances between state capitals, as the observed weight matrix.
In addition to endogenous spatial effects captured by $WY$, the covariates include state-level characteristics such as per-capita income, the unemployment rate, and the proportions of young and elderly populations, without and with their contextual covariates $WX$ in Panel A and B respectively. We also include individual and time fixed effects in our specifications. Lagged neighbor income and lagged neighbor unemployment are used as instruments as in the original paper by \cite{Besley1995}.
We consider three specifications, which yield a purely sparse matrix estimate $\widehat{S}_2$, a purely low‑rank matrix estimate $\widehat{L}_2$, and a low‑rank plus sparse estimate $\widehat{W}_2$. The purely sparse estimate $\widehat{S}_2$ captures the strong links and retains 90 nonzero entries. The estimate $\widehat{L}_2$ is a zero matrix under the selected penalty, suggesting that the data do not provide enough evidence of a low-rank component. The combined estimator $\widehat{W}_2$ coincides with the sparse‑only estimator $\widehat{S}_2$, and hence we only report results about $\widehat{W}_2$ in the subsequent analysis. The supervised variants $W_s$ in Panels A and B are very close to those of the unsupervised estimate $\widehat{W}_2$.
\begin{figure}[htbp]
\includegraphics[width=1\textwidth]{matrix_heatmap_raw_wg_raw_w2_est_w2.pdf}
\caption{\footnotesize Heatmaps of adjacency matrices associated with 48 states. From left to right, this figure provides the heatmaps associated with the row-normalized 0-1 geographic adjacency matrix $W_g$, the row-normalized inverse squared distance matrix $W_2$, the estimated low-rank and sparse matrix $\widehat W_2$.\label{heatmap48}}
\end{figure}
\begin{figure}[htbp]
\includegraphics[width=1\textwidth]{tenet_qgraph_raw_wg_raw_w2_est_w2.pdf}
\caption{\footnotesize Graphical representations of different weight matrices. From left to right, this figure provides the network representation for the row-normalized neighborhood weight matrix $W_g$, the row-normalized inverse squared distance matrix $W_2$ and the estimated low-rank and sparse weight matrix $\widehat W_2$.
\label{fig:three_graphs2}}
\end{figure}
Figures~\ref{heatmap48} and~\ref{fig:three_graphs2} provide heatmaps and graphical representations of the raw geographic weight matrices $W_g$, $W_2$ and the estimate $\widehat W_2$, respectively. We find that the raw inverse geographic distance matrix $W_2$ is much denser than the raw geographic adjacency matrix $W_g$. It may be more vulnerable to measurement errors, and thus exaggerate the overall degree of network dependence and obscure the dominant transmission channels. By contrast, the estimated network discards many weak, noise‑dominated connections, thus providing a clearer visualization of the significant links and improving interpretability.
Table~\ref{tab:geo_ranktop} provides a more detailed characterization of the state-level network structure. States are sorted according to their column sums, which measure the aggregate strength of their outward connections. Massachusetts is identified as the most influential state under all three weight matrices. Its out-degree is 1.70 under the row-normalized adjacency matrix $W_g$, 1.61 under the row normalized inverse squared distance matrix $W_2$, and 0.72 under the estimate $\widehat W_2$. The estimated network therefore preserves the leading position of Massachusetts while substantially reducing the magnitude of its estimated outward influence.
This suggests that our estimate does not fundamentally alter the ranking of the major network hubs, but removes potentially noisy links that may inflate the overall strength of network connections rather than incorporating them into the core neighborhood topology. By contrast, the ranking based on $W_g$ differs more noticeably because the geographical adjacency matrix only captures direct neighboring relationships, whereas $W_2$ allows for distance-decaying connections between all states.
\begin{table}[ht]
\centering
\caption{Top 4 States Ranked by Out-degree}
\label{tab:geo_ranktop}
\begin{tabular}{r|lr|lr|lr}
\toprule
& \multicolumn{2}{c|}{Raw $W_g$} & \multicolumn{2}{c|}{Raw $W_2$} & \multicolumn{2}{c}{$\widehat W_2$}\\
Rank & State & Out-degree & State & Out-degree & State & Out-degree\\
\midrule
1 & MA & 1.70 & MA & 1.61 & MA & 0.72\\
2 & GA & 1.62 & MD & 1.57 & MD & 0.64\\
3 & TN & 1.58 & RI & 1.53 & RI & 0.63\\
4 & ID, NH & 1.53 & DE & 1.52 & DE & 0.63\\
\bottomrule
\end{tabular}
\caption*{\footnotesize Notes: We consider three choices of the weight matrix, including the raw row-normalized adjacency matrix $W_g$, the raw row-normalized inverse squared distance matrix $W_2$, and the estimate $\widehat W_2$. For each matrix, we compute the out-degree (column sums) and report the four states with the largest values in descending order. The abbreviations MA, GA, TN, ID, NH, MD, RI, DE stand for Massachusetts, Georgia, Tennessee, Idaho, New Hampshire, Maryland, Rhode Island and Delaware, respectively.
}
\end{table}
\begin{table}[htbp]
\centering
\caption{Spillover Estimates and Influence Analysis}
\label{tab:geo-spillovers}
\begin{tabular}{lrrlrr}
\toprule
Weight matrix & $\widehat\lambda$ & $\Delta\widehat\lambda$ & Key region & Region impact & $\Delta$ Region impact\\
\midrule
\multicolumn{6}{l}{\textit{Panel A: without contextual effects}}\\\\
Raw $W_g$ & 0.658 & 30.3\% & Mountain & 0.181 & 3.9\%\\
Raw $W_2$ & 0.505 & 0.0\% & SA & 0.174 & 0.0\% \\
$\widehat W_2$ & 0.220 & -56.5\% & SA & 0.168 & -3.3\%\\
$\widehat W_s$ & 0.220 & -56.5\% & SA & 0.168 & -3.3\%\\
\addlinespace
\midrule
\multicolumn{6}{l}{\textit{Panel B: with contextual effects}}\\\\
Raw $W_g$ & 1.006 & -11.0\% & NA & NA & NA\\
Raw $W_2$ & 1.130 & 0.0\% & NA & NA & NA\\
$\widehat W_2$ & 0.811 & -28.3\% & SA & 0.177 & NA\\
$\widehat W_s$ & 0.811 & -28.2\% & SA & 0.177 & NA\\
\bottomrule
\end{tabular}
\caption*{\footnotesize Notes: We report the estimates of the spillover coefficients obtained using different spatial weight matrices. $W_g$ and $W_2$ denote the row-normalized 0-1 adjacency matrix and the row-normalized inverse squared distance weight matrix respectively. $\widehat W_2$ and $\widehat W_s$ are the estimator and the supervised estimator constructed from the raw weight matrix $W_2$. If $|\widehat \lambda|<1$, region impact is computed by first obtaining the column sums of the Leontief inverse,
and then aggregating them according to the nine geographical divisions. The divisions with the largest total impact are reported as the key region. The relative changes of the spillover and the region impact are reported using the results associated with the raw $W_2$ matrix as the baseline when available. }
\end{table}
Table~\ref{tab:geo-spillovers} reports the estimated spillover coefficient $\widehat{\lambda}$ and the regional influence analysis using different weight matrices. We consider two scenarios, Panel A (without $WX$) and Panel B (with $WX$), respectively. In both panels, two raw weight matrices $W_g$ and $W_2$ produce quite similar results, and they both exhibit much larger or even explosive values of $\widehat{\lambda}$ in Panel B. Such large values are likely driven by measurement error or noise in the raw adjacency matrices, yielding exaggerated spillover estimates that could be unreliable and potentially misleading. In contrast, the estimators, including the low-rank-plus-sparse estimator $\widehat{W}$, and its supervised variants, produce substantially smaller and more stable $\widehat{\lambda}$ values, and consistently identify the South Atlantic as the key region. Recall that each weight matrix is row-normalized prior to estimation.
By the Perron–Frobenius theorem, the spectral radius of a row-normalized non‑negative matrix is 1, and hence the standard Leontief-stability condition reduces to $|\widehat{\lambda}|<1$ \citep{LeSage2009}. Explosive spillover estimates can cause some elements of the Leontief inverse
to become negative and hard to interpret.
These results indicate that our approach might yield more plausible and interpretable spillover effects, avoiding the inflated influence suggested by raw network matrices.
\begin{table}[ht]
\centering
\caption{Top Four States Most Strongly Influenced by Massachusetts}
\label{tab:n520-app-gmm-ma-influence}
\begin{tabular}{lllll}
\toprule
Weight matrix & 1st & 2nd & 3rd & 4th\\
\midrule
\multicolumn{5}{l}{\textit{Panel A: Without contextual effects}}\\\\
Raw $W_g$ & MA (1.30) & RI (0.59) & CT (0.48) & VT (0.45)\\
Raw $W_2$ & MA (1.13) & RI (0.34) & NH (0.24) & CT (0.17)\\
$\widehat W_2$ & MA (1.04) & RI (0.19) & NH (0.17) & ME (0.06)\\
$\widehat W_s$ & MA (1.04) & RI (0.19) & NH (0.17) & ME (0.06)\\
\midrule
\multicolumn{5}{l}{\textit{Panel B: With contextual effects}}\\\\
$\widehat W_2$ & MA (2.51) & RI (1.88) & NH (1.82) & ME (1.48)\\
$\widehat W_s$ & MA (2.52) & RI (1.88) & NH (1.82) & ME (1.48)\\
\bottomrule
\end{tabular}
\caption*{\footnotesize Notes: We evaluate the column corresponding to Massachusetts in the Leontief inverse
and interpret its entries as the total influence that a shock originating from Massachusetts exerts. We report the four largest values in parentheses along with the corresponding states. Four choices of the weight matrix are considered: the raw adjacency matrix $W_g$, the raw row-normalized inverse squared distance matrix $W_2$, the estimator $\widehat W_2$ and the supervised estimator $\widehat W_s$ . The abbreviations MA, RI, CT, VT, NH, ME stand for Massachusetts, Rhode Island, Connecticut, Vermont, New Hampshire and Maine, respectively.
We do not include the results for Raw $W_{2}$ and Raw $W_{g}$ in Panel B because the estimated spillover effects do not satisfy the Leontief-stable condition.
}
\end{table}
Table~\ref{tab:n520-app-gmm-ma-influence} further examines the propagation of shocks originating from Massachusetts. We calculate the Leontief Inverse
and examine the four largest elements in the column that corresponding to Massachusetts. The two states most strongly influenced by a shock originating from Massachusetts are Massachusetts itself and Rhode Island. However, the magnitude of the estimated network influence is attenuated.
When contextual effects $WX$ are incorporated, the Leontief influence values increase substantially. Under the estimate $\widehat W_2$, the influence measure for Massachusetts rises from 1.04 to 2.51, while the corresponding values for Rhode Island, New Hampshire, and Maine increase to 1.88, 1.82, and 1.48, respectively. Overall, the results consistently identify Massachusetts as an important regional hub and indicate that contextual effects could considerably amplify the propagation of state-level shocks. Our approach retains stronger local transmission channels while suppressing weaker long-range associations that may be sensitive to measurement error.
\section{Conclusions}
\label{Sec:Conclusion}
This paper introduces a robust framework for studying interaction effects in social and spatial networks under noisy adjacency matrices. By leveraging the low-rank and sparse structure of real-world networks, our de-noising approach, combining LASSO and nuclear norm penalization, significantly improves estimation accuracy. We propose two estimation procedures—a two-step and a one-step supervised GMM estimator—that outperform standard GMM by $50-80\%$ in RMSE under noise and endogeneity.
Applying our proposed methodology to examine the international spillover of economic growth and the tax competition across U.S. states, we find pronounced discrepancy in estimating the spillover effects, which highlights the essential role of accurate network denoising. Failing to distinguish genuine connectivity from measurement noise can systematically distort spillover coefficients as well as the strength and direction of cross-unit interactions. Moreover, our decomposition framework might provide a new perspective on complex spatial dependencies by disentangling global and local components. This decomposition not only sharpens the economic interpretation but also improve econometric inference, thereby contributing to credible spatial econometric analysis.
\section{Declaration of generative AI and AI-assisted technologies in the writing process}
During the preparation of this work the authors used AI to suggest edits to some lengthy passages for concision and clarity. After using these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.
\clearpage
\bibliographystyle{apalike}
\bibliography{LowRank_new2}