Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
79,680 characters · 14 sections · 52 citation commands
Limit Theorems for Network Data without Metric Structure
\global\long\global\long\global\long\global\long\global\long\global\long\global\long\foreignlanguage{english}{ \global\long}
Keywords: functional dependence, network data, law of large numbers, central limit theorem, concentration inequality, spatial autoregressive model
Network-indexed dependence arises in a broad class of statistical and econometric models. To study estimators and test statistics in such settings, one needs probabilistic tools, such as moment inequalities, concentration inequalities, laws of large numbers, and central limit theorems, to accommodate dependence structures that are neither sequentially ordered nor naturally organized by a metric. This challenge is especially acute in nonlinear network models, where martingale methods available for linear specifications are typically not applicable.
A substantial body of literature has studied weak dependence concepts for spatial and network-indexed data. For linear spatial autoregressive (SAR) models, martingale difference array arguments play a central role h._kelejian_asymptotic_2001,lee2004asymptotic,lee_gmm_2007,lee2022qml. For more general dependent arrays, the literature develops strong-mixing and $\phi$-mixing spatial random fields jenish2009central, spatial near-epoch dependence (NED) jenish2012spatial, spatial functional dependence wu2023application, spatial mixingales based on model-dependent random metrics kuersteiner2019limit, and network adaptations of $\psi$-dependence kojevnikov_limit_2021; see also kuersteiner_limit_2013,KuersteinerPrucha2020,leung_weak_2019,leung_normal_2019. These contributions have substantially advanced asymptotic theory for dependent data and have also been applied in a variety of settings; see, for example, xu2015maximum,gao2025Causal,kojevnikov2021bootstrap,leung2022Causal,XU201896,qu_estimating_2015,liu2022robust. However, most of them rely on a metric, a geodesic distance, or a model-dependent random metric. Accordingly, dependence is characterized through distance-based shells and decay with distance.
This reliance on metric structure constitutes a major bottleneck for network data. Many network models employed in statistics and econometrics do not possess a natural embedding in Euclidean space. Even when a metric can be imposed, it may not accurately reflect the actual propagation of dependence. The problem is especially acute in networks with short graph distances, high connectivity, or moderate diameter. For example, consider an Erd\H{o}s-R\'{e}nyi model with mean degree $D_{n}\propto n^{\alpha}$ ($0<\alpha<1)$. Classical results on the diameter of Erd\H{o}s-R\'{e}nyi graphs imply that graph diameter is of constant order with probability approaching one in this setting klee1981diameters. In such settings, there are only a few relevant distance shells, so weak dependence cannot be naturally organized via gradual decay over expanding metric neighborhoods. What is needed, therefore, is a notion of weak dependence that does not require any metric structure, and can be directly verified from primitive conditions.
This restriction is substantive, not merely technical. In many network applications, it is unclear whether the nodes admit any meaningful Euclidean embedding. For instance, xu2022dynamic applied spatial NED theory to quantile regression in a dynamic network model for stocks traded on the NYSE and NASDAQ, even though it is not obvious that such assets should be viewed as points in a Euclidean space. Likewise, latent-space representations for social networks hoff2002Latent,athreya2018Statistical,smith2019Geometry provide useful modeling devices, but the underlying metric is typically imposed for tractability rather than being dictated by the dependence mechanism itself. More generally, in online and social networks, nodes that are far apart in any Euclidean sense may still interact strongly. These examples indicate that metric-based weak dependence conditions can impose a substantive modeling restriction in applications where no natural metric structure is available.
There is also a methodological mismatch between estimation and asymptotic theory. Many standard procedures, including quasi-maximum likelihood estimation and the generalized method of moments, can be formulated for network data without any metric structure. However, available asymptotic justifications often rely on metric-based weak dependence conditions. For instance, the robust estimator for SAR models proposed by liu2022robust ultimately invokes the spatial NED framework of jenish2012spatial. A partial step in this direction is kuersteiner2019limit, who replaces a fixed metric with a model-dependent random metric. This underscores the need for a notion of weak dependence that is both metric-free and directly verifiable from the model.
We address this problem by extending the concept of functional dependence measure (FDM) to network data. This measure is defined by perturbing the underlying innovations and quantifying the resulting effect on each component of the system. The construction is entirely node-based and does not rely on Euclidean distance, geodesic distance, or any other metric. It extends the functional dependence concept of wu2005nonlinear, ELMACHKOURI20131, and wu2023application from time series and spatial processes to general network dependence. Unlike strong mixing and related notions, the proposed FDM avoids suprema over sub-$\sigma$-fields and is therefore substantially easier to verify in specific network models.
The contributions of this paper are fourfold. First, we introduce a metric-free functional dependence framework for network data. This extends the functional dependence paradigm to dependence structures that are intrinsic to networks and not organized by any metric. Second, under this framework we develop a general probabilistic toolkit for network data. In particular, we establish a moment inequality, a concentration inequality, a weak law of large numbers, and central limit theorems. Third, for the CLT theory, we introduce a second-order functional dependence measure and derive sufficient conditions directly in terms of first- and second-order dependence measures. These conditions are invariant under relabeling of the nodes and are designed for nonlinear systems, thereby going beyond martingale-based arguments available for linear models. Fourth, we show that the proposed dependence concept is verifiable and useful in concrete models. We compute the dependence measures for nonlinear SAR models, derive primitive sufficient conditions for LLNs and CLTs, and verify these conditions under several network designs, including dominant-unit structures and random graph models. In addition, we establish the consistency of the maximum likelihood estimator for the SAR Tobit model using the tools developed in this paper.
The remainder of the paper is organized as follows. Section (ref) introduces the functional dependence measure and the associated notion of $(L^{p},q)$-functional dependence. Section (ref) develops the general theoretical tools, including a moment inequality, a concentration inequality, a weak law of large numbers, and central limit theorems; it also introduces a second-order functional dependence measure for the CLT analysis. Section (ref) applies the framework to nonlinear SAR models, derives primitive sufficient conditions for the LLN and CLTs, and verifies these conditions under several network designs; it also studies the SAR Tobit model. Section (ref) investigates how the functional dependence measure behaves under common transformations. Section (ref) concludes. Proofs of the main results are collected in Appendix (ref), and additional proofs are deferred to the supplementary appendix.
Notation. The set of positive integers is $\mathbb{N}\equiv\{1,2,\cdots\}$, $\mathbb{Z}$ denotes the set of integers, and $\mathbb{Z}^{d}$ is the $d$-dimensional integer lattice. For any column vector $x=\left(x_{1},x_{2},\ldots,x_{d}\right)'\in\mathbb{R}^{d}$, where $\mathbb{R}^{d}$ is the $d$-dimensional Euclidean space, $\left\Vert x\right\Vert =(x'x)^{1/2}$ denotes its Euclidean norm, $\left\Vert x\right\Vert _{\infty}=\max_{1\leq k\leq d}\left|x_{k}\right|$ represents its infinity norm, and $\left\Vert x\right\Vert _{1}=\sum^{d}_{k=1}\left|x_{k}\right|$ is its 1-norm. For any matrix $A=\left(a_{ij}\right)_{n\times m}$, its Frobenius norm is $\left\Vert A\right\Vert \equiv\sqrt{\sum_{ij}a^{2}_{ij}}$, its 1-norm (also called column sum norm) is defined as $\left\Vert A\right\Vert _{1}=\max_{1\leq j\leq m}\sum^{n}_{i=1}\left|a_{ij}\right|$, and its infinity norm (also called row sum norm) is defined as $\left\Vert A\right\Vert _{\infty}=\max_{1\leq i\leq n}\sum^{m}_{j=1}\left|a_{ij}\right|$. For any symmetric matrix $A$, $\min\mathrm{eig}(A)$ denotes its minimum eigenvalue. For any two sequences $a_{n}$ and $b_{n}$, $a_{n}\sim b_{n}$ if and only if $a_{n}/b_{n}\to1$ as $n\to\infty$. Let $\left(\Omega,\mathcal{F},\mathbb{P}\right)$ be a probability space. Denote probability and expectation by $\mathbb{P}$ and $\mathbb{E}$, respectively. For any random vector (matrix) $X$, its $L^{p}$ norm is defined as $\left\Vert X\right\Vert _{L^{p}}\equiv\left[\mathbb{E}\left(\left\Vert X\right\Vert ^{p}\right)\right]^{1/p}$ for any constant $p\geq1$. The symbols “$\xrightarrow{\mathbb{P}}$” and “$\ensuremath{\xrightarrow{d}}$” denote convergence in probability and convergence in distribution, respectively. For two positive non-random sequences $a_{n},b_{n}$ and random vector sequence $X_{n}$, $X_{n}=o_{\mathbb{P}}(a_{n})$ means $\mathbb{P}\left(\left\Vert X_{n}\right\Vert >\epsilon a_{n}\right)\to0$ as $n\to\infty$ for any $\epsilon>0$ and $X_{n}=O_{\mathbb{P}}(a_{n})$ means for any $\epsilon>0$, there exists a constant $M>0$ such that $\limsup_{n\to\infty}\mathbb{P}\left(\left\Vert X_{n}\right\Vert \geq Ma_{n}\right)<\epsilon$. For any sub-$\sigma$-field $\mathcal{C}$ of $\mathcal{F}$, denote the conditional probability, conditional expectation, and conditional variance by $\mathbb{P}_{\mathcal{C}}\left(\cdot\right)\equiv\mathbb{P}\left(\cdot\mid\mathcal{C}\right)$, $\mathbb{E}_{\mathcal{C}}\left(\cdot\right)\equiv\mathbb{E}\left(\cdot\mid\mathcal{C}\right)$, and $\mathrm{Var}_{\mathcal{C}}\left(\cdot\right)\equiv\mathrm{Var}\left(\cdot|\mathcal{C}\right)$, respectively. Besides, for a sub-$\sigma$-field $\mathcal{G}$ of $\mathcal{F}$, we denote $\mathbb{E}_{\mathcal{C}}\left(\cdot|\mathcal{G}\right)\equiv\mathbb{E}\left(\cdot\mid\mathcal{G}\bigvee\mathcal{C}\right)$, where $\mathcal{G}\bigvee\mathcal{C}$ denotes the $\sigma$-field generated by $\mathcal{G}$ and $\mathcal{C}$. For a random vector (matrix) $X$, let $\left\Vert X\right\Vert _{L^{p},\mathcal{C}}\equiv\left[\mathbb{E}_{\mathcal{C}}\left(\left\Vert X\right\Vert ^{p}\right)\right]^{1/p}$.
Consider a network with $n$ nodes (also referred to as individuals or units). For simplicity, name these $n$ nodes as $1,2,\cdots,n$, and denote $[n]=\{1,2,\cdots,n\}$. Although we use numeric labels, the order is arbitrary: random variables associated with individuals 1 and 2 may be independent, while those associated with individuals 1 and $n$ could be strongly correlated.
Let $\left(\Omega,\mathcal{F},\mathbb{P}\right)$ be the underlying probability space, and $\mathcal{C}_{n}$ be a sub-$\sigma$-field of $\mathcal{F}$. Associated with each $i\in[n]$ is a random vector $e_{i,n}\in\mathbb{R}^{p_{e}}$, where $p_{e}\in\mathbb{N}$ is a constant. Suppose that $e_{i,n}$'s are conditionally independent given $\mathcal{C}_{n}$, but they need not be identically distributed. Denote $e_{n}=(e_{1,n},e_{2,n},\cdots,e_{n,n})$. Suppose that $Y_{1,n},\cdots,Y_{n,n}$ are $n$ random vectors generated by $e_{n}$:
where the functions $F_{1,n},\ldots,F_{n,n}$ can be random functions (e.g. associated with a random network weights matrix, see Section (ref)) and we assume these functions are measurable with respect to $\mathcal{C}_{n}$. The triangular array $\{Y_{j,n}:1\leq j\leq n,\ n\geq1\}$ exhibits network dependence. Although $Y_{j,n}$ can be vector-valued, we restrict attention to the real-valued case for simplicity and without loss of generality. See Remark (ref) for details.
Suppose that conditional on $\mathcal{C}_{n}$, $e^{*}_{i,n}$ is an independent and identically distributed (i.i.d.) copy of $e_{i,n}$, and $e^{*}_{i,n}$ is independent of $e_{j,n}$ for all $j\neq i$. Denote $Y_{j,n,i}$ as the coupled version of $Y_{j,n}$ with $e_{i,n}$ replaced by $e^{*}_{i,n}$, i.e., $Y_{j,n,i}\equiv F_{j,n}\left(e_{1,n},\cdots,e_{i-1,n},e^{*}_{i,n},e_{i+1,n},\cdots,e_{n,n}\right)$. Now, we are ready to introduce the definition of functional dependence measure.
Comparison with existing weak dependence concepts. We now compare our FDM with several well-known notions of weak dependence in the literature.
(1) The index sets considered in wu2005nonlinear,wu2023application,ELMACHKOURI20131,jenish2009central,jenish2012spatial,kojevnikov_limit_2021 are subsets of $\mathbb{Z}^{d}$, $\mathbb{R}^{d}$, or a general metric space. In contrast, our definition does not require the index $i$ to belong to any metric space. This makes our approach applicable to network data, where individuals are not necessarily embedded in any metric space.
(2) Compared with strong mixing, the FDM is easier to compute because the coupled version $Y_{j,n,i}$ is constructed explicitly, and the relevant $L^{p}$-norm is easier to evaluate. In contrast, the strong mixing coefficient requires the calculation of a supremum over two $\sigma$-fields, and thus quite challenging (wu2005nonlinear and xu2021Conditions).
(3) Compared to the spatial NED, FDM is more conveniently to calculate under any $L^{p}$-norm, while the $L^{p}$-NED property is often tractable only for $p=2$, because it typically relies on the fact that the conditional expectation is the best predictor under $L^{2}$-distance.
In this section, we establish a moment inequality, a concentration inequality, laws of large numbers (LLN), and central limit theorems (CLTs) using the FDM. These tools are essential for deriving asymptotic theory in statistical and econometric models.
The basic idea of most proofs is to express $\sum^{n}_{j=1}\left(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n}\right)$ as the summation of a martingale difference array (MDA) and then apply the theories of MDA. Thus, we first define an increasing sequence of $\sigma$-fields. For $i\in[n]$, let $\mathcal{F}_{i,n}\equiv\sigma\left(e_{j,n}:j\leq i,j\in[n]\right)$ be the $\sigma$-field generated by $e_{1,n},\cdots,e_{i-1,n},e_{i,n}$, and let $\mathcal{F}_{0,n}\equiv\left\{ \Omega,\emptyset\right\} $.
In this subsection, we establish a moment inequality in Theorem (ref) for $\sum^{n}_{j=1}(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n})$, which directly yield a LLN and are useful for theoretical analysis in statistics and econometrics. We begin with a crucial lemma.
With Lemma (ref) and the Burkholder\textquoteright s inequality for martingales (Lemma (ref)), we obtain the following moment inequality.
A direct result of Theorem (ref) is a law of large numbers for functionally dependent network random variables\@.
Concentration inequalities, also called exponential inequalities in the literature, play an indispensable role in empirical process theory, as well as in semi-parametric, non-parametric, and high-dimensional statistics (wainwright2019HighDimensional). wu2005nonlinear and wu2016wu establish two concentration inequalities for functionally dependent stationary time series. In the literature, although there are some concentration inequalities for spatial data on irregular lattices (e.g., XU201896, yuan2025Bernsteintype, and wu2023application), the concentration inequalities for network data are rare. Hence, in this subsection, we follow the strategies in wu2016wu to establish a concentration inequality for network data.
We now illustrate when the condition ((ref)) holds. Consider the SAR model from Section (ref). From Proposition (ref), we have $\delta_{p,n}\left(j,i\right)\leq2\left\Vert \epsilon\right\Vert _{L^{p}}S^{+}_{ji,n}$.\footnote{For simplicity, here we assume that $\mathcal{C}_{n}$ is the trivial $\sigma$-field $\{\emptyset,\Omega\}$.} As a result, we obtain
If $\sup_{n\geq1}\frac{1}{n}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}S^{+}_{ji,n}\right]^{2}<C^{2}<\infty$ for some constant $C>0$, then \[ \sup_{p\geq2}\sup_{n\geq1}p^{-\nu}\sqrt{n}\Delta^{1/2}_{p,2}\leq2C\sup_{p\geq2}p^{-\nu}\left\Vert \epsilon\right\Vert _{L^{p}}. \] Thus, the condition ((ref)) holds with $\nu=1$ if $\epsilon_{i,n}$'s are uniformly sub-exponential, with $\nu=\frac{1}{2}$ if $\epsilon_{i,n}$'s are uniformly sub-Gaussian, and with $\nu=0$ if $\epsilon_{i,n}$'s are uniformly bounded. Compared to the Hoeffding's inequality (wainwright2019HighDimensional), we see that when $\nu=0$, the decay rate on the right-hand-side of ((ref)) with respect to $x$ is the same as in the independent case. When $\nu>0$, the decay rate is slower.
In this section, we establish two central limit theorems (CLTs) based on the FDM. Denote $Z_{i,n}\equiv\sum^{n}_{k=1}P_{i}Y_{k,n}$ for $i\in[n]$, where $P_{i}Y_{k,n}=\mathbb{E}_{\mathcal{C}_{n}}(Y_{k,n}|\mathcal{F}_{i,n})-\mathbb{E}_{\mathcal{C}_{n}}(Y_{k,n}|\mathcal{F}_{i-1,n})$. By construction, $\{Z_{i,n},\mathcal{F}_{i,n}\}^{n}_{i=1}$ is an MDA under both $\mathbb{P}_{\mathcal{C}_{n}}$ and $\mathbb{P}$. Therefore, we will prove our CLTs by applying a martingale CLT (Lemma (ref)) to the normalized sum $\sigma^{-1}_{n}\sum^{n}_{i=1}Z_{i,n}$. Before stating the CLTs, we introduce a second-order functional dependence measure, which is tailored to the analysis of the MDA increments $\{Z_{i,n}\}$ and, to the best of our knowledge, is new in the literature.
Fix $i\neq j$. Conditional on $\mathcal{C}_{n}$, let $e^{*}_{i,n}$ and $e^{*}_{j,n}$ be independent copies of $e_{i,n}$ and $e_{j,n}$, respectively, such that $(e_{1,n},\dots,e_{n,n},e^{*}_{i,n},e^{*}_{j,n})$ are mutually independent given $\mathcal{C}_{n}$. Define $Y_{k,n,\{i,j\}}$ as the coupled version of $Y_{k,n}$ obtained by replacing $e_{i,n}$ and $e_{j,n}$ with $e^{*}_{i,n}$ and $e^{*}_{j,n}$: \[ Y_{k,n,\{i,j\}}\equiv F_{k,n}(e_{1,n},\dots,e_{i-1,n},e^{*}_{i,n},e_{i+1,n},\dots,e_{j-1,n},e^{*}_{j,n},e_{j+1,n},\dots,e_{n,n}) \] for $i<j$. When $j<i$, the same definition applies with the roles of $i$ and $j$ swapped. Equivalently, the above display is understood as replacing the $i$th and $j$th coordinates simultaneously.
Why is this second-order FDM needed? To apply the martingale CLT, we need to establish the LLN for the squared increments $\{Z^{2}_{i,n}\}$. A direct way to show this is to invoke Theorem (ref), which requires controlling the (first-order) FDM of $\{Z^{2}_{i,n}\}$ (viewed as functions of the innovations $e_{i,n}$). By applying the difference of squares formula, controlling the FDM of $\{Z^{2}_{i,n}\}$ reduces to controlling the (first-order) FDM of $\{Z_{i,n}\}$. It turns out that this FDM can then be bounded in terms of the second-order FDM of the original field $\{Y_{i,n}\}$, as summarized below.
Lemma (ref) can be viewed as a second-order analogue of Lemma (ref). Indeed, Lemma (ref) implies $\left\Vert Z_{i,n}\right\Vert _{L^{p},\mathcal{C}_{n}}\leq\sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})$; that is, the $L^{p}$ norm of $Z_{i,n}$, which can be interpreted as a “zeroth-order” FDM for $\{Z_{i,n}\}$, is controlled by the (first-order) FDM of $\{Y_{i,n}\}$. In contrast, Lemma (ref) shows that the first-order FDM of $\{Z_{i,n}\}$ is controlled by the second-order FDM of $\{Y_{i,n}\}$. With this lemma in hand, we are ready to state our CLTs.
We next discuss how to verify the high-level conditions in Theorem (ref). First note that conditions ((ref)) and ((ref)) are invariant under relabeling of the nodes $1,\ldots,n$. This is a natural requirement because the validity of the CLT should depend on the dependence structure of the network, not on an arbitrary ordering.
Condition ((ref)) is based on the (first-order) FDM and relatively easy to verify. Condition ((ref)) involves the second\nobreakdash-order FDM and is more demanding. The next two lemmas provide two different ways to control the second-order FDM. Lemma (ref) exploits smoothness of the functions $F_{j,n}$ (recall that $Y_{j,n}=F_{j,n}(e_{n})$), whereas Lemma (ref) does not require smoothness and thus applies more generally.
Lemma (ref) is particularly convenient for smooth models. In Section (ref), we apply it to the SAR model (see Lemma (ref) and Proposition (ref)). The next lemma provides an alternative route that avoids smoothness assumptions.
Condition ((ref)) requires that, for each node $k$, the first-order dependence measures $\delta_{p,n}\left(k,\pi_{k}(r),\mathcal{C}_{n}\right)$ decay at a polynomial rate when ordered from largest to smallest. In other words, although a node may be affected by many other nodes, the strength of these influences must decline sufficiently fast across the ordered neighbors.
We now compare our CLT with existing results in the literature.
(1) The CLTs in wu2005nonlinear,wu2023application,ELMACHKOURI20131,jenish2009central,jenish2012spatial,kojevnikov_limit_2021 all require individuals to be located in some metric space. Our CLT imposes no such metric structure, making it applicable to network data, where a natural embedding may be absent.
(2) The key condition for CLTs in the literature (e.g., jenish2009central,jenish2012spatial,kojevnikov_limit_2021,wu2023application) typically takes the form $\sum^{\infty}_{s=1}\nu_{s}\theta_{s}<\infty$, where $s$ denotes distance, $\nu_{s}$ is the number of individuals at distance approximately equal to $s$ from a given individual, and $\theta_{s}$ is a weak dependence coefficient at distance $s$ (e.g., a strong mixing, NED, or $\psi$-dependence coefficient). For this condition to hold, one typically needs $\theta_{s}$ to decay sufficiently fast as $s$ increases, and the growth of $\nu_{s}$ must also be controlled. As a result, these conditions are most natural in settings where dependence weakens gradually over many distance shells. This requirement can be restrictive for network data, since many relevant networks exhibit small-world behavior and hence have small graph diameter.\footnote{For example, in the Erd\H{o}s-R\'{e}nyi model in Example (ref), if the mean degree satisfies $D_{n}\propto n^{\alpha}$ ($0<\alpha<1)$, then results in klee1981diameters imply that the graph diameter (based on the geodesic distance) is bounded by a constant with probability approaching one. Consequently, the graph radius is also of constant order. In such settings, existing weak dependence concepts based on decay with network distance become less compelling, because only finitely many distance shells are asymptotically relevant.} In contrast, our condition ((ref)) only requires that, for any node $i$, the magnitude of the impact it receives from other nodes (i.e., the FDM $\delta_{p,n}\left(k,\pi_{k}(r),\mathcal{C}_{n}\right)$) decreases as a power function of $r$ when these impacts are ordered from largest to smallest. This allows our CLT to accommodate networks with small or moderate diameters, a common feature of social and financial networks.\footnote{Financial networks, ranging from interbank markets to asset holding networks, typically exhibit small diameters. For example, consider the holding-based network of Chinese mutual funds. Funds $i$ and $j$ are regarded as connected if they allocate at least 5% of their portfolios to the same stocks. Using data from the second half of 2015, jiang2026NAVaR find that this network has a diameter of only 4.}
(3) kuersteiner2019limit develops a CLT for spatial mixingale processes using a model\nobreakdash-dependent random metric, relaxing the assumption of a fixed metric. kuersteiner2019limit assumes the network to be sparse that “rules out a buildup of a mass of nodes with very similar features is captured by a summability condition of the probabilities that two nodes are close in an appropriate sense”. In contrast, our approach does not require any form of underlying metric and allows for networks with small diameters.
(4) Going beyond the conventional notion of weak dependence (e.g., mixing and NED), leung_normal_2019 propose the \textquotedblleft stabilization\textquotedblright conditions that impose weak dependence on node degrees. While their motivation is similar, they focus on strategic network formation, whereas our concept applies to various types of networks.
(5) lee2022qml study a CLT for a linear-quadratic form of independent random variables in the presence of dominant units (also called popular units in their paper). Their CLT is mainly used for linear SAR type models, but ours is applicable to nonlinear models and some nonlinear estimators (e.g., Huber or quantile regression estimators) of linear SAR models.
Using the Cram\'er-Wold device, we can generalize Theorem (ref) to multivariate case. Now, $Y_{j,n}$'s are random vectors taking values in $\mathbb{R}^{p_{Y}}$ ($p_{Y}\geq1$). Denote $\Sigma_{n}\equiv\mathrm{Var}_{\mathcal{C}_{n}}\left(\sum^{n}_{j=1}\text{\ensuremath{Y_{j,n}}}\right)$, and let $I_{p_{Y}}$ be the $p_{Y}\times p_{Y}$ identity matrix.
To illustrate the application of the FDM, in this section, we calculate the FDM for SAR models and establish the consistency of the maximum likelihood estimator for the SAR Tobit model using FDM.
Let $F:\mathbb{R}\to\mathbb{R}$ be a Lipschitz function, i.e., $\left|F\left(x\right)-F\left(y\right)\right|\leq L\left|x-y\right|$ for some constant $L>0$ and all $x,y\in\mathbb{R}$. A (possibly nonlinear) SAR model can be written as
for $j=1,\ldots,n$, where $w_{j\cdot,n}$ is the $j$th row of a non-zero spatial/network weights matrix $W_{n}=\left(w_{ji,n}\right)_{n\times n}$, $Y_{n}=\left(Y_{1,n},Y_{2,n},\cdots,Y_{n,n}\right)'$, $X_{j,n}\in\mathbb{R}^{p}$ is the exogenous regressor, $\epsilon_{j,n}$ is the disturbance term, and $\lambda$ and $\beta$ are model parameters. When $F\left(x\right)=x$, Eq.(ref) is the standard (linear) SAR model; when $F\left(x\right)=\max\left(0,x\right)$, Eq.(ref) becomes a SAR Tobit model. Now, we investigate the FDM of SAR model. The following two assumptions on $F(\cdot)$ and $\{\epsilon_{j,n}\}$ are needed.
Note that Assumption (ref) allows the weights matrix $W_{n}$ to be random. Let $e_{j,n}=X'_{j,n}\beta+\epsilon_{j,n}$ in the SAR model (ref). Then $Y_{i,n}$'s are generated by the underlying random variables $e_{1,n},\ldots,e_{n,n}$. From Assumption (ref), $e_{j,n}$'s are conditionally independent on $\mathcal{C}_{n}$. Denote $\left|W_{n}\right|\equiv(|w_{ij,n}|)_{n\times n}$ and
where $I_{n}$ denotes the $n\times n$ identity matrix.
As a direct consequence of Proposition (ref) and Theorem (ref), we have the following LLN for the dependent variable $Y_{i,n}$ generated by the SAR model ((ref)).
Proposition (ref) directly follows from Proposition (ref) and Theorem (ref), and we omit the proof. Since $\left\Vert S^{+}_{n}\right\Vert _{1}\geq\mu_{\min\{p,2\}}$, a sufficient condition for LLN is $\left\Vert S^{+}_{n}\right\Vert _{1}=o_{\mathbb{P}}(n^{\frac{\min\{p,2\}-1}{\min\{p,2\}}})$.
Next, we bound the second-order FDM for SAR models in order to establish CLT for the dependent variable $Y_{i,n}$ generated by the SAR model ((ref)). To do so, we make the following assumption on the smoothness of the function $F$.
Relying on Lemma (ref), we have the following lemma on bounding the second-order FDM for SAR models.
Based on Lemma (ref), we impose the following assumption to verify condition ((ref)).
Notice that $\mu_{q}$ increases to $\mu_{\infty}$ as $q\to\infty$. We provide a sufficient condition for Assumption (ref). If $\mu_{\infty}=o_{\mathbb{P}}(n^{(\tilde{p}-1)/(3\tilde{p}-1)})$, we have $\mu^{2\tilde{p}}_{2\tilde{p}}\leq\mu^{2\tilde{p}}_{\infty}=o(n^{2\tilde{p}(\tilde{p}-1)/(3\tilde{p}-1)})=o_{\mathbb{P}}(n^{\tilde{p}-1}).$ Besides, by letting $u=\infty$ and $v=1$, we have $\mu^{(2\tilde{p}-1)+1/u}_{u(2\tilde{p}-1)+1}\cdot\mu^{\tilde{p}}_{v\tilde{p}}\leq\mu^{2\tilde{p}-1}_{\infty}\mu^{\tilde{p}}_{1}\leq\mu^{2\tilde{p}-1}_{\infty}\mu^{\tilde{p}}_{\infty}=o_{\mathbb{P}}(n^{\tilde{p}-1}).$ Thus, a sufficient condition for Assumption (ref) is $\|S^{+}_{n}\|_{1}=o_{\mathbb{P}}\bigl(n^{(\tilde{p}-1)/(3\tilde{p}-1)}\bigr)$. An important special case is $\sup_{n}\|S^{+}_{n}\|_{1}=O_{\mathbb{P}}(1)$, which is widely imposed in the literature; see, for example, lee2004asymptotic,lee_gmm_2007,yu2008quasi,xu2015maximum. Moreover, since $\frac{\tilde{p}-1}{3\tilde{p}-1}\le\frac{\min\{p,2\}-1}{\min\{p,2\}},$ the condition $\|S^{+}_{n}\|_{1}=o_{\mathbb{P}}(n^{(\tilde{p}-1)/(3\tilde{p}-1)})$ is stronger than the corresponding condition used for the LLN. Hence, by Proposition (ref), it also implies the LLN.
Next, we state the CLT result for the SAR model.
We next verify Assumption (ref) in several commonly used network settings. We start with a deterministic design that allows for a small number of highly influential units. We then consider three random graph models: the Erd\H{o}s-R\'{e}nyi model, the triangle model, and the stochastic block model. Together, these examples show that Assumption (ref) holds under a range of network structures of practical interest.
Within this framework, it is useful to distinguish two topological scenarios according to whether non-dominant units can feed back into the dominant ones.
The next proposition makes this comparison precise by characterizing when Assumption (ref) holds under each of the two dominant-unit configurations.
Proposition (ref) shows that the strict dominant structure permits a wider range of $(\eta_{m},\delta)$ than the general case. This reflects the fact that ruling out feedback from non-dominant units to dominant units makes the concentration of column influence easier to control. In particular, under the strict dominant structure, when $\eta_{m}=0$ we require $\delta<(3-\sqrt{5})/2\approx0.382$, whereas when $\eta_{m}=1/5$ we require $\delta<(5-3\sqrt{2})/5\approx0.152$.
We next consider network weights matrices generated by random graph models.
The ER model provides a convenient baseline for the random-graph setting. The next proposition shows that Assumption (ref) generally holds for the ER model.
Proposition (ref) shows that Assumption (ref) holds for the Erd\H{o}s-R\'{e}nyi model without any further restriction on the growth rate of $D_{n}$. This highlights that the proposed dependence concept can accommodate highly connected networks with small or moderate diameter.
This result shows that Assumption (ref) continues to hold when the network exhibits local clustering, provided that the overall degree level remains sufficiently controlled. We next turn to a different source of network heterogeneity, namely latent community structure.
This result shows that Assumption (ref) continues to hold in the presence of community structure. It is also worth noting that, beyond the requirement that the edge probabilities lie in $[0,1]$, the proposition does not impose a separate growth condition on $D_{\mathrm{wb},n}$ or $D_{\mathrm{bb},n}$.
Taken together, the above examples show that Assumption (ref) holds in a broad class of network settings, including deterministic designs with influential units as well as standard random graph models. As a result, the CLT developed in Proposition (ref) applies under a range of dependence patterns commonly used in empirical and theoretical work on networks.
In this section, we investigate the consistency of the maximum likelihood estimator (MLE) for the SAR Tobit model in xu2015maximum. Specifically, for each node $i\in[n]$, the model is defined as: \[ Y_{i,n}=\max\{0,Y^{*}_{i,n}\},\qquad Y^{*}_{i,n}=\lambda_{0}w_{i\cdot,n}Y_{n}+X^{\prime}_{i,n}\beta_{0}+\epsilon_{i,n}, \] where the exogenous regressors $X_{i,n}$ and the network weights matrix $W_{n}$ generate the sub-$\sigma$-field $\mathcal{C}_{n}\equiv\bigvee^{n}_{j=1}\sigma\left(X_{j,n}\right)\bigvee\sigma(W_{n})$. Conditional on $\mathcal{C}_{n}$, the innovations $\epsilon_{1,n},\dots,\epsilon_{n,n}$ are assumed to be i.i.d. following a normal distribution $N(0,\sigma^{2}_{0})$. Fix a parameter value $\theta=(\lambda,\beta,\sigma)$. Let $z_{i,n}(\theta)=(Y_{i,n}-\lambda w_{i\cdot,n}Y_{n}-X^{\prime}_{i,n}\beta)/\sigma$ and $G_{n}(Y_{n})=\text{diag}(1(Y_{1,n}>0),\dots,1(Y_{n,n}>0))$. The corresponding log-likelihood function can be written as
where $\Phi(\cdot)$ is the cumulative distribution function of the standard normal distribution. Our goal here is not to re-establish the full consistency argument, but rather to show that the key pointwise convergence step in that argument can be obtained without relying on an underlying metric space. In the framework of xu2015maximum, the metric-space assumption is used to establish the pointwise convergence of $\frac{1}{n}\ell_{n}(\theta)$. Accordingly, we focus on proving that $\frac{1}{n}\ell_{n}(\theta)=\frac{1}{n}\mathbb{E}_{\mathcal{C}_{n}}\ell_{n}(\theta)+o_{\mathbb{P}}(1)$, which is one ingredient in the consistency proof. The following assumptions are required for the subsequent analysis.
See xu2015maximum for primitive conditions for Assumption (ref)(v). We then obtain the following result.
Proposition (ref) shows that this pointwise convergence of the log-likelihood can be established within the FDM framework, without imposing any underlying metric-space structure. This provides one of the probabilistic ingredients used in the consistency analysis of the SAR Tobit MLE.
To study the asymptotic properties of an estimator, we usually need to deal with various functions of random variables, e.g., $Y^{2}_{i,n}$ or $\Phi(Y_{i,n})$, where $\Phi(\cdot)$ is a distribution function of some random variable. Therefore, to broaden the applicability of our theory, we study the FDM under different transformations, and investigate whether the important inequalities and the CLTs remain valid. As in Section (ref), denote $Y_{j,n}=F_{j,n}(e_{n})\in\mathbb{R}$, $Y_{j,n,i}\equiv F_{j,n}\left(e_{1,n},\cdots,e_{i-1,n},e^{*}_{i,n},e_{i+1,n},\cdots,e_{n,n}\right)$, and $\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\equiv\left\Vert Y_{j,n}-Y_{j,n,i}\right\Vert _{L^{p},\mathcal{C}_{n}}$ as the FDM of $\left\{ Y_{j,n}\right\} $ conditional on $\mathcal{C}_{n}$. We add the superscript $``(Y)"$ in $\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ to differentiate it from the FDM of other random variables.
We first consider Lipschitz-type functions $H_{j,n}:\mathbb{R}\rightarrow\mathbb{R}$. Suppose for all $\left(y,y^{\bullet}\right)\in\mathbb{R}\times\mathbb{R}$ and all $j$ and $n\geq1$:
In Propositions (ref)-(ref), we denote $Z_{j,n}\equiv H_{j,n}(Y_{j,n})$.
Obviously, if the FDM of $\{Y_{j,n}\}$ satisfies the conditions in Theorems (ref), (ref), or (ref), then the same holds for the FDM of $\{Z_{j,n}\}$.
Next, we consider unbounded $B_{j,n}\left(y,y^{\bullet}\right)$, e.g., $H_{j,n}\left(y\right)=y^{2}$.
Hence, if $\frac{1}{n^{\widetilde{q}}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta^{(Y)}_{q,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\widetilde{q}}=o_{\mathbb{P}}(1)$ and $\left\Vert Y\right\Vert _{L^{ar},\mathcal{C}_{n}}<\infty$ a.s., then \[ \frac{1}{n^{\widetilde{q}}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\widetilde{q}}=o_{\mathbb{P}}(1). \] So Theorem (ref) is applicable. Similarly, condition ((ref)) in Theorem (ref) can be verified for $\{Z_{j,n}\}$. In Proposition (ref), there is a trade-off between $p$, $q$ and $r$. If a larger $p$ is desired, then larger $q$ or $r$ is required. By imposing further conditions on $\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$, we can avoid the trade-off between $p$, $q$ and $r$, which is presented in the next proposition.
If $\frac{1}{n^{q}}\sum^{n}_{i=1}\left\{ \sum^{n}_{j=1}\left[\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\left(q-ap-p\right)/\left(pq-ap-p\right)}\right\} ^{q}=o_{\mathbb{P}}(1)$ , then, by Proposition (ref), \[ \frac{1}{n^{q}}\sum^{n}_{i=1}\left\{ \sum^{n}_{j=1}\delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right\} ^{q}=o_{\mathbb{P}}(1) \] So Theorem (ref) is applicable. Moreover, condition ((ref)) in Theorem (ref) can be verified using similar arguments.
In addition to Lipschitz transformation, another important nonlinear transformation is $1\left(y>0\right)$, which is useful in discrete choice and censor data model. See, e.g., xu2015maximum,XU201896.
Moreover, we are interested in the sum or the product of two random variables. Let $\left\{ Y_{j,n}\right\} $ and $\left\{ Z_{j,n}\right\} $ be two sets of random variables, and they are both functions of some conditionally independent random vectors $e_{i,n}$ given $\mathcal{C}_{n}$, where $i\in[n]$. Denote the FDMs of $\left\{ Y_{j,n}+Z_{j,n}\right\} $ and $\left\{ Y_{j,n}Z_{j,n}\right\} $ by $\delta^{(Y+Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ and $\delta^{(YZ)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$, respectively. For summation, the conclusion is a direct result of Minkowski's inequality, and we summarize it below.
The FDM of product is more complicated. For simplicity, we assume the original random fields are real-valued, since if $Y$ is functionally dependent, then all the elements of $Y$ are still functionally dependent and vice versa. Analogous to Propositions (ref) and (ref), we have two versions of conclusions, presented below.
Let us examine whether $\{Y_{j,n}Z_{j,n}\}$ satisfies the conditions in Theorem (ref). Denote $\Delta^{(YZ)}_{p,2}\left(\mathcal{C}_{n}\right)\equiv\frac{1}{n^{2}}\sum^{n}_{i=1}[\sum^{n}_{j=1}\delta^{(YZ)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)]^{2}$, and $\Delta^{(Y)}_{q_{1},2}\left(\mathcal{C}_{n}\right)$ and $\Delta^{(Z)}_{q_{2},2}\left(\mathcal{C}_{n}\right)$ are defined similarly. Then by Proposition (ref),
where the second inequality follows from the fact that $\left(a+b\right)^{2}\leq2a^{2}+2b^{2}$ for arbitrary $a$, $b\in\mathbb{R}$. Consequently, when $\max\left\{ \Delta^{(Y)}_{q_{1},2}\left(\mathcal{C}_{n}\right),\Delta^{(Z)}_{q_{2},2}\left(\mathcal{C}_{n}\right)\right\} =o_{\mathbb{P}}(1)$, $\{Y_{j,n}Z_{j,n}\}$ also satisfies the conditions in Theorem (ref).
Similar to Proposition (ref), there is a trade-off between $p$, $q_{1},r_{1}$and $p,q_{2},r_{2}$. Since $q_{1}>p$ and $q_{2}>p$, when three or more random fields are multiplied, larger values for $q_{1}$ and $q_{2}$ are required, which may not always be feasible. However, by introducing additional conditions, it may be possible to avoid such trade-offs, as demonstrated below.
Using Proposition (ref) together with the inequality $(a+b)^{r}\leq\max\{1,2^{r-1}\}(a^{r}+b^{r})$ for any $a,b,r\geq0$, we can show that if \[ \frac{1}{n^{2}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\left(\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right)^{\left(q-2p\right)/\left(pq-2p\right)}\right]^{2}=o_{\mathbb{P}}(1) \] and \[ \frac{1}{n^{2}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\left(\delta^{(Z)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right)^{\left(q-2p\right)/\left(pq-2p\right)}\right]^{2}=o_{\mathbb{P}}(1), \] then \[ \Delta^{(YZ)}_{p,2}\left(\mathcal{C}_{n}\right)=\frac{1}{n^{2}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta^{(YZ)}_{p,n}\left(i,j\right)\right]^{2}=o_{\mathbb{P}}(1). \] Then, Theorem (ref) is applicable. We can verify condition ((ref)) in Theorem (ref) under similar conditions.
In practice, one may use either Proposition (ref) or (ref) to study the FDM of $\{Y_{j,n}Z_{j,n}\}$. Proposition (ref) requires a stronger moment condition, while Proposition (ref) requires that most of $\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$ are smaller, as $\frac{\left(q-2p\right)}{pq-2p}<1$ implies that $\left[\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right]^{\left(q-2p\right)/\left(pq-2p\right)}>\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$.\footnote{Note that most of $\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$ are smaller than 1 asymptotically. }
This paper develops limit theorems for random variables with network structure, without requiring individuals to be located in a Euclidean or other metric space. This sets our approach apart from most existing limit theories in network statistics and econometrics, which typically rely on weak dependence concepts such as strong mixing, near-epoch dependence, or $\psi$-dependence. To establish these results, we generalize the functional dependence measure proposed by wu2005nonlinear. Using this framework, we derive several inequalities, a law of large numbers, and central limit theorems. We also demonstrate the applicability of these results by verifying their conditions for spatial autoregressive models, which are widely used in network data analysis. Finally, we study the behavior of the functional dependence measure under various transformations commonly encountered in applications. Extending this framework to panel data with network dependence is a natural direction for future research.