EconBase
← Back to paper

Limit Theorems for Network Data without Metric Structure

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

79,680 characters · 14 sections · 52 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Limit Theorems for Network Data without Metric Structure

\global\long\global\long\global\long\global\long\global\long\global\long\global\long\foreignlanguage{english}{ \global\long}

abstractThis paper develops limit theorems for random variables with network dependence, without requiring the individuals in the network to be located in a Euclidean or metric space. This distinguishes our approach from most existing limit theorems in network statistics and econometrics, which are based on weak dependence concepts such as strong mixing, near-epoch dependence, or $\psi$-dependence. All these weak dependence concepts presuppose an underlying metric. By relaxing the assumption of an underlying metric space, our theorems can be applied to a broader range of network data, including financial and social networks. To derive the limit theorems, we generalize the concept of functional dependence (also known as physical dependence) from time series to random variables with network dependence. Using this framework, we establish several inequalities, a law of large numbers, and central limit theorems. Furthermore, we demonstrate the verifiability of our high-level conditions by deriving primitive sufficient conditions for spatial autoregressive models, which are widely used in network data analysis.

Keywords: functional dependence, network data, law of large numbers, central limit theorem, concentration inequality, spatial autoregressive model

Introduction

Network-indexed dependence arises in a broad class of statistical and econometric models. To study estimators and test statistics in such settings, one needs probabilistic tools, such as moment inequalities, concentration inequalities, laws of large numbers, and central limit theorems, to accommodate dependence structures that are neither sequentially ordered nor naturally organized by a metric. This challenge is especially acute in nonlinear network models, where martingale methods available for linear specifications are typically not applicable.

A substantial body of literature has studied weak dependence concepts for spatial and network-indexed data. For linear spatial autoregressive (SAR) models, martingale difference array arguments play a central role h._kelejian_asymptotic_2001,lee2004asymptotic,lee_gmm_2007,lee2022qml. For more general dependent arrays, the literature develops strong-mixing and $\phi$-mixing spatial random fields jenish2009central, spatial near-epoch dependence (NED) jenish2012spatial, spatial functional dependence wu2023application, spatial mixingales based on model-dependent random metrics kuersteiner2019limit, and network adaptations of $\psi$-dependence kojevnikov_limit_2021; see also kuersteiner_limit_2013,KuersteinerPrucha2020,leung_weak_2019,leung_normal_2019. These contributions have substantially advanced asymptotic theory for dependent data and have also been applied in a variety of settings; see, for example, xu2015maximum,gao2025Causal,kojevnikov2021bootstrap,leung2022Causal,XU201896,qu_estimating_2015,liu2022robust. However, most of them rely on a metric, a geodesic distance, or a model-dependent random metric. Accordingly, dependence is characterized through distance-based shells and decay with distance.

This reliance on metric structure constitutes a major bottleneck for network data. Many network models employed in statistics and econometrics do not possess a natural embedding in Euclidean space. Even when a metric can be imposed, it may not accurately reflect the actual propagation of dependence. The problem is especially acute in networks with short graph distances, high connectivity, or moderate diameter. For example, consider an Erd\H{o}s-R\'{e}nyi model with mean degree $D_{n}\propto n^{\alpha}$ ($0<\alpha<1)$. Classical results on the diameter of Erd\H{o}s-R\'{e}nyi graphs imply that graph diameter is of constant order with probability approaching one in this setting klee1981diameters. In such settings, there are only a few relevant distance shells, so weak dependence cannot be naturally organized via gradual decay over expanding metric neighborhoods. What is needed, therefore, is a notion of weak dependence that does not require any metric structure, and can be directly verified from primitive conditions.

This restriction is substantive, not merely technical. In many network applications, it is unclear whether the nodes admit any meaningful Euclidean embedding. For instance, xu2022dynamic applied spatial NED theory to quantile regression in a dynamic network model for stocks traded on the NYSE and NASDAQ, even though it is not obvious that such assets should be viewed as points in a Euclidean space. Likewise, latent-space representations for social networks hoff2002Latent,athreya2018Statistical,smith2019Geometry provide useful modeling devices, but the underlying metric is typically imposed for tractability rather than being dictated by the dependence mechanism itself. More generally, in online and social networks, nodes that are far apart in any Euclidean sense may still interact strongly. These examples indicate that metric-based weak dependence conditions can impose a substantive modeling restriction in applications where no natural metric structure is available.

There is also a methodological mismatch between estimation and asymptotic theory. Many standard procedures, including quasi-maximum likelihood estimation and the generalized method of moments, can be formulated for network data without any metric structure. However, available asymptotic justifications often rely on metric-based weak dependence conditions. For instance, the robust estimator for SAR models proposed by liu2022robust ultimately invokes the spatial NED framework of jenish2012spatial. A partial step in this direction is kuersteiner2019limit, who replaces a fixed metric with a model-dependent random metric. This underscores the need for a notion of weak dependence that is both metric-free and directly verifiable from the model.

We address this problem by extending the concept of functional dependence measure (FDM) to network data. This measure is defined by perturbing the underlying innovations and quantifying the resulting effect on each component of the system. The construction is entirely node-based and does not rely on Euclidean distance, geodesic distance, or any other metric. It extends the functional dependence concept of wu2005nonlinear, ELMACHKOURI20131, and wu2023application from time series and spatial processes to general network dependence. Unlike strong mixing and related notions, the proposed FDM avoids suprema over sub-$\sigma$-fields and is therefore substantially easier to verify in specific network models.

The contributions of this paper are fourfold. First, we introduce a metric-free functional dependence framework for network data. This extends the functional dependence paradigm to dependence structures that are intrinsic to networks and not organized by any metric. Second, under this framework we develop a general probabilistic toolkit for network data. In particular, we establish a moment inequality, a concentration inequality, a weak law of large numbers, and central limit theorems. Third, for the CLT theory, we introduce a second-order functional dependence measure and derive sufficient conditions directly in terms of first- and second-order dependence measures. These conditions are invariant under relabeling of the nodes and are designed for nonlinear systems, thereby going beyond martingale-based arguments available for linear models. Fourth, we show that the proposed dependence concept is verifiable and useful in concrete models. We compute the dependence measures for nonlinear SAR models, derive primitive sufficient conditions for LLNs and CLTs, and verify these conditions under several network designs, including dominant-unit structures and random graph models. In addition, we establish the consistency of the maximum likelihood estimator for the SAR Tobit model using the tools developed in this paper.

The remainder of the paper is organized as follows. Section (ref) introduces the functional dependence measure and the associated notion of $(L^{p},q)$-functional dependence. Section (ref) develops the general theoretical tools, including a moment inequality, a concentration inequality, a weak law of large numbers, and central limit theorems; it also introduces a second-order functional dependence measure for the CLT analysis. Section (ref) applies the framework to nonlinear SAR models, derives primitive sufficient conditions for the LLN and CLTs, and verifies these conditions under several network designs; it also studies the SAR Tobit model. Section (ref) investigates how the functional dependence measure behaves under common transformations. Section (ref) concludes. Proofs of the main results are collected in Appendix (ref), and additional proofs are deferred to the supplementary appendix.

Notation. The set of positive integers is $\mathbb{N}\equiv\{1,2,\cdots\}$, $\mathbb{Z}$ denotes the set of integers, and $\mathbb{Z}^{d}$ is the $d$-dimensional integer lattice. For any column vector $x=\left(x_{1},x_{2},\ldots,x_{d}\right)'\in\mathbb{R}^{d}$, where $\mathbb{R}^{d}$ is the $d$-dimensional Euclidean space, $\left\Vert x\right\Vert =(x'x)^{1/2}$ denotes its Euclidean norm, $\left\Vert x\right\Vert _{\infty}=\max_{1\leq k\leq d}\left|x_{k}\right|$ represents its infinity norm, and $\left\Vert x\right\Vert _{1}=\sum^{d}_{k=1}\left|x_{k}\right|$ is its 1-norm. For any matrix $A=\left(a_{ij}\right)_{n\times m}$, its Frobenius norm is $\left\Vert A\right\Vert \equiv\sqrt{\sum_{ij}a^{2}_{ij}}$, its 1-norm (also called column sum norm) is defined as $\left\Vert A\right\Vert _{1}=\max_{1\leq j\leq m}\sum^{n}_{i=1}\left|a_{ij}\right|$, and its infinity norm (also called row sum norm) is defined as $\left\Vert A\right\Vert _{\infty}=\max_{1\leq i\leq n}\sum^{m}_{j=1}\left|a_{ij}\right|$. For any symmetric matrix $A$, $\min\mathrm{eig}(A)$ denotes its minimum eigenvalue. For any two sequences $a_{n}$ and $b_{n}$, $a_{n}\sim b_{n}$ if and only if $a_{n}/b_{n}\to1$ as $n\to\infty$. Let $\left(\Omega,\mathcal{F},\mathbb{P}\right)$ be a probability space. Denote probability and expectation by $\mathbb{P}$ and $\mathbb{E}$, respectively. For any random vector (matrix) $X$, its $L^{p}$ norm is defined as $\left\Vert X\right\Vert _{L^{p}}\equiv\left[\mathbb{E}\left(\left\Vert X\right\Vert ^{p}\right)\right]^{1/p}$ for any constant $p\geq1$. The symbols “$\xrightarrow{\mathbb{P}}$” and “$\ensuremath{\xrightarrow{d}}$” denote convergence in probability and convergence in distribution, respectively. For two positive non-random sequences $a_{n},b_{n}$ and random vector sequence $X_{n}$, $X_{n}=o_{\mathbb{P}}(a_{n})$ means $\mathbb{P}\left(\left\Vert X_{n}\right\Vert >\epsilon a_{n}\right)\to0$ as $n\to\infty$ for any $\epsilon>0$ and $X_{n}=O_{\mathbb{P}}(a_{n})$ means for any $\epsilon>0$, there exists a constant $M>0$ such that $\limsup_{n\to\infty}\mathbb{P}\left(\left\Vert X_{n}\right\Vert \geq Ma_{n}\right)<\epsilon$. For any sub-$\sigma$-field $\mathcal{C}$ of $\mathcal{F}$, denote the conditional probability, conditional expectation, and conditional variance by $\mathbb{P}_{\mathcal{C}}\left(\cdot\right)\equiv\mathbb{P}\left(\cdot\mid\mathcal{C}\right)$, $\mathbb{E}_{\mathcal{C}}\left(\cdot\right)\equiv\mathbb{E}\left(\cdot\mid\mathcal{C}\right)$, and $\mathrm{Var}_{\mathcal{C}}\left(\cdot\right)\equiv\mathrm{Var}\left(\cdot|\mathcal{C}\right)$, respectively. Besides, for a sub-$\sigma$-field $\mathcal{G}$ of $\mathcal{F}$, we denote $\mathbb{E}_{\mathcal{C}}\left(\cdot|\mathcal{G}\right)\equiv\mathbb{E}\left(\cdot\mid\mathcal{G}\bigvee\mathcal{C}\right)$, where $\mathcal{G}\bigvee\mathcal{C}$ denotes the $\sigma$-field generated by $\mathcal{G}$ and $\mathcal{C}$. For a random vector (matrix) $X$, let $\left\Vert X\right\Vert _{L^{p},\mathcal{C}}\equiv\left[\mathbb{E}_{\mathcal{C}}\left(\left\Vert X\right\Vert ^{p}\right)\right]^{1/p}$.

Definition of Functional Dependence Measure

Consider a network with $n$ nodes (also referred to as individuals or units). For simplicity, name these $n$ nodes as $1,2,\cdots,n$, and denote $[n]=\{1,2,\cdots,n\}$. Although we use numeric labels, the order is arbitrary: random variables associated with individuals 1 and 2 may be independent, while those associated with individuals 1 and $n$ could be strongly correlated.

Let $\left(\Omega,\mathcal{F},\mathbb{P}\right)$ be the underlying probability space, and $\mathcal{C}_{n}$ be a sub-$\sigma$-field of $\mathcal{F}$. Associated with each $i\in[n]$ is a random vector $e_{i,n}\in\mathbb{R}^{p_{e}}$, where $p_{e}\in\mathbb{N}$ is a constant. Suppose that $e_{i,n}$'s are conditionally independent given $\mathcal{C}_{n}$, but they need not be identically distributed. Denote $e_{n}=(e_{1,n},e_{2,n},\cdots,e_{n,n})$. Suppose that $Y_{1,n},\cdots,Y_{n,n}$ are $n$ random vectors generated by $e_{n}$:

equation[equation omitted — 54 chars of source]

where the functions $F_{1,n},\ldots,F_{n,n}$ can be random functions (e.g. associated with a random network weights matrix, see Section (ref)) and we assume these functions are measurable with respect to $\mathcal{C}_{n}$. The triangular array $\{Y_{j,n}:1\leq j\leq n,\ n\geq1\}$ exhibits network dependence. Although $Y_{j,n}$ can be vector-valued, we restrict attention to the real-valued case for simplicity and without loss of generality. See Remark (ref) for details.

Suppose that conditional on $\mathcal{C}_{n}$, $e^{*}_{i,n}$ is an independent and identically distributed (i.i.d.) copy of $e_{i,n}$, and $e^{*}_{i,n}$ is independent of $e_{j,n}$ for all $j\neq i$. Denote $Y_{j,n,i}$ as the coupled version of $Y_{j,n}$ with $e_{i,n}$ replaced by $e^{*}_{i,n}$, i.e., $Y_{j,n,i}\equiv F_{j,n}\left(e_{1,n},\cdots,e_{i-1,n},e^{*}_{i,n},e_{i+1,n},\cdots,e_{n,n}\right)$. Now, we are ready to introduce the definition of functional dependence measure.

defn[Functional dependence measure] For $p\geq1$, define the (first-order) functional dependence measure (FDM), also called the physical dependence measure, as \begin{equation} \delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\equiv\left\Vert Y_{j,n}-Y_{j,n,i}\right\Vert _{L^{p},\mathcal{C}_{n}}. \end{equation} When $\mathcal{C}_{n}$ is the trivial $\sigma$-field $\left\{ \Omega,\emptyset\right\} $, we simplify the notation to $\delta_{p,n}\left(j,i\right)\equiv\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$.
remThe quantity $\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ in Eq.((ref)) measures the influence of $e_{i,n}$ on $Y_{j,n}$: conditional on $\mathcal{C}_{n}$, if $e_{i,n}$ is replaced by its i.i.d. version $e^{*}_{i,n}$, the magnitude of the resulting change of $Y_{j,n}$ is $\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ under the norm $\left\Vert \cdot\right\Vert _{L^{p},\mathcal{C}_{n}}$. Although this definition might superficially resemble a metric, it is not one because it might not satisfy the triangle inequality. To illustrate, consider two nodes $i$ and $j$ with weak dependence. Under a metric-based concept of dependence, no node $k$ could be strongly correlated with both. However, this configuration is entirely possible under the FDM framework.
remWhen $Y_{j,n}=\left(Y^{(1)}_{j,n},\cdots,Y^{(p_{Y})}_{j,n}\right)^{\prime}$ is a $p_{Y}$-dimensional random vector, since $\|y\|\leq\|y\|_{1}$, for all $i=1,\cdots,n$, \begin{align*} & \left\Vert Y_{j,n}-Y_{j,n,i}\right\Vert _{L^{p},\mathcal{C}_{n}}\leq\left\Vert \sum^{p_{Y}}_{k=1}\left|Y^{(k)}_{j,n}-Y^{(k)}_{j,n,i}\right|\right\Vert _{L^{p},\mathcal{C}_{n}}\leq\sum^{p_{Y}}_{k=1}\left\Vert Y^{(k)}_{j,n}-Y^{(k)}_{j,n,i}\right\Vert _{L^{p},\mathcal{C}_{n}}, \end{align*} where $Y_{j,n,i}=\left(Y^{(1)}_{j,n,i},\cdots,Y^{(p_{Y})}_{j,n,i}\right)^{\prime}$ is the coupled version of $Y_{j,n}$ with $e_{i,n}$ replaced by its i.i.d. copy $e^{*}_{i,n}$. Thus, without loss of generality, it suffices to study the FDM for real-valued random variables.

Comparison with existing weak dependence concepts. We now compare our FDM with several well-known notions of weak dependence in the literature.

(1) The index sets considered in wu2005nonlinear,wu2023application,ELMACHKOURI20131,jenish2009central,jenish2012spatial,kojevnikov_limit_2021 are subsets of $\mathbb{Z}^{d}$, $\mathbb{R}^{d}$, or a general metric space. In contrast, our definition does not require the index $i$ to belong to any metric space. This makes our approach applicable to network data, where individuals are not necessarily embedded in any metric space.

(2) Compared with strong mixing, the FDM is easier to compute because the coupled version $Y_{j,n,i}$ is constructed explicitly, and the relevant $L^{p}$-norm is easier to evaluate. In contrast, the strong mixing coefficient requires the calculation of a supremum over two $\sigma$-fields, and thus quite challenging (wu2005nonlinear and xu2021Conditions).

(3) Compared to the spatial NED, FDM is more conveniently to calculate under any $L^{p}$-norm, while the $L^{p}$-NED property is often tractable only for $p=2$, because it typically relies on the fact that the conditional expectation is the best predictor under $L^{2}$-distance.

defnFor any constants $p\geq1$ and $q\geq1$, $\left\{ Y_{j,n}\right\} $ is said to be $\left(L^{p},q\right)$-functionally dependent on $\left\{ e_{i,n}\right\} $ given $\mathcal{C}_{n}$ if \[ \Delta_{p,q}\left(\mathcal{C}_{n}\right)\equiv\frac{1}{n^{q}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{q}=o_{\mathbb{P}}(1) \] as $n\to\infty$. When $\mathcal{C}_{n}$ is the trivial $\sigma$-field $\left\{ \Omega,\emptyset\right\} $, we simplify the notation as $\Delta_{p,q}\equiv\Delta_{p,q}\left(\mathcal{C}_{n}\right)$. When $\Delta_{p,q}=o(1)$, we simply say that $\left\{ Y_{j,n}\right\} $ is $\left(L^{p},q\right)$-functionally dependent on $\left\{ e_{i,n}\right\} $.
remThe term $\Delta_{p,q}\left(\mathcal{C}_{n}\right)$ plays a pivotal role throughout this paper. Since $\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ describes the impact of $e_{i,n}$ on $Y_{j,n}$, $\sum^{n}_{j=1}\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ is the total impact of $e_{i,n}$ on all $Y_{j,n}$'s, which can be regarded as the “influence power” of $e_{i,n}$ on $Y_{j,n}$'s. And $\frac{1}{n}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{q}$ can be interpreted as the “average” $q$th power of the influence powers of all $e_{i,n}$'s. Thus, if the average $q$th power of the influence powers of all $e_{i,n}$'s on $Y_{j,n}$'s is almost surely bounded given $\mathcal{C}_{n}$ for some $q>1$, then $\Delta_{p,q}\left(\mathcal{C}_{n}\right)=o_{\mathbb{P}}(1)$. Critically, we allow some (but not all) individuals to have large influence powers, i.e., we allow $\sup_{i\in[n]}\sum^{n}_{j=1}\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ to increase to $\infty$ as $n\to\infty$. As illustrated in Section (ref), for SAR models, $\sum^{n}_{j=1}\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ is proportional to the summation of the absolute values of the elements in the $i$th column of the matrix $L(I_{n}-L\left|\lambda W_{n}\right|)^{-1}$, where $W_{n}$ is the network weights matrix, i.e., $\sum^{n}_{j=1}\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\propto\sum^{n}_{j=1}\left|L\left[(I_{n}-L\left|\lambda W_{n}\right|)^{-1}\right]_{ji}\right|$. Therefore, $\Delta_{p,q}\left(\mathcal{C}_{n}\right)=o_{\mathbb{P}}(1)$ generalizes a standard assumption in many spatial econometric papers (h._kelejian_asymptotic_2001,lee2004asymptotic,lee_gmm_2007,yu2008quasi, among others): $\sup_{n}\left\Vert L(I_{n}-L\left|\lambda W_{n}\right|)^{-1}\right\Vert _{1}<\infty$, i.e., the column sum norm of $L(I_{n}-L\left|\lambda W_{n}\right|)^{-1}$ is uniformly bounded in $n$. The condition $\Delta_{p,q}\left(\mathcal{C}_{n}\right)=o_{\mathbb{P}}(1)$ excludes the case that all $Y_{j,n}$'s are mainly affected by the same very few $e_{i,n}$'s. Consider an extreme case that $Y_{j,n}=e_{1,n}$ for all $j=1,\dots,n$. Then $\sum^{n}_{j=1}\delta_{p,n}\left(j,1\right)\propto n$, and thus $\Delta_{p,q}=\frac{1}{n^{q}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta_{p,n}\left(j,i\right)\right]^{q}\propto1$. So $\left\{ Y_{j,n}\right\} $ is not $\left(L^{p},q\right)$-functionally dependent on $\left\{ e_{i,n}\right\} $ for any $p\geq1$ and $q\geq1$.

Properties of Functional Dependence

In this section, we establish a moment inequality, a concentration inequality, laws of large numbers (LLN), and central limit theorems (CLTs) using the FDM. These tools are essential for deriving asymptotic theory in statistical and econometric models.

The basic idea of most proofs is to express $\sum^{n}_{j=1}\left(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n}\right)$ as the summation of a martingale difference array (MDA) and then apply the theories of MDA. Thus, we first define an increasing sequence of $\sigma$-fields. For $i\in[n]$, let $\mathcal{F}_{i,n}\equiv\sigma\left(e_{j,n}:j\leq i,j\in[n]\right)$ be the $\sigma$-field generated by $e_{1,n},\cdots,e_{i-1,n},e_{i,n}$, and let $\mathcal{F}_{0,n}\equiv\left\{ \Omega,\emptyset\right\} $.

Moment inequality and weak law of large numbers

In this subsection, we establish a moment inequality in Theorem (ref) for $\sum^{n}_{j=1}(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n})$, which directly yield a LLN and are useful for theoretical analysis in statistics and econometrics. We begin with a crucial lemma.

lemLet $P_{i}Y_{j,n}\equiv\mathbb{E}_{\mathcal{C}_{n}}(Y_{j,n}|\mathcal{F}_{i,n})-\mathbb{E}_{\mathcal{C}_{n}}(Y_{j,n}|\mathcal{F}_{i-1,n})$. Then $\left\Vert P_{i}Y_{j,n}\right\Vert _{L^{p},\mathcal{C}_{n}}\leq\delta_{p,n}(j,i,\mathcal{C}_{n})$ almost surely (a.s.).
remThe quantity $P_{i}Y_{j,n}$ can be regarded as the change in the prediction of $Y_{j,n}$ when we have the new information $e_{i,n}$, conditional on $\mathcal{F}_{i-1,n}$ and $\mathcal{C}_{n}$. Notice that $\left\{ P_{i}Y_{j,n},\mathcal{F}_{i,n}\right\} $ is a MDA under both the conditional expectation $\mathbb{E}_{\mathcal{C}_{n}}$ and the unconditional expectation $\mathbb{E}$,\footnote{$\mathbb{E}_{\mathcal{C}_{n}}(P_{i}Y_{j,n}|\mathcal{F}_{i-1,n})=\mathbb{E}_{\mathcal{C}_{n}}\{[\mathbb{E}_{\mathcal{C}_{n}}(Y_{j,n}|\mathcal{F}_{i,n})-\mathbb{E}_{\mathcal{C}_{n}}(Y_{j,n}|\mathcal{F}_{i-1,n})]|\mathcal{F}_{i-1,n}\}=[\mathbb{E}_{\mathcal{C}_{n}}(Y_{j,n}|\mathcal{F}_{i-1,n})-\mathbb{E}_{\mathcal{C}_{n}}(Y_{j,n}|\mathcal{F}_{i-1,n})]=0$. Hence, $\mathbb{E}(P_{i}Y_{j,n}|\mathcal{F}_{i-1,n})=\mathbb{E}\mathbb{E}_{\mathcal{C}_{n}}(P_{i}Y_{j,n}|\mathcal{F}_{i-1,n})=0$.} and $(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n})=\sum^{n}_{i=1}P_{i}Y_{j,n}$.

With Lemma (ref) and the Burkholder\textquoteright s inequality for martingales (Lemma (ref)), we obtain the following moment inequality.

thmLet $C_{p}\equiv\sqrt{p-1}$ when $p\geq2$ and $C_{p}\equiv\frac{1}{p-1}$ when $p\in(1,2)$. Then for any constant $p>1$, $\frac{1}{n}\left\Vert \sum^{n}_{j=1}(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n})\right\Vert _{L^{p},\mathcal{C}_{n}}\leq C_{p}\left\{ \Delta_{p,\min\{p,2\}}\left(\mathcal{C}_{n}\right)\right\} ^{1/\min\{p,2\}}$ a.s.

A direct result of Theorem (ref) is a law of large numbers for functionally dependent network random variables\@.

thmIf $\Delta_{p,\min\{p,2\}}\left(\mathcal{C}_{n}\right)=o_{\mathbb{P}}(1)$ for some constant $p>1$, then $\frac{1}{n}\sum^{n}_{j=1}(Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n})\xrightarrow{\mathbb{P}}0$.
remWe compare our LLN with that of kojevnikov_limit_2021. A key assumption in their LLN (Assumption 3.2) implicitly requires that the average dependence between two individuals decays to zero as their geodesic distance grows, and that the geodesic distances of most nodes are large. As they note in the discussion following Assumption 3.2, this assumption rules out networks with small diameters, which are common in social and financial networks. In contrast, our LLN accommodates networks with small diameters. As discussed in Remark (ref), our assumption $\Delta_{p,\min\{p,2\}}\left(\mathcal{C}_{n}\right)=o_{\mathbb{P}}(1)$ only requires that the average influence power of $e_{i,n}$'s on $Y_{j,n}$'s be finite, a condition that is feasible even for networks with small diameters (see Example (ref)).

Concentration inequality

Concentration inequalities, also called exponential inequalities in the literature, play an indispensable role in empirical process theory, as well as in semi-parametric, non-parametric, and high-dimensional statistics (wainwright2019HighDimensional). wu2005nonlinear and wu2016wu establish two concentration inequalities for functionally dependent stationary time series. In the literature, although there are some concentration inequalities for spatial data on irregular lattices (e.g., XU201896, yuan2025Bernsteintype, and wu2023application), the concentration inequalities for network data are rare. Hence, in this subsection, we follow the strategies in wu2016wu to establish a concentration inequality for network data.

thmLet $Z_{n}\equiv\frac{1}{\sqrt{n}}\sum^{n}_{j=1}Y_{j,n}$. Assume $\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n}=0$ for all $j=1,2,\ldots,n$ and $n\geq1$. Assume that $\sup_{n\geq1}n\Delta_{p,2}\left(\mathcal{C}_{n}\right)<\infty$ a.s. for all $p\geq2$ and $\sup_{n\geq1}\sqrt{n}\Delta^{1/2}_{p,2}\left(\mathcal{C}_{n}\right)$ increases slower than $O(p^{\nu})$ for some constant $\nu\geq0$ in the sense that \begin{equation} \sup_{p\geq2}\sup_{n\geq1}p^{-\nu}\sqrt{n}\Delta^{1/2}_{p,2}\left(\mathcal{C}_{n}\right)\leq\gamma_{0} a.s. \end{equation} for some finite constant $\gamma_{0}>0$. Let $\alpha=\frac{2}{1+2\nu}$. Then for $t\in[0,t_{0})$, \[ m\left(t\right)\equiv\mathbb{E}\left[\exp\left(t\left|Z_{n}\right|^{\alpha}\right)\mid\mathcal{C}_{n}\right]\leq1+c_{\alpha}\left(1-\frac{t}{t_{0}}\right)^{-1/2}\frac{t}{t_{0}}\text{ a.s.}, \] where $t_{0}=\left(e\alpha\gamma^{\alpha}_{0}\right)^{-1}$, $c_{\alpha}$ is a constant only depending on $\alpha$. Consequently, by letting $t=\frac{t_{0}}{2}$, we have for all $x>0$, \begin{equation} \mathbb{P}\left(\left|Z_{n}\right|\geq x\mid\mathcal{C}_{n}\right)\leq\exp\left(-tx^{\alpha}\right)m\left(t\right)\leq\left(1+\frac{\sqrt{2}c_{\alpha}}{2}\right)\exp\left(-\frac{x^{\alpha}}{2e\alpha\gamma^{\alpha}_{0}}\right),\quada.s. \end{equation}

We now illustrate when the condition ((ref)) holds. Consider the SAR model from Section (ref). From Proposition (ref), we have $\delta_{p,n}\left(j,i\right)\leq2\left\Vert \epsilon\right\Vert _{L^{p}}S^{+}_{ji,n}$.\footnote{For simplicity, here we assume that $\mathcal{C}_{n}$ is the trivial $\sigma$-field $\{\emptyset,\Omega\}$.} As a result, we obtain

align*[align* omitted — 268 chars of source]

If $\sup_{n\geq1}\frac{1}{n}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}S^{+}_{ji,n}\right]^{2}<C^{2}<\infty$ for some constant $C>0$, then \[ \sup_{p\geq2}\sup_{n\geq1}p^{-\nu}\sqrt{n}\Delta^{1/2}_{p,2}\leq2C\sup_{p\geq2}p^{-\nu}\left\Vert \epsilon\right\Vert _{L^{p}}. \] Thus, the condition ((ref)) holds with $\nu=1$ if $\epsilon_{i,n}$'s are uniformly sub-exponential, with $\nu=\frac{1}{2}$ if $\epsilon_{i,n}$'s are uniformly sub-Gaussian, and with $\nu=0$ if $\epsilon_{i,n}$'s are uniformly bounded. Compared to the Hoeffding's inequality (wainwright2019HighDimensional), we see that when $\nu=0$, the decay rate on the right-hand-side of ((ref)) with respect to $x$ is the same as in the independent case. When $\nu>0$, the decay rate is slower.

Central limit theorems

In this section, we establish two central limit theorems (CLTs) based on the FDM. Denote $Z_{i,n}\equiv\sum^{n}_{k=1}P_{i}Y_{k,n}$ for $i\in[n]$, where $P_{i}Y_{k,n}=\mathbb{E}_{\mathcal{C}_{n}}(Y_{k,n}|\mathcal{F}_{i,n})-\mathbb{E}_{\mathcal{C}_{n}}(Y_{k,n}|\mathcal{F}_{i-1,n})$. By construction, $\{Z_{i,n},\mathcal{F}_{i,n}\}^{n}_{i=1}$ is an MDA under both $\mathbb{P}_{\mathcal{C}_{n}}$ and $\mathbb{P}$. Therefore, we will prove our CLTs by applying a martingale CLT (Lemma (ref)) to the normalized sum $\sigma^{-1}_{n}\sum^{n}_{i=1}Z_{i,n}$. Before stating the CLTs, we introduce a second-order functional dependence measure, which is tailored to the analysis of the MDA increments $\{Z_{i,n}\}$ and, to the best of our knowledge, is new in the literature.

Fix $i\neq j$. Conditional on $\mathcal{C}_{n}$, let $e^{*}_{i,n}$ and $e^{*}_{j,n}$ be independent copies of $e_{i,n}$ and $e_{j,n}$, respectively, such that $(e_{1,n},\dots,e_{n,n},e^{*}_{i,n},e^{*}_{j,n})$ are mutually independent given $\mathcal{C}_{n}$. Define $Y_{k,n,\{i,j\}}$ as the coupled version of $Y_{k,n}$ obtained by replacing $e_{i,n}$ and $e_{j,n}$ with $e^{*}_{i,n}$ and $e^{*}_{j,n}$: \[ Y_{k,n,\{i,j\}}\equiv F_{k,n}(e_{1,n},\dots,e_{i-1,n},e^{*}_{i,n},e_{i+1,n},\dots,e_{j-1,n},e^{*}_{j,n},e_{j+1,n},\dots,e_{n,n}) \] for $i<j$. When $j<i$, the same definition applies with the roles of $i$ and $j$ swapped. Equivalently, the above display is understood as replacing the $i$th and $j$th coordinates simultaneously.

defn[Second-order functional dependence measure] For $p\geq1$, define the second-order functional dependence measure (second-order FDM) of the random field $\left\{ Y_{i,n}\right\} $ as \begin{equation} \mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\equiv\left\Vert \sum^{n}_{k=1}\left\{ Y_{k,n}-Y_{k,n,i}-Y_{k,n,j}+Y_{k,n,\{i,j\}}\right\} \right\Vert _{L^{p},\mathcal{C}_{n}} for i\neq j \end{equation} and \[ \mathfrak{d}_{p,n}\left(i,i,\mathcal{C}_{n}\right)\equiv\left\Vert \sum^{n}_{k=1}\left\{ Y_{k,n}-Y_{k,n,i}\right\} \right\Vert _{L^{p},\mathcal{C}_{n}}. \] When $\mathcal{C}_{n}$ is the trivial $\sigma$-field $\left\{ \Omega,\emptyset\right\} $, we simplify the notation as $\mathfrak{d}_{p,n}\left(i,j\right)\equiv\mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$.

Why is this second-order FDM needed? To apply the martingale CLT, we need to establish the LLN for the squared increments $\{Z^{2}_{i,n}\}$. A direct way to show this is to invoke Theorem (ref), which requires controlling the (first-order) FDM of $\{Z^{2}_{i,n}\}$ (viewed as functions of the innovations $e_{i,n}$). By applying the difference of squares formula, controlling the FDM of $\{Z^{2}_{i,n}\}$ reduces to controlling the (first-order) FDM of $\{Z_{i,n}\}$. It turns out that this FDM can then be bounded in terms of the second-order FDM of the original field $\{Y_{i,n}\}$, as summarized below.

lemLet $Z_{i,n}=\sum^{n}_{j=1}P_{i}Y_{j,n}$ for $i\in[n]$ and denote the FDM of $\{Z_{i,n}\}$ on $\{e_{i,n}\}$ as $\delta^{\sharp}_{p,n}(i,j,\mathcal{C}_{n})$ for $i,j\in[n]$. Then, $\delta^{\sharp}_{p,n}(i,j,\mathcal{C}_{n})\leq\mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)=\mathfrak{d}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ a.s. for all $i,j\in[n]$.

Lemma (ref) can be viewed as a second-order analogue of Lemma (ref). Indeed, Lemma (ref) implies $\left\Vert Z_{i,n}\right\Vert _{L^{p},\mathcal{C}_{n}}\leq\sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})$; that is, the $L^{p}$ norm of $Z_{i,n}$, which can be interpreted as a “zeroth-order” FDM for $\{Z_{i,n}\}$, is controlled by the (first-order) FDM of $\{Y_{i,n}\}$. In contrast, Lemma (ref) shows that the first-order FDM of $\{Z_{i,n}\}$ is controlled by the second-order FDM of $\{Y_{i,n}\}$. With this lemma in hand, we are ready to state our CLTs.

A univariate CLT

thmDenote $\sigma^{2}_{n}\equiv\mathrm{Var}_{\mathcal{C}_{n}}\left(\sum^{n}_{j=1}Y_{j,n}\right)$. Let $p>2$ be a constant. Suppose that (1) $\sigma^{-2}_{n}=O_{\mathbb{P}}(n^{-1})$, (2) \begin{equation} \frac{1}{n^{p/2}}\sum^{n}_{i=1}\left\{ \sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})\right\} ^{p}=o_{\mathbb{P}}(1), \end{equation} and (3) \begin{equation} \frac{1}{n^{\min\{2,p/2\}}}\sum^{n}_{i=1}\left[\left\{ \sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})\right\} \sum^{n}_{j=1}\mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right]^{\min\{2,p/2\}}=o_{\mathbb{P}}(1) \end{equation} as $n\to\infty$. Then \[ \frac{\sum^{n}_{j=1}\left\{ Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n}\right\} }{\sigma_{n}}\xrightarrow{d}N\left(0,1\right). \]
remWhile Theorem (ref) establishes asymptotic normality, feasible inference additionally requires a consistent estimator of $\sigma^{2}_{n}$. At this level of generality, we do not develop a universal HAC-type estimator for the asymptotic variance under network dependence. Rather, variance estimation is expected to be model-specific. For example, for likelihood-based estimators, one may exploit the information equality and estimate the variance using the Hessian matrix. More generally, in parametric settings where the error distribution belongs to a known family, simulation-based methods may be used to approximate the asymptotic variance. A systematic treatment of feasible variance estimation under the present abstract framework is left for future work.

We next discuss how to verify the high-level conditions in Theorem (ref). First note that conditions ((ref)) and ((ref)) are invariant under relabeling of the nodes $1,\ldots,n$. This is a natural requirement because the validity of the CLT should depend on the dependence structure of the network, not on an arbitrary ordering.

Condition ((ref)) is based on the (first-order) FDM and relatively easy to verify. Condition ((ref)) involves the second\nobreakdash-order FDM and is more demanding. The next two lemmas provide two different ways to control the second-order FDM. Lemma (ref) exploits smoothness of the functions $F_{j,n}$ (recall that $Y_{j,n}=F_{j,n}(e_{n})$), whereas Lemma (ref) does not require smoothness and thus applies more generally.

lemConsider the system ((ref)). Suppose that $e_{1,n},\ldots,e_{n,n}$ are univariate, i.e., $p_{e}=1$. Let $\left\Vert e\right\Vert _{L^{p},\mathcal{C}_{n}}\equiv\sup_{i\in[n]}\left\Vert e^{*}_{i,n}-e_{i,n}\right\Vert _{L^{p},\mathcal{C}_{n}}$, where $e^{*}_{i,n}$ is an i.i.d. copy of $e_{i,n}$ conditional on $\mathcal{C}_{n}$. Assume the (random) functions $F_{1,n}(\cdot),\ldots,F_{n,n}(\cdot)$ are measurable with respect to $\mathcal{C}_{n}$ and twice continuously differentiable a.s. Then it holds that $\mathfrak{d}_{p,n}\left(i,i,\mathcal{C}_{n}\right)\leq\sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n}),\ \forall i\in[n],$ and \[ \mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\leq\left\Vert e\right\Vert ^{2}_{L^{p},\mathcal{C}_{n}}\sup_{e_{n}}\left|\sum^{n}_{k=1}\frac{\partial^{2}F_{k,n}(e_{n})}{\partial e_{i,n}\partial e_{j,n}}\right|,\ \forall i\neq j\text{ a.s.} \]

Lemma (ref) is particularly convenient for smooth models. In Section (ref), we apply it to the SAR model (see Lemma (ref) and Proposition (ref)). The next lemma provides an alternative route that avoids smoothness assumptions.

lemConsider the system ((ref)). \begin{enumerate}[label=(\roman*)] • For any $i,j\in[n]$, it holds that \[ \mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\leq2\sum^{n}_{k=1}\min\left\{ \delta_{p,n}(k,i,\mathcal{C}_{n}),\delta_{p,n}(k,j,\mathcal{C}_{n})\right\} . \] • For each $k\in[n]$, let $\pi_{k}(1),\ldots,\pi_{k}(n)$ be a permutation of $1,\ldots,n$ such that \[ \delta_{p,n}\left(k,\pi_{k}(1),\mathcal{C}_{n}\right)\ge\delta_{p,n}\left(k,\pi_{k}(2),\mathcal{C}_{n}\right)\ge\cdots\ge\delta_{p,n}\left(k,\pi_{k}(n),\mathcal{C}_{n}\right). \] If $\sup_{i\in[n]}\sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})=O_{\mathbb{P}}(1),$and there exist two constants $\alpha>\frac{\min\{2,p/2\}}{\min\{2,p/2\}-1}>1$ and $C>0$ such that \begin{equation} \delta_{p,n}\left(k,\pi_{k}(r),\mathcal{C}_{n}\right)\le Cr^{-\alpha},\qquad\forall\,k,r\in[n], \end{equation} with probability approaching one, then condition ((ref)) holds. \end{enumerate}

Condition ((ref)) requires that, for each node $k$, the first-order dependence measures $\delta_{p,n}\left(k,\pi_{k}(r),\mathcal{C}_{n}\right)$ decay at a polynomial rate when ordered from largest to smallest. In other words, although a node may be affected by many other nodes, the strength of these influences must decline sufficiently fast across the ordered neighbors.

Comparison with the CLTs in the literature

We now compare our CLT with existing results in the literature.

(1) The CLTs in wu2005nonlinear,wu2023application,ELMACHKOURI20131,jenish2009central,jenish2012spatial,kojevnikov_limit_2021 all require individuals to be located in some metric space. Our CLT imposes no such metric structure, making it applicable to network data, where a natural embedding may be absent.

(2) The key condition for CLTs in the literature (e.g., jenish2009central,jenish2012spatial,kojevnikov_limit_2021,wu2023application) typically takes the form $\sum^{\infty}_{s=1}\nu_{s}\theta_{s}<\infty$, where $s$ denotes distance, $\nu_{s}$ is the number of individuals at distance approximately equal to $s$ from a given individual, and $\theta_{s}$ is a weak dependence coefficient at distance $s$ (e.g., a strong mixing, NED, or $\psi$-dependence coefficient). For this condition to hold, one typically needs $\theta_{s}$ to decay sufficiently fast as $s$ increases, and the growth of $\nu_{s}$ must also be controlled. As a result, these conditions are most natural in settings where dependence weakens gradually over many distance shells. This requirement can be restrictive for network data, since many relevant networks exhibit small-world behavior and hence have small graph diameter.\footnote{For example, in the Erd\H{o}s-R\'{e}nyi model in Example (ref), if the mean degree satisfies $D_{n}\propto n^{\alpha}$ ($0<\alpha<1)$, then results in klee1981diameters imply that the graph diameter (based on the geodesic distance) is bounded by a constant with probability approaching one. Consequently, the graph radius is also of constant order. In such settings, existing weak dependence concepts based on decay with network distance become less compelling, because only finitely many distance shells are asymptotically relevant.} In contrast, our condition ((ref)) only requires that, for any node $i$, the magnitude of the impact it receives from other nodes (i.e., the FDM $\delta_{p,n}\left(k,\pi_{k}(r),\mathcal{C}_{n}\right)$) decreases as a power function of $r$ when these impacts are ordered from largest to smallest. This allows our CLT to accommodate networks with small or moderate diameters, a common feature of social and financial networks.\footnote{Financial networks, ranging from interbank markets to asset holding networks, typically exhibit small diameters. For example, consider the holding-based network of Chinese mutual funds. Funds $i$ and $j$ are regarded as connected if they allocate at least 5% of their portfolios to the same stocks. Using data from the second half of 2015, jiang2026NAVaR find that this network has a diameter of only 4.}

(3) kuersteiner2019limit develops a CLT for spatial mixingale processes using a model\nobreakdash-dependent random metric, relaxing the assumption of a fixed metric. kuersteiner2019limit assumes the network to be sparse that “rules out a buildup of a mass of nodes with very similar features is captured by a summability condition of the probabilities that two nodes are close in an appropriate sense”. In contrast, our approach does not require any form of underlying metric and allows for networks with small diameters.

(4) Going beyond the conventional notion of weak dependence (e.g., mixing and NED), leung_normal_2019 propose the \textquotedblleft stabilization\textquotedblright conditions that impose weak dependence on node degrees. While their motivation is similar, they focus on strategic network formation, whereas our concept applies to various types of networks.

(5) lee2022qml study a CLT for a linear-quadratic form of independent random variables in the presence of dominant units (also called popular units in their paper). Their CLT is mainly used for linear SAR type models, but ours is applicable to nonlinear models and some nonlinear estimators (e.g., Huber or quantile regression estimators) of linear SAR models.

A multivariate CLT

Using the Cram\'er-Wold device, we can generalize Theorem (ref) to multivariate case. Now, $Y_{j,n}$'s are random vectors taking values in $\mathbb{R}^{p_{Y}}$ ($p_{Y}\geq1$). Denote $\Sigma_{n}\equiv\mathrm{Var}_{\mathcal{C}_{n}}\left(\sum^{n}_{j=1}\text{\ensuremath{Y_{j,n}}}\right)$, and let $I_{p_{Y}}$ be the $p_{Y}\times p_{Y}$ identity matrix.

thmSuppose that (1) $\{\min\mathrm{eig}(\Sigma_{n})\}^{-1}=O_{\mathbb{P}}(n^{-1})$, (2) $\frac{1}{n^{p/2}}\sum^{n}_{i=1}\left\{ \sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})\right\} ^{p}=o_{\mathbb{P}}(1),$ and (3) $\frac{1}{n^{\min\{2,p/2\}}}\sum^{n}_{i=1}\left[\left\{ \sum^{n}_{k=1}\delta_{p,n}(k,i,\mathcal{C}_{n})\right\} \sum^{n}_{j=1}\mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right]^{\min\{2,p/2\}}=o_{\mathbb{P}}(1)$ as $n\to\infty$. Then, $\Sigma^{-1/2}_{n}\sum^{n}_{j=1}\left\{ Y_{j,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{j,n}\right\} \xrightarrow{d}N\left(0,I_{p_{Y}}\right).$

Application of functional dependence to SAR models

To illustrate the application of the FDM, in this section, we calculate the FDM for SAR models and establish the consistency of the maximum likelihood estimator for the SAR Tobit model using FDM.

SAR models

Let $F:\mathbb{R}\to\mathbb{R}$ be a Lipschitz function, i.e., $\left|F\left(x\right)-F\left(y\right)\right|\leq L\left|x-y\right|$ for some constant $L>0$ and all $x,y\in\mathbb{R}$. A (possibly nonlinear) SAR model can be written as

equation[equation omitted — 112 chars of source]

for $j=1,\ldots,n$, where $w_{j\cdot,n}$ is the $j$th row of a non-zero spatial/network weights matrix $W_{n}=\left(w_{ji,n}\right)_{n\times n}$, $Y_{n}=\left(Y_{1,n},Y_{2,n},\cdots,Y_{n,n}\right)'$, $X_{j,n}\in\mathbb{R}^{p}$ is the exogenous regressor, $\epsilon_{j,n}$ is the disturbance term, and $\lambda$ and $\beta$ are model parameters. When $F\left(x\right)=x$, Eq.(ref) is the standard (linear) SAR model; when $F\left(x\right)=\max\left(0,x\right)$, Eq.(ref) becomes a SAR Tobit model. Now, we investigate the FDM of SAR model. The following two assumptions on $F(\cdot)$ and $\{\epsilon_{j,n}\}$ are needed.

assumption$F$ is a Lipschitz function with a Lipschitz constant $L>0$, and $\zeta\equiv L\left|\lambda\right|\sup_{n}\left\Vert W_{n}\right\Vert _{\infty}<1$.
remAssumption (ref) is also considered by wu2023application, and it ensures the existence and uniqueness of the solution of Eq.(ref). It is analogous to the stationarity condition for autoregressive models in time series.
assumptionDenote $\mathcal{C}_{n}\equiv\bigvee^{n}_{j=1}\sigma\left(X_{j,n}\right)\bigvee\sigma(W_{n})$. $\epsilon_{j,n}$'s are conditionally independent given $\mathcal{C}_{n}$ and $\left\Vert \epsilon\right\Vert _{L^{p},\mathcal{C}_{n}}\equiv\sup_{j,n}\left\Vert \epsilon_{j,n}\right\Vert _{L^{p},\mathcal{C}_{n}}<\infty$ a.s. for some $p>1$.

Note that Assumption (ref) allows the weights matrix $W_{n}$ to be random. Let $e_{j,n}=X'_{j,n}\beta+\epsilon_{j,n}$ in the SAR model (ref). Then $Y_{i,n}$'s are generated by the underlying random variables $e_{1,n},\ldots,e_{n,n}$. From Assumption (ref), $e_{j,n}$'s are conditionally independent on $\mathcal{C}_{n}$. Denote $\left|W_{n}\right|\equiv(|w_{ij,n}|)_{n\times n}$ and

equation[equation omitted — 126 chars of source]

where $I_{n}$ denotes the $n\times n$ identity matrix.

propUnder Assumptions (ref)-(ref), $\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq2\left\Vert \epsilon\right\Vert _{L^{p},\mathcal{C}_{n}}S^{+}_{ji,n}$ a.s.

As a direct consequence of Proposition (ref) and Theorem (ref), we have the following LLN for the dependent variable $Y_{i,n}$ generated by the SAR model ((ref)).

propSuppose Assumptions (ref)-(ref) hold. Let $\mu_{\min\{p,2\}}\equiv\left\{ \frac{1}{n}\sum^{n}_{i=1}\left\Vert S^{+}_{\cdot i,n}\right\Vert ^{\min\{p,2\}}_{1}\right\} ^{\frac{1}{\min\{p,2\}}}$, where $S^{+}_{\cdot i,n}$ denotes the $i$th column of $S^{+}_{n}$. If $\mu_{\min\{p,2\}}=o_{\mathbb{P}}(n^{\frac{\min\{p,2\}-1}{\min\{p,2\}}})$, then for the SAR model ((ref)), it holds that $\frac{1}{n}\sum^{n}_{i=1}(Y_{i,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{i,n})\xrightarrow{\mathbb{P}}0$.

Proposition (ref) directly follows from Proposition (ref) and Theorem (ref), and we omit the proof. Since $\left\Vert S^{+}_{n}\right\Vert _{1}\geq\mu_{\min\{p,2\}}$, a sufficient condition for LLN is $\left\Vert S^{+}_{n}\right\Vert _{1}=o_{\mathbb{P}}(n^{\frac{\min\{p,2\}-1}{\min\{p,2\}}})$.

Next, we bound the second-order FDM for SAR models in order to establish CLT for the dependent variable $Y_{i,n}$ generated by the SAR model ((ref)). To do so, we make the following assumption on the smoothness of the function $F$.

assumptionThe function $F:\mathbb{R}\to\mathbb{R}$ is twice continuously differentiable. Moreover, its second derivative is uniformly bounded, i.e., $\|F''\|_{\infty}\equiv\sup_{x\in\mathbb{R}}|F''(x)|<\infty.$

Relying on Lemma (ref), we have the following lemma on bounding the second-order FDM for SAR models.

lemUnder Assumptions (ref)--(ref), the second-order FDM for $\{Y_{i,n}\}$ under the SAR model ((ref)) satisfies $\mathfrak{d}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\leq\frac{4\|F''\|_{\infty}}{L^{3}}\left\Vert \epsilon\right\Vert ^{2}_{L^{p},\mathcal{C}_{n}}\sum^{n}_{k=1}\left\Vert S^{+}_{\cdot k,n}\right\Vert _{1}S^{+}_{ki,n}S^{+}_{kj,n}\ \text{ for }i\neq j$ and $\mathfrak{d}_{p,n}\left(i,i,\mathcal{C}_{n}\right)\leq2\left\Vert \epsilon\right\Vert _{L^{p},\mathcal{C}_{n}}\sum^{n}_{k=1}S^{+}_{ki,n}\ \text{ for }i\in[n]\text{ a.s.}$

Based on Lemma (ref), we impose the following assumption to verify condition ((ref)).

assumptionFor any $q>1$, define $\mu_{q}\equiv\left\{ \frac{1}{n}\sum^{n}_{i=1}\left\Vert S^{+}_{\cdot i,n}\right\Vert ^{q}_{1}\right\} ^{1/q}$ and $\mu_{\infty}\equiv\left\Vert S^{+}_{n}\right\Vert _{1}$, where $S^{+}_{\cdot i,n}$ denotes the $i$th column of $S^{+}_{n}$. Let $\tilde{p}\equiv\min\{2,p/2\}$. The matrix $S^{+}_{n}$ satisfies (i) $\mu^{2\tilde{p}}_{2\tilde{p}}=o_{\mathbb{P}}(n^{\tilde{p}-1})$ and (ii) there exists a pair of conjugate indices $u,v\in[1,\infty]$ such that $1/u+1/v=1$ (with the convention $1/\infty=0$) and $\mu^{(2\tilde{p}-1)+1/u}_{u(2\tilde{p}-1)+1}\cdot\mu^{\tilde{p}}_{v\tilde{p}}=o_{\mathbb{P}}(n^{\tilde{p}-1}).$

Notice that $\mu_{q}$ increases to $\mu_{\infty}$ as $q\to\infty$. We provide a sufficient condition for Assumption (ref). If $\mu_{\infty}=o_{\mathbb{P}}(n^{(\tilde{p}-1)/(3\tilde{p}-1)})$, we have $\mu^{2\tilde{p}}_{2\tilde{p}}\leq\mu^{2\tilde{p}}_{\infty}=o(n^{2\tilde{p}(\tilde{p}-1)/(3\tilde{p}-1)})=o_{\mathbb{P}}(n^{\tilde{p}-1}).$ Besides, by letting $u=\infty$ and $v=1$, we have $\mu^{(2\tilde{p}-1)+1/u}_{u(2\tilde{p}-1)+1}\cdot\mu^{\tilde{p}}_{v\tilde{p}}\leq\mu^{2\tilde{p}-1}_{\infty}\mu^{\tilde{p}}_{1}\leq\mu^{2\tilde{p}-1}_{\infty}\mu^{\tilde{p}}_{\infty}=o_{\mathbb{P}}(n^{\tilde{p}-1}).$ Thus, a sufficient condition for Assumption (ref) is $\|S^{+}_{n}\|_{1}=o_{\mathbb{P}}\bigl(n^{(\tilde{p}-1)/(3\tilde{p}-1)}\bigr)$. An important special case is $\sup_{n}\|S^{+}_{n}\|_{1}=O_{\mathbb{P}}(1)$, which is widely imposed in the literature; see, for example, lee2004asymptotic,lee_gmm_2007,yu2008quasi,xu2015maximum. Moreover, since $\frac{\tilde{p}-1}{3\tilde{p}-1}\le\frac{\min\{p,2\}-1}{\min\{p,2\}},$ the condition $\|S^{+}_{n}\|_{1}=o_{\mathbb{P}}(n^{(\tilde{p}-1)/(3\tilde{p}-1)})$ is stronger than the corresponding condition used for the LLN. Hence, by Proposition (ref), it also implies the LLN.

Next, we state the CLT result for the SAR model.

propSuppose Assumptions (ref)--(ref) hold with $p>2$. Then conditions ((ref))--((ref)) hold for the dependent variable $Y_{i,n}$ generated by the SAR model ((ref)). In addition, if $\sigma^{-2}_{n}\equiv\left\{ \mathrm{Var}_{\mathcal{C}_{n}}\left(\sum^{n}_{j=1}Y_{j,n}\right)\right\} ^{-1}=O_{\mathbb{P}}(n^{-1})$, then $\sigma^{-1}_{n}\sum^{n}_{i=1}\left\{ Y_{i,n}-\mathbb{E}_{\mathcal{C}_{n}}Y_{i,n}\right\} \xrightarrow{d}N\left(0,1\right)$.

We next verify Assumption (ref) in several commonly used network settings. We start with a deterministic design that allows for a small number of highly influential units. We then consider three random graph models: the Erd\H{o}s-R\'{e}nyi model, the triangle model, and the stochastic block model. Together, these examples show that Assumption (ref) holds under a range of network structures of practical interest.

example[Dominant popular units]We first consider the “dominant popular units” structure for the network weights matrix $W_{n}$ proposed by lee2022qml. In this example, the network matrix is nonrandom and may contain a small group of highly influential units. This setting serves as a useful benchmark because it isolates, in a deterministic way, the effect of column concentration generated by dominant units. Following lee2022qml, partition the network weights matrix as \begin{equation} W_{n}=\begin{pmatrix}W_{n,11} & W_{n,1B}\\ W_{n,B1} & W_{n,BB} \end{pmatrix}, \end{equation} where the first block corresponds to the $m_{n}$ dominant units, with $m_{n}=O(n^{\eta_{m}})$ for some $0\le\eta_{m}<1$. We assume that the column sums associated with the dominant units in $W_{n}$ are of order $O(n^{\delta})$ for some $0\le\delta<1$; that is, the column sums of $\left(W'_{n,11},W'_{n,B1}\right)'$ have magnitude $O(n^{\delta})$.

Within this framework, it is useful to distinguish two topological scenarios according to whether non-dominant units can feed back into the dominant ones.

casenv• General dominant network ($W_{n,1B}\neq0$). In this case, non-dominant units may influence the dominant units. As emphasized in Proposition 3(a) of lee2022qml, this feedback channel can amplify propagation through the hubs and therefore makes Assumption (ref) harder to satisfy. • Strict Dominant Structure ($W_{n,1B}=0$). In this case, non-dominant units do not influence the dominant units. This removes the feedback-amplification channel and allows a sharper bound on the column sum of $S^{+}_{n}$; see Proposition 3(b) of lee2022qml.

The next proposition makes this comparison precise by characterizing when Assumption (ref) holds under each of the two dominant-unit configurations.

prop[Dominant popular units] Under the dominant popular units structure in Example (ref), suppose that $\eta_{m}+\delta<1$ and that $\left\Vert \left(I_{n-m_{n}}-L\left|\lambda W_{n,BB}\right|\right)^{-1}\right\Vert _{1}$ and $\left\Vert \left(I_{n-m_{n}}-L\left|\lambda W_{n,BB}\right|\right)^{-1}\right\Vert _{\infty}$ are uniformly bounded. Then: \begin{enumerate}[label=(\roman*)] • for the general dominant network ($W_{n,1B}\neq0$), Assumption (ref) holds with $p=4$ if $0\leq\eta_{m}+\delta<1/5$; • for the strict dominant structure ($W_{n,1B}=0$), Assumption (ref) holds with $p=4$ if $(\eta_{m},\delta)\in\{(\eta_{m},\delta):0\leq\eta_{m}<\frac{1}{3},0\leq\delta<\frac{3-5\eta_{m}-\sqrt{(1-\eta_{m})(5-7\eta_{m})}}{2}\}$. \end{enumerate}

Proposition (ref) shows that the strict dominant structure permits a wider range of $(\eta_{m},\delta)$ than the general case. This reflects the fact that ruling out feedback from non-dominant units to dominant units makes the concentration of column influence easier to control. In particular, under the strict dominant structure, when $\eta_{m}=0$ we require $\delta<(3-\sqrt{5})/2\approx0.382$, whereas when $\eta_{m}=1/5$ we require $\delta<(5-3\sqrt{2})/5\approx0.152$.

We next consider network weights matrices generated by random graph models.

example[Erd\H{o}s-R\'{e}nyi model]Suppose that the network is generated by an undirected Erd\H{o}s-R\'{e}nyi graph without self-links. For each unordered pair of distinct individuals $\{i,j\}$, a link is formed independently with probability $\frac{D_{n}}{n-1}$, where $D_{n}\in[0,n-1]$ is a deterministic sequence that parameterizes the expected degree. This normalization is natural because each node has $n-1$ potential neighbors, so if $d_{i,n}$ denotes the degree of node $i$, then $d_{i,n}\sim\mathrm{Binomial}(n-1,\frac{D_{n}}{n-1})$ and hence $\mathbb{E}[d_{i,n}]=D_{n}.$ Let $A_{n}=(A_{ij,n})_{n\times n}$ denote the resulting adjacency matrix, with $A_{ii,n}=0$ and $A_{ij,n}=A_{ji,n}$. The network weights matrix $W_{n}=(w_{ij,n})_{n\times n}$ is defined by row-normalizing $A_{n}$: \[ w_{ij,n}=\begin{cases} A_{ij,n}/d_{i,n}, & d_{i,n}>0,\\ 0, & d_{i,n}=0. \end{cases} \]

The ER model provides a convenient baseline for the random-graph setting. The next proposition shows that Assumption (ref) generally holds for the ER model.

prop[Erd\H{o}s-R\'{e}nyi model] Recall that $S^{+}_{n}=L\left(I_{n}-L\left|\lambda W_{n}\right|\right)^{-1}$. In Example (ref), if $L|\lambda|<1$, then Assumption (ref) holds.

Proposition (ref) shows that Assumption (ref) holds for the Erd\H{o}s-R\'{e}nyi model without any further restriction on the growth rate of $D_{n}$. This highlights that the proposed dependence concept can accommodate highly connected networks with small or moderate diameter.

example[Triangle model]We next consider a random graph model with local clustering. Related subgraph-based random graph models with links and triangles are studied by chandrasekhar2025Network. Suppose that each unordered triple of distinct individuals forms a triangle with probability $T_{n}/\binom{n}{3}$, where $T_{n}$ denotes the expected number of triangles in the network. Let $A^{(\mathrm{T})}_{n}=(A^{(\mathrm{T})}_{ij,n})_{n\times n}$ denote the adjacency matrix generated by these triangles, where \[ A^{(\mathrm{T})}_{ij,n}=\begin{cases} 1, & \text{if nodes }i\text{ and }j\text{ belong to at least one selected triangle},\\ 0, & \text{otherwise}. \end{cases} \] In addition, let $A^{(\mathrm{ER})}_{n}$ be an independent Erd\H{o}s-R\'{e}nyi adjacency matrix with link probability $\frac{D_{n}}{n-1}$, so that $D_{n}$ denotes the expected degree contributed by the background links. We define the overall adjacency matrix by $A_{n}=(A_{ij,n})=(\max\{A^{(\mathrm{T})}_{ij,n},A^{(\mathrm{ER})}_{ij,n}\})_{n\times n},$ and let $W_{n}$ be the associated row-normalized weights matrix. The following proposition verifies Assumption (ref) for this triangle network design under a simple restriction on the overall level of expected degree. For $1\le i<j<k\le n$, let $B_{\{i,j,k\},n}\equiv1$ when the triangle generating process generates a triangle $\{i,j,k\}$, and $B_{\{i,j,k\},n}\equiv0$ otherwise.
prop[Triangle model] In Example (ref), suppose that the triangle indicators $B_{\{i,j,k\},n}$ are i.i.d. $\mathrm{Bernoulli}\!\left(\frac{T_{n}}{\binom{n}{3}}\right)$, and the Erd\H{o}s-R\'{e}nyi edges $\{A^{(\mathrm{ER})}_{ij,n}:1\le i<j\le n\}$ are i.i.d. $\mathrm{Bernoulli}\!\left(\frac{D_{n}}{n-1}\right)$, independent of the triangle indicators. If $L|\lambda|<1$ and \begin{equation} D_{n}+\frac{6T_{n}}{n}\le C_{0}\log n \end{equation} for some constant $C_{0}>0$ and all sufficiently large $n$, then Assumption (ref) holds.

This result shows that Assumption (ref) continues to hold when the network exhibits local clustering, provided that the overall degree level remains sufficiently controlled. We next turn to a different source of network heterogeneity, namely latent community structure.

example[Stochastic block model]We finally consider the stochastic block model (SBM), which allows link probabilities to differ within and across latent communities. The SBM is widely used in the analysis of network community structure; see, for example, abbe2018community and wu2023distributed. Suppose there are $M_{n}$ blocks in total, and each individual $i$ is independently assigned to a latent block $g_{i,n}\in\{1,\ldots,M_{n}\}$ with probability $\mathbb{P}(g_{i,n}=m)=1/M_{n}$, $m=1,\ldots,M_{n}$. Let $D_{\mathrm{wb},n}$ and $D_{\mathrm{bb},n}$ denote the within-block and between-block expected degrees, respectively. Conditional on the block assignments, links are formed independently, with \[ \mathbb{P}(A_{ij,n}=1\mid g_{i,n}=g_{j,n})=\frac{D_{\mathrm{wb},n}}{n/M_{n}-1},\qquad\mathbb{P}(A_{ij,n}=1\mid g_{i,n}\neq g_{j,n})=\frac{D_{\mathrm{bb},n}}{n-n/M_{n}}, \] for $1\le i<j\le n$, and $A_{ii,n}=0$, $A_{ij,n}=A_{ji,n}$. The normalization is chosen so that $D_{\mathrm{wb},n}$ and $D_{\mathrm{bb},n}$ correspond to the expected numbers of within-block and between-block neighbors, respectively. Let $A_{n}=(A_{ij,n})_{n\times n}$ denote the resulting adjacency matrix, and let $W_{n}$ be the associated row-normalized weights matrix. The next proposition verifies Assumption (ref) under this SBM specification.
prop[Stochastic block model] In Example (ref), suppose that $M_{n}\ge2$, $\frac{D_{\mathrm{wb},n}}{n/M_{n}-1}\in[0,1]$, and $\frac{D_{\mathrm{bb},n}}{n-n/M_{n}}\in[0,1]$. If $M_{n}\log M_{n}=o(n)$ and $L|\lambda|<1$, then Assumption (ref) holds.

This result shows that Assumption (ref) continues to hold in the presence of community structure. It is also worth noting that, beyond the requirement that the edge probabilities lie in $[0,1]$, the proposition does not impose a separate growth condition on $D_{\mathrm{wb},n}$ or $D_{\mathrm{bb},n}$.

Taken together, the above examples show that Assumption (ref) holds in a broad class of network settings, including deterministic designs with influential units as well as standard random graph models. As a result, the CLT developed in Proposition (ref) applies under a range of dependence patterns commonly used in empirical and theoretical work on networks.

Application to the MLE of the SAR Tobit model

In this section, we investigate the consistency of the maximum likelihood estimator (MLE) for the SAR Tobit model in xu2015maximum. Specifically, for each node $i\in[n]$, the model is defined as: \[ Y_{i,n}=\max\{0,Y^{*}_{i,n}\},\qquad Y^{*}_{i,n}=\lambda_{0}w_{i\cdot,n}Y_{n}+X^{\prime}_{i,n}\beta_{0}+\epsilon_{i,n}, \] where the exogenous regressors $X_{i,n}$ and the network weights matrix $W_{n}$ generate the sub-$\sigma$-field $\mathcal{C}_{n}\equiv\bigvee^{n}_{j=1}\sigma\left(X_{j,n}\right)\bigvee\sigma(W_{n})$. Conditional on $\mathcal{C}_{n}$, the innovations $\epsilon_{1,n},\dots,\epsilon_{n,n}$ are assumed to be i.i.d. following a normal distribution $N(0,\sigma^{2}_{0})$. Fix a parameter value $\theta=(\lambda,\beta,\sigma)$. Let $z_{i,n}(\theta)=(Y_{i,n}-\lambda w_{i\cdot,n}Y_{n}-X^{\prime}_{i,n}\beta)/\sigma$ and $G_{n}(Y_{n})=\text{diag}(1(Y_{1,n}>0),\dots,1(Y_{n,n}>0))$. The corresponding log-likelihood function can be written as

align*[align* omitted — 261 chars of source]

where $\Phi(\cdot)$ is the cumulative distribution function of the standard normal distribution. Our goal here is not to re-establish the full consistency argument, but rather to show that the key pointwise convergence step in that argument can be obtained without relying on an underlying metric space. In the framework of xu2015maximum, the metric-space assumption is used to establish the pointwise convergence of $\frac{1}{n}\ell_{n}(\theta)$. Accordingly, we focus on proving that $\frac{1}{n}\ell_{n}(\theta)=\frac{1}{n}\mathbb{E}_{\mathcal{C}_{n}}\ell_{n}(\theta)+o_{\mathbb{P}}(1)$, which is one ingredient in the consistency proof. The following assumptions are required for the subsequent analysis.

assumption\begin{enumerate}[label=(\roman*)] • Conditional on $\mathcal{C}_{n}$, $\epsilon_{1,n},\dots,\epsilon_{n,n}$ are i.i.d. normal. • $0<|\lambda_{0}|<1$ and $\|W_{n}\|_{\infty}\le1$ a.s. for all $n$. • $\mu_{2}=\left(\frac{1}{n}\sum^{n}_{i=1}\left\Vert S^{+}_{\cdot i,n}\right\Vert ^{2}_{1}\right)^{\frac{1}{2}}=o_{\mathbb{P}}(\sqrt{n})$, where $S^{+}_{n}\equiv(I_{n}-|\lambda_{0}W_{n}|)^{-1}$. • Conditional on $\mathcal{C}_{n}$, elements in $X_{i,n}$ are uniformly bounded for all $i$, $n$. • $\sup_{i,n}\sup_{y}f_{Y^{*}_{i,n}\mid\mathcal{C}_{n}}(y)\le C_{f}<\infty$ a.s. for some constant $C_{f}>0$, where $f_{Y^{*}_{i,n}\mid\mathcal{C}_{n}}(y)$ is the conditional density of $Y^{*}_{i,n}$. \end{enumerate}

See xu2015maximum for primitive conditions for Assumption (ref)(v). We then obtain the following result.

propUnder Assumption (ref), for any fixed $\theta$ such that $\left|\lambda\right|<1$ and $\sigma>0$, we have that $\frac{1}{n}\ell_{n}(\theta)=\frac{1}{n}\mathbb{E}_{\mathcal{C}_{n}}\ell_{n}(\theta)+o_{\mathbb{P}}(1)$.

Proposition (ref) shows that this pointwise convergence of the log-likelihood can be established within the FDM framework, without imposing any underlying metric-space structure. This provides one of the probabilistic ingredients used in the consistency analysis of the SAR Tobit MLE.

Functional Dependence Measure Under Transformations

To study the asymptotic properties of an estimator, we usually need to deal with various functions of random variables, e.g., $Y^{2}_{i,n}$ or $\Phi(Y_{i,n})$, where $\Phi(\cdot)$ is a distribution function of some random variable. Therefore, to broaden the applicability of our theory, we study the FDM under different transformations, and investigate whether the important inequalities and the CLTs remain valid. As in Section (ref), denote $Y_{j,n}=F_{j,n}(e_{n})\in\mathbb{R}$, $Y_{j,n,i}\equiv F_{j,n}\left(e_{1,n},\cdots,e_{i-1,n},e^{*}_{i,n},e_{i+1,n},\cdots,e_{n,n}\right)$, and $\delta_{p,n}\left(j,i,\mathcal{C}_{n}\right)\equiv\left\Vert Y_{j,n}-Y_{j,n,i}\right\Vert _{L^{p},\mathcal{C}_{n}}$ as the FDM of $\left\{ Y_{j,n}\right\} $ conditional on $\mathcal{C}_{n}$. We add the superscript $``(Y)"$ in $\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ to differentiate it from the FDM of other random variables.

We first consider Lipschitz-type functions $H_{j,n}:\mathbb{R}\rightarrow\mathbb{R}$. Suppose for all $\left(y,y^{\bullet}\right)\in\mathbb{R}\times\mathbb{R}$ and all $j$ and $n\geq1$:

equation[equation omitted — 167 chars of source]

In Propositions (ref)-(ref), we denote $Z_{j,n}\equiv H_{j,n}(Y_{j,n})$.

propSuppose $H_{j,n}\left(\cdot\right)$ satisfies Eq.((ref)) with $\sup_{n,j}\sup_{y,y^{\bullet}}B_{j,n}\left(y,y^{\bullet}\right)\leq C<\infty$ for some constant $C$. Then \[ \delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq C\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\ a.s. \]

Obviously, if the FDM of $\{Y_{j,n}\}$ satisfies the conditions in Theorems (ref), (ref), or (ref), then the same holds for the FDM of $\{Z_{j,n}\}$.

Next, we consider unbounded $B_{j,n}\left(y,y^{\bullet}\right)$, e.g., $H_{j,n}\left(y\right)=y^{2}$.

propSuppose $H_{j,n}\left(\cdot\right)$ satisfies Eq.((ref)) with $B_{j,n}\left(y,y^{\bullet}\right)\leq C_{1}(\left|y\right|^{a}+\left|y^{\bullet}\right|^{a}+1)$ for some finite constants $C_{1}>0$ and $a\geq1$, constants $p,q,r\geq1$ satisfying $p^{-1}=q^{-1}+r^{-1}$, and $\left\Vert Y\right\Vert _{L^{ar},\mathcal{C}_{n}}\equiv\sup_{n,j}\left\Vert Y_{j,n}\right\Vert _{L^{ar},\mathcal{C}_{n}}<\infty$ a.s. Then \[ \delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq C_{1}(2\left\Vert Y\right\Vert ^{a}_{L^{ar},\mathcal{C}_{n}}+1)\delta^{(Y)}_{q,n}\left(j,i,\mathcal{C}_{n}\right)\ a.s. \]

Hence, if $\frac{1}{n^{\widetilde{q}}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta^{(Y)}_{q,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\widetilde{q}}=o_{\mathbb{P}}(1)$ and $\left\Vert Y\right\Vert _{L^{ar},\mathcal{C}_{n}}<\infty$ a.s., then \[ \frac{1}{n^{\widetilde{q}}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\widetilde{q}}=o_{\mathbb{P}}(1). \] So Theorem (ref) is applicable. Similarly, condition ((ref)) in Theorem (ref) can be verified for $\{Z_{j,n}\}$. In Proposition (ref), there is a trade-off between $p$, $q$ and $r$. If a larger $p$ is desired, then larger $q$ or $r$ is required. By imposing further conditions on $\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$, we can avoid the trade-off between $p$, $q$ and $r$, which is presented in the next proposition.

propSuppose $H_{j,n}\left(\cdot\right)$ satisfies condition ((ref)) with $B_{j,n}\left(y,y^{\bullet}\right)\leq C_{1}(\left|y\right|^{a}+\left|y^{\bullet}\right|^{a}+1)$ for some finite constants $C_{1}>0$ and $a\geq1$. If $\left\Vert Y\right\Vert _{L^{q},\mathcal{C}_{n}}\equiv\sup_{n,j}\left\Vert Y_{j,n}\right\Vert _{L^{q},\mathcal{C}_{n}}<\infty$ a.s. for some $q>\max\left(\frac{ap}{p-1},ap+p\right)$. Then \[ \delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq C_{2}\left(\mathcal{C}_{n}\right)\left[\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\left(q-ap-p\right)/\left(pq-ap-p\right)}\ a.s., \] where $C_{2}\left(\mathcal{C}_{n}\right)<\infty$ a.s.

If $\frac{1}{n^{q}}\sum^{n}_{i=1}\left\{ \sum^{n}_{j=1}\left[\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\left(q-ap-p\right)/\left(pq-ap-p\right)}\right\} ^{q}=o_{\mathbb{P}}(1)$ , then, by Proposition (ref), \[ \frac{1}{n^{q}}\sum^{n}_{i=1}\left\{ \sum^{n}_{j=1}\delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right\} ^{q}=o_{\mathbb{P}}(1) \] So Theorem (ref) is applicable. Moreover, condition ((ref)) in Theorem (ref) can be verified using similar arguments.

In addition to Lipschitz transformation, another important nonlinear transformation is $1\left(y>0\right)$, which is useful in discrete choice and censor data model. See, e.g., xu2015maximum,XU201896.

propDenote $Z_{j,n}\equiv1\left(Y_{j,n}>0\right)$. Denote the density function of $Y_{j,n}$ conditional on $\mathcal{C}_{n}$ by $f_{j,n}(y\mid\mathcal{C}_{n})$. If $\sup_{n\geq1,1\leq j\leq n}\sup_{y}f_{j,n}(y\mid\mathcal{C}_{n})<C_{1}<\infty$ a.s. for some constant $C_{1}>0$ , then \[ \delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq C\left[\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{1/\left(p+1\right)}, \] for some constant $C$ not depending on $i$, $j$, nor $n$.

Moreover, we are interested in the sum or the product of two random variables. Let $\left\{ Y_{j,n}\right\} $ and $\left\{ Z_{j,n}\right\} $ be two sets of random variables, and they are both functions of some conditionally independent random vectors $e_{i,n}$ given $\mathcal{C}_{n}$, where $i\in[n]$. Denote the FDMs of $\left\{ Y_{j,n}+Z_{j,n}\right\} $ and $\left\{ Y_{j,n}Z_{j,n}\right\} $ by $\delta^{(Y+Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ and $\delta^{(YZ)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$, respectively. For summation, the conclusion is a direct result of Minkowski's inequality, and we summarize it below.

prop$\delta^{(Y+Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)+\delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)$ a.s.

The FDM of product is more complicated. For simplicity, we assume the original random fields are real-valued, since if $Y$ is functionally dependent, then all the elements of $Y$ are still functionally dependent and vice versa. Analogous to Propositions (ref) and (ref), we have two versions of conclusions, presented below.

propSuppose constants $p,q_{1},q_{2},r_{1},r_{2}>1$ satisfy $p^{-1}=q^{-1}_{1}+r^{-1}_{1}=q^{-1}_{2}+r^{-1}_{2}$, $\left\Vert Y\right\Vert _{L^{r_{2}},\mathcal{C}_{n}}\equiv\sup_{n,j}\left\Vert Y_{j,n}\right\Vert _{L^{r_{2}},\mathcal{C}_{n}}<\infty$ a.s. and $\left\Vert Z\right\Vert _{L^{r_{1}},\mathcal{C}_{n}}\equiv\sup_{n,j}\left\Vert Z_{j,n}\right\Vert _{L^{r_{1}},\mathcal{C}_{n}}<\infty$ a.s. Then \[ \delta^{(YZ)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\leq\left\Vert Z\right\Vert _{L^{r_{1}},\mathcal{C}_{n}}\delta^{(Y)}_{q_{1},n}\left(j,i,\mathcal{C}_{n}\right)+\left\Vert Y\right\Vert _{L^{r_{2}},\mathcal{C}_{n}}\delta^{(Z)}_{q_{2},n}\left(j,i,\mathcal{C}_{n}\right)\ a.s. \]

Let us examine whether $\{Y_{j,n}Z_{j,n}\}$ satisfies the conditions in Theorem (ref). Denote $\Delta^{(YZ)}_{p,2}\left(\mathcal{C}_{n}\right)\equiv\frac{1}{n^{2}}\sum^{n}_{i=1}[\sum^{n}_{j=1}\delta^{(YZ)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)]^{2}$, and $\Delta^{(Y)}_{q_{1},2}\left(\mathcal{C}_{n}\right)$ and $\Delta^{(Z)}_{q_{2},2}\left(\mathcal{C}_{n}\right)$ are defined similarly. Then by Proposition (ref),

align*[align* omitted — 937 chars of source]

where the second inequality follows from the fact that $\left(a+b\right)^{2}\leq2a^{2}+2b^{2}$ for arbitrary $a$, $b\in\mathbb{R}$. Consequently, when $\max\left\{ \Delta^{(Y)}_{q_{1},2}\left(\mathcal{C}_{n}\right),\Delta^{(Z)}_{q_{2},2}\left(\mathcal{C}_{n}\right)\right\} =o_{\mathbb{P}}(1)$, $\{Y_{j,n}Z_{j,n}\}$ also satisfies the conditions in Theorem (ref).

Similar to Proposition (ref), there is a trade-off between $p$, $q_{1},r_{1}$and $p,q_{2},r_{2}$. Since $q_{1}>p$ and $q_{2}>p$, when three or more random fields are multiplied, larger values for $q_{1}$ and $q_{2}$ are required, which may not always be feasible. However, by introducing additional conditions, it may be possible to avoid such trade-offs, as demonstrated below.

propSuppose $\left\Vert Y\right\Vert _{L^{q},\mathcal{C}_{n}}\equiv\sup_{n,j}\left\Vert Y_{j,n}\right\Vert _{L^{q},\mathcal{C}_{n}}<\infty$ a.s., $\left\Vert Z\right\Vert _{L^{q},\mathcal{C}_{n}}\equiv\sup_{n,j}\left\Vert Z_{j,n}\right\Vert _{L^{q},\mathcal{C}_{n}}<\infty$ a.s. for some $q>\max(p/(p-1),2p)$ and $p>1$. Then \[ \delta^{(YZ)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\le C_{1}\left(\mathcal{C}_{n}\right)\left[\delta^{(Y)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\frac{q-2p}{pq-2p}}+C_{2}\left(\mathcal{C}_{n}\right)\left[\delta^{(Z)}_{p,n}\left(j,i,\mathcal{C}_{n}\right)\right]^{\frac{q-2p}{pq-2p}} \] a.s., where $C_{1}\left(\mathcal{C}_{n}\right)<\infty$ a.s. and $C_{2}\left(\mathcal{C}_{n}\right)<\infty$ a.s.

Using Proposition (ref) together with the inequality $(a+b)^{r}\leq\max\{1,2^{r-1}\}(a^{r}+b^{r})$ for any $a,b,r\geq0$, we can show that if \[ \frac{1}{n^{2}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\left(\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right)^{\left(q-2p\right)/\left(pq-2p\right)}\right]^{2}=o_{\mathbb{P}}(1) \] and \[ \frac{1}{n^{2}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\left(\delta^{(Z)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right)^{\left(q-2p\right)/\left(pq-2p\right)}\right]^{2}=o_{\mathbb{P}}(1), \] then \[ \Delta^{(YZ)}_{p,2}\left(\mathcal{C}_{n}\right)=\frac{1}{n^{2}}\sum^{n}_{i=1}\left[\sum^{n}_{j=1}\delta^{(YZ)}_{p,n}\left(i,j\right)\right]^{2}=o_{\mathbb{P}}(1). \] Then, Theorem (ref) is applicable. We can verify condition ((ref)) in Theorem (ref) under similar conditions.

In practice, one may use either Proposition (ref) or (ref) to study the FDM of $\{Y_{j,n}Z_{j,n}\}$. Proposition (ref) requires a stronger moment condition, while Proposition (ref) requires that most of $\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$ are smaller, as $\frac{\left(q-2p\right)}{pq-2p}<1$ implies that $\left[\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)\right]^{\left(q-2p\right)/\left(pq-2p\right)}>\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$.\footnote{Note that most of $\delta^{(Y)}_{p,n}\left(i,j,\mathcal{C}_{n}\right)$ are smaller than 1 asymptotically. }

Conclusion

This paper develops limit theorems for random variables with network structure, without requiring individuals to be located in a Euclidean or other metric space. This sets our approach apart from most existing limit theories in network statistics and econometrics, which typically rely on weak dependence concepts such as strong mixing, near-epoch dependence, or $\psi$-dependence. To establish these results, we generalize the functional dependence measure proposed by wu2005nonlinear. Using this framework, we derive several inequalities, a law of large numbers, and central limit theorems. We also demonstrate the applicability of these results by verifying their conditions for spatial autoregressive models, which are widely used in network data analysis. Finally, we study the behavior of the functional dependence measure under various transformations commonly encountered in applications. Extending this framework to panel data with network dependence is a natural direction for future research.