Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,391 characters · 20 sections · 60 citation commands
In this paper, we consider cross-sectional dependence arising because of observations' interdependence in a network. Datasets exhibiting such forms of dependence are common in economics and other disciplines, and the results derived in this paper will allow the researcher to formally argue the consistency and asymptotic normality of estimators with network dependent data. Moreover, to facilitate inference with network dependent data, we derive conditions for the consistency of the network heteroskedasticity and autocorrelation consistent (HAC) robust variance estimator. The estimator can be used for construction of standard errors robust to general forms of network dependence.
The main results of this paper are three-fold: the Law of Large Numbers (LLN), the Central Limit Theorem (CLT), and the consistency of HAC estimators.\footnote{ Conley:99:JOE proposed a HAC estimator in a spatial random field model. See Kelejian/Prucha:07 and Kim/Sun:11:JOE for spatial HAC estimators. Leung:19:WP considers spatial and network HAC estimators in models of discrete choice with social interactions. Kojevnikov:19:WP develops bootstrap-based alternatives to network HAC estimation. } We provide a unified condition for the LLN and CLT when the network is formed in a generic way such that the links are formed independently conditional on observed or unobserved variables. This includes various network formation models proposed and used in the literature. Our condition reveals an explicit tradeoff between the extensiveness of the cross-sectional dependence and the denseness of the network permitted. The condition is also simple, as it involves only the average of conditional link formation probabilities. The paper also provides generic high level conditions that can accommodate random fields on graphs, conditional dependency graphs, and large functional-causal systems of equations.
To model network dependence, we adopt the approach of $\psi$-dependence proposed by Doukhan/Louhichi:99, and extend the notion to accommodate common shocks. The notion of $\psi$-dependence is simple and intuitive. Roughly speaking, $\psi$-dependence measures the strength of dependence between two sets of random variables in terms of the covariance between nonlinear functions of random variables.
A primary benefit of modeling through $\psi$-dependence comes when dependence among the variables is produced through a system of causal equations in which sharing of exogenous shocks creates cross-sectional dependence among the variables of interest. We give four broad classes of such examples, including those where the random variables are generated from primitive random variables through a nonlinear transform. These classes cover many sub-examples that are used in statistics and econometrics. In such examples, a traditional approach of modeling through various mixing properties is cumbersome, because it is hard to find primitive conditions that guarantee the mixing properties for the variables of interest. On the other hand, one often can write the covariance bounds of those variables in terms of the primitive exogenous shocks using the causal equations. This flexibility of the $\psi$-dependence notion, however, carries a cost. The $\psi$-dependence of a nonlinearly transformed $\psi$-dependent random variables is not necessarily ensured, if the nonlinear transform does not belong to the class in the original definition. This paper provides several auxiliary results for such situations.
Network models have been used to capture a complex form of interdependence among cross-sectional observations. These observations may represent actions by people or firms, or outcomes from industry sectors, assets or products. Random fields indexed by points in a lattice in a Euclidean space have often been adopted as a model of spatial dependence in econometrics and statistics. Conley:99:JOE proposed using random field modeling to specify the cross-sectional dependence of observations in the context of GMM estimation. More recent contributions include Jenish/Prucha:09:JOE and Jenish/Prucha:12:JOE. See Jia:08:Eca for an application in entry decisions in retail markets, and Boucher/Mourifie:17:EJ for an inference problem for a network formation model. For limit theorems for such random fields in statistics, see Comets/Janzura:98:JAP and the references therein.
When the dependence ordering arises from geographic distances or their analogues, using such random fields appears natural. However, the dependence ordering often stems from pairwise relations among the sample units, which can be viewed as a form of a network. To apply the random field modeling, one would first need to transform these relations into a random field on a lattice in a Euclidean space using methods such as multidimensional scaling.\footnote{ See, e.g., Borg/Groenen:05:MDS. See also footnote 16 of Conley:99:JOE on page 15. }
However, embedding of a network into a lattice can distort the dependence ordering. In fact, we show in Section (ref) that network dependence is not necessarily embedded as a random field indexed by a lattice in the Euclidean space with a fixed dimension, when the network has a maximum clique whose size increases as the network grows. Networks with a growing maximum clique size often arise from those with a power-law degree distribution and high clustering coefficients. These features are typically shared by social networks that are observed in practice. In this paper, we directly use a network as a model of dependence ordering, so that such an embedding is not required when dependence ordering comes from pairwise relations.
Associating dependence patterns with networks has been previously used in the literature. Stein:72:BerkeleySymp introduced a notion of dependency graphs in studying the normal approximation of a sum of random variables which are allowed to be dependent only when they are adjacent in a given network. See also Janson:88:AP, Baldi/Rinott:89:AP, Chen/Shao:04:AP, and Rinott/Rotar:96:JMA for various results for normal approximation for variables with related local dependence structures, and Aronow&Samii:17, Leung:20:ReStat, and Song:17:ReStat for recent applications of dependency graphs to network data. Modeling based on dependency graphs has drawbacks. In particular, it requires independence between variables that are not adjacent in the network, and hence is not adequate to model more extensive forms of dependence.
A closely related strand of the literature studies various models of Markov random fields and spatial autoregressive models. Markov random fields constitute an alternative class of models of dependence which imposes conditional independence restrictions based on the network structure.\footnote{See, e.g., Lauritzen:96:GraphicalModels and Pearl:09:Causality. Recently, Lee/Song:18:Bernoulli established a central limit theorem using a more general local dependence notion that encompasses both dependency graphs and a class of Markov random fields. See also Chapter 19 of Murphy:12:ML for applications in the literature of machine learning.} Spatial autoregressive models specify cross-sectional dependence through the weight matrix in linear simultaneous equations, and have been extensively studied in econometrics. See, among others, Lee:04:Eca and Lee/Liu/Lin:10:EJ and references therein. Also see Gaetan/Guyon:10:SpatialStatistics for an extensive review of spatial modeling and limit theorems.
There is a line of recent research that pursues a general form of limit theorems in a situation where the dependence structure itself is generated through a stochastic mechanism. Kuersteiner/Pucha:15:WP embed dependence along a network as a martingale model. Similarly, Kuersteiner:19:WP adopted a conditional spatial mixingale modeling of cross-sectional dependence, and established limit theorems which accommodate various network formation models. Leung/Moon:19:WPb focus on the normal approximation of network statistics when the network is formed according to a generalized version of a random geometric graph.
In contrast to the dependency graph modeling, and similarly to the recent strand of literature mentioned above, our approach permits dependence between random variables that are only indirectly linked through intermediary variables. In fact a dependency graph model can be viewed as a special case of our network dependence modeling. The approach in this paper is also distinct from Markov random fields modeling. Markov random fields are based on conditional independence restrictions among the variables. While limit theorems on Markov random fields rely on independence restrictions that come from conditioning on certain random variables, our modeling expresses the degree of stochastic dependence in terms of the distance in the network.
The rest of the paper is organized as follows. In Section (ref) of the paper, we define network dependence of stochastic processes and provide examples. In particular, Section (ref) describes a class of network formation models that our approach can accommodate. Condition NF in that section describes the restrictions one needs to impose in order to apply our results to data generated by network formation models. The condition ties link formation probabilities with network dependence patterns, and requires that dependence between nodes decays with network distance at a rate depending on the link formation probabilities.
In Section (ref), we present the main results of the paper: the LLN and CLT. Condition ND in that section provides a unifying high level assumption for the asymptotic results. It also demonstrates the tradeoffs between how fast dependence decays with the network distance and network's denseness. Lemma (ref) establishes the connection between Conditions ND and NF.
Section (ref) is devoted to deriving conditions for the consistency of HAC estimators. As with time series, the consistency of HAC estimators requires truncation of network autocovariances corresponding to large distances. The amount of truncation is determined by a bandwidth parameter. Our proposed bandwidth selection rule is given in equation (ref). While in the time series case the bandwidth parameter is typically proportional to a fractional power of the sample size, it is logarithmic in our case. More aggressive truncation (than in the time series case) is due to the fact that the time series dependence structure can be viewed as a sparse network with a fixed number of neighbors at any distance. However, in our case the number of neighbors at any distance can grow with the sample size, which can result in fast accumulation of errors in HAC estimation.
Using Monte Carlo simulations, we evaluate the finite sample performance of our HAC estimator in Section (ref). We find that HAC-based inference is accurate even in relatively dense networks. At the same, the performance of HAC-based confidence intervals can deteriorate with networks' denseness and the amount of network dependence.
The Supplemental Note to this paper contains additional proofs and simulation results.
Let $N_n=\{1,2,\ldots,n\}$ be the set of cross-sectional unit indices. Modeling cross-sectional dependence usually assumes a certain metric on $N_n$. In some examples, this distance can be motivated by geographic distances or economic distances measured in terms of economic outcomes. This paper focuses on the pattern of cross-sectional dependence that is shaped along a given network.
Suppose that we observe an undirected network $G_n$ on $N_n$, where $G_n=(N_n,E_n)$, and $E_n\subseteq \{\{i,j\}: i,j \in N_n, i \ne j\}$ denotes the set of links. For $i,j\in N_n $, we define $d_n(i,j)$ to be the distance between $i$ and $j$ in $G_n$, i.e., the length of the shortest path between nodes $i$ and $j$ given $G_n$. The distance $d_n$ defines a metric on the set $N_n$. We refer to network dependence as a stochastic dependence pattern of random variables governed by the distance $d_n$ in $G_n$.
Let $N_n(i;s)$ denote the set of the nodes that are within the distance $s$ from node $i$, and let $N_n^\partial(i;s)$ denote the set of the nodes that are exactly at the distance $s$ from node $i$. That is,
Our first focus is on the relation between modeling dependence through network topology and that through random fields indexed by the elements of a finite subset of a metric space $(\mathcal{X},d_{\mathcal{X}})$. We denote the equilateral dimension of $\mathcal{X}$, i.e., the maximum number of equidistant points in $\mathcal{X}$ with respect to the distance $d_{\mathcal{X}}$, as $e(\mathcal{X})$. The main question here is whether any given connected network is embeddable in $\mathcal{X}$.\footnote{ A network/graph is connected if there is a path between every pair of nodes. } The following definition makes the notion of embedding precise.
When such an isometry exists, it means that modeling cross-sectional dependence using a network topology can be viewed as a special case of modeling a random field on a finite subset of $\mathcal{X}$. The following result shows that this is not always possible when the clique number $\omega(G_n)$ of $G_n$, i.e., the number of nodes in a maximum clique in $G_n$, is large enough.\footnote{ A clique of a graph $G$ is a subset of nodes such that every two distinct nodes are adjacent. }
Proposition (ref) gives only a necessary condition for isometric embedding. Consider, for example, $\R^k$ equipped with the Euclidean distance, which has the equilateral dimension of $k+1$. Figure (ref) provides an example of a network with the maximum clique size of two that cannot be embedded into the Euclidean $\R^2$ space, which has the equilateral distance of three. Figure (ref) provides an example with a non-Euclidean space. It shows a network with the maximal clique size of four that cannot be imbedded into $\R^2$ equipped with the $L_\infty$ distance, which has the equilateral dimension of four.
An important consequence of Proposition (ref) is that when the size of the maximum cliques in the network $G_n$ grows to infinity as $n \rightarrow \infty$, the sequence of networks cannot be embedded into a metric space having a finite equilateral dimension. Examples of such spaces include a $k$-dimensional normed space $M^k$ and a sphere $\mathbb{S}^{k}$ equipped with the usual distance because $e(M^k)\le 2^{k}$ and $e(\mathbb{S}^k)= k+2$ Petty:71. As a consequence, the random field models used in Conley:99:JOE with the Euclidean distance and in Jenish/Prucha:09:JOE with the Chebychev distance cannot include a network dependence model when the maximum clique size of the networks increases with the sample size. Indeed, there are random graphs whose degree distribution takes the form of a power law and the size of the maximum cliques grows to infinity as $n \to \infty$ Blasius/Friedrich/Krohmer:17:Algorithmica. Such models accommodate both dense and sparse graphs, and are often motivated as a model of many real networks that we observe in practice.
The asymptotic results developed in this paper can accommodate network generating processes with the maximum clique size increasing with the sample size. However, our results impose certain restrictions on the rate of growth of the maximum clique size.
One may consider “approximating" the network dependence ordering by a lattice in a finite dimensional Euclidean space. Multidimensional scaling (MDS) provides various ways to achieve such an approximation Borg/Groenen:05:MDS. The dependence ordering obtained through MDS is itself dependent on the data, and is stochastic. Hence, it is generally different from the true dependence ordering of the data. Proposition (ref) tells us that there is no guarantee that the approximation error of the MDS-based dependence ordering will be small with a large sample size.
Suppose that we are given a triangular array of $\R^v$-valued random vectors, $Y_{n,i}, i \in N_n$, which are laid on a network $G_n$ whose agency matrix we denote by $A_n$. That is, the $(i,j)$-th entry $A_{n,ij}$ of matrix $A_n$ is one if $i$ and $j$ are adjacent in $G_n$ and zero otherwise. The $(i,i)$-th entry of $A_n$ is zero for $i\in N_n$. We adapt the $\psi$-dependence notion of Doukhan/Louhichi:99 to our setup. We define $\N = \{1,2,3,\ldots\}$, and for any $v,a \in \N$, we endow $\R^{v \times a}$ with the distance
where $\vec{x}=(x_1,\ldots,x_a)$ and $\vec{y}=(y_1,\ldots,y_a)$ are points in $\R^{v\times a}$, and $\| \cdot \|$ denotes the Euclidean norm, i.e., $\|w\| = \sqrt{w^{\top}w}$, for $w \in \R^a$. Let
where $\L_{v,a}$ denotes the collection of bounded Lipschitz real functions on $\R^{v \times a}$, i.e.,
with $\Lip(f)$ denoting the Lipschitz constant of $f$,\footnote{ The Lipschitz constant for a function $f:\R^{v \times a} \rightarrow \R$ is the smallest constant $C$ such that $|f(\vec{x}) - f(\vec{y})| \le C \mathss{d}_a(\vec{x},\vec{y})$, for all $\vec{x},\vec{y} \in \R^{v \times a}$. } and $\|\cdot\|_\infty$ the sup-norm of $f$, i.e., $\|f\|_\infty = \sup_x|f(x)|$. For any positive integers $a,b,s$, consider two sets of nodes (of size $a$ and $b$) with distance between each other of at least $s$. Let $\P_n(a,b;s)$ denote the collection of all such pairs:
where
and $d_n(i,i')$ denotes the distance between nodes $i$ and $i'$ in $G_n$, i.e., the length of the shortest path between $i$ and $i'$ in $G_n$. For each set $A$ of positive integers, we write
We take $\{\C_n\}_{n \ge 1}$ to be a given sequence of $\sigma$-fields such that for each $n \ge 1$, the adjacency matrix $A_n$ of graph $G_n$ is $\C_n$-measurable. Below we introduce a notion of conditional $\psi$-dependence for a triangular array $\{Y_{n,i}\}_{i \in N_n}, n \ge 1, Y_{n,i} \in \R^v$. From here on, we write triangular arrays simply as $\{Y_{n,i}\}$, and sequences $\{\C_n\}_{n \ge 1}$ as $\{\C_n\}$.
In a typical set-up that we consider in this paper, $\{\theta_{n,s}\}$ approaches zero as $s$ grows. The $\sigma$-field $\C_n$ can be thought of as a “common shock" such that when we condition on it, the cross-sectional dependence of triangular array $\{Y_{n,i}\}$ becomes substantially weaker. However, we do not have to think of $\C_n$ as being originated from a variable that affects every node in the network. In many network set-ups, $\C_n$ can be thought of as having been generated by some characteristics or actions of multiple central nodes which affect many other nodes through their many links. For example, consider a star network, where node $1$ is adjacent to the other $n-1$ nodes. Suppose that $Y_{n,1}=U_1$ corresponds to the central node, and for the remaining nodes ($i\geq 2$), \[ Y_{n,i}=U_1+U_i, \] where $\{U_i:i=1,\ldots,n\}$ are independent. In that case, we can take $\C_n=\sigma(U_1)$. Then, conditionally on $\C_n$, $Y_{n,2},\ldots, Y_{n,n}$ are i.i.d., and $\PM\{\theta_{n,2}=0\mid \C_n\}=1$.
Unlike the unconditional version of $\psi$-dependence of Doukhan/Louhichi:99, in our definition the dependence coefficients $\{\theta_n\}$ are random, due to our accommodation of the common shocks, $\C_n$. We make the following assumption.
Assumption (ref) will be maintained throughout the paper. It is shown to be satisfied by all the examples we present in the next subsection. The following lemma shows that $\psi$-dependence of random vectors carries over to linear combinations of their elements.
A result similar to Lemma (ref) holds for nonlinear transforms of random variables, under certain conditions for the nonlinear transforms. See Appendix (ref) for details.
Limit theorems in this paper focus on the asymptotic behavior of the following sum
where $\C_n$ is a certain $\sigma$-field. The asymptotic behavior of such a sum often arises in the network-based interactions models. For example, let us consider the following linear interaction model: \[ y_i = \beta \overline{y}_i + \gamma' X_i + v_i, \] where
and $N_n^\partial(i;1)$ is defined in (ref). Such a model has been widely studied and used in the literature on social interactions Blume/Brock/Durlauf/Jayaraman:15:JPE. This literature usually assumes that the error term $v_i$ in the outcome equation is uncorrelated with the network $G_n$. In other words, the network $G_n$ is exogenously formed. A recent paper by Johnsson/Moon:19:ReStatForthComing extends the framework to accommodate a situation where $G_n$ is endogenously formed, by introducing an explicit yet generic network formation model and proposing a control function approach. Their network formation model is given as follows: \[ j \in N_n^\partial(i;1) \qtext{if and only if}\quad f_{ij}(X,T_i,T_j) \ge u_{ij}, \] where $f_{ij}$ is a nonstochastic map, $X = (X_i)_{i \in N_n}$, $T_i$'s are unobserved individual heterogeneity affecting the network formation process, and $u_{ij}$'s are link-specific error terms. Johnsson/Moon:19:ReStatForthComing introduce a set of assumptions which imply the following conditions:
where $\C_n$ is the $\sigma$-field generated by $(X,T,u)$, $X = (X_i)_{i \in N_n}$, $T = (T_i)_{i \in N_n}$, and $u = (u_{ij})_{i,j \in N_n}$. Then the normal approximation of the distribution of estimators for $\beta$ stems from the limit distribution of the sum of the form: \[ \sum_{i \in N_n} (v_i - \E[v_i\mid T_i])\varphi_i, \] where $\varphi_i$ is a random variable that is constructed as a function of $X$ and $G_n$ (such as instrumental variables). This sum becomes (ref) if we take $Y_{n,i} = (v_i - \E[v_i\mid T_i])\varphi_i$, and by Condition B, we have $\E[Y_{n,i}\mid \C_n] = 0$. Our limit theorems in this paper can be used to relax the conditional i.i.d. assumption in Condition A to accommodate the case where the terms $v_i - \E[v_i\mid T_i]$ exhibit network dependence along $G_n$, for instance, through a data generating process as in one of the examples in Section 2.4.
In contrast to time series dependence, limit theorems for network dependent processes depend not only on the strength of the dependence but also on the shape of the network itself. In this section, we consider a class of network formation models and give a sufficient condition for the network formation process. First, consider a generic network $G_n=(N_n,E_n)$ for which the link between each pair of nodes $\{i,j\}$, $i\ne j$, is realized randomly as follows: $i$ and $j$ are linked if and only if $A_{n,ij} = 1$, where
$\varphi_{n,ij}$'s and $\varepsilon_{ij}$'s are random variables such that $\varphi_{n,ij}=\varphi_{n,ji}$, $\varepsilon_{ij}=\varepsilon_{ji}$, and $\{\varepsilon_{ij}:i<j\}$ are i.i.d. and independent of $\varphi_n=(\varphi_{n,ij})_{i<j}$. This random graph model can be viewed as a generalization of the Erd\"{o}s--R\'{e}nyi graph model in the sense that conditional on $\varphi_n$, the link formation probabilities can be heterogeneous across all pairs of nodes. Many network formation models used in the literature take this form, where \[ \varphi_{n,ij} = f_n(W_{ij},T_i,T_j), \] for some function $f_n$, $W_{ij}$ is observable, and $T_i$ is an unobservable node-specific component. For example, Graham:17 specified $f_n$ as follows: \[ f_n(W_{ij},T_i,T_j) = W_{ij}'\beta_0 + T_i + T_j. \] Ridder/Sheng:19:WP considered an endogenous network formation model where the payoff depends not only on the neighbors' characteristics and the characteristics of their 2-neighbors. They find that the model yields the following “reduced-form" for the formation of the network with \[ f_n(W_{ij},T_i,T_j) = V_{n,ij}(W_{ij}), \] where we take $W_{ij}$ to be the covariates $(X_i)_{i \in N_n}$ and $V_{n,ij}$ is a nonstochastic map. Leung:19:JOE studied a network model which is contained in a graph generated through \[ f_n(W_{ij},T_i,T_j) = \sup_s V(r_n^{-1}\|X_i - X_j\|,s,W_{ij}), \] where $(X_i,X_j)$ is a part of the vector $W_{ij}$, $V$ is a map, and $r_n = (\kappa/n)^{1/d}$, with $d$ representing the dimension of $X_i$ and $\kappa$ is a positive constant. This latter graph can be viewed as a generalized version of a random geometric graph.
Suppose that $\{Y_{n,i}\}$ is conditionally $\psi$-dependent given $\{\C_n\}$ with the dependence coefficients $\{\theta_{n}\}$, such that
for some $p > 4$, and the functional $\psi_{a,b}$ satisfies Assumption (ref)(a). Our main interest is in the LLN and CLT of the following form:
where $\sigma_n^2 = \Var(\sum_{i \in N_n} Y_{n,i}\mid \C_n)$. The sparsity of the graph $G_n$ necessary for the limit theorems in this paper is summarized by the asymptotic behavior of the maximal expected degree $\pi_n$, where
In this paper we show that these limit theorems hold if the following condition holds.
Condition NF. There exist $\varepsilon>0$ and $q > \max\{p/(p-4),3p/(p-1)\}$ for $p >4$ in (ref) and a positive random variable $M$ such that for all $n \ge 1$, \[ \theta_{n,s} \le M((\pi_n \vee 1) + \varepsilon)^{-qs}, \quad \quad 1 \le s \le n, \] holds eventually with probability one.\footnote{ A sequence of events $\{E_n\}_{n\ge 1}$ holds eventually with probability one if $\PM(\bigcup_{n\ge 1}\bigcap_{m\ge n} E_m)=1$. }
For example, Condition NF is satisfied if there exists $\gamma \in (0,1)$ such that $\theta_{n,s} \le \gamma^s, 1 \le s \le n$, and $\gamma < ((\pi_n \vee 1) + \varepsilon)^{-q}$ for some $q > \max\{p/(p-4),3p/(p-1)\}$.
In this section, we consider four broad classes of examples of conditionally $\psi$-dependent random vectors.
Let $(\Omega,\mathcal{F},\PM)$ be an underlying probability space. For sub $\sigma$-fields $\mathcal{G}$, $\mathcal{H}$, $\C$ of $\mathcal{F}$, let \[ \alpha(\mathcal{G},\mathcal{H}\mid \C)\eqdef\sup_{G\in\mathcal{G},H\in \mathcal{H}}\abs{\Cov(\ind_G,\ind_H\mid \C)}. \] For a triangular array $\{Y_{n,i}\}$ and a sequence of $\sigma$-fields $\{\C_n\}$ we define the strong mixing coefficients by\footnote{ These coefficients are different from those given in Jenish/Prucha:09:JOE because our $\theta_n$ coefficients do not depend on $\abs{A}$ and $\abs{B}$. }
The proposition below provides a conditional covariance inequality that is due to Theorem 9 of PrakasaRao:13.
Hence, the array $\{Y_{n,i}\}$ is conditionally $\psi$-dependent given $\{\C_n\}$ with $\psi_{a,b}(f,g)=4\norm{f}_{\infty}\norm{g}_{\infty}$, and the dependence coefficients $\{\theta_{n,s}\}_{s \ge 1}$ are given by the strong mixing coefficients $\{\alpha_{n,s}\}_{s\ge 1}$.
The proof of Proposition (ref) follows by adapting the proof of Theorem A.5. of HallHeyde:80:MLT to the conditional settings and noticing that the strong mixing coefficients can be equivalently defined by replacing $\alpha(\mathcal{G},\mathcal{H}\mid \C)$ with $\alpha(\mathcal{G}\vee \C,\mathcal{H}\vee \C\mid \C)$.
Suppose that $\{Y_{n,i}\}$ is a given collection of random vectors and $G_n = (N_n,E_n)$ is a graph on the index set $N_n$. Let $\C_n$ be a given $\sigma$-field. We say that $\{Y_{n,i}\}$ has $G_n$ as a conditional dependency graph given $\C_n$, if for any set $A \subset N_n$, $Y_{n,A}$ and $\{Y_{n,i}: i \in N_n\setminus N_n(A)\}$ are conditionally independent given $\C_n$, where $N_n(A) = \bigcup_{i\in A}N_n(i;1)$. The notion of a conditional dependency graph is a conditional variant of a dependency graph introduced by Stein:72:BerkeleySymp. It is not hard to see that when $\{Y_{n,i}\}_{i\in N_n}$ has $G_n$ as a conditional dependency graph given $\C_n$ for each $n\ge 1$, the array $\{Y_{n,i}\}$ is conditionally $\psi$-dependent given $\{\C_n\}$ with \[ \psi_{a,b}(f,g) = 4 \|f\|_\infty \|g\|_\infty, \] and $\theta_n$ is such that $\theta_{n,s} = 0$ for all $s \ge 1$.
Consider a triangular array of $\R^k$-valued random vectors $\{\varepsilon_{n,i}\}_{i \in N_n}$ which is row-wise independent given $\C_n$. For $\R^v$-valued measurable functions $\{\bm\phi_{n,i}\}_{i \in N_n}$, let \[ Y_{n,i}\eqdef\bm\phi_{n,i}(\varepsilon_n), \quad i\in N_n, \] where $\varepsilon_n = (\varepsilon_{n,j} : j\in N_n)$. Further, define a modified version of $Y_{n,i}$, which replaces too distant shocks $\varepsilon_{n,j}$ with zeros: \[ Y_{n,i}^{(s)}\eqdef\bm\phi_{n,i}\big(\varepsilon_n^{(s)}\big), \] where $\varepsilon_n^{(s,i)} = (\varepsilon_{n,j}\ind\{j\in N_n(i;s)\} : j\in N_n)$ and $N_n(i;s)$ is defined in (ref).\footnote{ Zero can be replaced with another constant if the functions $\bm\phi_{n,i}$ are undefined at zero. } Now, for any $A,B\subset N_n$ with $d_n(A,B)>2s$, $Y_{n,A}^{(s)}$ and $Y_{n,B}^{(s)}$ are conditionally independent given $\C_n$.
It follows from Proposition (ref) that $\{Y_{n,i}\}$ is conditionally $\psi$-dependent given $\{\C_n\}$, where the $\psi$ function is given by \[ \psi_{a,b}(f,g) = a\norm{g}_{\infty}\Lip(f)+b\norm{f}_{\infty}\Lip(g). \] This functional $\psi_{a,b}(f,g)$ satisfies Assumption (ref)(a).
Proposition (ref) can be extended to the case where $\varepsilon_n = (\varepsilon_{n,j}: j \in N_n)$ is $\psi$-dependent, as shown below. Let $\phi_{n,ir}$ denote the $r$-th component of $\bm\phi_{n,i}$, and we endow the domain $\R^{k \times n}$ of $\phi_{n,ir}$ with the norm $\mathss{d}_n$ defined in (ref) with $a = n$, so that $\Lip(\phi_{n,ir})$ represents the Lipschitz constant of $\phi_{n,ir}$ with respect to $\mathss{d}_n$.
As a concrete example, consider a simple linear case in which \[ Y_{n,i}=\sum_{m\ge 0}\gamma_{m,n}\sum_{j\in N_n^{\partial}(i;m)}\varepsilon_{n,j}, \] where $\varepsilon_{n,j}$'s are $\psi$-dependent as in Proposition (ref) such that $D_{n}^2(s) \theta_{n,s}^\varepsilon \rightarrow_{a.s.} 0$, and $N_n^\partial(i;m)$ is as defined in (ref). Let us take $\|\cdot \|$ to be the Euclidean norm. Since \[ \normin{Y_{n,i}-Y_{n,i}^{(s)}}\le \sum_{m>s}\abs{\gamma_{m,n}}\sum_{j\in N_n^{\partial}(i;m)}\normin{\varepsilon_{n,j}}, \] setting $\alpha_n\eqdef \max_{i\in N_n}\E[\norm{\varepsilon_{n,i}}\mid \C_n]$, we find that
As for the functional $\psi_{a,b}$ in the $\psi$-dependence of $\{Y_{n,i}\}_{i \in N_n}$, observe that for any $x, \tilde x \in \R^{n \times k}$ (with their $(j,k)$-th entries denoted by $x_{j,r}$ and $\tilde x_{j,r}$),
where $x_j = (x_{j,r})_{r=1}^k$ and $\tilde x_j = (\tilde x_{j,r})_{r=1}^k$. Hence, \[ \sum_{r=1}^k \Lip(\phi_{ni,r}) =\sqrt{k} \sum_{m \ge 0}|\gamma_{m,n}|. \] Thus, we can find the functional $\psi$ in the $\psi$-dependence of $\{Y_{n,i}\}$ as $\psi_{a,b}$ in (ref) with \[ \bar \phi = \sqrt{k} \sup_{n \ge 1} \sum_{m \ge 0}|\gamma_{m,n}|. \] The functional $\psi_{a,b}$ in (ref) satisfies Assumption (ref)(a) if $\bar \phi < \infty$ in this model.
Let us consider the following process: \[ Y_{n,i}= \varphi_{n,i}(\varepsilon_n), \] where $\varepsilon_n = (\varepsilon_{n,i} : i\in N_n)$ is a positively associated process, $\varepsilon_{n,i} \in \R$, conditional on certain $\sigma$-field $\C_n$, i.e., for all coordinatewise non-decreasing real-valued measurable functions $f$ and $g$ and all finite subsets $A$ and $B$ of $N_n$, \[ \Cov(f(\varepsilon_{n,A}), g(\varepsilon_{n,B})\mid \C_n) \ge 0 \qtext{a.s.} \] When the above inequality is reversed for all finite subsets $A$ and $B$ of $N_n$, we say that $\{\varepsilon_{n,i}\}_{i \in N_n}$ is negatively associated. When a set of random variables is positively or negatively associated, independence between two random variables in the set is equivalent to their being uncorrelated. The following result follows as a consequence of a covariance inequality due to Theorem 3.1 of Birkel:88:AP and Lemma 19 of Doukhan/Louhichi:99.
The above proposition clearly shows that the dependence structure of $\{Y_{n,i}\}$ is determined by the (conditional) local dependence structure of $\varepsilon_{n,i}$'s and $\varphi_{n,i}$'s. In the special case where $\varepsilon_{n,i}$'s are all conditionally independent given $\C_n$, the sequence $\theta_{n,s}$ is reduced to the following: \[ \theta_{n,s} = \max_{(A',B') \in \mathcal{P}_n(a,b;s)} \max_{k_1 \in A',k_2 \in B'} \left\|\frac{\partial \varphi_{n,k_1}}{\partial \varepsilon_{n,i}}\right\|_{\infty} \left\|\frac{\partial \varphi_{n,k_2}}{\partial \varepsilon_{n,i}}\right\|_{\infty}. \] Suppose further that $N_n$ is endowed with a graph $G_n$ such that $\partial \varphi_{n,k}/\partial \varepsilon_{n,i} = 0$, whenever $i$ is at least $m$-edges away from $k$ in $G_n$. Then $Y_{n,i}$'s have a graph $G_n'$ as a conditional dependency graph given $\C_n$, where $i$ and $j$ are adjacent in $G_n'$ if and only if $i$ and $j$ are within $2m$ edges away. Hence, it follows that $\theta_{n,s} = 0$, for all $s \ge 2m$.
The following corollary shows that the array $\{Y_{n,i}\}$ is conditionally $\psi$-dependent given $\{\C_n\}$.
In this section, we provide a sufficient condition for the shape of the network that ensures our limit theorems (i.e., the LLN and CLT) hold. The crucial aspect of the network which matters for the limit theorem is the properties of the neighborhood shells. For the limit theorems to hold, the number of the neighbors at distance $s$ should not grow too fast as $s$ increases. The precise condition for such neighborhood shells depends on the dependence coefficients $\theta_{n,s}$, so that if $\theta_{n,s}$ decreases fast as $s$ increases, the requirement for the neighborhood shells can be weakened.
To introduce sufficient conditions, let
where $N_n^\partial(i;s)$ is defined in (ref). When $k=1$, we simply write $\delta_n^\partial (s;1) = \delta_n^\partial(s)$. This quantity measures the denseness of a network. Let us introduce further notation. Define
where $N_n(i;s)$ is defined in (ref), and we take $N_n(j;s-1) = \varnothing$ if $s = 0$. We also define
The quantity $c_n(s,m;k)$ is easy to compute when a network is given, and it captures the network properties that are relevant for the limit theorems. It consists of two components: $\Delta_n(s,m;k\alpha)$ and $\delta_n^\partial(s;\alpha/(\alpha-1))$. They capture the denseness of the network through the average neighborhood sizes and the average neighborhood shell size. We summarize a sufficient condition for the network and the weak dependence coefficient as follows.
Condition ND. There exist $p>4$ and a sequence $m_n \to \infty$ such that
Later we show that Condition ND is sufficient for the LLN and CLT in (ref). Note that $\Delta_n(s,m_n;k)$ tends to decrease fast to zero as $s$ goes beyond a certain level, because the set $N_n(j;s-1)$ quickly becomes large.
The following lemma shows that in the case of the network formation model in (ref), Condition NF implies Conditions ND(a) and (b).
Thus, Condition NF is a sufficient condition on the network formation model for the LLN and CLT in (ref) for $\{Y_{n,i}\}$. The proof of Lemma (ref) is found in the Supplemental Note to this paper. The proof is built on a bound on the tail probability of $\delta_n^\partial(s;k)$. This bound is obtained using similar arguments in Chung/Lu:01:AAM for the case of Erd\"{o}s--R\'{e}nyi graphs.
Let $\{Y_{n,i}\}$ be conditionally $\psi$\hypdependent given $\{\C_n\}$. Since a LLN can be applied element-by-element in the vector case, without loss of generality we can assume that $Y_{n,i}\in\R$ in this section, i.e., $v=1$.
Let $\norm{Y_{n,i}}_{\C_n,p}=(\E[ \vert Y_{n,i} \vert^p \mid \C_n])^{1/p}$. We assume the following moment condition.
The next assumption puts a restriction on the denseness of the network and the rate of decay of dependence with the network distance.
Note that the above assumption is implied by Condition ND(b) for $k=2$: $\theta_{n,s}\leq \theta_{n,s}^{1-4/p}$ for $p>4$, and $ \delta_n^\partial(s) \leq 4c_n(s,m;2)$.\footnote{The second inequality follows from (ref) and (ref) in the supplement.}
Assumption (ref) can fail, for example, if there is a node connected to almost every other node in the network as in the following example. Consider a network with the star topology, which has a central node or hub connected to every other node. In this case, the distance between any two nodes does not exceed 2: $\delta_n^\partial(1)=2(n-1)/n$, $\delta_n^\partial(2)=(n-2)(n-1)/n$, and $\delta_n^\partial(s)=0$ for $s\geq 3$. Hence, Assumption (ref) fails for a star network, unless $\theta_{n,2}=0$, i.e., unless there is no network dependence at the distance $s>1$.
Alternatively, consider a network with the ring topology, where nodes are connected in a circular fashion to form a loop, see Figure (ref)(A) in (ref). In that case, $\delta_n^\partial(s)\leq 2$, and Assumption (ref) holds when $n^{-1}\sum_{s\ge 1}\theta_{n,s}\to_{a.s.} 0$.
The following theorem establishes a conditional LLN.
An unconditional version of the result, which replaces the conditional norm in Theorem (ref) with the unconditional norm, can be established in a similar manner by replacing the conditional moment in Assumption (ref) with the unconditional moment.
Next, we discuss LLNs for nonlinear functions of $\{Y_{n,i}\}$. When $f\in \L_{v,1}$, a LLN for a nonlinear transformation $f(Y_{n,i})$ follows immediately from the definition of the $\psi$-dependence in Definition (ref).\footnote{ Note that compositions of bounded Lipschitz functions are also bounded and Lipschitz; hence, $f(Y_{n,i})$ is also $\psi$-dependent when $f\in\L_v$. } In that case, \[ \norm{\frac{1}{n}\sum_{i \in N_n} \left(f(Y_{n,i})-\E[f(Y_{n,i}) \mid \C_n ]\right)}_{\C_n,2}^2 \leq \frac{2}{n} \norm{ f}_{\infty}^2 +\psi_{1,1}(f,f) \frac{1}{n} \sum_{s\ge 1}\delta^\partial_n(s)\theta_{n,s}. \] We have the following result.
However, in general nonlinear transformations of $\psi$-dependent processes are not necessarily $\psi$-dependent. In such cases, LLNs for nonlinear transformations can be established using the covariance inequalities for transformation functions presented in Appendix (ref) in this paper. For example, suppose that the assumptions of Corollary (ref) in the appendix hold for some nonlinear function $h(\cdot)$ of a $\psi$-dependent process $\{Y_{n,i}\}$, and that $\theta_{n,s}$ is bounded by a constant uniformly over $s \ge 1$ and $n \ge 1$. In that case for some constants $C>0$ and $p>2$, the conditional covariance given $\C_n$ between $h(Y_{n,i})-\E[h(Y_{n,i}) \mid \C_n]$ and $h(Y_{n,j})-\E[h(Y_{n,j}) \mid \C_n]$ is bounded by \[ C\cdot \sup_{n,i}\norm{h(Y_{n,i})}_{\C_n,p}^2 \cdot \theta_{n,d_n(i,j)}^{1-\frac{2}{p}}. \] Therefore, as $n\to\infty$, \[ \norm{\frac{1}{n}\sum_{i \in N_n}(h(Y_{n,i})-\E[h(Y_{n,i}) \mid \C_n])}_{\C_n,2} \to_{a.s.} 0, \] provided that $\sup_{n,i}\norm{h(Y_{n,i})}_{\C_n,p}<\infty$ a.s., and a condition similar to that in Assumption (ref) holds: \[ \frac{1}{n} \sum_{s\ge 1} \delta_n^\partial(s) \theta_{n,s}^{1-\frac{2}{p}}=o_{a.s.}(1). \] Cases not covered by Corollary (ref) can be handled in a similar manner using the covariance inequality of Theorem (ref) in Appendix (ref). We use such a strategy to show the consistency of the HAC estimator in Section (ref).
In this section, we study the CLT for a sum of random variables that are conditionally $\psi$-dependent. Define
where $S_n = \sum_{i \in N_n} Y_{n,i}$. The assumption below presents a moment condition.
While the moment condition in Assumption (ref) is more restrictive than those conditions known for the CLT for special cases of $\psi$-dependence, such a moment condition is widely used in many models in practice. The following assumption limits the extent of the cross-sectional dependence of the random variables through restrictions on the network.
It is not hard to see that Condition ND is a sufficient condition for this assumption, when $\sigma_n \ge c \sqrt{n}$ with probability one, for some constant $c>0$ that does not depend on $n$. The latter condition is satisfied if the “long-run variance", $\Var(S_n\mid \C_n)/n$ is bounded away from $c^2>0$ for all $n \ge 1$.
The theorem below establishes the CLT for the normalized sum $S_n/\sigma_n$.
The proof of the CLT uses Stein's Lemma Stein:86. The CLT immediately gives a stable convergence of a normalized sum of random variables under appropriate conditions. More specifically, suppose that \[ \sigma_n^2/(nv^2) \rightarrow_{a.s.} 1, \] where $v^2$ is a random variable that is $\C$-measurable and $\C$ is a sub $\sigma$-field of $\C_n$ for all $n \ge 1$. Then it follows that $S_n/\sqrt{n}$ converges stably to a mixture normal random variable.
In this section, we develop network HAC estimation of the conditional variance of $S_n/\sqrt{n}$ given $\C_n$, where $S_n\eqdef\sum_{i\in N_n}Y_{n,i}$. First, we assume that $\E[Y_{n,i}\mid \C_n]=0$ a.s.\ for all $i\in N_n$. Let
Then the conditional variance of $S_n/\sqrt{n}$ given $\C_n$ is given by
Similarly to the time-series case, the asymptotic consistency of an estimator of $V_n$ requires a restriction on weights given to the estimated “autocovariance" terms $\Omega_n(\csdot)$. Consider a kernel function $\omega:\bar{\R}\to [-1,1]$ such that $\omega(0)=1$, $\omega(z)=0$ for $\abs{z}>1$, and $\omega(z)=\omega(-z)$ for all $z \in \bar{\R}$.
Let $b_n$ denote the bandwidth or the lag truncation parameter. Then the kernel HAC estimator of $V_n$ is given by
where $\omega_n(s)\eqdef\omega(s/b_n)$, and
The weight given for each sample covariance term $\tilde\Omega_n(s)$ is a function of distance $s$ implied by the structure of a network. Also notice that if nodes $i$ and $j$ are disconnected then $d_n(i,j)=\infty$ so that $\omega_n(d_n(i,j))=0$.
Unlike the time series case, the number of terms included in the double sum in (ref) depends on the shape of the network. Hence, if there are many empty neighborhood shells, a large value of the bandwidth can still produce a HAC estimator that performs well in finite samples.
Next, assume that $\E[Y_{n,i}\mid \C_n]=\Lambda_n$ a.s.\ for all $i\in N_n$ and the sequence of common conditional expectations $\{\Lambda_n\}$ is unknown.\footnote{If random vectors $\{Y_{n,i}\}_{i\in N_n}$ do not share a common expectation, it is hard to justify plugging the sample mean into $\hat{\Omega}_n(\csdot)$ because $\bar{Y}_n$ is not a consistent estimator of $\E[Y_{n,i}\mid \C_n]$.} By Theorem (ref), $\bar Y_n\eqdef S_n/n$ is a consistent estimator of $\Lambda_n$ in the sense that $\E[\normin{\bar{Y}_n-\Lambda_n}\mid \C_n]\to_{a.s.}0$. We redefine the kernel HAC estimator given in (ref) as follows:
where
We establish the consistency of the estimators (ref) and (ref) by imposing suitable conditions on the moments of the array $\{Y_{n,i}\}$, the denseness of a sequence of networks, and the rate of growth of the bandwidth parameter.
The assumption demonstrates the tradeoff between the conditional moments of $\{Y_{n,i}\}$ given $\{\C_n\}$ and the magnitude of the network dependence. For a given sequence of networks, a stronger network dependence requires the finiteness of higher conditional moments, i.e., a larger value of $p$. On the other hand, sparse networks allow for either weaker moments conditions or a stronger dependence along the network. Note that Assumptions (ref)(i) and (iii) are implied by Condition ND.
Assumption (ref)(ii) is a high-level condition, which requires that the kernel weights $\omega_n(s)$ converge to one sufficiently fast as $n\to\infty$. Proposition (ref) below provides primitive conditions for Assumption (ref)(ii) in the case of models satisfying Condition NF.
Assumption (ref)(iii) determines the admissible rate of growth of the sequence of bandwidths $\{b_n\}$. In particular, it strongly depends on the network topology. In case of models satisfying Condition NF, the following bandwidth selection rule is motivated by equation (ref) in the proof of Lemma (ref) in the Supplemental Note:
where “avg.deg" is the average degree $\delta_n^\partial(1)$ of the observed network and used to approximate $\pi_n$ in Condition NF. For example, in the case of the Parzen kernel, we found through extensive Monte Carlo simulations that setting the constant in (ref) to $2.0$ and $\varepsilon = 0.05$ works well, see Section (ref) for the details.
We define
Note that $\delta_n(b_n)$ measures the denseness of a network in terms of the average size of $b_n$-neighborhoods. Let $\norm{\csdot}_F$ denote the Frobenius norm.\footnote{ For a real matrix $A$, $\norm{A}_F\eqdef \sqrt{\trace(A^{\top}A)}$. }
In the second part of the proposition, the non-increasing in $s$ condition for $\{\theta_{n,s}/s^{p/(p-4)}\}$ is a mild requirement consistent with the notion of weak dependence. For example, in the linear model below Proposition (ref) with conditionally independent $\varepsilon_{n,i}$'s, we can take $\theta_{n,s}$ to be the bound on the right hand side of (ref) to satisfy this monotonicity condition.
Next, we provide primitive conditions for Assumption (ref)(ii) in the case of networks satisfying Condition NF.
The bandwidth condition in Proposition (ref) is consistent with the bandwidth selection rule in (ref). The condition in (ref) is satisfied by many commonly used kernels such as the truncated kernel $\omega(x)=1\{\abs{x}\leq 1\}$, Parzen, and Tukey--Hanning kernels Andrews:91. However, (ref) does not hold for the Bartlett kernel.
While according to Proposition (ref) the proposed HAC estimators are consistent, they are not necessarily positive semidefinite. The following example provides a simple case in which positive definiteness of the kernel function does not automatically imply positive semidefiniteness of the estimated covariance matrix.
In the context of spatial models, Conley/Molinary:07:JOE and Kim/Sun:11:JOE show that the network HAC estimator can be consistent despite measurement errors in locations. Below, we show that our network HAC estimators have a similar property when the network is only partially observed.
Suppose that the true network is given by $G_n=(N_n,E_n)$. However, the econometrician observes $G^*_n=(N_n,E^*_n)$, where the observed set of links $E^*_n$ is a subset of the true set of links $E_n$. Thus, the links are only partially observed by the econometrician. We continue to use $d_n(i,j)$ to denote the distance between $i$ and $j$ in $G_n$. Let $d^*_n(i,j)$ denote the distance between $i$ and $j$ in $G^*_n$. The immediate consequence of $E^*_n\subset E_n$ is that $d^*_n(i,j)\geq d_n(i,j)$. Hence, because some of the links are unobserved, the true network can be denser than the observed one. The HAC estimator is now defined similarly to (ref), however, with $\tilde \Omega_n(s)$ replaced by $\tilde \Omega^*_n(s)$, where \[ \tilde \Omega^*_n(s) =n^{-1} \sum_{i\in N_n} \sum_{j\in N^{*\partial}_n(i,s)} Y_{n,i}Y_{n,j}^\top, \] and $N_n^{*\partial}(i;s)=\{j\in N_n: d^*_n(i,j)=s\}$ is the set of nodes of distance $s$ from $i$ according to the observed network $G^*_n$. We denote the resulting HAC estimator as \[ \tilde V^*_n=\sum_{s\ge 0}\omega_n(s)\tilde \Omega^*_n(s). \]
The implications for the HAC estimator are two-fold: (a) some terms $Y_{n,i}Y_{n,j}^{\top}$ would appear in the covariance term $\tilde{\Omega}^*_n(s)$ with a larger distance $s$ than the true distance; (b) some terms $Y_{n,i}Y_{n,j}^{\top}$ would be missing from the estimator because there is no observed path between $i$ and $j$ in $G^*_n$. The direct consequence of (a) is that such terms would be assigned smaller weights $\omega_n(s)$ compared to those one would assign if the true network was observed. However, since the weights must converge to one according to Assumption (ref)(ii), the effect of (a) would be asymptotically negligible provided that the true unobserved network satisfies the rest of the conditions in Assumption (ref). From the expression for $V_n$ in (ref), one can also see that the effect of (b) is asymptotically negligible if the number of missing $Y_{n,i}Y_{n,j}^{\top}$ terms in the HAC estimator is of a smaller order than $n$.
We define $\delta_n^\partial(s\mid d^*_n=\infty)$ as the average number of $s$-neighbors that are isolated in the partially observed network: \[ \delta_n^\partial(s\mid d^*_n=\infty) = \frac{1}{n} \sum_{i\in N_n} | \{j\in N_n^\partial(i;s): d^*_n(i,j)=\infty\}|. \]
Assumption (ref) controls the share of nodes that appear isolated due to missing links. For example, the assumption holds if the total number of such nodes is $o_{a.s.}(n)$. For the consistency of the HAC estimator with partially observed networks, we also assume that the true network satisfies the conditions in Assumption (ref).
In the case of non-zero means, the estimator is defined similarly to (ref):
Let $N_n^{*}(i;s)=\{j\in N_n: d^*_n(i,j)\leq s\}$ be the set of nodes within distance $s$ from $i$ according to the observed network $G^*_n$, and define $\delta_n^*(s)=n^{-1}\sum_{i\in N_n} \vert N_n^{*}(i;s) \vert$. We have the following result.
The monotonicity condition for $\vert \omega(s)-1\vert $ holds, for example, for the truncated, Parzen, and Tukey-Hanning kernels.
For our simulation study, we use a version of the network formation model described in Section (ref). For each sample size $n=500,1000$, and $5000$, we randomly sample $n$ points, $\{X_1,\ldots,X_n\}$ from the uniform distribution on $[0,1]^2$. These points represent the nodes of a random graph $G_n$. Two nodes $i,j\in N_n$ become connected with probability that is inversely proportional to the Euclidean distance between $X_i$ and $X_j$, that is, \[ \PM(\{i,j\}\in E_n \mid X_i,X_j)=\exp\left(-\norm{X_i-X_j} \sqrt{2\pi n/\lambda}\right), \] where $\lambda$ is a positive constant that determines the average degree of the resulting graph. To reduce the dependence of the results on a particular realization of the latent process $\{X_1,\ldots,X_n\}$, a new network was drawn in each Monte Carlo repetition. We use the values $\lambda$=$1$, $2$, $3$, $4$, and $5$.
To generate $\{Y_{n,i}\}$, we consider a special case of the network dependent processes presented in Section (ref). Specifically, we generate samples using the following linear model:
where $\{\varepsilon_{n,i}\}$ are independent $\mathcal{N}(0,1)$ random variables, and $\gamma=0.0$, $0.1$, $0.2$, $0.3$, $0.4$, and $0.5$.
We use the $\hat V_n$ version of the network HAC estimator that does not assume a known mean. To compute the HAC estimator, the bandwidth is chosen according to the rule in (ref) with $\varepsilon=0.05$ and the constant equal to $1.7$, $1.8$, $1.9$, $2.0$, $2.1$, and $2.2$. We use the Parzen kernel given by \[ \omega(x)\eqdef
\]
The number of Monte Carlo repetitions is set to 10,000.\footnote{ The simulations were performed on Compute Canada clusters in Julia using 640 CPUs and 3GB of memory per CPU. The total computation time was 10.5 hours. } In each Monte Carlo repetition, we compute the average $\bar Y_{n}$, the HAC estimator $\hat V_n$, and construct the 95% asymptotic confidence interval for the mean of $Y_{n,i}$'s as $\bar Y_n \pm z_{0.975} \times (\hat V_n/n)^{1/2}$, where $z_{0.975} $ is the $0.975$-th percentile of the standard normal distribution.
According to our simulations with the Parzen kernel, setting the bandwidth constant to $2.0$ provides the most accurate coverage in terms of the average squared distance from the nominal coverage probability of $0.95$, where the average is computed across the all considered data generating processes. The distance exhibits a U-shape pattern across the considered constant values.
In Table (ref) we report the simulated coverage probabilities obtained with the bandwidth constant equal to $2.0$. The results for the other bandwidth constants are reported in Appendix (ref) of the Supplemental Note. The Supplemental Note also reports the simulated rejection probabilities of the corresponding HAC-based $t$-test.
Table (ref) also reports networks statistics such as the diameter, average degree, maximum degree, and average connected distance. One can see from the table that the average degree of the simulated networks is very close to the value of $\lambda$, and that larger values of $\lambda$ correspond to denser networks. Note also that, for example, in the case of $\lambda=3$, $n=1000$, and the bandwidths constant equal to $2.0$, our bandwidth selection rule (ref) produces the bandwidth of approximately $12.58$. In this case, the average simulated diameter is $41.70$, and the average connected distance is $15.89$. Hence, a non-trivial amount of truncation is applied when computing the HAC estimator (except for $\lambda=1$).
Note that when $\lambda=1$, the simulated coverage probabilities do not vary with the constant in the bandwidth selection rule in (ref), see Tables (ref) and (ref) in the Supplemental Note. As reported in Table (ref), in that case the simulated average degree is below one, and the bandwidth values resulting from (ref) exceed the diameters of the simulated networks even for the smallest considered value of the constant. Nevertheless, the coverage of the confidence intervals remains accurate because the networks generated with $\lambda=1$ are sparse with the average connected distance of $2.75-3.01$.
Table (ref) shows that while in the majority of the cases the simulated coverage probabilities are close to the nominal coverage of $0.95$, the performance of the HAC-based confidence intervals deteriorates for larger values of the denseness parameter $\lambda$ and the dependence parameter $\gamma$. Nevertheless, the coverage improves with the sample size. For example, when $n=5000$ the simulated coverage probabilities are between $0.931-0.949$ even for denser graphs with $\lambda=5$ as long as the dependence parameter $\gamma$ does not exceed $0.3$.
\putbib[network_dep]