Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
29,444 characters · 14 sections · 16 citation commands
GMM and M Estimation under Network Dependence
In recent years, asymptotic analysis of network-dependent data has garnered significant attention in econometrics \citep*[e.g.,][]{kuersteiner2019limit,leung2019normal,kuersteiner2020dynamic,KMS2021}.\footnote{See also jenish2012spatial for related work in which the dependence is embedded in Euclidean space. This set of references focuses on work related to weakly dependent structures, and thus excludes another important branch of the literature on network asymptotics, namely, the literature on exchangeable arrays, because its strong dependence structure differs substantially from the focus of the present paper, both in terms of network configuration and asymptotic theory. For the convenience of readers, however, I list some key theoretical contributions in this area: GRAHAM202023, davezies2021inference, menzel2021bootstrap, chiang2023inference, and graham2024kernel, among others. A related notion of dependence has also been used for time series babii2022machine.} In the most recent one of these, \citet*[KMS,][]{KMS2021} establish limit theorems and develop a robust variance estimator for a general class of dependent processes that encompass dependency-graph models in particular. Their framework, grounded in a conditional $\psi$-dependence concept adopted from doukhan1999new, offers powerful tools for handling network data and has spurred further research in related fields. Furthermore, the theory and methods introduced by KMS have been widely applied in econometric studies of network models leung2022_dependence,gao_ding2023_network,hoshino_yanagi2024_noncompliance.
Many applications, particularly those involving nonlinear models like limited dependent variable models, demand uniform convergence results. In the context of general classes of M estimators (including maximum likelihood estimators) and generalized method of moments (GMM) estimators, a uniform law of large numbers (ULLN) is crucial for ensuring that the empirical criterion function converges uniformly to its population counterpart. This uniform convergence is fundamental for establishing the consistency, and subsequently the asymptotic normality, of these estimators, as detailed in standard references such as the handbook chapter by NeweyMcFadden1994.
Although KMS offer elegant pointwise limit theorems under network dependence, their results do not directly yield the uniform law of large numbers (ULLN) required for nonlinear estimation. Achieving uniform convergence necessitates controlling not only the individual moments of network-dependent observations but also the fluctuations of the entire process uniformly across the parameter space.
The main contribution of this paper is to bridge this gap by establishing a novel ULLN under network dependence. The results build on the framework of \citet*{KMS2021}, which leverages model restrictions based on conditional $\psi$-dependence and decay rates of network dependence, concepts that will be briefly reviewed in Section (ref). To extend pointwise convergence to uniform convergence, I impose additional regularity conditions including the uniform equicontinuity. The resulting ULLN is then applied to establish the consistency and asymptotic normality of the GMM and M estimators.
This paper was motivated by a practical question raised by a graduate student: “Can the results of KMS be extended to nonlinear GMM estimation?” In exploring this question, I identified a critical gap, namely, the lack of a ULLN in the KMS framework, as mentioned above. The purpose of this paper is to address that gap and help bridge the elegant theoretical work of KMS with practical applications in empirical research. It is important to emphasize that the developments presented here rely heavily on the foundational contributions of KMS. While this paper provides a step toward applying their theory to GMM and M estimation, I encourage readers using the results in the present paper to give primary credit to KMS for laying the essential groundwork.
The remainder of the paper is organized as follows. In Section (ref), I introduce the setup. Section (ref) presents the ULLN. Sections (ref) and (ref) introduce M and GMM estimators, respectively, and their asymptotic properties. These two sections also provide practical guidelines. Section (ref) concludes. Mathematical proofs of all the theoretical results are provided in the appendix.
This section introduces the econometric framework.
First, I introduce some basic notations. Let \( v, a \in \mathbb{N} \). For any function \( f: \mathbb{R}^{v \times a} \to \mathbb{R} \), define \[ \|f\|_\infty = \sup_{x \in \mathbb{R}^{v \times a}} |f(x)| \quad\text{and}\quad \operatorname{Lip}(f) = \sup_{x \neq y} \frac{|f(x)-f(y)|}{d(x,y)}, \] where \( d(x,y) \) is a metric on \( \mathbb{R}^{v \times a} \). With these definitions, we introduce the class of uniformly bounded Lipschitz functions: \[ L_{v,a} = \Bigl\{ f : \mathbb{R}^{v \times a} \to \mathbb{R} \, : \, \|f\|_\infty < \infty \text{ and } \operatorname{Lip}(f) < \infty \Bigr\}. \]
This subsection provides a concise overview of the baseline model introduced in \citet*[KMS,][]{KMS2021} and the notational conventions used in the KMS framework; for a more detailed exposition, please refer to the original paper by KMS.
For each \(n \in \mathbb{N}\), let \(N_n = \{1,2,\ldots,n\}\) denote the set of indices corresponding to the nodes in the network \(G_n\) with the adjacency matrix \(A_n\) whose elements are $0$ and $1$. A link between nodes \(i\) and \(j\) exists if and only if the \((i,j)\)-th entry of \(A_n\) equals one. For each \(n \in \mathbb{N}\), let \(\mathcal{C}_n\) be the \(\sigma\)-algebra with respect to which the adjacency matrix \(A_n\) is measurable. Let \(d_n(i,j)\) denote the network distance between nodes \(i\) and \(j\) in \(N_n\), defined as the length of the shortest path connecting \(i\) and \(j\) in \(G_n\).
For \(a,b \in \mathbb{N}\) and a positive real number \(s\), define \[ P_n(a,b;s) = \Bigl\{ (A,B) \, : \, A,B \subset N_n,\; |A| = a,\; |B| = b,\; \text{and}\; d_n(A,B) \ge s \Bigr\}, \] where \[ d_n(A,B) = \min\{ d_n(i,j) : i \in A,\; j \in B \}. \] Thus, each element of \(P_n(a,b;s)\) is a pair of node sets of sizes \(a\) and \(b\) with a distance of at least \(s\) between them.
Consider a triangular array \(\{Y_{n,i}\}_{i \in N_n}\) of random vectors in \(\mathbb{R}^v\). The following definition introduces the notion of conditional \(\psi\)-dependence as provided in KMS.
As emphasized in KMS, it is important to note that the decay coefficients are generally random, allowing one to accommodate the “common shocks” \(\mathcal{C}_n\) present in the network. I now present the following two key assumptions from KMS, which will be employed throughout the present paper.
For each node $i\in N_n$ in the network for each row $n$ and $s\ge 1$, define \[ N_{n}(i;s) = \{ j\in N_n : d_n(i,j) \leq s\} \quad\text{and}\quad N_{n}^\partial(i;s) = \{ j\in N_n : d_n(i,j)=s\}, \] representing the number of nodes within and at a distance $s$, respectively. Then, define the average shell size \[ \delta_{n}^\partial(s) = \frac{1}{n}\sum_{i\in N_n} |N_{n}^\partial(i;s)|. \] With this notation, the following assumption restricts the denseness of the network and the decay rate of dependence with the network distance.
I refer readers to the original paper by KMS for detailed discussions of these assumptions, as they are excerpted from KMS. Under these assumptions, along with additional moment and regularity conditions, KMS establish the pointwise law of large numbers -- see Proposition 3.1 in their paper.
This subsection introduces a parameter-indexed class of functions and imposes additional restrictions to establish the uniform law of large numbers.
Let $\Theta\subset\mathbb{R}^d$ denote a parameter space. For each $\theta\in\Theta$, let \[ f(\cdot,\theta): \mathbb{R}^v \to \mathbb{R} \] be a measurable function. I impose the following conditions on the parameter space $\Theta$ and the function class $\{f(\cdot,\theta):\theta\in\Theta\}$.
For $p>0$, let $\|f(Y_{n,i},\theta)\|_{\mathcal{C}_n,p}$ denote the conditional $L^p$ norm defined by \[ \|f(Y_{n,i},\theta)\|_{\mathcal{C}_n,p} = \bigl(E\bigl(|f(Y_{n,i},\theta)|^p \mid \mathcal{C}_n\bigr)\bigr)^{1/p}. \] With this notation, the following assumption imposes conditions on the function class.
Assumption (ref) (i) is the bounded moment condition required by Assumption 3.1 of KMS with $f(Y_{n,i},\theta)$ treated as an observation in place of $Y_{n,i}$. Besides, Assumption (ref) (ii) imposes the uniform bound and Lipschitz conditions on each function $f(\cdot,\theta)$ in the class. Taken together, these two components impose restrictions on the behavior of \( f(Y_{n,i}; \theta) \) for `each' \( \theta \), without placing any constraint on the effects of \( \theta \) on it.
For `each' \(\theta \in \Theta\), the pointwise law of large numbers, as stated in Proposition 3.1 of KMS, holds under Assumptions (ref), (ref), and (ref). I will leverage this pointwise result by KMS as an auxiliary step in establishing the uniform law of large numbers, which requires the following uniform equicontinuity condition in addition.
Assumption (ref), together with Assumption (ref), will allow me to have a finite-net approximation of $f(Y_{n,i},\theta)$ for all $\theta \in \Theta$, as a way to establish the uniform result.
I now state the uniform law of large numbers for network-dependent data.
The next two sections demonstrate how this result can be applied to establish the consistency and asymptotic normality of GMM and M estimators. From this point onward, I focus on the case of a trivial sigma-field $\mathcal{C}_n$ and omit conditioning on it, following the convention in the existing literature leung2022_dependence,gao_ding2023_network,hoshino_yanagi2024_noncompliance, which actually applies the large-sample theory developed by KMS.
To proceed, I introduce few additional notations. Following KMS (Section 3.1), define $$ c_n(s,m;k) = \inf_{\alpha>1}[\Delta_n(s,m;k\alpha)]^{1/\alpha}[\delta_n^\partial(s;\alpha/(\alpha-1))]^{1-1/\alpha} $$ to control the network dependence at distance $s$, where
Recall that $N_n(i;s)$ and $N_n^\partial(i;s)$ are defined in Section (ref). I refer readers to KMS (Section 3.1) for discussions of these objects and the roles which they play. Finally, let $\lambda_{\min}(A)$ denote the minimum eigenvalue of square matrix $A$.
Let $Q(\cdot)$ and $Q_n(\cdot)$ be the population and sample criterion functions for M estimation, defined on $\Theta$ by $$ Q(\theta) = E\bigl(f(Y_{n,i},\theta)\bigr) \quad\text{and}\quad Q_n(\theta) = \frac{1}{n}\sum_{i \in N_n} f(Y_{n,i},\theta), $$ respectively.\footnote{I consider the case where the population criterion is independent of $n$, but a slight modification of the assumptions can accommodate settings where the population criterion depends on $n$.} The M estimator is defined as \[ \hat{\theta}_{M} \in \arg\max_{\theta\in\Theta} Q_n(\theta). \] In the pesudo maximum likelihood estimation (PMLE) framework, $f(Y_{n,i},\theta)$ corresponds to the logarithm of the marginal density function of $Y_{i,n}$ given the parameter $\theta$.
Suppose that the population criterion satisfies the following condition.
With this identification condition, the standard argument based on NeweyMcFadden1994, for example, yields the consistency $\hat\theta_{M} \stackrel{p}{\rightarrow} \theta_0$ by the uniform law of large numbers (my Theorem (ref)). Let me state this conclusion formally as the following corollary to Theorem (ref).
To establish the asymptotic normality, the following two assumptions are used in addition.
This assumption is invoked to directly obtain the CLT of KMS (their Theorem 3.2) for establishing the asymptotic normality of $\sqrt{n} a^\top \nabla_\theta Q_n(\theta_0)$ for a vector $a$ such that $\|a\|=1$.\footnote{While KMS consider multivariate random variables, their CLT result is stated for univariate cases.} With our focus on the trivial sigma-field $\mathcal{C}_n$, part (i) of Assumption (ref) implies Assumption 3.3 of KMS, part (ii) implies Assumption 2.1 (b) of KMS, and part (iii) implies Assumption 3.4 of KMS by Rayleigh quotient. Part (iv) requires that the vairance of the mean score in the griangular array converges. We refer readers to KMS for discussions of these conditions.
Parts (i)--(ii) of Assumption (ref), together with Assumptions (ref), (ref), and (ref), are used to invoke the uniform law of large numbers (Theorem (ref)) on the Hessian: $\sup_{\theta \in \Theta} |\nabla_{\theta\theta} Q_n(\theta) - \nabla_{\theta\theta} Q(\theta)| \rightarrow 0$ a.s., where the equicontinuity in part (i) and the $L^1$ dominance in part (ii) allow the dominated convergence theorem to yield $\nabla_{\theta\theta} Q(\theta) = E\bigl(\nabla_{\theta\theta} f(Y_{n,i},\theta)\bigr)$, which is guaranteed to be a continuous function of $\theta$. Further, part (iii) ensures that its limit is invertible at $\theta_0$.
Now, combining the CLT of KMS (their Theorem 3.2) with my Theorem (ref) and Corollary (ref), we obtain the following asymptotic normality result through checking the conditions of NeweyMcFadden1994.
The current section presents the practical procedure to implement an M estimation under network dependence.
First, obtain the estimate \[ \hat{\theta}_{M} \in \arg\max_{\theta\in\Theta} \frac{1}{n}\sum_{i \in N_n} f(Y_{n,i},\theta). \]
Second, adapting the network HAC estimation procedure of KMS (Section 4) to the present framework of M estimation, compute the network-robust variance estimate \[ \hat \Sigma = \sum_{s \geq 0} \omega(s/b_n) \cdot \frac{1}{n} \sum_{i \in N_n} \sum_{j \in N_n^\partial(i;s)} \left( \nabla_\theta f(Y_{n,i},\hat\theta) \right) \left( \nabla_\theta f(Y_{n,i},\hat\theta) \right)^\top \] for the score, where $\omega(\cdot)$ denotes a kernel function\footnote{The kernel $\omega: \mathbb{R} \rightarrow [-1,1]$ satisfies $\omega(0) = 1$, $\omega(z) = 0$ for $|z| > 1$, and $\omega(z) = \omega(-z)$ for all $z \in \mathbb{R}$.} and $b_n$ is a bandwidth parameter.
For example, using the Parzen kernel, \[ \omega(u) =
\] KMS demonstrate that the following bandwidth choice performs well in simulations:\footnote{That said, the optimal choice of bandwidth should remain an important direction for future research.} \[ b_n = \frac{2\log n}{\log\left( \max\{\hat\delta_n^\partial(1), 1.05\} \right)}, \] where $\hat\delta_n^\partial(1)$ denotes the average degree of the observed network.
Finally, compute the Hessian estimator \[ \hat H = \frac{1}{n} \sum_{i \in N_n} \nabla_{\theta\theta} f(Y_{n,i},\hat\theta). \] Note that even in the PMLE framework, the information equality (which is established under i.i.d. sampling) may not hold in general under network dependence.
Let $f(\cdot,\cdot)$ denote the moment function such that the true parameter vector $\theta_0 \in \Theta$ satisfies the moment equality $$ E\bigl(f(Y_{n,i},\theta_0)\bigr) = 0. $$ Define the sample moment function by \[ \bar{f}_n(\theta) = \frac{1}{n}\sum_{i=1}^n f(Y_{n,j}, \theta). \] For any sequence \(W_n\) of positive definite weighting matrices (which may depend on the data) converging in probability to a positive definite matrix $W$, the GMM estimator is defined as
We can define the population criterion by\footnote{A similar remark to Footnote (ref) applies here.} \[ Q(\theta) = E\bigl(f(Y_{n,i},\theta)\bigr)^\top W E\bigl(f(Y_{n,i},\theta)\bigr). \]
Suppose that the population moment satisfies the following condition.
With this identification condition, the standard argument based on NeweyMcFadden1994, for example, yields the consistency $\hat\theta_{GMM} \stackrel{p}{\rightarrow} \theta_0$ by the uniform law of large numbers (my Theorem (ref)). Let me state this conclusion formally as the following corollary to Theorem (ref).
To establish the asymptotic normality, the following two assumptions are used in addition.
Assumptions (ref) and (ref) are analogous to Assumptions (ref) and (ref), respectively, and hence similar discussions apply, which are omitted here to avoid repetitions.
Combining the CLT of KMS (their Theorem 3.2) with my Theorem (ref) and Corollary (ref), we obtain the following asymptotic normality result through checking the conditions of NeweyMcFadden1994.
The current section presents the practical procedure to implement an GMM estimation under network dependence.
First, obtain the estimate
Second, adapting the network HAC estimation procedure of KMS (Section 4) to the present framework of GMM estimation, compute the network-robust variance estimate \[ \hat \Omega = \sum_{s \geq 0} \omega(s/b_n) \cdot \frac{1}{n} \sum_{i \in N_n} \sum_{j \in N_n^\partial(i;s)} f(Y_{n,i},\hat\theta) f(Y_{n,i},\hat\theta)^\top, \] where $\omega(\cdot)$ is a kernel function and $b_n$ is a bandwidth parameter. See Section (ref) for further discussions of $\omega(\cdot)$ and $b_n$.
Finally, compute the gradient estimator \[ \hat G = \frac{1}{n} \sum_{i \in N_n} D_\theta f(Y_{n,i},\hat\theta). \] As usual, one may iterate the above procedure to implement the two-step GMM estimation.
This paper establishes the asymptotic properties of GMM and M estimators under network dependence. As a key step toward this goal, I extend the law of large numbers from \citet*[Proposition 3.1]{KMS2021} to a novel uniform law of large numbers (ULLN), stated in Theorem (ref). Since the consistency of nonlinear estimators, such as GMM and M estimators, requires uniform convergence of the criterion functions, this result lays the foundation for proving their consistency and, subsequently, their asymptotic normality. For completeness, Sections (ref) and (ref) present full sets of assumptions under which these asymptotic properties hold for the M and GMM estimators, respectively.
As already mentioned in the introduction, this paper originated from a practical question posed by a graduate student: “Can the results of KMS be applied to nonlinear GMM estimation?” In addressing this question, I identified a key gap, namely, the absence of a ULLN in KMS as discussed above. This paper was written to bridge that gap and connect the elegant theory of KMS with the needs of empirical practitioners. That said, the results presented here build heavily on KMS, and much of the foundational work and theoretical development should be credited to their contribution. Accordingly, even if readers use the results presented in this paper in the context of GMM and M estimation, I strongly encourage them to give primary credit to KMS, whose work has done most of the heavy lifting.