The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
107,344 characters
\hypersetup{linkcolor=black}
\begin{center}
{\Large \bfseries Moment Restrictions for Dyadic Network Formation Models with Nontransferable Utility\par}
\vspace{1.2em}
{\large Hanping Chen\textsuperscript{a}\footnote{Email: [email removed]}.}\quad Zeqi Wu\textsuperscript{b}\footnote{Email: [email removed]}.}\par}
\vspace{0.6em}
{\small \textsuperscript{a}\,School of Management and Economics, The Chinese University of Hong Kong, Shenzhen\par}
{\small \textsuperscript{b}\,Institute of Statistics and Big Data, Renmin University of China\par}
\end{center}
\setcounter{footnote}{0}
\hypersetup{linkcolor=blue}
\vspace{1em}
\global\long\global\long\global\long\def\1{\mathds{1}}
\global\long\global\long\global\long\def\bs#1{\boldsymbol{#1}}
\global\long\def\bf#1{\mathbf{#1}}
\global\long\global\long\global\long\global\long\global\long\global\long\global\long
\begin{abstract}
This paper investigates the construction of moment restrictions in dyadic
network formation models with unobserved individual heterogeneity under nontransferable utility.
Using observed links and covariates from five-node pentads, we construct moment restrictions
that do not depend on individual fixed effects.
For a broad class of covariate specifications, the construction is minimal
in the sense that it uses the least possible number of nodes and dyads.
Based on these moment restrictions, we propose the pentad-GMM estimator.
We establish asymptotic normality of the pentad-GMM estimator in dense, sparse,
and ultra-sparse network regimes, with regime-specific convergence rates and asymptotic variances. These results provide a basis for inference across all
three regimes.
To make our method computationally efficient, we develop an algorithm that reduces the
computational cost of the estimator from na\"ive \(O(N^5)\) to \(O(N^3)\).
We apply the proposed method to an academic-discussion network, and find academic
homophily and a positive association between link formation and a potential partner's openness
to different perspectives.
\end{abstract}
\noindent \textit{JEL Classification:} C10, C13, C25, D85
\noindent \textit{Keywords:} Network Formation, Nontransferable Utility, Fixed Effects, Sparse Networks, Functional Differencing
\section{Introduction}
Networks are pervasive in both economic and social contexts. The relationships that
individuals form with each other reflect their characteristics and preferences. Lenders, for
example, decide whom to lend to based on borrowers' credit and repayment histories,
while researchers and firms seek collaborators with complementary expertise
and resources. Understanding how these characteristics and preferences shape
link formation is central to explaining social and economic
interactions.
An important distinction in network formation modeling is whether utility
is transferable between the two agents.
Under transferable utility (TU), utilities or payoffs can
be redistributed through compensating mechanisms, and a link forms when
the joint surplus of the two individuals is nonnegative.
The transferable-utility assumption, however, is restrictive and may not
hold in many applications. In friendship networks, for instance, a link requires
mutual recognition by both individuals, but there is typically no
transfer that can compensate one party for
entering a relationship they would otherwise decline.
This setting is more naturally described by
non-transferable-utility (NTU) models. Under NTU, each individual makes a
unilateral decision about whether to form the relationship, and a link
is established if and only if both agents agree
(\citealp{myerson1991game}, \citealp{jackson1996strategic}).
These linking decisions depend not only on observed characteristics
but also on unobserved traits such as sociability, productivity, and risk tolerance.
A highly extroverted individual, for example,
forms relationships across a broad range of observed types, a pattern that
could otherwise be attributed to weak homophily \citep{graham2017econometric}.
Moreover, ignoring such heterogeneity when it is correlated with observed
characteristics can bias estimates of homophily and other covariate effects.
One way to account for this heterogeneity is to estimate the
parameter of interest and the individual fixed effects by joint maximum likelihood.
Treating individual heterogeneity as fixed effects to be estimated, however,
leads to the well-known incidental-parameter problem of \citet{neyman1948consistent},
which can induce asymptotic bias in the estimate of the parameter of interest.
Several methods have been proposed to correct incidental-parameter bias in
nonlinear panel models with individual fixed effects
(e.g., \citealp{hahn2004jackknife}; \citealp{dhaene2015split};
\citealp{fernandez2016individual}), and related work extends these methods to
network models with individual fixed effects
(e.g., \citealp{hughes2022estimating,li2024estimation}).
However, a limitation of these methods is their reliance on
compactness restrictions on the individual fixed effects.
With bounded fixed effects, such restrictions typically exclude
sparse-network sequences in which link
probabilities vanish as the network grows.
Yet sparsity is a common feature of empirical networks
\citep{graham2024sparse}.
An alternative way is to eliminate and difference out the individual fixed effects.
For nonlinear panel models, approaches to eliminating fixed effects
include methods based on sufficient statistics, functional differencing,
and maximum score restrictions
(e.g., \citealp{chamberlain1980analysis}; \citealp{manski1987semiparametric};
\citealp{bonhomme2012functional}).
Analogous approaches are available for TU network formation
(e.g., \citealp{graham2017econometric}; \citealp{toth2017semiparametric};
\citealp{bonhomme2023functional}).
The joint-surplus structure in TU models preserves additive separability in
the fixed effects, making these differencing methods comparatively
straightforward to apply.
However, these arguments do not carry over directly to NTU models.
The bilateral structure of NTU models poses a different challenge.
Under NTU, researchers usually observe only the realized undirected link,
not the two underlying unilateral decisions. A realized link reveals mutual
consent, whereas an absence of a link does not reveal whether one or both individuals
refuse to link. Moreover, the two individual fixed effects enter the observed link probability
through separate asymmetric latent utilities. As emphasized by
\citet{gao2023logical} and \citet{li2024estimation}, this asymmetry, combined
with partial observability, breaks the additive separability exploited by
standard sufficient-statistic and differencing arguments.
To address these challenges, we propose a pentad differencing
strategy for logistic NTU network formation models with individual
unobserved heterogeneity. The construction is analogous in spirit to
functional differencing in nonlinear panel models
(\citealp{bonhomme2012functional,honore2024moment,bonhomme2025feedback}).
The main idea is to fix two pivotal nodes and form three triangles sharing
the pivotal dyad, using three peripheral nodes and seven dyads in total.
Eliminating each peripheral node's fixed effect gives a homogeneous linear equation in two pivotal nodes'
fixed effects and a constant term. Together, stacking the
three equations implies that a \(3\times3\) coefficient matrix is singular, and thus has zero
determinant. Using this zero-determinant restriction, we construct a moment
function involving only observed links and covariates that has zero conditional
expectation at \(\bs\beta_0\) and forms the basis for GMM estimation.
With the individual fixed effects eliminated, the method does not
require compactness restrictions on their support and accommodates sparse
networks in which link probabilities vanish as the network size grows.
We also show that, for a broad class of covariate specifications, five
nodes and seven dyads are necessary for constructing a nontrivial fixed-effect-free
moment restriction. Among constructions using five nodes and seven dyads,
ours is unique up to relabeling, and the resulting moment restrictions are unique up to rescaling.
We then establish asymptotic normality of the resulting pentad-GMM estimator across
dense, sparse, and ultra-sparse networks.
Conditional on the observed characteristics and unobserved individual
heterogeneity, we apply a Hoeffding decomposition over dyads
(\citealp{hoeffding1948class}) and show that the leading projection varies
with network sparsity.
These three sparsity regimes are characterized by how the average
link probability \(\rho_N\) scales with network size \(N\).
In dense and mildly sparse networks, with
\(N\rho_N\to\infty\), the one-dyad projection dominates. In the
sparse regime, with
\(\rho_N\asymp N^{-1}\), several graph projections contribute
at the same order. In the ultra-sparse regime, with
\(N\rho_N\to0\) and \(N^5\rho_N^4\to\infty\),
the leading variance contribution comes from projections associated
with some five-node, four-edge trees.
The change in the leading projection contrasts with the tetrad-logit estimator of
\citet{graham2017econometric}, whose asymptotic representation is
driven by dyad-level projections under both the dense and sparse
networks considered there.
To accommodate the different leading projections across sparsity regimes, we
propose a tractable covariance estimator. The estimator is
consistent under the appropriate normalization in all three regimes and
requires no prior knowledge of \(\rho_N\) or the sparsity regime.
For implementation, we exploit the determinant structure to reduce the
computational cost of the estimator from na\"ive \(O(N^5)\)
to \(O(N^3)\) operations.
The empirical application demonstrates the proposed method using data from an
academic-discussion network among Master of Social Work (MSW) students.
The asymmetric NTU estimates reveal academic homophily and a tendency to
form discussion links within the same cohort. They also suggest that students
who are more open to different perspectives are more attractive discussion
partners.
\textit{Related Literature}. This paper contributes to the literature
on econometric models of dyadic link formation in a single large network with
unobserved heterogeneity. A major
strand of this literature studies dyadic link formation models in which
unobserved heterogeneity enters through individual-specific fixed effects.
An incomplete list includes
\citet{chatterjee2011random}, \citet{yan2013central},
\citet{graham2017econometric},
\citet{charbonneau2017multiple}, \citet{toth2017semiparametric},
\citet{jochmans2018semiparametric}, \citet{dzemski2019empirical},
\citet{gao2020nonparametric}, \citet{gao2023logical},
\citet{hughes2022estimating}, \citet{candelaria2024semiparametric} and
\citet{zeleneev2020identification}.
For comprehensive reviews, see
\citet{de2020econometric} and \citet{graham2020network}.
Much of this work focuses on transferable-utility (TU) network formation
models with individual fixed effects. Under logistic errors, fixed
effects can be eliminated by conditioning on sufficient statistics such
as the degree sequence (e.g., \citealp{graham2017econometric};
\citealp{charbonneau2017multiple}). Semiparametric approaches also
eliminate fixed effects without imposing a parametric distribution on the
idiosyncratic shocks. For example, \citet{toth2017semiparametric} uses rank-based arguments,
while \citet{candelaria2024semiparametric} combines a special-regressor
transformation with arithmetic differencing.
Related work also studies fixed-effect estimation and bias correction in
nonlinear dyadic models, including the jackknife approach developed by
\citet{hughes2022estimating}. More recently, \citet{yan2026penalized} develop a penalized-likelihood
approach that allows the fixed effects to diverge at a logarithmic rate and
accommodates sparse degree sequences. Their primary specification observes
both directed links within each dyad.
Our paper differs from this work in focusing on undirected NTU link formation
with individual fixed effects,
giving rise to partial observability.
The papers closest to our setting are \citet{gao2023logical} and \citet{li2024estimation}.
\citet{gao2023logical} study undirected network formation under
nontransferable utility (NTU) and propose \textit{logical differencing},
which uses logical contraposition to eliminate fixed effects without imposing
parametric assumptions on idiosyncratic shocks (see also
\citealp{gao2020robust}). Their analysis focuses on identification and
consistent estimation rather than large-network inference. Furthermore, their analysis is based
on dense-network asymptotics, with conditional linking probabilities
remaining nonvanishing as the network grows.
\citet{li2024estimation} develop
estimation and inference methods for NTU models with individual fixed effects
by combining a joint method-of-moments initial estimator, a Le Cam one-step
refinement, and a split-network jackknife with bagging. Their procedure
corrects incidental-parameter bias and delivers asymptotic normality for the
homophily estimator. Nevertheless, they impose compactness restrictions on the fixed
effects, which excludes sparse networks. We instead construct moment conditions that
eliminate the fixed effects and develop asymptotic normality for both dense and
sparse networks without imposing compactness on the fixed effects.
Our work also relates to the literature on sparse network asymptotics
and inference under dyadic dependence. A key insight from this literature
is that the leading variance component can change with network sparsity.
\citet{graham2024sparse} shows that, in logistic regression with dyadic
dependence, variance components that are negligible in dense networks can
become first order in sparse networks. Relatedly,
\citet{chandrasekhar2025network} study subgraph generated models and
develop asymptotic theory for estimators based on higher-order network
structures. This paper shares the emphasis on variance decompositions
generated by overlapping network configurations and extends it to
regime-adaptive inference for NTU link formation models with fixed effects.
The paper also builds on the panel data literature on eliminating fixed
effects in nonlinear models. In static and dynamic logit panels, fixed
effects can be eliminated by conditioning on sufficient statistics or
state-switching events, as in \citet{chamberlain1980analysis} and
\citet{honore2000panel}. Functional differencing provides a more general
route by constructing moment restrictions that are free of fixed effects
(\citealp{bonhomme2012functional}). Recent contributions extend and
systematize this logic in binary choice and dynamic panel settings,
including \citet{honore2019panel}, \citet{honore2021identification},
\citet{kitazawa2022transformations},
\citet{davezies2023fixed}, \citet{dano2023transition},
\citet{bonhomme2023functional}, \citet{honore2024moment},
\citet{bonhomme2025feedback}, and \citet{dobronyi2021identification}. This paper
follows the same broad principle by constructing moment restrictions that
do not contain the fixed effects.
While our approach builds on the idea of functional differencing, we develop
an explicit and minimal construction for NTU networks, which is, to the best of our knowledge, new to the literature.
Another line of the network formation literature studies
strategic link interdependence,
where the payoff from a link may depend on other links in the network.
This literature often models the observed network as an equilibrium object using pairwise stability or related network-game concepts,
with pairwise stability tracing back to \citet{jackson1996strategic}
(see, e.g., \citealp{miyauchi2016structural}, \citealp{mele2017structural},
\citealp{de2018identifying}, \citealp{sheng2020structural},
and \citealp{menzel2024strategic}).
Recent work also incorporates unobserved heterogeneity into
strategic network formation models; for example,
\citet{gao2026tractable} study strategic link interdependence
with individual fixed effects and develop tractable
subnetwork-based identifying restrictions under TU models.
The present paper abstracts from strategic link interdependence
and focuses instead on the fixed-effect problem in NTU link formation.
The remainder of the paper is organized as follows.
Sections~\ref{sec:model} and~\ref{sec:identification} present the NTU model,
the pentad moment function, and the GMM estimator. Section~\ref{sec:asymptotics}
develops the regime-specific asymptotic theory, while
Sections~\ref{sec:implementation} and~\ref{sec:empirical} present the
implementation details and empirical application. Section~\ref{sec:conclusion}
concludes. Appendix~\ref{app:graph_notation} introduces graph notation and terminology.
Appendices~\ref{sec:feasible_covariance} and~\ref{sec:simulation}
develop a tractable covariance estimator for inference and report the simulations, respectively.
The remaining appendices contain all proofs.
\section{Network Formation Model}
\label{sec:model} In this paper, the researcher
observes $(\bf L,\bf X)$, where $\bf L$ is the $N\times N$
adjacency matrix with binary $(i,j)$th entry $L_{ij}\in\{0,1\}$
indicating the presence of a link between individuals $i$ and $j$, and
$\bf X:=(\bs X_{1},\ldots,\bs X_{N})^\top$ is the $N\times d$
matrix of observed characteristics, with $\bs X_{i}\in\mathbb{R}^{d}$
for all $i\in[N]:=\{1,\ldots,N\}$. We focus on an undirected
non-transferable utility (NTU) network, so that $L_{ij}=L_{ji}$,
formed among $N$ individuals (nodes) in a single large network. Self-links
are ruled out by convention (i.e., $L_{ii}=0$). The parameter of interest,
$\bs{\beta}_{0}\in\mathbb{R}^{d}$, captures how observed dyadic covariates
affect each individual's unilateral linking decision.
Let $\bs\Gamma:=(\Gamma_{1},\ldots,\Gamma_{N})^{\top}\in \mathbb{R}^N$
denote the vector of individual-specific unobserved effects, which capture
heterogeneity in individuals' latent propensities to form links. Let the binary random variable $G_{ij}\in\{0,1\}$
indicate whether individual $i$ proposes to form a link with
individual $j$. Individual \(i\)'s unilateral linking decision is specified as:
\[
G_{ij}=\1\left\{ \Gamma_{i}+\bs X_{ij}^\top\bs{\beta}_{0}-\epsilon_{ij}\geq 0\right\} ,\quad1\leq i\neq j\leq N.
\]
In this specification, individual $i$'s latent utility depends on individual-specific
unobserved heterogeneity $\Gamma_{i}$ and observed dyad-level
covariates $\bs X_{ij}=w(\bs X_i,\bs X_j)$, where
$w:\mathbb R^d\times\mathbb R^d\to\mathbb R^d$. The map $w$ need not be
symmetric, so in general $\bs X_{ij}\neq\bs X_{ji}$. These
dyadic covariates capture homophily or complementarity between
individuals $i$ and $j$ and may enter their respective latent utilities
asymmetrically, even though the realized link is undirected.
In addition to the observed and unobserved
covariates, individual $i$'s utility also depends on the idiosyncratic
shock $\epsilon_{ij}$. Under the NTU model, link formation
requires the consent of both individuals:
\[
L_{ij}=\underbrace{G_{ij}}_{\text{\ensuremath{i} proposes to \ensuremath{j}}}\cdot\underbrace{G_{ji}}_{\text{\ensuremath{j} proposes to \ensuremath{i}}},\quad1\leq i\neq j\leq N.
\]
That is, a symmetric link $L_{ij}=L_{ji}$ is formed between $i$
and $j$ if and only if both individuals propose to link
and each obtains nonnegative utility from the connection. Therefore,
the link indicator $L_{ij}$ is the product of two binary indicators.
Each indicator equals one when that individual’s latent utility is nonnegative:
\begin{equation}
L_{ij}=\1\left\{ \Gamma_{i}+\bs X_{ij}^\top\bs{\beta}_{0}-\epsilon_{ij}\geq0\right\}
\cdot\1\left\{ \Gamma_{j}+\bs X_{ji}^\top\bs{\beta}_{0}-\epsilon_{ji}\geq0\right\} .\label{mod:mod2}
\end{equation}
The mutual-consent structure in \eqref{mod:mod2} captures the key economic
distinction between NTU and TU. It parallels the bilateral-consent requirement
in the canonical pairwise-stability framework, under which both agents must
favor a new link (\citealp{jackson1996strategic}). This distinction matters when
consent cannot be replaced by compensation. Friendships, peer
ties, and research collaborations often lack a price or complete contract that
can freely redistribute the gains from the relationship, so a large benefit to
one party does not offset a negative payoff to the other. By contrast, models
with transfers allow players to bargain over side payments
(\citealp{bloch2007formation}), and econometric TU models reduce link formation
to the sign of joint surplus (e.g., \citealp{graham2017econometric}). The NTU
formulation therefore better reflects relationships in which consent is
individual and compensation is limited.
We complete the model with the following baseline assumption.
\begin{assumption}[Random Sampling and Conditional Independence]\label{ass:dyad_ind}
For each \(N\), the node characteristics $\{(\bs X_i,\Gamma_i)\}_{i=1}^N$
are i.i.d. across \(i\), and their common distribution may depend on \(N\).
The conditional likelihood of the network $\bf L$, given
$\bf X$ and $\bs\Gamma$, is
\[
\mathbb{P}\left(\bf L=\bf l\mid\bf X,\bs\Gamma\right)=\prod_{i<j}\mathbb{P}\left(L_{ij}=l_{ij}\mid\bs X_{i},\bs X_{j},\Gamma_{i},\Gamma_{j}\right),
\]
with
\[
\begin{aligned}
\mathbb{P}\left(L_{ij}=l_{ij}\mid\bf X,\bs\Gamma\right)
&=
\left[
F\left(\Gamma_{i}+\bs X_{ij}^\top\bs{\beta}_{0}\right)
F\left(\Gamma_{j}+\bs X_{ji}^\top\bs{\beta}_{0}\right)
\right]^{l_{ij}}\\
&\quad\times
\left[
1-F\left(\Gamma_{i}+\bs X_{ij}^\top\bs{\beta}_{0}\right)
F\left(\Gamma_{j}+\bs X_{ji}^\top\bs{\beta}_{0}\right)
\right]^{1-l_{ij}}
\end{aligned}
\]
for all $i<j$, where $F(\cdot)$ is the standard logistic cumulative distribution function (cdf).
\end{assumption}
Assumption~\ref{ass:dyad_ind} imposes the random sampling and conditional
dyadic independence structure of the model. First, within each \(N\), the
node characteristics \((\bs X_i,\Gamma_i)\) are sampled i.i.d. across
individuals. This condition applies to the joint distribution of
\((\bs X_i,\Gamma_i)\) and does not require \(\bs X_i\) and \(\Gamma_i\)
to be independent. Second,
conditional on the full collection of node characteristics, unordered dyads are
independent. For each \(i<j\), the conditional probability of a realized
link has the mutual-consent product-logit form
\[
P_{ij}(\bs\beta_0)
:= \mathbb{P}(L_{ij}=1\mid \bf X, \bs\Gamma)=
\mathbb{P}(L_{ij}=1\mid \bs X_i,\bs X_j,\Gamma_i,\Gamma_j)
=
F(\Gamma_i+\bs X_{ij}^{\top}\bs\beta_0)
F(\Gamma_j+\bs X_{ji}^{\top}\bs\beta_0).
\]
The product-logit form is consistent with conditionally independent standard
logistic shocks \(\epsilon_{ij}\) and \(\epsilon_{ji}\) for the two directed
decisions within each dyad, while the likelihood factorization separately
imposes conditional independence across unordered dyads.
This sampling and conditional link-independence structure parallels
Assumptions 1 and 3 in \citet{graham2017econometric} and
\citet{li2024estimation}.
\section{Moment Restrictions and Estimation}
\label{sec:identification}
\subsection{Differencing via Pentads}
The mutual-consent rule in \eqref{mod:mod2} captures the bilateral nature of
link formation across a broad range of networks, as discussed in
Section~\ref{sec:model}. However, this structure makes the individual fixed
effects $\bs\Gamma$ difficult to difference out. As implied by the model in \eqref{mod:mod2},
both agents' decisions are revealed only when $L_{ij}=1$, in which case
$\left(G_{ij},G_{ji}\right)=(1,1)$. An absent link ($L_{ij}=0$), however,
can arise from three possible outcomes:
$\left(G_{ij},G_{ji}\right)\in\{(0,0),(1,0),(0,1)\}$.
The researcher therefore cannot tell whether $i$, $j$, or both declined to
form the link. If the directed proposal decisions $G_{ij}$ were observed, the
individual fixed effects could be eliminated using arguments analogous to
those used in TU models
(\citealp{graham2017econometric}).
Moreover, under Assumption~\ref{ass:dyad_ind}, the fixed effects
\(\Gamma_i\) and \(\Gamma_j\) do not enter the conditional link probability
\(P_{ij}(\bs\beta_0)\) through a single additive index
\citep{gao2023logical,li2024estimation}.
This further complicates the direct application of arithmetic-differencing arguments
developed for TU models (e.g., \citealp{toth2017semiparametric};
\citealp{candelaria2024semiparametric}).
We therefore develop a new differencing argument for the NTU model. Under
Assumption~\ref{ass:dyad_ind}, we construct a nontrivial moment function\footnote{A
moment function \(\psi(\bf L,\bf X;\bs{\beta})\) is trivial if
\(\psi(\cdot,\bf X;\bs\beta_0)\equiv0\) almost surely.}
$\psi(\bf L,\bf X;\bs{\beta})$ of the observed links and covariates that
satisfies
\begin{equation}
\mathbb{E}\big[\psi(\bf L,\bf X;\bs{\beta}_{0})\,\big|\,\bf X,\bs\Gamma\big]=0.
\label{eq:cond_moment_A-1}
\end{equation}
By the law of iterated expectations, \eqref{eq:cond_moment_A-1} implies
\(\mathbb{E}[\psi(\bf L,\bf X;\bs{\beta}_{0})\mid\bf X]=0\).
The resulting restriction, which conditions only on $\bf X$, forms the basis
for GMM estimation without requiring an estimator of $\bs\Gamma$. The
conditional moment in \eqref{eq:cond_moment_A-1} is in the spirit of
functional differencing (\citealp{bonhomme2012functional}), but the
construction developed here for the NTU model is new. It uses a five-node,
seven-dyad configuration with two pivot nodes, \(i,j\), and three peripheral
nodes, \(k,l,m\), as illustrated in
Figure~\ref{fig:five-node-triangles}.
\begin{figure}[htbp!]
\centering
\begin{tikzpicture}[scale=0.4]
\node[circle, draw=blue, fill=blue!15, text=blue, minimum size=0.8cm] (l) at (90:2.5) {$l$};
\node[circle, draw=blue, fill=blue!15, text=blue, minimum size=0.8cm] (m) at (162:2.5) {$m$};
\node[circle, draw=red, fill=red!15, text=red, minimum size=0.8cm] (i) at (234:2.5) {$i$};
\node[circle, draw=red, fill=red!15, text=red, minimum size=0.8cm] (j) at (306:2.5) {$j$};
\node[circle, draw=blue, fill=blue!15, text=blue, minimum size=0.8cm] (k) at (18:2.5) {$k$};
\draw (i) -- (k);
\draw (i) -- (l);
\draw (i) -- (m);
\draw (i) -- (j);
\draw (j) -- (k);
\draw (j) -- (l);
\draw (j) -- (m);
\end{tikzpicture}
\caption{Five-node pivot-peripheral configuration.}
\label{fig:five-node-triangles}
\end{figure}
We first derive a linear relation from one pivot--peripheral triangle and
then combine the three relations generated by \(k,l,m\). Throughout this
construction, probabilities are conditional on \((\bf X,\bs\Gamma)\) and
evaluated at \(\bs{\beta}_{0}\). Write \(\alpha_u:=e^{-\Gamma_u}\) for a
node \(u\), and write
\(W_{uv}(\bs{\beta}_{0}):=\exp(-\bs X_{uv}^\top\bs{\beta}_{0})\). For any dyad
\((u,v)\) among these five nodes, write
\(P_{uv}:=P_{uv}(\bs{\beta}_{0})\) for the conditional link probability
defined in Section~\ref{sec:model}.
Under Assumption~\ref{ass:dyad_ind}, the product-logit probability satisfies
\[
\begin{aligned}
P_{uv}
&=\frac{1}
{(1+W_{uv}(\bs{\beta}_{0})\alpha_u)
(1+W_{vu}(\bs{\beta}_{0})\alpha_v)},\\
\frac{1-P_{uv}}{P_{uv}} &=W_{uv}(\bs{\beta}_{0})W_{vu}(\bs{\beta}_{0})\alpha_u\alpha_v
+W_{uv}(\bs{\beta}_{0})\alpha_u
+W_{vu}(\bs{\beta}_{0})\alpha_v.
\end{aligned}
\]
For the triangle \(\{i,j,k\}\), multiplying \((1-P_{uv})/P_{uv}\) by
\(P_{uv}\) for \((u,v)\in\{(i,k),(j,k),(i,j)\}\) gives
\begin{equation}
\label{eq:triangle-p-identities}
\begin{aligned}
1-P_{ik}
&=P_{ik}W_{ik}(\bs{\beta}_{0})W_{ki}(\bs{\beta}_{0})\alpha_i\alpha_k
+P_{ik}W_{ik}(\bs{\beta}_{0})\alpha_i
+P_{ik}W_{ki}(\bs{\beta}_{0})\alpha_k,\\
1-P_{jk}
&=P_{jk}W_{jk}(\bs{\beta}_{0})W_{kj}(\bs{\beta}_{0})\alpha_j\alpha_k
+P_{jk}W_{jk}(\bs{\beta}_{0})\alpha_j
+P_{jk}W_{kj}(\bs{\beta}_{0})\alpha_k,\\
1-P_{ij}
&=P_{ij}W_{ij}(\bs{\beta}_{0})W_{ji}(\bs{\beta}_{0})\alpha_i\alpha_j
+P_{ij}W_{ij}(\bs{\beta}_{0})\alpha_i
+P_{ij}W_{ji}(\bs{\beta}_{0})\alpha_j.
\end{aligned}
\end{equation}
Solving the identities for dyads \((i,k)\) and \((j,k)\) in
\eqref{eq:triangle-p-identities} for the peripheral fixed effect
\(\alpha_k\) and equating the two expressions eliminates \(\alpha_k\).
Substituting the identity for dyad \((i,j)\) then eliminates the product
\(\alpha_i\alpha_j\). The resulting equation is linear in
\(P_{ij}\alpha_i\), \(P_{ij}\alpha_j\), and a constant term.
The following lemma records the resulting relation for any
pivot--peripheral triangle.
\begin{lem}
\label{lem:lem1}
Suppose Assumption~\ref{ass:dyad_ind} holds. For any distinct nodes
\(i,j,r\), there exist coefficients \(b_{ijr}(\bs{\beta}_{0})\),
\(c_{ijr}(\bs{\beta}_{0})\), and \(d_{ijr}(\bs{\beta}_{0})\), given
explicitly in \eqref{eq:triangle-coefficients-app}, such that
\(
b_{ijr}(\bs{\beta}_{0})P_{ij}\alpha_i
+c_{ijr}(\bs{\beta}_{0})P_{ij}\alpha_j
+d_{ijr}(\bs{\beta}_{0})=0.
\)
\end{lem}
Applying Lemma~\ref{lem:lem1} to \(r=k,l,m\) gives three linear
relations in \(P_{ij}\alpha_i\), \(P_{ij}\alpha_j\), and a constant term.
Collecting the three relations gives
\[
\mathbf M_{ijklm}(\bs{\beta}_{0})
\begin{bmatrix}
P_{ij}\alpha_i\\
P_{ij}\alpha_j\\
1
\end{bmatrix}
=0,
\qquad
\mathbf M_{ijklm}(\bs{\beta}_{0}):=
\begin{bmatrix}
b_{ijk}(\bs{\beta}_{0}) & c_{ijk}(\bs{\beta}_{0}) & d_{ijk}(\bs{\beta}_{0})\\
b_{ijl}(\bs{\beta}_{0}) & c_{ijl}(\bs{\beta}_{0}) & d_{ijl}(\bs{\beta}_{0})\\
b_{ijm}(\bs{\beta}_{0}) & c_{ijm}(\bs{\beta}_{0}) & d_{ijm}(\bs{\beta}_{0})
\end{bmatrix}.
\]
Since \([P_{ij}\alpha_i,P_{ij}\alpha_j,1]^\top\) is nonzero,
\(\mathbf M_{ijklm}(\bs{\beta}_{0})\) is singular, so
\(\det(\mathbf M_{ijklm}(\bs{\beta}_{0}))=0\).
The equality \(\det(\mathbf M_{ijklm}(\bs{\beta}_{0}))=0\) is not yet a
feasible moment condition, since the entries of
\(\mathbf M_{ijklm}(\bs{\beta}_{0})\) still involve the unobserved
conditional link probabilities \(P_{uv}\).
For a generic parameter value \(\bs\beta\), define the feasible matrix
\(\widetilde{\mathbf M}_{ijklm}(\bs\beta)\) by the explicit entries in
\eqref{eq:feasible-coefficients-app} and \eqref{eq:feasible-matrix-app}.
At \(\bs{\beta}_{0}\), this construction replaces \(P_{uv}\) and
\(1-P_{uv}\) in \(\mathbf M_{ijklm}(\bs{\beta}_{0})\) with
\(L_{uv}\) and \(1-L_{uv}\), respectively.
To verify that this substitution preserves the conditional mean of the
determinant, Lemma~\ref{lem:det-row-supports} in
Appendix~\ref{app:detconstruction} shows that each dyad enters every
monomial in the determinant expansion at most once, so the expansion is
multilinear in the dyad probabilities.
Assumption~\ref{ass:dyad_ind} gives
\(\mathbb{E}[L_{uv}\mid\bf X,\bs\Gamma]=P_{uv}\) and, for distinct dyads \((u,v)\)
and \((a,b)\),
\(\mathbb{E}[L_{uv}L_{ab}\mid\bf X,\bs\Gamma]=P_{uv}P_{ab}\). More generally, the
same factorization holds for finite products over distinct dyads; the
corresponding identities for \(1-L_{uv}\) follow by linearity.
Consequently,
\[
\mathbb{E}\!\left[\det(\widetilde{\mathbf M}_{ijklm}(\bs{\beta}_{0}))
\mid \bf X,\bs\Gamma\right]
=\det(\mathbf M_{ijklm}(\bs{\beta}_{0}))
=0.
\]
The following theorem summarizes the resulting fixed-effect-free moment
restriction.
\begin{thm}[Fixed-Effect-Free Moment Restriction]\label{thm:moment_restriction}
For distinct nodes \(i,j,k,l,m\), define the pentad moment function
\(\psi_{ijklm}(\bs\beta):=\det(\widetilde{\mathbf M}_{ijklm}(\bs\beta))\).
Under Assumption~\ref{ass:dyad_ind}, the following conditional moment
restrictions hold at the true value \(\bs{\beta}_0\):
\begin{equation}
\label{eq:moment_cond-free}
\begin{aligned}
\mathbb{E}\Bigl[\psi_{ijklm}(\bs{\beta}_{0})
\,\big|\,\bf X,\bs\Gamma\Bigr]&=0,\quad
\mathbb{E}\Bigl[\psi_{ijklm}(\bs{\beta}_{0})
\,\big|\,\bf X\Bigr]&=0.
\end{aligned}
\end{equation}
\end{thm}
Theorem~\ref{thm:moment_restriction} gives an exact fixed-effect-free moment
restriction for every five-node configuration. The conditional restriction
in \eqref{eq:moment_cond-free} accommodates unrestricted support for
\(\bs\Gamma\) and arbitrary dependence between \(\bf X\) and
\(\bs\Gamma\). The pentad construction thus provides a new algebraic method
for obtaining fixed-effect-free moment restrictions in NTU network formation
models.
The key idea in our construction is to difference out fixed effects by
exploiting the singularity of a linear-system coefficient matrix. This idea
extends beyond the NTU model. As a simple illustration, consider the TU logit
model in \citet{graham2017econometric}, written in our notation as
\begin{equation}
\label{eq:tu-model}
L_{uv}
=\1\left\{\Gamma_u+\Gamma_v+\bs X_{uv}^{\top}\bs{\beta}_{0}
-\epsilon_{uv}\geq0\right\},
\qquad 1\leq u<v\leq N.
\end{equation}
Here \(\bs X_{uv}=\bs X_{vu}\) is a symmetric vector of observed dyadic
covariates and \(\epsilon_{uv}=\epsilon_{vu}\). Conditional on
\((\bf X,\bs\Gamma)\), the shocks \(\{\epsilon_{uv}:u<v\}\) are independent
and have the standard logistic cdf \(F\). The conditional link probability is
therefore
\(
P_{uv}:=\mathbb{P}(L_{uv}=1\mid\bf X,\bs\Gamma)
=F\!\left(\Gamma_u+\Gamma_v
+\bs X_{uv}^{\top}\bs{\beta}_{0}\right)
\).
Set \(\alpha_u:=e^{-\Gamma_u}\) and
\(W_{uv}(\bs{\beta}_{0}):=\exp(-\bs X_{uv}^\top\bs{\beta}_{0})\). Then
\(
(1-P_{uv})/P_{uv}
=W_{uv}(\bs{\beta}_{0})\alpha_u\alpha_v.
\)
Fixing the pivotal pair \((i,j)\), each peripheral node eliminates its own
fixed effect and yields a homogeneous linear equation in
\((\alpha_i,\alpha_j)^\top\). For example, the dyads \((i,k)\) and \((j,k)\) imply
\(\alpha_k=(1-P_{ik})/[P_{ik}W_{ik}(\bs{\beta}_{0})\alpha_i]\) and
\(\alpha_k=(1-P_{jk})/[P_{jk}W_{jk}(\bs{\beta}_{0})\alpha_j]\), respectively;
equating these expressions eliminates \(\alpha_k\) and produces the first row
of the system below. The two peripheral nodes \(k\) and \(l\) therefore give
\begin{equation}
\label{eq:tu-homogeneous-system}
\underbrace{\begin{bmatrix}
P_{ik}(1-P_{jk})W_{ik}(\bs{\beta}_{0}) &
-P_{jk}(1-P_{ik})W_{jk}(\bs{\beta}_{0})\\
P_{il}(1-P_{jl})W_{il}(\bs{\beta}_{0}) &
-P_{jl}(1-P_{il})W_{jl}(\bs{\beta}_{0})
\end{bmatrix}}_{\mathbf M_{ijkl}}
\underbrace{\begin{bmatrix}
\alpha_i\\[4pt]
\alpha_j
\end{bmatrix}}_{\bs\alpha}
=
\begin{bmatrix}
0\\[4pt]
0
\end{bmatrix}.
\end{equation}
Since \(\bs\alpha\) is nonzero, \(\mathbf M_{ijkl}\) is singular, implying
\[
P_{ik}P_{jl}(1-P_{jk})(1-P_{il})
W_{ik}(\bs{\beta}_{0})W_{jl}(\bs{\beta}_{0})
-P_{jk}P_{il}(1-P_{ik})(1-P_{jl})
W_{jk}(\bs{\beta}_{0})W_{il}(\bs{\beta}_{0})=0.
\]
Multiplying \(\det(\mathbf M_{ijkl})=0\) by
\(\exp\{(\bs X_{ik}+\bs X_{jk})^\top\bs{\beta}_{0}\}\) and using the
conditional independence of the four dyads gives
\(
\mathbb{E}[\psi_{ijkl}(\bf L,\bf X;\bs{\beta}_{0})
\mid\bf X,\bs\Gamma]=0,
\)
where
\[
\begin{aligned}
\psi_{ijkl}(\bf L,\bf X;\bs\beta)
&=\1\{L_{ik}=1,L_{il}=0,L_{jk}=0,L_{jl}=1\}
\exp\big((\bs X_{jk}-\bs X_{jl})^\top\bs\beta\big)\\
&\quad-\1\{L_{ik}=0,L_{il}=1,L_{jk}=1,L_{jl}=0\}
\exp\big((\bs X_{ik}-\bs X_{il})^\top\bs\beta\big).
\end{aligned}
\]
This construction uses the same four-node tetrad as
\citet{graham2017econometric}, but it produces a different moment function.
The TU tetrad also clarifies why our NTU construction requires one more node.
Both constructions fix the pivotal pair \((i,j)\), with each peripheral
node supplying one equation after its fixed effect is eliminated.
Under TU, these equations are homogeneous in
\((\alpha_i,\alpha_j)^\top\), so two peripheral nodes give the
\(2\times2\) system and tetrad restriction in
\eqref{eq:tu-homogeneous-system}. Under NTU, eliminating the peripheral
fixed effect and then using the pivotal-dyad identity in
\eqref{eq:triangle-p-identities} to remove \(\alpha_i\alpha_j\) leaves
an equation in \(P_{ij}\alpha_i\) and \(P_{ij}\alpha_j\) with a constant
term; see Lemma~\ref{lem:lem1}. Two such equations generally determine
the two nuisance terms. Substituting these values into the equation from
a third peripheral node yields a restriction that no longer contains
either nuisance term.
Thus, augmenting the two nuisance terms with the constant coordinate \(1\)
converts the three affine equations into a \(3\times3\) homogeneous system
in \((P_{ij}\alpha_i,P_{ij}\alpha_j,1)^\top\), giving the
five-node determinant restriction in
Theorem~\ref{thm:moment_restriction}.
We next show that our pentad construction is \emph{minimal} in the sense that it uses
the minimum possible numbers of dyads and nodes for a nontrivial
fixed-effect-free moment and, among five-node, seven-dyad constructions, is
unique up to relabeling and scale. To state this result, we first introduce
notation for moments based on an arbitrary finite graph.
For a finite simple graph \(G=(V,E)\) with node set \(V\subseteq[N]\) and
dyad set \(E\), write
\(\mathbf L_G:=(L_{uv})_{(u,v)\in E}\) and
\(\mathbf X_G:=(\bs X_u)_{u\in V}\). For each such graph \(G\), consider a real-valued
measurable function
\(\psi_G(\mathbf L_G,\mathbf X_G;\bs\beta)\) such that
\begin{equation}
\mathbb{E}\!\left[\psi_G(\mathbf L_G,\mathbf X_G;\bs\beta_0)
\mid\bf X,\bs\Gamma\right]=0\quad\text{almost surely.}
\label{eq:graph-local-moment}
\end{equation}
In this notation, the pentad moment in
Theorem~\ref{thm:moment_restriction} corresponds to
\(V=\{i,j,k,l,m\}\),
\(E=\{(i,j),(i,k),(j,k),(i,l),(j,l),(i,m),(j,m)\}\).
\begin{thm}[Minimality of the Pentad Moment]
\label{thm:pentad_minimality}
Recall that \(\bs X_{ij}=w(\bs X_i,\bs X_j)\), and suppose that
\(w(\bs x,\bs y)
=\bigl(w_1(x_1,y_1),\allowbreak\ldots,\allowbreak
w_d(x_d,y_d)\bigr)^\top\).
Each \(w_r:\mathbb R^2\to\mathbb R\) satisfies one of the following:
\textnormal{(i)} \(w_r\) is real analytic on
\(\mathbb R^2\), with \(\frac{\partial w_r}{\partial y_r}(x_r^\circ,x_r^\circ)\neq0\) for some
\(x_r^\circ\in\mathbb R\); or \textnormal{(ii)} there exist
\(\varepsilon_r>0\), \(\kappa_r\in\{1,2\}\), and real-analytic functions
\(f_{r,+},f_{r,-}:(-\varepsilon_r,\infty)\to\mathbb R\) such that
\(w_r(s,t)=f_{r,+}(t-s)\) for \(s\leq t\),
\(w_r(s,t)=f_{r,-}(s-t)\) for \(s>t\),
\(f_{r,+}(0)=f_{r,-}(0)\),
\(f_{r,+}^{(\kappa_r)}(0)f_{r,-}^{(\kappa_r)}(0)\neq0\),
and, if \(\kappa_r=2\), \(f'_{r,+}(0)=f'_{r,-}(0)=0\).
For any \(\bs\beta_0\in\mathbb R^d\setminus\{\mathbf 0\}\) and any
data-generating process satisfying Assumption~\ref{ass:dyad_ind} with an
absolutely continuous distribution of \((\bs X_i,\Gamma_i)\), the
following statements hold for every finite simple graph \(G=(V,E)\) and every measurable \(\psi_G\) satisfying
\eqref{eq:graph-local-moment}:
\begin{enumerate}[label=\textnormal{(\alph*)},leftmargin=2em]
\item If \(\lvert E\rvert<7\) or \(\lvert V\rvert<5\), then
\(\psi_G\) is trivial, i.e., \(\psi_G(\cdot,\mathbf X_G;\bs\beta_0)\equiv0\) almost surely.
\item Suppose \(\lvert V\rvert=5\) and \(\lvert E\rvert=7\). Then
\(\psi_G\) is nontrivial only if there exist distinct
\(i,j,k,l,m\in V\) such that
\(E=\{(i,j),(i,k),(j,k),\allowbreak(i,l),(j,l),\allowbreak
(i,m),(j,m)\}\). For each such ordered tuple \((i,j,k,l,m)\),
\(\psi_G(\mathbf L_G,\mathbf X_G;\bs\beta_0)
\allowbreak=C(\mathbf X_G)\psi_{ijklm}(\bs\beta_0)\) almost surely for
some measurable function \(C\), where
\(\psi_{ijklm}(\bs\beta_0)\) denotes the pentad moment function defined in
Theorem~\ref{thm:moment_restriction}; moreover,
\(\psi_{ijklm}(\bs\beta_0)\) is nontrivial.
\end{enumerate}
\end{thm}
Theorem~\ref{thm:pentad_minimality} establishes that our pentad construction
is minimal in both dyads and nodes for a broad class of covariate
specifications. The restrictions on \(w\) are convenient sufficient
conditions for minimality, but are not necessary. Their role is to exclude
special algebraic degeneracies among the dyadic covariates \(\bs X_{ij}\).
We also exclude the degenerate case \(\bs\beta_0=\mathbf0\), in which all
covariate indices \(\bs X_{ij}^{\top}\bs\beta_0\) vanish.
The conditions on \(w\) cover symmetric maps such as
\(w_r(x_r,y_r)=\lvert x_r-y_r\rvert\) and
\(w_r(x_r,y_r)=\lvert x_r-y_r\rvert^2\), as well as asymmetric maps such as
\(w_r(x_r,y_r)=y_r\) and
\(w_r(x_r,y_r)=2\lvert x_r-y_r\rvert\1\{x_r>y_r\}
\allowbreak+\lvert x_r-y_r\rvert\1\{x_r\leq y_r\}\).
The minimality conclusion may also hold for other choices of \(w\),
but the proof of Theorem~\ref{thm:pentad_minimality} needs to be checked
for each such choice.
\subsection{Pentad-GMM Estimator}\label{sec:pentad_gmm}
We estimate \(\bs{\beta}_{0}\) by GMM using the fixed-effect-free pentad
moment restrictions in \eqref{eq:moment_cond-free}. The sample
moments used below keep the five roles \((i,j,k,l,m)\) ordered. Define
\(\mathcal S_N:=\{(i,j,k,l,m)\in[N]^5:i,j,k,l,m\text{ are all distinct}\}\);
then \(\left\lvert \mathcal S_N \right\rvert=N(N-1)(N-2)(N-3)(N-4)\).
For \(s=(i,j,k,l,m)\in\mathcal S_N\), let
\(\psi_s(\bs\beta):=\psi_{ijklm}(\bs\beta)\). To obtain enough moment
conditions for a multidimensional parameter, let $q\ge d$ be fixed and let
\(\bs Z_s:=z(\bf X_s)\in\mathbb{R}^{q}\) be an instrument
vector constructed from the observed
covariates in configuration $s=(i,j,k,l,m)$, where $\bf X_s$ collects
the node covariates in the order \((i,j,k,l,m)\).
The instrument may depend on the ordered roles, including the
order of the peripheral labels \(k,l,m\). Define the pentad moment function
\(\bs\Psi_s(\bs{\beta}):= \bs Z_s\psi_s(\bs{\beta})\) and the
population moment \(\bs h(\bs{\beta}):=\mathbb{E}[\bs\Psi_s(\bs{\beta})]\). The
fixed-effect-free conditional moment restriction \eqref{eq:moment_cond-free}
and the law of iterated expectation imply
\(
\bs h(\bs{\beta}_{0})
=\mathbb{E}\!\left[\bs Z_s\,\psi_s(\bs{\beta}_{0})\right]
=0.
\)
The sample moment vector is then given by
\begin{equation}
\bs h_{N}(\bs{\beta})\;:=\;\frac{1}{\left\lvert \mathcal S_N \right\rvert}\sum_{s\in\mathcal S_N}\bs\Psi_s(\bs{\beta})\;\in\;\mathbb{R}^{q}.\label{eq:sample_moment-multi}
\end{equation}
The pentad-GMM estimator for \(\bs\beta_0\) is then defined as the minimizer:
\begin{equation}
\widehat{\bs{\beta}}_{\mathrm{GMM}}\in
\arg\min_{\bs{\beta}\in\mathcal B}
\bs h_N(\bs\beta)^\top\widehat{\bf W}_N\bs h_N(\bs\beta),
\label{eq:beta_gmm}
\end{equation}
where \(\widehat{\bf W}_{N}\in\mathbb R^{q\times q}\) is a feasible
symmetric positive definite weight matrix, and the parameter space
\(\mathcal B\subset\mathbb R^d\) is compact, excludes the origin
\(\mathbf 0\notin\mathcal B\), and satisfies
\(\bs{\beta}_{0}\in\operatorname{int}(\mathcal B)\),
where \(\operatorname{int}(\mathcal B)\) denotes the interior of \(\mathcal B\). We exclude the origin
because \(\psi_s(\mathbf 0)=0\) for every \(s\in\mathcal S_N\) and every
realization of the network. Hence,
\(\bs h_N(\mathbf 0)=\bs h(\mathbf 0)=0\), making \(\mathbf 0\) a degenerate
root of both the sample and population moment equations. The GMM objective in
\eqref{eq:beta_gmm} therefore always equals zero at the origin.\footnote{The
need to exclude the origin arises from the pentad-GMM criterion rather than
from the underlying network model.
When \(\bs\beta_0=\mathbf 0\), Assumption~\ref{ass:dyad_ind} implies that the
conditional link probability factors as
\(\mathbb{P}(L_{ij}=1\mid\bs X_i,\bs X_j)
=\mathbb{E}[F(\Gamma_i)\mid\bs X_i]\mathbb{E}[F(\Gamma_j)\mid\bs X_j]\).
This factorization provides a testable implication for examining the zero
coefficient vector separately. The distinction between zero and nonzero
coefficient vectors also arises in \citet{davezies2023fixed}, who formally
characterize the zero vector in binary choice panels through a different
conditional probability restriction.}
\section{Asymptotic Analysis}
\label{sec:asymptotics}
In this section, we derive asymptotic properties for the sample moment $\bs{h}_N(\bs \beta_0)$ and for the pentad-GMM estimator
\(\widehat{\bs{\beta}}_{\mathrm{GMM}}\) based on
\eqref{eq:sample_moment-multi}. We begin with basic regularity conditions for
the asymptotic analysis. The first imposes boundedness on the dyadic
design and instruments. Similar restrictions on covariates and
instruments are common in the network-formation literature;
see, for example, \citet{graham2017econometric},
\citet{gao2023logical}, and \citet{li2024estimation}.
\begin{assumption}[Bounded dyadic design and instruments]
\label{ass:bounded_design}
The parameter dimension \(d\) and instrument dimension \(q\) are fixed.
The induced dyadic covariates and instruments satisfy
\(\sup_{i\neq j}\left\lVert \bs X_{ij} \right\rVert\le C_X\) and
\(\sup_{s\in\mathcal S_N}\left\lVert \bs Z_s \right\rVert\le C_Z\) a.s. for
deterministic constants \(C_X,C_Z<\infty\).
\end{assumption}
Throughout this paper, let $\mathcal{E}_{N}:=\{(u,v):1\le u<v\le N\}$ denote the set of all unordered dyads in the network.
For an unordered dyad $e=(a,b)\in\mathcal{E}_{N}$, recall
that $P_{e}(\bs{\beta}_{0}):=\mathbb{E}[L_{e}\mid\bf X,\bs\Gamma]$
is the conditional link probability.
\begin{assumption}[Average link probability and relative expected degree]
\label{ass:sparse_envelope}
Let \(\rho_N:=\mathbb{E}[P_{12}(\bs\beta_0)]\). For each \(i\in[N]\), define
individual \(i\)'s relative expected degree by
\(\Lambda_{i,N}:=\rho_N^{-1}\mathbb{E}[P_{ij}(\bs\beta_0)\mid \bs X_i,\Gamma_i]\),
where \(j\ne i\) is arbitrary.
Assume that \(\sup_N\mathbb{E}[\Lambda_{i,N}^{16}]<\infty\).
\end{assumption}
Assumption~\ref{ass:sparse_envelope} defines \(\rho_N\) as the average
link probability and \(\Lambda_{i,N}\) as individual \(i\)'s relative
expected degree, namely, the expected degree conditional on
\(\bs X_i,\Gamma_i\) divided by the average expected degree: indeed,
\(\mathbb{E}[\Lambda_{i,N}]=1\) and
\(\mathbb{E}[\sum_{j\ne i}L_{ij}\mid \bs X_i,\Gamma_i]
=(N-1)\rho_N\Lambda_{i,N}\). The moment restriction
\(\sup_N\mathbb{E}[\Lambda_{i,N}^{16}]<\infty\) permits some individuals to have
expected degrees far above the average, provided that individuals with such
large relative expected degrees are sufficiently rare. This
restriction is the single-population counterpart of the third-moment
conditions used by \citet[Assumption~2(iii)]{graham2024sparse} to rule
out extreme skewness in the consumer and product degree distributions.
Although the moment restriction \(\sup_N\mathbb{E}[\Lambda_{i,N}^{16}]<\infty\)
restricts the tail behavior of the relative expected degree, it does not
restrict the asymptotic scale of the average expected degree
\((N-1)\rho_N\).
Since \(\rho_N\) is the average link probability, its
asymptotic scale determines network sparsity and the growth of the average
expected degree: a nonvanishing \(\rho_N\) yields a dense network sequence,
whereas \(\rho_N\to0\) yields a sparse sequence whose degree of sparsity
depends on how quickly \(\rho_N\) vanishes (see
\citealp{graham2020network,graham2024sparse}). Our analysis covers three sparsity
regimes: dense or mildly sparse networks with
\(N\rho_N\to\infty\); sparse networks with finite average expected degree,
\(N\rho_N\asymp1\); and ultra-sparse networks with
\(N\rho_N\to0\) and \(N^5\rho_N^4\to\infty\).
We organize the asymptotic analysis through the projection structure of
the sample moment. Conditional on \(\bf X\) and \(\bs\Gamma\), applying the
Hoeffding decomposition over dyads (obtained by Möbius inversion on the
lattice of dyad subsets; see \citealp{rota1964foundations}) to
\(\bs h_N(\bs{\beta}_{0})\) gives
\begin{equation}
\label{eq:main_projection_decomp}
\bs h_N(\bs{\beta}_{0})
=\sum_{A\subseteq\mathcal E_N:1\le\left\lvert A \right\rvert\le7}\bs\phi_{A,N},
\quad
\bs\phi_{A,N}:=\sum_{B\subseteq A}(-1)^{\left\lvert A \right\rvert-\left\lvert B \right\rvert}
\mathbb{E}[\bs h_N(\bs{\beta}_{0})\mid \bf X,\bs\Gamma,\{L_e\}_{e\in B}].
\end{equation}
Here \(A\) is a labeled dyad set. For
\(s=(i,j,k,l,m)\in\mathcal S_N\), write
\[
E(s):=\{(i,j),(i,k),(i,l),(i,m),(j,k),(j,l),(j,m)\}
\]
for the seven dyads entering the five-node configuration. The five-node
kernel \(\bs\Psi_s(\bs{\beta}_{0})\) admits the finite polynomial representation
\begin{equation}
\label{eq:kernel_poly_rep}
\bs\Psi_s(\bs{\beta}_{0})
=\sum_{\nu=1}^{\nu_{\max}}\bs c_{s,\nu}(\bf X)m_{s,\nu}({\bf L}),
\end{equation}
Here the constant \(\nu_{\max}<\infty\) does not depend on \(N\),
\(\bs c_{s,\nu}(\bf X)\in\mathbb R^q\) is a coefficient vector, and
\(m_{s,\nu}({\bf L})=\prod_{e\in E_{s,\nu}}L_e\) is a square-free monomial
with support \(E_{s,\nu}\subseteq E(s)\); see Lemma~\ref{lem:poly}. Applying the
projection formula in
\eqref{eq:main_projection_decomp} term by term to
\eqref{eq:kernel_poly_rep} gives
\[
\bs\phi_{A,N}
=\left\lvert \mathcal S_N \right\rvert^{-1}\bs C_{A,N}(\bf X,\bs\Gamma)\prod_{e\in A}\xi_e,
\]
where $\xi_{e}:=L_{e}-P_{e}(\bs{\beta}_{0})$
is the centered link indicator and
\[
\bs C_{A,N}(\bf X,\bs\Gamma)
:=
\sum_{s\in\mathcal S_N:\,A\subseteq E(s)}
\sum_{\nu:A\subseteq E_{s,\nu}}
\bs c_{s,\nu}(\bf X)
\prod_{f\in E_{s,\nu}\setminus A}P_f(\bs{\beta}_{0}).
\]
Thus the dyads in \(A\) enter the projection through the product of the
centered link indicators \(\{\xi_e:e\in A\}\). All other dyads appearing
in the same square-free monomial are replaced by their conditional link
probabilities. Since \(\{\xi_e:e\in\mathcal E_N\}\) are conditionally
independent, distinct dyad-set projections are conditionally orthogonal:
\[
\mathbb{E}[\bs\phi_{A,N}\bs\phi_{A',N}^{\top}\mid\bf X,\bs\Gamma]=0
\qquad\text{whenever }A\neq A'.
\]
Together with \eqref{eq:main_projection_decomp}, this orthogonality
implies that the conditional covariance, and hence the variance scale,
of \(\bs h_N(\bs{\beta}_{0})\) is obtained by adding the covariance
contributions of the dyad-set projections \(\bs\phi_{A,N}\). We therefore
work with the dyad-set projections rather than the five-node summands
directly. Since the variance order of \(\bs\phi_{A,N}\) depends on the
unlabeled graph shape represented by \(A\), rather than by the
particular node labels, we group the projections according to that
shape.
For a finite dyad set \(A\subseteq\mathcal E_N\), let
\(V(A)\) be the set of nodes that appear in the dyads in \(A\), and let
\(G_A:=(V(A),A)\) be the corresponding labeled graph.
Let \([G_A]\) denote its isomorphism class, equivalently the unlabeled
graph shape determined by the dyads present among the nodes in \(A\).
Following the graph-isomorphism terminology in \citet{graham2020network},
each such isomorphism class can be used as an element of a graph class:
two dyad sets \(A\) and \(A'\) have the same isomorphism class if and only if
\(G_A\) and \(G_{A'}\) are isomorphic (see Appendix~\ref{app:graph_notation}).
For example, all singleton dyad sets are labeled copies of the one-dyad
class, and all two-edge paths are labeled copies of the two-star
class.
More generally, we use the term \emph{graph class} for a finite
collection of such unlabeled graph shapes. For such a class \(g\), let
\(\mathscr C_N(g):=\{A\subseteq\mathcal E_N:[G_A]\in g\}\) denote the
collection of all labeled dyad sets whose unlabeled isomorphism class is
one of the elements of \(g\). For example, the graph class
\(g_{\mathrm{3tree}}\) contains the three-edge path
(Figure~\ref{fig:projection_classes_main}(c)) and the three-star
(Figure~\ref{fig:projection_classes_main}(d)), so
\(\mathscr C_N(g_{\mathrm{3tree}})\) collects all labeled dyad sets with
either shape. For such a graph class \(g\), define the projection sum
\(
\bs\Pi_N(g):=\sum_{A\in\mathscr C_N(g)}\bs\phi_{A,N},
\)
which collects the projection terms whose dyad sets have a shape in \(g\).
The graph classes that can arise from the five-node determinant kernel
are characterized in Lemma~\ref{lem:kerneltaxonomy} and summarized in
Table~\ref{tab:regimemap}.
We now define four graph classes whose projection sums can be asymptotically
leading under the sparsity regimes considered in this paper.
The one-dyad class is
\(g_{\mathrm{dyad}}:=\{[G_{\{(1,2)\}}]\},\)
so that \(\bs\Pi_N(g_{\mathrm{dyad}})\) is the one-dyad projection sum.
The two-star class is
\(g_{\text{2-star}}:=\{[G_{\{(1,2),(1,3)\}}]\},\)
the two-edge path on three nodes. We also use two combined graph
classes. Let
\(g_{\mathrm{3tree}}
:=
\{
[G_{\{(1,2),(2,3),(3,4)\}}],
[G_{\{(1,2),(1,3),(1,4)\}}]
\big\},\)
which collects the three-edge path on four nodes and the three-star on
four nodes. Let
\(g_{\mathrm{4tree}}
:=
\big\{
[G_{\{(1,2),(2,3),(3,4),(4,5)\}}],
[G_{\{(1,2),(1,3),(1,4),(2,5)\}}]
\big\},\)
which collects the four-edge path on five nodes and the fork-shaped
four-edge tree with degree sequence \((3,2,1,1,1)\). The associated
projection sums are
\(\bs\Pi_N(g_{\mathrm{dyad}})\), \(\bs\Pi_N(g_{\text{2-star}})\),
\(\bs\Pi_N(g_{\mathrm{3tree}})\), and \(\bs\Pi_N(g_{\mathrm{4tree}})\), each
defined by summing \(\bs\phi_{A,N}\) over the labeled copies in the
corresponding class. Figure~\ref{fig:projection_classes_main} provides a
graphical illustration of these classes.
\begin{figure}[!htbp]
\centering
\captionsetup{font=small}
\captionsetup[subfigure]{font=small}
\begin{subfigure}[t]{0.30\textwidth}
\centering
\begin{tikzpicture}[scale=0.60,graphNode/.style={circle,fill=black,inner sep=1.35pt},graphEdge/.style={line width=0.6pt}]
\node[graphNode] (d1) at (0,0) {};
\node[graphNode] (d2) at (1.6,0) {};
\draw[graphEdge] (d1)--(d2);
\end{tikzpicture}
\caption{\(g_{\mathrm{dyad}}\)}
\end{subfigure}
\begin{subfigure}[t]{0.30\textwidth}
\centering
\begin{tikzpicture}[scale=0.60,graphNode/.style={circle,fill=black,inner sep=1.35pt},graphEdge/.style={line width=0.6pt}]
\node[graphNode] (w1) at (0.80,0.75) {};
\node[graphNode] (w2) at (0,0) {};
\node[graphNode] (w3) at (1.60,0) {};
\draw[graphEdge] (w1)--(w2);
\draw[graphEdge] (w1)--(w3);
\end{tikzpicture}
\caption{\(g_{\text{2-star}}\)}
\end{subfigure}
\hfill
\begin{subfigure}[t]{0.30\textwidth}
\centering
\begin{tikzpicture}[scale=0.60,graphNode/.style={circle,fill=black,inner sep=1.35pt},graphEdge/.style={line width=0.6pt}]
\node[graphNode] (p1) at (0,0) {};
\node[graphNode] (p2) at (0.72,0) {};
\node[graphNode] (p3) at (1.44,0) {};
\node[graphNode] (p4) at (2.16,0) {};
\draw[graphEdge] (p1)--(p2)--(p3)--(p4);
\end{tikzpicture}
\caption{\(g_{\mathrm{3tree}}\): three-edge path}
\end{subfigure}
\vspace{0.9em}
\begin{subfigure}[t]{0.30\textwidth}
\centering
\begin{tikzpicture}[scale=0.60,graphNode/.style={circle,fill=black,inner sep=1.35pt},graphEdge/.style={line width=0.6pt}]
\node[graphNode] (s1) at (0.80,0.45) {};
\node[graphNode] (s2) at (0.80,1.10) {};
\node[graphNode] (s3) at (0,0) {};
\node[graphNode] (s4) at (1.60,0) {};
\draw[graphEdge] (s1)--(s2);
\draw[graphEdge] (s1)--(s3);
\draw[graphEdge] (s1)--(s4);
\end{tikzpicture}
\caption{\(g_{\mathrm{3tree}}\): three-star}
\end{subfigure}
\hfill
\begin{subfigure}[t]{0.30\textwidth}
\centering
\begin{tikzpicture}[scale=0.60,graphNode/.style={circle,fill=black,inner sep=1.35pt},graphEdge/.style={line width=0.6pt}]
\node[graphNode] (q1) at (0,0) {};
\node[graphNode] (q2) at (0.62,0) {};
\node[graphNode] (q3) at (1.24,0) {};
\node[graphNode] (q4) at (1.86,0) {};
\node[graphNode] (q5) at (2.48,0) {};
\draw[graphEdge] (q1)--(q2)--(q3)--(q4)--(q5);
\end{tikzpicture}
\caption{\(g_{\mathrm{4tree}}\): four-edge path}
\end{subfigure}
\hfill
\begin{subfigure}[t]{0.30\textwidth}
\centering
\begin{tikzpicture}[scale=0.60,graphNode/.style={circle,fill=black,inner sep=1.35pt},graphEdge/.style={line width=0.6pt}]
\node[graphNode] (t1) at (0.70,0.45) {};
\node[graphNode] (t2) at (1.35,0.45) {};
\node[graphNode] (t3) at (0.10,1.00) {};
\node[graphNode] (t4) at (0.10,-0.10) {};
\node[graphNode] (t5) at (1.95,0.45) {};
\draw[graphEdge] (t1)--(t2);
\draw[graphEdge] (t1)--(t3);
\draw[graphEdge] (t1)--(t4);
\draw[graphEdge] (t2)--(t5);
\end{tikzpicture}
\caption{\(g_{\mathrm{4tree}}\): fork-shaped four-edge tree}
\end{subfigure}
\caption{Graph classes whose projection sums can be asymptotically leading
under the sparsity regimes considered in this paper. Panels~(a) and~(b)
display \(g_{\mathrm{dyad}}\) and \(g_{\text{2-star}}\), respectively;
panels~(c) and~(d) display the two shapes in \(g_{\mathrm{3tree}}\); and
panels~(e) and~(f) display the two shapes in \(g_{\mathrm{4tree}}\).}
\label{fig:projection_classes_main}
\end{figure}
The projection construction above expresses
\(\bs h_N(\bs\beta_0)\) in terms of conditionally orthogonal projection
sums \(\bs\Pi_N(g)\). To derive the asymptotic distribution of
\(\bs h_N(\bs\beta_0)\), we first identify which of these sums form the
leading approximation by comparing their second-moment orders.
\subsection{Regime-Specific Asymptotics of the Sample Moment}
We partition the unlabeled graph shapes arising from dyad subsets
\(A\subseteq E_{s,\nu}\subseteq E(s)\) into graph classes according to
the features that determine the second-moment orders of their projection
sums. Two shapes
belong to the same graph class exactly when they
have the same number of dyads, the same number of distinct nodes, and
the same minimum of \(\left\lvert E_{s,\nu} \right\rvert\) over the square-free monomials
\(m_{s,\nu}({\bf L})\) in \eqref{eq:kernel_poly_rep} for which the
support \(E_{s,\nu}\) contains a labeled copy of the shape. Fix one such
graph class \(g\), and let \(r(g)\) and \(v(g)\) denote, respectively,
the common number of dyads and the common number of distinct nodes among
its shapes. Define the common minimum degree of a square-free monomial
whose support contains a labeled copy of a shape in \(g\) by
\(
m(g):=\min\{\left\lvert E_{s,\nu} \right\rvert:\exists A\subseteq E_{s,\nu}
\text{ with }[G_A]\in g\},
\)
where the minimum ranges over the square-free monomials
\(m_{s,\nu}({\bf L})\) in \eqref{eq:kernel_poly_rep}.
Proposition~\ref{prop:classbound} in
Appendix~\ref{sec:regime_complete_asymptotic} shows that
\(\mathbb{E}\|\bs\Pi_N(g)\|^2
=O\big(N^{-v(g)}\rho_N^{2m(g)-r(g)}\big)\).
Table~\ref{tab:regimemap} lists all graph classes $g$ arising from dyad
subsets \(A\subseteq E_{s,\nu}\) and the corresponding
values of \(r(g)\), \(v(g)\), \(m(g)\), and
\(N^{-v(g)}\rho_N^{2m(g)-r(g)}\).
\begin{table}[H]
\centering
\caption{Complete list of graph classes and variance scales induced by dyad
subsets \(A\subseteq E_{s,\nu}\) in the projection decomposition. Here
\(E_{s,\nu}\) is the dyad support of the square-free monomial
\(m_{s,\nu}({\bf L})\) in \eqref{eq:kernel_poly_rep}.}
\label{tab:regimemap}
{\scriptsize
\setlength{\tabcolsep}{2.2pt}
\begin{tabularx}{\textwidth}{@{}>{\raggedright\arraybackslash}m{2.05cm}*{7}{>{\centering\arraybackslash}X}@{}}
\toprule
Graph class \(g\) &
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (1,0) {}; \draw[taxEdge] (a)--(b);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.7,.65) {}; \node[taxNode] (c) at (1.4,0) {}; \draw[taxEdge] (a)--(b)--(c);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,.35) {}; \node[taxNode] (b) at (.9,.35) {}; \node[taxNode] (c) at (0,-.35) {}; \node[taxNode] (d) at (.9,-.35) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (c)--(d);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,.75) {}; \node[taxNode] (b) at (-.65,0) {}; \node[taxNode] (c) at (.65,0) {}; \draw[taxEdge] (a)--(b)--(c)--(a);
\end{tikzpicture}
&
\mbox{
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.7,0) {}; \node[taxNode] (c) at (1.4,0) {}; \node[taxNode] (d) at (2.1,0) {}; \draw[taxEdge] (a)--(b)--(c)--(d);
\end{tikzpicture}
\hspace{0.18em}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (0,.8) {}; \node[taxNode] (c) at (-.7,-.45) {}; \node[taxNode] (d) at (.7,-.45) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d);
\end{tikzpicture}
} &
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,.35) {}; \node[taxNode] (b) at (.55,.75) {}; \node[taxNode] (c) at (1.1,.35) {}; \node[taxNode] (d) at (.1,-.35) {}; \node[taxNode] (e) at (1,-.35) {}; \draw[taxEdge] (a)--(b)--(c); \draw[taxEdge] (d)--(e);
\end{tikzpicture}
&
\mbox{
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (-.55,-.45) {}; \node[taxNode] (b) at (.55,.45) {}; \node[taxNode] (c) at (-.55,.45) {}; \node[taxNode] (d) at (.55,-.45) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (b)--(c); \draw[taxEdge] (c)--(d); \draw[taxEdge] (d)--(a);
\end{tikzpicture}
\hspace{0.18em}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,.75) {}; \node[taxNode] (b) at (-.65,0) {}; \node[taxNode] (c) at (.65,0) {}; \node[taxNode] (d) at (1.35,0) {}; \draw[taxEdge] (a)--(b)--(c)--(a); \draw[taxEdge] (c)--(d);
\end{tikzpicture}
} \\
\midrule
\(r(g)\) & \(1\) & \(2\) & \(2\) & \(3\) & \(3\) & \(3\) & \(4\) \\
\(v(g)\) & \(2\) & \(3\) & \(4\) & \(3\) & \(4\) & \(5\) & \(4\) \\
\(m(g)\) & \(4\) & \(4\) & \(4\) & \(5\) & \(4\) & \(4\) & \(5\) \\
variance scale & \(N^{-2}\rho_N^7\) & \(N^{-3}\rho_N^6\) &
\(N^{-4}\rho_N^6\) & \(N^{-3}\rho_N^7\) & \(N^{-4}\rho_N^5\) &
\(N^{-5}\rho_N^5\) & \(N^{-4}\rho_N^6\) \\
\bottomrule
\end{tabularx}
\vspace{1.2ex}
\begin{tabularx}{\textwidth}{@{}>{\raggedright\arraybackslash}m{2.05cm}*{6}{>{\centering\arraybackslash}X}@{}}
\toprule
Graph class \(g\) &
\mbox{
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.55,0) {}; \node[taxNode] (c) at (1.1,0) {}; \node[taxNode] (d) at (1.65,0) {}; \node[taxNode] (e) at (2.2,0) {}; \draw[taxEdge] (a)--(b)--(c)--(d)--(e);
\end{tikzpicture}
\hspace{0.18em}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.65,0) {}; \node[taxNode] (c) at (-.55,.55) {}; \node[taxNode] (d) at (-.55,-.55) {}; \node[taxNode] (e) at (1.25,0) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d); \draw[taxEdge] (b)--(e);
\end{tikzpicture}
} &
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (0,.85) {}; \node[taxNode] (c) at (0,-.85) {}; \node[taxNode] (d) at (-.85,0) {}; \node[taxNode] (e) at (.85,0) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d); \draw[taxEdge] (a)--(e);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (-.75,0) {}; \node[taxNode] (b) at (0,.65) {}; \node[taxNode] (c) at (.75,0) {}; \node[taxNode] (d) at (0,-.65) {}; \draw[taxEdge] (a)--(b)--(c)--(d)--(a); \draw[taxEdge] (a)--(c);
\end{tikzpicture}
&
\mbox{
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (-.55,-.45) {}; \node[taxNode] (b) at (.55,.45) {}; \node[taxNode] (c) at (-.55,.45) {}; \node[taxNode] (d) at (.55,-.45) {}; \node[taxNode] (e) at (1.2,-.45) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (b)--(c); \draw[taxEdge] (c)--(d); \draw[taxEdge] (d)--(a); \draw[taxEdge] (d)--(e);
\end{tikzpicture}
\hspace{0.18em}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,.65) {}; \node[taxNode] (b) at (-.55,0) {}; \node[taxNode] (c) at (.55,0) {}; \node[taxNode] (d) at (-.45,1.2) {}; \node[taxNode] (e) at (.45,1.2) {}; \draw[taxEdge] (a)--(b)--(c)--(a); \draw[taxEdge] (a)--(d); \draw[taxEdge] (a)--(e);
\end{tikzpicture}
\hspace{0.18em}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (0,.65) {}; \node[taxNode] (b) at (-.55,0) {}; \node[taxNode] (c) at (.55,0) {}; \node[taxNode] (d) at (-1.1,0) {}; \node[taxNode] (e) at (1.1,0) {}; \draw[taxEdge] (a)--(b)--(c)--(a); \draw[taxEdge] (b)--(d); \draw[taxEdge] (c)--(e);
\end{tikzpicture}
} &
\mbox{
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (-.45,-.45) {}; \node[taxNode] (b) at (.45,-.45) {}; \node[taxNode] (c) at (-.9,.55) {}; \node[taxNode] (d) at (0,.9) {}; \node[taxNode] (e) at (.9,.55) {}; \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d); \draw[taxEdge] (a)--(e); \draw[taxEdge] (b)--(c); \draw[taxEdge] (b)--(d); \draw[taxEdge] (b)--(e);
\end{tikzpicture}
\hspace{0.18em}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (-.65,0) {}; \node[taxNode] (b) at (0,.6) {}; \node[taxNode] (c) at (.65,0) {}; \node[taxNode] (d) at (0,-.6) {}; \node[taxNode] (e) at (-1.15,0) {}; \draw[taxEdge] (a)--(b)--(c)--(d)--(a); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(e);
\end{tikzpicture}
} &
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
]
\node[taxNode] (a) at (-.45,-.45) {}; \node[taxNode] (b) at (.45,-.45) {}; \node[taxNode] (c) at (-.9,.55) {}; \node[taxNode] (d) at (0,.9) {}; \node[taxNode] (e) at (.9,.55) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d); \draw[taxEdge] (a)--(e); \draw[taxEdge] (b)--(c); \draw[taxEdge] (b)--(d); \draw[taxEdge] (b)--(e);
\end{tikzpicture}
\\
\midrule
\(r(g)\) & \(4\) & \(4\) & \(5\) & \(5\) & \(6\) & \(7\) \\
\(v(g)\) & \(5\) & \(5\) & \(4\) & \(5\) & \(5\) & \(5\) \\
\(m(g)\) & \(4\) & \(5\) & \(6\) & \(5\) & \(6\) & \(7\) \\
variance scale & \(N^{-5}\rho_N^4\) & \(N^{-5}\rho_N^6\) &
\(N^{-4}\rho_N^7\) & \(N^{-5}\rho_N^5\) & \(N^{-5}\rho_N^6\) &
\(N^{-5}\rho_N^7\) \\
\bottomrule
\end{tabularx}
}
\end{table}
The variance scales in Table~\ref{tab:regimemap} show that only the four
graph classes displayed in Figure~\ref{fig:projection_classes_main} can
enter a leading approximation in at least one of the three sparsity regimes.
Using the one-dyad scale
\(N^{-2}\rho_N^7\) for \(\bs\Pi_N(g_{\mathrm{dyad}})\) as the benchmark,
the scales for
\(\bs\Pi_N(g_{\text{2-star}})\), \(\bs\Pi_N(g_{\mathrm{3tree}})\), and
\(\bs\Pi_N(g_{\mathrm{4tree}})\) are, respectively,
\((N\rho_N)^{-1}\), \((N\rho_N)^{-2}\), and
\((N\rho_N)^{-3}\) times the benchmark. Thus, when
\(N\rho_N\to\infty\), only
the one-dyad projection is retained in the leading approximation; when
\(\rho_N\asymp N^{-1}\), all four variance scales are of order
\(N^{-9}\), so all four projections are retained; and when
\(N\rho_N\to0\), the four-tree scale dominates the other three.
Accordingly, define the leading projection sums for the three cases by
\(\bs\Pi_N^{\mathrm{D}}:=\bs\Pi_N(g_{\mathrm{dyad}})\),
\(\bs\Pi_N^{\mathrm{S}}:=\bs\Pi_N(g_{\mathrm{dyad}})
+\bs\Pi_N(g_{\text{2-star}})+\bs\Pi_N(g_{\mathrm{3tree}})
+\bs\Pi_N(g_{\mathrm{4tree}})\), and
\(\bs\Pi_N^{\mathrm{US}}:=\bs\Pi_N(g_{\mathrm{4tree}})\).
Under Assumptions~\ref{ass:dyad_ind}, \ref{ass:bounded_design}, and
\ref{ass:sparse_envelope}, the order bounds in
Table~\ref{tab:regimemap} yield the leading approximations.
If \(N\rho_N\to\infty\), then
\(N\rho_N^{-7/2}\bs h_N(\bs\beta_0)
=N\rho_N^{-7/2}\bs\Pi_N^{\mathrm{D}}+o_p(1)\).
If \(\rho_N\asymp N^{-1}\), then
\(N^{9/2}\bs h_N(\bs\beta_0)
=N^{9/2}\bs\Pi_N^{\mathrm{S}}+o_p(1)\).
If \(N\rho_N\to0\), then
\(N^{5/2}\rho_N^{-2}\bs h_N(\bs\beta_0)
=N^{5/2}\rho_N^{-2}\bs\Pi_N^{\mathrm{US}}+o_p(1)\).
For these approximations to yield nontrivial Gaussian limits, the normalized
leading projection sum in each case must have a finite, nonzero limiting
covariance. Assumption~\ref{ass:regimeSigma} requires convergence of the
normalized conditional covariance of the relevant projection sum. The
positive-trace condition rules out the zero covariance matrix as a limit, but
does not require the limiting covariance matrix to be positive definite.
\begin{assumption}
\label{ass:regimeSigma}
\textnormal{(i)} If \(N\rho_N\to\infty\), there exists a positive
semidefinite matrix \(\bf\Sigma_{\mathrm{D}}\) such that
\(\mathrm{Var}(N\rho_N^{-7/2}\bs\Pi_N^{\mathrm{D}}\mid\bf X,\bs\Gamma)\xrightarrow{P}\bf\Sigma_{\mathrm{D}}\) and
\(\operatorname{tr}(\bf\Sigma_{\mathrm{D}})>0\). \textnormal{(ii)} If
\(\rho_N\asymp N^{-1}\), there exists a positive
semidefinite matrix \(\bf\Sigma_{\mathrm{S}}\) such that
\(\mathrm{Var}(N^{9/2}\bs\Pi_N^{\mathrm{S}}\mid\bf X,\bs\Gamma)\xrightarrow{P}\bf\Sigma_{\mathrm{S}}\) and
\(\operatorname{tr}(\bf\Sigma_{\mathrm{S}})>0\). \textnormal{(iii)} If \(N\rho_N\to0\),
there exists a positive semidefinite matrix
\(\bf\Sigma_{\mathrm{US}}\) such that
\(\mathrm{Var}(N^{5/2}\rho_N^{-2}\bs\Pi_N^{\mathrm{US}}\mid\bf X,\bs\Gamma)\xrightarrow{P}\bf\Sigma_{\mathrm{US}}\) and
\(\operatorname{tr}(\bf\Sigma_{\mathrm{US}})>0\).
\end{assumption}
\begin{thm}[Asymptotic normality of the sample moment]
\label{thm:regimeCLT}
Under Assumptions~\ref{ass:dyad_ind}, \ref{ass:bounded_design},
\ref{ass:sparse_envelope}, and \ref{ass:regimeSigma}, the sample moment
\(\bs h_N(\bs{\beta}_{0})\) has the following Gaussian limits.
\textnormal{(i)} In the dense or mildly sparse regime
(\(N\rho_N\to\infty\)),
\(N\rho_N^{-7/2}\bs h_N(\bs{\beta}_{0})\xrightarrow{D} N(0,\bf\Sigma_{\mathrm{D}})\).
\textnormal{(ii)} In the sparse regime (\(\rho_N\asymp N^{-1}\)),
\(N^{9/2}\bs h_N(\bs{\beta}_{0})\xrightarrow{D} N(0,\bf\Sigma_{\mathrm{S}})\).
\textnormal{(iii)} In the ultra-sparse regime
(\(N\rho_N\to0\) and \(N^5\rho_N^4\to\infty\)),
\(N^{5/2}\rho_N^{-2}\bs h_N(\bs{\beta}_{0})
\xrightarrow{D} N(0,\bf\Sigma_{\mathrm{US}})\).
\end{thm}
The leading projection approximations and
Theorem~\ref{thm:regimeCLT} show that network sparsity determines both the
normalization of the sample moment and the graph classes entering its limiting
covariance. Table~\ref{tab:boundaryregimes} summarizes the corresponding
sparsity conditions, leading graph classes, and variance orders.
\begin{table}[!ht]
\centering
\caption{Leading graph classes and variance orders in the three Gaussian
regimes. Graph labels follow Figure~\ref{fig:projection_classes_main}.}
\label{tab:boundaryregimes}
{\small
\setlength{\extrarowheight}{1.5pt}
\setlength{\tabcolsep}{0.45pt}
\begin{tabularx}{\textwidth}{@{}>{\centering\arraybackslash}m{3.50cm}
@{\hspace{0.10cm}}>{\centering\arraybackslash}m{4.15cm}
>{\centering\arraybackslash}m{2.30cm}
>{\centering\arraybackslash}X
>{\centering\arraybackslash}m{1.50cm}@{}}
\toprule
Regime & Condition & Leading graph classes & Graph shape(s) & \shortstack{Variance\\order} \\
\midrule
\mbox{Dense/mildly sparse} &
\(N\rho_N\to\infty\) &
\(g_{\mathrm{dyad}}\) &
\makebox[\linewidth][c]{
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (1,0) {}; \draw[taxEdge] (a)--(b);
\end{tikzpicture}
} &
\(N^{-2}\rho_N^7\) \\
\addlinespace[0.55ex]
\parbox[c][2.15cm][c]{\linewidth}{\centering Sparse} &
\parbox[c][2.15cm][c]{\linewidth}{\centering \(\rho_N\asymp N^{-1}\)} &
\parbox[c][2.15cm][c]{\linewidth}{\centering
\shortstack[c]{\(g_{\mathrm{dyad}}\); \(g_{\text{2-star}}\)\\[1.75ex]
\(g_{\mathrm{3tree}}\); \(g_{\mathrm{4tree}}\)}} &
\parbox[c][2.15cm][c]{\linewidth}{\centering
\raisebox{1.50ex}{
\begin{tabular}[c]{@{}c@{\hspace{0.12em}}c@{\hspace{0.12em}}c@{}}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (1,0) {}; \draw[taxEdge] (a)--(b);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.7,.65) {}; \node[taxNode] (c) at (1.4,0) {}; \draw[taxEdge] (a)--(b)--(c);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.7,0) {}; \node[taxNode] (c) at (1.4,0) {}; \node[taxNode] (d) at (2.1,0) {}; \draw[taxEdge] (a)--(b)--(c)--(d);
\end{tikzpicture}
\\[1.05ex]
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (0,.8) {}; \node[taxNode] (c) at (-.7,-.45) {}; \node[taxNode] (d) at (.7,-.45) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.55,0) {}; \node[taxNode] (c) at (1.1,0) {}; \node[taxNode] (d) at (1.65,0) {}; \node[taxNode] (e) at (2.2,0) {}; \draw[taxEdge] (a)--(b)--(c)--(d)--(e);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.65,0) {}; \node[taxNode] (c) at (-.55,.55) {}; \node[taxNode] (d) at (-.55,-.55) {}; \node[taxNode] (e) at (1.25,0) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d); \draw[taxEdge] (b)--(e);
\end{tikzpicture}
\end{tabular}}} &
\parbox[c][2.15cm][c]{\linewidth}{\centering \(N^{-9}\)} \\
\addlinespace[0.55ex]
Ultra-sparse &
\(N\rho_N\to0,\;N^5\rho_N^4\to\infty\) &
\(g_{\mathrm{4tree}}\) &
\makebox[\linewidth][c]{
\begin{tabular}[c]{@{}c@{\hspace{0.30em}}c@{}}
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.55,0) {}; \node[taxNode] (c) at (1.1,0) {}; \node[taxNode] (d) at (1.65,0) {}; \node[taxNode] (e) at (2.2,0) {}; \draw[taxEdge] (a)--(b)--(c)--(d)--(e);
\end{tikzpicture}
&
\begin{tikzpicture}[
baseline=(current bounding box.center),
scale=0.33,
taxNode/.style={circle,fill=black,inner sep=0.9pt,outer sep=0pt},
taxEdge/.style={line width=0.6pt},
scale=0.90
]
\node[taxNode] (a) at (0,0) {}; \node[taxNode] (b) at (.65,0) {}; \node[taxNode] (c) at (-.55,.55) {}; \node[taxNode] (d) at (-.55,-.55) {}; \node[taxNode] (e) at (1.25,0) {}; \draw[taxEdge] (a)--(b); \draw[taxEdge] (a)--(c); \draw[taxEdge] (a)--(d); \draw[taxEdge] (b)--(e);
\end{tikzpicture}
\end{tabular}} &
\(N^{-5}\rho_N^4\) \\
\bottomrule
\end{tabularx}
}
\end{table}
\FloatBarrier
The condition \(N^5\rho_N^4\to\infty\) in
Theorem~\ref{thm:regimeCLT}(iii) has an effective sample size interpretation.
The leading variation in this regime comes from configurations of five nodes
connected by four links (see Figure~\ref{fig:projection_classes_main}(e)--(f)).
The number of possible configurations is of order
\(N^5\), and a proportion of order \(\rho_N^4\) contains the required link
pattern. Thus, \(N^5\rho_N^4\) is the effective number of informative
configurations, and \(N^5\rho_N^4\to\infty\) is the corresponding effective
sample size condition for asymptotic normality in the ultra-sparse regime.
In the sparse Erd\H{o}s--R\'enyi model, the classical central limit condition
of \citet{rucinski1988when} for the count of a fixed tree with five nodes and
four edges reduces to \(N\rho_N^{4/5}\to\infty\), which is equivalent to
\(N^5\rho_N^4\to\infty\).
\subsection{Asymptotic Properties of the Pentad-GMM Estimator}
\label{sec:gmm_asymptotics}
We now use the limit theory for the sample moment to obtain the asymptotic
distribution of the pentad-GMM estimator. The following conditions impose a
population limit, identification, the rank of the limiting Jacobian, and
weight convergence.
\begin{assumption}
\label{ass:consistency}
(i) There exists a deterministic continuous function
\(\bs h_0:\mathcal B\to\mathbb{R}^{q}\) such that
\(\rho_N^{-4}\mathbb{E}[\bs h_N(\bs{\beta})]\to\bs h_0(\bs{\beta})\)
for every \(\bs{\beta}\in\mathcal B\), and
\(\bs h_0(\bs{\beta})=0\iff \bs{\beta}=\bs{\beta}_{0}\).
(ii) The matrix
\(\dot{\bs h}_0:=
\partial\bs h_0(\bs\beta_0)/\partial\bs\beta^\top
\in\mathbb{R}^{q\times d}\)
has full column rank.
(iii) The GMM weight matrix $\widehat{\bf W}_N$ used in \eqref{eq:beta_gmm} satisfies
$\widehat{\bf W}_N\xrightarrow{P} \bf W$, where $\bf W$ is deterministic and positive definite.
\end{assumption}
Assumption~\ref{ass:consistency} collects standard GMM
regularity conditions; see \citet{hansen1982large} and
\citet{newey1994large}. Assumption~\ref{ass:consistency}(i) identifies
\(\bs\beta_0\) as the unique zero of the limiting moment. For the five-node
determinant moments entering \eqref{eq:sample_moment-multi}, the smallest
nonzero link degree is four, so the population moment has the natural scale
\(\rho_N^4\), which explains the factor \(\rho_N^{-4}\) in
Assumption~\ref{ass:consistency}(i). Assumption~\ref{ass:consistency}(ii), together with
the positive definiteness in
Assumption~\ref{ass:consistency}(iii), makes
\(\dot{\bs h}_0^\top\bf W\dot{\bs h}_0\) nonsingular for the local
linearization. Assumption~\ref{ass:consistency}(iii) accommodates a data-dependent
weight matrix. In our implementation, we use the inverse empirical-Gram weight
\(\widehat{\bf W}_N=
\{\left\lvert \mathcal S_N \right\rvert^{-1}\sum_{s\in\mathcal S_N}
\bs Z_s\bs Z_s^\top\}^{-1}\). If
\(\mathbb{E}[\bs Z_s\bs Z_s^\top]\) is a fixed positive definite matrix, a law of
large numbers for \(U\)-statistics
(see \citealp[Section~5.4, Theorem~A]{serfling1980approximation}) implies
that the inverse
empirical-Gram weight satisfies Assumption~\ref{ass:consistency}(iii).
\begin{thm}[Asymptotic properties of Pentad-GMM]
\label{thm:gmm}
Suppose Assumptions~\ref{ass:dyad_ind}, \ref{ass:bounded_design},
\ref{ass:sparse_envelope}, \ref{ass:regimeSigma}, and
\ref{ass:consistency} hold, and
\(N^5\rho_N^4\to\infty\). For the regime-specific normalization and
covariance, set
\[
(a_N,\bf\Sigma)
=
\begin{cases}
(N\rho_N^{-7/2},\bf\Sigma_{\mathrm{D}}),
& N\rho_N\to\infty,\\
(N^{9/2},\bf\Sigma_{\mathrm{S}}),
& \rho_N\asymp N^{-1},\\
(N^{5/2}\rho_N^{-2},\bf\Sigma_{\mathrm{US}}),
& N\rho_N\to0 \ \text{and}\ N^5\rho_N^4\to\infty,
\end{cases}
\]
where \(\bf\Sigma_{\mathrm{D}},\bf\Sigma_{\mathrm{S}}\), and
\(\bf\Sigma_{\mathrm{US}}\) are the limits in
Assumption~\ref{ass:regimeSigma}. Let
\(\widehat{\bs{\beta}}_{\mathrm{GMM}}\) denote the minimizer in
\eqref{eq:beta_gmm}.
Define
\(
\mathbf{V}_0
:=
\big(\dot{\bs h}_0^\top\bf W\dot{\bs h}_0\big)^{-1}
\dot{\bs h}_0^\top\bf W\bf\Sigma\bf W\dot{\bs h}_0
\big(\dot{\bs h}_0^\top\bf W\dot{\bs h}_0\big)^{-1}
\). Then
\[
a_N\rho_N^4(\widehat{\bs{\beta}}_{\mathrm{GMM}}-\bs{\beta}_{0})
\xrightarrow{D}N(0,\mathbf{V}_0).
\]
\end{thm}
The normalization in Theorem~\ref{thm:gmm} implies that
\(
\widehat{\bs{\beta}}_{\mathrm{GMM}}-\bs{\beta}_{0}
=
O_p\big((N^2\rho_N)^{-1/2}\big)
\)
in the dense/mildly sparse and sparse regimes, whereas
\(
\widehat{\bs{\beta}}_{\mathrm{GMM}}-\bs{\beta}_{0}
=
O_p\big((N^5\rho_N^4)^{-1/2}\big)
\)
in the ultra-sparse regime. Since the average link probability is \(\rho_N\),
the average expected degree is \((N-1)\rho_N\), and the expected total number
of links is of order \(N^2\rho_N\). Thus, in both the dense or mildly sparse
regime and the sparse regime, the convergence rate \(N\rho_N^{1/2}\) is of the
same order as the square root of the expected number of links. It is of order
\(N\) when \(\rho_N\) is bounded away from zero and \(\sqrt{N}\) when
\(\rho_N\asymp N^{-1}\), with the intermediate rate \(N\rho_N^{1/2}\) in mildly
sparse networks. The \(N\)-rate in dense
networks and the \(\sqrt{N}\)-rate in sparse networks coincide with those
obtained for the TU tetrad-logit estimator by \citet{graham2017econometric}.
Although the convergence rates coincide, the leading variance components differ.
In the sparse regime, the tetrad-logit estimator in \cite{graham2017econometric} is driven by the
one-dyad projection, whereas our pentad-GMM estimator generally requires
higher-order projections as well.
In the ultra-sparse regime, the rate
\(N^{5/2}\rho_N^2=(N^5\rho_N^4)^{1/2}\) is the square root of the effective
sample size \(N^5\rho_N^4\), which corresponds to the number of informative
five-node, four-link configurations discussed after
Theorem~\ref{thm:regimeCLT}.
Although Theorem~\ref{thm:gmm} yields different normalizations and covariance
limits across the three regimes, the same inference procedure can be used in
all three, provided that a single covariance estimator is consistent under
every regime.
\begin{cor}[Unified inference]
\label{cor:generic_gmm_inference}
Under the conditions of Theorem~\ref{thm:gmm}, let
\(\widehat{\bf\Omega}_N\) be a covariance estimator that does
not depend on the sparsity regime and satisfies, under each of the three
regimes in Theorem~\ref{thm:gmm},
\(a_N^2\widehat{\bf\Omega}_N\xrightarrow{P}\bf\Sigma\).
Set \(\widehat{\dot{\bs h}}_N
:=\partial\bs h_N(\widehat{\bs\beta}_{\mathrm{GMM}})/
\partial\bs\beta^\top\) and
\(\widehat{\mathbf{V}}_N
:=(\widehat{\dot{\bs h}}_N^\top\widehat{\bf W}_N
\widehat{\dot{\bs h}}_N)^{-1}\allowbreak
\widehat{\dot{\bs h}}_N^\top\widehat{\bf W}_N\allowbreak
\widehat{\bf\Omega}_N\allowbreak
\widehat{\bf W}_N\widehat{\dot{\bs h}}_N\allowbreak
(\widehat{\dot{\bs h}}_N^\top\widehat{\bf W}_N
\widehat{\dot{\bs h}}_N)^{-1}\).
Then \(a_N^2\rho_N^8\widehat{\mathbf{V}}_N\xrightarrow{P}\mathbf{V}_0\). Consequently,
for every fixed \(\bs c\in\mathbb R^d\) such that
\(\bs c^\top \mathbf{V}_0\bs c>0\),
\[
\frac{\bs c^\top(\widehat{\bs\beta}_{\mathrm{GMM}}-\bs\beta_0)}
{\sqrt{\bs c^\top\widehat{\mathbf{V}}_N\bs c}}
\xrightarrow{D}N(0,1).
\]
\end{cor}
Corollary~\ref{cor:generic_gmm_inference} is stated for a generic covariance
estimator. Appendix~\ref{sec:feasible_covariance} provides a tractable
regime-adaptive covariance estimator
\(\widehat{\bf\Omega}_{N}^{\mathrm{univ}}(\widehat{\bs\beta}_{\mathrm{GMM}})\) and verifies the required
covariance consistency under all three regimes. This yields unified
inference without preliminary estimation of \(\rho_N\) or regime selection.
\section{Implementation}
\label{sec:implementation}
This section describes the practical implementation of our estimator,
including the construction of the instruments, the normalization of the GMM
criterion using pivotal-pair moments, and an algorithm that reduces the
computational cost of evaluating the pentad moments.
The conditional restriction in \eqref{eq:moment_cond-free} implies, by iterated
expectations, that any bounded instrument \(\bs Z_s=z(\bf X_s)\) is valid:
\(\mathbb{E}[\bs Z_s\psi_s(\bs\beta_0)]=\mathbf 0\).\footnote{The usual optimal
instrument based on the conditional expectation of the derivative of the moment function
\citep{chamberlain1987asymptotic} requires knowledge of the conditional
distribution of the individual fixed effects given the observed covariates.
We avoid estimating this distribution. Substituting preliminary estimates of
all \(N\) individual fixed effects into the derivative would instead produce a
generated instrument and require a separate asymptotic analysis.}
We construct an instrument \(\bs Z_s\) to exploit the relabeling properties of
the pentad kernel
\(\psi_{ijklm}(\bs\beta)=\det(\widetilde{\mathbf M}_{ijklm}(\bs\beta))\).
Lemma~\ref{lem:pivsym} shows that this kernel is
invariant under interchange of the pivotal nodes \(i\) and \(j\) and
alternating under permutations of the peripheral nodes \(k,l,m\). Because
\(\bs h_N(\bs\beta)\) averages over all ordered pentads, the contributions from
odd and even peripheral permutations cancel whenever \(\bs Z_s\) is invariant
under permutations of \(k,l,m\), so that
\(\bs h_N(\bs\beta)=\mathbf 0\) for every \(\bs\beta\). Such an instrument cannot
satisfy the identification condition in Assumption~\ref{ass:consistency}(i). More
generally, only the part of an instrument that is alternating under peripheral
permutations contributes to \(\bs h_N(\bs\beta)\); all remaining parts cancel in
the ordered average. We therefore choose \(\bs Z_s\) to be alternating under
permutations of \(k,l,m\) and invariant under interchange of \(i\) and \(j\). A
determinant-based construction naturally provides both properties, as shown
below.
For the construction below, suppose \(d\geq2\), and let
\(X_{uv,1},\ldots,X_{uv,d}\) denote the coordinates of
\(\bs X_{uv}\). To include interactions between covariate coordinates, for
each pair \(1\leq a<b\leq d\), define their normalized sum and difference by
\[
X_{uv}^{ab,+}:=
\frac{X_{uv,a}+X_{uv,b}}{\sqrt 2},
\qquad
X_{uv}^{ab,-}:=
\frac{X_{uv,a}-X_{uv,b}}{\sqrt 2}.
\]
For two scalar dyadic arrays \(\mathbf{A}=(A_{uv})_{u\ne v}\) and
\(\mathbf{B}=(B_{uv})_{u\ne v}\), define the determinant difference
\[
\mathcal D_{ij;klm}(\mathbf{A},\mathbf{B})
:=
\det
\begin{pmatrix}
1 & A_{ik} & B_{jk}\\
1 & A_{il} & B_{jl}\\
1 & A_{im} & B_{jm}
\end{pmatrix}
-
\det
\begin{pmatrix}
1 & B_{ik} & A_{jk}\\
1 & B_{il} & A_{jl}\\
1 & B_{im} & A_{jm}
\end{pmatrix}.
\]
Interchanging any two of \(k,l,m\) changes the sign of both determinants and
hence of \(\mathcal D_{ij;klm}(\mathbf{A},\mathbf{B})\).
After interchanging \(i\) and \(j\), the first determinant equals the negative
of the original second determinant, while the second equals the negative of
the original first. Hence
\(\mathcal D_{ji;klm}(\mathbf{A},\mathbf{B})=\mathcal D_{ij;klm}(\mathbf{A},\mathbf{B})\).
For either sign, write \(\mathbf{X}^{ab,\pm}:=(X_{uv}^{ab,\pm})_{u\ne v}\) for the
corresponding scalar dyadic array and define its pointwise square by
\((\mathbf{X}^{ab,\pm})^2:=((X_{uv}^{ab,\pm})^2)_{u\ne v}\).
For each \(1\leq a<b\leq d\), define the three-dimensional block
\[
\bs Z_s^{ab}
:=
\Bigl(
\mathcal D_{ij;klm}\bigl(\mathbf{X}^{ab,+},(\mathbf{X}^{ab,+})^2\bigr),
\mathcal D_{ij;klm}\bigl(\mathbf{X}^{ab,-},(\mathbf{X}^{ab,-})^2\bigr),
\mathcal D_{ij;klm}\bigl(\mathbf{X}^{ab,+},\mathbf{X}^{ab,-}\bigr)
\Bigr)^\top.
\]
Stacking these blocks gives
\begin{equation}
\label{eq:covariate_only_instrument}
\bs Z_s
:=
\Bigl(
(\bs Z_s^{12})^\top,\ldots,(\bs Z_s^{1d})^\top,
(\bs Z_s^{23})^\top,\ldots,(\bs Z_s^{2d})^\top,
\ldots,(\bs Z_s^{(d-1)d})^\top
\Bigr)^\top
\in\mathbb R^{3d(d-1)/2}.
\end{equation}
Because each entry of \(\bs Z_s\) is a determinant difference of the form
\(\mathcal D_{ij;klm}(\cdot,\cdot)\), \(\bs Z_s\) changes sign whenever two of
\(k,l,m\) are interchanged and is unchanged when \(i\) and \(j\) are
interchanged, as required.
Within each block \(\bs Z_s^{ab}\), the first two entries pair the normalized
sum and difference with their respective squares, while the third forms a
cross term between the two. Thus
\eqref{eq:covariate_only_instrument} contains three instruments for each of the
\(\binom d2\) coordinate pairs, for a total of \(3d(d-1)/2\).
To improve finite-sample stability, we normalize the GMM criterion by the
average weighted squared norm of the pivotal-pair moments, with a ridge term
that keeps the denominator positive.
For \(i<j\), define the average moment for pivotal pair \(\{i,j\}\) by
\begin{equation}
\bs m_{ij,N}(\bs\beta)
:=\frac{1}{(N-2)(N-3)(N-4)}
\sum_{k,l,m\in[N]\setminus\{i,j\}:\,k,l,m\text{ distinct}}
\bs\Psi_{(i,j,k,l,m)}(\bs\beta).
\label{eq:pivotal_pair_moment}
\end{equation}
Set \(\widehat{\rho}_N:=\binom{N}{2}^{-1}\sum_{i<j}L_{ij}+N^{-2}\) and
\(\widehat c_N:=\widehat{\rho}_N^7+N^{-1}\widehat{\rho}_N^6+N^{-2}\widehat{\rho}_N^5+N^{-3}\widehat{\rho}_N^4\).
Here \(\widehat{\rho}_N\) is the empirical link density with an \(N^{-2}\) adjustment
that keeps it strictly positive even when no links are observed, and
\(\widehat c_N\) is a strictly positive data-driven ridge term constructed from
\(\widehat{\rho}_N\). The normalized GMM criterion is
\begin{equation}
Q_N^\dagger(\bs\beta)
:=
\frac{\bs h_N(\bs\beta)^\top\widehat{\bf W}_N\bs h_N(\bs\beta)}
{\binom{N}{2}^{-1}\sum_{i<j}
\bs m_{ij,N}(\bs\beta)^\top\widehat{\bf W}_N
\bs m_{ij,N}(\bs\beta)
+\widehat c_N}.
\label{eq:pair_normalized_criterion}
\end{equation}
Because \(\bs\Psi_s(\mathbf 0)=0\) for every pentad,
\(\bs h_N(\mathbf 0)=\bs m_{ij,N}(\mathbf 0)=0\). Consequently, the
unnormalized criterion in \eqref{eq:beta_gmm} may become artificially small
near the excluded origin $\bs 0$ and lead the numerical optimizer toward values close
to zero. The average weighted squared norm of the pivotal-pair moments in
the denominator of
\eqref{eq:pair_normalized_criterion} also tends to zero as
\(\bs\beta\) approaches \(\mathbf 0\) and, like the numerator, is quadratic in
averages of the same pentad moments. This normalization reduces the tendency of the criterion to fall simply
because the pentad moments shrink near the origin, thereby improving
numerical stability.
The ridge \(\widehat c_N\) adds a strictly
positive term to the denominator in finite samples. Its magnitude is calibrated
to the sparsity-dependent scale of
\(\binom{N}{2}^{-1}\sum_{i<j}
\bs m_{ij,N}(\bs\beta)^\top\widehat{\bf W}_N
\bs m_{ij,N}(\bs\beta)\), keeping the normalization well behaved across the
three sparsity regimes; see the proof of
Lemma~\ref{lem:pair_normalization} for the derivation.
Related rescalings appear
in fixed-effects binary-choice estimation. \citet{honore2024moment} normalize
each fixed-effect-free moment to bound the moment and its gradient uniformly
and improve small-sample GMM performance; the
conditional-likelihood scores in \citet{honore2000panel} have a similar form.
\begingroup
\emergencystretch=2em
The estimator used in our simulation study and empirical application is
\(\widehat{\bs\beta}^{\dagger}\in
\arg\min_{\bs\beta\in\mathcal B}Q_N^\dagger(\bs\beta)\).
Lemma~\ref{lem:pair_normalization} shows that the resulting estimator
\(\widehat{\bs\beta}^{\dagger}\) is asymptotically equivalent to the
pentad-GMM estimator \(\widehat{\bs\beta}_{\mathrm{GMM}}\) in all three regimes, i.e.,
\(a_N\rho_N^4(\widehat{\bs\beta}^{\dagger}
-\widehat{\bs\beta}_{\mathrm{GMM}})\xrightarrow{P}0\), where \(a_N\) is the regime-specific normalization factor defined in
Theorem~\ref{thm:gmm}. Consequently, under each regime, the scaled
estimators \(a_N\rho_N^4(\widehat{\bs\beta}^{\dagger}-\bs\beta_0)\) and
\(a_N\rho_N^4(\widehat{\bs\beta}_{\mathrm{GMM}}-\bs\beta_0)\) both
converge in distribution to \(N(0,\mathbf{V}_0)\).
In the simulation study and empirical application, we refer to
\(\widehat{\bs\beta}^{\dagger}\) as the pentad-GMM estimator.
\par\endgroup
Brute-force evaluation of \eqref{eq:sample_moment-multi} over all ordered
five-node configurations requires \(O(N^5)\) operations. The resulting
computational burden increases rapidly with network size \(N\). However, exploiting
the determinant structure reduces the computational complexity from
\(O(N^5)\) to \(O(N^3)\). The reduction begins by fixing an ordered pivotal
pair \((i,j)\) and ordering the remaining nodes by their labels.
Fix \(t\in\{1,\ldots,q\}\). The \(t\)-th coordinate of
\(\bs\Psi_s(\bs\beta)\) is \(Z_{s,t}\psi_s(\bs\beta)\). Both factors are
alternating under permutations of \(k,l,m\), so their product is invariant.
Hence, for any three distinct peripheral nodes, the six ways of assigning them
to the roles \((k,l,m)\) produce the same value of
\(Z_{s,t}\psi_s(\bs\beta)\). The sum over ordered peripheral triples therefore
equals six times the sum over triples with \(k<l<m\). After expanding the
determinants defining \(Z_{s,t}\) and \(\psi_s(\bs\beta)\), the contribution of
\((i,j)\) to the \(t\)th coordinate of \(\bs h_N(\bs\beta)\) is a linear
combination of a fixed number of sums of the form
\begin{equation}
\label{eq:cumulative_triple_sum}
\sum_{k<l<m}x_k y_l z_m
=
\sum_l\left(\sum_{k<l}x_k\right)y_l
\left(\sum_{m>l}z_m\right).
\end{equation}
Because \(Z_{s,t}\) does not depend on \(\bs\beta\), differentiating
\(\psi_s(\bs\beta)\) with respect to \(\bs\beta\) likewise expresses every
entry in the \(t\)th row of \(\dot{\bs h}_N(\bs\beta)\) as a linear
combination of a fixed number of sums \(\sum_{k<l<m}x_k y_l z_m\).
All prefix sums \(\sum_{k<l}x_k\) can be computed in \(O(N)\) operations, and
all suffix sums \(\sum_{m>l}z_m\) can likewise be computed in \(O(N)\)
operations; computing both collections therefore costs \(O(N)\) operations in
total. Substituting these precomputed sums into
\eqref{eq:cumulative_triple_sum} leaves a single sum over \(l\), which requires
another \(O(N)\) operations. Thus, each triple sum of the form in
\eqref{eq:cumulative_triple_sum} can be evaluated in \(O(N)\) operations.
Since \(q\) and \(d\) are fixed and each coordinate
involves only a fixed number of such terms, the contributions of a fixed
ordered pair \((i,j)\) to \(\bs h_N(\bs\beta)\) and
\(\dot{\bs h}_N(\bs\beta)\) can both be evaluated in \(O(N)\) operations.
There are \(N(N-1)=O(N^2)\) ordered pivotal pairs \((i,j)\), so evaluating both
\(\bs h_N(\bs\beta)\) and \(\dot{\bs h}_N(\bs\beta)\) requires
\(O(N^3)\) operations.
The resulting \(O(N^3)\) algorithm substantially improves the practical
feasibility of pentad-GMM for large networks.
Appendix~\ref{sec:simulation} evaluates the finite-sample behavior of the
implementation described in this section. Across dense and sparse networks
with symmetric and asymmetric dyadic covariates, root mean squared error (RMSE) declines as \(N\)
increases, while marginal coverage based on the universal covariance estimator
is broadly close to its nominal level. These findings support the practical
feasibility of the proposed estimation and inference procedure, although
sampling variability remains larger in the sparse designs.
\section{Empirical Illustration}
\label{sec:empirical}
This section applies the pentad-GMM estimator to study the
\emph{Peer Relationships among Master of Social Work Students} data
\citep{mauldin2020peer}, a longitudinal study of one entering master of social work (MSW) class.
The data combine four waves of network surveys with demographic and academic
records. We focus on the wave-4 academic-discussion network and examine how
academic similarity, shared cohort membership, and a potential partner's
universal--diverse orientation are associated with link formation after allowing
for individual-specific unobserved heterogeneity. We compare the estimates under NTU logit models
with those under TU logit models (\citealp{graham2017econometric}).
\subsection{Data and Variables}
The data cover MSW students at a large public university in the southern
United States. Students were surveyed at orientation in July--August 2014,
twice during their first semester in October and November--December 2014, and
at the end of the Fall 2015 semester in December 2015. At waves 2--4,
respondents used a roster of 145 students to report academic-discussion,
friendship, and professional-influence links. We use the wave-4 academic
discussion network as the outcome. Students separately nominate classmates
with whom they report studying, discussing coursework, exchanging feedback on
assignments, or asking questions about homework. We construct an
undirected measure of these bilateral interactions. For each pair \(i<j\), we
record \(L_{ij}=1\) if at least one student reports academic discussion with
the other and set \(L_{ij}=0\) otherwise. The data show only who reported
the interaction, but we do not observe who initiated or proposed it.
The final sample contains 109 students across six educational cohorts with complete
wave-4 academic-discussion information, grade point average (GPA), and scores on the
15-item short form of the Miville--Guzman Universality--Diversity Scale
(M-GUDS-S; hereafter MGUDS). These students form
5,886 unordered pairs, of which 1,246 are academic-discussion links, yielding a
link density of 0.2117.
GPA is calculated over the most recent 60 credit hours and measures prior
academic performance. MGUDS scores measure universal--diverse orientation.
The scale covers diversity of contact,
appreciation of similarities and differences, and comfort with differences.
Higher scores indicate a stronger universal--diverse orientation, reflecting
greater openness to engaging with people from different backgrounds and
perspectives \citep{fuertes2000factor}. Raw GPA and MGUDS values are measured on different
scales and are difficult to compare directly. In the subsequent analysis, we therefore use inverse-normal
rank scores \citep[see][for a similar transformation]{karadja2017richer}, which preserve the ordering of
observations within each measure and map the two measures onto a common
normal-score scale.
Table~\ref{tab:msw-estimation-summary} reports descriptive statistics for the
variables used in the analysis.
\begin{table}[H]
\centering
\caption{Sample summary statistics}
\label{tab:msw-estimation-summary}
\footnotesize
\begin{threeparttable}
\begin{tabular}{lrrrrr}
\toprule
Variable & Observations & Mean & SD & Min & Max \\
\midrule
\multicolumn{6}{l}{\textit{Panel A: Student characteristics}}\\
Raw GPA & 109 & 3.459 & 0.330 & 2.630 & 4.000 \\
Raw MGUDS score & 109 & 72.128 & 6.948 & 54.000 & 89.000 \\
GPA inverse-normal rank & 109 & -0.004 & 0.985 & -2.605 & 1.919 \\
MGUDS inverse-normal rank & 109 & 0.000 & 0.996 & -2.605 & 2.605 \\
\addlinespace
\multicolumn{6}{l}{\textit{Panel B: Network and dyadic characteristics}}\\
Academic-discussion link & 5,886 & 0.212 & 0.409 & 0.000 & 1.000 \\
GPA inverse-normal rank distance & 5,886 & 1.123 & 0.824 & 0.000 & 4.524 \\
MGUDS inverse-normal rank distance & 5,886 & 1.130 & 0.840 & 0.000 & 5.211 \\
Same cohort & 5,886 & 0.170 & 0.375 & 0.000 & 1.000 \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{\linewidth}
\footnotesize
\textit{Notes:} SD denotes standard deviation. ``Raw'' and ``inverse-normal rank'' denote values before and after the inverse-normal rank transformation, respectively.
\end{minipage}
\end{threeparttable}
\end{table}
\subsection{Results and Discussion}
We consider three specifications of the TU model in \eqref{eq:tu-model}
and the NTU model in \eqref{mod:mod2}, all of which include
GPA inverse-normal rank distance and a same-cohort indicator.
The TU specification and the symmetric
NTU specification use MGUDS inverse-normal rank distance to examine whether students with similar universal--diverse
orientations are more likely to form academic-discussion links.
We also examine whether students with higher MGUDS are more attractive
discussion partners. In the TU model, however, additive MGUDS terms
are absorbed by the individual fixed effects, so their coefficient
cannot be separately identified. This limitation motivates the
asymmetric NTU specification, in which partner MGUDS inverse-normal
rank \(\mathrm{MGUDS}_j\) enters student \(i\)'s utility.
Since \(\mathrm{MGUDS}_j\) varies across potential partners,
it cannot be absorbed by student \(i\)'s fixed effect \(\Gamma_i\).
The resulting regressors are
\begingroup
\[
\begin{aligned}
\bs X_{ij}^{\mathrm{TU}}=\bs X_{ij}^{\mathrm{NTU,S}}
&:=\begin{pmatrix}
\left\lvert \mathrm{MGUDS}_i-\mathrm{MGUDS}_j \right\rvert\\
\left\lvert \mathrm{GPA}_i-\mathrm{GPA}_j \right\rvert\\
\1\{C_i=C_j\}
\end{pmatrix},\
\bs X_{ij}^{\mathrm{NTU,A}}
:=\begin{pmatrix}
\mathrm{MGUDS}_j\\
\left\lvert \mathrm{GPA}_i-\mathrm{GPA}_j \right\rvert\\
\1\{C_i=C_j\}
\end{pmatrix},
\end{aligned}
\]
\endgroup
where \(\mathrm{GPA}_i\) and \(\mathrm{MGUDS}_i\) denote student \(i\)'s
inverse-normal rank scores, and \(C_i\) denotes the student's educational cohort.
For the TU specification, we use the tetrad-logit estimator of
\citet{graham2017econometric} and its dyad-projection covariance. We estimate
both NTU specifications using the procedures in Section~\ref{sec:implementation}.
With three regressors,
\eqref{eq:covariate_only_instrument} stacks three determinant blocks into nine
moment conditions. As the GMM objective function may have local optima, we
follow \citet{davezies2023fixed}, use 200 random starting values, and retain the
parameter vector that yields the lowest value of
\eqref{eq:pair_normalized_criterion}.
\begin{table}[!b]
\centering
\caption{TU and NTU estimates of academic-discussion link formation}
\label{tab:empirical-pentad-gmm}
\footnotesize
\begin{threeparttable}
\begin{tabular}{l@{\hspace{1.5em}}ccc}
\toprule
Variable & TU tetrad logit & \multicolumn{2}{c}{NTU pentad-GMM} \\
\cmidrule(lr){2-2}\cmidrule(l){3-4}
& (1) & (2) & (3) \\
\midrule
Partner MGUDS inverse-normal rank
& -- & -- & \(0.352^{***}\) \\
& & & \((0.052)\) \\
\addlinespace[0.2em]
MGUDS inverse-normal rank distance
& \(0.036\) & \(-0.134\) & -- \\
& \((0.076)\) & \((0.188)\) & \\
\addlinespace[0.2em]
GPA inverse-normal rank distance
& \(-0.091\) & \(0.366\) & \(-0.339^{**}\) \\
& \((0.078)\) & \((0.321)\) & \((0.144)\) \\
\addlinespace[0.2em]
Same cohort
& \(3.538^{***}\) & \(3.298^{***}\) & \(2.508^{**}\) \\
& \((0.159)\) & \((0.522)\) & \((1.134)\) \\
\bottomrule
\end{tabular}
\par\smallskip
\begin{minipage}{\linewidth}
\footnotesize
\textit{Notes:} Standard errors are in parentheses. GPA and MGUDS ranks use the
inverse-normal transformation. NTU standard errors use the
positive-semidefinite projection of the universal covariance estimator in
\eqref{eq:universalOmega}; TU standard errors use Graham's dyad-projection
covariance estimator. Significance levels are denoted by \(^{***}p<0.01\),
\(^{**}p<0.05\), and \(^{*}p<0.10\).
\end{minipage}
\end{threeparttable}
\end{table}
Columns (1) and (2) of Table~\ref{tab:empirical-pentad-gmm} compare the TU
and symmetric NTU models using the same regressors. Both show that students
in the same cohort are more likely to form academic-discussion links.
Their same-cohort coefficients are
3.538 and 3.298,
respectively, both significant at the \(1\%\) level, whereas neither GPA nor
MGUDS rank distance is statistically significant. In column (3), however,
the asymmetric NTU specification provides evidence of strong academic homophily
alongside a positive association between partner MGUDS rank and willingness
to form discussion links. These findings emerge when partner MGUDS rank
replaces MGUDS rank distance. The coefficient on partner MGUDS rank is 0.352,
significant at the \(1\%\) level. The GPA
rank-distance and same-cohort coefficients are \(-0.339\) and 2.508,
respectively, both significant at the \(5\%\) level.
Together, these findings suggest that students value both a partner's openness
to different perspectives and similarity in academic achievement. Students who are more
receptive to different perspectives may be more approachable partners for
asking questions and exchanging feedback. This interpretation accords with
evidence linking MGUDS scores to teamwork interest and aptitude among college
students \citep{kottke2011additional}.
\section{Conclusion}
\label{sec:conclusion}
This paper develops a pentad differencing strategy for logistic NTU network
formation models with individual fixed effects. Using observed links and
covariates from five-node pentads, we construct moment restrictions that do
not depend on these fixed effects. We then show that, for a broad class of
covariate specifications, this five-node, seven-dyad construction is minimal
and unique up to relabeling, and the corresponding moment restrictions are unique up to rescaling.
These moment restrictions form the basis of the pentad-GMM estimator. A
Hoeffding decomposition over dyads shows how the leading projection changes
with network sparsity, giving rise to different convergence rates and
asymptotic variances. Under the stated conditions, we establish asymptotic
normality in dense, sparse, and ultra-sparse networks. We also propose a
tractable covariance estimator that is consistent under the appropriate
normalization in all three regimes, allowing for unified inference without prior knowledge of sparsity regimes.
For implementation, we provide a computationally efficient algorithm to reduce
the computational cost of the estimator from na\"ive \(O(N^5)\) to \(O(N^3)\)
operations.
Finally, we apply our method to an academic-discussion network among Master
of Social Work students. The asymmetric NTU estimates show academic homophily and a positive association between partner openness to
different perspectives and link formation. These findings illustrate the
empirical value of our method. Future work could extend the proposed differencing strategy to network models
with strategic link interdependence and to other economic settings.
\clearpage