The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
31,946 characters
Rate-Agnostic Wald Inference for Dyadic Regressions
\begin{bibunit}[jpe]
\spacingset{1}
\if10
{
\title{\bf Rate-Agnostic Wald Inference for Dyadic Regressions\thanks{The authors declare the use of generative AI. The research question, the modeling framework, and the initial roadmap of results are the authors' own. Under the authors' direction and review, AI agents (Claude Opus 5 and Claude Fable 5 by Anthropic) helped build the numerical exercise and prove the results reported here, while various separate agent sessions (Gemini 3.1 Pro by Google and GPT 5.6 Soltice by OpenAI through \href{https://coarse.ink/}{https://coarse.ink/}) independently verified the mathematical correctness of all derivations. The authors have adjudicated all AI-assisted content and take full responsibility for the accuracy, integrity, and originality of the final manuscript.}}
\author{Anonymous}
\date{\today}
\maketitle
} \fi
\if00
{
\title{\bf Rate-Agnostic Wald Inference for Dyadic Regressions\thanks{The authors declare the use of generative AI. The research question, the modeling framework, and the initial roadmap of results are the authors' own. Under the authors' direction and review, AI agents (Claude Opus 5 and Claude Fable 5 by Anthropic) helped build the numerical exercise and prove the results reported here, while various separate agent sessions (Gemini 3.1 Pro by Google and GPT 5.6 Soltice by OpenAI through \href{https://coarse.ink/}{https://coarse.ink/}) independently verified the mathematical correctness of all derivations. The authors have adjudicated all AI-assisted content and take full responsibility for the accuracy, integrity, and originality of the final manuscript.}}
\author{Benjamin O. Harrison\thanks{Department of Economics, Emory University,
Rich Building 306, 1602 Fishburne Dr., Atlanta, GA 30322-2240, USA.
\faEnvelopeO: [email removed].}
\and David T. Jacho-Ch\'{a}vez\thanks{Corresponding author. Department of Economics,
Emory University, Rich Building 306, 1602 Fishburne Dr., Atlanta, GA 30322-2240, USA.
\faEnvelopeO: [email removed].}}
\date{\today}
\maketitle
} \fi
\begin{abstract}
\normalsize
\noindent This paper develops Wald inference for least-squares estimation of linear regression models on dyadic data, accommodating configurations where multiple observations share the same pair of units (e.g., directed flows, multilayer networks, and dyadic panels). We establish that the dyadic-robust Wald statistic is asymptotically $\chi^2_q$ for an arbitrary nonrandom sequence of full-rank restrictions, under a single condition on the accumulation of dependence. Throughout, no convergence rate is assumed or estimated, permitting the condition number of the score's variance matrix to diverge. We further propose a delete-one-unit jackknife alternative that is positive semidefinite by construction. This jackknife statistic attains the same asymptotic limit under one additional condition on dyad multiplicity and remains asymptotically conservative when that condition fails. A supplement contains all proofs, Monte Carlo experiments featuring estimated coefficients that converge at heterogeneous rates, and an empirical gravity application to bilateral trade.
\end{abstract}
\noindent
{\it Keywords:} Cluster-robust inference; Dyadic data; Jackknife; Networks; Wald test \\
\noindent
{\it JEL codes:} C12; C31.
\spacingset{1.15}
\section{Introduction}
\label{sec:intro}
In this paper we provide asymptotically valid Wald inference for least squares in linear regression models where each observation is attached to a \emph{pair} of units, a dyad. Several observations may be attached to the \emph{same} pair. Examples of this setting include the two directed flows between two countries, the layers of a multilayer network, or the dates of a dyadic panel. Our framework encompasses all these settings at once by allowing the map from observations to dyads to be non-injective. Dyadic observations may be arbitrarily dependent when their underlying dyads share at least one common unit; however, observations with disjoint dyads are assumed to be independent. In particular, this accommodates reciprocity by leaving the correlation between the two orientations of a given pair unrestricted.
Inference in this setting is routinely conducted with the dyadic-robust variance
estimator of \citet{fafchamps2007formation} and \citet{aronow2015cluster}, surveyed in \citet{graham2020network}. \citet{tabordmeehan2019inference} provides the corresponding limit theory for dyadic least squares inference, establishing conditions under which the dyadic-robust $t$-statistic is asymptotically standard normal. However, the theory relies on a common convergence rate across the estimator's components. By contrast, \citet[p.~277]{hansen2019asymptotic} develop a rate-free inference framework whose Theorem~9 permits different linear combinations of an estimator to converge at different rates, but their reliance on mutually independent clusters precludes its application to the dyadic configurations studied here.
Our contribution bridges these two frameworks. We retain the dyadic dependence structure formalized by \citet{tabordmeehan2019inference}, but dispense with the common-rate requirement by adapting the rate-free logic of \citet{hansen2019asymptotic} to a setting without mutually independent clusters. Theorem \ref{thm:main} establishes asymptotically valid Wald inference without imposing a common scalar convergence rate. The Wald statistic is asymptotically $\chi^2_q$ for an \emph{arbitrary} nonrandom sequence of full-rank restriction matrices, subject to a single condition on the accumulation of dependence. This condition assumes no scalar convergence rate and permits different linear combinations of the estimator to converge at heterogeneous rates. This matters in practice, since heterogeneous rates are the norm on dyadic designs often used in the trade literature: a coefficient on an exporter attribute and one on a trading-pair attribute accumulate dependence at different speeds, and a researcher running a gravity regression cannot be asked to know, or estimate, either rate.
While Theorem \ref{thm:main} establishes the asymptotic properties of the dyadic-robust variance estimator, our second contribution addresses its finite-sample behavior. Because its inner matrix is a difference of two positive semidefinite matrices, the estimator is not guaranteed to be positive semidefinite at a given sample size, which can render the resulting Wald statistic undefined. We therefore study a jackknife alternative that deletes one \emph{unit}, and with it every observation that unit touches, at a time. The construction is that of
\citet{snijders1999nonparametric}, rescaled to consistency for a dyadic M-estimator in \citet{graham2020network}; see also \citet{bell2002bias} and
\citet{lin2020theoretical} for similar constructions in other settings. It is a sum of outer products, and is therefore positive semidefinite. Because each observation belongs to exactly two deletion groups, however, the classical argument over disjoint independent clusters is unavailable. Theorem \ref{thm:jack} shows that the jackknife Wald statistic attains the same $\chi^2_q$ limit under one further condition on dyad multiplicity, and Proposition \ref{prop:onesided} shows that without that condition the failure is one-sided, in the sense that the test never over-rejects asymptotically.
Finally, we provide three technical contributions that might be of independent interest. First, we establish a central limit theorem for the least squares score standardized by its exact variance rather than by a rate; this accommodates triangular arrays where the tested direction changes with the sample size. Second, we prove that the dyadic-robust variance estimator is consistent relative to the true variance uniformly over the restrictions tested, permitting its linear combinations to converge at heterogeneous rates. Third, we derive an exact finite-sample identity for the effect of deleting one unit on the least squares estimator. Because each observation belongs to exactly two deletion groups, this identity reveals that the first-order terms sum exactly to zero, rendering the difference between the jackknife mean and the full-sample estimator purely of second order.
The remainder of the paper is organized as follows. Section \ref{sec:model} presents the model, the dyadic configuration notation, and the two variance estimators. Section \ref{sec:assumptions} states the main results and compares our assumptions with the existing literature, and Section \ref{sec:conclusion} concludes. All proofs are deferred to the supplemental materials. The supplement also provides three worked-out illustrations and a numerical exercise evaluating a gravity equation across three dyadic configurations: directed Erd\H{o}s--R\'{e}nyi configurations of varying density, multilayer configurations featuring heterogeneous convergence rates, and an empirical application to the bilateral trade data of \citet{santossilva2006log}.
\section{Model and Estimators}
\label{sec:model}
Units are indexed $g=1,\dots,G$ and all limits are taken as $G\to\infty$. Observations are
indexed $n=1,\dots,N$ with $N=N_G$. Each carries a support $\psi(n)\subset
\{1,\dots,G\}$ with $|\psi(n)|=2$, called its dyad, and $\psi$ is \emph{not} assumed
injective, so several observations may carry the same dyad. Write $M_g :=
\#\{n : g\in\psi(n)\}$ for the number of observations touching unit $g$ (observations, not
distinct partners), $\mathcal{M}^{H} := \max_g M_g$ and $\mathcal{M}^{L} := \min_g M_g$. Dually, for a dyad $d$, $L_d := \#\{n : \psi(n)=d\}$, $\mathcal{L} := \max_d L_d$, and
$\mathcal{J}_N := \sum_d L_d^2$, the number of ordered pairs of observations carrying the
same dyad. Since each observation contributes to exactly two of the $M_g$ and $\sum_d L_d
= N$,
\begin{equation}\label{eq:degrees}
\frac{\mathcal{M}^{L}G}{2} \;\le\; N \;\le\; \frac{\mathcal{M}^{H}G}{2},
\qquad\qquad
N \;\le\; \mathcal{J}_N \;\le\; \mathcal{L}N, \qquad \mathcal{L}\le\mathcal{M}^{H}\le N .
\end{equation}
Let $\mathbf{1}_{nm} := \mathbf{1}\{\psi(n)\cap\psi(m)\ne\emptyset\}$ and $\mathcal{A} :=
\{(n,m) : \mathbf{1}_{nm}=1\}$. For matrices, $\|\cdot\|$ and $\|\cdot\|_F$ denote the
spectral and Frobenius norms, $\lambda_{\min}$ and $\lambda_{\max}$ extreme eigenvalues,
and for symmetric positive definite $A$, $A^{1/2}$ is the symmetric square root with
$A^{-1/2} := (A^{1/2})^{-1}$. Finally, throughout, $C := C_x^2C_u^2$ denotes the constant
built from Assumption \ref{as:bound} below.
The linear model is $y_n = \boldsymbol{\beta}'\mathbf{x}_n + u_n$, with
$\mathbf{x}_n\in\mathbb{R}^{K}$ and $K$ fixed, and we estimate it by ordinary least
squares. Let $\mathcal{E}_N$ denote the event that $\sum_n \mathbf{x}_n\mathbf{x}_n'$ is
invertible, and write $\mathcal{E}_N^{c}$ for its complement. On $\mathcal{E}_N$ we set
$\widehat{\boldsymbol{\beta}} := (\sum_n\mathbf{x}_n\mathbf{x}_n')^{-1}
\sum_n\mathbf{x}_ny_n$ and $\widehat{u}_n := y_n -
\widehat{\boldsymbol{\beta}}'\mathbf{x}_n$, while on $\mathcal{E}_N^{c}$ we set
$\widehat{\boldsymbol{\beta}} := \mathbf{0}$ and $\widehat{u}_n := y_n$. Define the score,
its variance, and the design matrices $S_N := \sum_n\mathbf{x}_nu_n$, $\Omega_N :=
\operatorname{Var}(S_N)$, $Q_N := N^{-1}\sum_n\mathbb{E}[\mathbf{x}_n\mathbf{x}_n']$, and
$\widehat{Q}_N := N^{-1}\sum_n\mathbf{x}_n\mathbf{x}_n'$. The \emph{dyadic-robust variance
estimator} of $\widehat{\boldsymbol{\beta}}$ is built
from the feasible and infeasible meats
\begin{equation}\label{eq:meat}
\widehat{\Omega}_N \;:=\; \sum_{n=1}^{N}\sum_{m=1}^{N} \mathbf{1}_{nm}\,
\widehat{u}_n\widehat{u}_m\,\mathbf{x}_n\mathbf{x}_m', \qquad
\widetilde{\Omega}_N \;:=\; \sum_{n=1}^{N}\sum_{m=1}^{N} \mathbf{1}_{nm}\,
u_nu_m\,\mathbf{x}_n\mathbf{x}_m',
\end{equation}
through the sandwich
\begin{equation}\label{eq:vhat}
\widehat{V}_N \;:=\; \frac{1}{N^{2}}\,\widehat{Q}_N^{-1}\widehat{\Omega}_N
\widehat{Q}_N^{-1}, \qquad\qquad V_N \;:=\; \frac{1}{N^{2}}\,Q_N^{-1}\Omega_NQ_N^{-1} .
\end{equation}
Since $\widehat{Q}_N$ is singular on $\mathcal{E}_N^{c}$, we set $\widehat{V}_N:=\mathbf{0}$ there. Alternatively, one can use a Moore--Penrose inverse \citep[see, e.g.,][p.~144]{graham2020network}; either approach is valid because
$\mathbb{P}(\mathcal{E}_N)\to1$ under Assumptions \ref{as:dep}--\ref{as:config}.
Here $\widehat{V}_N$ estimates $V_N$, which standardizes
$\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta}$ in Theorem \ref{thm:main} below.
The second variance estimator is the aforementioned
delete-one-unit jackknife, i.e., $\widehat{V}^{J}$. Write $A := \sum_n\mathbf{x}_n\mathbf{x}_n'$,
$\mathcal{N}_g := \{n : g\in\psi(n)\}$, $A_g := \sum_{n\in\mathcal{N}_g}
\mathbf{x}_n\mathbf{x}_n'$ and $A_{(g)} := A - A_g$; on the event $\mathcal{E}^{J}_N$ that
every $A_{(g)}$ is invertible, let $\widehat{\boldsymbol{\beta}}_{(g)} :=
A_{(g)}^{-1}\sum_{n\notin\mathcal{N}_g}\mathbf{x}_ny_n$,
$\bar{\boldsymbol{\beta}} := G^{-1}\sum_g\widehat{\boldsymbol{\beta}}_{(g)}$, and
\begin{equation}\label{eq:vjack}
\widehat{V}^{J} \;:=\; \frac{G-1}{G}\sum_{g=1}^{G}
\big(\widehat{\boldsymbol{\beta}}_{(g)} - \bar{\boldsymbol{\beta}}\big)
\big(\widehat{\boldsymbol{\beta}}_{(g)} - \bar{\boldsymbol{\beta}}\big)' ,
\end{equation}
with $\widehat{V}^{J} := \mathbf{0}$ on $(\mathcal{E}^{J}_N)^{c}$ as well. Deleting unit $g$
removes all $M_g$ observations touching it at once. Two features of \eqref{eq:vjack}
matter below. First, $\widehat{V}^{J}$ is positive semidefinite by construction, which
$\widehat{\Omega}_N$ is not (see, e.g., \citealp{mackinnon2023cluster}, for further
details). Second, each observation belongs to exactly two of the $G$ deletion groups, so
the classical disjoint-cluster argument \citep{cameron2011robust, hansen2019asymptotic}
cannot be used in the mathematical proofs.
\section{Assumptions and Main Results}
\label{sec:assumptions}
We consider the following regularity conditions.
\begin{assumption}[Dyadic dependence]\label{as:dep}
For any two disjoint
index sets $S_1,S_2\subset\{1,\dots,N\}$, the families
$\{(\mathbf{x}_n,u_n)\}_{n\in S_1}$ and $\{(\mathbf{x}_m,u_m)\}_{m\in S_2}$ are
independent whenever $\psi(n)\cap\psi(m)=\emptyset$ for every $n\in S_1$ and $m\in S_2$.
\end{assumption}
\begin{assumption}[Bounded support]\label{as:bound}
There are finite constants $C_x,C_u$ with $\|\mathbf{x}_n\|\le C_x$ and
$|u_n|\le C_u$ almost surely for all $n$ and $N$.
\end{assumption}
\begin{assumption}[Identification]\label{as:ident}
$\mathbb{E}[\mathbf{x}_nu_n]=\mathbf{0}$ for all $n$, and
$\lambda_{\min}(Q_N)\ge c_Q>0$ for all $N$.
\end{assumption}
\begin{assumption}[Variance accumulation]\label{as:acc}
$\Omega_N$ is positive definite and\newline
$\delta_N := N(\mathcal{M}^{H})^{3}\big/\lambda_{\min}(\Omega_N)^{2}\longrightarrow 0$.
\end{assumption}
\begin{assumption}[Configuration]\label{as:config}
$K$ is fixed and $\mathcal{M}^{L}\ge1$ as $G\to\infty$.
\end{assumption}
Assumption \ref{as:dep} is the dyadic analogue of independent clustering, and it is
a weaker version of Assumption 2.1 in \citet[p.~672]{tabordmeehan2019inference} in that the support map need not be injective and the observations need not be identically distributed. Note that only the independence of families
indexed by node-disjoint dyads is ever used below. This assumption therefore imposes restrictions only on observations that do not share a unit. Consequently, any two observations sharing a common unit, whether they represent two orientations of a single pair, distinct network layers, or different time periods, may be arbitrarily dependent. In particular, reciprocity and cross-layer correlation are left unrestricted. Assumption \ref{as:dep} also generalizes the sampling scheme of \citet[Section 2, p.~269]{hansen2019asymptotic} because it does not require clusters to be disjoint. In their framework, clusters are mutually independent by construction. In contrast, the natural clusters $\{n:g\in\psi(n)\}$ in our setting overlap, with each observation belonging to exactly two of them. Their scheme can be recovered by placing each mutually independent cluster on its own isolated pair of units. However, Assumption \ref{as:dep} does rule out dependence across node-disjoint dyads. Consequently, any shock common to all dyads must be absorbed into the mean, either by conditioning on its realization or by including it as a nonrandom shifter, rather than subsumed in the error term.
Assumption \ref{as:ident} is standard in the linear regression literature; see, for example, the conditions in \citet[p.~672]{tabordmeehan2019inference}. It only requires the regressors and the error term to be uncorrelated, rather than imposing strict mean independence $\mathbb{E}[u_n\mid\mathbf{x}_n]=0$. Thus, $\boldsymbol{\beta}$ is defined as a projection parameter rather than a conditional-mean coefficient. It is important to note, however, that Assumption \ref{as:ident} imposes orthogonality observation by observation. When the data are not identically distributed, this condition is stronger than the aggregate restriction $\sum_n\mathbb{E}[\mathbf{x}_nu_n]=\mathbf{0}$ implied by a pooled population projection. These individual restrictions are precisely what Lemma \ref{lem:struct}(a) uses to ensure the raw cross-products coincide with their covariances and that node-disjoint cross-products vanish, thereby identifying $\mathbb{E}[\widetilde{\Omega}_N]$ with $\Omega_N$. The estimand is therefore a common coefficient satisfying pair-specific orthogonality, and it is imposed as such in the empirical application in Appendix \ref{sec:empirical}.
Similarly, Assumption \ref{as:bound} is the bounded support condition of Assumption 3.1 in
\citet[p.~675]{tabordmeehan2019inference}. It is stronger than the moment conditions
under which cluster-robust limit theory is usually developed, and it is used here to bound
the summands of $\Omega_N$ uniformly, which with the pair count of Lemma
\ref{lem:struct}(b) gives $\lambda_{\max}(\Omega_N)=O(N\mathcal{M}^{H})$. That cap is what permits the condition number of $\Omega_N$ to
diverge in Theorem \ref{thm:main} below. When clusters are disjoint the same control
follows from bounded moments alone, and it does not when they overlap. Notice that Assumption
\ref{as:bound} restricts the support of $(\mathbf{x}_n,u_n)$ and nothing else about its
distribution, so it imposes neither homoskedasticity nor any parametric form. Assumption
\ref{as:config} requires that no unit be isolated and that $K$ stay fixed as $G$ grows.
That is, it rules out node fixed effects and slope vectors whose dimension grows with
$G$; see Illustration
3 in \ref{sec:smb}.
Assumption \ref{as:acc} is more substantive. It replaces Assumption 2.6 in
\citet[p.~675]{tabordmeehan2019inference}, which requires $(NG^{r})^{-1}\Omega_N$ to
converge to a positive-definite limit for some exponent $r\in[0,1]$. No such rate is
imposed here,
and no normalization of $\Omega_N$ is required to converge. What is asked instead is
positive definiteness at each $N$ together with $\delta_N\to0$. Dispensing with the limit
is what lets Theorem \ref{thm:main} below test joint restrictions whose directions
accumulate dependence at genuinely different orders, and Appendix \ref{sec:mc2} reports a
design in which they do. Its second part, namely
$\delta_N\to0$, is similar to the eigenvalue condition in Theorem 9 of
\citet[p.~277]{hansen2019asymptotic}, where the smallest eigenvalue of a normalized
cluster variance is bounded away from zero, but stronger on dense configurations, in the sense that
$\lambda_{\min}(\Omega_N)$ must outgrow $\{N(\mathcal{M}^{H})^{3}\}^{1/2}$, a bound that
exceeds the order $N$ asked for there exactly when $(\mathcal{M}^{H})^{3}>N$. The numerator
$N(\mathcal{M}^{H})^{3}$ counts the overlapping index quadruples that control the
variance of the meat (see Lemma \ref{lem:struct}(c)), a count that is of order $N$ when
clusters are bounded and independent, and grows with the maximum degree once they
overlap.
Notice that Assumption \ref{as:acc} asks nothing about the rate at which
$\widehat{\boldsymbol{\beta}}$ converges, and no such rate is either assumed or estimated
anywhere below. The matrix $\Omega_N$ need not converge under any normalization, and its
condition number $\kappa(\Omega_N):=\lambda_{\max}(\Omega_N)/\lambda_{\min}(\Omega_N)$ is
free to diverge, so that different linear combinations of $\widehat{\boldsymbol{\beta}}$
may converge at different rates, as the multilayer designs of Illustration 2 in
\ref{sec:smb} exhibit. The spread of those rates is nevertheless
controlled, since $\delta_N\to0$ implies
$\kappa(\Omega_N)=o\big((N/\mathcal{M}^{H})^{1/2}\big)$. That is, \emph{rate-agnostic}
refers to what has to be known in order to conduct inference, and not to how far apart
the underlying rates may lie. If the design is nonrandom and the errors satisfy
$\operatorname{Var}(\mathbf{u})\succeq\underline{\sigma}^{2}I_N$ for some
$\underline{\sigma}^{2}>0$, then $\Omega_N\succeq\underline{\sigma}^{2}\sum_n
\mathbf{x}_n\mathbf{x}_n'$ and hence $\lambda_{\min}(\Omega_N)\ge
\underline{\sigma}^{2}c_QN$ by Assumption \ref{as:ident}, so that Assumption
\ref{as:acc} holds automatically whenever
$(\mathcal{M}^{H})^{3}=o(N)$, a requirement on the configuration alone which by
\eqref{eq:degrees} entails $\mathcal{M}^{H}=o(\sqrt{G})$, and is therefore available on
sparse designs. On denser
ones it is a joint restriction on the configuration, the design and the error structure. The following theorem establishes the consistency of the sandwich estimator in
\eqref{eq:vhat} and the limiting distribution of the implied Wald statistic.
\begin{theorem}\label{thm:main}
Let Assumptions \ref{as:dep}--\ref{as:config} hold, and let $\{R_N\}$ be an arbitrary
sequence of nonrandom $K\times q$ matrices of full column rank $q\le K$. Write
$\mathcal{V}_N := R_N'V_NR_N$ and $\widehat{\mathcal{V}}_N := R_N'\widehat{V}_NR_N$. Then, as $G\to\infty$:
(a) $\mathcal{V}_N^{-1/2}R_N'(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta})
\longrightarrow_d N(\mathbf{0},I_q)$;
(b) $\mathcal{V}_N^{-1/2}\widehat{\mathcal{V}}_N\mathcal{V}_N^{-1/2}\longrightarrow_p I_q$;
(c) $\mathbb{P}(\widehat{\mathcal{V}}_N\succ0)\to1$, and on that event
$\widehat{\mathcal{V}}_N^{-1/2}R_N'(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta})
\longrightarrow_d N(\mathbf{0},I_q)$;
(d) under $H_0: R_N'\boldsymbol{\beta}=\mathbf{r}$, the Wald statistic
\[
W_N := \big(R_N'\widehat{\boldsymbol{\beta}}-\mathbf{r}\big)'\widehat{\mathcal{V}}_N^{-1}
\big(R_N'\widehat{\boldsymbol{\beta}}-\mathbf{r}\big) \;\longrightarrow_d\; \chi^2_q .
\]
\end{theorem}
Note that parts (a) and (b) are statements on $\mathcal{E}_N$, whose probability tends to
one, and that part (b) delivers $\mathbb{P}(\widehat{\mathcal{V}}_N\succ0)\to1$, which is the event
on which parts (c) and (d) are stated. As previously pointed out, the sequence $\{R_N\}$
is arbitrary, and the bounds behind parts (a) and (b) do not involve it, so both
convergences hold uniformly over such sequences. Uniformity extends to the size statement
in part (d), over classes of data generating processes sharing the constants
$C_x$, $C_u$, $c_Q$ and $K$ of
Assumptions \ref{as:bound}, \ref{as:ident} and \ref{as:config} and a common vanishing
bound on $\delta_N$ (see the end of the proof of Theorem 1 in
\ref{sec:sma}).
Theorem \ref{thm:main} offers two advantages over available results. First, it dispenses
with independent clusters. Theorem 9 in \citet[p.~277]{hansen2019asymptotic} is also
rate-free, but it requires the observations to be partitioned into mutually independent
clusters, and a dyadic configuration need not admit a
growing number of them. Second, it tests joint restrictions
whose directions converge at different orders. The dyadic-robust $t$-statistic of
Theorem 3.1 in \citet[p.~676]{tabordmeehan2019inference} studentizes a single
coefficient under one common
convergence rate. Our proof of Theorem \ref{thm:main} uses the central limit theorem for dependency graphs of Theorem 2 in \citet[p.~307]{janson1988normal}. The following condition on dyad
multiplicity makes the jackknife estimator in
\eqref{eq:vjack} attain the same limit as the dyadic-robust one.
\begin{assumption}[Dyad multiplicity]\label{as:mult}
$\mathcal{J}_N\big/\lambda_{\min}(\Omega_N)\longrightarrow 0$.
\end{assumption}
Assumption \ref{as:mult} is introduced here for the jackknife alone. Since $\mathcal{J}_N\ge N$ by
\eqref{eq:degrees}, it imposes that $\Omega_N$ accumulate strictly faster than under
independent sampling, and it is the more demanding the more observations a single dyad
carries.
\begin{theorem}\label{thm:jack}
Let Assumptions \ref{as:dep}--\ref{as:mult} hold, and let $\{R_N\}$ be as in Theorem
\ref{thm:main}. With $\mathcal{V}_N:=R_N'V_NR_N$ and $\widehat{\mathcal{V}}^{J}_N := R_N'\widehat{V}^{J}R_N$,
as $G\to\infty$: (a) $\mathcal{V}_N^{-1/2}\widehat{\mathcal{V}}^{J}_N\mathcal{V}_N^{-1/2}\longrightarrow_p I_q$,
uniformly over restriction sequences as in Theorem \ref{thm:main}(b); (b)
$\widehat{\mathcal{V}}^{J}_N\succeq0$ on $\mathcal{E}^{J}_N$, and
$\mathbb{P}(\mathcal{E}^{J}_N)\to1$, $\mathbb{P}(\widehat{\mathcal{V}}^{J}_N\succ0)\to1$; (c) on
that event, $(\widehat{\mathcal{V}}^{J}_N)^{-1/2}R_N'(\widehat{\boldsymbol{\beta}}-\boldsymbol{\beta})
\longrightarrow_d N(\mathbf{0},I_q)$; (d) under $H_0: R_N'\boldsymbol{\beta}=\mathbf{r}$,
\[
W^{J}_N := \big(R_N'\widehat{\boldsymbol{\beta}}-\mathbf{r}\big)'
\big(\widehat{\mathcal{V}}^{J}_N\big)^{-1}\big(R_N'\widehat{\boldsymbol{\beta}}-\mathbf{r}\big)
\;\longrightarrow_d\; \chi^2_q .
\]
\end{theorem}
It is important to note that part (b) of Theorem \ref{thm:jack} is a finite-sample property rather than an asymptotic statement. Positive semidefiniteness holds at every $N$ by construction, following directly from the outer-product form of \eqref{eq:vjack}. In contrast, Theorem \ref{thm:main}(c) lacks this finite-sample guarantee, since
$\widehat{\Omega}_N$ is a difference of two positive semidefinite matrices (see Lemma \ref{lem:leaveout}(d) in the supplement), and Assumptions \ref{as:dep}--\ref{as:config} leave this difference unrestricted. Therefore it is instructive to precisely delineate the role of Assumption \ref{as:mult}. The positive semidefiniteness in part (b) does not rely on it, holding unconditionally at every $N$. What Assumption \ref{as:mult} actually delivers is the asymptotic convergence established in parts (a) and (d). Without this condition, the jackknife estimator need not symmetrically attain
$\mathcal{V}_N$, leaving in its place only the one-sided inequality described by Proposition \ref{prop:onesided} below. Consequently, the following proposition formally establishes that the error of the jackknife estimator is strictly one-sided.
\begin{proposition}\label{prop:onesided}
Let Assumptions \ref{as:dep}--\ref{as:config} hold and let $\{R_N\}$ be as in Theorem
\ref{thm:main}. Then there is a sequence $\varepsilon_N=o_p(1)$, not depending on $R_N$,
such that
$\lambda_{\min}\big(\mathcal{V}_N^{-1/2}\widehat{\mathcal{V}}^{J}_N\mathcal{V}_N^{-1/2}\big)\ge1-\varepsilon_N$.
Consequently, for every $\gamma\in(0,1)$,
$\limsup_{G\to\infty}\mathbb{P}\big(W^{J}_N>\chi^2_{q,1-\gamma}\big)\le\gamma$ under
$H_0:R_N'\boldsymbol{\beta}=\mathbf{r}$.
\end{proposition}
That a jackknife variance estimator errs upward is classical, going back to \citet[p.~590]{efron1981jackknife}, and it has been established for network functionals by \citet[p.~6107]{lin2020theoretical}. Note that the only difference between Proposition \ref{prop:onesided} and those results is that the error is signed in the positive semidefinite order, and that the bound holds uniformly over the restrictions tested. Proposition \ref{prop:onesided} does not impose Assumption \ref{as:mult}, which is what turns its inequality into the equality of Theorem \ref{thm:jack}(a). In particular, if Assumption \ref{as:mult} fails, a test that rejects when $W^{J}_N>\chi^2_{q,1-\gamma}$ still has asymptotic size no larger than $\gamma$.
Finally, note that Assumptions \ref{as:dep}--\ref{as:config} make $\mathbb{P}(\mathcal{E}_N)\to1$, but not $\mathcal{E}_N$ an almost-sure event at any fixed $N$. We therefore set $W_N:=0$ and $W^{J}_N:=0$ whenever the matrix required to compute each statistic is singular. Both statistics are then
defined on every sample, and the asymptotic size claims of Theorem \ref{thm:main}(d), Theorem \ref{thm:jack}(d) and Proposition \ref{prop:onesided} are ordinary probability statements; see, e.g., Remark 3.4 in \citet[p.~676]{tabordmeehan2019inference} for other alternatives.
\section{Discussion and Concluding Remarks}
\label{sec:conclusion}
In this paper we have considered Wald inference for linear regression on dyadic data in which several observations may share the same pair of units. We have shown that the test built from the dyadic-robust variance estimator is asymptotically valid for an arbitrary sequence of full-rank nonrandom joint restrictions, under a single condition on the accumulation of
dependence and with no convergence rate assumed or estimated. We have also shown that its delete-one-unit jackknife counterpart, which is positive semidefinite at every sample size, attains the same $\chi^2_q$ limiting distribution under one further condition on dyad multiplicity, and that it never over-rejects asymptotically when that condition fails. The
practical content is that the estimator applied researchers already compute
\citep{fafchamps2007formation,aronow2015cluster,graham2020network} supports joint hypothesis testing exactly as computed, on directed, multilayer, and panel dyadic designs alike, without the common-rate normalization that its existing formal justification requires. Appendix \ref{sec:mc2} puts that last point on a design in which the coefficients tested jointly do converge at three different rates, and finds the size of the test governed by the number of units rather than by the spread of those rates. The analogous rate-free result for clustered data is Theorem 9 of \citet[p.~277]{hansen2019asymptotic}, which relies fundamentally on mutually independent clusters. Since dyadic data inherently violate this independence, our Theorems \ref{thm:main} and \ref{thm:jack} substitute it with a dyadic dependency graph. The two sets of conditions are not nested: while their sampling scheme is a special case of our dyadic configuration, Assumptions \ref{as:bound} and \ref{as:acc} neither imply nor are implied by their cluster-size and moment regimes.
Three extensions are left for future research. The first is nonlinear estimation. The
Poisson pseudo-maximum-likelihood estimator of \citet{santossilva2006log} has a score that
is again a sum of dyad-indexed terms, so the route to Theorem \ref{thm:main} appears open.
The leave-unit-out algebra behind Theorem \ref{thm:jack} would there be replaced by an
asymptotic expansion. The second is
resampling. A bootstrap counterpart to the tests developed here would have to reproduce
dependence accumulating at heterogeneous and unknown rates, and its first-order
validity would require a separate argument. The third is growing
dimension. Node fixed effects, or layer-specific slopes with a growing number of layers,
violate the fixed-$K$ part of Assumption \ref{as:config}, and every bound in the
supplement would need to be made uniform in $K$. Finally, our analysis conditions on the observed network configuration throughout. Modeling the underlying network formation process represents a fundamentally distinct problem that remains beyond the scope of this paper.
\spacingset{1}
\putbib[references]
\end{bibunit}
\clearpage
\setcounter{page}{1}
\setcounter{section}{0}
\setcounter{equation}{0}
\setcounter{table}{0}
\setcounter{remark}{0}