EconBase
← Back to paper

Learning Dependence Structures for Econometric Inference

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

76,925 characters

Learning Dependence Structures for Econometric Inference




\maketitle

\begin{abstract}
We develop a framework for learning dependence structures from empirical dependence operators. Rather than treating cluster, factor, and sparse dependence as maintained assumptions, we represent them as covariance geometries in a common Hilbert space and summarize dependence through a low-dimensional dependence profile based on projection similarity scores. We establish identification under a principal-angle separation condition, prove consistency and asymptotic normality of the estimated profile, and derive finite-sample classification error bounds. We further show that when covariance-geometry tangent spaces overlap, no statistical procedure can distinguish the geometries at first order, providing a formal characterization of ambiguous dependence structures. Projection-residual diagnostics assess absolute goodness-of-fit and detect misspecified covariance dictionaries. Finally, we establish oracle adaptivity of profile-guided inference: dependence profiles can be used to select dependence-robust procedures in a data-driven manner, yielding inference that is asymptotically equivalent to an infeasible oracle that knows the dominant covariance geometry in advance.


\bigskip
\noindent
\medskip{}
\noindent \textbf{Keywords}:
Dependence learning;
covariance geometry;
dependence diagnostics;
dependence classification;
factor dependence;
multiway clustering.

\vspace{0.5cm}

\noindent \textbf{JEL Classification}:
C12, C13, C14, C38.
\end{abstract}


\section{Introduction}
\label{sec:introduction}

Robust inference under dependence is one of the central concerns of modern econometrics. Economic data frequently exhibit dependence arising from common institutional environments, aggregate shocks, social interactions, geographic proximity, production networks, financial interconnectedness, and repeated observations over time.

Prominent examples include cluster-robust inference \citep{White1980,LiangZeger1986,Arellano1987,CameronGelbachMiller2011,CameronMiller2015}, spatial and HAC-type dependence \citep{Conley1999,DriscollKraay1998}, factor-based methods \citep{ChamberlainRothschild1983,BaiNg2002,Bai2003}, and network or sparse dependence models \citep{BickelLevina2008,CaiLiu2011,Auerbach2019,Leung2022}. Although these approaches differ substantially in their assumptions, they share a common feature: the dependence structure is treated as known.

In practice, however, the underlying dependence architecture is rarely known. Firms may be exposed to common industry shocks, latent macroeconomic factors, and production-network spillovers. Regional economic outcomes may reflect state-level institutions, common national shocks, geographic spillovers, migration networks, and local interactions. The central challenge is often not merely robust inference under dependence but learning the dependence structure itself.

\paragraph{Motivating examples.}

Consider a panel of firms observed over time.
Residual dependence may arise because firms are exposed to common
industry shocks, latent aggregate factors, or production-network
linkages. These mechanisms imply very different covariance
geometries---cluster, factor, and sparse dependence---yet are often
simultaneously plausible in practice.

As a second example, consider regional economic outcomes.
Dependence may reflect state-level institutions, common national
shocks, geographic spillovers, or migration networks. Standard
robust inference procedures typically require the researcher to
specify one of these dependence structures in advance.

In both examples, the primary challenge is not merely to conduct
inference under dependence, but to learn which dependence geometry
is most consistent with the observed data.

This paper develops a framework for dependence structure learning. Rather than viewing cluster, factor, and sparse dependence as competing models, we view them as distinct geometric classes of covariance operators. Cluster dependence generates approximately structured-support covariance structures. Factor dependence generates approximately low-rank
covariance structures. Sparse dependence generates covariance structures in which dependence is concentrated among a relatively small subset of economically meaningful interactions.

The central object of interest is not a covariance estimator. Instead, the framework starts from an empirical dependence operator constructed from the sampling design. Depending on the application, this operator may correspond to a sample covariance matrix, a long-run covariance operator, a spatial dependence operator, a network dependence operator, or another empirical measure of dependence. The objective is to learn the geometry of dependence encoded in this operator.

The paper does not seek to estimate dependence by imposing a cluster, factor, or sparse model. Rather, it assumes the existence of a population dependence operator and studies what aspects of dependence geometry can be learned from a consistent estimate of that operator.

A key distinction between the present framework and existing approaches is that the object of interest is not the dependence operator itself. Even when a researcher can consistently estimate an appropriate dependence operator, the resulting object often remains high dimensional and difficult to interpret. The goal of dependence learning is therefore to recover the geometric organization encoded in the operator. The dependence profile provides a low-dimensional summary that quantifies the extent to which the operator resembles cluster, factor, and sparse covariance geometries. In this sense, dependence structures are treated as estimands rather than maintained assumptions.

Importantly, the objective is not necessarily to recover a unique dependence structure. Different covariance geometries may overlap locally, making ambiguity an inherent feature of the problem rather than a consequence of limited sample size. Dependence learning therefore seeks both to identify dominant dependence patterns and to diagnose situations in which multiple geometries provide similarly plausible descriptions of the observed dependence operator.

An analogy may help clarify the objective. As principal component analysis uses the covariance matrix to recover a low-dimensional representation of variation, our framework uses an empirical dependence operator to recover a low-dimensional representation of dependence geometry. The operator serves as an intermediate object through which the underlying dependence structure is learned.

An important feature of the framework is that it distinguishes between relative and absolute notions of dependence fit. A large similarity score indicates that one covariance geometry provides a better approximation than the alternatives under consideration, but it does not imply that the geometry provides a good approximation in an absolute sense. To address this issue, we introduce projection-residual diagnostics that measure the distance between the empirical dependence operator and the closest covariance geometry. These diagnostics allow the researcher to distinguish between clear geometric classifications, ambiguous dependence structures arising from overlapping covariance geometries, and situations in which none of the candidate geometries adequately captures the observed dependence pattern.

A central theme of the paper is the distinction between identifiable and ambiguous dependence structures. Identification is governed by the geometric separation of covariance classes. When covariance geometries are sufficiently separated, dependence profiles are locally identifiable and support reliable classification. When covariance geometries overlap, however, ambiguity may be intrinsic rather than a consequence of limited sample size. We show that the principal-angle condition, applied to the off-diagonal (dependence-relevant) part of each tangent space, provides a geometric identification criterion: positive principal angles ensure local distinguishability,
whereas overlapping tangent spaces generate local indistinguishability.  Theorem~\ref{thm:identification} characterizes the identifiable region of the covariance-geometry space, while Theorem~\ref{thm:two_geometry_indistinguishability} characterizes the boundary at which identification fails. Together, the two results delineate the boundary between identifiable and fundamentally ambiguous dependence structures.

The proposed dependence profile is also conceptually related to several familiar ideas in econometrics. In factor models, researchers often summarize variation through factor contribution measures. In forecast combination and model averaging, weights summarize the relative importance of competing predictive models \citep{BatesGranger1969,Hansen2007}. In semiparametric efficiency theory, projection operators characterize efficient influence functions through orthogonal projections onto tangent spaces \citep{Newey1994,BickelKlaassenRitovWellner1993}. The dependence profile serves a similar descriptive purpose: it does not represent an additive covariance decomposition, but summarizes the relative geometric proximity of the observed dependence operator to economically meaningful dependence classes.

Existing robust inference methods require the researcher to specify a dependence structure in advance. Our framework instead quantifies the extent to which alternative covariance geometries are supported by the data. To formalize this idea, we introduce projection operators onto cluster, factor, and sparse covariance classes, which induce a
dependence profile $(\omega_C,\omega_F,\omega_S)$. These classes should be viewed as illustrative benchmark geometries
rather than an exhaustive collection.

A central objective of the paper is not only to learn dependence structures, but also to use this information to guide econometric inference. We show that estimated dependence profiles can be used to construct profile-guided procedures that adapt to the dominant dependence geometry. Under suitable conditions, the resulting procedures are asymptotically equivalent to an infeasible oracle that knows the dominant covariance geometry in advance. Simulation evidence shows that profile-guided inference closely approximates the infeasible oracle procedure across a broad range of dependence environments.

The paper's theoretical contribution is organized around three
principal results. Theorem~\ref{thm:identification} establishes
local identification of dependence profiles under a principal-angle
separation condition. Theorem~\ref{thm:two_geometry_indistinguishability}
shows that when covariance geometries share common tangent
directions, no statistical procedure can distinguish them at first
order. Finally, Theorem~\ref{thm:oracle_adaptivity} establishes
oracle adaptivity of profile-guided inference, showing that
dependence learning can be used to construct inference procedures
that are asymptotically equivalent to an infeasible oracle that
knows the dominant covariance geometry in advance. The paper makes five contributions.

\textit{Geometric framework.}
We represent dependence structures as covariance geometries---closed subsets
of the Hilbert space of symmetric matrices---and characterize their local
structure through tangent spaces and principal angles.
Cluster, factor, and sparse dependence are treated as geometric classes
rather than competing maintained assumptions. The framework
itself is not tied to these particular classes and can accommodate other
covariance geometries, including spatial, network, long-memory, or other
application-specific dependence structures.

\textit{Identification and estimation.}
We establish local identification of the dependence profile under a
principal-angle separation condition, prove consistency, and derive
asymptotic normality of the estimated dependence profile.
Projection-residual diagnostics provide a complementary absolute
goodness-of-fit measure.

\textit{Classification theory.}
We establish consistency of dominant-geometry classification
(Theorem~\ref{thm:dominant_geometry_consistency}), derive finite-sample error
bounds based on the separation margin~$\Delta_\omega$
(Proposition~\ref{prop:classification_bound}), and characterize classification
probabilities under local near-tie alternatives
(Theorem~\ref{thm:local_ties}).

\textit{Geometric ambiguity.}
We show that when two covariance-geometry tangent spaces share a common
direction, no sequence of tests can distinguish the geometries at first order
(Theorem~\ref{thm:two_geometry_indistinguishability}). This impossibility result explains ambiguous dependence profiles and near-ties as consequences of local nonidentification rather than finite-sample noise.

\textit{Adaptive inference.} We establish oracle adaptivity of
profile-guided inference (Theorem~\ref{thm:oracle_adaptivity}).
Estimated dependence profiles can be used to select
dependence-robust procedures in a data-driven manner, and the
resulting inference is asymptotically equivalent to an infeasible
oracle that knows the dominant covariance geometry in advance.



\subsection{Related Literature}
\label{subsec:literature}

This paper contributes to several strands of econometrics. First, it relates to inference under dependence, including clustered dependence \citep{LiangZeger1986,Arellano1987,CameronGelbachMiller2011}, spatial dependence \citep{Conley1999,DriscollKraay1998}, network dependence \citep{Auerbach2019,Leung2022}, and factor dependence \citep{ChamberlainRothschild1983,BaiNg2002,Bai2003}. These methods generally assume that the relevant dependence architecture is known.

Second, the paper relates to covariance regularization and high-dimensional dependence modeling. Sparse covariance estimation exploits localized dependence structures \citep{BickelLevina2008,CaiLiu2011,FriedmanHastieTibshirani2008}, while factor models exploit low-rank covariance representations. Our framework does not seek to select a single covariance model, but to quantify geometric proximity to multiple covariance geometries.

Third, the paper is related to low-rank plus sparse decomposition and robust principal component analysis \citep{CandesLiMaWright2011,Chandrasekaran2011}. Those methods aim to recover structural matrix decompositions. By contrast, our objective is to characterize the dependence architecture through similarity scores, allowing multiple mechanisms to coexist.

Finally, the paper draws on projection theory, matrix geometry, and semiparametric projection methods \citep{Deutsch2001,AbsilMahonySepulchre2008,Newey1994,BickelKlaassenRitovWellner1993}. The principal-angle condition introduced below plays a central role in geometric identification and clarifies the sources of ambiguous dependence structures.

The paper is organized as follows. Section~\ref{sec:geometry} introduces the covariance geometry framework. Section~\ref{sec:identification} develops identification and ambiguity results. Sections~\ref{sec:estimation} and~\ref{sec:asymptotics} study estimation and asymptotic theory. Section~\ref{sec:classification} develops dependence diagnostics and classification. Section~\ref{sec:profile_guided_inference} presents profile-guided inference. Sections~\ref{sec:simulation} and~\ref{sec:empirical} report simulation and empirical evidence.



\section{Covariance Geometry Framework}
\label{sec:geometry}

Let
\[
u=(u_1,\ldots,u_n)'
\]
denote the disturbance vector and define
\[
\Sigma=E(uu').
\]
We view \(\Sigma\) as an element of the Hilbert space
\[
\mathbb H=\{A=A':\|A\|_F<\infty\},
\qquad
\langle A,B\rangle_F=\operatorname{tr}(A'B).
\]
Rather than assuming that $\Sigma$ belongs to a single dependence class, we characterize its dependence architecture
through its geometric proximity to several economically meaningful covariance classes.

\subsection{Covariance Classes}

Let
\[
\mathcal S_C,\qquad \mathcal S_F,\qquad \mathcal S_S
\]
denote cluster, factor, and sparse covariance classes.

A covariance matrix belongs to the cluster class
\(\mathcal S_C\)
if its support can be represented as the union of a finite number of
clustering dimensions.
Suppose there are \(M\) clustering dimensions and let \(g_m(i)\),
\(m=1,\ldots,M\), denote the cluster membership of observation \(i\)
along dimension \(m\).
Define the \emph{cluster-support matrix} \(M_C\) by
\[
(M_C)_{ij}
=
\mathbf{1}\Bigl\{g_m(i)=g_m(j)\text{ for at least one }m\in\{1,\ldots,M\}\Bigr\}.
\]
Thus \((M_C)_{ij}=1\) whenever observations \(i\) and \(j\) share at
least one cluster membership, and \((M_C)_{ij}=0\) otherwise.
A covariance matrix \(\Gamma\) belongs to \(\mathcal S_C\) if
\(\Gamma_{ij}=0\) whenever \((M_C)_{ij}=0\), equivalently if
\(\Gamma=M_C\odot\Gamma\).
This definition includes one-way clustering (\(M=1\)), two-way clustering
(\(M=2\)), and multiway clustering (\(M\geq 2\)) as special cases.

For a fixed or slowly growing rank \(r\), define the factor class as
\[
\mathcal S_F(r)
=
\Big\{
\Gamma\succeq0:
\Gamma=L+D,\;
\operatorname{rank}(L)\le r,\;
D \text{ diagonal}
\Big\}.
\]

For \(k_n=o(n^2)\), define the sparse class as
\[
\mathcal S_S
=
\Big\{
\Gamma\in\mathbb H:
|\operatorname{supp}(\Gamma)|\le k_n
\Big\}.
\]

The covariance dictionary considered in this paper consists of cluster,
factor, and sparse geometries. These serve as benchmark geometries for
illustration; the framework extends naturally to richer dictionaries.
\subsection{Projection Operators}
Throughout the paper, let
\[
\mathfrak D=\{C,F,S\}
\]
denote the set of covariance geometries, corresponding respectively to
cluster, factor, and sparse dependence. Define
\[
P_d(\Sigma)
=
\arg\min_{\Gamma\in\mathcal S_d}
\|\Sigma-\Gamma\|_F.
\]
$P_d(\Sigma)$ represents the closest approximation to $\Sigma$ within the  covariance class $\mathcal S_d$.
Although we use projection notation, \(P_d\) need not be a linear orthogonal projection because the covariance classes may be nonlinear or nonconvex.

\begin{lemma}[Existence of Projections]
\label{lem:projectionexistence}
If \(\mathcal S_d\) is closed in \(\mathbb H\), then \(P_d(\Sigma)\) exists for every \(\Sigma\in\mathbb H\).
\end{lemma}

\begin{lemma}[Continuity of Projection Operators]
\label{lem:projectioncontinuity}
Suppose \(P_d\) is locally unique at \(\Sigma\). If \(\Sigma_m\to\Sigma\), then
\[
P_d(\Sigma_m)\to P_d(\Sigma).
\]
\end{lemma}

\subsection{Why the Frobenius Geometry?}

The Frobenius norm is induced by the Hilbert inner product \(\langle A,B\rangle_F=\operatorname{tr}(A'B)\). It therefore provides a natural geometry for defining projections, tangent spaces, principal angles, and local separation between covariance classes. It is also computationally convenient: low-rank projections are based on spectral truncation, cluster projections on support masks, and sparse projections on thresholding or sparse approximation. Alternative metrics may be useful in other applications, but the Frobenius geometry gives the most transparent framework for the identification and asymptotic theory developed below.


\subsection{Dependence Similarity Scores}
Define
\[
S_d=\|P_d(\Sigma)\|_F^2,
\qquad d\in\mathfrak D.
\]

The normalized dependence similarity scores are
\[
\omega_d=\frac{S_d}{\sum_{d\in\mathfrak D}S_{d}},
\qquad d\in\mathfrak D.
\]
The vector
\[
\omega=(\omega_C,\omega_F,\omega_S)'
\]
is called the dependence profile.

\begin{remark}[No Variance Decomposition]
The quantities \(\omega_C,\omega_F,\omega_S\) should not be interpreted as fractions of covariance explained. Since the covariance geometries overlap and the projections are not generally orthogonal,
\[
\sum_{d\in\mathfrak D}S_{d}
\neq
\|\Sigma\|_F^2
\]
in general. The dependence profile measures relative geometric similarity, not additive covariance contributions.
\end{remark}



\section{Geometric Identification}
\label{sec:identification}
The dependence profile is identified through the local geometry of the
covariance classes. Identification depends on whether the cluster,
factor, and sparse geometries remain sufficiently separated in a
neighborhood of the population dependence operator.

For each geometry \(d\in\mathfrak D\), let \(T_d(\Sigma_d)\) denote the
corresponding tangent space at a regular point \(\Sigma_d\).
Since diagonal perturbations are common to all covariance geometries and
carry no information about cross-sectional dependence, identification is
based on the off-diagonal tangent spaces

\[
T_d^{\mathrm{off}}(\Sigma_d)
=
\left\{
H\in T_d(\Sigma_d)
:
H_{ii}=0
\ \text{for all } i
\right\},
\qquad
d\in\mathfrak D.
\]
\paragraph{Principal-Angle Condition.}
For two tangent spaces \(U\) and \(V\), define the smallest principal angle by
\[
\theta(U,V)
=
\inf_{u\in U,\;v\in V}
\arccos
\left(
\frac{\langle u,v\rangle_F}{\|u\|_F\|v\|_F}
\right).
\]

\begin{assumption}[Principal-Angle Identification]
\label{ass:principalangle}
There exists \(\theta_0>0\) such that
\[
\theta\bigl(T_i^{\mathrm{off}}(\Sigma_i),\,T_j^{\mathrm{off}}(\Sigma_j)\bigr)\ge \theta_0
\]
for all \(i\neq j\), \(i,j\in\mathfrak D\).
\end{assumption}
Formal definitions of the covariance geometries, regularity conditions,
and tangent spaces are given in the Online Appendix.


\begin{assumption}[Sparse Projection Regularity]
\label{ass:sparse_unique}

The population dependence operator \(\Gamma_0\) is a regular sparse
point: the \(k_n\)-th and \((k_n+1)\)-th largest absolute off-diagonal
entries of \(\Gamma_0\) are distinct.

\end{assumption}

\begin{lemma}[Principal Angle and Tangent-Space Transversality]
\label{lem:angle_transversality}
Let $U$ and $V$ be closed linear subspaces of $\mathbb H$. Then
$\theta(U,V)>0$ if and only if $U\cap V=\{0\}$.
\end{lemma}


\begin{theorem}[Local Identification of the Dependence Profile]
\label{thm:identification}
Suppose Assumptions \ref{ass:principalangle} and \ref{ass:sparse_unique} hold. Then the dependence profile
$(\omega_C,\omega_F,\omega_S)$ is locally identified.

Moreover, no nonzero off-diagonal perturbation direction lies
simultaneously in two distinct covariance geometries. Consequently,
local changes in the dependence profile admit a unique first-order
geometric interpretation.
\end{theorem}


\begin{remark}[Interpretation]
\label{rem:identification_interpretation}

Theorem~\ref{thm:identification} establishes that dependence geometries
are locally distinguishable whenever their off-diagonal tangent spaces
remain sufficiently separated. The principal-angle condition therefore
plays the role of an identification assumption for dependence
structure.

\end{remark}

\begin{assumption}[Local Asymptotic Normality]
\label{ass:lan}
The statistical model $\{P_\Gamma:\Gamma\in\mathbb H\}$ satisfies local
asymptotic normality (LAN) at $\Gamma_0$: for every bounded sequence
$H_n\to H$ in $\mathbb H$,
\[
\log\frac{dP_{\Gamma_0+n^{-1/2}H_n}^n}{dP_{\Gamma_0}^n}
=
\frac{1}{\sqrt n}\sum_{i=1}^n \ell_i(H)
-
\tfrac12\|H\|_{\mathcal I}^2
+o_p(1),
\]
where $\ell_i(H)$ is a zero-mean score with $E[\ell_i(H)^2]=\|H\|_{\mathcal I}^2$ for
a positive-definite inner product $\|\cdot\|_{\mathcal I}$ on $\mathbb H$.
\end{assumption}
Assumption~\ref{ass:lan} is standard.
For Gaussian observations with covariance $\Gamma_0$, LAN holds with the
Fisher information inner product induced by the Gaussian log-likelihood; see
\citet{LeCamYang2000}, Chapter~7.
\begin{theorem}[Two-Geometry Local Indistinguishability]
\label{thm:two_geometry_indistinguishability}
Let $\mathcal S_i$ and $\mathcal S_j$ be two covariance geometries with
$i\neq j$.
Suppose Assumption~\ref{ass:lan} holds at a regular point
$\Gamma_0\in\mathcal S_i\cap\mathcal S_j$, and that there exists a nonzero off-diagonal direction $
H \in
T_i^{off}(\Gamma_0)
\cap
T_j^{off}(\Gamma_0)$.

Let $\Gamma_{i,n}\in\mathcal S_i$ and $\Gamma_{j,n}\in\mathcal S_j$ satisfy
\[
\Gamma_{i,n}
=
\Gamma_0+\frac{1}{\sqrt n}H+o(n^{-1/2}),
\qquad
\Gamma_{j,n}
=
\Gamma_0+\frac{1}{\sqrt n}H+o(n^{-1/2}).
\]

Then, for any sequence of tests $\varphi_n\in[0,1]$,
\[
\limsup_{n\to\infty}
\bigl|E_{\Gamma_{i,n}}\varphi_n-E_{\Gamma_{j,n}}\varphi_n\bigr|
=0.
\]
Consequently, no test can have asymptotic size tending to zero and power
tending to one for distinguishing $\mathcal S_i$ from $\mathcal S_j$ along
these local sequences.
\end{theorem}

\begin{remark}[Identification versus Ambiguity]
\label{rem:identification_vs_ambiguity}
Theorems~\ref{thm:identification} and
\ref{thm:two_geometry_indistinguishability} characterize the boundary between identifiable and ambiguous dependence
structures. Positive principal angles imply local identification,
whereas overlapping tangent spaces imply local indistinguishability.

\end{remark}





\section{Statistical Learning of Dependence Profiles}
\label{sec:estimation}

This section describes how the population dependence profile is estimated
from an empirical dependence operator. Let \(\widehat\Gamma_n\) denote an empirical dependence operator
constructed from the data. The operator may be a covariance matrix, a
long-run covariance operator, a spatial covariance operator, a network
dependence operator, or another object summarizing the dependence
features relevant for the inferential problem. Let \(\Gamma_0\) denote
the corresponding population dependence operator.

\begin{assumption}[Consistency of the Empirical Dependence Operator]
\label{ass:operator_consistency}

There exists a population dependence operator \(\Gamma_0\in\mathbb H\)
such that

\[
\|
\widehat\Gamma_n-\Gamma_0
\|_F
=
o_p(1).
\]

\end{assumption}

\begin{remark}[Dependence Operator versus Dependence Structure]
\label{rem:operator_vs_structure}

Assumption~\ref{ass:operator_consistency} does not require the
dependence structure to be known. The assumption concerns the
population dependence operator \(\Gamma_0\), which plays the role of an
intermediate object. The estimands of interest are the dependence
profile, the dominant covariance geometry, and the associated
classification, all of which are functionals of \(\Gamma_0\).

The situation is analogous to principal component analysis, where the covariance matrix is consistently estimated while the principal components themselves remain unknown.

\end{remark}

The choice of \(\widehat\Gamma_n\) is application-specific. It may be a covariance, long-run covariance, spatial, network, or cluster-supported operator. The theory below is conditional on the chosen operator; a fuller discussion of operator choice and computational implementation is given in the Online Appendix.

\subsection{Choice of Dependence Operator}
\label{subsec:operatorchoice}
The empirical dependence operator should capture the form of
dependence relevant for the inferential problem. For contemporaneous cross-sectional dependence, a natural benchmark is
the residual covariance operator
\[
\widehat\Gamma_n
=
\frac{1}{T}
\sum_{t=1}^{T}
\widehat u_t\widehat u_t',
\qquad
\widehat u_t=(\widehat u_{1t},\ldots,\widehat u_{nt})'.
\]
For serially dependent data one may instead use a long-run covariance
operator. For spatial, network, or clustered data, the operator may
incorporate distances, adjacency relations, or cluster-support
restrictions.

The framework does not seek to select a unique dependence operator.
Rather, it studies the covariance geometry encoded in a chosen
population operator. The dependence profile is invariant to positive
rescalings of the operator; additional examples and technical details
are provided in the Online Appendix.

\subsection{Projection Residual Diagnostics}
\label{subsec:projection_residuals}

The dependence profile provides a relative measure of geometric fit.
To assess goodness-of-fit in an absolute sense, we introduce
projection-residual diagnostics.

To quantify absolute goodness-of-fit, define the normalized projection
residual
\[
\rho_d(\Gamma)
=
\frac{
\|
\Gamma-P_d(\Gamma)
\|_F
}
{
\|\Gamma\|_F
},
\qquad
d\in\mathfrak D.
\]
The quantity \(\rho_d(\Gamma)\) measures the fraction of the operator
that remains unexplained after projection onto geometry \(d\). Since
\(
0\in\mathcal S_d
\),
the definition of the projection implies

\[
\|\Gamma-P_d(\Gamma)\|_F
\le
\|\Gamma\|_F.
\]
Consequently,
\[
0
\le
\rho_d(\Gamma)
\le
1.
\]
Small values indicate that geometry \(d\) provides a good approximation
to the dependence operator, whereas large values indicate substantial
lack of fit.

Define the minimum residual
\[
\rho_{\min}(\Gamma)
=
\min_{d\in\mathfrak D}
\rho_d(\Gamma).
\]
The quantity \(\rho_{\min}(\Gamma)\) summarizes the distance from the
dependence operator to the closest covariance geometry in the
dictionary.

We say that a dependence operator exhibits a
``none-of-the-above'' pattern whenever
\[
\rho_{\min}(\Gamma)>\delta,
\]
for a prespecified threshold \(\delta\in(0,1)\).

Large values of \(\rho_{\min}(\Gamma)\) indicate that none of the
cluster, factor, or sparse geometries provides an adequate
approximation to the observed dependence structure. In such situations,
the dependence profile should be interpreted with caution, since the
operator may contain dependence features not represented in the
covariance dictionary.

\begin{proposition}
\label{prop:residual_consistency}
Suppose Assumption~\ref{ass:operator_consistency} hlds. Then
\[
\rho_d(\widehat\Gamma_n)
\overset{p}{\longrightarrow}
\rho_d(\Gamma_0),
\qquad
d\in\mathfrak D,
\]
and
\[
\rho_{\min}(\widehat\Gamma_n)
\overset{p}{\longrightarrow}
\rho_{\min}(\Gamma_0).
\]
\end{proposition}


\subsection{Population Dependence Profile}
\label{subsec:population_dependence_profile}

The population dependence profile is defined from projections of
\(\Gamma_0\) onto the covariance geometries. For each
\(d\in\mathfrak D\), define

\[
P_{d,0}
=
P_d(\Gamma_0),
\qquad
S_{d,0}
=
\|P_{d,0}\|_F^2.
\]
The population similarity score associated with geometry \(d\) is

\[
\omega_{d,0}
=
\frac{S_{d,0}}{\sum_{d\in\mathfrak D}S_{d,0}},
\qquad d\in\mathfrak D.
\]

The vector

\[
\omega_0
=
(
\omega_{C,0},
\omega_{F,0},
\omega_{S,0}
)'
\]
is the population dependence profile.

\subsection{Similarity-Score Estimation}
\label{subsec:similarity_score_estimation}
Define

\[
\widehat S_d
=
\|
\widehat P_d
\|_F^2,
\qquad
 d\in\mathfrak D.
\]
The estimated similarity score for geometry \(d\) is
\[
\widehat\omega_d
=
\frac{\widehat S_d}{\sum_{d\in\mathfrak D}
\widehat S_d},
\qquad
 d\in\mathfrak D.
\]
The estimated dependence profile is
\[
\widehat\omega
=
(
\widehat\omega_C,
\widehat\omega_F,
\widehat\omega_S
)'.
\]



\section{Asymptotic Theory}
\label{sec:asymptotics}

This section establishes consistency and asymptotic normality of the
estimated dependence profile.

\begin{assumption}[Asymptotic Linearity of the Dependence Operator]
\label{ass:operatorclt}

There exists a mean-zero random vector \(\psi_i\) such that

\[
\sqrt n\,
\operatorname{vec}
(
\widehat\Gamma_n-\Gamma_0
)
=
\frac{1}{\sqrt n}
\sum_{i=1}^n
\psi_i
+
o_p(1),
\]
and
\[
E(\psi_i\psi_i')
=
\Omega_\Gamma.
\]

\end{assumption}

\begin{lemma}[Uniform Consistency of Projection Estimators]
\label{lem:uniform_projection_consistency}

Suppose Assumptions \ref{ass:principalangle}, \ref{ass:sparse_unique}, and \ref{ass:operator_consistency} hold. Then

\[
\max_{d\in\mathfrak D}
\|
\widehat P_d-P_{d,0}
\|_F
=
o_p(1).
\]

\end{lemma}

\begin{theorem}[Consistency of the Dependence Profile]
\label{thm:profile_consistency}

Suppose Assumptions \ref{ass:principalangle}, \ref{ass:sparse_unique}, \ref{ass:operator_consistency} hold. If
\(\sum_{d\in\mathfrak D}S_{d,0}>0\), then

\[
\widehat\omega-\omega_0
=
o_p(1).
\]

\end{theorem}

To state the limiting distribution, define the score map

\[
\mathcal S(\Gamma)
=
\left(
\|P_C(\Gamma)\|_F^2,
\|P_F(\Gamma)\|_F^2,
\|P_S(\Gamma)\|_F^2
\right)'.
\]
Let
\[
s_0
=
\mathcal S(\Gamma_0)
=
(
S_{C,0},
S_{F,0},
S_{S,0}
)'.
\]
Let
\[
\pi(s)
=
\frac{s}{\mathbf 1's}
\]
denote the normalization map from similarity scores to the dependence
profile, where \(\mathbf 1=(1,1,1)'\). Thus, \(\omega_0=\pi(s_0)\).

\begin{corollary}[Differentiability of the Score Map]
\label{ass:score_differentiability}

Suppose Assumptions \ref{ass:principalangle} and \ref{ass:sparse_unique}
hold at $\Gamma_0$. Then the score map \(\mathcal S(\cdot)\) is Hadamard
differentiable at \(\Gamma_0\), with derivative \(\dot{\mathcal
S}_{\Gamma_0}\) given coordinatewise by
\[
\dot S_{d,\Gamma_0}[H]
=
2\,\bigl\langle P_d(\Gamma_0),\,\Pi_{T_d(\Gamma_0)}(H)\bigr\rangle_F,
\qquad d\in\mathfrak D.
\]
\end{corollary}


\begin{theorem}[Asymptotic Distribution of the Dependence Profile]
\label{thm:profile_clt}

Suppose Assumptions \ref{ass:principalangle},
\ref{ass:sparse_unique}, and \ref{ass:operatorclt} hold. (By
Corollary~\ref{ass:score_differentiability}, the score map $\mathcal
S(\cdot)$ is then automatically Hadamard differentiable at $\Gamma_0$,
with derivative $\dot{\mathcal S}_{\Gamma_0}$.) If
\(\sum_{d\in\mathfrak D}S_{d,0}>0\), then
\[
\sqrt n
(
\widehat\omega-\omega_0
)
\Rightarrow
N(0,\Xi),
\]
where $\Xi
=
J_\omega
\Omega_\Gamma
J_\omega', \text{ with }
J_\omega
=
\dot\pi_{s_0}
\dot{\mathcal S}_{\Gamma_0}$.

Here \(\dot\pi_{s_0}\) is the derivative of the normalization map
\(\pi(s)=s/(\mathbf 1's)\) evaluated at \(s_0\).

\end{theorem}



\section{Dependence Diagnostics and Classification}
\label{sec:classification}

The dependence profile provides a low-dimensional summary of the
dependence architecture encoded in the empirical dependence operator.
Large values of \(\widehat{\omega}_C\), \(\widehat{\omega}_F\), and
\(\widehat{\omega}_S\) indicate that the empirical dependence operator
is geometrically close to the cluster, factor, and sparse covariance
classes, respectively.

The dependence profile measures relative geometric similarity rather than additive covariance contributions.
A natural classification rule is
\begin{equation}
\widehat d
=
\operatorname*{arg\,max}_{d\in\mathfrak D}
\widehat{\omega}_d.
\label{eq:classification_rule}
\end{equation}
The estimated geometry \(\widehat d\) identifies the covariance class
with the largest estimated similarity score. The full dependence profile is often more informative than the
classification alone. For example,
\[
(0.60,0.30,0.10)
\]
suggests cluster dominance with meaningful factor dependence, whereas
\[
(0.35,0.35,0.30)
\]
suggests hybrid dependence and cautions against a purely binary
classification.

\subsection{Statistical Uncertainty for Dependence Scores}
\label{subsec:profile_variance_estimation}

Inference for the dependence profile is based on
Theorem~\ref{thm:profile_clt},
\[
\sqrt n
\left(
\widehat{\omega}-\omega_0
\right)
\Rightarrow
N(0,\Xi),
\]
where
\[
\omega_0
=
(\omega_{C,0},\omega_{F,0},\omega_{S,0})'.
\]

Let
\[
e_C=(1,0,0)',
\qquad
e_F=(0,1,0)',
\qquad
e_S=(0,0,1)'.
\]
For \(d\in\mathfrak D\), the asymptotic variance of
\(\widehat{\omega}_d\) is
\[
\sigma_d^2
=
e_d'\Xi e_d.
\]

\begin{assumption}[Consistent Covariance Estimation]
\label{ass:Xi_estimation}

There exists an estimator \(\widehat{\Xi}\) such that
\[
\widehat{\Xi}-\Xi=o_p(1).
\]

\end{assumption}

Under Assumption~\ref{ass:Xi_estimation},
\[
\widehat{\sigma}_d^2
=
e_d'\widehat{\Xi}e_d
\]
is a consistent estimator of \(\sigma_d^2\).

The construction of \(\widehat{\Xi}\) depends on the chosen empirical
dependence operator. When the operator admits an asymptotic linear
representation, a plug-in estimator based on the influence function and
the delta method may be used. The construction of a universal variance
estimator for all possible dependence operators is outside the scope of
the paper.

\subsection{Diagnostic Tests}
\label{subsec:score_diagnostics}

To test
\[
H_0:\omega_C=\omega_C^\ast,
\]
consider
\[
T_C
=
\frac{
\sqrt n
\left(
\widehat{\omega}_C-\omega_C^\ast
\right)
}
{\widehat{\sigma}_C},
\]
where
\[
\widehat{\sigma}_C^2
=
e_C'\widehat{\Xi}e_C.
\]
Under \(H_0\),
\[
T_C
\Rightarrow
N(0,1).
\]
Analogous statistics apply to the factor and sparse scores.

More generally, let \(R_{\omega}\) be a \(q\times 3\) matrix of full row
rank and consider
\[
H_0:
R_{\omega}\omega_0=r_{\omega}.
\]
Define
\[
W_{\omega}
=
n
\left(
R_{\omega}\widehat{\omega}-r_{\omega}
\right)'
\left(
R_{\omega}\widehat{\Xi}R_{\omega}'
\right)^{-1}
\left(
R_{\omega}\widehat{\omega}-r_{\omega}
\right).
\]
Under \(H_0\),
\[
W_{\omega}
\Rightarrow
\chi_q^2.
\]

These tests quantify the uncertainty associated with the estimated
dependence profile and should be interpreted as diagnostic tools.

\subsection{Dominant and Hybrid Dependence}
\label{subsec:dominant_hybrid_classification}

A dominant dependence geometry is suggested when one similarity score
substantially exceeds the others. For example,
\[
\widehat{\omega}_C
>
\max
\left\{
\widehat{\omega}_F,
\widehat{\omega}_S
\right\}
\]
indicates that the empirical dependence operator is closest to the
cluster covariance geometry.

Hybrid dependence is suggested when several scores are simultaneously
large. In such cases, the full dependence profile should be reported
rather than reducing the analysis to a single classification label.

This distinction is important because covariance geometries may overlap.
For example, cluster covariance matrices may also be sparse when cluster
sizes are small relative to the sample size. The dependence profile
should therefore be viewed as a continuous description of dependence
architecture rather than a rigid model-selection device.

\subsection{Dominant Dependence Geometry}
\label{subsec:dominant_geometry}

A central question is whether the estimated dependence profile can be
used to identify the dominant covariance geometry underlying the chosen
dependence operator.

\begin{assumption}[Unique Dominant Geometry]
\label{ass:unique_dominant_geometry}

There exists a unique geometry
\(d^{\star}\in\mathfrak D\)
such that

\[
\omega_{d^{\star}}
>
\max_{d\neq d^{\star}}
\omega_d.
\]
Equivalently, the separation margin
\[
\Delta_{\omega}
=
\omega_{d^{\star}}
-
\max_{d\neq d^{\star}}
\omega_d
\]
satisfies $
\Delta_{\omega}>0$.

\end{assumption}

Under Assumption~\ref{ass:unique_dominant_geometry}, define

\begin{equation}
d^{\star}
=
\operatorname*{arg\,max}_{d\in\mathfrak D}
\omega_d.
\label{eq:dominant_geometry}
\end{equation}

The quantity \(d^{\star}\) represents the population dominant
dependence geometry.

Assumption~\ref{ass:unique_dominant_geometry} rules out knife-edge
cases in which two or more covariance geometries have exactly the same
population similarity score. Similar separation conditions are common
in model selection and statistical classification problems.

\begin{proposition}[Classification Error Bound]
\label{prop:classification_bound}

Suppose Assumption
\ref{ass:unique_dominant_geometry}
holds. Then
\[
P(\widehat d\neq d^\star)
\leq
P
\left(
\max_{d\in\mathfrak D}
|
\widehat\omega_d-\omega_d
|
>
\frac{\Delta_\omega}{2}
\right),
\]
where
\[
\Delta_\omega
=
\omega_{d^\star}
-
\max_{d\neq d^\star}
\omega_d.
\]
\end{proposition}


\begin{theorem}[Consistency of Dominant-Geometry Classification]
\label{thm:dominant_geometry_consistency}

Suppose the conditions of
Theorem~\ref{thm:profile_consistency}
hold so that

\[
\widehat{\omega}_d
\overset{p}{\longrightarrow}
\omega_d,
\qquad
d\in\mathfrak D,
\]

and Assumption~\ref{ass:unique_dominant_geometry}
holds. Then

\[
P\!\left(
\widehat d=d^{\star}
\right)
\longrightarrow
1.
\]

\end{theorem}

Theorem~\ref{thm:dominant_geometry_consistency} shows that the
dependence profile consistently identifies the dominant covariance
geometry whenever the dominant similarity score is separated from the
remaining scores. Consequently, the dependence profile is not merely a
descriptive summary of dependence but also provides a statistically
consistent basis for dependence classification.

\begin{remark}[Relation to Statistical Classification]

Theorem~\ref{thm:dominant_geometry_consistency}
is analogous to consistency results in model selection and statistical
classification. The distinguishing feature here is that the objects
being classified are covariance geometries rather than parametric
models.

\end{remark}

The classification result provides a formal justification for using the
estimated dependence profile as a guide for procedure recommendation. The implications for inference are developed formally in Section \ref{sec:profile_guided_inference}.
\begin{theorem}[Local Alternatives and Near-Ties]
\label{thm:local_ties}

Suppose

\[
\sqrt n
\left(
\widehat\omega-\omega_n
\right)
\Rightarrow
N(0,\Xi),
\]
where
\[
\omega_n
=
(\omega_{C,n},\omega_{F,n},\omega_{S,n})'.
\]

Assume that two geometries,
\(d_1\) and \(d_2\),
satisfy

\[
\omega_{d_1,n}
-
\omega_{d_2,n}
=
\frac{c}{\sqrt n},
\]
for some constant \(c\in\mathbb R\), while all remaining scores remain
separated by positive constants. Then
\[
P(\widehat d=d_1)
\]
converges to a nondegenerate limit lying strictly between zero and one.

\end{theorem}
\begin{remark}[Near-Ties]
\label{rem:near_ties}

When the separation margin
\(\Delta_{\omega}\)
is close to zero, small sampling fluctuations may alter the dominant
geometry classification. In such situations, the full dependence
profile $
\widehat{\omega}
=
(\widehat{\omega}_C,
  \widehat{\omega}_F,
  \widehat{\omega}_S)'
$ is typically more informative than the discrete classifier
\(\widehat d\).
\end{remark}





\section{Profile-Guided Inference}
\label{sec:profile_guided_inference}

The dependence profile provides a low-dimensional summary of the
dependence geometry encoded in an empirical dependence operator. A
natural question is whether the estimated profile can be used to guide
econometric inference. The objective of this section is not to propose a new covariance
estimator. Rather, we show how dependence learning can be combined with
existing dependence-robust procedures in a systematic way.

\subsection{Profile-Guided Procedure Recommendation}
\label{subsec:procedure_recommendation}

Let $\widehat d$ denote the estimated dominant dependence geometry as defined in \eqref{eq:classification_rule}.
The interpretation is straightforward. When
\(
\widehat\omega_C
\)
dominates, cluster-robust procedures provide a natural benchmark. When
\(
\widehat\omega_F
\)
dominates, factor-robust or common-shock-adjusted procedures are
suggested. When
\(
\widehat\omega_S
\)
dominates, sparse, network, spatial, or local-dependence robust
procedures deserve particular attention.

The recommendation should be viewed as evidence-based rather than
deterministic. The dependence profile does not imply that one procedure
is universally correct. Rather, it summarizes which covariance geometry
appears most consistent with the observed dependence operator.

When multiple similarity scores are substantial, the data provide
evidence for several dependence mechanisms simultaneously. In such
situations, reporting inference results from multiple
dependence-robust procedures may be preferable. Similarly, when the
projection-residual diagnostics indicate poor fit, the entire
covariance dictionary should be viewed with caution.

\subsection{Why Learn the Geometry?}
\label{subsec:why_learn_geometry}

A natural alternative is to use a single broadly robust procedure, such
as multiway clustering, HAC inference, or a general sandwich estimator.
Dependence learning remains useful for three reasons.

First, broad robustness is not costless. Procedures designed to be valid
under large classes of dependence may be conservative or poorly sized
when the realized dependence geometry differs from the structure for
which the procedure is best suited. Second, such procedures still require
the researcher to specify a dependence class, such as clustering
dimensions, a spatial metric, a bandwidth, or a network. These choices
are themselves assumptions about dependence. Third,
Theorem~\ref{thm:oracle_adaptivity} shows that profile-guided selection
is asymptotically equivalent to an infeasible oracle that knows the
dominant covariance geometry.

Thus, the value of dependence learning is not only diagnostic. It
provides a data-driven way to choose among geometry-specific inference
procedures while retaining oracle-equivalent first-order behavior.

\subsection{Procedure Confidence Index}
\label{subsec:procedure_confidence}

The reliability of a profile-based procedure recommendation depends on
two considerations. First, the dominant geometry should be sufficiently
separated from the competing geometries. Define the estimated separation
margin

\[
\widehat\Delta_\omega
=
\max_{d\in\mathfrak D}
\widehat\omega_d
-
\max_{d\neq \widehat d}
\widehat\omega_d.
\]

Second, the covariance dictionary should provide a satisfactory
approximation to the empirical dependence operator. Recall the minimum
projection residual

\[
\widehat\rho_{\min}
=
\min_{d\in\mathfrak D}
\widehat\rho_d.
\]

Combining these quantities yields the procedure confidence index

\[
\widehat\kappa
=
\bigl(
1-\widehat\rho_{\min}
\bigr)
\widehat\Delta_\omega.
\]

Large values of
\(
\widehat\kappa
\)
indicate that one geometry is clearly dominant and that the covariance
dictionary provides a good approximation to the observed dependence
operator. Small values indicate either weak separation, poor geometric
fit, or both.

The index therefore summarizes the strength of the empirical evidence
supporting a profile-based procedure recommendation. Let
\[
\kappa_0
=
\bigl(
1-\rho_{\min,0}
\bigr)
\Delta_{\omega,0}, \qquad
\rho_{\min,0}
=
\min_{d\in\mathfrak D}
\rho_d,
\qquad
\Delta_{\omega,0}
=
\max_{d\in\mathfrak D}
\omega_d
-
\max_{d\neq d^\star}
\omega_d, \text{ with }
d^\star
=
\arg\max_{d\in\mathfrak D}
\omega_d.
\]

\begin{proposition}[Consistency of the Procedure Confidence Index]
\label{prop:kappa_consistency}

Suppose the conditions of
Theorem~\ref{thm:profile_consistency}
hold. Then

\[
\widehat\kappa
\overset{p}{\longrightarrow}
\kappa_0.
\]
\end{proposition}


\subsection{Profile-Guided Variance Estimation}
\label{subsec:profile_guided_variance}
Suppose a scalar parameter \(\theta_0\) is estimated by \(\widehat\theta_n\).
For each geometry \(d\in\mathfrak D\), let \(\widehat V_d\) be a measurable
variance estimator and let \(\tau_{n,d}\to\infty\) be a deterministic
normalization such that
\[
\tau_{n,d}^{2}Var(\widehat\theta_n)\longrightarrow V_d,
\qquad 0<V_d<\infty.
\]


Define
\[
\widehat d
=
\arg\max_{d\in\mathfrak D}\widehat\omega_d,
\qquad
\widehat V^{*}
=
\widehat V_{\widehat d},
\]
where ties in the argmax are broken by a fixed deterministic rule. Thus, inference is based on the variance estimator associated with the
estimated dominant dependence geometry.

\begin{proposition}[Consistency of the Profile-Guided Variance Estimator]
\label{prop:vcov_consistency}

Let
\(
\widehat V_d
\)
denote a geometry-specific variance estimator for
\(
d\in\mathfrak D
\).
Suppose the conditions of
Theorem~\ref{thm:dominant_geometry_consistency}
hold and

\[
\widehat V_{d^\star}
\overset{p}{\longrightarrow}
V_{d^\star},
\]

where

\[
d^\star
=
\arg\max_{d\in\mathfrak D}
\omega_d.
\]

Then

\[
\widehat V^{*}
\overset{p}{\longrightarrow}
V_{d^\star}.
\]

\end{proposition}


Proposition~\ref{prop:vcov_consistency} shows that profile-guided
selection preserves consistency whenever the dominant geometry is
correctly classified. The next result establishes a stronger oracle
property.

\begin{theorem}[Oracle adaptivity of profile-guided inference]
\label{thm:oracle_adaptivity}

Suppose the conditions of
Theorem~\ref{thm:dominant_geometry_consistency}
hold. Let
\(
d^\star
=
\arg\max_{d\in\mathfrak D}
\omega_d
\)
denote the dominant covariance geometry, and let
\(
\widehat V_d
\)
be a variance estimator satisfying

\[
\widehat V_d
\overset{p}{\longrightarrow}
V_d,
\qquad
d\in\mathfrak D.
\]

Define

\[
\widehat d
=
\arg\max_{d\in\mathfrak D}
\widehat\omega_d,
\qquad
\widehat V^\star
=
\widehat V_{\widehat d}.
\]

Then

\[
\widehat {V}^{*}
-
\widehat V_{d^\star}
=
o_p(1).
\]

Consequently,

\[
\widehat {V}^{*}
\overset{p}{\longrightarrow}
V_{d^\star}.
\]

Thus, the profile-guided variance estimator is asymptotically equivalent
to the infeasible oracle estimator that knows the dominant covariance
geometry in advance.

\end{theorem}
When the asymptotic normalization depends on the covariance geometry,
comparison with the oracle statistic is naturally made on the event
\(\{\widehat d=d^\star\}\). On this event the selected rate
\(\tau_{n,\widehat d}\) and variance estimator \(\widehat V^*\)
coincide exactly with \(\tau_{n,d^\star}\) and
\(\widehat V_{d^\star}\). Since this event occurs with probability
approaching one, the profile-guided and oracle statistics are
asymptotically equivalent.

\begin{corollary}[Oracle-equivalent Wald inference]
\label{cor:oracle_wald}

Let
\[
T_n^{*}
=
\frac{\tau_{n,\widehat {d}}\,(\widehat\theta_n-\theta_0)}
{\sqrt{\widehat V^{*}}}
\]
and
\[
T_n^{oracle}
=
\frac{\tau_{n,{d}^\star}\,(\widehat\theta_n-\theta_0)}
{\sqrt{\widehat V_{d^\star}}}.
\]
Assume furthermore that the oracle statistic satisfies

\[
T_n^{oracle}
\Rightarrow
N(0,1).
\]
Under the assumptions of
Theorem \ref{thm:oracle_adaptivity},

\[
T_n^{*}
-
T_n^{oracle}
=
o_p(1).
\]

Consequently,

\[
T^{*}_{n}
\Rightarrow N(0,1)
\]

whenever the oracle statistic is asymptotically standard normal.
\end{corollary}

Theorem~\ref{thm:oracle_adaptivity} and Corollary~\ref{cor:oracle_wald} provide the main econometric payoff of
dependence learning. Estimated profiles can be used to select dependence-robust procedures in a data-driven way, and the resulting inference is asymptotically equivalent to the infeasible oracle procedure that knows the dominant dependence geometry in advance.

\subsection{Profile-Weighted Inference}
\label{subsec:profile_weighted}

The profile-guided estimator selects a single dominant geometry. An
alternative approach is to combine multiple procedures using the
estimated dependence profile itself. Define

\[
\widehat V^{\mathrm{avg}}
=
\sum_{d\in\mathfrak D}
\widehat\omega_d
\widehat V_d.
\]

The resulting estimator may be interpreted as a dependence-weighted
combination of geometry-specific variance estimators, analogous to
forecast combination and model averaging \citep{BatesGranger1969,Hansen2007}. Developing efficiency theory
and optimal weighting schemes for profile-weighted inference is an
important direction for future research.






\section{Simulation Evidence}
\label{sec:simulation}

This section summarizes Monte Carlo evidence for the dependence-learning
framework. The simulations focus on three questions. First, can the
estimated dependence profile recover the dominant covariance geometry?
Second, how does classification behave under hybrid and near-tie
dependence? Third, does profile-guided inference behave like the
infeasible oracle procedure that knows the dominant dependence geometry
in advance?

The detailed data-generating processes, parameter calibrations,
projection algorithms, and additional robustness exercises are reported
in the Online Appendix. The baseline designs include pure cluster, pure
factor, sparse network, cluster--factor hybrid, cluster--sparse hybrid,
factor--sparse hybrid, all-three hybrid, and two-way cluster dependence.

\subsection{Dependence Learning}
\label{subsec:simulation_learning}

Table~\ref{tab:simulation_profiles} reports the population dependence
profiles and the Monte Carlo behavior of the estimated profiles. The
profile recovers the dominant geometry in the benchmark cluster and
factor designs and records substantial sparse affinity under designs in
which support restrictions are important, including two-way clustering.
The procedure confidence index \(\kappa_0\) is largest when one geometry
is clearly separated and the covariance dictionary provides a good
absolute fit.

\begin{table}[!ht]
\centering
\caption{Estimated Dependence Profiles}
\label{tab:simulation_profiles}
\resizebox{\textwidth}{!}{

  \IfFileExists{results/tables/table_simulation_profiles.tex}{}{
    \begin{tabular}{lccc}\toprule
    Output & Cluster & Factor & Sparse\\
    \midrule
    Placeholder & -- & -- & --\\
    \bottomrule
    \end{tabular}
  }

}
\begin{minipage}{0.92\textwidth}
\footnotesize
\emph{Notes:} The table reports population similarity scores, Monte Carlo
means and standard deviations of the estimated scores, the root mean
squared error of the estimated profile, and the population procedure
confidence index \(\kappa_0\). Similarity scores measure relative
geometric proximity and should not be interpreted as additive variance
shares.
\end{minipage}
\end{table}

Figure~\ref{fig:hybrid_profiles} shows the average estimated profile
along a cluster--factor hybrid path. The profile changes smoothly as the
relative strength of factor dependence increases, illustrating that the
method summarizes hybrid dependence rather than forcing a binary
classification.

\begin{figure}[!ht]
\centering
\includegraphicsIfExists[width=.72\textwidth]{results/figures/figure_hybrid_profiles.pdf}
\caption{Dependence Profiles along a Cluster--Factor Hybrid Path}
\label{fig:hybrid_profiles}
\begin{minipage}{0.86\textwidth}
\footnotesize
\emph{Notes:} The figure plots average estimated similarity scores along
a path moving from cluster-dominant to factor-dominant dependence.
\end{minipage}
\end{figure}

\subsection{Classification and Ambiguity}
\label{subsec:simulation_classification}

The classification theory predicts that dominant-geometry classification
is reliable when the separation margin is large and remains probabilistic
under local near-ties. Table~\ref{tab:near_ties} reports classification
frequencies for local cluster--factor alternatives, and
Figure~\ref{fig:classification_margin} plots the relationship between
classification error and the population separation margin.

\begin{table}[!ht]
\centering
\caption{Near-Ties and Ambiguous Dependence}
\label{tab:near_ties}

  \IfFileExists{results/tables/table_near_ties.tex}{}{
    \begin{tabular}{lccc}\toprule
    Output & Cluster & Factor & Sparse\\
    \midrule
    Placeholder & -- & -- & --\\
    \bottomrule
    \end{tabular}
  }

\begin{minipage}{0.88\textwidth}
\footnotesize
\emph{Notes:} The table reports Monte Carlo classification frequencies
under local cluster--factor alternatives. Near-ties generate
nondegenerate classification probabilities, as predicted by the local
alternatives theory.
\end{minipage}
\end{table}

\begin{figure}[!ht]
\centering
\includegraphicsIfExists[width=.72\textwidth]{results/figures/figure_classification_margin.pdf}
\caption{Classification Error and Separation Margin}
\label{fig:classification_margin}
\begin{minipage}{0.86\textwidth}
\footnotesize
\emph{Notes:} The figure plots Monte Carlo misclassification frequency
against the population separation margin. Larger margins are associated
with lower classification error.
\end{minipage}
\end{figure}

\subsection{Oracle-Equivalent Inference}
\label{subsec:simulation_oracle}

The main econometric implication of Section~\ref{sec:profile_guided_inference}
is that profile-guided inference should be asymptotically equivalent to
an infeasible oracle procedure. Table~\ref{tab:oracle_equivalence}
reports finite-sample coverage probabilities for the oracle procedure,
the profile-guided procedure, and a deliberately misspecified procedure.

\begin{table}[!ht]
\centering
\caption{Oracle Equivalence of Profile-Guided Inference}
\label{tab:oracle_equivalence}

  \IfFileExists{results/tables/table_oracle_equivalence.tex}{}{
    \begin{tabular}{lccc}\toprule
    Output & Cluster & Factor & Sparse\\
    \midrule
    Placeholder & -- & -- & --\\
    \bottomrule
    \end{tabular}
  }

\begin{minipage}{0.88\textwidth}
\footnotesize
\emph{Notes:} The table reports empirical coverage probabilities for
nominal 95\% confidence intervals. The oracle procedure uses the
variance estimator associated with the true dominant covariance geometry.
The profile-guided procedure uses the estimated dominant geometry. The
``Wrong Procedure'' column reports a deliberately misspecified benchmark.
\end{minipage}
\end{table}

The profile-guided procedure is nearly indistinguishable from the oracle
procedure across the benchmark designs, while the misspecified procedure
can exhibit severe undercoverage. These results provide finite-sample
support for Theorem~\ref{thm:oracle_adaptivity} and
Corollary~\ref{cor:oracle_wald}. They also give the main practical
message of the paper: learning the dependence profile can guide the
choice of robust inference procedure in a way that tracks the infeasible
oracle benchmark.

\begin{figure}[!ht]
\centering
\includegraphicsIfExists[width=.72\textwidth]{results/figures/figure_oracle_tracking.pdf}
\caption{Oracle Tracking along a Cluster--Factor Dominance Sweep}
\label{fig:oracle_tracking}
\begin{minipage}{0.86\textwidth}
\footnotesize
\emph{Notes:} The figure compares rejection rates for fixed procedures,
the profile-guided procedure, and the infeasible oracle along a
cluster--factor dominance sweep. The profile-guided and oracle procedures
closely coincide when classification is reliable.
\end{minipage}
\end{figure}


\section{Empirical Illustration}
\label{sec:empirical}

This section illustrates how the proposed dependence profile and
projection-residual diagnostics can be used in an empirical setting. The
goal is not to provide a new asset-pricing analysis, but to show how the
framework can be applied to a publicly available dataset in which several
dependence mechanisms are plausible.

We use the Fama--French 49 industry portfolios and the Fama--French
three factors from the data library of \citet{FamaFrench1993} and
\citet{FamaFrenchDataLibrary}. The data contain value-weighted and
equal-weighted returns for U.S. industry portfolios and are available at
daily, monthly, and annual frequencies. The monthly and daily series
begin in July 1926. The Data Library also provides the standard
asset-pricing factors used in empirical finance.

The empirical setting is well suited to the proposed framework. Industry
portfolio returns may exhibit cluster-like dependence because industries
can be grouped into broader sectors. They may exhibit factor dependence
because returns load on common market-wide and macro-financial shocks.
They may also exhibit sparse dependence because some industries are more
closely linked through input-output relationships, supply chains, or
common demand shocks than others.

\subsection{Data and Regression Specification}
\label{subsec:empirical_data}

The empirical analysis uses monthly excess returns on the 49 industry
portfolios together with the market, SMB, and HML factors from the
Kenneth R. French Data Library \citep{FamaFrenchDataLibrary}. Let $ R_{it}$ denote the excess return on industry portfolio \(i\) in month \(t\), where $i=1,\ldots,N$, and $t=1,\ldots,T$.
In the baseline illustration, \(N=49\). We estimate the factor model

\begin{equation}
R_{it}
=
\alpha_i
+
\beta_{iM} MKT_t
+
\beta_{iS} SMB_t
+
\beta_{iH} HML_t
+
u_{it},
\label{eq:empirical_factor_model}
\end{equation}
where \(MKT_t\), \(SMB_t\), and \(HML_t\) denote the Fama--French market,
size, and value factors. The residual \(u_{it}\) captures the component
of industry returns not explained by the observed factors.

Stack residuals across industries as $
\widehat u_t
=
(\widehat u_{1t},\ldots,\widehat u_{Nt})'$.

The baseline empirical dependence operator is the residual covariance
matrix
\begin{equation}
\widehat\Gamma
=
\frac{1}{T}
\sum_{t=1}^{T}
\widehat u_t\widehat u_t'.
\label{eq:empirical_operator}
\end{equation}

We also consider the residual correlation operator

\begin{equation}
\widehat\Gamma^{COR}
=
\widehat D^{-1/2}
\widehat\Gamma
\widehat D^{-1/2},
\label{eq:empirical_correlation_operator}
\end{equation}
where $\widehat D=\operatorname{diag}(\widehat\Gamma)$. The covariance operator preserves differences in residual volatility across industries, whereas the correlation operator focuses on dependence geometry after normalizing scale.

\subsection{Dependence Classes and Projections}
\label{subsec:empirical_classes}

We compute projections of the empirical dependence operator onto the
cluster, factor, and sparse covariance geometries.

\paragraph{Cluster Geometry.}

The cluster geometry introduced in this paper encompasses one-way,
two-way, and multiway clustering through a structured-support
representation. In the present empirical application, we use a one-way clustering
specification based on the Fama--French 10-sector economic classification
(Consumer Non-Durables, Consumer Durables, Manufacturing, Energy, Hi-Tech,
Telecom, Shops, Health, Utilities, and Other).

Let
\[
g(i)\in\{1,\ldots,G\}
\]
denote the sector membership of industry portfolio \(i\). Define the
cluster-support matrix

\[
(M_C)_{ij}
=
\mathbf{1}\{g(i)=g(j)\}.
\]

The cluster projection is

\[
\widehat P_C
=
M_C
\odot
\widehat\Gamma.
\]

Thus, the empirical implementation treats industry sectors as the
relevant clustering dimension and measures the extent to which the
residual dependence operator conforms to the resulting cluster-support
structure. If additional clustering dimensions were available, the
cluster-support matrix could be constructed according to the multiway
definition in Section~\ref{sec:geometry}.

\paragraph{Factor Geometry.}

The factor projection \(\widehat P_F\) is computed by the
alternating-projection algorithm described in Online Appendix,
applied to the empirical dependence operator \(\widehat\Gamma\).  The
output decomposes as \(\widehat P_F = L^* + D^*\), where \(L^*\) is a
rank-\(r\) positive-semidefinite matrix and \(D^*\) is diagonal.  In the
baseline implementation, \(r\) is selected using an eigenvalue-ratio
criterion applied to \(\widehat\Gamma-\widehat D\), where \(\widehat D\)
is the diagonal of \(\widehat\Gamma\).  As robustness checks, we also
consider fixed ranks \(r\in\{1,2,3\}\).

\paragraph{Sparse Geometry.}

The sparse projection is computed by retaining the \(k_N\) largest
off-diagonal entries of \(\widehat\Gamma\) in absolute value and setting
all remaining off-diagonal entries to zero. The diagonal entries are
retained. Formally,
\[
\widehat P_S
=
\operatorname*{arg\,min}_{\Gamma\in\mathcal S_S}
\|
\widehat\Gamma-\Gamma
\|_F^2.
\]
This projection provides the closest sparse approximation to the
empirical dependence operator under Frobenius loss.

\subsection{Estimated Dependence Profiles and Residual Diagnostics}
\label{subsec:empirical_profile}

We apply the estimated dependence profile and projection-residual
diagnostics defined in Sections~\ref{subsec:similarity_score_estimation}
and~\ref{subsec:projection_residuals}, with \(\widehat\Gamma\) taken to
be the residual covariance operator in \eqref{eq:empirical_operator}.

The dependence profile and the projection residuals answer different
questions. The profile describes which geometry provides the strongest
relative fit among the candidate classes. The residual diagnostics assess
absolute fit. A large value of \(\widehat\rho_{\min}\) indicates that
none of the candidate geometries provides a close approximation to the
empirical dependence operator, even if one geometry receives the largest
relative similarity score.

A large \(\widehat{\omega}_C\) indicates that residual dependence is
well described, in relative terms, by a cluster-support structure. In
the present application this corresponds to dependence concentrated
within broad industry sectors. A large \(\widehat{\omega}_F\) suggests
that the residual covariance still contains low-rank common components
not fully absorbed by the observed Fama--French factors. A large
\(\widehat{\omega}_S\) indicates that dependence is concentrated among a
relatively small subset of industry pairs.

\subsection{Conventional Robust Inference Benchmarks}
\label{subsec:empirical_inference}

To connect the dependence profile with conventional robust inference, we
compare standard errors for a pooled regression of industry excess
returns on the Fama--French factors. Specifically, estimate
\begin{equation}
R_{it}
=
\alpha
+
\beta_M MKT_t
+
\beta_S SMB_t
+
\beta_H HML_t
+
e_{it}.
\label{eq:empirical_pooled}
\end{equation}

We report heteroskedasticity-robust standard errors, standard errors
clustered by broad industry sector, and two-way clustered standard
errors by industry and month. We also report a common-shock benchmark
that removes leading principal components from the residual covariance
operator before computing standard errors.

The purpose of this comparison is descriptive. If the estimated profile
assigns substantial weight to the cluster geometry, clustered standard
errors should be viewed as empirically relevant. If the factor score is
large, then procedures that account for common shocks or factor
dependence deserve attention. If the sparse score is large, then
network- or sparse-dependence robust procedures may be important. The
dependence profile therefore serves as a diagnostic summary of the
covariance architecture underlying the robust inference procedures
considered in practice.

\subsection{Empirical Results}
\label{subsec:empirical_results}

Table~\ref{tab:empirical_profile} reports the estimated dependence
profiles and projection-residual diagnostics for the residual covariance
and correlation operators.

\begin{table}[!ht]
\centering
\caption{Dependence Profiles and Projection-Residual Diagnostics for Industry Portfolio Residuals}
\label{tab:empirical_profile}

  \IfFileExists{results/tables/table_empirical_profile.tex}{}{
    \begin{tabular}{lccc}\toprule
    Output & Cluster & Factor & Sparse\\
    \midrule
    Placeholder & -- & -- & --\\
    \bottomrule
    \end{tabular}
  }

\begin{minipage}{0.92\textwidth}
\footnotesize
\emph{Notes:}
\vspace{0.4em}
The table reports estimated similarity scores
\((\widehat\omega_C,\widehat\omega_F,\widehat\omega_S)\)
and projection-residual diagnostics
\((\widehat\rho_C,\widehat\rho_F,\widehat\rho_S,\widehat\rho_{\min})\)
using residuals from the industry-level Fama--French regression in
\eqref{eq:empirical_factor_model}. The covariance operator preserves
scale differences across industries, while the correlation operator
normalizes industry-specific residual volatility. The similarity scores
are relative geometric proximity measures and should not be interpreted
as additive variance shares. The residual diagnostics measure absolute
distance from the empirical operator to each covariance geometry.
\end{minipage}
\end{table}

Figure~\ref{fig:empirical_profile} displays the estimated dependence
profiles graphically.

\begin{figure}[!ht]
\centering
\includegraphicsIfExists[width=.75\textwidth]{results/figures/figure_empirical_profile.pdf}
\caption{Estimated Dependence Profile for Industry Portfolio Residuals}
\label{fig:empirical_profile}
\begin{minipage}{0.92\textwidth}
\footnotesize
\vspace{0.4em}
\emph{Notes:}
The figure displays the estimated dependence profile based on the
residual dependence operator. Each bar reports the relative geometric
affinity of the empirical operator to the cluster, factor, and sparse
covariance geometries. The corresponding projection-residual diagnostics
are reported in Table~\ref{tab:empirical_profile}.
\end{minipage}
\end{figure}

Table~\ref{tab:empirical_se} compares conventional robust standard
errors for the pooled factor regression.

\begin{table}[!ht]
\centering
\caption{Conventional Robust Standard Errors}
\label{tab:empirical_se}

  \IfFileExists{results/tables/table_empirical_standard_errors.tex}{}{
    \begin{tabular}{lccc}\toprule
    Output & Cluster & Factor & Sparse\\
    \midrule
    Placeholder & -- & -- & --\\
    \bottomrule
    \end{tabular}
  }

\begin{minipage}{0.92\textwidth}
\footnotesize
\vspace{0.4em}
\emph{Notes:}
The table reports coefficient estimates and standard errors for the
pooled factor regression in \eqref{eq:empirical_pooled}. The columns
compare heteroskedasticity-robust standard errors, sector-clustered
standard errors, two-way clustered standard errors, and common-shock
adjusted standard errors. The comparison illustrates how the estimated
dependence profile can guide the interpretation of conventional robust
inference procedures.
\end{minipage}
\end{table}


Table~\ref{tab:empirical_profile} reveals a hybrid dependence structure.

\textit{Covariance operator.}
The factor and sparse scores are tied at $\widehat\omega_F=\widehat\omega_S=0.445$,
with a small cluster score $\widehat\omega_C=0.110$.
The minimum projection residual $\widehat\rho_{\min}=0.008$ indicates
that the sparse projection provides a near-perfect absolute fit to the
residual covariance operator.
Taken together, these findings indicate that residual co-movement among
the 49 industry portfolios---after removing the three Fama--French
factors---is concentrated in a sparse set of industry pairs rather than
being driven by broad industry-sector clustering.
The small $\widehat\rho_{\min}$ confirms that the sparse geometry
adequately captures this residual dependence in an absolute sense.
The non-negligible factor score suggests that some low-rank common
component remains in the residuals despite the three-factor adjustment;
this is consistent with known limitations of the Fama--French model as
a complete description of common risk.

\textit{Correlation operator.}
Normalizing by industry volatility shifts the profile toward sparse
dominance ($\widehat\omega_S=0.397$) with a moderate cluster score
($\widehat\omega_C=0.255$) and $\widehat\rho_{\min}=0.320$.
The larger minimum residual indicates that the correlation structure is
less well captured by the candidate geometries than the covariance
structure, suggesting that industry-level volatility differences
mask some localized dependence that only emerges after scale normalization.

\textit{Standard errors.}
Table~\ref{tab:empirical_se} shows that sector-clustered standard errors
are roughly $2.3\times$ the White standard errors for the market factor
($0.051$ vs.\ $0.022$), and two-way clustered standard errors are
slightly larger still ($0.060$). The common-shock adjusted standard
errors are substantially smaller ($0.007$), reflecting the removal of
the residual common factor component. These patterns are consistent with
the estimated profile: the hybrid factor--sparse dependence structure
implies that sector clustering overstates the within-cluster dependence,
while the residual common component---identified by the factor score---is
better absorbed by a common-shock adjustment than by clustering.


\textit{Procedure confidence index and inference recommendation.}
Table~\ref{tab:empirical_inference} reports the procedure confidence
index $\widehat\kappa$ and the profile-guided inference recommendation
for each dependence operator.

\begin{table}[!ht]
\centering
\caption{Profile-Guided Inference Recommendation for Industry Portfolio Residuals}
\label{tab:empirical_inference}
\resizebox{\textwidth}{!}{

  \IfFileExists{results/tables/table_empirical_inference.tex}{}{
    \begin{tabular}{lccc}\toprule
    Output & Cluster & Factor & Sparse\\
    \midrule
    Placeholder & -- & -- & --\\
    \bottomrule
    \end{tabular}
  }

}
\begin{minipage}{0.99\textwidth}
\footnotesize
\vspace{0.4em}
\emph{Notes:}
The table reports the separation margin
$\widehat\Delta_\omega=\max_d\widehat\omega_d - \max_{d\neq\widehat d}\widehat\omega_d$,
the minimum projection residual $\widehat\rho_{\min}$, the procedure
confidence index $\widehat\kappa=(1-\widehat\rho_{\min})\widehat\Delta_\omega$,
the estimated dominant geometry, and the implied profile-guided
recommendation. A small $\widehat\kappa$ indicates either near-tie
separation or poor dictionary fit; the recommended response is to report
inference from multiple procedures.
\end{minipage}
\end{table}

Both operators yield small values of $\widehat\kappa$.
For the covariance operator, the factor and sparse scores are exactly
tied at $0.445$, so $\widehat\Delta_\omega=0$ and $\widehat\kappa=0$:
the profile provides no statistical basis for preferring any single
robust procedure.
For the correlation operator, the sparse score marginally dominates
($\widehat\omega_S=0.397$), but the separation margin is only
$\widehat\Delta_\omega=0.050$ and the minimum residual is
$\widehat\rho_{\min}=0.320$, giving $\widehat\kappa=0.034$.

These low confidence-index values have a clear practical implication:
they correctly indicate that the empirical dependence structure is
genuinely hybrid, and that no single inference procedure can be
recommended with high confidence based on the data alone.
The appropriate response---which the framework prescribes directly---is
to report inference results from multiple procedures and document
sensitivity to the choice.
Table~\ref{tab:empirical_se} operationalises this recommendation: the
common-shock adjusted standard errors (appropriate when factor dependence
is present) and the sector-clustered standard errors (appropriate when
cluster dependence dominates) differ by a factor of approximately
$7\times$ for the market factor, confirming that the choice of
procedure is empirically consequential.

This example illustrates a key advantage of the proposed framework
over a priori procedure selection.
A researcher who assumes cluster dependence and uses sector-clustered
standard errors is implicitly claiming $\widehat\kappa\approx 1$ for
the cluster geometry---a claim the data refute.
A researcher who learns the profile first discovers the hybrid
structure, obtains $\widehat\kappa\approx 0$, and is correctly
directed to report multiple procedures rather than to commit to one.
The framework thus transforms an informal sensitivity analysis into a
statistically grounded diagnostic.

\subsection{Replication Details}
\label{subsec:empirical_replication}

The empirical illustration uses publicly available industry portfolio
returns and factor data from Kenneth French's Data Library. The
replication package downloads the data, constructs excess returns,
estimates \eqref{eq:empirical_factor_model}, computes the empirical
dependence operators in \eqref{eq:empirical_operator} and
\eqref{eq:empirical_correlation_operator}, estimates dependence profiles
and projection-residual diagnostics, and generates
Tables~\ref{tab:empirical_profile} and \ref{tab:empirical_se} and
Figure~\ref{fig:empirical_profile}.

All data transformations, sector mappings, parameter choices, and random
seeds used in the empirical illustration are documented in the
replication files.







\section{Conclusion}
\label{sec:conclusion}

This paper develops a framework for learning dependence structures from empirical measures of dependence. Rather than treating dependence structures as maintained assumptions, we view cluster, factor, and sparse dependence as covariance geometries and measure their proximity to an empirical dependence operator.

The proposed dependence profile provides a low-dimensional summary of
high-dimensional dependence. We establish identification and estimation
results for dependence profiles, introduce projection-residual diagnostics
that assess the adequacy of the covariance dictionary, and develop a
statistical theory for dependence classification. The analysis also characterizes ambiguous dependence structures through local alternatives and an indistinguishability result based on tangent-space overlap, showing that distinct dependence mechanisms may generate asymptotically indistinguishable similarity scores, thereby inducing classification uncertainty.

The framework delivers a concrete applied payoff. Proposition~\ref{prop:vcov_consistency}
shows that the profile-guided variance estimator $\widehat{V}^* = \widehat{V}_{\widehat{d}}$
is consistent under the dominant geometry, and the simulation experiment confirms that following the profile's
recommendation substantially reduces over-rejection relative to a default
White HC estimator---from approximately $80\%$ to near-nominal levels for
the cluster DGP and substantially below $80\%$ for the factor DGP.
The profile therefore answers the applied question directly: a researcher who
learns $(\widehat\omega_C,\widehat\omega_F,\widehat\omega_S)$ and applies the
corresponding robust procedure achieves better-sized inference than one who
applies an arbitrary correction.

More broadly, the paper suggests a geometric perspective in which dependence
structures become objects of statistical learning rather than maintained
assumptions, and dependence learning is governed by a geometric
identification problem: positive principal angles permit learning, whereas
tangent-space overlap creates fundamental ambiguity.

Future research may extend the covariance dictionary to additional dependence
geometries, develop optimal profile-weighted combinations of robust procedures
in the spirit of \citet{Hansen2007}, and establish the formal rate properties
of the profile-guided estimator under local misspecification.



\bibliographystyle{ecta_fallback_plainnat}
\bibliography{references}

\clearpage