Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
76,925 characters · 37 sections · 19 citation commands
Learning Dependence Structures for Econometric Inference
Robust inference under dependence is one of the central concerns of modern econometrics. Economic data frequently exhibit dependence arising from common institutional environments, aggregate shocks, social interactions, geographic proximity, production networks, financial interconnectedness, and repeated observations over time.
Prominent examples include cluster-robust inference White1980,LiangZeger1986,Arellano1987,CameronGelbachMiller2011,CameronMiller2015, spatial and HAC-type dependence Conley1999,DriscollKraay1998, factor-based methods ChamberlainRothschild1983,BaiNg2002,Bai2003, and network or sparse dependence models BickelLevina2008,CaiLiu2011,Auerbach2019,Leung2022. Although these approaches differ substantially in their assumptions, they share a common feature: the dependence structure is treated as known.
In practice, however, the underlying dependence architecture is rarely known. Firms may be exposed to common industry shocks, latent macroeconomic factors, and production-network spillovers. Regional economic outcomes may reflect state-level institutions, common national shocks, geographic spillovers, migration networks, and local interactions. The central challenge is often not merely robust inference under dependence but learning the dependence structure itself.
\paragraph{Motivating examples.}
Consider a panel of firms observed over time. Residual dependence may arise because firms are exposed to common industry shocks, latent aggregate factors, or production-network linkages. These mechanisms imply very different covariance geometries---cluster, factor, and sparse dependence---yet are often simultaneously plausible in practice.
As a second example, consider regional economic outcomes. Dependence may reflect state-level institutions, common national shocks, geographic spillovers, or migration networks. Standard robust inference procedures typically require the researcher to specify one of these dependence structures in advance.
In both examples, the primary challenge is not merely to conduct inference under dependence, but to learn which dependence geometry is most consistent with the observed data.
This paper develops a framework for dependence structure learning. Rather than viewing cluster, factor, and sparse dependence as competing models, we view them as distinct geometric classes of covariance operators. Cluster dependence generates approximately structured-support covariance structures. Factor dependence generates approximately low-rank covariance structures. Sparse dependence generates covariance structures in which dependence is concentrated among a relatively small subset of economically meaningful interactions.
The central object of interest is not a covariance estimator. Instead, the framework starts from an empirical dependence operator constructed from the sampling design. Depending on the application, this operator may correspond to a sample covariance matrix, a long-run covariance operator, a spatial dependence operator, a network dependence operator, or another empirical measure of dependence. The objective is to learn the geometry of dependence encoded in this operator.
The paper does not seek to estimate dependence by imposing a cluster, factor, or sparse model. Rather, it assumes the existence of a population dependence operator and studies what aspects of dependence geometry can be learned from a consistent estimate of that operator.
A key distinction between the present framework and existing approaches is that the object of interest is not the dependence operator itself. Even when a researcher can consistently estimate an appropriate dependence operator, the resulting object often remains high dimensional and difficult to interpret. The goal of dependence learning is therefore to recover the geometric organization encoded in the operator. The dependence profile provides a low-dimensional summary that quantifies the extent to which the operator resembles cluster, factor, and sparse covariance geometries. In this sense, dependence structures are treated as estimands rather than maintained assumptions.
Importantly, the objective is not necessarily to recover a unique dependence structure. Different covariance geometries may overlap locally, making ambiguity an inherent feature of the problem rather than a consequence of limited sample size. Dependence learning therefore seeks both to identify dominant dependence patterns and to diagnose situations in which multiple geometries provide similarly plausible descriptions of the observed dependence operator.
An analogy may help clarify the objective. As principal component analysis uses the covariance matrix to recover a low-dimensional representation of variation, our framework uses an empirical dependence operator to recover a low-dimensional representation of dependence geometry. The operator serves as an intermediate object through which the underlying dependence structure is learned.
An important feature of the framework is that it distinguishes between relative and absolute notions of dependence fit. A large similarity score indicates that one covariance geometry provides a better approximation than the alternatives under consideration, but it does not imply that the geometry provides a good approximation in an absolute sense. To address this issue, we introduce projection-residual diagnostics that measure the distance between the empirical dependence operator and the closest covariance geometry. These diagnostics allow the researcher to distinguish between clear geometric classifications, ambiguous dependence structures arising from overlapping covariance geometries, and situations in which none of the candidate geometries adequately captures the observed dependence pattern.
A central theme of the paper is the distinction between identifiable and ambiguous dependence structures. Identification is governed by the geometric separation of covariance classes. When covariance geometries are sufficiently separated, dependence profiles are locally identifiable and support reliable classification. When covariance geometries overlap, however, ambiguity may be intrinsic rather than a consequence of limited sample size. We show that the principal-angle condition, applied to the off-diagonal (dependence-relevant) part of each tangent space, provides a geometric identification criterion: positive principal angles ensure local distinguishability, whereas overlapping tangent spaces generate local indistinguishability. Theorem (ref) characterizes the identifiable region of the covariance-geometry space, while Theorem (ref) characterizes the boundary at which identification fails. Together, the two results delineate the boundary between identifiable and fundamentally ambiguous dependence structures.
The proposed dependence profile is also conceptually related to several familiar ideas in econometrics. In factor models, researchers often summarize variation through factor contribution measures. In forecast combination and model averaging, weights summarize the relative importance of competing predictive models BatesGranger1969,Hansen2007. In semiparametric efficiency theory, projection operators characterize efficient influence functions through orthogonal projections onto tangent spaces Newey1994,BickelKlaassenRitovWellner1993. The dependence profile serves a similar descriptive purpose: it does not represent an additive covariance decomposition, but summarizes the relative geometric proximity of the observed dependence operator to economically meaningful dependence classes.
Existing robust inference methods require the researcher to specify a dependence structure in advance. Our framework instead quantifies the extent to which alternative covariance geometries are supported by the data. To formalize this idea, we introduce projection operators onto cluster, factor, and sparse covariance classes, which induce a dependence profile $(\omega_C,\omega_F,\omega_S)$. These classes should be viewed as illustrative benchmark geometries rather than an exhaustive collection.
A central objective of the paper is not only to learn dependence structures, but also to use this information to guide econometric inference. We show that estimated dependence profiles can be used to construct profile-guided procedures that adapt to the dominant dependence geometry. Under suitable conditions, the resulting procedures are asymptotically equivalent to an infeasible oracle that knows the dominant covariance geometry in advance. Simulation evidence shows that profile-guided inference closely approximates the infeasible oracle procedure across a broad range of dependence environments.
The paper's theoretical contribution is organized around three principal results. Theorem (ref) establishes local identification of dependence profiles under a principal-angle separation condition. Theorem (ref) shows that when covariance geometries share common tangent directions, no statistical procedure can distinguish them at first order. Finally, Theorem (ref) establishes oracle adaptivity of profile-guided inference, showing that dependence learning can be used to construct inference procedures that are asymptotically equivalent to an infeasible oracle that knows the dominant covariance geometry in advance. The paper makes five contributions.
Geometric framework. We represent dependence structures as covariance geometries---closed subsets of the Hilbert space of symmetric matrices---and characterize their local structure through tangent spaces and principal angles. Cluster, factor, and sparse dependence are treated as geometric classes rather than competing maintained assumptions. The framework itself is not tied to these particular classes and can accommodate other covariance geometries, including spatial, network, long-memory, or other application-specific dependence structures.
Identification and estimation. We establish local identification of the dependence profile under a principal-angle separation condition, prove consistency, and derive asymptotic normality of the estimated dependence profile. Projection-residual diagnostics provide a complementary absolute goodness-of-fit measure.
Classification theory. We establish consistency of dominant-geometry classification (Theorem (ref)), derive finite-sample error bounds based on the separation margin $\Delta_\omega$ (Proposition (ref)), and characterize classification probabilities under local near-tie alternatives (Theorem (ref)).
Geometric ambiguity. We show that when two covariance-geometry tangent spaces share a common direction, no sequence of tests can distinguish the geometries at first order (Theorem (ref)). This impossibility result explains ambiguous dependence profiles and near-ties as consequences of local nonidentification rather than finite-sample noise.
Adaptive inference. We establish oracle adaptivity of profile-guided inference (Theorem (ref)). Estimated dependence profiles can be used to select dependence-robust procedures in a data-driven manner, and the resulting inference is asymptotically equivalent to an infeasible oracle that knows the dominant covariance geometry in advance.
This paper contributes to several strands of econometrics. First, it relates to inference under dependence, including clustered dependence LiangZeger1986,Arellano1987,CameronGelbachMiller2011, spatial dependence Conley1999,DriscollKraay1998, network dependence Auerbach2019,Leung2022, and factor dependence ChamberlainRothschild1983,BaiNg2002,Bai2003. These methods generally assume that the relevant dependence architecture is known.
Second, the paper relates to covariance regularization and high-dimensional dependence modeling. Sparse covariance estimation exploits localized dependence structures BickelLevina2008,CaiLiu2011,FriedmanHastieTibshirani2008, while factor models exploit low-rank covariance representations. Our framework does not seek to select a single covariance model, but to quantify geometric proximity to multiple covariance geometries.
Third, the paper is related to low-rank plus sparse decomposition and robust principal component analysis CandesLiMaWright2011,Chandrasekaran2011. Those methods aim to recover structural matrix decompositions. By contrast, our objective is to characterize the dependence architecture through similarity scores, allowing multiple mechanisms to coexist.
Finally, the paper draws on projection theory, matrix geometry, and semiparametric projection methods Deutsch2001,AbsilMahonySepulchre2008,Newey1994,BickelKlaassenRitovWellner1993. The principal-angle condition introduced below plays a central role in geometric identification and clarifies the sources of ambiguous dependence structures.
The paper is organized as follows. Section (ref) introduces the covariance geometry framework. Section (ref) develops identification and ambiguity results. Sections (ref) and (ref) study estimation and asymptotic theory. Section (ref) develops dependence diagnostics and classification. Section (ref) presents profile-guided inference. Sections (ref) and (ref) report simulation and empirical evidence.
Let \[ u=(u_1,\ldots,u_n)' \] denote the disturbance vector and define \[ \Sigma=E(uu'). \] We view \(\Sigma\) as an element of the Hilbert space \[ \mathbb H=\{A=A':\|A\|_F<\infty\}, \qquad \langle A,B\rangle_F=\operatorname{tr}(A'B). \] Rather than assuming that $\Sigma$ belongs to a single dependence class, we characterize its dependence architecture through its geometric proximity to several economically meaningful covariance classes.
Let \[ \mathcal S_C,\qquad \mathcal S_F,\qquad \mathcal S_S \] denote cluster, factor, and sparse covariance classes.
A covariance matrix belongs to the cluster class \(\mathcal S_C\) if its support can be represented as the union of a finite number of clustering dimensions. Suppose there are \(M\) clustering dimensions and let \(g_m(i)\), \(m=1,\ldots,M\), denote the cluster membership of observation \(i\) along dimension \(m\). Define the cluster-support matrix \(M_C\) by \[ (M_C)_{ij} = \mathbf{1}\Bigl\{g_m(i)=g_m(j)\text{ for at least one }m\in\{1,\ldots,M\}\Bigr\}. \] Thus \((M_C)_{ij}=1\) whenever observations \(i\) and \(j\) share at least one cluster membership, and \((M_C)_{ij}=0\) otherwise. A covariance matrix \(\Gamma\) belongs to \(\mathcal S_C\) if \(\Gamma_{ij}=0\) whenever \((M_C)_{ij}=0\), equivalently if \(\Gamma=M_C\odot\Gamma\). This definition includes one-way clustering (\(M=1\)), two-way clustering (\(M=2\)), and multiway clustering (\(M\geq 2\)) as special cases.
For a fixed or slowly growing rank \(r\), define the factor class as \[ \mathcal S_F(r) = \Big\{ \Gamma\succeq0: \Gamma=L+D,\; \operatorname{rank}(L)\le r,\; D \text{ diagonal} \Big\}. \]
For \(k_n=o(n^2)\), define the sparse class as \[ \mathcal S_S = \Big\{ \Gamma\in\mathbb H: |\operatorname{supp}(\Gamma)|\le k_n \Big\}. \]
The covariance dictionary considered in this paper consists of cluster, factor, and sparse geometries. These serve as benchmark geometries for illustration; the framework extends naturally to richer dictionaries.
Throughout the paper, let \[ \mathfrak D=\{C,F,S\} \] denote the set of covariance geometries, corresponding respectively to cluster, factor, and sparse dependence. Define \[ P_d(\Sigma) = \arg\min_{\Gamma\in\mathcal S_d} \|\Sigma-\Gamma\|_F. \] $P_d(\Sigma)$ represents the closest approximation to $\Sigma$ within the covariance class $\mathcal S_d$. Although we use projection notation, \(P_d\) need not be a linear orthogonal projection because the covariance classes may be nonlinear or nonconvex.
The Frobenius norm is induced by the Hilbert inner product \(\langle A,B\rangle_F=\operatorname{tr}(A'B)\). It therefore provides a natural geometry for defining projections, tangent spaces, principal angles, and local separation between covariance classes. It is also computationally convenient: low-rank projections are based on spectral truncation, cluster projections on support masks, and sparse projections on thresholding or sparse approximation. Alternative metrics may be useful in other applications, but the Frobenius geometry gives the most transparent framework for the identification and asymptotic theory developed below.
Define \[ S_d=\|P_d(\Sigma)\|_F^2, \qquad d\in\mathfrak D. \]
The normalized dependence similarity scores are \[ \omega_d=\frac{S_d}{\sum_{d\in\mathfrak D}S_{d}}, \qquad d\in\mathfrak D. \] The vector \[ \omega=(\omega_C,\omega_F,\omega_S)' \] is called the dependence profile.
The dependence profile is identified through the local geometry of the covariance classes. Identification depends on whether the cluster, factor, and sparse geometries remain sufficiently separated in a neighborhood of the population dependence operator.
For each geometry \(d\in\mathfrak D\), let \(T_d(\Sigma_d)\) denote the corresponding tangent space at a regular point \(\Sigma_d\). Since diagonal perturbations are common to all covariance geometries and carry no information about cross-sectional dependence, identification is based on the off-diagonal tangent spaces
\[ T_d^{\mathrm{off}}(\Sigma_d) = \left\{ H\in T_d(\Sigma_d) : H_{ii}=0 \ \text{for all } i \right\}, \qquad d\in\mathfrak D. \] \paragraph{Principal-Angle Condition.} For two tangent spaces \(U\) and \(V\), define the smallest principal angle by \[ \theta(U,V) = \inf_{u\in U,\;v\in V} \arccos \left( \frac{\langle u,v\rangle_F}{\|u\|_F\|v\|_F} \right). \]
Formal definitions of the covariance geometries, regularity conditions, and tangent spaces are given in the Online Appendix.
Assumption (ref) is standard. For Gaussian observations with covariance $\Gamma_0$, LAN holds with the Fisher information inner product induced by the Gaussian log-likelihood; see LeCamYang2000, Chapter 7.
This section describes how the population dependence profile is estimated from an empirical dependence operator. Let \(\widehat\Gamma_n\) denote an empirical dependence operator constructed from the data. The operator may be a covariance matrix, a long-run covariance operator, a spatial covariance operator, a network dependence operator, or another object summarizing the dependence features relevant for the inferential problem. Let \(\Gamma_0\) denote the corresponding population dependence operator.
The choice of \(\widehat\Gamma_n\) is application-specific. It may be a covariance, long-run covariance, spatial, network, or cluster-supported operator. The theory below is conditional on the chosen operator; a fuller discussion of operator choice and computational implementation is given in the Online Appendix.
The empirical dependence operator should capture the form of dependence relevant for the inferential problem. For contemporaneous cross-sectional dependence, a natural benchmark is the residual covariance operator \[ \widehat\Gamma_n = \frac{1}{T} \sum_{t=1}^{T} \widehat u_t\widehat u_t', \qquad \widehat u_t=(\widehat u_{1t},\ldots,\widehat u_{nt})'. \] For serially dependent data one may instead use a long-run covariance operator. For spatial, network, or clustered data, the operator may incorporate distances, adjacency relations, or cluster-support restrictions.
The framework does not seek to select a unique dependence operator. Rather, it studies the covariance geometry encoded in a chosen population operator. The dependence profile is invariant to positive rescalings of the operator; additional examples and technical details are provided in the Online Appendix.
The dependence profile provides a relative measure of geometric fit. To assess goodness-of-fit in an absolute sense, we introduce projection-residual diagnostics.
To quantify absolute goodness-of-fit, define the normalized projection residual \[ \rho_d(\Gamma) = \frac{ \| \Gamma-P_d(\Gamma) \|_F } { \|\Gamma\|_F }, \qquad d\in\mathfrak D. \] The quantity \(\rho_d(\Gamma)\) measures the fraction of the operator that remains unexplained after projection onto geometry \(d\). Since \( 0\in\mathcal S_d \), the definition of the projection implies
\[ \|\Gamma-P_d(\Gamma)\|_F \le \|\Gamma\|_F. \] Consequently, \[ 0 \le \rho_d(\Gamma) \le 1. \] Small values indicate that geometry \(d\) provides a good approximation to the dependence operator, whereas large values indicate substantial lack of fit.
Define the minimum residual \[ \rho_{\min}(\Gamma) = \min_{d\in\mathfrak D} \rho_d(\Gamma). \] The quantity \(\rho_{\min}(\Gamma)\) summarizes the distance from the dependence operator to the closest covariance geometry in the dictionary.
We say that a dependence operator exhibits a “none-of-the-above” pattern whenever \[ \rho_{\min}(\Gamma)>\delta, \] for a prespecified threshold \(\delta\in(0,1)\).
Large values of \(\rho_{\min}(\Gamma)\) indicate that none of the cluster, factor, or sparse geometries provides an adequate approximation to the observed dependence structure. In such situations, the dependence profile should be interpreted with caution, since the operator may contain dependence features not represented in the covariance dictionary.
The population dependence profile is defined from projections of \(\Gamma_0\) onto the covariance geometries. For each \(d\in\mathfrak D\), define
\[ P_{d,0} = P_d(\Gamma_0), \qquad S_{d,0} = \|P_{d,0}\|_F^2. \] The population similarity score associated with geometry \(d\) is
\[ \omega_{d,0} = \frac{S_{d,0}}{\sum_{d\in\mathfrak D}S_{d,0}}, \qquad d\in\mathfrak D. \]
The vector
\[ \omega_0 = ( \omega_{C,0}, \omega_{F,0}, \omega_{S,0} )' \] is the population dependence profile.
Define
\[ \widehat S_d = \| \widehat P_d \|_F^2, \qquad d\in\mathfrak D. \] The estimated similarity score for geometry \(d\) is \[ \widehat\omega_d = \frac{\widehat S_d}{\sum_{d\in\mathfrak D} \widehat S_d}, \qquad d\in\mathfrak D. \] The estimated dependence profile is \[ \widehat\omega = ( \widehat\omega_C, \widehat\omega_F, \widehat\omega_S )'. \]
This section establishes consistency and asymptotic normality of the estimated dependence profile.
To state the limiting distribution, define the score map
\[ \mathcal S(\Gamma) = \left( \|P_C(\Gamma)\|_F^2, \|P_F(\Gamma)\|_F^2, \|P_S(\Gamma)\|_F^2 \right)'. \] Let \[ s_0 = \mathcal S(\Gamma_0) = ( S_{C,0}, S_{F,0}, S_{S,0} )'. \] Let \[ \pi(s) = \frac{s}{\mathbf 1's} \] denote the normalization map from similarity scores to the dependence profile, where \(\mathbf 1=(1,1,1)'\). Thus, \(\omega_0=\pi(s_0)\).
The dependence profile provides a low-dimensional summary of the dependence architecture encoded in the empirical dependence operator. Large values of \(\widehat{\omega}_C\), \(\widehat{\omega}_F\), and \(\widehat{\omega}_S\) indicate that the empirical dependence operator is geometrically close to the cluster, factor, and sparse covariance classes, respectively.
The dependence profile measures relative geometric similarity rather than additive covariance contributions. A natural classification rule is
The estimated geometry \(\widehat d\) identifies the covariance class with the largest estimated similarity score. The full dependence profile is often more informative than the classification alone. For example, \[ (0.60,0.30,0.10) \] suggests cluster dominance with meaningful factor dependence, whereas \[ (0.35,0.35,0.30) \] suggests hybrid dependence and cautions against a purely binary classification.
Inference for the dependence profile is based on Theorem (ref), \[ \sqrt n \left( \widehat{\omega}-\omega_0 \right) \Rightarrow N(0,\Xi), \] where \[ \omega_0 = (\omega_{C,0},\omega_{F,0},\omega_{S,0})'. \]
Let \[ e_C=(1,0,0)', \qquad e_F=(0,1,0)', \qquad e_S=(0,0,1)'. \] For \(d\in\mathfrak D\), the asymptotic variance of \(\widehat{\omega}_d\) is \[ \sigma_d^2 = e_d'\Xi e_d. \]
Under Assumption (ref), \[ \widehat{\sigma}_d^2 = e_d'\widehat{\Xi}e_d \] is a consistent estimator of \(\sigma_d^2\).
The construction of \(\widehat{\Xi}\) depends on the chosen empirical dependence operator. When the operator admits an asymptotic linear representation, a plug-in estimator based on the influence function and the delta method may be used. The construction of a universal variance estimator for all possible dependence operators is outside the scope of the paper.
To test \[ H_0:\omega_C=\omega_C^\ast, \] consider \[ T_C = \frac{ \sqrt n \left( \widehat{\omega}_C-\omega_C^\ast \right) } {\widehat{\sigma}_C}, \] where \[ \widehat{\sigma}_C^2 = e_C'\widehat{\Xi}e_C. \] Under \(H_0\), \[ T_C \Rightarrow N(0,1). \] Analogous statistics apply to the factor and sparse scores.
More generally, let \(R_{\omega}\) be a \(q\times 3\) matrix of full row rank and consider \[ H_0: R_{\omega}\omega_0=r_{\omega}. \] Define \[ W_{\omega} = n \left( R_{\omega}\widehat{\omega}-r_{\omega} \right)' \left( R_{\omega}\widehat{\Xi}R_{\omega}' \right)^{-1} \left( R_{\omega}\widehat{\omega}-r_{\omega} \right). \] Under \(H_0\), \[ W_{\omega} \Rightarrow \chi_q^2. \]
These tests quantify the uncertainty associated with the estimated dependence profile and should be interpreted as diagnostic tools.
A dominant dependence geometry is suggested when one similarity score substantially exceeds the others. For example, \[ \widehat{\omega}_C > \max \left\{ \widehat{\omega}_F, \widehat{\omega}_S \right\} \] indicates that the empirical dependence operator is closest to the cluster covariance geometry.
Hybrid dependence is suggested when several scores are simultaneously large. In such cases, the full dependence profile should be reported rather than reducing the analysis to a single classification label.
This distinction is important because covariance geometries may overlap. For example, cluster covariance matrices may also be sparse when cluster sizes are small relative to the sample size. The dependence profile should therefore be viewed as a continuous description of dependence architecture rather than a rigid model-selection device.
A central question is whether the estimated dependence profile can be used to identify the dominant covariance geometry underlying the chosen dependence operator.
Under Assumption (ref), define
The quantity \(d^{\star}\) represents the population dominant dependence geometry.
Assumption (ref) rules out knife-edge cases in which two or more covariance geometries have exactly the same population similarity score. Similar separation conditions are common in model selection and statistical classification problems.
Theorem (ref) shows that the dependence profile consistently identifies the dominant covariance geometry whenever the dominant similarity score is separated from the remaining scores. Consequently, the dependence profile is not merely a descriptive summary of dependence but also provides a statistically consistent basis for dependence classification.
The classification result provides a formal justification for using the estimated dependence profile as a guide for procedure recommendation. The implications for inference are developed formally in Section (ref).
The dependence profile provides a low-dimensional summary of the dependence geometry encoded in an empirical dependence operator. A natural question is whether the estimated profile can be used to guide econometric inference. The objective of this section is not to propose a new covariance estimator. Rather, we show how dependence learning can be combined with existing dependence-robust procedures in a systematic way.
Let $\widehat d$ denote the estimated dominant dependence geometry as defined in (ref). The interpretation is straightforward. When \( \widehat\omega_C \) dominates, cluster-robust procedures provide a natural benchmark. When \( \widehat\omega_F \) dominates, factor-robust or common-shock-adjusted procedures are suggested. When \( \widehat\omega_S \) dominates, sparse, network, spatial, or local-dependence robust procedures deserve particular attention.
The recommendation should be viewed as evidence-based rather than deterministic. The dependence profile does not imply that one procedure is universally correct. Rather, it summarizes which covariance geometry appears most consistent with the observed dependence operator.
When multiple similarity scores are substantial, the data provide evidence for several dependence mechanisms simultaneously. In such situations, reporting inference results from multiple dependence-robust procedures may be preferable. Similarly, when the projection-residual diagnostics indicate poor fit, the entire covariance dictionary should be viewed with caution.
A natural alternative is to use a single broadly robust procedure, such as multiway clustering, HAC inference, or a general sandwich estimator. Dependence learning remains useful for three reasons.
First, broad robustness is not costless. Procedures designed to be valid under large classes of dependence may be conservative or poorly sized when the realized dependence geometry differs from the structure for which the procedure is best suited. Second, such procedures still require the researcher to specify a dependence class, such as clustering dimensions, a spatial metric, a bandwidth, or a network. These choices are themselves assumptions about dependence. Third, Theorem (ref) shows that profile-guided selection is asymptotically equivalent to an infeasible oracle that knows the dominant covariance geometry.
Thus, the value of dependence learning is not only diagnostic. It provides a data-driven way to choose among geometry-specific inference procedures while retaining oracle-equivalent first-order behavior.
The reliability of a profile-based procedure recommendation depends on two considerations. First, the dominant geometry should be sufficiently separated from the competing geometries. Define the estimated separation margin
\[ \widehat\Delta_\omega = \max_{d\in\mathfrak D} \widehat\omega_d - \max_{d\neq \widehat d} \widehat\omega_d. \]
Second, the covariance dictionary should provide a satisfactory approximation to the empirical dependence operator. Recall the minimum projection residual
\[ \widehat\rho_{\min} = \min_{d\in\mathfrak D} \widehat\rho_d. \]
Combining these quantities yields the procedure confidence index
\[ \widehat\kappa = \bigl( 1-\widehat\rho_{\min} \bigr) \widehat\Delta_\omega. \]
Large values of \( \widehat\kappa \) indicate that one geometry is clearly dominant and that the covariance dictionary provides a good approximation to the observed dependence operator. Small values indicate either weak separation, poor geometric fit, or both.
The index therefore summarizes the strength of the empirical evidence supporting a profile-based procedure recommendation. Let \[ \kappa_0 = \bigl( 1-\rho_{\min,0} \bigr) \Delta_{\omega,0}, \qquad \rho_{\min,0} = \min_{d\in\mathfrak D} \rho_d, \qquad \Delta_{\omega,0} = \max_{d\in\mathfrak D} \omega_d - \max_{d\neq d^\star} \omega_d, \text{ with } d^\star = \arg\max_{d\in\mathfrak D} \omega_d. \]
Suppose a scalar parameter \(\theta_0\) is estimated by \(\widehat\theta_n\). For each geometry \(d\in\mathfrak D\), let \(\widehat V_d\) be a measurable variance estimator and let \(\tau_{n,d}\to\infty\) be a deterministic normalization such that \[ \tau_{n,d}^{2}Var(\widehat\theta_n)\longrightarrow V_d, \qquad 0<V_d<\infty. \]
Define \[ \widehat d = \arg\max_{d\in\mathfrak D}\widehat\omega_d, \qquad \widehat V^{*} = \widehat V_{\widehat d}, \] where ties in the argmax are broken by a fixed deterministic rule. Thus, inference is based on the variance estimator associated with the estimated dominant dependence geometry.
Proposition (ref) shows that profile-guided selection preserves consistency whenever the dominant geometry is correctly classified. The next result establishes a stronger oracle property.
When the asymptotic normalization depends on the covariance geometry, comparison with the oracle statistic is naturally made on the event \(\{\widehat d=d^\star\}\). On this event the selected rate \(\tau_{n,\widehat d}\) and variance estimator \(\widehat V^*\) coincide exactly with \(\tau_{n,d^\star}\) and \(\widehat V_{d^\star}\). Since this event occurs with probability approaching one, the profile-guided and oracle statistics are asymptotically equivalent.
Theorem (ref) and Corollary (ref) provide the main econometric payoff of dependence learning. Estimated profiles can be used to select dependence-robust procedures in a data-driven way, and the resulting inference is asymptotically equivalent to the infeasible oracle procedure that knows the dominant dependence geometry in advance.
The profile-guided estimator selects a single dominant geometry. An alternative approach is to combine multiple procedures using the estimated dependence profile itself. Define
\[ \widehat V^{\mathrm{avg}} = \sum_{d\in\mathfrak D} \widehat\omega_d \widehat V_d. \]
The resulting estimator may be interpreted as a dependence-weighted combination of geometry-specific variance estimators, analogous to forecast combination and model averaging BatesGranger1969,Hansen2007. Developing efficiency theory and optimal weighting schemes for profile-weighted inference is an important direction for future research.
This section summarizes Monte Carlo evidence for the dependence-learning framework. The simulations focus on three questions. First, can the estimated dependence profile recover the dominant covariance geometry? Second, how does classification behave under hybrid and near-tie dependence? Third, does profile-guided inference behave like the infeasible oracle procedure that knows the dominant dependence geometry in advance?
The detailed data-generating processes, parameter calibrations, projection algorithms, and additional robustness exercises are reported in the Online Appendix. The baseline designs include pure cluster, pure factor, sparse network, cluster--factor hybrid, cluster--sparse hybrid, factor--sparse hybrid, all-three hybrid, and two-way cluster dependence.
Table (ref) reports the population dependence profiles and the Monte Carlo behavior of the estimated profiles. The profile recovers the dominant geometry in the benchmark cluster and factor designs and records substantial sparse affinity under designs in which support restrictions are important, including two-way clustering. The procedure confidence index \(\kappa_0\) is largest when one geometry is clearly separated and the covariance dictionary provides a good absolute fit.
Figure (ref) shows the average estimated profile along a cluster--factor hybrid path. The profile changes smoothly as the relative strength of factor dependence increases, illustrating that the method summarizes hybrid dependence rather than forcing a binary classification.
The classification theory predicts that dominant-geometry classification is reliable when the separation margin is large and remains probabilistic under local near-ties. Table (ref) reports classification frequencies for local cluster--factor alternatives, and Figure (ref) plots the relationship between classification error and the population separation margin.
The main econometric implication of Section (ref) is that profile-guided inference should be asymptotically equivalent to an infeasible oracle procedure. Table (ref) reports finite-sample coverage probabilities for the oracle procedure, the profile-guided procedure, and a deliberately misspecified procedure.
The profile-guided procedure is nearly indistinguishable from the oracle procedure across the benchmark designs, while the misspecified procedure can exhibit severe undercoverage. These results provide finite-sample support for Theorem (ref) and Corollary (ref). They also give the main practical message of the paper: learning the dependence profile can guide the choice of robust inference procedure in a way that tracks the infeasible oracle benchmark.
This section illustrates how the proposed dependence profile and projection-residual diagnostics can be used in an empirical setting. The goal is not to provide a new asset-pricing analysis, but to show how the framework can be applied to a publicly available dataset in which several dependence mechanisms are plausible.
We use the Fama--French 49 industry portfolios and the Fama--French three factors from the data library of FamaFrench1993 and FamaFrenchDataLibrary. The data contain value-weighted and equal-weighted returns for U.S. industry portfolios and are available at daily, monthly, and annual frequencies. The monthly and daily series begin in July 1926. The Data Library also provides the standard asset-pricing factors used in empirical finance.
The empirical setting is well suited to the proposed framework. Industry portfolio returns may exhibit cluster-like dependence because industries can be grouped into broader sectors. They may exhibit factor dependence because returns load on common market-wide and macro-financial shocks. They may also exhibit sparse dependence because some industries are more closely linked through input-output relationships, supply chains, or common demand shocks than others.
The empirical analysis uses monthly excess returns on the 49 industry portfolios together with the market, SMB, and HML factors from the Kenneth R. French Data Library FamaFrenchDataLibrary. Let $ R_{it}$ denote the excess return on industry portfolio \(i\) in month \(t\), where $i=1,\ldots,N$, and $t=1,\ldots,T$. In the baseline illustration, \(N=49\). We estimate the factor model
where \(MKT_t\), \(SMB_t\), and \(HML_t\) denote the Fama--French market, size, and value factors. The residual \(u_{it}\) captures the component of industry returns not explained by the observed factors.
Stack residuals across industries as $ \widehat u_t = (\widehat u_{1t},\ldots,\widehat u_{Nt})'$.
The baseline empirical dependence operator is the residual covariance matrix
We also consider the residual correlation operator
where $\widehat D=\operatorname{diag}(\widehat\Gamma)$. The covariance operator preserves differences in residual volatility across industries, whereas the correlation operator focuses on dependence geometry after normalizing scale.
We compute projections of the empirical dependence operator onto the cluster, factor, and sparse covariance geometries.
\paragraph{Cluster Geometry.}
The cluster geometry introduced in this paper encompasses one-way, two-way, and multiway clustering through a structured-support representation. In the present empirical application, we use a one-way clustering specification based on the Fama--French 10-sector economic classification (Consumer Non-Durables, Consumer Durables, Manufacturing, Energy, Hi-Tech, Telecom, Shops, Health, Utilities, and Other).
Let \[ g(i)\in\{1,\ldots,G\} \] denote the sector membership of industry portfolio \(i\). Define the cluster-support matrix
\[ (M_C)_{ij} = \mathbf{1}\{g(i)=g(j)\}. \]
The cluster projection is
\[ \widehat P_C = M_C \odot \widehat\Gamma. \]
Thus, the empirical implementation treats industry sectors as the relevant clustering dimension and measures the extent to which the residual dependence operator conforms to the resulting cluster-support structure. If additional clustering dimensions were available, the cluster-support matrix could be constructed according to the multiway definition in Section (ref).
\paragraph{Factor Geometry.}
The factor projection \(\widehat P_F\) is computed by the alternating-projection algorithm described in Online Appendix, applied to the empirical dependence operator \(\widehat\Gamma\). The output decomposes as \(\widehat P_F = L^* + D^*\), where \(L^*\) is a rank-\(r\) positive-semidefinite matrix and \(D^*\) is diagonal. In the baseline implementation, \(r\) is selected using an eigenvalue-ratio criterion applied to \(\widehat\Gamma-\widehat D\), where \(\widehat D\) is the diagonal of \(\widehat\Gamma\). As robustness checks, we also consider fixed ranks \(r\in\{1,2,3\}\).
\paragraph{Sparse Geometry.}
The sparse projection is computed by retaining the \(k_N\) largest off-diagonal entries of \(\widehat\Gamma\) in absolute value and setting all remaining off-diagonal entries to zero. The diagonal entries are retained. Formally, \[ \widehat P_S = \operatorname*{arg\,min}_{\Gamma\in\mathcal S_S} \| \widehat\Gamma-\Gamma \|_F^2. \] This projection provides the closest sparse approximation to the empirical dependence operator under Frobenius loss.
We apply the estimated dependence profile and projection-residual diagnostics defined in Sections (ref) and (ref), with \(\widehat\Gamma\) taken to be the residual covariance operator in (ref).
The dependence profile and the projection residuals answer different questions. The profile describes which geometry provides the strongest relative fit among the candidate classes. The residual diagnostics assess absolute fit. A large value of \(\widehat\rho_{\min}\) indicates that none of the candidate geometries provides a close approximation to the empirical dependence operator, even if one geometry receives the largest relative similarity score.
A large \(\widehat{\omega}_C\) indicates that residual dependence is well described, in relative terms, by a cluster-support structure. In the present application this corresponds to dependence concentrated within broad industry sectors. A large \(\widehat{\omega}_F\) suggests that the residual covariance still contains low-rank common components not fully absorbed by the observed Fama--French factors. A large \(\widehat{\omega}_S\) indicates that dependence is concentrated among a relatively small subset of industry pairs.
To connect the dependence profile with conventional robust inference, we compare standard errors for a pooled regression of industry excess returns on the Fama--French factors. Specifically, estimate
We report heteroskedasticity-robust standard errors, standard errors clustered by broad industry sector, and two-way clustered standard errors by industry and month. We also report a common-shock benchmark that removes leading principal components from the residual covariance operator before computing standard errors.
The purpose of this comparison is descriptive. If the estimated profile assigns substantial weight to the cluster geometry, clustered standard errors should be viewed as empirically relevant. If the factor score is large, then procedures that account for common shocks or factor dependence deserve attention. If the sparse score is large, then network- or sparse-dependence robust procedures may be important. The dependence profile therefore serves as a diagnostic summary of the covariance architecture underlying the robust inference procedures considered in practice.
Table (ref) reports the estimated dependence profiles and projection-residual diagnostics for the residual covariance and correlation operators.
Figure (ref) displays the estimated dependence profiles graphically.
Table (ref) compares conventional robust standard errors for the pooled factor regression.
Table (ref) reveals a hybrid dependence structure.
Covariance operator. The factor and sparse scores are tied at $\widehat\omega_F=\widehat\omega_S=0.445$, with a small cluster score $\widehat\omega_C=0.110$. The minimum projection residual $\widehat\rho_{\min}=0.008$ indicates that the sparse projection provides a near-perfect absolute fit to the residual covariance operator. Taken together, these findings indicate that residual co-movement among the 49 industry portfolios---after removing the three Fama--French factors---is concentrated in a sparse set of industry pairs rather than being driven by broad industry-sector clustering. The small $\widehat\rho_{\min}$ confirms that the sparse geometry adequately captures this residual dependence in an absolute sense. The non-negligible factor score suggests that some low-rank common component remains in the residuals despite the three-factor adjustment; this is consistent with known limitations of the Fama--French model as a complete description of common risk.
Correlation operator. Normalizing by industry volatility shifts the profile toward sparse dominance ($\widehat\omega_S=0.397$) with a moderate cluster score ($\widehat\omega_C=0.255$) and $\widehat\rho_{\min}=0.320$. The larger minimum residual indicates that the correlation structure is less well captured by the candidate geometries than the covariance structure, suggesting that industry-level volatility differences mask some localized dependence that only emerges after scale normalization.
Standard errors. Table (ref) shows that sector-clustered standard errors are roughly $2.3\times$ the White standard errors for the market factor ($0.051$ vs.\ $0.022$), and two-way clustered standard errors are slightly larger still ($0.060$). The common-shock adjusted standard errors are substantially smaller ($0.007$), reflecting the removal of the residual common factor component. These patterns are consistent with the estimated profile: the hybrid factor--sparse dependence structure implies that sector clustering overstates the within-cluster dependence, while the residual common component---identified by the factor score---is better absorbed by a common-shock adjustment than by clustering.
Procedure confidence index and inference recommendation. Table (ref) reports the procedure confidence index $\widehat\kappa$ and the profile-guided inference recommendation for each dependence operator.
Both operators yield small values of $\widehat\kappa$. For the covariance operator, the factor and sparse scores are exactly tied at $0.445$, so $\widehat\Delta_\omega=0$ and $\widehat\kappa=0$: the profile provides no statistical basis for preferring any single robust procedure. For the correlation operator, the sparse score marginally dominates ($\widehat\omega_S=0.397$), but the separation margin is only $\widehat\Delta_\omega=0.050$ and the minimum residual is $\widehat\rho_{\min}=0.320$, giving $\widehat\kappa=0.034$.
These low confidence-index values have a clear practical implication: they correctly indicate that the empirical dependence structure is genuinely hybrid, and that no single inference procedure can be recommended with high confidence based on the data alone. The appropriate response---which the framework prescribes directly---is to report inference results from multiple procedures and document sensitivity to the choice. Table (ref) operationalises this recommendation: the common-shock adjusted standard errors (appropriate when factor dependence is present) and the sector-clustered standard errors (appropriate when cluster dependence dominates) differ by a factor of approximately $7\times$ for the market factor, confirming that the choice of procedure is empirically consequential.
This example illustrates a key advantage of the proposed framework over a priori procedure selection. A researcher who assumes cluster dependence and uses sector-clustered standard errors is implicitly claiming $\widehat\kappa\approx 1$ for the cluster geometry---a claim the data refute. A researcher who learns the profile first discovers the hybrid structure, obtains $\widehat\kappa\approx 0$, and is correctly directed to report multiple procedures rather than to commit to one. The framework thus transforms an informal sensitivity analysis into a statistically grounded diagnostic.
The empirical illustration uses publicly available industry portfolio returns and factor data from Kenneth French's Data Library. The replication package downloads the data, constructs excess returns, estimates (ref), computes the empirical dependence operators in (ref) and (ref), estimates dependence profiles and projection-residual diagnostics, and generates Tables (ref) and (ref) and Figure (ref).
All data transformations, sector mappings, parameter choices, and random seeds used in the empirical illustration are documented in the replication files.
This paper develops a framework for learning dependence structures from empirical measures of dependence. Rather than treating dependence structures as maintained assumptions, we view cluster, factor, and sparse dependence as covariance geometries and measure their proximity to an empirical dependence operator.
The proposed dependence profile provides a low-dimensional summary of high-dimensional dependence. We establish identification and estimation results for dependence profiles, introduce projection-residual diagnostics that assess the adequacy of the covariance dictionary, and develop a statistical theory for dependence classification. The analysis also characterizes ambiguous dependence structures through local alternatives and an indistinguishability result based on tangent-space overlap, showing that distinct dependence mechanisms may generate asymptotically indistinguishable similarity scores, thereby inducing classification uncertainty.
The framework delivers a concrete applied payoff. Proposition (ref) shows that the profile-guided variance estimator $\widehat{V}^* = \widehat{V}_{\widehat{d}}$ is consistent under the dominant geometry, and the simulation experiment confirms that following the profile's recommendation substantially reduces over-rejection relative to a default White HC estimator---from approximately $80\%$ to near-nominal levels for the cluster DGP and substantially below $80\%$ for the factor DGP. The profile therefore answers the applied question directly: a researcher who learns $(\widehat\omega_C,\widehat\omega_F,\widehat\omega_S)$ and applies the corresponding robust procedure achieves better-sized inference than one who applies an arbitrary correction.
More broadly, the paper suggests a geometric perspective in which dependence structures become objects of statistical learning rather than maintained assumptions, and dependence learning is governed by a geometric identification problem: positive principal angles permit learning, whereas tangent-space overlap creates fundamental ambiguity.
Future research may extend the covariance dictionary to additional dependence geometries, develop optimal profile-weighted combinations of robust procedures in the spirit of Hansen2007, and establish the formal rate properties of the profile-guided estimator under local misspecification.