Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,110 characters · 16 sections · 45 citation commands
The Spurious Factor Dilemma: Robust Inference in Heavy-Tailed Elliptical Factor Models
KEY WORDS: Elliptical distributions; Factor models; Heavy tails; Spurious factors.
\addtocontents{toc}{\setcounter{tocdepth}{2}}
Factor models serve as a cornerstone in the analysis of large-scale datasets across various disciplines, including economics, finance, genetics, and signal processing. Their power lies in the ability to parsimoniously capture complex dependencies and interactions among numerous variables by attributing them to a small number of latent common factors bai2002determining, stock2002forecasting, fan2018large. By effectively modeling these latent structures, factor analysis provides crucial tools for dimension reduction, feature extraction, and understanding the underlying drivers of observed phenomena fan2021recent.
A fundamental challenge in applying factor models is the determination of the correct number of common factors, denoted by $m$. This problem has received considerable attention, yet remains a subject of ongoing research due to its critical impact on subsequent analysis. Underestimating $m$ leads to the omission of significant systematic components, potentially resulting in biased estimates of factor loadings, inconsistent forecasting, and flawed structural interpretations bai2003inferential, baltagi2017identification. Conversely, overestimating $m$ introduces noise by fitting spurious factors, which can inflate estimation variance, reduce model interpretability, and increase computational costs barigozzi2020consistent.
Existing methodologies for estimating $m$ largely fall into two categories. The first, prevalent in econometrics, leverages the connection between factor analysis and Principal Component Analysis (PCA). Methods like the information criteria of bai2002determining and alessi2010improved, the eigenvalue ratio tests of ahn2013eigenvalue and lam2012factor, and the randomization tests trapani2018randomized, kong2020random rely on the assumption that eigenvalues associated with common factors diverge at a faster rate than those corresponding to idiosyncratic noise as dimensions grow. The second category employs Random Matrix Theory (RMT) to provide finer distinctions, particularly in high-dimensional settings where both the number of variables ($p$) and observations ($n$) are large. RMT-based tests, such as those by onatski2009testing, utilize the fact that the largest noise eigenvalues converge to the Tracy–Widom distribution, while factor-related eigenvalues appear as distinct outliers, often asymptotically Gaussian after proper scaling onatski2010determining, cai2020limiting, ke2023estimation.
However, both classes of methods face limitations when faced with data exhibiting heavy-tailed randomness. Heavy-tailed distributions, characterized by a higher probability of extreme events compared to Gaussian distributions, are ubiquitous in financial returns, climate data, and various other fields roy2021empirical,ke2023estimation. PCA-based methods, while relatively robust to certain types of noise dependence, can falter because heavy tails can generate large sample eigenvalues purely from noise, mimicking the signature of true factors. RMT-based methods typically rely on moment conditions or concentration properties that are violated by heavy-tailed noise, leading them to misinterpret large noise eigenvalues as signals. Figure (ref) provides a visual example of this issue, where a large spurious eigenvalue appears far from the bulk, potentially misleading standard selection criteria.
Furthermore, empirical data often exhibit non-linear dependencies alongside heavy tails. Standard factor models typically assume idiosyncratic errors or linear dependencies, potentially missing complex interaction patterns. Elliptical Factor Models (EFM), based on elliptical distributions, offer a flexible framework that naturally incorporates both heavy-tailedness (via the radial component) and non-linear dependencies (via the elliptical structure) chamberlain1982arbitrage,baltagi2017identification. Recently, bao2025signal has demonstrated that even heteroscedastic or cross-correlated noise with independent entries along the time dimension can produce misleading spikes, which traditional singular value methods may incorrectly interpret as signals. As a result, spurious eigenvalues are quite common in practice.
This paper addresses the critical issue of factor number overestimation in the EFM framework, specifically focusing on the confusion caused by heavy-tailed noise. We pose two central questions: Can we reliably detect spurious factors generated by heavy tails in elliptical models? Can we develop a robust factor selection procedure for such data?
Our primary contribution is a novel methodology with a rigorous theoretical basis for distinguishing between “real" factor signals and “spurious" noise-induced signals among the large sample eigenvalues. We introduce a novel algorithm called the fluctuation magnification algorithm, specifically designed for sample covariance matrices derived from EFM data. The core idea is that perturbing the data via an elaborated magnifier affects real and spurious signals differently. We theoretically establish that, under the fluctuation magnification, the (appropriately scaled) eigenvalues corresponding to true common factors exhibit stability, converging to a normal distribution or having small variance relative to their magnitude. In contrast, spurious eigenvalues generated by heavy tails display significantly larger fluctuations under the perturbation. This difference in stability provides a clear mechanism for detection.
Based on these distinct asymptotic behaviors (detailed in Sections (ref) and (ref)), we develop a test statistic that quantifies the fluctuation of each large sample eigenvalue under the magnification algorithm. This allows us to formally test for the presence of spurious factors (Section (ref)). As a key application, we integrate this detection mechanism into a two-step procedure to robustly estimate the number of common factors ($m$) in heavy-tailed EFMs. The first step uses the magnification algorithm to identify potential spurious signals among the leading eigenvalues. The second step employs existing criteria (e.g., from onatski2010determining) as a safeguard, primarily to handle cases where the noise might be light-tailed, ensuring consistency across different tail behaviors.
Our work can be viewed as being generated from high-dimensional resampling techniques lopes2019bootstrapping, han2018gaussian, ding2023extreme, ke2023estimation, yao2021rates, yu2024testing to the challenging setting of elliptical factor models with heavy tails, providing the first procedure, to our knowledge, specifically designed to detect spurious factors in this context. Numerical simulations demonstrate the superior performance of our method compared to established techniques, especially in heavy-tailed scenarios. Application to real financial data yields results consistent with financial theory and highlights the practical relevance of addressing spurious factors.
The remainder of the paper is organized as follows. Section (ref) formally introduces the Elliptical Factor Model and outlines the key assumptions. Section (ref) further illustrates the problem of spurious factors using examples. Section (ref) presents the main asymptotic theory that details the behavior of real and spurious eigenvalues. Section (ref) describes the fluctuation magnifier algorithm and the proposed testing and factor selection procedures. Section (ref) provides simulation results, and Section (ref) discusses the real data application. We provide a sketch for our proof strategy for the theoretical results in Section (ref). The conclusion is offered in Section (ref). All detailed technical proofs are deferred to Appendix (ref).
Let $\mathbb{C}_+$ denote the complex upper half-plane. We use $C > 0$ to represent a generic positive constant whose value may change from line to line. For sequences of positive deterministic values $\{a_n\}$ and $\{b_n\}$, $a_n = \mathrm{O}(b_n)$ means $a_n \leq C b_n$ for some $C > 0$. If $a_n = \mathrm{O}(b_n)$ and $b_n = \mathrm{O}(a_n)$, we write $a_n \asymp b_n$. We write $a_n = \mathrm{o}(b_n)$ if $a_n \leq c_n b_n$ for some positive sequence $c_n \downarrow 0$. For a sequence of random variables $\{x_n\}$ and positive real values $\{a_n\}$, $x_n = \mathrm{O}_{\mathbb{P}}(a_n)$ indicates that $x_n / a_n$ is stochastically bounded. $x_n = \mathrm{o}_{\mathbb{P}}(a_n)$ means $x_n / a_n$ converges to zero in probability. For a sequence of positive random variables $\{y_n\}$, $y_{(k)}$ denotes the $k$-th order statistic, $y_{(1)} \geq y_{(2)} \geq \cdots \geq y_{(n)} > 0$. Vectors are marked in bold.
Elliptical distributions provide a versatile class for modeling multivariate data, extending the normal distribution to accommodate heavy tails and capture specific dependence structures like tail dependence, making them particularly relevant in finance and other fields chamberlain1982arbitrage, fama1993common, baltagi2017identification. A $p$-dimensional random vector $\mathbf{y}$ follows a centered elliptical distribution, denoted $\mathbf{y} \sim EC_p(0, \Sigma, \xi)$, if it admits the stochastic representation:
where $\Sigma \in \mathbb{R}^{p \times p}$ is a positive definite matrix representing the population covariance matrix, $\xi \ge 0$ is a non-negative scalar random variable representing the “radius," and $\mathbf{u} \in \mathbb{R}^p$ is a random vector uniformly distributed on the unit sphere $\mathbb{S}^{p-1}$, independent of $\xi$. The variable $\xi$ governs the tail behavior of the distribution.
We integrate this structure with the standard linear factor model. Let $\mathbf{y}_1, \dots, \mathbf{y}_n$ be independent and identically distributed (i.i.d.) random vectors in $\mathbb{R}^p$. The factor model posits:
where $B \in \mathbb{R}^{p \times m}$ is the deterministic factor loading matrix, $\mathbf{f}_t \in \mathbb{R}^m$ is the vector of $m$ latent common factors, and $\mathbf{e}_t \in \mathbb{R}^p$ is the vector of idiosyncratic errors. We assume $m$ is fixed and much smaller than $p$ and $n$. Standard identification conditions often include $\frac{1}{n} \sum_{t=1}^n \mathbf{f}_t \mathbf{f}_t' \overset{\mathbb{P}}{\rightarrow} I_m$ as $p \to \infty$ fan2018large.
To define the Elliptical Factor Model (EFM), we assume that $\mathbf{f}_t$ and $\mathbf{e}_t$ are uncorrelated and the joint distribution of factors and errors follows an elliptical structure. Specifically, if $(\mathbf{f}_t', \mathbf{e}_t')'$ is elliptically distributed with mean zero and covariance matrix $\operatorname{diag}(I_m, \Sigma_{err})$, where $\operatorname{diag}(I_m, \Sigma_{err})$ means that
Then $(\mathbf{f}_t', \mathbf{e}_t')' \sim EC_{m+p}(\mathbf{0}, \operatorname{diag}(I_m, \Sigma_{err}), \xi_t)$. Consequently, the observation vector $\mathbf{y}_t$ also follows an elliptical distribution: \[ \mathbf{y}_t=
\sim EC_p(\mathbf{0}, \Sigma, \xi_t), \] where
is the population covariance matrix of $\mathbf{y}_t$ (assuming $\mathbb{E}[\xi_t^2]$ is finite and normalized appropriately). The stochastic representation for the observations becomes:
where $\{\xi_t\}_{t=1}^n$ are i.i.d. copies of the radius variable $\xi$, and $\{\mathbf{u}_t\}_{t=1}^n$ are i.i.d. uniform on $\mathbb{S}^{p-1}$, independent of $\xi$. The data matrix is $Y = (\mathbf{y}_1, \dots, \mathbf{y}_n)= \Sigma^{1/2} U D$, where $U = (\mathbf{u}_1, \dots, \mathbf{u}_n)$ and $D = \operatorname{diag}(\xi_1, \dots, \xi_n)$.
The key feature allowing for heavy tails is the random radius $\xi_t$. We impose the following assumption on its distribution.
We assume the population covariance matrix $\Sigma$ exhibits a spiked structure, which is common in factor analysis chamberlain1982arbitrage, yu2024testing.
Given observations $\{\mathbf{y}_t\}_{t=1}^n$ from the EFM (ref), we form the sample covariance matrix $S$ and its companion $\mathcal{S}$:
Note that $S$ and $\mathcal{S}$ share the same non-zero eigenvalues, conventionally denoted $\lambda_1 \ge \lambda_2 \ge \dots \ge \lambda_{\min(p,n)}$. Standard factor analysis methods aim to estimate $m$ by examining these top eigenvalues. The underlying assumption is that $\lambda_1, \dots, \lambda_m$ are primarily influenced by the population spikes $\sigma_1, \dots, \sigma_m$, while $\lambda_{m+1}, \dots$ reflect the noise structure $\Sigma_{err}$.
However, under the EFM with heavy tails (Assumption (ref)), this distinction becomes blurred. The random radii $\xi_t^2$ can take extremely large values. When a large $\xi_t^2$ aligns with certain directions in $U$ and $\Sigma^{1/2}$, it can inflate some sample eigenvalues dramatically, even those notionally corresponding to the noise part $\Sigma_{err}$. These inflated noise eigenvalues can become outliers, exceeding the bulk and potentially mixing with, or even surpassing, the eigenvalues generated by the true factors $B\mathbf{f}_t$. We term these large eigenvalues not directly associated with the $m$ population spikes as “spurious factors".
Figure (ref) provides a toy example. Standard methods relying on gaps or thresholding (e.g., ahn2013eigenvalue, onatski2010determining) would likely identify two factors ($m=2$) instead of the true $m=1$, mistaking the eigenvalue at 5.5 for a real factor signal. This overestimation poses a significant problem in practice. Therefore, the primary goal of this paper is to develop a methodology to detect these spurious factors. Let $\mathsf{f}$ denote the number of spurious factors among the top eigenvalues (i.e., the number of large eigenvalues $\lambda_k$ that do not correspond to the true population spikes $\sigma_1, \dots, \sigma_m$). We aim to test the hypothesis:
Rejecting $\mathbf{H}_0$ implies the presence of spurious factors among the leading eigenvalues. This detection is crucial for accurately estimating the true number of factors $m$. Our proposed approach, detailed in Section (ref), leverages the differential fluctuation of real and spurious signals under magnification perturbations.
This section provides the theoretical foundation for our detection method by characterizing the asymptotic behavior of both real factor signals and spurious noise components present in the sample eigenvalues $\{\lambda_k\}$. Due to the spiked structure of $\Sigma$, we decompose $\Sigma=\Sigma_1+\Sigma_2$, where $\Sigma_1$ contains the $m$ largest components of $\Sigma$, and $\Sigma_2$ captures the remaining ones (also recall that $\Sigma$ is diagonal from Assumption (ref)). As a consequence, we can rewrite
In the sequel, we focus on the eigenvalues of $\mathcal{S} = Y'Y$. Note that $S$ and $\mathcal{S}$ share the same non-zero eigenvalues.
The top $m$ eigenvalues $\lambda_1, \dots, \lambda_m$ are primarily influenced by the population spikes $\sigma_1, \dots, \sigma_m$. However, their exact location is perturbed by the noise component $\Sigma_{err}$ and the heavy-tailed radii $D^2$.
Following standard techniques for spiked models, we first establish a first-order approximation.
To obtain a more precise description for the locations of $\lambda_i$, we employ the methodology from yu2024testing and introduce the random quantities $\theta_i,1\le i\le m$ to trace the asymptotic behavior of each $\lambda_i$. We define $\theta_i$ as the unique solution to
DXYZspiked showed that under the elliptical model, $\theta_i$ is a closer approximation to $\lambda_i$ compared with $\sigma_i$. In our study, to address the effect of $D^2$, we need the second order description $\zeta_i$ to be the unique solution to
The existence and uniqueness of $\theta_i, \zeta_i$ are guaranteed under our assumptions for large $p, n$.
Using these observations, we obtain a more precise characterization of $\lambda_i$. Let $\mathbf{e}_k$ be the $k$-th standard basis vector.
Then, related to the results in Lemma (ref), Theorem (ref) essentially implies the improved asymptotically rate $\lambda_i/\theta_i-1=\mathrm{O}_{\mathbb{P}}\big(\mathsf{T}/\sigma_i^2+(\sqrt{n}\sigma_i)^{-1}+\frac{1}{\sqrt{n}}+\sqrt{\frac{\mathsf{T}}{n}}\times\mathbbm{1}(\alpha\in(1,2])\big)$, which can be seen as an improvement of Theorem (ref) especially for $\alpha\in(2,+\infty)$ and exponential decay tail. On the other hand, we notice that the error rate in Theorem (ref) will get large when the randomness of $\mathbf{y}$ exhibits heavier distributions (in case of the index $\alpha$). Then, it suggests that the extreme heavy-tailness will strongly influence the convergence rates of the spiked eigenvalues.
Specifically, in the case where $\xi^2$ either has a polynomial decay tail with $\alpha\in(2,+\infty)$ or has an exponential decay tail in Assumption (ref), the following central limit theorem holds for real signals from Theorem (ref),
When $\xi^2$ exhibits serious heavy-tailed decay with $\alpha\in(1,2]$, the asymptotic normality in Theorem (ref) will not hold, but demands a larger scaling. In practice, it is enough to only consider the asymptotic mean and variance for $\lambda_i/\theta_i$. We have the following results for real factors under the serious heavy-tailed scenario.
We first prepare several notations. Recalling the decomposition of $\mathcal{S}$ in (ref), we denote the matching matrices removing the spike structure from $\mathcal{S}$ (or $S$) as
Denote $\lambda_{i,2}$ as the $i$-th largest eigenvalue of $S_2$ or $\mathcal{S}_2$. Denote the quantity $\mathsf{q}$ as
for some sufficiently small constant $\epsilon>0$.
The following lemma indicates that the leading eigenvalues of $S$, despite the first $m$ spiked eigenvalues, are close to the leading eigenvalues of $\mathcal{S}_2$,
Typically, the largest eigenvalues of $\mathcal{S}_2$ are strongly influenced by the ordered statistics in $D^2$, whose phenomena is captured in the following theorem,
As a result of Theorem (ref), we may easily find that $\mathbb{E}(\lambda_{k,2})=\xi_{(k)}^2(\Bar{\sigma}+\mathrm{o}_{\mathbb{P}}(1))$, given any realization of $D^2$.
In this section, we develop testing strategies for the problem (ref), guided by the theoretical framework established in Sections (ref) and (ref). We begin by rigorously formalizing the fluctuation magnification mechanism from the idea of resampling techniques for $S$, followed by a systematic algorithm to implement the procedure to identify spurious factors.
Suppose we observe the data matrix $Y$. We define a sequential magnifier data matrices $\widetilde{Y}_j,1\le j\le K$ from $Y$ for some large fixed integer $K>0$. Let
where $w_t\sim w, 1\le t\le n$ are some properly chosen i.i.d. random variables whose details will be specified below. For $1\le j\le K$, we independently construct from $\widetilde{Y}_j$ with
Let $\lambda_k^{(j)}$ be the $k$-th largest eigenvalue of $\widetilde{S}^{(j)}$ (or $\widetilde{S}^{(j)}$) and $\lambda_{k,2}^{(j)}$ the $k$-th largest eigenvalue of $\widetilde{\mathcal{S}}_2^{(j)}$. In the sequel, we use the notation $\mathbb{P}^*$ to denote the probability measure conditional on the sample $Y$ by $\mathbb{P}^*(\cdot)=\mathbb{P}(\cdot|\mathcal{F}_Y)$ where $\mathcal{F}_Y=\sigma(Y)$ is the sigma algebra generated from $\{\mathbf{y}_1,\dots,\mathbf{y}_n\}$. Then, we define $\overset{d^*}{\rightarrow}$, $\mathrm{o}_{\mathbb{P}^*}$ and $\mathrm{O}_{\mathbb{P}^*}$ accordingly from $\mathbb{P}^*$.
We first concern ourselves with the real factors. We consider the two situations in Theorem (ref) and Theorem (ref). Applying these two theorems to each $\widetilde{S}^{(j)}$ for $1\le j\le K$, we have the following proposition.
As a consequence, we have the following proposition.
Next, we turn to the spurious factors perturbed by some magnifier matrices. Lemma (ref) and Theorem (ref) indicate that for $m+1\le i\le \hat{o}$,
To simplify the discussion, we restrict ourselves to $i=m+1$ while the other cases can be handled similarly. The key idea here is to choose a suitable $w$ such that the fluctuation of $\lambda^{(j)}_{m+1}$ will not degenerate with $n$. A good candidate is that $w$ is uniformly distributed on the interval $[a,b]$ (say $w\in U[a,b]$) with $(a+b)/2=1$ and $a\ge b\xi^2_{(2)}/\xi^2_{(1)}$, which satisfies the condition in Proposition (ref). Then, it is easy to find that $(\xi^2w)_{(1)}$ is uniformly distributed on $[a\xi^2_{(1)},b\xi^2_{(1)}]$ with
Since $\lambda_{m+1}^{(j)}=(\Bar{\sigma}+\mathrm{o}_{\mathbb{P}}(1))\cdot(\xi^2w)_{(1)}$ are i.i.d. across each fluctuation magnification procedure, we can apply CLT to obtain the following results.
Inspired by the theoretical observations in Section (ref), we propose the following statistics.
Recap the test problem (ref) in Section (ref). We propose a novel algorithm that leverages a fluctuation magnification mechanism to detect potentially spurious signals and thereby determine the true number of common factors. To achieve this, we conduct a two-step testing procedure. In the first step, we apply the fluctuation magnification mechanism to identify spurious signals from the sample matrix $S$ without distinguishing outliers in advance. As a result, some bulk eigenvalues (non-outliers) may also be flagged as spurious. In the second step, we verify whether any bulk components were mistakenly identified as spurious by employing standard methods such as onatski2010determining, thus refining the selection of common factors.
1. First step detection. We now proceed to the first step. Under Assumption (ref), it suffices to focus on the largest spurious signal under the alternative hypothesis $\mathbf{H}_a$. As noted in Remark (ref), it is crucial to choose an appropriate distribution for $w$ to ensure the validity of {Proposition (ref)}. Typically, $w\in U[a,b]$ satisfies the following conditions. $$\mbox{(1)}~(a+b)/2=1;~~\mbox{(2)}~ a\ge b\xi^2_{(2)}/\xi^2_{(1)};~~\mbox{(3)}~ \sigma_m\gg\mathsf{T}b.$$
However, in practice, $\xi^2_{(1)},\xi^2_{(2)}$ are unobservable, and $\sigma_m$ is difficult to determine. To solve this issue, we approximate the ratio $\xi^2_{(2)}/\xi^2_{(1)}$ using $\lambda_{m+2}/\lambda_{m+1}$, and $\sigma_m/\mathsf{T}$ using $\lambda_{m}/\lambda_{m+1}$. The validity of the first approximation is ensured by Theorem (ref), while Lemma (ref) and Theorem (ref) support the second. The algorithm is summarized in Algorithm (ref), which gives preliminary estimations for the number of spurious factors and the location of the largest potential spurious signal.
2. Second step detection. First step detection aims to identify the most prominent spurious signals. The locations of the detected spurious signals provide useful information for separating potential real signals from noise contamination. When the data do not contain heavy-tailed noise, the first step procedure may incorrectly classify certain bulk eigenvalues as spurious. To mitigate this issue, we subsequently apply a refined factor‑number detection method exclusively to the corresponding potential real signals, which yields more accurate estimates and avoids over‑estimation. In this paper, we adopt the detection procedures of onatski2010determining and Dobriban as representative examples to obtain a more accurate estimate of the number of common factors in elliptical factor models that may exhibit heavy-tailed randomness. Algorithm (ref) proposed below builds on the first-step results together with the robustness properties of established factor selection methods, thereby effectively addressing the challenges posed by heavy-tailed distributions.
In this section, we conduct simulation studies to validate two main aspects: (1) The overestimation of the number of common factors in EFM when noise exhibits heavy-tailed randomness, confirming the existence of spurious signals. (2) The effectiveness of our fluctuation magnification approach (referred to as the detection algorithm in Section 5.2) in distinguishing spurious signals and correctly identifying the true number of common factors in such cases. Consequently, two scenarios of heavy-tailed data $\mathbf{y}$ from the model (ref) are analyzed: (I) A typical heavy-tailed case, where $\mathbf{y}$ follows a distribution with a polynomial decay tail satisfying $\alpha \in (2, +\infty)$. (II) A serious heavy-tailed case, where $\mathbf{y}$ follows a distribution with a polynomial decay tail satisfying $\alpha \in (1, 2]$. We focus on the strength of the heavy tail that arises from polynomial decay tails rather than exponential decay tails, as the former allows easier adjustment of the strength of the heavy tail via the parameter $\alpha$.
We first establish the settings for the typical heavy-tailed case for $\mathbf{y}$ in (ref), characterized by a polynomial decay tail with $\alpha \in (2, +\infty)$. To represent various scenarios of heavy-tailed data, we employ multivariate t-distributions with degrees of freedom set at 4.3, 4.8, and 5.3 (i.e., $t(4.3)$, $t(4.8)$, and $t(5.3)$, respectively), all of which satisfy the condition $\alpha > 2$. We consider a high-dimensional framework with $p=1000$ (number of variables) and $n=1000$ (number of observations). Our analysis focuses on two population covariance matrices $\Sigma$:
Here, the larger diagonal components represent the distinct common factors. It is clear that $\Sigma_{I}$ encompasses two common factors induced by $\{16, 8\}$, while $\Sigma_{II}$ encompasses three common factors induced by $\{24, 16, 8\}$.
In our comparative study, we consider two baseline factor-number detection methods: the ED estimator of onatski2010determining (denoted as “Onta") and the DDPA+ procedure of Dobriban (denoted as “DDPA+"). Their versions enhanced by the Fluctuation Magnification technique proposed in Section (ref) are referred to as “Onta with MF” and “DDPA+ with MF”, respectively. In our simulations, to reduce computational load, we set the magnifiers $w_{i} \sim U[0.1, 1.9]$ in Algorithm (ref) and $K=1000$ in Algorithm (ref), which yielded sufficiently good results. Detailed experimental results are presented in Figure (ref), and Tables (ref) and (ref).
Specifically, Figure (ref) illustrates the performance of the methods for a representative case (Case (i)) using the “Onta" estimator under $t(4.3)$ as an example. Panel (a) displays the first 50 eigenvalues of the sample covariance matrix. Panel (b) shows the estimated factor number $k$ obtained from the standard “Onta" method. Panel (c) demonstrates the identification of spurious signals via our proposed statistics $\mathbb{T}_i$ from Algorithm (ref); the statistics corresponding to spurious signals are markedly larger than those associated with real signals. These observations highlight the effectiveness of Algorithm (ref) in detecting spurious signals.
To compare the performance of Onta and DDPA+ with and without MF technique under Case (i) and Case (ii), we conducted 500 repeated experiments for each setting. The effective sample size (ESS) reported in Tables (ref) and (ref) corresponds to the number of valid replications retained after filtering out cases where the eigenvalue ratio between the last real and first spurious signal is too close. Specifically, for Case (i) we exclude replications with \(\lambda_2/\lambda_3 < 1.1\) to avoid the second and third eigenvalues being too close; for Case (ii) we exclude those with \(\lambda_3/\lambda_4 < 1.1\) to avoid the third and fourth eigenvalues being too close. This filtering ensures a clear distinction between spurious and real signals, preventing ambiguity that could otherwise distort the comparison. The tables present the underestimation (Under), overestimation (Over), and correct estimation (Right) rates for three \(t\)-distributions with varying degrees of freedom.
From Figure (ref), and Tables (ref) and (ref), we can draw the following conclusions.
Next, we evaluate the detection behavior under serious heavy-tailed settings for $\mathbf{y}$ described in (ref), characterized by polynomially decayed tails with $\alpha \in (1, 2]$. The remaining settings align with those in the Typical Heavy-tailed Case; only the differences are described here. We use multivariate t-distributions with degrees of freedom 2.5, 3.0, and 3.5 (denoted as $t(2.5)$, $t(3.0)$, and $t(3.5)$, respectively) to model the polynomial decay tail with $\alpha \in (1, 2]$. Additionally, the covariance matrix $\Sigma$ is configured with larger spiked eigenvalues as follows:
Figure (ref) illustrates the performance of the methods for a representative case (Case (iii)) using the “Onta" estimator under $t(2.5)$ as an example.
Using the same evaluation framework as in the typical heavy-tailed case, we assess the underestimation, overestimation, and correct estimation rates for three $t$-distributions with varying degrees of freedom.
From Figure (ref) and Tables (ref) and (ref), we observe that the standalone DDPA+ method essentially fails under the Serious Heavy-tailed Case. Integrating our MF tool successfully restores its performance, yielding relatively accurate estimations. Moreover, the improvement achieved by employing the MF tools is even more pronounced in this case, highlighting their enhanced effectiveness under more challenging heavy-tailed settings.
In this section, the proposed methods are applied to the real-world FRED-MD dataset, which was previously studied in XIA2017235 and YU2019104543. This dataset, introduced in McCracken01102016, is publicly available on the St. Louis Fed's website: https://www.stlouisfed.org/research/economists/mccracken/fred-databases. It consists of monthly data for 128 macroeconomic variables. For this analysis, we focused on the 786 observations covering the period from March 1959 to June 2024. The raw data is non-stationary and contains missing values. The first step is to transform the series into a stationary one using the code provided on the aforementioned website. After this preprocessing, the first two observations were eliminated due to the application of differencing operators, resulting in a $784 \times 128$ panel. The website also offers code to replace outliers with “reasonable" values, but we chose to omit this step, as extreme observations are inherent in data drawn from heavy-tailed distributions. Missing values were imputed with the sample mean of the corresponding non-missing entries in the same column.
We first analyzed the entire panel to determine the number of common factors. Our detection algorithm estimates $\widehat{r} = 3$, consistent with the result in XIA2017235 and slightly smaller than the estimate $\widehat{r} = 4$ proposed by YU2019104543. We consider $\widehat{r} = 3$ more reasonable than $\widehat{r} = 4$, though both studies retain outliers and recognize heavy-tailed distributions (as noted by YU2019104543). This conclusion stems from temporal dynamics analysis: rolling 4-year windows (Feb 1992-Feb 2024) reveal that spurious factors intermittently emerge in certain subperiods. These spurious factors cause an estimate of $\widehat{r} = 4$ to potentially overestimate the true factor number. Thus $\widehat{r} = 3$ provides a more robust full-sample representation.
The choice to begin this time-varying analysis in February 1992 is necessitated by data availability: Prior to this date, datasets for ACOGNO, AMDMNOx, ANDENOx, and AMDMUOx were completely missing. This extensive data gap before 1992 presents a significant challenge for our 48-month window analysis: Filling such large amounts of missing data within these relatively small observation windows would inevitably introduce substantial inaccuracies. Table (ref) shows $\mathbb{G}_i$ (“On" estimation), $i=1,\ldots,10$ of every 48 months data matrices.
We set $\mathbb{G}_i>9$ as the threshold for testing the factor, and derive that the periods “1992-1996", “1996-2000", and “2012-2016" possess two common factors. However, the period “2000-2004" may possess either two or five factors, “2004-2008" may possess either two or six factors, “2008-2012" may possess either two or nine factors, “2016-2020" may possess either two or six factors, and “2020-2024" may possess either two or nine factors. Now we conduct a fluctuation magnification algorithm on these uncertain periods. Figure (ref) illustrates the variance after fluctuation magnification for the period 2000-2004. The repetition in the magnification algorithm for variance calculation is set to $K=200$.
From Figure (ref), we observe no significant fluctuations (increases) in variance at \(i=6\), indicating that there is no substantial change in the factors (i.e., no shift from spiked to bulk). In comparison, the variance shows a slight increase from \(i=2\) to \(i=3\). This suggests that the variance remains relatively stable (the variance remains close to 0.02), making the judgment of \(i=2\) more reasonable. This is not unexpected, as fluctuation magnification for variance primarily serves as an auxiliary tool to confirm the choice of \(i=2\), and the observed stability further supports this decision. However, when we examine the period “2004-2008" in Figure (ref), we observe an unusual pattern that diverges from the previous trends.
From Figure (ref), we observe significant fluctuations (from 0.02 to 0.28) in variance at \(i=6\), while the fluctuations at other values remain relatively small. This suggests that \(i=6\) may represent a spurious signal (noting that spikes larger than the first spurious signal are real signals), and the true number of factors is likely five. This outcome is not surprising, as the large-scale financial crisis of 2008 led to substantial disruptions in the global economy. Such economic shocks can result in changes in the underlying structure of the data, influencing a number of factors. The following Figure (ref) shows the change in the number of factors from 1992 to 2024.
In this section, we outline the main ideas of the proof and full technical details are deferred to the Appendix. Our proof proceeds in three steps.
This paper has confronted the critical “spurious factor dilemma" in EFMs, where heavy-tailed randomness, prevalent in economic and financial data, can generate noise-induced eigenvalues that masquerade as real factors. We introduce a novel theory-based fluctuation magnification algorithm, which uniquely leverages the differential stability of real versus spurious factors under targeted perturbations: real factor signals exhibit resilience, while spurious ones betray their noisy origins through amplified volatility. This advancement significantly enhances robust factor analysis, offering a more reliable foundation for modeling and forecasting in high-dimensional, heavy-tailed distribution, thereby improving the fidelity of economic and financial decision-making.
While classical RMT often relies on assumptions of finite higher-order moments, which can be challenged by heavy-tailed distributions, this paper indicates that the conceptual framework and analytical power of RMT remain remarkably insightful and adaptable for understanding complex systems. Indeed, ongoing research continues to extend RMT's reach, developing new results and specialized techniques that successfully characterize the spectral properties of matrices with heavy-tailed entries. These advancements, by providing a deeper understanding of how eigenvalues and eigenvectors behave under non-standard conditions, are proving instrumental in developing robust methodologies capable of distinguishing real signals from noise and making reliable inferences from heavy-tailed data, thereby underscoring RMT's enduring effectiveness and its evolving role in the modern high-dimensional world.