Jiang Hu, Jiahui Xie, Yangchun Zhang, Wang Zhou
arXiv 5 Jun 2025 · Statistics — Methodology
arXiv:2506.05116 · PDF · DOI · OpenAlex · Extracted main text
Factor models are essential tools for analyzing high-dimensional data, particularly in economics and finance. However, standard methods for determining the number of factors often overestimate the true number when data exhibit heavy-tailed randomness, misinterpreting noise-induced outliers as genuine factors. This paper addresses this challenge within the framework of Elliptical Factor Models (EFM), which accommodate both heavy tails and potential non-linear dependencies common in real-world data. We demonstrate theoretically and empirically that heavy-tailed noise generates spurious eigenvalues that mimic true factor signals. To distinguish these, we propose a novel methodology based on a fluctuation magnification algorithm. We show that under magnifying perturbations, the eigenvalues associated with real factors exhibit significantly less fluctuation (stabilizing asymptotically) compared to spurious eigenvalues arising from heavy-tailed effects. This differential behavior allows the identification and detection of the true and spurious factors. We develop a formal testing procedure based on this principle and apply it to the problem of accurately selecting the number of common factors in heavy-tailed EFMs. Simulation studies and real data analysis confirm the effectiveness of our approach compared to existing methods, particularly in scenarios with pronounced heavy-tailedness.
appendix boundary found by appendix_command · 33% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | A. Onatski (2010) Determining the number of factors from empirical distribution of eigenvalues | 0.969 | 11 | 5 | 91% |
| 2 | L. Yu, P. Zhao, and W. Zhou (2025) Testing the number of common factors by bootstrapped sample covariance matrix in high-dimensional factor models | 0.928 | 5 | 4 | 80% |
| 3 | E. Dobriban and A. B. Owen (2018) Deterministic parallel analysis: An improved method for selecting factors and principal components | 0.874 | 6 | 2 | 100% |
| 4 | B. H. Baltagi, C. Kao, and F. Wang (2017) Identification and estimation of a large factor model with structural instability | 0.737 | 3 | 2 | 100% |
| 5 | G. Chamberlain and M. Rothschild (1983) Arbitrage, factor structure, and mean-variance analysis on large asset markets | 0.737 | 3 | 2 | 100% |
| 6 | J. Fan, H. Liu, and W. Wang (2018) Large covariance estimation through elliptical factor models | 0.737 | 3 | 2 | 100% |
| 7 | S. C. Ahn and A. R. Horenstein (2013) Eigenvalue ratio test for the number of factors | 0.644 | 2 | 2 | 100% |
| 8 | T. T. Cai, X. Han, and G. Pan (2020) Limiting laws for divergent spiked eigenvalues and largest nonspiked eigenvalue of sample covariance matrices | 0.644 | 2 | 2 | 100% |
| 9 | L. Yu, Y. He, and X. Zhang (2019) Robust factor number specification for large-dimensional elliptical factor model | 0.585 | 3 | 1 | 100% |
| 10 | Z. T. Ke, Y. Ma, and X. Lin (2023) Estimation of the number of spiked eigenvalues in a covariance matrix by bulk eigenvalue matching analysis | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 38 scored citations.