arXiv 21 Aug 2025 · Statistics — Methodology
arXiv:2508.15675 · PDF · DOI · OpenAlex · Extracted main text
Principal component analysis (PCA) is arguably the most widely used approach for large-dimensional factor analysis. While it is effective when the factors are sufficiently strong, it can be inconsistent when the factors are weak and/or the noise has complex dependence structure. We argue that the inconsistency often stems from bias and introduce a general approach to restore consistency. Specifically, we propose a general weighting scheme for PCA and show that with a suitable choice of weighting matrices, it is possible to deduce consistent and asymptotic normal estimators under much weaker conditions than the usual PCA. While the optimal weight matrix may require knowledge about the factors and covariance of the idiosyncratic noise that are not known a priori, we develop an agnostic approach to adaptively choose from a large class of weighting matrices that can be viewed as PCA for weighted linear combinations of auto-covariances among the observations. Theoretical and numerical results demonstrate the merits of our methodology over the usual PCA and other recently developed techniques for large-dimensional approximate factor models.
appendix boundary found by appendix_command · 21% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Jianqing Fan, Yuling Yan, and Yuheng Zheng (2024) When can weak latent factors be statistically inferred? | 0.961 | 9 | 4 | 89% |
| 2 | Jungjun Choi and Ming Yuan (2024) High dimensional factor analysis with weak factors self | 0.811 | 4 | 2 | 100% |
| 3 | Joshua Agterberg, Zachary Lubberts, and Carey E Priebe (2022) Entrywise estimation of singular vectors of low-rank matrices with heteroskedasticity and dependence | 0.737 | 5 | 4 | 40% |
| 4 | Yuling Yan, Yuxin Chen, and Jianqing Fan (2021) Inference for heteroskedastic pca with missing data | 0.737 | 3 | 3 | 67% |
| 5 | Clifford Lam, Qiwei Yao, and Neil Bathia (2011) Estimation of latent factors for high-dimensional time series | 0.737 | 3 | 2 | 100% |
| 6 | Clifford Lam and Qiwei Yao (2012) Factor modeling for high-dimensional time series: inference for the number of factors | 0.737 | 3 | 2 | 100% |
| 7 | Alexei Onatski (2012) Asymptotics of the principal components estimator of large factor models with weakly influential factors | 0.644 | 2 | 2 | 100% |
| 8 | Dong Xia (2021) Normal approximation and confidence region of singular subspaces | 0.585 | 5 | 3 | 20% |
| 9 | Sainan Jin, Ke Miao, and Liangjun Su (2021) On factor models with random missing: Em estimation, inference, and cross validation | 0.585 | 3 | 3 | 33% |
| 10 | Hang Zhou, Dongyi Wei, and Fang Yao (2022) Theory of functional principal component analysis for discretely observed data | 0.585 | 3 | 1 | 100% |
Showing the top 10 of 52 scored citations.