Liyuan Cui, Guanhao Feng, Yuefeng Han, Jiayan Li
arXiv 29 Dec 2025 · Statistics — Methodology
arXiv:2512.23567 · PDF · DOI · OpenAlex · Extracted main text
We tackle the challenge of estimating grouping structures and factor loadings in asset pricing models, where traditional regressions struggle due to sparse data and high noise. Existing approaches, such as those using fused penalties and multi-task learning, often enforce coefficient homogeneity across cross-sectional units, reducing flexibility. Clustering methods (e.g., spectral clustering, Lloyd's algorithm) achieve consistent recovery under specific conditions but typically rely on a single data source. To address these limitations, we introduce the Panel Coupled Matrix-Tensor Clustering (PMTC) model, which simultaneously leverages a characteristics tensor and a return matrix to identify latent asset groups. By integrating these data sources, we develop computationally efficient tensor clustering algorithms that enhance both clustering accuracy and factor loading estimation. Simulations demonstrate that our methods outperform single-source alternatives in clustering accuracy and coefficient estimation, particularly under moderate signal-to-noise conditions. Empirical application to U.S. equities demonstrates the practical value of PMTC, yielding higher out-of-sample total $R^2$ and economically interpretable variation in factor exposures.
appendix boundary found by appendix_command · 90% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Zhang, Anderson Y and Zhou, Harrison Y (2024) Leave-one-out singular subspace perturbation analysis for spectral clustering | 1.000 | 5 | 4 | 100% |
| 2 | Zhang, Anru and Xia, Dong (2018) Tensor SVD: Statistical and computational limits | 1.000 | 5 | 3 | 100% |
| 3 | Han, Rungang and Luo, Yuetian and Wang, Miaoyan and Zhang, Anru R (2022) Exact clustering in tensor block model: Statistical optimality and computational limit | 0.935 | 11 | 5 | 82% |
| 4 | Gao, Chao and Zhang, Anderson Y (2022) Iterative algorithm for discrete structure recovery | 0.874 | 5 | 2 | 100% |
| 5 | Fama, Eugene F. and French, Kenneth R (1992) The Cross-Section of Expected Stock Returns | 0.644 | 2 | 2 | 100% |
| 6 | Cong, Lin William and Feng, Guanhao and He, Jingyu and Li, Junye (2023) Sparse modeling under grouped heterogeneity with an application to asset pricing self | 0.644 | 2 | 2 | 100% |
| 7 | Cong, Lin William and Feng, Guanhao and He, Jingyu and He, Xin (2025) Growing the efficient frontier on panel trees self | 0.644 | 2 | 2 | 100% |
| 8 | Han, Yuefeng and Chen, Rong and Yang, Dan and Zhang, Cun-Hui (2024) Tensor factor model estimation by iterative projection self | 0.644 | 2 | 2 | 100% |
| 9 | Patton, Andrew J and Weller, Brian M (2022) Risk price variation: The missing half of empirical asset pricing | 0.644 | 2 | 2 | 100% |
| 10 | Su, Liangjun and Shi, Zhentao and Phillips, Peter CB (2016) Identifying latent structures in panel data | 0.644 | 2 | 2 | 100% |
Showing the top 10 of 49 scored citations.