EconBase
← All papers

Divide-and-Conquer: A Distributed Hierarchical Factor Approach to Modeling Large-Scale Time Series Data

Zhaoxing Gao, Ruey S. Tsay

arXiv 26 Mar 2021 · Statistics — Methodology · publishedJournal of the American Statistical Association (2022) · 13 citations (OpenAlex)

arXiv:2103.14626 · PDF · DOI · OpenAlex · Extracted main text

Abstract

This paper proposes a hierarchical approximate-factor approach to analyzing high-dimensional, large-scale heterogeneous time series data using distributed computing. The new method employs a multiple-fold dimension reduction procedure using Principal Component Analysis (PCA) and shows great promises for modeling large-scale data that cannot be stored nor analyzed by a single machine. Each computer at the basic level performs a PCA to extract common factors among the time series assigned to it and transfers those factors to one and only one node of the second level. Each 2nd-level computer collects the common factors from its subordinates and performs another PCA to select the 2nd-level common factors. This process is repeated until the central server is reached, which collects common factors from its direct subordinates and performs a final PCA to select the global common factors. The noise terms of the 2nd-level approximate factor model are the unique common factors of the 1st-level clusters. We focus on the case of 2 levels in our theoretical derivations, but the idea can easily be generalized to any finite number of hierarchies. We discuss some clustering methods when the group memberships are unknown and introduce a new diffusion index approach to forecasting. We further extend the analysis to unit-root nonstationary time series. Asymptotic properties of the proposed method are derived for the diverging dimension of the data in each computing unit and the sample size $T$. We use both simulated data and real examples to assess the performance of the proposed method in finite samples, and compare our method with the commonly used ones in the literature concerning the forecastability of extracted factors.

Citation extraction

45
references
122
in-text mentions
45
distinct cited
6
self-citations
15,211
main-text words

appendix boundary found by appendix_titled_section at “Appendix: Proofs” · 71% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lam, C. and Yao, Q (2012) Factor modeling for high-dimensional time series: inference for the number of factors1.00083100%
2Ahn, S. C., and Horenstein, A. R (2013) Eigenvalue ratio test for the number of factors1.00053100%
3Bai, J. and Ng, S (2002) Determining the number of factors in approximate factor models0.94118583%
4Bai J (2003) Inferential theory for factor models of large dimensions0.9285480%
5Gao, Z. and Tsay, R. S (2020) Modeling high-dimensional unit-root time series self0.92843100%
6Stock, J. H., and Watson, M. W (2002) Macroeconomic forecasting using diffusion indexes0.874102100%
7Fan, J., Wang, D., Wang, K., and Zhu, Z (2019) Distributed estimation of principal eigenspaces0.87462100%
8Fan, J., Liao, Y., and Mincheva, M. (2013). Large covariance estimat… Journal of the Royal Statistical Society, Series B, 75(4), 603–6800.8307357%
9Alonso, A. M., Galeano, P., and Peña, D (2020) A robust procedure to build dynamic factor models with cluster structure0.73732100%
10Gao, Z. and Tsay, R. S (2020) Modeling high-dimensional time series: a factor model with dynamically dependent factors and diverging eigenvalues self0.71411336%

Showing the top 10 of 45 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Sparse Asymptotic PCA: Identifying Sparse Latent Factors Across Time Horizon in High-Dimensional Time Series0.84333
2Determination of the effective cointegration rank in high-dimensional time-series predictive regressions0.40511
3Supervised Dynamic PCA: Linear Dynamic Forecasting with Many Predictors0.40511
4High-Dimensional Matrix-Variate Diffusion Index Models for Time Series Forecasting0.40511
5High-Dimensional Spatial Arbitrage Pricing Theory with Heterogeneous Interactions0.40511