EconBase
← All papers

How well can we learn large factor models without assuming strong factors?

Yinchu Zhu

arXiv 23 Oct 2019 · Mathematics — Statistics Theory · 2 citations (OpenAlex)

arXiv:1910.10382 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In this paper, we consider the problem of learning models with a latent factor structure. The focus is to find what is possible and what is impossible if the usual strong factor condition is not imposed. We study the minimax rate and adaptivity issues in two problems: pure factor models and panel regression with interactive fixed effects. For pure factor models, if the number of factors is known, we develop adaptive estimation and inference procedures that attain the minimax rate. However, when the number of factors is not specified a priori, we show that there is a tradeoff between validity and efficiency: any confidence interval that has uniform validity for arbitrary factor strength has to be conservative; in particular its width is bounded away from zero even when the factors are strong. Conversely, any data-driven confidence interval that does not require as an input the exact number of factors (including weak ones) and has shrinking width under strong factors does not have uniform coverage and the worst-case coverage probability is at most 1/2. For panel regressions with interactive fixed effects, the tradeoff is much better. We find that the minimax rate for learning the regression coefficient does not depend on the factor strength and propose a simple estimator that achieves this rate. However, when weak factors are allowed, uncertainty in the number of factors can cause a great loss of efficiency although the rate is not affected. In most cases, we find that the strong factor condition (and/or exact knowledge of number of factors) improves efficiency, but this condition needs to be imposed by faith and cannot be verified in data for inference purposes.

Citation extraction

55
references
81
in-text mentions
55
distinct cited
4
self-citations
10,880
main-text words

appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Moon, H. R. and Weidner, M (2015) Linear regression for panel with unknown number of factors as interactive fixed effects0.87472100%
2Bai, J (2003) Inferential theory for factor models of large dimensions0.81142100%
3Candès, E. J. and Plan, Y (2011) Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements0.7373367%
4Rohde, A. and Tsybakov, A. B (2011) Estimation of high-dimensional low-rank matrices0.7373367%
5Bai, J (2009) Panel data models with interactive fixed effects0.69351100%
6Bai, J. and Ng, S (2019) Matrix completion, counterfactuals, and factor analysis of missing data0.64441100%
7Vershynin, R (2010) Introduction to the non-asymptotic analysis of random matrices0.5113233%
8Abadie, A., Diamond, A., and Hainmueller, J (2015) Comparative politics and the synthetic control method0.51121100%
9Chernozhukov, V., Newey, W., Robins, J., and Singh, R (2018) Double/de-biased machine learning of global and local parameters using regularized riesz representers0.51121100%
10Onatski, A (2012) Asymptotics of the principal components estimator of large factor models with weakly influential factors0.51121100%

Showing the top 10 of 55 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Robust Estimation and Inference in Panels with Interactive Fixed Effects0.90983