arXiv 23 Oct 2019 · Mathematics — Statistics Theory · 2 citations (OpenAlex)
arXiv:1910.10382 · PDF · DOI · OpenAlex · Extracted main text
In this paper, we consider the problem of learning models with a latent factor structure. The focus is to find what is possible and what is impossible if the usual strong factor condition is not imposed. We study the minimax rate and adaptivity issues in two problems: pure factor models and panel regression with interactive fixed effects. For pure factor models, if the number of factors is known, we develop adaptive estimation and inference procedures that attain the minimax rate. However, when the number of factors is not specified a priori, we show that there is a tradeoff between validity and efficiency: any confidence interval that has uniform validity for arbitrary factor strength has to be conservative; in particular its width is bounded away from zero even when the factors are strong. Conversely, any data-driven confidence interval that does not require as an input the exact number of factors (including weak ones) and has shrinking width under strong factors does not have uniform coverage and the worst-case coverage probability is at most 1/2. For panel regressions with interactive fixed effects, the tradeoff is much better. We find that the minimax rate for learning the regression coefficient does not depend on the factor strength and propose a simple estimator that achieves this rate. However, when weak factors are allowed, uncertainty in the number of factors can cause a great loss of efficiency although the rate is not affected. In most cases, we find that the strong factor condition (and/or exact knowledge of number of factors) improves efficiency, but this condition needs to be imposed by faith and cannot be verified in data for inference purposes.
appendix boundary found by appendix_command · 51% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Moon, H. R. and Weidner, M (2015) Linear regression for panel with unknown number of factors as interactive fixed effects | 0.874 | 7 | 2 | 100% |
| 2 | Bai, J (2003) Inferential theory for factor models of large dimensions | 0.811 | 4 | 2 | 100% |
| 3 | Candès, E. J. and Plan, Y (2011) Tight oracle inequalities for low-rank matrix recovery from a minimal number of noisy random measurements | 0.737 | 3 | 3 | 67% |
| 4 | Rohde, A. and Tsybakov, A. B (2011) Estimation of high-dimensional low-rank matrices | 0.737 | 3 | 3 | 67% |
| 5 | Bai, J (2009) Panel data models with interactive fixed effects | 0.693 | 5 | 1 | 100% |
| 6 | Bai, J. and Ng, S (2019) Matrix completion, counterfactuals, and factor analysis of missing data | 0.644 | 4 | 1 | 100% |
| 7 | Vershynin, R (2010) Introduction to the non-asymptotic analysis of random matrices | 0.511 | 3 | 2 | 33% |
| 8 | Abadie, A., Diamond, A., and Hainmueller, J (2015) Comparative politics and the synthetic control method | 0.511 | 2 | 1 | 100% |
| 9 | Chernozhukov, V., Newey, W., Robins, J., and Singh, R (2018) Double/de-biased machine learning of global and local parameters using regularized riesz representers | 0.511 | 2 | 1 | 100% |
| 10 | Onatski, A (2012) Asymptotics of the principal components estimator of large factor models with weakly influential factors | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 55 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.
| Citing paper | Intensity | Mentions | Sections | |
|---|---|---|---|---|
| 1 | Robust Estimation and Inference in Panels with Interactive Fixed Effects | 0.909 | 8 | 3 |