EconBase
← All papers

Factor-Based Imputation of Missing Values and Covariances in Panel Data of Large Dimensions

Ercument Cahan, Jushan Bai, Serena Ng

arXiv 4 Mar 2021 · Econometrics · publishedJournal of Econometrics (2022) · 8 citations (OpenAlex)

arXiv:2103.03045 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Economists are blessed with a wealth of data for analysis, but more often than not, values in some entries of the data matrix are missing. Various methods have been proposed to handle missing observations in a few variables. We exploit the factor structure in panel data of large dimensions. Our \textsc{tall-project} algorithm first estimates the factors from a \textsc{tall} block in which data for all rows are observed, and projections of variable specific length are then used to estimate the factor loadings. A missing value is imputed as the estimated common component which we show is consistent and asymptotically normal without further iteration. Implications for using imputed data in factor augmented regressions are then discussed. To compensate for the downward bias in covariance matrices created by an omitted noise when the data point is not observed, we overlay the imputed data with re-sampled idiosyncratic residuals many times and use the average of the covariances to estimate the parameters of interest. Simulations show that the procedures have desirable finite sample properties.

Citation extraction

23
references
50
in-text mentions
23
distinct cited
1
self-citations
11,472
main-text words

appendix boundary found by appendix_titled_section at “Appendix” · 77% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Jin, Miao and Su (2021) On Factor Models with Random Missing: EM Estimation, Inference, and Cross Validation, Journal of Econometrics 222:1, Part C, 745…0.87452100%
2Xiong and Pelger (2019) Large Dimensional Latent Factor Modeling with Missing Observations and Applications to Causal Inference, SSRN Working Paper 3465…0.87452100%
3Bai and Ng (2021) Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data, Journal of the American Statistical Association0.77313546%
4Bai and Ng (2006) Confidence Intervals for Diffusion Index Forecasts and Inference with Factor-Augmented Regressions, Econometrica 74:4, 1133–11500.73732100%
5Bai (2003) Inferential Theory for Factor Models of Large Dimensions, Econometrica 71:1, 135–172 self0.64441100%
6Stock and Watson (2016) Factor Models and Structural Vector Autoregressions in Macroeconomics, in J. B0.64422100%
7Stock and Watson (1998) Diffusion Indexes, NBER Working Paper 67020.51121100%
8Athey, Bayati, Doudchenko, Imbens and Khosravi (2018) Matrx Completion Methods for Causal Panel Data Methods, arXiv:1710.10251v20.40511100%
9Bai and Ng (2002) Determining the Number of Factors in Approximate Factor Models, Econometrica 70:1, 191–2210.40511100%
10Banbura and Modugno (2014) Maximum Likelihood Estimation of Factor Models on Datasets with Arbitrary Pattern of Missing Data, Journal of Applied Econometri…0.40511100%

Showing the top 10 of 23 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Target PCA: Transfer Learning Large Dimensional Panel Data0.81142
2Large Dimensional Latent Factor Modeling with Missing Observations and Applications to Causal Inference0.73732
32401.136650.73732
4MATRIX COMPLETION, COUNTERFACTUALS, AND FACTOR ANALYSIS OF MISSING DATA0.64422
5Matrix Completion When Missing Is Not at Random and Its Applications in Causal Panel Data Models0.51121
6Inference for Regression with Variables Generated by AI or Machine Learning0.40511
7Estimation of large approximate dynamic matrix factor models based on the EM algorithm and Kalman filtering0.40511
8Causal Forecasting in Panel Data: A Two-Way Synthetic Forecasting Approach0.40511
90.5cmLow-Rank Estimation of Nonlinear Panel Data Models0.00011