EconBase
← All papers

Leave-out estimation of variance components

Patrick Kline, Raffaele Saggio, Mikkel Sølvsten

arXiv 5 Jun 2018 · Econometrics · publishedEconometrica (2020) · 180 citations (OpenAlex)

arXiv:1806.01494 · PDF · DOI · OpenAlex · Extracted main text

Abstract

We propose leave-out estimators of quadratic forms designed for the study of linear models with unrestricted heteroscedasticity. Applications include analysis of variance and tests of linear restrictions in models with many regressors. An approximation algorithm is provided that enables accurate computation of the estimator in very large datasets. We study the large sample properties of our estimator allowing the number of regressors to grow in proportion to the number of observations. Consistency is established in a variety of settings where plug-in methods and estimators predicated on homoscedasticity exhibit first-order biases. For quadratic forms of increasing rank, the limiting distribution can be represented by a linear combination of normal and non-central $\chi^2$ random variables, with normality ensuing under strong identification. Standard error estimators are proposed that enable tests of linear restrictions and the construction of uniformly valid confidence intervals for quadratic forms of interest. We find in Italian social security records that leave-out estimates of a variance decomposition in a two-way fixed effects model of wage determination yield substantially different conclusions regarding the relative contribution of workers, firms, and worker-firm sorting to wage inequality than conventional methods. Monte Carlo exercises corroborate the accuracy of our asymptotic approximations, with clear evidence of non-normality emerging when worker mobility between blocks of firms is limited.

Citation extraction

83
references
169
in-text mentions
83
distinct cited
1
self-citations
22,101
main-text words

appendix boundary found by appendix_command · 47% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Cattaneo, M. D., M. Jansson, and W. K. Newey (2018) Inference in linear regression models with many covariates and heteroscedasticity1.00083100%
2Abowd, J. M., F. Kramarz, and D. N. Margolis (1999) High wage workers and high wage firms1.00073100%
3Andrews, M. J., L. Gill, T. Schank, and R. Upward (2008) High wage workers and low wage firms: negative assortative matching or limited mobility bias?1.00063100%
4Anatolyev, S (2012) Inference in regression models with many regressors1.00053100%
5Newey, W. K. and J. R. Robins (2018) Cross-fitting and fast remainder rates for semiparametric estimation0.92843100%
6Searle, S. R., G. Casella, and C. E. McCulloch (2009) Variance components, Volume 3910.92843100%
7Card, D., J. Heining, and P. Kline (2013) Workplace heterogeneity and the rise of west german wage inequality0.8307657%
8Achlioptas, D (2003) Database-friendly random projections: Johnson-lindenstrauss with binary coins0.7373367%
9Dhaene, G. and K. Jochmans (2015) Split-panel jackknife estimation of fixed-effect models0.7373367%
10Hahn, J. and W. Newey (2004) Jackknife and analytical bias reduction for nonlinear panel models0.7373367%

Showing the top 10 of 83 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Cluster-Robust Inference for Quadratic Forms1.000124
2Adjustments with Many Regressors under Covariate-Adaptive Randomizations1.00053
3Testing Many Restrictions Under Heteroskedasticity0.937174
4A Dimension-Agnostic Bootstrap Anderson-Rubin Test For Instrumental Variable Regressions0.92843
5Branching Fixed Effects: A Proposal for Communicating Uncertainty0.87452
6Finite-Population Inference for Heterogeneity in Many-Group Synthetic Difference-in-Differences0.84343
7Assumption-lean Falsification Tests of Rate Double-Robustness of Double-Machine-Learning Estimators0.84333
8Automatic Inference for Value-Added Regressions0.84333
9Ridge Estimation of High Dimensional Two-Way Fixed Effect Regression0.81142
10Robust Inference on Infinite and Growing Dimensional Time Series Regression0.73732