EconBase
← All papers

Panel Data with Unknown Clusters

Yong Cai

arXiv 10 Jun 2021 · Econometrics

arXiv:2106.05503 · PDF · DOI · OpenAlex · Extracted main text

Abstract

Clustered standard errors and approximate randomization tests are popular inference methods that allow for dependence within observations. However, they require researchers to know the cluster structure ex ante. We propose a procedure to help researchers discover clusters in panel data. Our method is based on thresholding an estimated long-run variance-covariance matrix and requires the panel to be large in the time dimension, but imposes no lower bound on the number of units. We show that our procedure recovers the true clusters with high probability with no assumptions on the cluster structure. The estimated clusters are independently of interest, but they can also be used in the approximate randomization tests or with conventional cluster-robust covariance estimators. The resulting procedures control size and have good power.

Citation extraction

17
references
43
in-text mentions
17
distinct cited
2
self-citations
8,235
main-text words

appendix boundary found by appendix_command · 64% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Bai, J., S. H. Choi, and Y. Liao (2021) Standard errors for panel data models with unknown clusters1.000125100%
2Canay, I. A., J. P. Romano, and A. M. Shaikh (2017) Randomization Tests under an Approximate Symmetry Assumption0.9285380%
3Abadie, A., S. Athey, G. Imbens, and J. Wooldridge (2017) When Should You Adjust Standard Errors for Clustering?0.8746367%
4Bonhomme, S. and E. Manresa (2015) Grouped Patterns of Heterogeneity in Panel Data0.64422100%
5Hansen, B. and S. Lee (2019) Asymptotic Theory for Clustered Samples0.5112250%
6Cai, Y (2021) A Modified Randomization Test for the Level of Clustering self0.51121100%
7Cameron, A. C., J. B. Gelbach, and D. L. Miller (2008) Bootstrap-Based Improvements for Inference with Clustered Errors0.51121100%
8Ibragimov, R. and U. K. Müller (2016) Inference with Few heterogeneous Clusters0.51121100%
9MacKinnon, J. G., M. A. Nielsen, and M. D. Webb (2020) Testing for the Appropriate Level of Clustering in Linear Regression Models0.51121100%
10Bai, J (2009) Panel Data Models With Interactive Fixed Effects0.40511100%

Showing the top 10 of 17 scored citations.