EconBase
← All papers

When Should You Adjust Standard Errors for Clustering?

Alberto Abadie, Susan Athey, Guido Imbens, Jeffrey Wooldridge

arXiv 9 Oct 2017 · Mathematics — Statistics Theory · publishedThe Quarterly Journal of Economics (2022) · 1,392 citations (OpenAlex)

arXiv:1710.02926 · PDF · DOI · OpenAlex · Extracted main text

Abstract

In empirical work it is common to estimate parameters of models and report associated standard errors that account for "clustering" of units, where clusters are defined by factors such as geography. Clustering adjustments are typically motivated by the concern that unobserved components of outcomes for units within clusters are correlated. However, this motivation does not provide guidance about questions such as: (i) Why should we adjust standard errors for clustering in some situations but not others? How can we justify the common practice of clustering in observational studies but not randomized experiments, or clustering by state but not by gender? (ii) Why is conventional clustering a potentially conservative "all-or-nothing" adjustment, and are there alternative methods that respond to data and are less conservative? (iii) In what settings does the choice of whether and how to cluster make a difference? We address these questions using a framework of sampling and design inference. We argue that clustering can be needed to address sampling issues if sampling follows a two stage process where in the first stage, a subset of clusters are sampled from a population of clusters, and in the second stage, units are sampled from the sampled clusters. Then, clustered standard errors account for the existence of clusters in the population that we do not see in the sample. Clustering can be needed to account for design issues if treatment assignment is correlated with membership in a cluster. We propose new variance estimators to deal with intermediate settings where conventional cluster standard errors are unnecessarily conservative and robust standard errors are too small.

Citation extraction

23
references
43
in-text mentions
24
distinct cited
3
self-citations
35,484
main-text words

appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Alberto Abadie, Susan Athey, Guido W Imbens, and Jeffrey M Wooldridge (2020) Sampling-based versus design-based uncertainty in regression analysis self0.69371100%
2Manuel Arellano (1987) Practitioners corner: Computing robust standard errors for within-groups estimators0.58531100%
3A Colin Cameron and Douglas L Miller (2015) A practitioner's guide to cluster-robust inference0.58531100%
4Marianne Bertrand, Esther Duflo, and Sendhil Mullainathan (2004) How much should we trust differences-in-differences estimates?0.51121100%
5Min-Te Chao and Shaw-Hwa Lo (1985) A bootstrap method for finite population0.51121100%
6F. Eicker (1963) Asymptotic normality and consistency of the least squares estimators for families of linear regressions0.51121100%
7Peter J Huber (1967) The behavior of maximum likelihood estimates under nonstandard conditions0.51121100%
8Kung-Yee Liang and Scott L Zeger (1986) Longitudinal data analysis using generalized linear models0.51121100%
9James G MacKinnon, Morten rregaard Nielsen, and Matthew Webb (2021) Cluster-robust inference: A guide to empirical practice0.51121100%
10Brent R Moulton (1986) Random group effects and the precision of regression estimates0.51121100%

Showing the top 10 of 24 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1Clustering with Potential Multidimensionality: Inference and Practice1.000104
2The Exact Variance of the Average Treatment Effect Estimator in Cluster Randomized Controlled Trials1.000103
3Inference for Group Interaction Experiments0.92843
41 Panel Data with Unknown Clusters0.87463
5Recent Developments in Inference: Practicalities for Applied Economics0.87462
6Design-Based Uncertainty for Quasi-Experiments0.81142
7Design-Based Multi-Way Clustering0.81142
8On Policy Evaluation With Aggregate Time-Series Instruments0.73732
9Misspecified regressions with mixed regressors: robust inference and causal interpretation0.73732
10Inference in Difference-in-Differences: How Much Should We Trust in Independent Clusters?0.64452