Alberto Abadie, Susan Athey, Guido Imbens, Jeffrey Wooldridge
arXiv 9 Oct 2017 · Mathematics — Statistics Theory · publishedThe Quarterly Journal of Economics (2022) · 1,392 citations (OpenAlex)
arXiv:1710.02926 · PDF · DOI · OpenAlex · Extracted main text
In empirical work it is common to estimate parameters of models and report associated standard errors that account for "clustering" of units, where clusters are defined by factors such as geography. Clustering adjustments are typically motivated by the concern that unobserved components of outcomes for units within clusters are correlated. However, this motivation does not provide guidance about questions such as: (i) Why should we adjust standard errors for clustering in some situations but not others? How can we justify the common practice of clustering in observational studies but not randomized experiments, or clustering by state but not by gender? (ii) Why is conventional clustering a potentially conservative "all-or-nothing" adjustment, and are there alternative methods that respond to data and are less conservative? (iii) In what settings does the choice of whether and how to cluster make a difference? We address these questions using a framework of sampling and design inference. We argue that clustering can be needed to address sampling issues if sampling follows a two stage process where in the first stage, a subset of clusters are sampled from a population of clusters, and in the second stage, units are sampled from the sampled clusters. Then, clustered standard errors account for the existence of clusters in the population that we do not see in the sample. Clustering can be needed to account for design issues if treatment assignment is correlated with membership in a cluster. We propose new variance estimators to deal with intermediate settings where conventional cluster standard errors are unnecessarily conservative and robust standard errors are too small.
appendix boundary found by none_found · 100% of the source is main text. Read the extracted text to check this.
The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.
| Reference | Intensity | Mentions | Sections | Main text | |
|---|---|---|---|---|---|
| 1 | Alberto Abadie, Susan Athey, Guido W Imbens, and Jeffrey M Wooldridge (2020) Sampling-based versus design-based uncertainty in regression analysis self | 0.693 | 7 | 1 | 100% |
| 2 | Manuel Arellano (1987) Practitioners corner: Computing robust standard errors for within-groups estimators | 0.585 | 3 | 1 | 100% |
| 3 | A Colin Cameron and Douglas L Miller (2015) A practitioner's guide to cluster-robust inference | 0.585 | 3 | 1 | 100% |
| 4 | Marianne Bertrand, Esther Duflo, and Sendhil Mullainathan (2004) How much should we trust differences-in-differences estimates? | 0.511 | 2 | 1 | 100% |
| 5 | Min-Te Chao and Shaw-Hwa Lo (1985) A bootstrap method for finite population | 0.511 | 2 | 1 | 100% |
| 6 | F. Eicker (1963) Asymptotic normality and consistency of the least squares estimators for families of linear regressions | 0.511 | 2 | 1 | 100% |
| 7 | Peter J Huber (1967) The behavior of maximum likelihood estimates under nonstandard conditions | 0.511 | 2 | 1 | 100% |
| 8 | Kung-Yee Liang and Scott L Zeger (1986) Longitudinal data analysis using generalized linear models | 0.511 | 2 | 1 | 100% |
| 9 | James G MacKinnon, Morten rregaard Nielsen, and Matthew Webb (2021) Cluster-robust inference: A guide to empirical practice | 0.511 | 2 | 1 | 100% |
| 10 | Brent R Moulton (1986) Random group effects and the precision of regression estimates | 0.511 | 2 | 1 | 100% |
Showing the top 10 of 24 scored citations.
arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.