EconBase
← All papers

Cluster-Robust Prediction-Powered Inference

David Broska, Michael Howes

arXiv 7 Oct 2026 · Statistics — Methodology

arXiv:2610.09601 · PDF · Extracted main text

Abstract

Data collection is often costly or logistically demanding, limiting both the questions researchers can pursue and how precisely they can answer them. Prediction-powered inference (PPI) can reduce the amount of data needed for precise parameter estimation by combining labeled data with machine learning predictions. However, ignoring dependence within clusters can produce confidence intervals that cover the true parameter less often than their nominal rate. We introduce Cluster-Robust PPI++, which provides standard errors in closed form and asymptotically valid confidence intervals under arbitrary dependence within independent clusters, requiring no bootstrap or resampling. Our central contribution is to accommodate partially labeled clusters, a common empirical setting in which clusters contain both labeled and unlabeled units. As units are dependent within clusters, partially labeled clusters violate the independence assumption of PPI++. We also show how precision increases depend on the labeling design, and derive a cluster-aware power tuning rule that minimizes asymptotic variance. In an application to television news, standard PPI++ confidence intervals have coverage below 60%, whereas Cluster-Robust PPI++ can achieve nominal 95% coverage.

Citation extraction

24
references
49
in-text mentions
24
distinct cited
2
self-citations
6,231
main-text words

appendix boundary found by appendix_command · 44% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Angelopoulos, A. N., Duchi, J. C., and Zrnic, T (2024) PPI++: Efficient prediction-powered inference0.89718472%
2Broska, D., Howes, M., and van Loon, A (2025) The mixed subjects design: Treating large language models as potentially informative observations self0.73732100%
3Boussalis, C., Coan, T. G., Holman, M. R., and Müller, S (2021) Gender, candidate emotional expression, and voter reactions during televised debates0.64422100%
4Rister Portinari Maranca, A., Chung, J., Hinck, M., Wolsky, A. D., E… (2025) Correcting the measurement errors of AI-assisted labeling in image analysis using design-based supervised learning0.64422100%
5Girbau, A., Kobayashi, T., Renoust, B., Matsui, Y., and Satoh, S (2024) Face detection, tracking, and classification from large-scale news archives for analysis of key political figures0.5112250%
6Kluger, D. M., Lu, K., Zrnic, T., Wang, S., and Bates, S (2025) Prediction-powered inference with imputed covariates and nonuniform sampling0.51121100%
7Angelopoulos, A. N., Bates, S., Fannjiang, C., Jordan, M. I., and Zr… (2023) Prediction-powered inference0.40511100%
8Arel-Bundock, V., Briggs, R. C., Doucouliagos, H., Aviña, M. M., and… (2026) Quantitative political science research is greatly underpowered0.40511100%
9Cameron, A. C. and Miller, D. L (2015) A practitioner's guide to cluster-robust inference0.40511100%
10Egami, N., Hinck, M., Stewart, B. M., and Wei, H (2023) Using imperfect surrogates for downstream inference: Design-based supervised learning for social science applications of large l…0.40511100%

Showing the top 10 of 24 scored citations.