EconBase
← All papers

Spatially Robust Inference with Predicted and Missing at Random Labels

Stephen Salerno, Zhenke Wu, Tyler McCormick

arXiv 11 Mar 2026 · Statistics — Machine Learning

arXiv:2603.11368 · PDF · DOI · OpenAlex · Extracted main text

Abstract

When outcome data are expensive or onerous to collect, scientists increasingly substitute predictions from machine learning and AI models for unlabeled cases, a process which has consequences for downstream statistical inference. While recent methods provide valid uncertainty quantification under independent sampling, real-world applications involve missing at random (MAR) labeling and spatial dependence. For inference in this setting, we propose a doubly robust estimator with cross-fit nuisances. We show that cross-fitting induces fold-level correlation that distorts spatial variance estimators, producing unstable or overly conservative confidence intervals. To address this, we propose a jackknife spatial heteroscedasticity and autocorrelation consistent (HAC) variance correction that separates spatial dependence from fold-induced noise. Under standard identification and dependence conditions, the resulting intervals are asymptotically valid. Simulations and benchmark datasets show substantial improvement in finite-sample calibration, particularly under MAR labeling and clustered sampling.

Citation extraction

35
references
64
in-text mentions
35
distinct cited
6
self-citations
6,360
main-text words

appendix boundary found by appendix_command · 63% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Zrnic, Tijana and Candès, Emmanuel J (2024) Cross-prediction-powered inference1.00064100%
2Conley, Timothy G (1999) GMM estimation with cross sectional dependence1.00053100%
3Chandrasekhar, Arun G and Jackson, Matthew O and McCormick, Tyler H… (2023) General covariance-based conditions for central limit theorems with dependent triangular arrays self0.92844100%
4Bester, C Alan and Conley, Timothy G and Hansen, Christian B (2011) Inference with dependent data using cluster covariance estimators0.92843100%
5Chernozhukov, Victor and Chetverikov, Denis and Demirer, Mert and Du… (2018) Double/debiased machine learning for treatment and structural parameters0.8434375%
6Bang, Heejung and Robins, James M (2005) Doubly robust estimation in missing data and causal inference models0.7373367%
7Angelopoulos, Anastasios N and Duchi, John C and Zrnic, Tijana (2023) PPI++: Efficient prediction-powered inference0.64422100%
8Bullock, Eric L and Woodcock, Curtis E and Souza Jr, Carlos and Olof… (2020) Satellite-based estimates reveal widespread forest degradation in the Amazon0.64422100%
9Cameron, A Colin and Gelbach, Jonah B and Miller, Douglas L (2011) Robust inference with multiway clustering0.64422100%
10Jenish, Nazgul and Prucha, Ingmar R (2009) Central limit theorems and uniform laws of large numbers for arrays of random fields0.64422100%

Showing the top 10 of 35 scored citations.

Cited by, within the corpus

arXiv econ.EM papers that cite this one, ranked by how heavily they lean on it.

Citing paperIntensityMentionsSections
1When Surveys Become Conversations: Adaptive Matrix Validation for AI-Assisted Interviews0.00011